Reproducible benchmark harness with live leaderboard for planners
OpenPlan Bench provides a reproducible harness for evaluating classical PDDL and multi‑agent path‑finding planners. It runs each planner on a curated suite, records timeouts, errors, and validated solutions, and publishes a living leaderboard. Researchers can compare planners under consistent budgets and see detailed provenance. Unlike ad‑hoc benchmarks, it logs every run, ensuring transparency and reproducibility.
View on GitHub →openplan-labs/openplan-bench