Stele Bench: reproducible benchmark suite for coding-agent memory evaluation
Stele Bench offers a complete, open‑source framework to assess how coding agents retain and use memory across tasks. It includes contracts, fixtures, scripts, and detailed result bundles, enabling researchers to run, verify, and extend the experiments reliably. The suite is aimed at AI developers and researchers who need a rigorous way to measure memory effects in LLM‑based coding assistants. Its thorough documentation and automated verification set it apart from ad‑hoc benchmarks.
View on GitHub →Stele-Dev/stele-bench