hatchmoment. scored by care · not by stars

stele-bench

Stele Bench: reproducible benchmark suite for coding-agent memory evaluation

Stele Bench offers a complete, open‑source framework to assess how coding agents retain and use memory across tasks. It includes contracts, fixtures, scripts, and detailed result bundles, enabling researchers to run, verify, and extend the experiments reliably. The suite is aimed at AI developers and researchers who need a rigorous way to measure memory effects in LLM‑based coding assistants. Its thorough documentation and automated verification set it apart from ad‑hoc benchmarks.

ai-agentsbenchmarkreproducibilityresearch
View on GitHub →

Stele-Dev/stele-bench