hatchmoment. scored by care · not by stars

journeyman

Journeyman: CLI benchmark that profiles LLM agents' process quality

Journeyman lets you drop any OpenAI‑compatible agent into seven simulated tasks and measures how it behaves, not just whether it finishes. It returns a nine‑axis profile showing strengths and blind spots, helping developers understand and improve their agents. The tool is pure Python, zero‑dependency, and runs locally without touching real files. It's useful for AI developers who need systematic, reproducible agent evaluation.

agent-evaluationagentsbenchmarkllmllm-evaluation
View on GitHub →

codechu/journeyman