Journeyman: CLI benchmark that profiles LLM agents' process quality
Journeyman lets you drop any OpenAI‑compatible agent into seven simulated tasks and measures how it behaves, not just whether it finishes. It returns a nine‑axis profile showing strengths and blind spots, helping developers understand and improve their agents. The tool is pure Python, zero‑dependency, and runs locally without touching real files. It's useful for AI developers who need systematic, reproducible agent evaluation.
View on GitHub →codechu/journeyman