Dependency‑free TypeScript library for calibrated LLM evaluation certificates
Sigil tackles the hard problem of proving how trustworthy an LLM judge is by computing calibrated metrics such as expected calibration error, Brier score, and conformal certificates. It also builds a Pareto frontier of quality‑cost‑latency, signs deterministic reports, and watches for drift with e‑processes. The tool is aimed at ML engineers, risk‑management teams, and anyone needing auditable, offline‑only evaluation of language models. Its zero‑dependency design and exhaustive statistical suite make it a uniquely portable alternative to heavyweight evaluation frameworks.
View on GitHub →apatureai/sigil