JudgeShift CX: evaluates bilingual LLM judges for customer‑support replies
JudgeShift CX provides a reproducible benchmark to measure when an LLM judge can be trusted to rank customer‑support responses in English and Portuguese. It ships a curated golden‑dataset, a deterministic Python inference pipeline, and a TypeScript dashboard that visualises coverage, agreement, and abstention rates. Designed for developers building automated support agents, it helps decide when to let the model act and when to defer to a human. Compared to ad‑hoc scripts, it offers systematic bias checks, verbosity stress tests, and cross‑language consistency metrics.
View on GitHub →ruimiguelyo/judgeshift-cx