hatchmoment. scored by care · not by stars

judgeshift-cx

JudgeShift CX: evaluates bilingual LLM judges for customer‑support replies

JudgeShift CX provides a reproducible benchmark to measure when an LLM judge can be trusted to rank customer‑support responses in English and Portuguese. It ships a curated golden‑dataset, a deterministic Python inference pipeline, and a TypeScript dashboard that visualises coverage, agreement, and abstention rates. Designed for developers building automated support agents, it helps decide when to let the model act and when to defer to a human. Compared to ad‑hoc scripts, it offers systematic bias checks, verbosity stress tests, and cross‑language consistency metrics.

braintrustcustomer-supportllm-as-a-judgellm-evaluationmultilingualnlppythontypescript
View on GitHub →

ruimiguelyo/judgeshift-cx