Testing AI Applications · June 10, 2026 Why your LLM-as-a-Judge is “Too nice” (and how to fix it) When building LLM-as-a-judge pipelines, I've noticed a recurring failure mode that quietly destroys the reliability of evaluation metrics: the judge is simply too polite.