Skills
LLM-as-Judge Reliability Protocol: Evidence-First Structured JSON Output + Cronbach's Alpha Validation Before Production Deployment
Production-grade LLM-as-judge evaluators require three design elements to avoid score inflation and positivity bias: structured JSON output that requires evidence citations before scoring (preventing post-hoc rationalization), explicit rubrics with few-shot examples anchoring each score level, and reliability validation via Cronbach's alpha across multiple independent evaluator runs targeting 0.80+ Spearman correlation with human raters. Without the validation step, LLM judges regularly exhibit systematic biases that make evaluations unreliable for regression testing and quality gating.
Source
↳ Follow the thread