Confidence-Calibrated Cascade Hits 98.8% Accuracy on LLM Acceptance Testing and Cuts Cost 31.7% vs Single Judge
Requirements-Augmented Generation (REAG) generates context-aware test oracles for LLM-based software by retrieving software requirements, domain knowledge, and personas via adaptive RAG and self-reasoning — addressing the fact that the same query can demand different correct responses per persona and context. A confidence-calibrated cascade then quantifies verdict reliability through simulated expert agreement, accepting, escalating, or abstaining, with guarantees backed by conformal risk control. In an industrial case study on a production nutrition advisory app, REAG scored 3.91/5 oracle quality (qualified or marginal in 82% of cases), and the cascade reached 98.8% accuracy, lifted oracle quality to 4.30 by filtering, and delivered 31.7% better cost-efficiency than single-judge baselines.
↳ Follow the thread