Research
A Self-Improving Inference Cascade Reads a Flat 3% Error While True Error Swings to 32%
Measuring the standard cost-saving pattern of a cheap student escalating a hard tail to a frontier verifier, the authors find the verifier's blind spot grows with student capability (0.12 to 0.55 as the student scales 0.5B to 32B), so it is worst in exactly the cheap-student, cheap-verifier regime cascades exist to create. Buying it away returns the saving: a frontier verifier cuts the blind spot to about 0.05 but escalates on 46% of hard-MATH queries against a 39% true error rate. Fine-tuning the student on verifier rejections degraded and eventually collapsed it across every teacher tried, and every metric computed through the verifier read a flat 3% error while delivered error moved to 32%.
↳ Follow the thread