Voices
Thinking Machines Lab and UIUC beat the human text-to-SQL benchmark for the first time, with no pipeline engineering
An August 27 post from Thinking Machines Lab with Yuxuan Zhu, Tengjun Jin, Yoojin Choi and Daniel Kang describes ReViSQL-K2.6, an RLVR run on Kimi-K2.6 using VeriEQL semantic-equivalence verification plus rule-based process rewards on expert-curated data. With 16-sample self-consistency it reaches 92.97% on Arcwise-Plat-SQL at $0.56/task, edging the 92.96% human benchmark, and greedy decoding gets 91.37% at $0.035/task. The load-bearing claim for builders is that it outperforms scaffolded baselines by 8 to 22 points while deleting the orchestration layer entirely.
↳ Follow the thread