Skills
Optimizing the context artifacts beat optimizing the retrieval harness by roughly 2x on a 5,176-query production benchmark
Working on enterprise text-to-SQL, this paper argues the bottleneck is what context reaches the model rather than the model or the retrieval loop, and it distills historical query profiles into reusable SQL reference cards. On an internal benchmark of 5,176 production queries from a major online retailer, optimizing the context artifacts yielded roughly 12 to 25% AST similarity gains against roughly 3 to 12% from optimizing the harness. The public BEAVER result is weaker and honestly flagged, 9.00% versus 6.33% at p=0.12 on a held-out N=300, so treat the effect size as directional.
↳ Follow the thread