Agents
Snowflake's HybridDeepResearch shows frontier models hit only ~50-54% Pass@8 when an answer needs both SQL and web search
arXiv 2609.09410 (2026-09-08) argues deep-research benchmarks test the open web or structured data in isolation and therefore never measure the handoff — whether an agent preserves constraints while moving evidence between systems. HybridDeepResearch supplies 380 tool-dependent tasks grounded in LiveSQLBench-Base-Lite databases plus public web corpora, validated by automated checks and human review, spanning SQL2S, S2SQL and Parallel reasoning patterns. GLM-5.2, Claude Sonnet 4.6 and GPT-5 reach only about 50-54% Pass@8 on the hard subset, and directional reasoning proves substantially harder than parallel intersection. Code and data are public on GitHub and Hugging Face.
Source
↳ Follow the thread