Sources
DataSpace Finds Your Agent Harness Is Worth 15.36 Points — More Than the Gap Between Some Frontier Models
DataSpace (arXiv 2608.03451, submitted August 4) benchmarks data agents on 410 cross-language tasks over 7,439 artifacts totaling 15.01GB spanning CSV, JSON, SQLite, Markdown, PDF and video, validated by 11 domain experts. Testing six frontier multimodal models across five agent frameworks, the best accuracy reached only 66.34% — and harness choice alone produced a 15.36-point spread independent of the underlying model. The consistent failure mode across all six backbones was multimodal evidence integration and joins, which is exactly the workload most 'chat with your data' products claim to solve.
↳ Follow the thread