Best LLM Reliability Score on Structured Data Analysis Is 24.21%; Models Rarely Refuse When No Valid Evidence Path Exists
TrustDABench asks two diagnostic questions about spreadsheet and CSV analysis that accuracy benchmarks skip: can the model refuse or ask for clarification when no valid path exists from question to evidence, and does it preserve the correct analysis when the same evidence is reshaped into a different table form. Built from an evidence-path view with 19 perturbation operators and 2,340 human-verified perturbed instances across eight models, the best reliability result is 24.21% average MRS from GPT-5.5 and the best robustness still shows 9.10% average ASR from Claude-Sonnet-5. Failures are systematic rather than random: models rarely detect conflicting evidence and keep walking executable but unsupported analysis paths.
↳ Follow the thread