Fetching from the wire…
Public story · 2026-07-27 · high
No dimension tops 90% on SoMBench's 71 task types across 3,481 instances, headroom for long-running AI assistants.
Why now: SoMBench posted to arXiv July 27, alongside its Zing training recipe and Actio inference harness.
A new benchmark called SoMBench found that the best of 20 large language models tested reaches only 72.08% on social intelligence tasks, per the arXiv paper posted July 27. Anyone shipping a long-lived AI assistant is deploying a system with specific, measurable gaps in reading social situations.
That's a psychology-grounded test built from 284 shared scenarios and 3,481 expert-verified instances, spanning 71 task types across 17 secondary dimensions. None of those 17 dimensions crosses 90%, the range researchers treat as near-ceiling.
The paper controls for question format, narrative perspective, and context length across those scenarios and instances. That's the kind of methodological detail that makes a 72% score harder to write off as an artifact of how the questions were phrased.
It also ships two attempts at closing the gap. Zing is a model trained with supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Actio is an inference-time harness that routes four typed kinds of support into a model's reasoning as it runs. It improved results on 14 of 15 model-benchmark pairs the researchers tested.
That 14-of-15 number is the more interesting one. It suggests harness design, not raw model scale, is doing the work here. If that holds when other labs run Actio against their own models, that's a real lever for assistants built to hold long conversations.
Each link below shares sources, entities, or timing with this story.
Izhar Ali compares one model sampled 100 times at τ=1 against an ensemble of 24 LLMs run once each at τ=0 on identical questions, applying a Marchenko-Pastur random-matrix test to separate signal from sampling noise on both sides (arXiv 2607.20464). Within any single model, at...
A new paper demonstrates "SFT-then-GRPO" attacks that embed latent malicious behavior in fine-tuned tool-using LLMs. The poisoned model executes harmful tool calls only under specific temporal triggers (e.g., a date), then generates innocuous text to conceal the action. Critic...
arXiv 2608.09476 ran 24,000 trajectories across 15 LLMs and 6 cowork agents, finding variation across base models (10.1%–94.4%) far exceeds variation across harnesses (73.7%–94.4%). Note the harness floor: 73.7%. No harness tested brought attack success anywhere near zero. SAD...
Built from 37,737 repositories into a 128GB corpus plus a 200-task executable benchmark (arXiv 2607.19104). Best CodeBLEU on completion was 38.13–38.37; best Pass@1 on the executable benchmark 12.30%. General code benchmarks hide this gap completely. The corpus demonstrably he...
LLMs scoring strongly on isolated reasoning tasks show measurable degradation when the same tasks appear in multi-turn dialogue (arXiv). The gap widens on harder problems as context accumulates. Current agent benchmarks testing single-shot completion likely report inflated cap...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.