Fetching from the wire…
Public story · 2026-07-27 · high
No dimension tops 90% on SoMBench's 71 task types across 3,481 instances, headroom for long-running AI assistants.
Why now: SoMBench posted to arXiv July 27, alongside its Zing training recipe and Actio inference harness.
A new benchmark called SoMBench found that the best of 20 large language models tested reaches only 72.08% on social intelligence tasks, per the arXiv paper posted July 27. Anyone shipping a long-lived AI assistant is deploying a system with specific, measurable gaps in reading social situations.
That's a psychology-grounded test built from 284 shared scenarios and 3,481 expert-verified instances, spanning 71 task types across 17 secondary dimensions. None of those 17 dimensions crosses 90%, the range researchers treat as near-ceiling.
The paper controls for question format, narrative perspective, and context length across those scenarios and instances. That's the kind of methodological detail that makes a 72% score harder to write off as an artifact of how the questions were phrased.
It also ships two attempts at closing the gap. Zing is a model trained with supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Actio is an inference-time harness that routes four typed kinds of support into a model's reasoning as it runs. It improved results on 14 of 15 model-benchmark pairs the researchers tested.
That 14-of-15 number is the more interesting one. It suggests harness design, not raw model scale, is doing the work here. If that holds when other labs run Actio against their own models, that's a real lever for assistants built to hold long conversations.
Each link below shares sources, entities, or timing with this story.
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage / Tension
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, dimension).
Shared entities / Same source domain / Earlier coverage
Both cover LLMs, SFT; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-03-05.
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, best).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, context).
Shared entity: LLMs / Same source domain / Earlier coverage / Downstream implication
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-19.
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-19.
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-14.
Shared entity: LLMs / Same source domain / Earlier coverage / Tension
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-06-09.