SoMBench: The Best of 20 LLMs Scores Only 72.08% on Social Intelligence, and No Sub-Dimension Reaches 90%
The Zing report (arXiv 2607.23740) introduces SoMBench, a psychology-grounded benchmark of 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms, controlling question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Evaluating 20 representative LLMs, the best model reaches just 72.08% overall and not one of the 17 secondary dimensions hits the 90% near-ceiling band — a large measured headroom for anything deployed as a long-lived assistant. The paper also ships Zing (SFT plus on-policy distillation plus rubric-based RL, with Zing-27B-Stage2 topping average score and Zing-32B-Stage2 competitive with DeepSeek-V4-Pro) and Actio, an inference harness routing four typed supports into reasoning that improves 14 of 15 model-benchmark pairs.
↳ Follow the thread