Fetching from the wire…
Public story · 2026-09-06 · high
TabLDM trains only on synthetic causal data, no real spreadsheets, yet still beats fine-tuned rivals on regression benchmarks.
Why now: Xiaomi posted the TabLDM paper to arXiv in September 2026.
Xiaomi built Xiaomi-TabLDM, a tabular model trained only on synthetic causal data, with no real spreadsheets in its training set. Tabular data, pricing tables, sensor logs, transaction records, usually can't leave the company that owns it. A model that never needs real examples sidesteps that privacy problem entirely.
It ranks first on OpenML-CTR23 and second on regression across three separate benchmark suites: TALENT, TabArena and BCCO. On TabArena regression, it posts the second-highest Elo score behind TabFM. Getting there costs 82% less training time and 68% less prediction time than TabFM needs.
The paper doesn't say whether this generalizes past benchmark tables shaped by clean causal structure. Structural causal models are a specific bet about how spreadsheet columns relate, and a second-place finish on TabArena regression means Xiaomi hasn't beaten TabFM outright. Whether the same synthetic-only approach holds up on tables with messier, non-causal relationships, the kind production data usually has, is still open.
Each link below shares sources, entities, or timing with this story.
The stealth model is Xiaomi's MiMo V2 Flash: 309B MoE with 15B active parameters, 256K context, hybrid-thinking toggle. 73.4% SWE-Bench Verified — top open-source globally — approaching GPT-5-High at roughly 3.5% of the cost. A successor model was teased in the same OpenClaw P...
Xiaomi-Robotics-1 (arXiv 2607.15330) reports 74.5% average success on RoboCasa, beating RLDX-1, Cosmos Policy, GR00T N1.6, Pi-0.5 and Pi-0-FAST, plus a new SOTA 57.6% on RoboCasa365 against a prior best of 46.6%. 100K hours of *real* manipulation data is the moat, not the arch...
Across 30 models from three families, verbalized confidence compared against logits-based confidence on 8 classification tasks and semantic entropy on 2 generation tasks: instance-level association is weak on average, improving only on easier items and stronger base models. In...
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
ContinualSkillBench (arXiv 2608.03874) tests continual skill learning across five domains of 100 interconnected subtasks each, ordered by rising difficulty with deliberate reuse opportunities. Sequential execution helps, but in-context learning performs comparably to explicit...
SARC-DQ found competent agents converted freshness/lineage/provenance defects into costly actions about 60% of the time, with both data-quality flags and the agents' own hedging detecting them at chance. The conversion rate was flat across four model tiers spanning a 15x price...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.