Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.04570 tests 150 personas across 6 tasks. Every one of 12 models over-inferred user attributes on 35-49% of claims, ranging 27-59% by task type. The damning result is the self-monitoring inversion: models rating themselves as over-inferring least ranked as fabricating most (rho = -0.60, p = 0.044). Inferred attributes accumulate roughly linearly across multi-turn interactions with almost no revision. Any memory layer trusting the model's own confidence is compounding fabricated profile data.
Each link below shares sources, entities, or timing with this story.
Across 30 models from three families, verbalized confidence compared against logits-based confidence on 8 classification tasks and semantic entropy on 2 generation tasks: instance-level association is weak on average, improving only on easier items and stronger base models. In...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
A paper submitted July 23 benchmarks open-weight LLMs as coding agents across a consumer-grade deployment spectrum on 20 longitudinal data-preparation tasks producing 102 variables, reporting that current 31-35B models "almost saturated the benchmark" with average task complet...
Diffusion LMs decode many tokens per step but pay to interact with all suffix tokens every step, and existing fixes just keep a local window while re-initializing suffix tokens identically each timestep (arXiv 2608.23167). This method splits the suffix into local, middle and t...
Multiple independent models train against each other with peer-derived rewards and no ground-truth labels, gaining 3.0-8.6% across seven text benchmarks and 2.3-7.2% across four multimodal ones. The mechanism claim matters more than the numbers: varying architectures, model si...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.