Fetching from the wire…
Research2026-08-02 · source-backed
Alexi Gladstone, Yilun Du and Heng Ji generate K candidate matches between output and real data at each training step and train only on the best: exploration inside the training loop rather than as RL post-training. 6.2x sample, 4.1x FLOP, 47% better parameter efficiency, with the advantage widening at scale. An Explorative Policy matches Diffusion Policy on robotics with 1 forward pass instead of 100. Single-source blog post, no peer review yet.
Each link below shares sources, entities, or timing with this story.
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
arXiv 2608.05424 shows ImageNet- and LAION-scale pretrained encoders pick up metadata traces tied to camera and image-processing properties. Deliberately injecting metadata-semantics correlations during pretraining produces systematically higher metadata sensitivity and larger...
Saivineeth147/lora-speedrun ranks approaches by actual elapsed training time instead of the FLOP counts and loss curves that dominate published comparisons. 114 points on HN, thread focused on whether wall-clock is the honest metric for people iterating on a single GPU. It is....
Poolside AI released two models that change the math on local coding agents. Laguna M.1 is a 225B total / 23B active MoE model scoring 72.5% on SWE-bench Verified. Laguna XS.2 is a 33B total / 3B active model scoring 68.2% on the same benchmark, 44.5% on SWE-bench Pro, and 30....
DSEffi-Bench covers 1,000 instances across 10+ libraries with stress-testing harnesses and human-validated references, evaluated on 16 models (arXiv 2608.30248). GPT-5.4 leads correctness at 66.9% Pass but its 71.7% efficiency score barely beats GPT-5.4-mini's 71.6% despite so...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.