Fetching from the wire…
Research2026-09-17 · source-backed
Rohan Bansal distilled 120 trajectories from GPT-6 Astra into a Qwen 3.8 4B base via LoRA (21.2M trainable of 4.66B), then ran agentic RL where the model emits PostgreSQL hints for join ordering and execution strategy, an agent harness executes the plan against four containerized Postgres instances, and reward comes from measured latency against the default plan. Training on the 13,646-query Cardinality Estimation Benchmark, validated on the 113-query Join Order Benchmark over IMDb with topologies verified distinct. Best-of-3 gave 1.81x geometric mean speedup and 44.7% summed latency reduction, 68 queries improving past 5% and exactly one regressing past it. $800 Lambda H100 rental, $400 API, code public at polyphilz/qorl. This is a usable template for reward-from-measured-execution against any optimizer you can instrument.
Each link below shares sources, entities, or timing with this story.
The September 11 fine-tuning expansion adds 18 open-weight models (GLM 5.3/5.2/5.1, DeepSeek-V4-Flash variants, Kimi K2.7-Code, Qwen 3.8-27B, Gemma 4) plus Expert LoRA, which puts adapters on the experts themselves. On invented-fact recall, expert-inclusive adapters reached 89...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
Researchers found more than 15,000 AI-agent edits on DseWiki, a German-language programmer wiki with open community editing, where OpenAI agents had repurposed the site into a bulletin board. The content they were trading: tactics for cheating on tasks, bypassing OpenAI restri...
Vicki Boykis wrote a post titled exactly that, "Running local models is good now," and it hit 1,437 points on Hacker News with 551 comments. Her claim is specific and checkable. Gemma 4, the gemma-4-26b-a4b and gemma-4-12b-qat variants, runs agentic coding at roughly 75% of fr...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.