Fetching from the wire…
Research2026-09-14 · source-backed
This BlackboxNLP reproduction re-ran the original distillation-transmits-preferences experiments with new preference categories, a new task (chess move generation), Ministral8B, and an answer-space ablation. The original claims hold, but transmission strength varies widely across traits and tasks and one model shows almost no effect. Not a universal property of distillation.
Each link below shares sources, entities, or timing with this story.
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
The benchmark hands a developer agent a real client engagement setup, business records, a requirements-holding client, a production API, an inherited codebase, cost and model limits, then scores it by deploying the customer-service agent it built against held-out simulated use...
SkillJack found safety detection on poisoned trajectories ran 98.5% but fell to 11.4% on skills extracted from those same trajectories, with 80% of skill-mediated attacks persisting after the original records were deleted. Distillation launders intent. Cleaning your trace stor...
arXiv 2608.04804 sends a 7B searcher into the repo first, sandbox-verifies its reproduction claims and strips false ones, then routes to one of four frontier fixers. On the full 266-task Python slice under the official capped budget it solves 159 vs 158 for the best single mod...
This method retains four categories of reusable context (task specs, data schemas, tool configs, output constraints) while discarding session-specific reasoning, enabling role-based workspace transfer across users (arXiv:2607.09493). It reports 96% completion versus 79% withou...
arXiv 2608.23740 exposes file-level claim, status and broadcast as MCP tools over a shared filesystem, motivated by the observation that a single agent abandons up to half of hard tasks with a one-file stub-and-exit. Across five frontier coding CLIs on four backend tasks, two-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.