Fetching from the wire…
Research2026-07-28 · source-backed
MAPD inserts structured JSON protocols (task types, reasoning plans, grounding facts) as an intermediate representation distilled from proprietary models into open ones, decomposing queries, retrieving evidence, repairing failed searches and converting exploration traces into protocols that guide a privileged training branch. 39.4% success on Qwen3-1.7B, 44.4% on Qwen3-4B across seven QA benchmarks, while suppressing style drift and verbosity degeneration. A credible recipe for moving agentic search onto hardware you own. (arXiv 2607.24280)
Each link below shares sources, entities, or timing with this story.
SecOPD fine-tunes a defense using token-level feedback during on-policy distillation rather than the sequence-level signal prior work used. Against PISmith adaptive injections on Qwen3.6-27B it reports 9.0% attack success where Meta-SecAlign, the previous state of the art, sit...
A 2.4-trillion-parameter MoE multimodal model with 1M-token context, aimed at long-horizon autonomous software work. Alibaba reports 86.1 against 83.2 for GPT-5.6 Sol Max and 85.0 for Fable 5, the first credible claim that a Chinese lab leads on GUI-driving agentic benchmarks....
Visual on-policy distillation normally needs a stronger teacher or privileged supervision (arXiv 2608.14144). S²VOPD inverts the asymmetry: the same model acts as teacher on the original image and student on a strongly augmented view, so the signal comes free of annotations, r...
arXiv 2607.27146 attacks from-scratch program synthesis, where agents get only natural-language docs and an execute-only binary as oracle. The pipeline auto-converts open-source command-line programs into source-free training environments and uses GLM-5.2 as teacher for synthe...
Quesma ran the model across GPQA Diamond, IFBench and Terminal-Bench 2.1 (89 agentic coding tasks) on L40S, H100 and H200 via Modal. Q4_K_M at 17 GB matched BF16 at 55 GB within a point on all three. UD-Q2_K_XL at 10.7 GB held instruction-following but dropped Terminal-Bench f...
The comparison is against GB300 NVL72, with 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark using real recorded agentic coding sessions with context growth, tool calls and sub-agent spawning preserved (NVIDIA). DeepSeek V4 Pro and Qwen3.5 were...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.