Fetching from the wire…
Agents2026-09-09 · source-backed
Evaluating task-progress reporting on τ²-bench and StageIF, reliability depends on which stage the task has reached: most deployed models lose accuracy once work is under way and recover once done, while the newest generation closes the mid-task dip and instead under-reports completion at the finish line. arXiv 2609.08589 The paper's explicit conclusion is to stop gating control flow on the model's own state reports.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.25479 shows a malicious model provider can embed dormant steering logic in the architecture definition itself via a trigger-gated additive modification of an intermediate representation. No data poisoning, no control of downstream fine-tuning, no deployment-time pro...
ARCHER is a test-driven multi-agent program-synthesis harness for building-compliance checking. Evaluating six harnesses of increasing agentic sophistication across four backbone models, deterministic orchestration won for *every* backbone, improving mean union accuracy 82% ov...
GSQ uses Gumbel-Softmax sampling to match the accuracy of QTIP and AQLM while keeping the deployment simplicity of GPTQ/AWQ. If you're quantizing models for local inference, this eliminates the accuracy-vs-complexity tradeoff.
Poisoned entries in persistent memory force unintended tool selection during retrieval — even against explicit user instructions. Unlike prompt injection targeting input, MCFA targets the memory store, making it persistent and harder to detect. If your agent has long-term memo...
Rather than a separate router model, PyroDash internalizes the escalation policy inside the small model: mid-generation the SLM emits a control token, and a Collaborate Engine hands the query plus partial reasoning trace to a frozen LLM for a single completion. No LLM retraini...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.