Fetching from the wire…
Public story · 2026-09-06 · high
In a 300-instance test, recall on unresolved decisions went from zero to 0.667, then to 1.0 once verification results were added.
Why now: The paper posted to arXiv in September 2026.
DNative-Twin records an agentic decision as a typed trajectory: observed state, the path taken, and the authority behind it. Then it replays that trajectory alone, outside the rest of the system.
That matters for teams shipping agents that act without a human checking each step. In a 300-instance test, the system started at zero recall on decisions it couldn't fully explain.
The paper names its own limit. Graph structure can localize which parts of a decision changed, but it can't determine what a tool did in a state nobody observed.
Adding replay-contract state raised recall on those unresolved cases to 0.667. Adding verification results on top of that pushed it to 1.0.
Verification results account for most of the gain. Without them, DNative-Twin misses roughly one in three of the decisions it can't fully explain.
Each link below shares sources, entities, or timing with this story.
Twin has a frontier coding agent write a simulator of an unknown grid game, and the harness refuses to let the agent act until that program reproduces every observed transition, with each mismatch becoming a counterexample used to repair the model (arXiv 2608.14490). It clears...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
When an agent consolidates an external observation into long-term memory, attach platform-controlled metadata recording the source's trust level, then gate tool execution by matching action risk against supporting-memory authority. Laundered memories hit a 1.000 attack success...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
New training paradigm that teaches agents *why* actions succeed by contrasting successful against suboptimal alternatives. Treats action quality judgment as a first-class training objective rather than afterthought reflection. arXiv 2603.08706
When authority, resource, and evidence gates run together, a remediation applied by one control changes the action or context another control already judged. The paper's two implemented operators, evidence substitution and resource-budget downroute, do not commute. arXiv Their...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.