Fetching from the wire…
Research2026-09-03 · source-backed
This paper defines verifier information density V_d = k/C, the fraction of an agent's causal chain whose per-turn correctness the verifier exposes, and shows terminal-state verifiers sit deep in the low-V_d regime where targeting credit onto important turns cannot help. In shared-rollout comparisons on tau²-bench, continuous dense reward spread uniformly beats sparse binary outcome reward, while concentrating the same advantage on progress turns is exactly as harmful as concentrating it on random turns. The mechanism is coverage: terminal-state verification collapses signal to a single final-write turn (k=1 in 98% of rollouts) while success needs 5-8 steps. Measured V_d is ~0.15 on tau²-bench against a phase boundary at ~0.8.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own protocol, but no implementation of the multi-shell decoder existed. This paper supplies one for the full 301-class codebook with a fused dequantize-plus-matvec kernel and measures batch-...
A controlled study ran five Qwen models over eight cases against a DWSIM simulator, 120 slots per arm, with one instruction as the only difference: request a fresh simulation after a substantive modification. No hard gate. Re-verification happened in 94 of 120 guided slots aga...
Ockhamareto (arXiv 2608.24473) reinforces a unit-test rollout only when it's non-dominated on both mutation-killing and test count, then ties each test's killing power back to specific source tokens. Against MIST-RL that's a 3.4x better per-test trade-off, plus 30 to 35 percen...
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% pea...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.