Fetching from the wire…
Research2026-07-26 · source-backed
Jiao, Wu, Sun and colleagues identify Joint Probability Dependence Error as why dLLM parallel decoding is stuck with overly conservative confidence thresholds, forcing redundant denoising iterations (arXiv 2607.20467). DC-Leap adds Dynamic Contiguous Verification folding strictly-ordered causal constraints into parallel decoding to neutralize JPDE, plus draft-guided leaping across multiple tokens for look-ahead context. 53.19x on MBPP for long-sequence generation, 105.02x combined with KV-cache. Training-free means it applies to existing dLLM checkpoints today.
Each link below shares sources, entities, or timing with this story.
Ockhamareto (arXiv 2608.24473) reinforces a unit-test rollout only when it's non-dominated on both mutation-killing and test count, then ties each test's killing power back to specific source tokens. Against MIST-RL that's a 3.4x better per-test trade-off, plus 30 to 35 percen...
Applying a harmless code-grammar constraint during decoding raises malicious-code attack success by 30+ percentage points across 10 popular models (arXiv). The natural-language refusal stays intact while the constrained decoder produces the payload anyway. If you use structure...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
The standard protocol ablates a latent and measures effect at the token where it fires hardest, but that token is chosen by the dictionary under evaluation. Two dictionaries get compared at different places. Training six autoencoders from one initialization showed 7.6% and 11....
arXiv 2608.12253 shows the standard practice of training a policy against a single LLM simulating the user fails because the simulator is itself mode-collapsed, so the policy learns to exploit its dominant mode. Verbalized Sampling recovers up to 9% held-out success; Populatio...
RETRACE has a verifier infer what problem the patch appears to solve using only the patch and trajectory, then compares that inference against the real issue. Training-free, lifted Pass@1 by 7.0% and 3.6% on mini-SWE-agent over SWE-bench Verified. The information-hiding trick...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.