Fetching from the wire…
Research2026-09-05 · source-backed
arXiv 2609.04172 trains OPD with a single query and finds it keeps improving for hundreds of steps across task domains and model families. Measuring state coverage, the fraction of full-data states a query set's rollouts reach, one query hits 71.5% and 16 semantically distinct queries reach 98.9%. Alignment slows at the same rate either way. The authors summarize it as OPD being data-overfed and algorithm-starved, and content-light templates plus off-domain WildChat queries also approached the real-query baseline (arXiv).
Each link below shares sources, entities, or timing with this story.
Prior work fuses OPD's dense token-level supervision with RLVR's sparse reward in a single step, either weighted-additive or teacher-modulated advantage rescaling. A plain two-stage OPD-then-RL scheme beats pure OPD, pure RLVR and all joint baselines across logic and math benc...
Measuring teacher supervision during on-policy distillation shows substantial noise that worsens as the teacher scales, yet the student converges comparably whether that supervision is kept or stripped (arXiv 2608.31046). Learning concentrates on low log-probability tokens, an...
Four preregistered studies (1,584 multi-agent simulations, 16 languages, 3 model families) prove that alignment interventions reducing harmful outputs in English actively amplify them in Japanese and 14 other languages. Alignment-induced dissociation correlates with Power Dist...
Measuring 8,234 revisions over nine years, with suppression detected semantically and validated against blinded hand labelling at 0.828 precision and 0.911 recall, exclusions were added 1,642 times and withdrawn 304 (arXiv 2608.31062). Per individual rule the ratio climbs to 1...
Measuring expert time on two datacenter GPU generations shows it is linear in neither token count (EPLB, LPLB, UltraEP) nor activated-expert count (METRO). Below roughly 156 to 168 tokens, HBM weight streaming dominates so cost attaches to activated replicas. Above it, grouped...
arXiv 2608.11878 replaces the handful of manually implemented injection-testing environments with an Environment Simulator, Attacker Agent, and User Simulator that generate executable stateful environments and discover viable injection points automatically. Injection timing an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.