SourcesVESPO — Stable Off-Policy LLM Training 152 HF UpvotesarXiv·high signalXBlueskyLinkedInCopy linkVariational sequence-level soft policy optimization. Stable training at 64x staleness ratio. Top paper on HuggingFace Feb 23.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerMAGS routes coding-agent output through Dafny and reports 100% success at producing verified programs on 220 tasksarXivPolicy dependency / Stack layerHarness-Layer Auto-Research Cut Agent Token Traffic 44.7-49.0% at Equal Task PerformancearXiv 2609.20519Policy dependency / Stack layerWrapping an LLM in a five-stage deterministic control loop took constraint satisfaction from ~26% to ~96%arXiv 2609.19710Policy dependency / Stack layerWhen2Think replaces uniform length penalties with instance-level difficulty control, ending the efficiency taxarXiv / HuggingFace Daily PapersPolicy dependency / Stack layerWhen2Think Rewards a Model for Answering Easy Questions Directly Without Losing Accuracy on Hard OnesarXiv 2609.19671Policy dependency / Stack layerICML position paper: every agent framework has reimplemented an operating system badly, so build the OS layerarXivPolicy dependency / Stack layerJames Mickens argues chain-of-thought monitoring can never be a sound security controlarXivPolicy dependency / ContrastMetacognitive Feedback Cut Answer Offloading to an LLM Assistant Roughly in Half (N=704)arXiv 2609.20143