Stack layer / Contrast
Continuity Kernel: a transactional commit protocol for long-lived agent state, model-checked over 2.8M states
arXiv
Stack layer / Threat pattern
HARD: let the agent evolve its own runtime defenses instead of hand-writing guardrails
arXiv
Policy dependency / Contrast
Best models score 70.4% on single-hop API calls but 2.4% on knowing when a tool-use policy makes a question unanswerable
arXiv 2608.12282
Stack layer / Follow-up thread
Faraday: a 27B agent trained specifically to replicate research beats Claude Opus 4.8 and GPT-5.5 on held-out replication
arXiv
Stack layer / Update thread
PlayWorld Puts Multi-Modal 'Agent Players' Inside World Models Across 171 Scenarios — and Finds Long-Horizon State Persistence Still Broken
arXiv (via HuggingFace Daily Papers)
Stack layer / Contrast
Coins: Evaluating LLM Formal Specifications by Instantiating Them on Trusted Test Cases Instead of Proving Equivalence
arXiv 2608.13077
Stack layer / Follow-up thread
Mechanist: An Agentic System for Autonomous Interpretability Research, Built on a 13,000-Paper Knowledge Graph Over 43M Papers
arXiv 2608.12036
Policy dependency / Contrast
Simulator Collapse: RL Against a Single LLM User-Simulator Overfits, and Population Co-Training Recovers 14% of Held-Out Success
arXiv 2608.12253