Stack layer / Contrast
'How Do Agents Fail on AutoResearch': 800 Trajectories Across 8 Harness-Model Combos Find One Failure Every Model Shares — They Never Check Output Against Evidence
arXiv (via HuggingFace Daily Papers)
Stack layer / Threat pattern
How you lay out your repo changes prompt-injection success: highly modular workspaces measurably lower attack success rate
arXiv 2608.14876
Policy dependency / Stack layer
Compose agent guardrails as an algebra instead of a rule list: 94.8% of policy-violating events intercepted while keeping 86.9% task completion
arXiv 2608.16402
Stack layer / Follow-up thread
S²VOPD Gets Distillation Gains With No Teacher and No Labels — by Degrading the Student's View Instead of Privileging the Teacher's
arXiv 2608.14144
Stack layer / Follow-up thread
'Beyond Final Scores' Instruments What Agents Actually Do Across 36 Long-Horizon R&D Tasks — Verdict: Engineering Optimizers, Not Researchers
arXiv (via HuggingFace Daily Papers)
Stack layer / Follow-up thread
TDD-Agent Turns Generated Tests Into Evolving Reasoning Artifacts Instead of Post-Hoc Validators
arXiv 2608.16742
Stack layer / Follow-up thread
Alibaba's CPI-Bench Splits Image-Editing Evaluation Into General, Practical and Intelligent Subsets — and Claims the Closest Alignment to the Arena Image Edit Leaderboard
arXiv (via HuggingFace Daily Papers)
Stack layer / Follow-up thread
Standard demonstrations teach agents hindsight, not exploration — SAFARI synthesizes exploration-rich trajectories instead
arXiv 2608.14339