ResearchConstraint Decay: LLM Agents Lose 30 Percentage Points When Backend Code Must Follow Architectural RulesarXiv·high signalXBlueskyLinkedInCopy linkResearchers quantify 'constraint decay' — as structural requirements accumulate (ORM patterns, API conventions, DB schemas), LLM agent code-gen performance drops ~30pp in assertion pass rates. Weaker configs approach zero. Flask handled well; Django and FastAPI trigger failures. Data-layer defects (wrong queries, ORM violations) are the leading root cause.SourceSource pagearXiv↳ Follow the threadShared entity / Stack layerSRE-Bench: 5,000 expert hours to build the first contamination-free reverse-engineering benchmark, and frontier agents solve only 31.5% of itarXivShared entity / Stack layerVero: frontier coding agents solve only 27 of 43 formally verified repository tasks, and zero specs on the hardest reposarXivShared entity / Stack layerAgent Behavioral Contracts II: agents co-fail on 90% of missions, so composing reliability bounds under independence is wrongarXivShared entity / Stack layerConvergent Detour Hijacking: an attack that leaves the task correct and just burns 67% more tokensarXivShared entity / Policy dependency'Practice Makes Unsafe': every one of 21 self-evolving agent configurations wrote unsafe skills that outlived the attack that caused themarXivShared entity / Stack layer'Labels Are Not Endpoints': an audit reclassifies 58 MCP agent attack-success labels to benign, dropping verified attacks to zeroarXivShared entity / Stack layer'Catastrophic remembering': CLAUDE.md files grow +226% over their lifetime and almost never shrink — and prompt comments fix itarXivPolicy dependency / Stack layerTopoIntent Compiles Natural-Language Security Intent Into Runnable Mininet Topologies With iptables ACLsarXiv 2608.13389