Skills
Your CLAUDE.md may not be doing anything: 288 controlled runs across Claude Code and Codex found context strategy does not measurably move correctness
A two-agent ablation ran 17 real tasks from 3 repositories through 288 gold-test-evaluated runs, varying only the persistent context file (AGENTS.md / CLAUDE.md), and equivalence-tested the effect down to a ≤10–15 percentage-point band. Failures traced to implementation deficiencies — design choices, pattern selection, precise code wiring — not missing repository knowledge, and agent-specific difficulty (Spearman 0.75 on borderline tasks) explains why earlier studies disagreed. Actionable read: stop tuning context files to buy pass rate, and spend that effort on worked patterns and wiring examples the agent can copy instead of prose describing the repo.
↳ Follow the thread