Fetching from the wire…
Public story · 2026-08-25 · high
Six frontier agents got the official patch notes anyway, and the best one still missed 30% of the stale content.
Why now: Four separate papers on skill-file evaluation surfaced together, and none of them grade skill files the way current libraries do.
Repo2Skill-Evo tested 105 real release transitions across 57 repositories, and every single one broke part of the AI skill file built for it.
For any team keeping a skill file for a dependency, every major version bump has a 100% chance of breaking it, not an occasional edge case.
Skill files work by naming the exact script, API call, or flag a repo needs. That precision is what makes them useful, and a version bump breaks it without anyone noticing.
The Repo2Skill-Evo paper handed six frontier agents the official patch notes explaining what changed, then asked them to repair the stale skill file. The best agent reached 69.7% on a metric that scores catching stale content against over-editing. Even with the diff in hand, it left 30% of the rot in place.
NVIDIA ran a different test the same week. It asked whether grading a skill file's structure predicts whether the skill helps, and the correlation came back at 0.14, per coverage from AINews. Their fix is to run each task twice, once with the skill loaded and once without, and score the difference in completed work.
A third paper, SkillAlchemy, generated skills straight from source material instead of having humans write them. It beat no-skill runs by 19.9 points and reached parity with human-curated skills across 87 tasks.
A fourth study on long-horizon web agents found verified execution experience beat distilled summaries by 8.7 to 15.5 points across seven models.
There's a problem here nobody's solved yet. If skills decay on every release and agents can't reliably repair them even with the diff, a skill library's maintenance cost scales with how often its dependencies change. At some point that costs more than the skills save, and none of the four studies say where that point sits.
Each link below shares sources, entities, or timing with this story.
NVIDIA uses Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (NVIDIA uses Claude Code); both cover SKILL, Their, When; reported by the same outlet (arxiv.org).
NVIDIA uses Claude Code / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA uses Claude Code); both cover Their, Then, When; reported by the same outlet (arxiv.org).
NVIDIA uses Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (NVIDIA uses Claude Code); both cover Skill, Then, When; overlapping topics (agent, point, skill).
NVIDIA uses Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA uses Claude Code); both cover Skill, Spearman, Then; reported by the same outlet (arxiv.org).
NVIDIA uses Claude Code / Shared entities / Earlier coverage
Linked by a graph relationship (NVIDIA uses Claude Code); both cover Their, Then, There, When; earlier Their coverage from 2026-03-25.
NVIDIA uses Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (NVIDIA uses Claude Code); both cover Then, There, When; overlapping topics (agent, version).
NVIDIA invested in OpenAI / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (NVIDIA invested in OpenAI); both cover There, When; reported by the same outlet (arxiv.org, latent.space).
NVIDIA uses Claude Code / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA uses Claude Code); both cover Even, When; overlapping topics (even, point).