Fetching from the wire…
Vibe Coding2026-09-07 · source-backed
An online skill-evolution framework turns interaction trajectories and evaluator feedback into a persistent versioned library, with each iteration executing against a frozen snapshot so evidence-guided updates only reach later iterations and no model parameters change (arXiv 2609.04869). Against a configuration-matched empty-library control across four OSWorld domains with identical action-generation and grounding stacks, the evolving library won all four post-warm-up runs. Provenance analysis in GIMP found skills retrieved across task-of-origin boundaries and revision churn where repeated accepted edits failed to recover the originating task, so the gains are not monotone. Code at Skill-Evo4GUI.
Each link below shares sources, entities, or timing with this story.
Skill-α (arXiv 2608.01678) reframes skill generation as RL over sequential edits, decomposing skill construction into individually evaluable changes. The novel signal is a rollback reward that scores each modification by comparing downstream task execution using the original s...
1. Build a Private Claude Code Plugin Marketplace (intermediate, vibe-coding) — Bundle skills, agents, hooks, MCP servers into installable team plugins via GitHub repos. Docs 2. Google ADK TypeScript Multi-Agent Orchestration (intermediate, agent-patterns) — Code-first agent f...
Everyone writing SKILL.md files has absorbed the same folklore. Keep the top file thin. Push detail into reference files. Let the agent walk the tree as needed. More layers, more context efficiency. A controlled study submitted July 20 tested that across InfiniteBench, three a...
Today's release means an agent no longer carries the full tool surface in context, and CLI tool access becomes permission-scoped. A Tape Trace Inspector records agent runs, provider requests, tool calls and Skill usage, Skills unify into a shared library with per-agent enablem...
21 out of 21. Not most. All of them. arXiv 2608.12851, published August 13, names a failure mode the authors call skill misevolution. An agent that learns from its own successful trajectories will turn an unsafe success into reusable policy, and that policy persists after the...
EVOHARNESSBENCH does something I haven't seen a benchmark do: it holds the task stream fixed and evolves the harness (arXiv 2609.04280). Seventeen multi-stage streams built from 802 tasks, 520 tools, 42 skills and 62 agents. The finding is that harness expansion alone degrades...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.