Fetching from the wire…
Skills2026-08-10 · source-backed
Field Aware Agent Skill Retrieval computes sparse and dense similarity separately for each field and combines the scores, hitting 77.95 Recall@10 on SkillRet and 83.78 on SRA-Bench, with the margin widening as the skill bank grows. If your library crossed a few hundred entries and the right skill stopped surfacing, that's a retrieval-layer bug, not a prompt bug.
Each link below shares sources, entities, or timing with this story.
Skill-α (arXiv 2608.01678) reframes skill generation as RL over sequential edits, decomposing skill construction into individually evaluable changes. The novel signal is a rollback reward that scores each modification by comparing downstream task execution using the original s...
The first defect state-aware multi-round review benchmark: 2,269 real tasks across five languages, each annotated with defect description, type and severity plus cross-round state labels tracking a defect's full trajectory (arXiv 2608.27442). Mainstream models degrade signific...
GameASG-Bench builds 47 browser game-generation tasks across 12 genres, each with an evaluation interface declared before generation. Across nine agent stacks the highest mean runtime check pass rate is 93.2%, but the highest strict task success, requiring every applicable che...
Per-trace debugging experience in agentic code translation doesn't accumulate into reusable knowledge (arXiv 2609.15381). TRAIL runs two agents adversarially, a Translator and a Challenger, distilling trace-specific experience into generalizable translation rules. Against the...
Skill Issue points out that auto-synthesized skill documents get optimized against tasks a capable agent already solves with no document at all, leaving the optimizer nothing to measure. The fix is mining harder tasks by reverting merged PRs at a frozen base commit. On three K...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.