Fetching from the wire…
Tools2026-09-21 · source-backed
briefd indexes a git repo of Markdown and exposes one MCP call, compile_bundle, returning a deduplicated bundle under a token budget you set. On its benchmark of 44 documents and 47 tasks, a 2,000-token bundle costs 1,800 tokens and contains the answering section 96% of the time, against 12,869 tokens at 100% for pasting everything and 4,946 tokens at under 50% for a hand-curated file. The benchmark is the author's own but reproducible with make bench.
Each link below shares sources, entities, or timing with this story.
The project stores project memory as Markdown with YAML frontmatter in a knowledge/ directory, searches it with in-memory BM25, and exposes it through an embedded MCP server, sitting between ad-hoc CLAUDE.md files and black-box vector stores. It claims 80% token reduction via...
MCP-Atlas has 1,000 human-verified tasks across 36 real MCP servers and 220 tools, now with a 100-tool-call budget instead of a 20-turn limit. Current leaders: Gemini 3.5 Flash at 83.6%, Claude Opus 4.8 at 82.2%. Tool-Decathlon runs 108 long-horizon tasks in isolated container...
Alibaba International's Accio team open-sourced 107 tasks (53 CLI, 28 browser, 16 file, 10 API/MCP) running against fourteen offline replicas of real business software in a fresh container per task, with verifiers inspecting mock-service state rather than the transcript. Claud...
This one rearranged my week. An essay published August 4 walks through Databricks' independent benchmark of coding harnesses against its own multi-million-line codebase. Pi, a harness with four built-in tools and a system prompt under 1,000 tokens, paired with Opus 4.8 at xhig...
It captures agent sessions against your server across Claude, ChatGPT and other clients, surfacing intent, reasoning, every tool call, and success scores, then groups sessions by use case ranked by volume and success rate and clusters failures by root cause. $50 per additional...
The v2.1.205 release turned /doctor into a full setup audit that flags unused skills, MCP, and plugins against their context cost, deduplicates local vs checked-in CLAUDE.md, and flags slow hooks (Releasebot). A typical 5-server, 58-tool MCP setup burns ~55k tokens before your...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.