Fetching from the wire…
Vibe Coding2026-09-14 · source-backed
Three projects inside seven days: daidocs with plain-text .dai files (claiming 92% on LongMemEval-S with Claude Fable 5, 22.40 points over the same model with no memory, at ~10x fewer tokens per question), Baron Munchausen at public alpha 0.6.1 with zero runtime dependencies, and Slowave on Show HN. What separates the first two from the pile is that both attach a judgment to the answer: Baron returns grounded, partial or ungrounded and names the unbacked sentences, and daidocs ships per-question judge verdicts with a sha256 manifest. Evaluating one of these, the retrieval claim matters less than whether it tells you when it's making things up.
Each link below shares sources, entities, or timing with this story.
Barry Zhang and Mahesh Murag, the engineers who built Claude Skills at Anthropic, published a talk and engineering post that's gotten 14K+ likes and is reshaping how I think about agent development. The core argument: most agent approaches fail because they lack domain experti...
MemPalace (57,821 stars, v3.6.0) reports 96.6% raw recall@5 on LongMemEval with no LLM required, 98.4% with hybrid v4 on a held-out 450 questions, LoCoMo R@10 rising 60.3% → 88.9%, ConvoMem 92.9%, MemBench 80.3%, while explicitly refusing head-to-head comparison against Mem0,...
If you wrote an MCP server before July, it's on a protocol shape the maintainers have already removed. Not deprecated-with-a-migration-window. Removed from the spec. MCP lead maintainers David Soria Parra and Den Delimarsky published an updated roadmap on August 22, and the re...
Released September 3, v2.38.0 adds context_window to ModelProfile and context_window_used to RunContext (#4611), giving agent code a first-party way to read remaining context instead of estimating from token counts. It also adds a VLLMProvider for self-hosted vLLM servers, Cla...
Hindsight by Vectorize.io builds a knowledge graph from agent interactions rather than storing raw text, modeling how human long-term memory works. 91% on LongMemEval (state-of-the-art). The MCP server makes it a drop-in memory backend for Claude, Cursor, and Windsurf via a si...
Two facts sit next to each other and neither cancels the other out. Anthropic published on September 4 that an internal general-purpose research model, roughly comparable to Claude Fable 5.1, formalized Fermat's Last Theorem in Lean over 11 days working largely autonomously. T...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.