Fetching from the wire…
Tools2026-07-23 · source-backed
Three frontier models shipped in a single week this month, and teams with a standing eval harness had a routing decision in hours. Anthropic's own agent-eval guidance says 20-50 tasks drawn from your real usage and real failures is enough to detect issues (DeepEval). DeepEval 4.0 (~17k stars, 8M+ monthly PyPI downloads) added a local harness aimed at coding agents specifically; Promptfoo tagged v0.121.19 on July 14 (MIT, 23.3k stars), now under OpenAI while staying open source. Mine your own failure log for 20 tasks today. Don't wait to build a proper benchmark, you won't.
Each link below shares sources, entities, or timing with this story.
Go look at your ~/.claude/CLAUDE.md right now. Mine has internal package names, a build command with a host in it, and notes about which credentials live where. I wrote it assuming exactly one reader. RuntimeWire published traced request captures on August 9 showing Muse Code...
Blaizzy/nativ (1,163 stars, Swift, MIT, macOS 26+) comes from the mlx-vlm author and bundles that server into a SwiftUI app that discovers MLX models already in your HF cache. It exposes OpenAI-compatible chat, Responses, image, audio and model endpoints plus Anthropic Message...
63,012 stars, 12,400 forks, MIT, scanning 100+ pre-configured companies including Anthropic, OpenAI, ElevenLabs, Retool and n8n plus 55+ job board providers, scoring each listing 1.0-5.0 across role summary, CV match, compensation research and personalization strategy. The des...
If you have a CLAUDE.md, you're in scope. Today. arXiv 2607.14611 (cs.CR, filed July 16) evaluates prompt injection planted in the persistent memory files that agentic coding systems write and re-read across sessions. The researchers tested both Anthropic's Claude Code and Ope...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
Your Claude subscription is about to get a lot more expensive if you're running agents programmatically. Starting June 15, Anthropic is decoupling all programmatic usage (Agent SDK, claude -p, Claude Code terminal) from the interactive subscription pool. Instead of eating from...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.