Fetching from the wire…
Tools2026-07-26 · source-backed
The July 22 post describes an installable skill that reads your repo to map the agent surface (prompts, models, tools, skills, hooks), pulls production traces via langsmith-cli, then interviews you to converge on evals you approve one at a time (LangChain). Output is a containerized Harbor task per eval under evals/<task-id>/ with instruction.md, an environment/ Dockerfile, and tests/ verifier logic. The reported lesson: single-pass generation produced weak evals, and the good ones came from inspecting both the agent's and the verifier's trajectories to catch reward-hacked shortcuts.
Each link below shares sources, entities, or timing with this story.
LangChain released LangGraph / Shared entity: LangChain / Same source domain
Linked by a graph relationship (LangChain released LangGraph); both cover LangChain; reported by the same outlet (langchain.com).
LangChain released LangGraph / Shared entity: LangChain / Earlier coverage
Linked by a graph relationship (LangChain released LangGraph); both cover LangChain; earlier LangChain coverage from 2026-03-19.
Linked by a graph relationship (LangChain released LangGraph); both cover LangChain; earlier LangChain coverage from 2026-06-02.
output uses Claude Code / Shared entity: LangChain / Earlier coverage / Tension
Linked by a graph relationship (output uses Claude Code); both cover LangChain; earlier LangChain coverage from 2026-03-25.
LangChain released LangGraph
Linked by a graph relationship (LangChain released LangGraph).
Linked by a graph relationship (LangChain released LangGraph).
Linked by a graph relationship (LangChain released LangGraph).
Linked by a graph relationship (LangChain released LangGraph).