Fetching from the wire…
Tools2026-05-11 · source-backed
Instead of end-to-end black-box testing, you can now assess individual retrievers, tool calls, generators, and agent interactions within a traced pipeline. The pattern: build golden datasets from real production failures (200-500 examples), not synthetic data.
Each link below shares sources, entities, or timing with this story.
LLM uses OpenAI / Shared entity: DeepEval / Same source domain / Shared topic / What happened next / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover DeepEval; reported by the same outlet (deepeval.com).
Simon Willison released LLM / Shared entity: LLM / Shared topic / What happened next / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, call).
LLM uses OpenAI / Shared entity: LLM / Shared topic / What happened next / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; overlapping topics (agent, pipeline).
Simon Willison released LLM / Shared entity: LLM / Shared topic / What happened next
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, call).
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, interaction).
Simon Willison released LLM / Shared entity: LLM / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, pattern).
LLM uses OpenAI / Shared entity: LLM / What happened next / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; picks up the LLM thread on 2026-07-27.
Simon Willison released LLM / Shared entity: LLM / What happened next / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-06-19.