Fetching from the wire…
Skills2026-03-19 · source-backed
Add a cross-encoder after vector retrieval: each (query, chunk) pair scored independently rather than by embedding similarity. 50–200ms latency overhead offset by passing fewer, better chunks to the LLM. At scale, generation savings exceed re-ranker cost. Best options: Cohere Rerank v3.5, Jina Reranker v2, BGE-Reranker-v2-m3. Abhishek Gautam
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / What happened next / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-08-16.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / What happened next
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-08-21.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-08-17.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-08-12.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; picks up the LLM thread on 2026-08-07.