Fetching from the wire…
Skills2026-07-30 · source-backed
CoRT rescores each response twice, with and without the rubric in context, and uses the counterfactual likelihood contrast as a proxy for which tokens actually depend on the rubric, giving token-level credit with no auxiliary scoring model to train. 4.4 percentage points over response-level GRPO. The trick generalizes past RL: anywhere you're running LLM-judge or rubric-scored evaluation, that delta tells you what the rubric is actually doing.
Each link below shares sources, entities, or timing with this story.
LLM uses OpenAI / Shared entity: GRPO / Same source domain / Earlier coverage
Linked by a graph relationship (LLM uses OpenAI); both cover GRPO; reported by the same outlet (arxiv.org).
LLM uses OpenAI / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-07-27.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
LLM uses OpenAI / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-06-19.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-14.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-22.