Fetching from the wire…
Top 5 · 2026-04-02 · source-backed
What if the chain-of-thought isn't driving the answer? What if it's a post-hoc story the model tells itself?
A new paper on arXiv titled "Therefore I Am. I Think" ran linear probes on reasoning model internals and found something uncomfortable. Tool-calling decisions are detectable from pre-generation activations with high confidence. Sometimes before a single reasoning token is produced. The model has already made up its mind. The chain-of-thought that follows is shaped by that prior decision, not the other way around.
This matters because the entire pricing model for "thinking" LLMs is built on the assumption that more reasoning tokens equals better output. OpenAI charges more for o-series models. Anthropic's extended thinking burns 10 to 50x more tokens than standard responses. The reasoning is supposed to be doing work. If it's not, if it's rationalization rather than computation, then a meaningful chunk of what we're paying for is theater.
I want to be careful here. The paper doesn't prove that CoT is useless. It proves that for certain decision types (specifically tool calling), the outcome is already encoded before reasoning begins. That could mean the reasoning is confirmatory rather than exploratory. It doesn't necessarily mean removing CoT would produce the same results. It could be serving a different function than we think, like constraint verification or consistency checking, even if it's not the primary decision mechanism.
But connect this to the Stack Overflow trust data (story #4 below). 84% of developers use AI tools. 3% strongly trust the output. Maybe that distrust is well-calibrated. We don't just lack trust in what these models produce. We may lack understanding of how they produce it. If the reasoning trace isn't what's driving quality, then we can't use it to evaluate quality either. The explanation isn't explaining.
For builders using reasoning models: don't assume more thinking tokens means better results. Benchmark your specific use case with and without extended reasoning. If the quality gap is small, you might be burning tokens on rationalization. For agent builders doing tool calling specifically, this paper suggests the decision is already made by the time you see the reasoning. Your prompt engineering should focus on what goes into the model's context, not on coaxing better reasoning chains out.
Each link below shares sources, entities, or timing with this story.
Anthropic partners with OpenAI / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Anthropic partners with OpenAI); both cover Anthropic, CoT, OpenAI, Tool; overlapping topics (model, reasoning, tool).
Anthropic released MCP / Shared entities / Same source domain / Shared topic / What happened next / Tension
Linked by a graph relationship (Anthropic released MCP); both cover Anthropic, OpenAI; reported by the same outlet (arxiv.org).
Mozilla partners with Anthropic / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Mozilla partners with Anthropic); both cover Anthropic, OpenAI, Tool; overlapping topics (model, tool).
Anthropic partners with OpenAI / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Anthropic partners with OpenAI); both cover OpenAI, Think; overlapping topics (already, calling, model, token, tool).
Anthropic released Claude Code / Shared entities / Shared topic / Tension
Linked by a graph relationship (Anthropic released Claude Code); both cover Stack Overflow, Therefore, Think; overlapping topics (tool, trust).
Opus built by Anthropic / Shared entities / Shared topic / What happened next / Tension
Linked by a graph relationship (Opus built by Anthropic); both cover Anthropic, Think; overlapping topics (better, decision, model).
Anthropic partners with OpenAI / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Anthropic partners with OpenAI); both cover Anthropic, OpenAI; reported by the same outlet (arxiv.org).
Anthropic partners with Google / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Anthropic partners with Google); both cover Anthropic, OpenAI; reported by the same outlet (arxiv.org).