EchoCoT Extracts Hidden Reasoning Traces Near-Verbatim Through a Replay Surface Between Tool Calls
arXiv 2608.20055 identifies a previously unnoticed reasoning-replay surface that opens between tool calls and builds EchoCoT, a multi-step attack that iteratively pulls hidden chain-of-thought out of black-box reasoning models using API-returned fidelity signals, with an LLM-driven search for a universal injection trajectory. On open-source LRMs it reaches 66.4% near-verbatim extraction, with trace length within 10% of the target and at least 90% of tokens matching exactly, and the same trajectory generalizes to unseen datasets at up to 80%. On Gemini-2.5 it extracted 33,463 tokens against a 32,948-token target, and a substantial fraction of traces pulled from five frontier proprietary models matched provider-reported reasoning lengths and CoT summaries.
↳ Follow the thread