Measuring CoT Faithfulness Depends on How You Measure: Classifier Sensitivity Exposed
arXiv·medium signal
Reported chain-of-thought faithfulness metrics—such as 'DeepSeek-R1 acknowledges hints 39% of the time'—are highly sensitive to the choice of classifier and evaluation framing, making cross-paper comparisons unreliable. Single aggregate faithfulness numbers mask substantial variance depending on classifier architecture and prompting. Practitioners evaluating reasoning models need to treat published CoT faithfulness scores with significant skepticism and run their own multi-classifier evaluations.