arXiv: Many-Shot CoT-ICL — In-Context Learning Breaks Down for Reasoning Tasks, Curvilinear Demonstration Selection Proposed
arXiv·medium signal
New paper (2605.13511) shows that established rules of many-shot in-context learning fail for reasoning tasks: similarity-based retrieval doesn't work because question similarity doesn't ensure procedural compatibility, and performance variance grows with more demonstrations (order-scaling effect). The authors reframe effective many-shot CoT as in-context test-time learning and propose Curvilinear Demonstration Selection (CDS), yielding up to 5.42 percentage-point gains. Directly relevant to anyone building RAG or few-shot pipelines.