Introspective X Training: NVIDIA's Feedback-Conditioned Scaling Method Tested at 18 Trillion Tokens
arXiv·medium signal
From NVIDIA/UW/CMU/UCSD, proposes IXT inspired by offline reward-conditioned RL: a thinking reward model annotates training data with natural language critique feedback, enabling quality-aware training from the earliest pipeline stages. Validated on 7.5-12B dense transformers trained up to 18T tokens seen. Applicable to any training stage (pretraining, SFT, RL), not just post-training.