CreativeInstruct restores base-model diversity with [StartCreativity] spans — and it makes downstream RL measurably better
Submitted 2026-08-07 by Sahu, Bansal, and Stengel-Eskin (arXiv:2608.07460), CreativeInstruct is an instruction-tuning method that teaches a model to inject special [StartCreativity] spans biasing generation toward base-model-like variety, plus a structural diversity metric based on graph edit distance that catches narrative variation lexical and semantic metrics miss. Annotators rated its output more creative than the post-trained model's in 70.3% of cases with no quality loss and no multi-model inference. The agent-relevant result is downstream: GRPO applied to a CreativeInstruct checkpoint gained ~4% on AMC and ~5 points on MATH over the same training on the post-trained checkpoint — diversity as RL substrate, not decoration.
Source
↳ Follow the thread