RL Training Compute Follows Power Law with Reasoning Depth — Expressiveness Is the Key Variable
arXiv·medium signal
ScaleLogic benchmark reveals RL training compute scales as T proportional to D^gamma with reasoning depth, where the exponent gamma ranges from 1.04 to 2.60 depending on logical expressiveness. Training on more expressive logical systems yields up to 10.66-point downstream improvement — what a model trains on, not just how much, determines transfer. Curriculum training further improves scaling efficiency.