Reasoning Core Releases 50 Procedural Problem Generators and Shows Semantic Validity Alone Doesn't Make Training Data Useful
arXiv 2608.05148·medium signal
Reasoning Core is a public collection of 50 procedural generators spanning mathematics, logic, planning, state tracking, formal languages, structured data, games, causality, and code, each with semantic scorers, difficulty controls, and task evaluators, aimed at completion-supervised fine-tuning rather than the RL settings procedural data usually serves. Under a matched protocol across four base-model settings and multiple training durations, the 3B primary comparison beats the no-procedural-data baseline and all three alternative collections — Procedural Warmup, Reasoning Gym, and SynLogic — on mean DROP, LogiQA, and ARC-Challenge. Their audits, combining model-assisted review, human adjudication, and regression testing, found subtle mismatches among generation, rendering, targets, and scoring in these collections, reinforcing that a generator being semantically valid does not make its output good training data.