SmolDataEnvs: 5,000+ verified data-science RL environments for training sub-10B agent models on one GPU
Hugging Face (FineEnvs)·low signal
The FineEnvs SmolDataEnvs release on Hugging Face has a 5,000-task training suite, a 144-task validation suite sized for frequent checks, and a 250-task test suite weighted toward hard problems. The tasks come from real Kaggle notebooks in the jupyter-agent dataset over 471 datasets. Each question-answer pair was verified by having strong agents reproduce the gold answer in a live sandbox under deterministic grading. SFT and GRPO 2B baseline checkpoints are published, so small-model RL for data analysis can be tried on a single GPU.