Hugging Face TRL v1.0: Post-Training Library Hits Major Milestone with SFT, GRPO, and DPO Support
Hugging Face Blog·medium signal
Hugging Face's Transformer Reinforcement Learning (TRL) library reaches v1.0, marking a stable release for post-training foundation models using Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO). The library is built on the Transformers ecosystem and supports multi-modal architectures across various hardware setups. Migration from v0.x to v1.0 is minimal for users already on v0.29.