Sources
DR-Venus-4B: Edge-Scale Deep Research Agent Trained on Only 10K Open Data Points
A new arXiv paper (2604.19859) introduces DR-Venus-4B, a 4-billion-parameter deep research agent built entirely on ~10K open-data samples using a two-stage recipe: agentic supervised fine-tuning followed by agentic reinforcement learning with turn-level rewards. It significantly outperforms prior sub-9B agentic models on deep research benchmarks and narrows the gap to 30B-class systems. Models, code, and training recipes are open-sourced — a strong signal that small models can do serious autonomous research with the right training recipe.
Source
↳ Follow the thread