AREX: BAAI Trains a Recursively Self-Improving Deep-Research Agent That Compresses Its Own History Into an 'Improvement State'
AREX (arXiv 2607.21461, submitted 23 July, 24 authors led by Shuqi Lu, 124 upvotes on HuggingFace Daily Papers) alternates between gathering evidence and drafting provisional answers, then audits those answers constraint-by-constraint. Its distinguishing mechanism is a learned autonomous context-update tool that compresses growing interaction history into a compact improvement state without calling an external model — self-improvement inside the context loop rather than via a critic model. The team trained a 4B dense and a 122B-A10B MoE variant on synthetic tasks and trajectories, reporting wins over comparable-scale baselines on BrowseComp, WideSearch, DeepSearchQA and Humanity's Last Exam while staying competitive with models using far more activated parameters.
↳ Follow the thread