AREX-2: Long-Horizon Reflection for Self-Improving LLM Agents
AREX-2 advances self-improving LLM agents by introducing reflection and long-horizon execution, demonstrating significant gains on both algorithmic and research tasks through iterative solution refinement.
AREX-2 is a new approach to self-improving large language model (LLM) agents, focusing on the agent's ability to refine its solutions iteratively at test time. This method introduces two key capabilities: reflection (the ability to improve upon current solutions) and long-horizon execution (the ability to sustain effective iteration over many rounds). AREX-2 is trained on data synthesized from machine learning and algorithmic tasks, leveraging scenarios with verifiable feedback to drive its self-improvement.

Key Capabilities: Reflection and Long-Horizon Execution
Reflection enables the agent to generate successively better solutions, while long-horizon execution ensures these improvements persist over extended interaction. The research shows these capabilities are domain-agnostic and effective for a wide range of tasks.
Evaluation Results
AREX-2, built on Qwen3.8-27B, achieves high scores on various benchmarks: 81.8 on MLE-bench Lite, 70.7 on Frontier-CS, as well as strong results in deep research domains, such as 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA. Performance continues to improve as the agent is allowed to iterate over more rounds.
