CompoWorld: Compositional Environment Scaling for General Agents Paper • 2609.33665 • Published 14 days ago • 41
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks Paper • 2609.38288 • Published 12 days ago • 137
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 11 days ago • 64
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published Sep 4 • 119
CUADebug: Diagnosing and Repairing Computer-Use Agent Failures Paper • 2608.02643 • Published Sep 6 • 3
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Paper • 2608.08621 • Published Aug 9 • 22
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published Jul 30 • 75
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales Paper • 2606.12736 • Published Jun 10 • 6
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Paper • 2509.25454 • Published Sep 29, 2025 • 148