WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 14 days ago • 142
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 19 days ago • 274
Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs Paper • 2505.11277 • Published May 16, 2025 • 64
NeuroCogMap Reveals Cognitive Organization of Large Language Models Paper • 2607.00397 • Published Jul 1 • 12
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems Paper • 2606.18837 • Published Jun 17 • 59