When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 4 days ago • 96
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 7 days ago • 209
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published 17 days ago • 114
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 10 days ago • 42
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 10 days ago • 256
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 7 days ago • 239
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 16 days ago • 163
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models Paper • 2609.12641 • Published 10 days ago • 70
SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution Paper • 2609.05594 • Published 17 days ago • 36