RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents Paper • 2606.19047 • Published Jun 17 • 4
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification Paper • 2606.01476 • Published May 31 • 9
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 239
PANDO: Efficient Multimodal AI Agents via Online Skill Distillation Paper • 2605.24785 • Published May 26 • 11
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration Paper • 2605.17423 • Published May 17 • 34
Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs Paper • 2505.11277 • Published May 16, 2025 • 29
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Paper • 2605.21467 • Published May 20 • 207