AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents Paper • 2610.05140 • Published 4 days ago • 23
When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models Paper • 2610.05719 • Published 3 days ago • 14
Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers Paper • 2610.00531 • Published 8 days ago • 57
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 8 days ago • 117
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 8 days ago • 117
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Paper • 2609.40340 • Published 8 days ago • 109
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 11 days ago • 46
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge Paper • 2609.34327 • Published 10 days ago • 41
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 22 days ago • 67
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents Paper • 2609.17632 • Published 23 days ago • 46
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published about 1 month ago • 147
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 30 days ago • 84
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published Aug 29 • 29
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Paper • 2608.07169 • Published Aug 7 • 50
Environment-free Synthetic Data Generation for API-Calling Agents Paper • 2607.16900 • Published Jul 18 • 21
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks Paper • 2606.29082 • Published Jun 27 • 44
One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications Paper • 2606.25621 • Published Jun 24 • 20
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 167
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents Paper • 2604.24005 • Published Apr 27 • 9