view article Article Your Inference Server is Secretly a Learner: Reef Infrastructure for Continual Self-Improving Agents quao627 • 5 days ago • 15
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis Paper • 2609.15309 • Published 6 days ago • 12
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis Paper • 2609.15309 • Published 6 days ago • 12
Benchmarking Visual State Tracking in Multimodal Video Understanding Paper • 2606.03920 • Published Jun 2 • 53
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Paper • 2605.14445 • Published May 14 • 21
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Paper • 2605.08678 • Published May 9 • 9