WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Paper • 2608.27454 • Published 19 days ago • 33
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 22 days ago • 207
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 26 days ago • 66
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Paper • 2608.18580 • Published 27 days ago • 121
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 26 days ago • 275
Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published Aug 14 • 170
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published Aug 15 • 446
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search Paper • 2608.15669 • Published about 1 month ago • 63
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 29 days ago • 151
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published Aug 13 • 63
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 111
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published Jul 31 • 113
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published Aug 10 • 779
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Paper • 2608.08311 • Published Aug 8 • 92
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 343
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published Aug 10 • 136