Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published Aug 8 • 30
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 46
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 46
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Paper • 2606.14502 • Published Jun 12 • 118
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations Paper • 2602.05885 • Published Feb 5 • 28
Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics Paper • 2601.14027 • Published Jan 20 • 15
How Far Are We from Genuinely Useful Deep Research Agents? Paper • 2512.01948 • Published Dec 1, 2025 • 58
How Far Are We from Genuinely Useful Deep Research Agents? Paper • 2512.01948 • Published Dec 1, 2025 • 58
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity Paper • 2511.08487 • Published Nov 11, 2025 • 4