Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 16 days ago • 77
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 24 days ago • 376
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 23 days ago • 327
Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration Paper • 2609.01072 • Published 29 days ago • 13
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 28 days ago • 186
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published Aug 19 • 100
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis Paper • 2608.18940 • Published Aug 19 • 35
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 286
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Paper • 2607.29211 • Published Jul 31 • 14
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries Paper • 2608.05604 • Published Aug 6 • 81
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published Aug 3 • 161
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 158