RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 9 days ago • 277
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models Paper • 2610.03665 • Published 8 days ago • 62
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop? Paper • 2609.35025 • Published 12 days ago • 9
Org-Agent: Beyond Personal Assistants Towards Organizational Agents Paper • 2609.34392 • Published 12 days ago • 7
Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics Paper • 2609.36608 • Published 11 days ago • 8
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 14 days ago • 326
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models Paper • 2609.32607 • Published 14 days ago • 155
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression Paper • 2609.36322 • Published 12 days ago • 115
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 11 days ago • 103