MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use Paper • 2608.20202 • Published 14 days ago • 34
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence Paper • 2608.12036 • Published 22 days ago • 86
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies Paper • 2510.11389 • Published Oct 13, 2025
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published Jul 30 • 34
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published Jul 30 • 39
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published Jul 30 • 34
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published Jul 30 • 39
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published Jul 27 • 33
A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces Paper • 2602.03442 • Published Feb 3 • 22
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 62
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 62
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Paper • 2606.26058 • Published Jun 24 • 67
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Paper • 2606.26058 • Published Jun 24 • 67
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills Paper • 2605.24117 • Published May 22 • 23
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills Paper • 2605.24117 • Published May 22 • 23
CARPE: Context-Aware Image Representation Prioritization via Ensemble for Large Vision-Language Models Paper • 2601.13622 • Published Jan 20 • 1
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization Paper • 2604.07343 • Published Apr 8 • 13
Learning to Self-Verify Makes Language Models Better Reasoners Paper • 2602.07594 • Published Feb 7 • 3