WorldSculpt: Generating Compositional Worlds from Grounded Videos Paper • 2609.05416 • Published 6 days ago • 23
Marionette: Predicting World States, Rendering Geometry, Painting Appearance Paper • 2608.14530 • Published 27 days ago • 34
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Paper • 2608.13546 • Published 28 days ago • 166
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow Paper • 2607.28362 • Published Jul 30 • 22
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published Jul 20 • 61
AlayaWorld: Long-Horizon and Playable Video World Generation Paper • 2607.06291 • Published Jul 7 • 92
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published Jul 2 • 70
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference Paper • 2603.25730 • Published Mar 26 • 53
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents Paper • 2508.13186 • Published Aug 14, 2025 • 21
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams Paper • 2508.06851 • Published Aug 9, 2025