World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation Paper • 2608.05369 • Published 7 days ago • 25
Invisible Shortcuts: Why Vision Encoders Know Your Camera Paper • 2608.05424 • Published 7 days ago • 16
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 6 days ago • 43
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 6 days ago • 58
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 8 days ago • 90
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing Paper • 2607.29025 • Published 12 days ago • 16
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 13 days ago • 182