arxiv:2508.03320
Hongyang Wei
nonwhy
AI & ML interests
multimodal large language models & low-level vision
Recent Activity
upvoted a paper 3 days ago
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds upvoted a paper 3 days ago
EchoWM: Open and Enterable Omnimodal World Models upvoted a paper 21 days ago
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing