Submitted by KAI LIU 20 JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation JavisVerse 69 3