OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published Jul 26 • 27
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Paper • 2606.07639 • Published Jun 1 • 6
MOSS Transcribe Collection A unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription. • 4 items • Updated Jul 11 • 13
OpenMOSS-Team/MOSS-Transcribe-preview-2B Automatic Speech Recognition • 2B • Updated Jun 26 • 902 • 52