-
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
Paper • 2603.25319 • Published • 32 -
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Paper • 2603.25040 • Published • 126 -
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
Paper • 2603.22458 • Published • 134 -
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
Paper • 2603.21986 • Published • 121
Cai
Mona8834
AI & ML interests
None yet
Recent Activity
updated a collection 10 days ago
LLM Papers updated a collection 14 days ago
LLM Papers updated a collection 14 days ago
LLM PapersOrganizations
None yet