multimodal-vision-ai/mvai-doctag-r3-v1
Updated • 306
Multimodal AI, Vision-Language Models, Document Intelligence, OCR, Handwriting Recognition, Compact VLMs, Multimodal Post-Training, SFT, GRPO
Hohai University
The Multimodal Vision AI Lab focuses on research in multimodal artificial intelligence, vision-language models, document intelligence, OCR, handwriting recognition, compact multimodal models, and multimodal post-training.
We develop open datasets, models, benchmarks, evaluation tools, and educational resources for multimodal vision and document intelligence research.
Projects, datasets, models, and benchmarks will be released progressively through this organization.