The datasets behind our Qwen3-VL-8B video RLVR runs: the base 24f/100k mixture and the HopChain v5 multi-hop corpora.
Nguyen Quang Trung
ngqtrung
AI & ML interests
None yet
Recent Activity
liked a Space about 23 hours ago
HuggingFaceTB/smol-training-playbook upvoted a paper 7 days ago
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal
Training updated a collection 8 days ago
Video RLVR — final training dataOrganizations
Qwen3-VL-8B RLVR — Models (v1)
Qwen3-VL-8B GRPO RLVR checkpoints from a token-dropout exploration study. OMR ppexplore=winner (0.714); video ~0.485 dead-heat.
VMAR — Raw & Source
Raw multi-style distilled traces, the pre-distillation template seed, and the never-trained real-audio eval.
videorl
Qwen3-VL-8B video RLVR - GRPO checkpoint ladder
Every published checkpoint of the 8B video-QA GRPO campaign, ranked by core-3 mean_accuracy (5,645 rows). Untrained base = 0.4426.
-
ngqtrung/video-8b-grpo-n16-warm-explore
Video-Text-to-Text • 9B • Updated • 14 -
ngqtrung/video-8b-grpo-base
Image-Text-to-Text • 9B • Updated • 6 -
ngqtrung/video-8b-grpo-ppexplore-n16k8
Image-Text-to-Text • 9B • Updated • 4 -
ngqtrung/video-8b-grpo-ppexplore
Image-Text-to-Text • 9B • Updated • 4
Qwen3-VL-8B RLVR — Datasets (v1)
Curated SFT + GRPO RL datasets (video MC-QA, OMR math-image, OpenMMReasoner-RL, Vero) for Qwen3-VL-8B post-training.
VMAR — Train-Ready (SFT + RL + Eval)
Train-ready VMAR datasets: teacher-distilled SFT corpus, RL prompt set, and the curated in-loop eval benchmark.
Video RLVR — final training data
The datasets behind our Qwen3-VL-8B video RLVR runs: the base 24f/100k mixture and the HopChain v5 multi-hop corpora.
Qwen3-VL-8B video RLVR - GRPO checkpoint ladder
Every published checkpoint of the 8B video-QA GRPO campaign, ranked by core-3 mean_accuracy (5,645 rows). Untrained base = 0.4426.
-
ngqtrung/video-8b-grpo-n16-warm-explore
Video-Text-to-Text • 9B • Updated • 14 -
ngqtrung/video-8b-grpo-base
Image-Text-to-Text • 9B • Updated • 6 -
ngqtrung/video-8b-grpo-ppexplore-n16k8
Image-Text-to-Text • 9B • Updated • 4 -
ngqtrung/video-8b-grpo-ppexplore
Image-Text-to-Text • 9B • Updated • 4
Qwen3-VL-8B RLVR — Models (v1)
Qwen3-VL-8B GRPO RLVR checkpoints from a token-dropout exploration study. OMR ppexplore=winner (0.714); video ~0.485 dead-heat.
Qwen3-VL-8B RLVR — Datasets (v1)
Curated SFT + GRPO RL datasets (video MC-QA, OMR math-image, OpenMMReasoner-RL, Vero) for Qwen3-VL-8B post-training.
VMAR — Raw & Source
Raw multi-style distilled traces, the pre-distillation template seed, and the never-trained real-audio eval.
VMAR — Train-Ready (SFT + RL + Eval)
Train-ready VMAR datasets: teacher-distilled SFT corpus, RL prompt set, and the curated in-loop eval benchmark.
videorl