FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
Abstract
FIRM-Video uses checklist-driven verification of temporal visual evidence to build reliable reward models for text-to-video evaluation and alignment.
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
Community
FIRM-Video introduces a checklist-driven, check-before-score framework for reliable and efficient text-to-video reward modeling. It decomposes evaluation into verifiable criteria for instruction following, world coherence, and perceptual quality, and builds FIRM-Video-90K and FIRM-Video-Bench. FIRM-Video-8B achieves the best overall MAE on the benchmark and consistently improves Best-of-8 video selection across multiple generators.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing (2026)
- VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation (2026)
- Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models (2026)
- Evidence-Backed Video Question Answering (2026)
- Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation (2026)
- LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models (2026)
- DynEval: Holistic Evaluations of T2I Generative Models in the Wild (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.21839 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper