CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
Abstract
CLIP-CC-Bench evaluates long-form video description using expert paragraph references and ensemble LLM embeddings to benchmark 17 video-language models.
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.
Community
Can a video-language model write a faithful paragraph about a minute of film?
Video benchmarks still lean on short clips, one-sentence captions, and multiple-choice QA โ none of
which test long-form description. CLIP-CC-Bench does.
๐ฌ 200 movie clips (~90 s each, 5 h total, 140+ films from 1959โ2024), each paired with a
human-narrated reference paragraph averaging ~400 words. Narrators describe only what is on screen โ
no proper nouns โ so recognising the movie earns a model nothing.
โ๏ธ 17 video-language models scored by an ensemble of 5 embedding judges, combining
paragraph-level (coarse) and sentence-level (fine) matching into one harmonic mean, aggregated with
Borda count.
๐ Models get the story, not the details. Coarse beats fine for essentially every model, and the
best system reaches a mean HM-CF of only 0.67 โ long-form video description is far from solved.
๐ We also stress-test the protocol itself: the five judges agree on rankings at mean Spearman
0.98, and across 1,000 bootstrap resamples of the clip set the top model stays #1 every
time โ so the leaderboard is a property of the task, not of which 200 clips we happened to pick.
Dataset, per-clip scores for all 17 models, and the full evaluation pipeline are released.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models (2026)
- Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning (2026)
- Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models (2026)
- CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales (2026)
- An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations (2026)
- Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval (2026)
- ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.04302 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper