CG Bench
CG-Bench is a benchmark designed to assess the in-depth understanding of long-form videos via questions and answers anchored to specific clues in the videos.
1,219 manually annotated videos, 12,129 question and answer pairs, standard video formats
MIT
Description
CG-Bench contains 1,219 long videos annotated with 12,129 question-answer pairs, covering a variety of categories (14 primary, 171 secondary, 638 tertiary). The dataset focuses on questions that require retrieving relevant clues from the videos in order to assess the real understanding of the models.
What is this dataset for?
- Test and improve video comprehension of multimodal language models (MLLMs)
- Evaluate the ability of models to reason over long video sequences based on specific clues
- Develop robust video question and answer systems beyond simple multiple choices
Can it be enriched or improved?
Yes, it's possible to add new annotations, extend video categories, or refine hints for more granular ratings. The community can also propose additional data or alternative annotation formats to enrich the benchmark.
🔎 In summary
🧠 Recommended for
- Computer vision researchers
- MLLMS developers
- Video QA projects
🔧 Compatible tools
- PyTorch
- TensorFlow
- Hugging Face Transformers
- Video frameworks like OpenCV
💡 Tip
Focus on effective video preprocessing to optimize speed during training phases.
Frequently Asked Questions
Is this dataset adapted to commercial and open-source models?
Yes, CG-Bench compares the performance of the two types of models and is designed to improve the capabilities of both.
Do annotations only include multiple choice questions?
No, unlike traditional benchmarks, CG-Bench uses open-ended questions that require a thorough understanding.
Does the dataset include videos of varying lengths and categories?
Yes, with over 1,200 videos covering 14 main categories and multiple sub-categories to ensure significant diversity.




