MSR-VTT
MSR-VTT is a video dataset for the automatic generation of descriptions in natural language. It contains 10,000 video clips covering 20 varied categories, each annotated with 20 sentences in English, providing a rich corpus for training and evaluating video captioning models.
10,000 video clips, around 200,000 descriptive sentences, common video formats (mp4, avi...)
Free, download available via Microsoft
Description
MSR-VTT is a large annotated video corpus for the automatic captioning task. Each video clip is accompanied by twenty sentences in English describing the content. The videos cover a wide variety of everyday categories.
What is this dataset for?
- Train video captioning models in natural language
- Assess multimodal understanding (video + text)
- Develop intelligent assistants that can describe video scenes
Can it be enriched or improved?
Yes, you can add additional annotations such as action tags, named entities, or multilingual translations to enrich the use of the dataset.
🔎 In summary
🧠 Recommended for
- Vision and language researchers
- Multimodal AI developers
- Students
🔧 Compatible tools
- PyTorch
- TensorFlow
- Hugging Face Transformers
- OpenCV
💡 Tip
Pre-process videos to standardize the pre-workout format.
Frequently Asked Questions
What is the average length of video clips in MSR-VTT?
Approximately 10 to 30 seconds per clip, covering a variety of scenes.
Are the annotations all in English?
Yes, all descriptions are in English.
Can MSR-VTT be used for multimodal translation?
Yes, when combined with translation datasets, it can be used as a basis for multi-modal, multilingual tasks.




