OpenS2V-5M
Massive multimodal dataset combining subject, text, and video (720p) to train advanced video generation models.
5 million subject-text-video triples in 720p, JSON for metadata and annotations (masks, bboxes in RLE)
CC-BY 4.0
Description
The dataset OpenS2V-5M contains five million triples composed of a subject, text, and video in high definition 720p. It includes detailed metadata such as resolution, frames per second, and aesthetic scores, as well as masks and boxes encompassing topics in RLE format. This dataset is ideal for training models to generate videos based on text or subject inputs.
What is this dataset for?
- Train video generation models based on topics or text descriptions
- Develop AI applications for multi-view video synthesis
- Experiment with multimodal models combining text, subject, and video
Can it be enriched or improved?
Yes, this dataset can be customized by dynamically extracting images of faces to improve the quality of human subjects. It is possible to add additional annotations or to modify the quality thresholds via the scores provided to adjust the quality/quantity compromise. Scripts and tools are available to facilitate these enhancements.
🔎 In summary
🧠 Recommended for
- Video generation researchers
- Multimodal developers
- Advanced AI projects
🔧 Compatible tools
- PyTorch
- TensorFlow
- Diffusers
- Video generation frameworks
💡 Tip
Use aesthetic scores to filter videos and optimize the quality of training data.
Frequently Asked Questions
What type of video content does this dataset contain?
It contains 720p HD videos associated with extracted human subjects (faceless heads) and text descriptions.
Is it possible to customize the annotations in the dataset?
Yes, thanks to the JSON files and the scripts provided, you can extract or modify the masks and surrounding boxes according to your needs.
Is this dataset suitable for beginners in video generation?
No, its volume and technical complexity make it more suitable for experienced users.




