Text2Video — Pika 2.2 Human Preferences
A massive dataset of human video evaluations generated from textual descriptions, useful for training preference systems or refining video models.
756,000 responses (JSON), low resolution GIF videos, and links to HD videos
Apache 2.0
Description
Text-2-Video Human Preferences — Pika 2.2 is a data set containing approximately 756,000 human responses collected in 24 hours from 29,000 annotators. These responses compare videos generated by the Pika 2.2 model from textual descriptions, in a pair format (video1 vs video2), making it possible to measure quality, fidelity, and alignment with the prompt.
What is this dataset for?
- Objectively evaluate text-based video generation models
- Train preference or alignment models (e.g. reward models in RLHF)
- Analyze the current limitations of broadcast models or transform into video
Can it be enriched or improved?
Yes. You can integrate more metadata (duration, content category, perceived quality) or cross-reference responses with automatic metrics. It is also possible to add natural language layers (justifications, comments) or to extend the benchmark to other models.
🔎 In summary
🧠 Recommended for
- RLHF developers
- Video generation researchers
- Evaluation of multimodal models
🔧 Compatible tools
- PyTorch
- Hugging Face Datasets
- Weights & Biases
- PEFT
- VLLM
💡 Tip
Use weighted responses (“weighted_results”) to train a custom scoring model.
Frequently Asked Questions
Does this dataset make it possible to evaluate models other than Pika?
Yes, videos can be replaced by videos from other models to reuse the assessment structure.
Are the videos downloadable?
Lightweight GIFs are included, and links to HD videos are provided separately in the dataset.
Can this dataset be used to train a preference or reward model?
Absolutely, it is one of its main uses, especially in RLHF or assisted human evaluation pipelines.




