VIKI-R
Dataset designed for training visual reasoning models. It combines videos with complex questions with answers and explanations.
Description
VIKI-R is a multi-modal dataset with over 22,000 examples, where each video is associated with an instruction, a complex question, an answer, and a rationale. It specifically targets multimodal reasoning in dynamic contexts, via annotated videos. The Parquet format allows efficient and fast handling.
What is this dataset for?
- Train models that can answer complex questions by analyzing video sequences
- Test the ability to follow instructions and visual causal reasoning
- Evaluate Video-LLM models on multimodal comprehension tasks
Can it be enriched or improved?
Yes, the dataset can be enriched with additional videos or new types of questions (e.g. spatial, temporal, inferential reasoning). It is also possible to add finer linguistic annotations or audio if not present.
🔎 In summary
🧠 Recommended for
- Multimodal AI researchers
- Video-LLMS developers
- Complex visual comprehension projects
🔧 Compatible tools
- Hugging Face Transformers
- PyTorch
- Vid2Seq
- Video-LLM
- MMAction2
💡 Tip
For best results, standardize pre-workout video lengths, and pre-process key frames.
Frequently Asked Questions
Does this dataset contain dialogs?
No, it contains instructions, questions, answers, and justifications, but no dialogue between agents.
Are the videos directly included in the dataset?
No, only metadata and annotations are included. The videos are referenced but must be obtained separately if required.
Can it be used for simple classification tasks?
This is not its primary purpose, but it can be adapted for this by simplifying the tasks based on the annotations provided.




