By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Multimodal

VIKI-R

Dataset designed for training visual reasoning models. It combines videos with complex questions with answers and explanations.

Download dataset
Size

22,637 examples, 813 MB, Parquet format

Licence

Apache 2.0

Description

‍

VIKI-R is a multi-modal dataset with over 22,000 examples, where each video is associated with an instruction, a complex question, an answer, and a rationale. It specifically targets multimodal reasoning in dynamic contexts, via annotated videos. The Parquet format allows efficient and fast handling.

‍

‍

What is this dataset for?

‍

  • Train models that can answer complex questions by analyzing video sequences
  • Test the ability to follow instructions and visual causal reasoning
  • Evaluate Video-LLM models on multimodal comprehension tasks

‍

‍

Can it be enriched or improved?

‍

Yes, the dataset can be enriched with additional videos or new types of questions (e.g. spatial, temporal, inferential reasoning). It is also possible to add finer linguistic annotations or audio if not present.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Easy to load and well-structured)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (No major cleaning needed)
🏷️ Annotation richness⭐⭐⭐⭐⭐ (Includes answer + justification)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Better for users with multimodal vision background
🔁 Fine-tuning ready🎯 Yes, for VideoQA and reasoning
🌍 Cultural diversity⚠️ Medium – depends on video content

‍

‍

🧠 Recommended for

  • Multimodal AI researchers
  • Video-LLMS developers
  • Complex visual comprehension projects

‍

‍

🔧 Compatible tools

  • Hugging Face Transformers
  • PyTorch
  • Vid2Seq
  • Video-LLM
  • MMAction2

‍

‍

💡 Tip

For best results, standardize pre-workout video lengths, and pre-process key frames.

Frequently Asked Questions

Does this dataset contain dialogs?

No, it contains instructions, questions, answers, and justifications, but no dialogue between agents.

Are the videos directly included in the dataset?

No, only metadata and annotations are included. The videos are referenced but must be obtained separately if required.

Can it be used for simple classification tasks?

This is not its primary purpose, but it can be adapted for this by simplifying the tasks based on the annotations provided.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.