By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
WM-ABench
Video

WM-ABench

WM-ABench is a benchmark intended to test the fine understanding of vision-language models on the modeling of physical dynamics through simulated videos. It covers 23 different dimensions of modeling the world, with complex perception and prediction tasks.

Download dataset
Size

Over 100,000 annotated video sequences with multiple choice questions from 6 physical simulators

Licence

Apache 2.0

Description

The dataset WM-ABench contains over 100,000 video instances from 6 physical simulation environments. Each example includes video sequences depicting object interactions, accompanied by multiple choice questions with similar but incorrect options, in order to assess the real physical understanding of vision-language models.

What is this dataset for?

  • Evaluate the spatial, temporal, and material perception abilities of vision-language models
  • Test the ability to predict physical dynamics (falls, collisions, complex actions)
  • Serve as a basis for fine-tuning or the validation of advanced models in vision and multimodal language

Can it be enriched or improved?

Yes, it is possible to add additional annotations such as finer object metadata, complexity labels, or to introduce new simulated scenes. The dataset can also be extended by remixing or combining with other video benchmarks to improve diversity.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Standard format with available API, requires multimedia knowledge)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – simulated and well-structured data)
🏷️ Annotation richness⭐⭐⭐⭐⭐ (Excellent – multiple-choice questions with "hard negatives" for fine evaluation)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Moderate – better with experience in computer vision and NLP
🔁 Fine-tuning ready🎯 Perfect for fine-tuning multimodal vision-language models
🌍 Cultural diversity⚠️ Limited – dataset focused on physical simulations, few cultural elements

🧠 Recommended for

  • Computer vision researchers
  • Multimodal model developers
  • AI teams in robotics

🔧 Compatible tools

  • Hugging Face Datasets
  • PyTorch
  • TensorFlow
  • Compatible physical simulators

💡 Tip

Use the smaller subsets for quick testing before starting the full training.

Frequently Asked Questions

What types of questions are asked to models in WM-Abench?

Multiple choice questions dealing with visual perception and the prediction of physical dynamics in simulated scenes.

Can this dataset be used for fine-tuning vision-language models?

Yes, it is perfectly suited for fine-tuning or evaluating multimodal models combining vision and language.

What is the approximate size of the dataset?

Over 100,000 annotated video instances, offering a wide variety of examples for training and evaluation.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.