WM-ABench
WM-ABench is a benchmark intended to test the fine understanding of vision-language models on the modeling of physical dynamics through simulated videos. It covers 23 different dimensions of modeling the world, with complex perception and prediction tasks.
Over 100,000 annotated video sequences with multiple choice questions from 6 physical simulators
Apache 2.0
Description
The dataset WM-ABench contains over 100,000 video instances from 6 physical simulation environments. Each example includes video sequences depicting object interactions, accompanied by multiple choice questions with similar but incorrect options, in order to assess the real physical understanding of vision-language models.
What is this dataset for?
- Evaluate the spatial, temporal, and material perception abilities of vision-language models
- Test the ability to predict physical dynamics (falls, collisions, complex actions)
- Serve as a basis for fine-tuning or the validation of advanced models in vision and multimodal language
Can it be enriched or improved?
Yes, it is possible to add additional annotations such as finer object metadata, complexity labels, or to introduce new simulated scenes. The dataset can also be extended by remixing or combining with other video benchmarks to improve diversity.
🔎 In summary
🧠 Recommended for
- Computer vision researchers
- Multimodal model developers
- AI teams in robotics
🔧 Compatible tools
- Hugging Face Datasets
- PyTorch
- TensorFlow
- Compatible physical simulators
💡 Tip
Use the smaller subsets for quick testing before starting the full training.
Frequently Asked Questions
What types of questions are asked to models in WM-Abench?
Multiple choice questions dealing with visual perception and the prediction of physical dynamics in simulated scenes.
Can this dataset be used for fine-tuning vision-language models?
Yes, it is perfectly suited for fine-tuning or evaluating multimodal models combining vision and language.
What is the approximate size of the dataset?
Over 100,000 annotated video instances, offering a wide variety of examples for training and evaluation.




