OmniSpatial — Multimodal Spatial Reasoning Benchmark
OmniSpatial is a benchmark designed to test the spatial reasoning abilities of vision-language models. The dataset includes multiple choice questions that cover four types of tasks: dynamic reasoning, spatial interaction, complex logic, and perspective taking. Each question is linked to an image and includes options with the correct answer shown.
Questions annotated in JSON format with structured schema, several thousand presumed items
Apache 2.0
Description
OmniSpatial proposes a set of questions structured in JSON, intended to assess the spatial reasoning of multimodal models combining vision and language. The tasks cover motion analysis, spatial interactions, complex logic, and perspective taking.
What is this dataset for?
- Test the spatial and dynamic understanding of VLM models
- Evaluate multi-stage reasoning ability in complex visual contexts
- Improve the robustness of models on advanced vision-language tasks
Can it be enriched or improved?
The dataset can be enriched by adding new images, scenarios, and question types, in particular by refining task subcategories. Community contribution can broaden the diversity of spatial and dynamic contexts.
🔎 In summary
🧠 Recommended for
- Vision-language researchers
- Multimodal model developers
- AI engineers
🔧 Compatible tools
- Python JSON parsers
- PyTorch
- TensorFlow
- VLM frameworks
💡 Tip
Leverage the JSON structure to automate the creation of custom test sets.
Frequently Asked Questions
What are the main categories of tasks covered by this dataset?
Dynamic reasoning, spatial interaction, complex logic, and perspective taking.
Does this dataset include original images?
Yes, each question is associated with an image used for spatial reasoning.
Can this dataset be used for fine-tuning multimodal models?
Yes, it is designed to train and evaluate models combining vision and language on complex tasks.




