By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
OmniSpatial — Multimodal Spatial Reasoning Benchmark
Multimodal

OmniSpatial — Multimodal Spatial Reasoning Benchmark

OmniSpatial is a benchmark designed to test the spatial reasoning abilities of vision-language models. The dataset includes multiple choice questions that cover four types of tasks: dynamic reasoning, spatial interaction, complex logic, and perspective taking. Each question is linked to an image and includes options with the correct answer shown.

Download dataset
Size

Questions annotated in JSON format with structured schema, several thousand presumed items

Licence

Apache 2.0

Description

OmniSpatial proposes a set of questions structured in JSON, intended to assess the spatial reasoning of multimodal models combining vision and language. The tasks cover motion analysis, spatial interactions, complex logic, and perspective taking.

What is this dataset for?

  • Test the spatial and dynamic understanding of VLM models
  • Evaluate multi-stage reasoning ability in complex visual contexts
  • Improve the robustness of models on advanced vision-language tasks

Can it be enriched or improved?

The dataset can be enriched by adding new images, scenarios, and question types, in particular by refining task subcategories. Community contribution can broaden the diversity of spatial and dynamic contexts.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Well structured, requires JSON and vision-language understanding)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low, well-annotated data)
🏷️ Annotation richness⭐⭐⭐⭐⭐ (Complete, with detailed categories and validated answers)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Moderate, useful for users with multimodal foundation
🔁 Fine-tuning ready🎯 Suitable for fine-tuning on spatial reasoning and VLM
🌍 Cultural diversity⚠️ Technical focus, visual diversity to develop

🧠 Recommended for

  • Vision-language researchers
  • Multimodal model developers
  • AI engineers

🔧 Compatible tools

  • Python JSON parsers
  • PyTorch
  • TensorFlow
  • VLM frameworks

💡 Tip

Leverage the JSON structure to automate the creation of custom test sets.

Frequently Asked Questions

What are the main categories of tasks covered by this dataset?

Dynamic reasoning, spatial interaction, complex logic, and perspective taking.

Does this dataset include original images?

Yes, each question is associated with an image used for spatial reasoning.

Can this dataset be used for fine-tuning multimodal models?

Yes, it is designed to train and evaluate models combining vision and language on complex tasks.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.