Robo2vlm Reasoning
Robo2VLM Reasoning is a dataset designed to test the ability of vision-language models to perform reasoning based on natural robotic scenes. It offers reasoning traces generated automatically from visual contexts, to enrich the analysis and understanding of object manipulation scenes.
Several thousand multimodal examples with images, questions and textual reasoning generated
Apache 2.0
Description
Robo2vlm Reasoning is an extension of the Robo2VLM dataset, focused on visual reasoning skills. It uses robotic in-the-wild scenes combined with visual questions and answers justified by generated reasoning. This benchmark makes it possible to assess the coherence and logic of the answers provided by multimodal models.
What is this dataset for?
- Test the visual reasoning skills of VLMs (Vision-Language Models)
- Improving the contextual understanding of robotic scenes by AIs
- Evaluate the ability to justify answers through coherent lines of reasoning
Can it be enriched or improved?
Yes, this dataset can be enriched by additional human annotations, linguistic variants, or by adding error cases to train robustness. It can also be extended to other types of tasks such as planning or anticipating robotic actions.
🔎 In summary
🧠 Recommended for
- Cognitive robotics researchers
- VLM developers
- Advanced VQA projects
🔧 Compatible tools
- Hugging Face Transformers
- OpenVQA
- Lava
- Gemini
- PyTorch
💡 Tip
Use this dataset with particular attention to the coherence of the reasoning chains, especially during the evaluation phase.
Frequently Asked Questions
What is the difference between Robo2VLM and Robo2VLM Reasoning?
Robo2VLM Reasoning adds automatically generated reasoning traces for each visual question, allowing for a finer evaluation.
What type of scenes are represented in this dataset?
These are scenes from real robotic manipulations, captured in-the-wild, with everyday objects in various contexts.
Is this dataset suitable for training?
Yes, although it is designed for evaluation, it can also be used to fine-tune multimodal VQA or reasoning models.




