PhyX Multimodal Physics Reasoning
A unique dataset to test physical reasoning at university level through complex questions illustrated by accurate images.
Description
PhyX is a next-generation benchmark designed to assess the skills of AI models in multimodal physical reasoning. It includes 3,000 questions covering six major areas of physics (mechanics, optics, thermodynamics, etc.), each question being accompanied by a realistic image and a simplified textual context. The whole forms a rigorous corpus to test the capacity for scientific inference beyond the simple memorization of facts.
What is this dataset for?
- Evaluate or train multimodal models based on advanced physical reasoning
- Measuring the ability of LLMs to integrate visual information and implicit laws
- Building demanding scientific AI benchmarks (university level)
Can it be enriched or improved?
Yes. You can add other modalities such as videos, enrich answer options, or create multilingual translated versions. Additional annotations on the types of reasoning or cognitive complexity would also be beneficial for fine-tuning tasks.
🔎 In summary
🧠 Recommended for
- Scientific AI researchers
- Benchmark designers
- Specialized LLM developers
🔧 Compatible tools
- Hugging Face Transformers
- OpenFlamingo
- VLMEvalkit
- PyTorch
💡 Tip
Use the “mini” subset for rapid prototyping before using the full set for benchmarks or fine-tuning.
Frequently Asked Questions
Does this dataset include videos or only static images?
It is based on static realistic images, not videos. Each question is linked to an illustrative image in PNG format.
Is this data set suitable for training or only for evaluation?
It is primarily designed for evaluation, but can also be used for fine-tuning with adapted models.
What types of reasoning are covered in the questions?
The questions cover 6 types of physical reasoning (e.g. causality, dynamics, space-time relationships, etc.) on 25 sub-domains.




