MedXpertqa
MedXpertQA is an expert medical benchmark combining question-answer tasks in text alone or with clinical images. It focuses on advanced reasoning, diagnosis, and understanding of body systems.
4,460 questions (text + images), JSONL and JPEG formats, divided into dev/test for textual and multimodal medical QA
MIT
Description
MedXpertqa is a medical dataset designed to assess the ability of artificial intelligence models to understand and reason about complex clinical cases. It consists of two subsets: one based on text (medical QA), and the other multimodal including images (radiologies, skin exams, etc.) and structured clinical pictures. The dataset covers various specialties such as dermatology, oncology, internal medicine, etc.
What is this dataset for?
- Test the ability of models to perform high-level medical reasoning
- Train multimodal models to understand medical cases illustrated by images and tabular data
- Serve as a benchmark to compare different LLMs approaches in medicine
Can it be enriched or improved?
Yes. It is possible to add new clinical cases, in particular from poorly represented specialties or from varied geographical regions. Annotations can also be enriched by adding justifications to answers or detailed clinical explanations. The dataset is structured to facilitate expansion and reuse.
🔎 In summary
🧠 Recommended for
- Medical AI researchers
- LLM developers for healthcare
- Assisted diagnostic projects
🔧 Compatible tools
- Hugging Face Transformers
- MedClip
- OpenBIOLLM
- Medical image analysis tools
💡 Tip
For better generalization, combine this dataset with multilingual sources and human medical explanations.
Frequently Asked Questions
Does this dataset contain medical imaging data?
Yes, the multimodal subset includes clinical images (cutaneous, radiological, etc.) in addition to textual questions.
Can this dataset be used to train medical LLMs?
Absolutely. It is ideal for fine-tuning and evaluating models on complex medical tasks requiring reasoning.
Is it suitable for commercial or clinical use?
It can be used for commercial purposes (MIT license), but remains an experimental benchmark without regulatory clinical validation.



