pathVQA — Medical questions on images
Medical VQA data set containing over 32,000 question/answer pairs based on images from pathology manuals. It contains open and binary questions, and is a valuable resource for training models that combine computer vision and natural language processing.
Description
PathVQA is a data set designed for the Visual Question Answering (VQA) task in the medical field. It contains 32,632 question/answer pairs annotated manually from pathology images extracted from specialized textbooks and a digital library. The dataset includes both open-ended questions and yes/no questions, covering a variety of topics related to anatomy, cell biology, or clinical pathology.
What is this dataset for?
- Train specialized VQA models in medical imaging
- Evaluate the visual and textual understanding of multimodal models
- Serve as a benchmark for visual QA tasks in a medical environment
Can it be enriched or improved?
Yes. It is possible to annotate the types of questions (anatomy, diagnosis, etc.) or to categorize the images by organ or pathology. The database can also be used to generate images via AI from QA or integrate it into automated medical assistance pipelines.
🔎 In summary
🧠 Recommended for
- Medical vision researchers
- VQA developers
- AI health projects
🔧 Compatible tools
- PyTorch
- OpenVQA
- BLIP
- Lava
- Hugging Face Transformers
💡 Tip
Filter questions by type (binary vs. open) to create specific subtasks during training.
Frequently Asked Questions
Can this dataset be used for clinical purposes?
No, it is designed for AI research and should not be used directly for diagnostic purposes without medical validation.
Are all images used in QA pairs?
No, only 4,289 images out of the 5,004 are actually associated with questions.
Is this data set multilingual?
No, all questions and answers are written in English. A multilingual version with translation could be considered.




