By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
pathVQA — Medical questions on images
Multimodal

pathVQA — Medical questions on images

Medical VQA data set containing over 32,000 question/answer pairs based on images from pathology manuals. It contains open and binary questions, and is a valuable resource for training models that combine computer vision and natural language processing.

Download dataset
Size

32,632 QA pairs on 4,289 images, image + text formats

Licence

MIT

Description

‍

PathVQA is a data set designed for the Visual Question Answering (VQA) task in the medical field. It contains 32,632 question/answer pairs annotated manually from pathology images extracted from specialized textbooks and a digital library. The dataset includes both open-ended questions and yes/no questions, covering a variety of topics related to anatomy, cell biology, or clinical pathology.

‍

‍

What is this dataset for?

‍

  • Train specialized VQA models in medical imaging
  • Evaluate the visual and textual understanding of multimodal models
  • Serve as a benchmark for visual QA tasks in a medical environment

‍

‍

Can it be enriched or improved?

‍

Yes. It is possible to annotate the types of questions (anatomy, diagnosis, etc.) or to categorize the images by organ or pathology. The database can also be used to generate images via AI from QA or integrate it into automated medical assistance pipelines.

‍

‍

🔎 In summary

Criterion Evaluation
🧩Ease of Use ⭐⭐⭐☆☆ (average – requires multimodal processing of image + text)
🧼Cleaning Required ⭐⭐⭐⭐⭐ (low – data ready to use)
🏷️Annotation Richness ⭐⭐⭐⭐☆ (good – open-ended and yes/no questions covering many cases)
📜Commercial License ✅ Yes (MIT)
👨‍💻Ideal for Beginners 🧪 Moderate – requires image+text handling
🔁Reusable for Fine-Tuning 🔥 Yes – perfect for medical VQA or CLIP models
🌍Cultural Diversity 🌍 Limited – content based on classic English-language manuals

‍

‍

🧠 Recommended for

  • Medical vision researchers
  • VQA developers
  • AI health projects

‍

‍

🔧 Compatible tools

  • PyTorch
  • OpenVQA
  • BLIP
  • Lava
  • Hugging Face Transformers

‍

‍

💡 Tip

Filter questions by type (binary vs. open) to create specific subtasks during training.

Frequently Asked Questions

Can this dataset be used for clinical purposes?

No, it is designed for AI research and should not be used directly for diagnostic purposes without medical validation.

Are all images used in QA pairs?

No, only 4,289 images out of the 5,004 are actually associated with questions.

Is this data set multilingual?

No, all questions and answers are written in English. A multilingual version with translation could be considered.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.