By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Scientists' First Exam: Cognitive Assessment of MLLMs
Multimodal

Scientists' First Exam: Cognitive Assessment of MLLMs

Benchmark designed to test the perception, understanding and scientific reasoning of MLLMs, through tasks covering 5 disciplines.

Download dataset
Size

66 multimodal tasks in scientific QA (images + questions), JSON format and visual files

Licence

MIT

Description

Scientists' First Exam (SFE) is a high quality benchmark for evaluating the scientific cognitive abilities of multimodal language models (MLLMs). It includes 66 tasks covering five disciplines: astronomy, chemistry, earth sciences, life sciences, and materials science. Each task combines scientific visual data (graphs, experimental images, etc.) with pairs of questions and answers formulated to test perception, understanding, and reasoning.

What is this dataset for?

  • Evaluate the ability of MLLMs to interpret complex scientific data
  • Test different layers of scientific reasoning in real contexts
  • Provide a standardized benchmark for artificial cognition research

Can it be enriched or improved?

Yes, SFE can be enriched by adding new disciplines or by translating question pairs into other languages. Researchers can also create task variants with varying levels of complexity or provide human ratings to better calibrate assessment results.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Provides loading and evaluation scripts)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (None – data ready for evaluation)
🏷️ Annotation richness⭐⭐⭐⭐⭐ (Very rich – expert and structured tasks)
📜 Commercial license✅ Yes (MIT)
👨‍💻 Beginner friendly⚠️ No – requires advanced scientific understanding
🔁 Fine-tuning ready🎯 Possible for targeted scientific VQA
🌍 Cultural diversity🌐 Bilingual English/Chinese – good global accessibility

🧠 Recommended for

  • AI research laboratories
  • MLLMS developers
  • AI cognition projects

🔧 Compatible tools

  • Hugging Face
  • LMMS-Eval
  • PyTorch
  • Gradio

💡 Tip

To refine your models, start with perception tasks before moving on to comparative reasoning.

Frequently Asked Questions

Can this benchmark be used to train AI models?

The dataset is designed for evaluation, but can be adapted to fine-tuning for specific scientific VQA tasks.

In what languages are the tasks available?

All tasks are available in English and Chinese, which encourages multi-lingual evaluation of models.

Is it necessary to have scientific expertise to exploit this dataset?

Yes, some tasks require a good understanding of the scientific disciplines concerned for a relevant assessment.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.