Scientists' First Exam: Cognitive Assessment of MLLMs
Benchmark designed to test the perception, understanding and scientific reasoning of MLLMs, through tasks covering 5 disciplines.
66 multimodal tasks in scientific QA (images + questions), JSON format and visual files
MIT
Description
Scientists' First Exam (SFE) is a high quality benchmark for evaluating the scientific cognitive abilities of multimodal language models (MLLMs). It includes 66 tasks covering five disciplines: astronomy, chemistry, earth sciences, life sciences, and materials science. Each task combines scientific visual data (graphs, experimental images, etc.) with pairs of questions and answers formulated to test perception, understanding, and reasoning.
What is this dataset for?
- Evaluate the ability of MLLMs to interpret complex scientific data
- Test different layers of scientific reasoning in real contexts
- Provide a standardized benchmark for artificial cognition research
Can it be enriched or improved?
Yes, SFE can be enriched by adding new disciplines or by translating question pairs into other languages. Researchers can also create task variants with varying levels of complexity or provide human ratings to better calibrate assessment results.
🔎 In summary
🧠 Recommended for
- AI research laboratories
- MLLMS developers
- AI cognition projects
🔧 Compatible tools
- Hugging Face
- LMMS-Eval
- PyTorch
- Gradio
💡 Tip
To refine your models, start with perception tasks before moving on to comparative reasoning.
Frequently Asked Questions
Can this benchmark be used to train AI models?
The dataset is designed for evaluation, but can be adapted to fine-tuning for specific scientific VQA tasks.
In what languages are the tasks available?
All tasks are available in English and Chinese, which encourages multi-lingual evaluation of models.
Is it necessary to have scientific expertise to exploit this dataset?
Yes, some tasks require a good understanding of the scientific disciplines concerned for a relevant assessment.




