SuperGPQA — LLM assessment on 285 disciplines
SuperGPQA is a massive QA benchmark covering 285 academic disciplines, designed to test the comprehension and reasoning skills of LLMs.
Description
SuperGPQA is a question-and-answer dataset covering a wide range of academic disciplines, ranging from law to medicine, mathematics, music and computer science. This benchmark aims to test the abilities of large language models (LLMs) to understand, reason and respond in a relevant manner in complex and specialized fields.
What is this dataset for?
- Evaluate the performance of LLMs on multi-domain QA tasks
- Serve as a basis for training or testing specialized models in general or academic knowledge
- Analyzing the robustness of models in the face of complex questions from 285 disciplines
Can it be enriched or improved?
Yes. It is possible to add other disciplines, to translate the content for multilingual use or to adapt it to different educational levels (bachelor's degree, high school, etc.). Additional annotations can be added (level of difficulty, justification of answers...).
🔎 In summary
🧠 Recommended for
- NLP researchers
- Benchmark designers
- LLM QA Engineers
🔧 Compatible tools
- LangChain
- Transformers
- Eval Gauntlet
- Weights & Biases
💡 Tip
Filter specific areas (e.g. law, medicine) to make targeted micro-evaluations of your model.
Frequently Asked Questions
Can SuperGPQA be used for fine tuning?
Yes, although designed for evaluation, its diversity makes it useful for adjusting a model for specialized QA tasks.
Is it necessary to respect other licenses in addition to the ODC-BY license?
Yes, some snippets come from other open source datasets. The associated conditions must be respected if you reuse them in isolation.
Is it suitable for use in French?
The dataset is in English. For use in French, automatic translation followed by manual validation is recommended.




