DROP — Benchmark of understanding with reasoning
DROP is an advanced text comprehension dataset, designed to assess the ability of a system to answer complex questions using discrete operations (counting, sorting, calculating) based on long paragraphs.
Approximately 86,000 examples, JSON format with paragraphs, questions, answers
CC-BY-SA 4.0
Description
DROP (Discrete Reasoning Over Paragraphs) is a reading benchmark composed of more than 86,000 questions in English. It challenges language models by asking them to read paragraphs and answer questions that require symbolic reasoning: counting, comparing, extracting entities, or adding. The data is the result of a collaborative and adversarial process in order to increase the difficulty of the dataset.
What is this dataset for?
- Evaluate or train complex, paragraph based QA models
- Test the ability of an LLM to perform discrete reasoning (number, dates, quantities, etc.)
- Strengthening contextual understanding in AI systems
Can it be enriched or improved?
Yes. The dataset can be translated or adapted to other languages. New types of questions (sorting, time calculations, reformulations) can be added to enrich the variety. Metadata (e.g. level of difficulty, type of question) can also refine its use in supervised or self-supervised learning.
🔎 In summary
🧠 Recommended for
- QA researchers
- LLMs fine-tuning
- Symbolic logic evaluation
🔧 Compatible tools
- Hugging Face Transformers
- T5
- DeBerta
- OpenAI GPT
💡 Tip
Use models capable of chaining several reasoning steps for better results (e.g. T5 with prompting in several rounds).
Frequently Asked Questions
How does DROP differ from other QA datasets?
Unlike simple QAs, DROP requires calculations, counting, or sorting from text, making it a more demanding benchmark.
Can we use this dataset for few-shot learning?
Yes, by selecting examples of various question types, DROP can be used in few-shot or chain-of-thought configurations.
Is it suitable for multilingual models?
The dataset is in English only, but can be translated or adapted to other languages for multi-lingual experiments.



