DuoRC — Questions/Answers about movie synopsis
DuoRC is a QA dataset in English based on movie summaries taken from Wikipedia and IMDb. It contains two versions of questions: one aligned with the source passage (SelFrc), the other asked on a parallel passage (paraphraserC), promoting a deep understanding of the text and paraphrasing.
187,213 examples, JSON text format, two subsets (selFrc, paraphraserc)
MIT
Description
DuoRC is a QA corpus made up of more than 187,000 question/answer pairs built from movie scenarios. The data is divided into two parts: SelFrc (questions and answers from the same Wikipedia text) and ParaphraserC (questions from Wikipedia and answers from IMDb). This voluntary misalignment makes it possible to test the ability of a model to reason, paraphrase, or synthesize a response.
What is this dataset for?
- Form models by Extractive Question Answering or abstractive
- Evaluate text comprehension through parallel passages (Wikipedia vs IMDb)
- Strengthen the paraphrase and generative response capacities of a model
Can it be enriched or improved?
Yes, duORC can be extended to other narrative sources (series, books) or to other languages. It is also possible to classify questions by types (causal, factual, inferential) or to add complexity or cinematographic genre annotations for specialized fine-tuning.
🔎 In summary
🧠 Recommended for
- NLP researchers
- Fine-tuning QA models
- Text summary projects
🔧 Compatible tools
- Hugging Face Transformers
- BERT QA
- T5
- BART
- Haystack
💡 Tip
To assess generalization, train on SelFrc and test on Paraphraserc: this measures robustness to semantic misalignment.
Frequently Asked Questions
What is the difference between SelFrc and Paraphraserc?
SelFrc uses questions and answers drawn from the same passage, while Paraphraserc offers answers from different but related passages.
Is the dataset suitable for generating long responses?
Yes, especially with the paraphrasERC part, which favors generative and paraphrased responses rather than extractive ones.
Does duORC contain movie or genre metadata?
No, the dataset is based on the synopses only and does not contain additional movie metadata.




