By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
DuoRC — Questions/Answers about movie synopsis
Text

DuoRC — Questions/Answers about movie synopsis

DuoRC is a QA dataset in English based on movie summaries taken from Wikipedia and IMDb. It contains two versions of questions: one aligned with the source passage (SelFrc), the other asked on a parallel passage (paraphraserC), promoting a deep understanding of the text and paraphrasing.

Download dataset
Size

187,213 examples, JSON text format, two subsets (selFrc, paraphraserc)

Licence

MIT

Description

‍

DuoRC is a QA corpus made up of more than 187,000 question/answer pairs built from movie scenarios. The data is divided into two parts: SelFrc (questions and answers from the same Wikipedia text) and ParaphraserC (questions from Wikipedia and answers from IMDb). This voluntary misalignment makes it possible to test the ability of a model to reason, paraphrase, or synthesize a response.

‍

‍

What is this dataset for?

‍

  • Form models by Extractive Question Answering or abstractive
  • Evaluate text comprehension through parallel passages (Wikipedia vs IMDb)
  • Strengthen the paraphrase and generative response capacities of a model

‍

‍

Can it be enriched or improved?

‍

Yes, duORC can be extended to other narrative sources (series, books) or to other languages. It is also possible to classify questions by types (causal, factual, inferential) or to add complexity or cinematographic genre annotations for specialized fine-tuning.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Directly usable with common QA models)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – data well-structured in JSON)
🏷️ Annotation richness⭐⭐⭐✩✩ (Two QA variants - aligned / misaligned)
📜 Commercial license✅ Yes (MIT)
👨‍💻 Beginner friendly⚠️ Accessible, but requires good mastery of QA tasks
🔁 Fine-tuning ready🎯 Excellent base for BERT, T5, BART-type models
🌍 Cultural diversity⚠️ Centered on US movies, but extrapolatable

‍

‍

🧠 Recommended for

  • NLP researchers
  • Fine-tuning QA models
  • Text summary projects

‍

‍

🔧 Compatible tools

  • Hugging Face Transformers
  • BERT QA
  • T5
  • BART
  • Haystack

‍

‍

💡 Tip

To assess generalization, train on SelFrc and test on Paraphraserc: this measures robustness to semantic misalignment.

Frequently Asked Questions

What is the difference between SelFrc and Paraphraserc?

SelFrc uses questions and answers drawn from the same passage, while Paraphraserc offers answers from different but related passages.

Is the dataset suitable for generating long responses?

Yes, especially with the paraphrasERC part, which favors generative and paraphrased responses rather than extractive ones.

Does duORC contain movie or genre metadata?

No, the dataset is based on the synopses only and does not contain additional movie metadata.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.