By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Ambig QA
Text

Ambig QA

Ambig QA dataset including 14,042 questions taken from NQ-Open, annotated to reflect various ambiguities detected in the open questions. This dataset helps train models that can handle questions with multiple or ambiguous answers in broad contexts.

Download dataset
Size

Approximately 14,000 questions in JSON with various annotations, structured format for open QA

Licence

CC BY-SA 3.0

Description

‍

Ambig QA is a dataset of approximately 14,000 questions from NQ-Open, annotated specifically for the ambiguities that appear in open-ended questions in natural language. These annotations make it possible to better understand the different types of ambiguity and to prepare the models for responses adapted to the multiple possible interpretations.

‍

‍

What is this dataset for?

‍

  • Train question-answering models to manage ambiguity in open questions
  • Evaluate the robustness and the ability to disambiguate QA systems
  • Improve chatbots to better interpret inaccurate queries

‍

‍

Can it be enriched or improved?

‍

Yes, the dataset can be enriched with additional annotations concerning the nature of ambiguities, or extended to other languages and domains to improve the diversity of use cases. Remixing with other QA datasets is also possible.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Standard JSON format, requires understanding of annotations)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – well-cleaned and structured dataset)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Rich – detailed annotations on ambiguity)
📜 Commercial license⚡ Yes, but under Share-Alike (CC BY-SA 3.0)
👨‍💻 Beginner friendly⚠️ Moderate – useful for those with basic NLP QA knowledge
🔁 Fine-tuning ready🎯 Yes, perfect for specialized QA training
🌍 Cultural diversity⚠️ English, varied open-ended questions

‍

‍

🧠 Recommended for

  • QA researchers
  • Chatbot developers
  • NLP engineers

‍

‍

🔧 Compatible tools

  • Hugging Face Transformers
  • SpacY
  • AllennLP
  • JSON annotation tools

‍

‍

💡 Tip

Use ambiguity annotations to create models that can generate multiple relevant responses.

Frequently Asked Questions

What makes this dataset unique for open-ended questions?

It is specifically annotated to capture and represent different forms of ambiguity in open-ended questions, which is rare in other QA datasets.

Can this dataset be used for multilingual tasks?

No, it is only in English, but the format can be adapted for other languages via translation and annotation.

What license applies and what are the constraints?

CC BY-SA 3.0 license, which means that any derivative use must be shared under the same license, promoting open sharing.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.