Ambig QA
Ambig QA dataset including 14,042 questions taken from NQ-Open, annotated to reflect various ambiguities detected in the open questions. This dataset helps train models that can handle questions with multiple or ambiguous answers in broad contexts.
Approximately 14,000 questions in JSON with various annotations, structured format for open QA
CC BY-SA 3.0
Description
Ambig QA is a dataset of approximately 14,000 questions from NQ-Open, annotated specifically for the ambiguities that appear in open-ended questions in natural language. These annotations make it possible to better understand the different types of ambiguity and to prepare the models for responses adapted to the multiple possible interpretations.
What is this dataset for?
- Train question-answering models to manage ambiguity in open questions
- Evaluate the robustness and the ability to disambiguate QA systems
- Improve chatbots to better interpret inaccurate queries
Can it be enriched or improved?
Yes, the dataset can be enriched with additional annotations concerning the nature of ambiguities, or extended to other languages and domains to improve the diversity of use cases. Remixing with other QA datasets is also possible.
🔎 In summary
🧠 Recommended for
- QA researchers
- Chatbot developers
- NLP engineers
🔧 Compatible tools
- Hugging Face Transformers
- SpacY
- AllennLP
- JSON annotation tools
💡 Tip
Use ambiguity annotations to create models that can generate multiple relevant responses.
Frequently Asked Questions
What makes this dataset unique for open-ended questions?
It is specifically annotated to capture and represent different forms of ambiguity in open-ended questions, which is rare in other QA datasets.
Can this dataset be used for multilingual tasks?
No, it is only in English, but the format can be adapted for other languages via translation and annotation.
What license applies and what are the constraints?
CC BY-SA 3.0 license, which means that any derivative use must be shared under the same license, promoting open sharing.




