By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Text

SelQa

SelQA is a benchmark for the question answering task based on the selection of the correct answer among several candidates. The dataset contains questions with context and candidate answers, often paraphrased, intended to assess the accuracy of QA models.

Download dataset
Size

Approximately 404,776 examples, 68 MB, JSON/Parquet format

Licence

Apache 2.0

Description

The dataset SelQa includes over 400,000 examples in English, each example containing a question, a context (section or article), and a set of candidate answers from which the correct one must be selected. Data is provided in JSON and Parquet format.

What is this dataset for?

  • Train question answering models based on the selection of answers
  • Evaluate the ability of models to identify the correct response in a given context
  • Improving QA systems for practical applications like virtual assistants

Can it be enriched or improved?

Yes, it is possible to add additional annotations on the types of questions, to integrate more varied paraphrases, or to extend the corpus to other languages to reinforce the robustness of the models.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Well-structured format, easy to integrate)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – clean and well-organized data)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good, includes paraphrases and varied contexts)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly🌟 Yes, accessible even to QA novices
🔁 Fine-tuning ready🎯 Perfect for fine-tuning QA models
🌍 Cultural diversity⚡ English, with thematic diversity

🧠 Recommended for

  • NLP researchers
  • Virtual assistant developers
  • AI students

🔧 Compatible tools

  • Hugging Face Transformers
  • PyTorch
  • TensorFlow
  • Jupyter notebooks

💡 Tip

Use the available paraphrases to train the robustness of the models in the face of linguistic variation.

Frequently Asked Questions

What is the size of the SelQA dataset?

The dataset includes approximately 404,776 examples, with a file size of approximately 68 MB.

Can SelQa be used for open questions or multiple choice questions only?

It is primarily designed for selecting the correct answer from multiple candidates, not for open-ended questions.

Is this dataset suitable for commercial use?

Yes, the Apache 2.0 license allows unrestricted commercial use.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.