By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
WebInstruct Verified
Text

WebInstruct Verified

Dataset of varied and verified questions and answers, designed to strengthen the reasoning skills of large language models in multiple scientific and human fields.

Download dataset
Size

Over 230,000 verified questions and answers, structured text format

Licence

Apache 2.0

Description

The dataset WebInstruct Verified contains more than 230,000 questions and answers collected and verified from the web. It covers a broad disciplinary spectrum ranging from mathematics to physics, chemistry, finance, and humanities. Each question was filtered to ensure the verifiability of the answers, including various formats such as floats, tables, matrices, or latex.

What is this dataset for?

  • Train language models to develop general and specific reasoning skills
  • Test and improve the performance of LLMs on complex and diversified issues
  • Building AI systems capable of contextual validation and multi-domain reasoning

Can it be enriched or improved?

This dataset can be enriched by adding new specific fields or by detailed annotation of the types of reasoning required. The integration of additional annotations, such as the complexity of the questions or the resolution strategies, can also improve its use for training.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Requires some preprocessing for integration)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low, data has been rigorously verified)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good level, with verified answers and varied format)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ Moderate, recommended for users with intermediate experience
🔁 Fine-tuning ready🎯 Excellent for fine-tuning on reasoning tasks
🌍 Cultural diversity⚠️ Wide disciplinary spectrum, but few indications on linguistic diversity

🧠 Recommended for

  • NLP researchers
  • LLMs developers
  • AI R&D teams

🔧 Compatible tools

  • Hugging Face Transformers
  • LangChain
  • Jupyter Notebooks

💡 Tip

Use cross-checking with a Gemini-like model to filter complex questions prior to training.

Frequently Asked Questions

Does this dataset only cover math questions?

No, it covers a wide range of fields including physics, chemistry, finance, humanities, and more.

Have the answers been verified?

Yes, each response was verified using a generative model and rigorous validation methods.

Is this dataset suitable for fine-tuning language models?

Yes, it is especially designed to train and refine models with advanced multi-domain reasoning capabilities.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.