WebInstruct Verified
Dataset of varied and verified questions and answers, designed to strengthen the reasoning skills of large language models in multiple scientific and human fields.
Over 230,000 verified questions and answers, structured text format
Apache 2.0
Description
The dataset WebInstruct Verified contains more than 230,000 questions and answers collected and verified from the web. It covers a broad disciplinary spectrum ranging from mathematics to physics, chemistry, finance, and humanities. Each question was filtered to ensure the verifiability of the answers, including various formats such as floats, tables, matrices, or latex.
What is this dataset for?
- Train language models to develop general and specific reasoning skills
- Test and improve the performance of LLMs on complex and diversified issues
- Building AI systems capable of contextual validation and multi-domain reasoning
Can it be enriched or improved?
This dataset can be enriched by adding new specific fields or by detailed annotation of the types of reasoning required. The integration of additional annotations, such as the complexity of the questions or the resolution strategies, can also improve its use for training.
🔎 In summary
🧠 Recommended for
- NLP researchers
- LLMs developers
- AI R&D teams
🔧 Compatible tools
- Hugging Face Transformers
- LangChain
- Jupyter Notebooks
💡 Tip
Use cross-checking with a Gemini-like model to filter complex questions prior to training.
Frequently Asked Questions
Does this dataset only cover math questions?
No, it covers a wide range of fields including physics, chemistry, finance, humanities, and more.
Have the answers been verified?
Yes, each response was verified using a generative model and rigorous validation methods.
Is this dataset suitable for fine-tuning language models?
Yes, it is especially designed to train and refine models with advanced multi-domain reasoning capabilities.




