Reasoning V1 20M
Massive synthetic dataset of questions and answers with textual reasoning traces on various non-mathematical subjects, intended to train models to reason.
Description
Reasoning V1 20M is a gigantic synthetic dataset comprising more than 22 million questions and answers accompanied by traces of reasoning. These examples cover a wide variety of fields such as social sciences, natural sciences, education, creative writing, or general conversations. The format includes a reasoning section surrounded by tags <think>before the final answer.
What is this dataset for?
- Train language models to replicate complex and diverse reasoning skills
- Improving the performance of smaller models via supervised fine-tuning (SFT) with explicit traces
- Evaluate the ability of models to follow step-by-step reasoning on a variety of topics
Can it be enriched or improved?
Yes, although traces and responses are not checked for accuracy, it is possible to add human or automatic validation. You can also annotate the quality of the reasoning or add corrections to refine the quality of the dataset.
🔎 In summary
🧠 Recommended for
- Advanced NLP researchers
- Reasoning model developers
- Explainable AI projects
🔧 Compatible tools
- Hugging Face Datasets
- PyTorch Lightning
- DeepSpeed
- LoRa
💡 Tip
Validate or filter examples to improve the quality of fine-tuning.
Frequently Asked Questions
Is this dataset fully checked for the accuracy of the answers?
No, the answers and lines of reasoning are not verified and may contain errors.
What is the approximate size of the dataset?
Over 22 million examples with around 35.8 billion tokens.
Can this dataset be used to train smaller models?
Yes, it is specially designed for fine-tuning smaller models with reasoning supervision.




