By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Reasoning V1 20M
Text

Reasoning V1 20M

Massive synthetic dataset of questions and answers with textual reasoning traces on various non-mathematical subjects, intended to train models to reason.

Download dataset
Size

Over 22 million rows, 35.8 billion tokens, JSON/parquet format

Licence

Apache 2.0

Description

Reasoning V1 20M is a gigantic synthetic dataset comprising more than 22 million questions and answers accompanied by traces of reasoning. These examples cover a wide variety of fields such as social sciences, natural sciences, education, creative writing, or general conversations. The format includes a reasoning section surrounded by tags <think>before the final answer.

What is this dataset for?

  • Train language models to replicate complex and diverse reasoning skills
  • Improving the performance of smaller models via supervised fine-tuning (SFT) with explicit traces
  • Evaluate the ability of models to follow step-by-step reasoning on a variety of topics

Can it be enriched or improved?

Yes, although traces and responses are not checked for accuracy, it is possible to add human or automatic validation. You can also annotate the quality of the reasoning or add corrections to refine the quality of the dataset.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Large but standard format, requires resources)
🧼 Need for cleaning⭐⭐⭐⭐✩ (Moderate: no quality validation of answers)
🏷️ Annotation richness⭐⭐⭐✩✩ (Reasoning traces present but not validated)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly⚠️ No, requires good resources and expertise
🔁 Fine-tuning ready🎯 Excellent for fine-tuning on supervised reasoning
🌍 Cultural diversity⚠️ Mostly English, high thematic diversity

🧠 Recommended for

  • Advanced NLP researchers
  • Reasoning model developers
  • Explainable AI projects

🔧 Compatible tools

  • Hugging Face Datasets
  • PyTorch Lightning
  • DeepSpeed
  • LoRa

💡 Tip

Validate or filter examples to improve the quality of fine-tuning.

Frequently Asked Questions

Is this dataset fully checked for the accuracy of the answers?

No, the answers and lines of reasoning are not verified and may contain errors.

What is the approximate size of the dataset?

Over 22 million examples with around 35.8 billion tokens.

Can this dataset be used to train smaller models?

Yes, it is specially designed for fine-tuning smaller models with reasoning supervision.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.