By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Cotton67k
Text

Cotton67k

Corpus of assistant/user dialogues structured according to the ShareGPT format, highlighting detailed step-by-step reasoning (Chain-of-Thought).

Download dataset
Size

67,844 examples, 813 MB, JSON/shareGPT format

Licence

Apache 2.0

Description

Cotton-67k is a conversational dataset specialized in natural language reasoning. It contains 67,844 examples of user-model exchanges in ShareGPT format. Each example involves explicit reasoning in several steps (Chain-of-Thought), with a great diversity of tasks (general, STEM, conditional). The data is derived from open-source models (Qwen, DeepSeek, QwQ, etc.) and has been cleaned to ensure a clear and consistent structure.

What is this dataset for?

  • Train LLMs to explain their reasoning step by step
  • Improve the quality of answers on complex tasks (math, code, science)
  • Test or evaluate model inference capabilities in open contexts

Can it be enriched or improved?

Yes, this dataset can be enriched by adding variants of answers, by translating the exchanges into other languages, or by adding an annotation of the types of reasoning. It can also be combined with specialized datasets such as OpenCodeReasoning for multimodal or programmation-oriented training.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Very simple to exploit - ShareGPT format)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Already cleaned and structured)
🏷️ Annotation richness⭐⭐⭐⭐⭐ (Complete reasoning structured by turn)
📜 Commercial license✅ Yes (Apache 2.0)
👨‍💻 Beginner friendly🌟 Yes, easily usable for fine-tuning
🔁 Fine-tuning ready🎯 Ideal for strengthening CoT reasoning
🌍 Cultural diversity⚠️ Moderate – generic STEM-oriented content

🧠 Recommended for

  • LLM trainers
  • NLP researchers
  • AI educational projects

🔧 Compatible tools

  • LoRa
  • OpenChat
  • Hugging Face Transformers
  • VLLM

💡 Tip

Combine this dataset with your own real cases to refine the reasoning according to a target domain.

Frequently Asked Questions

Is the dataset multilingual?

No, it's only in English at the moment. However, it is possible to translate it for multilingual use.

Do the dialogues all follow the same pattern?

Yes, each example follows a user/assistant format, centered on a question and comprehensive step-by-step reasoning.

Is it suitable for a medium size model (7B or less)?

Yes, the dataset is used in particular to train or refine 7B models and works very well for this scale.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.