Cotton67k
Corpus of assistant/user dialogues structured according to the ShareGPT format, highlighting detailed step-by-step reasoning (Chain-of-Thought).
Description
Cotton-67k is a conversational dataset specialized in natural language reasoning. It contains 67,844 examples of user-model exchanges in ShareGPT format. Each example involves explicit reasoning in several steps (Chain-of-Thought), with a great diversity of tasks (general, STEM, conditional). The data is derived from open-source models (Qwen, DeepSeek, QwQ, etc.) and has been cleaned to ensure a clear and consistent structure.
What is this dataset for?
- Train LLMs to explain their reasoning step by step
- Improve the quality of answers on complex tasks (math, code, science)
- Test or evaluate model inference capabilities in open contexts
Can it be enriched or improved?
Yes, this dataset can be enriched by adding variants of answers, by translating the exchanges into other languages, or by adding an annotation of the types of reasoning. It can also be combined with specialized datasets such as OpenCodeReasoning for multimodal or programmation-oriented training.
🔎 In summary
🧠 Recommended for
- LLM trainers
- NLP researchers
- AI educational projects
🔧 Compatible tools
- LoRa
- OpenChat
- Hugging Face Transformers
- VLLM
💡 Tip
Combine this dataset with your own real cases to refine the reasoning according to a target domain.
Frequently Asked Questions
Is the dataset multilingual?
No, it's only in English at the moment. However, it is possible to translate it for multilingual use.
Do the dialogues all follow the same pattern?
Yes, each example follows a user/assistant format, centered on a question and comprehensive step-by-step reasoning.
Is it suitable for a medium size model (7B or less)?
Yes, the dataset is used in particular to train or refine 7B models and works very well for this scale.




