By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Coarse-Discourse - Annotated corpus of speeches on Reddit forums
Text

Coarse-Discourse - Annotated corpus of speeches on Reddit forums

The <strong>Coarse-Discourse</strong> dataset contains more than 116,000 comments extracted from Reddit discussions, manually annotated with speech acts (e.g. elaboration). The data also contains links between comments, post depth, and subreddit metadata.

Download dataset
Size

Approximately 116,000 annotated comments on Reddit, JSON format, size ~50MB

Licence

CC-BY 4.0

Description

‍

Coarse-Discourse is a large body of speech annotations in online discussions, taken from Reddit. Each comment is annotated with speech acts, making it easy to analyze conversational dynamics and understand discursive structures in forums.

‍

‍

What is this dataset for?

‍

  • Train models for understanding and classifying speech acts
  • Analyze conversational interactions on forums and social networks
  • Improve moderation systems and the detection of discursive behaviors

‍

‍

Can it be enriched or improved?

‍

Yes, it is possible to add additional annotations, for example finer levels of speech acts, or to integrate other platforms to enrich the diversity of interactions.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (JSON data, requires understanding of annotations)
🧼 Need for cleaning⭐⭐⭐✩✩ (Moderate, link and metadata handling)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good, speech acts manually validated)
📜 Commercial license✅ Yes (CC-BY 4.0)
👨‍💻 Beginner friendly⚠️ Moderate – useful for advanced conversational NLP projects
🔁 Fine-tuning ready🎯 Perfect for speech act classification and modeling
🌍 Cultural diversity⚠️ English, based on diverse Reddit community

‍

‍

🧠 Recommended for

  • NLP researchers
  • Chatbot developers
  • Social media analysts

‍

‍

🔧 Compatible tools

  • Python (pandas, json)
  • NLP frameworks (SpacY, Transformers)
  • Annotation tools

‍

‍

💡 Tip

Use depth and link metadata to reconstruct the thread during analysis.

Frequently Asked Questions

What types of speech acts are annotated in this dataset?

Mostly crude acts such as “elaboration”, “justification”, allowing to characterize the discursive function of comments.

How many comments does this dataset contain?

Approximately 116,000 annotated comments from over 9,000 Reddit threads.

Is this dataset suitable for commercial projects?

Yes, licensed under CC-BY 4.0, it can be used freely including in commercial projects while respecting the attribution.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.