Coarse-Discourse - Annotated corpus of speeches on Reddit forums
The <strong>Coarse-Discourse</strong> dataset contains more than 116,000 comments extracted from Reddit discussions, manually annotated with speech acts (e.g. elaboration). The data also contains links between comments, post depth, and subreddit metadata.
Approximately 116,000 annotated comments on Reddit, JSON format, size ~50MB
CC-BY 4.0
Description
Coarse-Discourse is a large body of speech annotations in online discussions, taken from Reddit. Each comment is annotated with speech acts, making it easy to analyze conversational dynamics and understand discursive structures in forums.
What is this dataset for?
- Train models for understanding and classifying speech acts
- Analyze conversational interactions on forums and social networks
- Improve moderation systems and the detection of discursive behaviors
Can it be enriched or improved?
Yes, it is possible to add additional annotations, for example finer levels of speech acts, or to integrate other platforms to enrich the diversity of interactions.
🔎 In summary
🧠 Recommended for
- NLP researchers
- Chatbot developers
- Social media analysts
🔧 Compatible tools
- Python (pandas, json)
- NLP frameworks (SpacY, Transformers)
- Annotation tools
💡 Tip
Use depth and link metadata to reconstruct the thread during analysis.
Frequently Asked Questions
What types of speech acts are annotated in this dataset?
Mostly crude acts such as “elaboration”, “justification”, allowing to characterize the discursive function of comments.
How many comments does this dataset contain?
Approximately 116,000 annotated comments from over 9,000 Reddit threads.
Is this dataset suitable for commercial projects?
Yes, licensed under CC-BY 4.0, it can be used freely including in commercial projects while respecting the attribution.




