Hate Speech Detection Dataset for Social Media
Synthetic dataset designed to train and evaluate NLP models for detecting hateful and offensive speech on social networks. It covers neutral, offensive, and hateful categories, with privacy-respecting content and no sensitive real data.
1,829 posts in CSV, with text fields, label, platform, date, and user ID
CC0: Public Domain
Description
The dataset Hate Speech Detection Dataset for Social Media includes 1,829 synthetic posts from platforms like Twitter and Reddit, annotated to detect hate, offensive, or neutral speech. The data also includes metadata such as time, platform, and user ID.
What is this dataset for?
- Train classification models for automatic moderation in real time
- Evaluating NLP pipelines for the detection of hate speech on social networks
- Contribute to research on the security and surveillance of online content
Can it be enriched or improved?
The dataset can be extended with real or larger examples, and improved with finer annotations (e.g. degrees of intensity, specific targets). The integration of multi-language contexts would also be beneficial.
🔎 In summary
🧠 Recommended for
- NLP researchers
- Moderation system developers
- Machine learning students
🔧 Compatible tools
- Scikit-learn
- Hugging Face Transformers
- SpacY
- TensorFlow
- PyTorch
💡 Tip
Train with complementary data sets to improve robustness in real conditions.
Frequently Asked Questions
Can this dataset be used to automatically detect hate speech in real time?
Yes, it is specifically designed to train and evaluate classification models in real time on social networks.
Does this dataset contain real or sensitive data?
No, it is a synthetic dataset guaranteeing confidentiality and the absence of truly sensitive content.
Does the dataset cover multiple languages?
No, the dataset is mainly in English and does not offer multilingual annotations.




