By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Hate Speech Detection Dataset for Social Media
Text

Hate Speech Detection Dataset for Social Media

Synthetic dataset designed to train and evaluate NLP models for detecting hateful and offensive speech on social networks. It covers neutral, offensive, and hateful categories, with privacy-respecting content and no sensitive real data.

Download dataset
Size

1,829 posts in CSV, with text fields, label, platform, date, and user ID

Licence

CC0: Public Domain

Description

The dataset Hate Speech Detection Dataset for Social Media includes 1,829 synthetic posts from platforms like Twitter and Reddit, annotated to detect hate, offensive, or neutral speech. The data also includes metadata such as time, platform, and user ID.

What is this dataset for?

  • Train classification models for automatic moderation in real time
  • Evaluating NLP pipelines for the detection of hate speech on social networks
  • Contribute to research on the security and surveillance of online content

Can it be enriched or improved?

The dataset can be extended with real or larger examples, and improved with finer annotations (e.g. degrees of intensity, specific targets). The integration of multi-language contexts would also be beneficial.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Simple CSV format, ready to use)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – well-structured synthetic data)
🏷️ Annotation richness⭐⭐⭐✩✩ (Three clear classes, no complex annotations)
📜 Commercial license✅ Yes (CC0)
👨‍💻 Beginner friendly✅ Yes, accessible for initial NLP projects
🔁 Fine-tuning ready📝 Suitable for fine-tuning detection models
🌍 Cultural diversity⚠️ Synthetic, no specific indication on linguistic diversity

🧠 Recommended for

  • NLP researchers
  • Moderation system developers
  • Machine learning students

🔧 Compatible tools

  • Scikit-learn
  • Hugging Face Transformers
  • SpacY
  • TensorFlow
  • PyTorch

💡 Tip

Train with complementary data sets to improve robustness in real conditions.

Frequently Asked Questions

Can this dataset be used to automatically detect hate speech in real time?

Yes, it is specifically designed to train and evaluate classification models in real time on social networks.

Does this dataset contain real or sensitive data?

No, it is a synthetic dataset guaranteeing confidentiality and the absence of truly sensitive content.

Does the dataset cover multiple languages?

No, the dataset is mainly in English and does not offer multilingual annotations.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.