By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Text

TLDR-17

Massive dataset of Reddit posts in English with content and associated summary, designed to train and evaluate automatic summarization models.

Download dataset
Size

Approximately 3.8 million posts with summary, JSON files, total size ~19 GB

Licence

CC-BY 4.0

Description

The dataset TLDR-17 contains more than 3.8 million Reddit posts in English, each with a short summary (on average 28 words). Data is preprocessed and provided in JSON format with multiple fields such as author, content, summary, and original subreddit.

What is this dataset for?

  • Train automatic summary models (abstractive summarization)
  • Evaluate the performance of NLP models on real social content
  • Develop content synthesis systems for forums and social networks

Can it be enriched or improved?

This dataset can be enriched by adding qualitative annotations on the quality of the summaries, or by creating multi-language versions. Selecting and cleaning can improve model performance.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Large volume, requires resources and preprocessing)
🧼 Need for cleaning⭐⭐⭐✩✩ (Moderate to high – data from Reddit with possible noise)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good, human summary provided)
📜 Commercial license✅ Yes (CC-BY 4.0)
👨‍💻 Beginner friendly⚠️ Moderate, rather for users with significant resources
🔁 Fine-tuning ready🎯 Perfect for fine-tuning summarization models
🌍 Cultural diversity⚡ English, broad content from Reddit communities

🧠 Recommended for

  • NLP researchers
  • AI social media teams
  • Chatbot developers and assistants

🔧 Compatible tools

  • Hugging Face Transformers
  • TensorFlow
  • PyTorch
  • SpacY
  • Jupyter notebooks

💡 Tip

Pre-filter the data to limit the noise associated with certain subreddits or off-topic posts.

Frequently Asked Questions

Does this dataset include metadata about the authors or the date of the posts?

Yes, fields like author, subreddit, and subreddit_id are included, but specific dates are not always present.

Can this dataset be used for tasks other than the automatic summary?

Mostly summary-oriented, but the content can be reused for other NLP tasks such as text classification.

What is the total size of the downloaded dataset?

Approximately 19 GB of data generated, with 3.14 GB of compressed files to download.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.