By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Text

SNLI

The SNLI dataset offers 570,000 pairs of manually annotated English sentences for the classification into nicking, contradiction, or neutral, essential for textual inference in NLP.

Download dataset
Size

570,000 phrase-premise/hypothesis pairs, JSON format

Licence

CC-BY-SA 4.0

Description

‍

SNLI is a major corpus for the Natural Language Inference (NLI) task. Each example contains a “premise” sentence and a “hypothesis” sentence annotated with a label indicating whether the second one implies the first one, contradicts it, or is neutral.

‍

‍

What is this dataset for?

‍

  • Train and evaluate textual inference models
  • Improving semantic understanding in NLP systems
  • Test the ability of models to reason about logical relationships between sentences

‍

‍

Can it be enriched or improved?

‍

Yes, by adding examples in other languages, by enriching the annotations, or by integrating explanations of decisions for more transparent learning.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Simple and widely supported format)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Very low – ready-to-use data)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good, with three balanced labels)
📜 Commercial license✅ Yes (CC-BY-SA 4.0)
👨‍💻 Beginner friendly🌟 Yes, often used as a reference dataset
🔁 Fine-tuning ready🎯 Perfect for fine-tuning on text comprehension tasks
🌍 Cultural diversity⚡ English, texts from Flickr users and Amazon Mechanical Turk

‍

‍

🧠 Recommended for

  • NLP researchers
  • AI developers
  • Deep learning students

‍

‍

🔧 Compatible tools

  • Hugging Face Transformers
  • PyTorch
  • TensorFlow
  • AllennLP

‍

‍

💡 Tip

Use SNLI in conjunction with other NLI datasets to improve model robustness.

Frequently Asked Questions

What type of annotations does the SNLI dataset contain?

Manual labels indicating whether the second sentence implies, contradicts, or is neutral to the first sentence.

Is the dataset suitable for commercial use?

Yes, the CC-BY-SA 4.0 license allows commercial use under the condition of sharing the same.

What is the average sentence size in SNLI?

On average, the “premise” sentence contains 14.1 tokens and the “hypothesis” 8.3 tokens.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.