By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
DAIGT Proper Train Dataset
Text

DAIGT Proper Train Dataset

Text dataset composed of several thousand examples designed to train models to detect text generated by AI. It includes texts generated by various major language models (ChatGPT, LLama, Falcon, Claude, Mistral) as well as official essays. This corpus is suitable for fine-tuning and evaluating AI detection models.

Download dataset
Size

Several thousand text essays in plain text or JSON format

Licence

MIT

Description

The DAIGT Proper Train Dataset is a rich textual corpus including several thousand examples of essays generated by various AI language models as well as human trials. It contains metadata such as the test ID and the generation prompt when available. This dataset is designed for AI-generated text detection competitions and projects.

What is this dataset for?

  • Train classification models to distinguish human text and AI text
  • Evaluate the robustness of AI detectors on a variety of data
  • Fine-tuning to improve detection in different contexts

Can it be enriched or improved?

Yes, it is possible to add texts generated by new models or to integrate additional annotations (e.g. difficulty, subject). The dataset can also be stratified by source or prompt for fine analyses.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Ready-to-use dataset, useful metadata)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – well-organized data)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good – metadata on source and prompts)
📜 Commercial license✅ Yes (MIT)
👨‍💻 Beginner friendly🌟 Suitable for beginner and advanced users
🔁 Fine-tuning ready🎯 Excellent for AI text detection and classification
🌍 Cultural diversity📖 Varied texts produced by multiple LLMs

🧠 Recommended for

  • NLP researchers
  • AI detection teams
  • Generative AI competitors

🔧 Compatible tools

  • Hugging Face Transformers
  • Scikit-learn
  • PyTorch
  • TensorFlow

💡 Tip

Use metadata to create balanced and varied training sets.

Frequently Asked Questions

Does this dataset only contain AI-generated text?

No, it also contains human trials to allow for balanced learning.

What is the license of the dataset?

The license is MIT, which allows free use including commercial use.

Can this dataset be used to train plagiarism or similarity detectors?

Yes, it can be used for these tasks, especially as a complement to other corpora.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.