AI vs Human Generated Dataset
This image dataset juxtaposes authentic Shutterstock content and AI-generated images to train synthetic content detection models.
Description
This corpus includes 85,500 images divided into two categories: authentic images from Shutterstock, and equivalents generated by cutting-edge AI models. It allows for rigorous comparative analysis, with a balanced structure (including images of human persons).
What is this dataset for?
- Develop models for detecting AI-generated images
- Create benchmarks to assess the veracity of visual content
- Study the perceptible differences between synthetic and real visuals
Can it be enriched or improved?
Yes, it is possible to add additional annotations (level of realism, type of generation, human detection success rate), or to increase the corpus with other AI generators. Enrichment through emotional or contextual labelling can also open up new lines of research.
🔎 In summary
🧠 Recommended for
- AI computer vision developers
- Visual cybersecurity researchers
- Kaggle competitions
🔧 Compatible tools
- PyTorch
- TensorFlow
- OpenCV
- Keras
- Label Studio
💡 Tip
To test your models, apply training on AI images and cross-test on generations from other tools (DALL·E, Midjourney, etc.)
Frequently Asked Questions
Are all images labelled according to their origin?
Yes, each image comes with metadata about whether it's authentic or AI-generated, making supervised training easy.
Can this dataset be used in a commercial context?
Yes, the Apache 2.0 license allows commercial exploitation as long as license conditions are met.
Is this dataset suitable for detecting deepfakes?
It is more suitable for the detection of artificially generated images (general synthesis), but can be combined with other specialized corpora for deepfake.




