Clothing1M
Clothing1M is a dataset of more than one million clothing images from e-commerce sites. The labels are noisy because they are extracted automatically, but a subset of 74,000 images is properly annotated for training, validation, and testing.
Over 1 million JPEG images divided into 14 classes, with noisy and clean annotated data
Free download, license not explicitly specified
Description
Clothing1M is a large-scale dataset containing more than one million clothing images, divided into 14 categories. The images were collected from multiple online shopping platforms, resulting in a high rate of erroneous labels. However, the dataset includes 50,000 properly annotated images for training, 14,000 for validation, and 10,000 for testing.
What is this dataset for?
- Test the robustness of classification models in the face of noisy data
- Experiment with automatic cleaning or re-labelling techniques
- Training AI models applied to e-commerce and clothing recognition
Can it be enriched or improved?
Yes, you can refine labels, group classes, or integrate product metadata (color, genre, style). It is also possible to cross this dataset with other e-commerce or social datasets to enrich visual representations.
🔎 In summary
🧠 Recommended for
- NLP/Vision robustness projects
- Label noise researchers
- E-commerce AI
🔧 Compatible tools
- PyTorch
- TensorFlow
- FastAI
- Cleanlab
💡 Tip
Use label correction methods (confident learning) to extract a more reliable subset.
Frequently Asked Questions
What is the main benefit of Clothing1M despite the noise in the data?
It makes it possible to assess the robustness of models in the face of erroneous annotations, a common case in real data.
Does the dataset contain a properly annotated subset?
Yes, 50,000 images for training, 14,000 for validation, and 10,000 for testing are properly labeled.
Can this dataset be used for a commercial project?
Since the license is not specified, it is recommended that you contact the authors for commercial use.




