By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Clothing1M
Image

Clothing1M

Clothing1M is a dataset of more than one million clothing images from e-commerce sites. The labels are noisy because they are extracted automatically, but a subset of 74,000 images is properly annotated for training, validation, and testing.

Download dataset
Size

Over 1 million JPEG images divided into 14 classes, with noisy and clean annotated data

Licence

Free download, license not explicitly specified

Description

‍

Clothing1M is a large-scale dataset containing more than one million clothing images, divided into 14 categories. The images were collected from multiple online shopping platforms, resulting in a high rate of erroneous labels. However, the dataset includes 50,000 properly annotated images for training, 14,000 for validation, and 10,000 for testing.

‍

‍

What is this dataset for?

‍

  • Test the robustness of classification models in the face of noisy data
  • Experiment with automatic cleaning or re-labelling techniques
  • Training AI models applied to e-commerce and clothing recognition

‍

‍

Can it be enriched or improved?

‍

Yes, you can refine labels, group classes, or integrate product metadata (color, genre, style). It is also possible to cross this dataset with other e-commerce or social datasets to enrich visual representations.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Large volume, requires prior filtering)
🧼 Need for cleaning⭐⭐✩✩✩ (High: massive presence of incorrect labels)
🏷️ Annotation richness⭐⭐⭐✩✩ (Low overall, clean on a subset)
📜 Commercial license⚠️ Not specified, verify for specific usage
👨‍💻 Beginner friendly⚠️ Moderate, good for experimentation but may be complex to exploit
🔁 Fine-tuning ready✅ Very suitable for fine-tuning robust models
🌍 Cultural diversity⚠️ Low: styles from Asian and Western e-commerce platforms

‍

‍

🧠 Recommended for

  • NLP/Vision robustness projects
  • Label noise researchers
  • E-commerce AI

‍

‍

🔧 Compatible tools

  • PyTorch
  • TensorFlow
  • FastAI
  • Cleanlab

‍

‍

💡 Tip

Use label correction methods (confident learning) to extract a more reliable subset.

Frequently Asked Questions

What is the main benefit of Clothing1M despite the noise in the data?

It makes it possible to assess the robustness of models in the face of erroneous annotations, a common case in real data.

Does the dataset contain a properly annotated subset?

Yes, 50,000 images for training, 14,000 for validation, and 10,000 for testing are properly labeled.

Can this dataset be used for a commercial project?

Since the license is not specified, it is recommended that you contact the authors for commercial use.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.