By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
CC 25M fondant
Image

CC 25M fondant

A vast dataset of 25 million image URLs extracted from the web, with Creative Commons license information, ideal for creating or refining computer generation and vision models.

Download dataset
Size

25M images (via URLs), JSON metadata, links + license

Licence

Creative Commons (individual images under CC licenses)

Description

CC 25M fondant is a massive open-source dataset containing 25 million image URLs collected from the web via Common Crawl. Each link is associated with a Creative Commons license, guaranteeing the reusability of the content. This dataset was built with the Fondant framework, allowing efficient reproducibility for large-scale pipelines.

What is this dataset for?

  • Preparing royalty-free visual datasets for model training
  • Evaluate models for generating or searching images in real conditions
  • Create thematic image corpora (via domain filtering, type, etc.)

Can it be enriched or improved?

Yes. This dataset is a raw database that can be filtered, sorted and enriched by classification, clustering or human validation. It is also possible to associate additional image vectors or metadata (categories, colors, source...).

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐✩✩✩ (Requires initial processing - download, filtering)
🧼 Need for cleaning⭐⭐✩✩✩ (High – links may be broken or outdated)
🏷️ Annotation richness⭐⭐✩✩✩ (Limited – only URL + license)
📜 Commercial license⚠️ Yes (according to each image’s CC license)
👨‍💻 Beginner friendly⚠️ No – requires automation tools
🔁 Fine-tuning ready✅ Yes, after filtering and processing
🌍 Cultural diversity🌐 High – global and varied web content

🧠 Recommended for

  • Image generation projects
  • Custom crawl
  • Free dataset constitution

🔧 Compatible tools

  • Fondant
  • DuckDB
  • FastDownload
  • Datasets Hugging Face
  • PyTorch

💡 Tip

Filter images by domain (gov, edu, org) to improve reliability and cultural diversity.

Frequently Asked Questions

Does this dataset contain the images themselves?

No, it only contains URLs to publicly available images, with their license. You have to download them yourself.

Can it be used for a commercial project?

Yes, as long as each image complies with the associated CC license. Filtering by license type is recommended.

Is it possible to filter the dataset before use?

Yes, the dataset is compatible with tools like Fondant or Pandas to sort by domain, image type, language, etc.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.