CC 25M fondant
A vast dataset of 25 million image URLs extracted from the web, with Creative Commons license information, ideal for creating or refining computer generation and vision models.
25M images (via URLs), JSON metadata, links + license
Creative Commons (individual images under CC licenses)
Description
CC 25M fondant is a massive open-source dataset containing 25 million image URLs collected from the web via Common Crawl. Each link is associated with a Creative Commons license, guaranteeing the reusability of the content. This dataset was built with the Fondant framework, allowing efficient reproducibility for large-scale pipelines.
What is this dataset for?
- Preparing royalty-free visual datasets for model training
- Evaluate models for generating or searching images in real conditions
- Create thematic image corpora (via domain filtering, type, etc.)
Can it be enriched or improved?
Yes. This dataset is a raw database that can be filtered, sorted and enriched by classification, clustering or human validation. It is also possible to associate additional image vectors or metadata (categories, colors, source...).
🔎 In summary
🧠 Recommended for
- Image generation projects
- Custom crawl
- Free dataset constitution
🔧 Compatible tools
- Fondant
- DuckDB
- FastDownload
- Datasets Hugging Face
- PyTorch
💡 Tip
Filter images by domain (gov, edu, org) to improve reliability and cultural diversity.
Frequently Asked Questions
Does this dataset contain the images themselves?
No, it only contains URLs to publicly available images, with their license. You have to download them yourself.
Can it be used for a commercial project?
Yes, as long as each image complies with the associated CC license. Filtering by license type is recommended.
Is it possible to filter the dataset before use?
Yes, the dataset is compatible with tools like Fondant or Pandas to sort by domain, image type, language, etc.




