Coyo 700M
Massive dataset of image-text pairs extracted from HTML documents, intended to train multimodal vision-language foundation models.
Approximately 747 million image+text pairs, 135 GB in Parquet files
CC-BY 4.0
Description
Coyo 700M is a very large data set comprising approximately 747 million pairs of images and associated texts extracted from HTML documents. Each pair is accompanied by metadata allowing a variety of use in training multimodal AI models.
What is this dataset for?
- Train multimodal vision-language models on a large scale
- Improving the visual and textual comprehension skills of multimodal LLMs
- Complete other datasets for the formation of foundation models
Can it be enriched or improved?
The dataset can be refined by filtering, cleaning, or reannotation to improve pair quality. The integration of additional metadata or the translation of texts can also enrich the corpus.
🔎 In summary
🧠 Recommended for
- Multimodal AI researchers
- Foundation Models Developers
- Advanced data scientists
🔧 Compatible tools
- Apache Spark
- Hugging Face Datasets
- TensorFlow
- PyTorch
- Dask
💡 Tip
Prioritize advanced filtering to improve the quality of the dataset before training.
Frequently Asked Questions
What is the exact size of the Coyo 700M dataset?
The dataset contains approximately 747 million image-text pairs, occupying approximately 135 GB in Parquet format.
Can I use this dataset for commercial use?
Yes, the CC-BY 4.0 license allows commercial use provided the source is cited.
Is this dataset suitable for machine learning beginners?
No, the size and complexity of the dataset require advanced resources and skills.




