By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Coyo 700M
Multimodal

Coyo 700M

Massive dataset of image-text pairs extracted from HTML documents, intended to train multimodal vision-language foundation models.

Download dataset
Size

Approximately 747 million image+text pairs, 135 GB in Parquet files

Licence

CC-BY 4.0

Description

Coyo 700M is a very large data set comprising approximately 747 million pairs of images and associated texts extracted from HTML documents. Each pair is accompanied by metadata allowing a variety of use in training multimodal AI models.

What is this dataset for?

  • Train multimodal vision-language models on a large scale
  • Improving the visual and textual comprehension skills of multimodal LLMs
  • Complete other datasets for the formation of foundation models

Can it be enriched or improved?

The dataset can be refined by filtering, cleaning, or reannotation to improve pair quality. The integration of additional metadata or the translation of texts can also enrich the corpus.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐✩✩✩ (Requires massive processing and significant resources)
🧼 Need for cleaning⭐⭐✩✩✩ (High, filtering needed for optimal quality)
🏷️ Annotation richness⭐⭐⭐✩✩ (Varied metadata, but basic annotations)
📜 Commercial license✅ CC-BY 4.0 – free commercial use
👨‍💻 Beginner friendly⚠️ No, volume and processing are complex
🔁 Fine-tuning ready✅ Excellent basis for multimodal models
🌍 Cultural diversity🌐 Wide diversity from global web sources

🧠 Recommended for

  • Multimodal AI researchers
  • Foundation Models Developers
  • Advanced data scientists

🔧 Compatible tools

  • Apache Spark
  • Hugging Face Datasets
  • TensorFlow
  • PyTorch
  • Dask

💡 Tip

Prioritize advanced filtering to improve the quality of the dataset before training.

Frequently Asked Questions

What is the exact size of the Coyo 700M dataset?

The dataset contains approximately 747 million image-text pairs, occupying approximately 135 GB in Parquet format.

Can I use this dataset for commercial use?

Yes, the CC-BY 4.0 license allows commercial use provided the source is cited.

Is this dataset suitable for machine learning beginners?

No, the size and complexity of the dataset require advanced resources and skills.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.