By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
OGC MEGA MultiDomain DocRetrieval
Image

OGC MEGA MultiDomain DocRetrieval

Massive dataset combining several sources to train multilingual visual documentary research models, with negative contrast support.

Download dataset
Size

Around 1.09 million examples with images, query texts, languages (fr, en, de, es, it), JSON formats + PNG/JPEG images

Licence

Apache 2.0

Description

OGC MEGA MultiDomain DocRetrieval is the largest open-source dataset dedicated to visual documentary research. It combines over a million examples from seven different sources, covering military, energy, geotechnical, and more. Each entry contains a text query, a main image, and up to 16 negative images that can be used to train altered models. The dataset is structured and multilingual, covering five main languages including French, English, German, Spanish and Italian.

What is this dataset for?

  • Train efficient visual literature search models
  • Applying contrastive learning approaches to improve the relevance of search systems
  • Strengthen the robustness of models on multilingual and varied documents

Can it be enriched or improved?

Yes. You can manually annotate the relevance of correspondence, add other languages, or even integrate metadata such as the type of document or its sectoral origin. The dataset is sufficiently modular to allow thematic or linguistic extension.

🔎 In summary

Criterion Evaluation
🧩Ease of Use ⭐⭐☆☆☆ (average – requires image and text preprocessing)
🧼Need for Cleaning ⭐☆ (low – well structured, but image file handling required)
🏷️Annotation Richness ⭐⭐⭐⭐☆ (solid: text, image, negative labels, language)
📜Commercial License ✅ Yes (Apache 2.0)
👨‍💻Beginner-Friendly ⚠️ Recommended with supervision – high volume and complexity
🔁Reusable for Fine-Tuning 💪 Perfect for contrastive or multimodal retrieval models
🌍Cultural Diversity 🌟 Excellent – multilingual and multi-domain

🧠 Recommended for

  • Multimodal NLP Laboratories
  • RAG engineers
  • Industrial documentation systems

🔧 Compatible tools

  • CLIP
  • BLIP-2
  • Hugging Face Transformers
  • Faiss
  • OpenAI Embeddings

💡 Tip

Remember to balance languages or domains in your batches to avoid training bias.

Frequently Asked Questions

Does the dataset contain manually annotated images?

No, the matches are implicit via the data structure. Negative examples are defined by field, but not annotated individually.

Do all languages have the same amount of data?

No, English is the majority (~ 64%), but French represents ~ 20%, followed by German, Spanish and Italian.

Is it compatible with models like CLIP or BLIP?

Yes, it is perfectly suited to contrasting or bimodal architectures such as CLIP, BLIP or OpenCLIP.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.