OGC MEGA MultiDomain DocRetrieval
Massive dataset combining several sources to train multilingual visual documentary research models, with negative contrast support.
Around 1.09 million examples with images, query texts, languages (fr, en, de, es, it), JSON formats + PNG/JPEG images
Apache 2.0
Description
OGC MEGA MultiDomain DocRetrieval is the largest open-source dataset dedicated to visual documentary research. It combines over a million examples from seven different sources, covering military, energy, geotechnical, and more. Each entry contains a text query, a main image, and up to 16 negative images that can be used to train altered models. The dataset is structured and multilingual, covering five main languages including French, English, German, Spanish and Italian.
What is this dataset for?
- Train efficient visual literature search models
- Applying contrastive learning approaches to improve the relevance of search systems
- Strengthen the robustness of models on multilingual and varied documents
Can it be enriched or improved?
Yes. You can manually annotate the relevance of correspondence, add other languages, or even integrate metadata such as the type of document or its sectoral origin. The dataset is sufficiently modular to allow thematic or linguistic extension.
🔎 In summary
🧠 Recommended for
- Multimodal NLP Laboratories
- RAG engineers
- Industrial documentation systems
🔧 Compatible tools
- CLIP
- BLIP-2
- Hugging Face Transformers
- Faiss
- OpenAI Embeddings
💡 Tip
Remember to balance languages or domains in your batches to avoid training bias.
Frequently Asked Questions
Does the dataset contain manually annotated images?
No, the matches are implicit via the data structure. Negative examples are defined by field, but not annotated individually.
Do all languages have the same amount of data?
No, English is the majority (~ 64%), but French represents ~ 20%, followed by German, Spanish and Italian.
Is it compatible with models like CLIP or BLIP?
Yes, it is perfectly suited to contrasting or bimodal architectures such as CLIP, BLIP or OpenCLIP.




