PD12M: Public Domain Image Collection
PD12M is a large dataset of public domain images, containing 12 million filtered items with active links. Each image is accompanied by EXIF metadata, an embedding vector optimized for vector search and indexing, facilitating advanced queries and computer vision applications.
12 million public images, JPEG/PNG format, 16-bit half-precision vectors (1152 dimensions) optimized for L2 indexing
Apache 2.0
Description
The dataset PD12M brings together a massive collection of public domain images, enriched with detailed metadata such as EXIF and lengthy descriptions. The images are preprocessed and equipped with high-precision vector embeddings, suitable for similarity search and efficient indexing.
What is this dataset for?
- Search for images and matches by similarity in large databases
- Training and evaluation of computer vision or multimodal models
- Development of AI applications using EXIF metadata and vector embeddings
Can it be enriched or improved?
Yes, it is possible to add custom annotations, extend metadata, or update the dataset regularly with new public images. The standardization of formats and the optimization of embeddings can also be adapted according to specific needs.
🔎 In summary
🧠 Recommended for
- Computer vision researchers
- Vector search developers
- Multimodal AI projects
🔧 Compatible tools
- PostgreSQL with vectorchord/pg_search extensions
- Python
- DO
- Annoy
💡 Tip
Use standardized embeddings to speed up similar searches and maximize the accuracy of results.
Frequently Asked Questions
Does this dataset only contain royalty-free images?
Yes, all images are in the public domain and are royalty-free for commercial and non-commercial use.
What type of metadata goes with images?
The dataset includes EXIF metadata, long descriptions, dimensions, and vector embeddings for each image.
Do you need specific technical skills to use PD12M?
Yes, knowledge of relational databases and vector indexing is recommended to fully exploit this dataset.




