By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
OlmoCR PES2O Dataset
Image

OlmoCR PES2O Dataset

Massive OCR corpus from scanned documents, used for optical character recognition and extracted text modeling.

Download dataset
Size

7.8 million lines, text files + structured metadata

Licence

ODC-BY

Description

‍

OlmoCR-PES2O-0225 is a massive open-source dataset designed for the tasks of optical character recognition (OCR). It contains millions of examples from scanned documents, each line representing an OCR instance with its extracted text and associated metadata. It is designed to be used for the training, evaluation, and performance analysis of modern OCR systems at the same time.

‍

‍

What is this dataset for?

‍

  • Train OCR models that can extract text from real documents
  • Test the robustness of models across a wide range of document types
  • Building AI-based document processing pipelines

‍

‍

Can it be enriched or improved?

‍

Yes. It is possible to add structural annotations (text areas, hierarchy), or to couple this corpus with the original images for visio-textual models. The addition of various languages or formats (e.g. manuscripts) is also possible.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Simple to parse, well-structured files)
🧼 Need for cleaning⭐⭐⭐⭐✩ (Low – data is consistent but OCR validation may be needed)
🏷️ Annotation richness⭐⭐⭐✩✩ (Medium – raw text + some metadata, no visual segmentation)
📜 Commercial license✅ Yes (ODC-BY)
👨‍💻 Beginner friendly⚠️ Yes, if OCR tools already mastered
🔁 Fine-tuning ready🎯 Perfect for fine-tuning OCR models like Tesseract, Donut, TrOCR
🌍 Cultural diversity⚠️ Not specified – mainly English content

‍

‍

🧠 Recommended for

  • OCR data engineers
  • Developers of automatic document reading engines
  • Digital archiving projects

‍

‍

🔧 Compatible tools

  • PyTesserAct
  • Hugging Face Transformers
  • doughnut
  • TroCR
  • LayoutLM

‍

‍

💡 Tip

Filter examples with low OCR confidence to improve the quality of training games.

Frequently Asked Questions

Does this dataset include the original images of the documents?

No, it only contains the extracted text and associated metadata. Source image is not provided.

Can it be used to train a complete OCR model?

Yes, but it will have to be coupled with similar images for the visual phase. The dataset is still useful for text preprocessing.

Does the content cover multiple languages or only English?

The dataset is mainly English. However, it can be enriched for multilingual applications.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.