OlmoCR PES2O Dataset
Massive OCR corpus from scanned documents, used for optical character recognition and extracted text modeling.
Description
OlmoCR-PES2O-0225 is a massive open-source dataset designed for the tasks of optical character recognition (OCR). It contains millions of examples from scanned documents, each line representing an OCR instance with its extracted text and associated metadata. It is designed to be used for the training, evaluation, and performance analysis of modern OCR systems at the same time.
What is this dataset for?
- Train OCR models that can extract text from real documents
- Test the robustness of models across a wide range of document types
- Building AI-based document processing pipelines
Can it be enriched or improved?
Yes. It is possible to add structural annotations (text areas, hierarchy), or to couple this corpus with the original images for visio-textual models. The addition of various languages or formats (e.g. manuscripts) is also possible.
🔎 In summary
🧠 Recommended for
- OCR data engineers
- Developers of automatic document reading engines
- Digital archiving projects
🔧 Compatible tools
- PyTesserAct
- Hugging Face Transformers
- doughnut
- TroCR
- LayoutLM
💡 Tip
Filter examples with low OCR confidence to improve the quality of training games.
Frequently Asked Questions
Does this dataset include the original images of the documents?
No, it only contains the extracted text and associated metadata. Source image is not provided.
Can it be used to train a complete OCR model?
Yes, but it will have to be coupled with similar images for the visual phase. The dataset is still useful for text preprocessing.
Does the content cover multiple languages or only English?
The dataset is mainly English. However, it can be enriched for multilingual applications.




