Topic: Finance QA Vision 100k
This dataset contains over 100,000 English question and answer pairs, generated from nearly 10,000 images of financial documents. It is designed to train or test visual question and answer (VQA) systems in the financial field.
Description
Topic: Finance QA Vision 100k is a massive data set combining visual financial documents and textual questions and answers. It includes 107,050 QA pairs from 9,801 images, representing various types of documents (reports, tables, invoices, etc.).
What is this dataset for?
- Develop VQA (Visual Question Answering) systems applied to complex documents
- Improving the accuracy of financial document analysis via OCR + semantic understanding
- Test or refine multimodal LLMs in a structured professional context
Can it be enriched or improved?
Yes, this dataset can be enriched with metadata (types of documents, level of complexity, origin). It is also possible to add translations, OCR annotations, or multiple responses depending on ambiguous cases.
🔎 In summary
🧠 Recommended for
- VQA researchers
- Smart OCR developers
- LLMs projects applied to finance
🔧 Compatible tools
- doughnut
- LMv3 layout
- Pix2Struct
- Tesseract
- Hugging Face Transformers
💡 Tip
Filter questions by structure or document type for more effective specialized training.
Frequently Asked Questions
Are questions and answers extracted automatically or manually?
They are generated automatically from the content of the documents, with a high level of consistency between visual and text.
Is this dataset suitable for models like LayoutLM or Pix2Struct?
Yes, it is fully compatible with documentary VQA models and structured word processing in images.
Can it be used in a multilingual context?
Currently, it is only in English, but can be enriched by translation or multilingual alignment.




