RefBlock
Multimodal dataset designed to detect and link text blocks containing scientific references in PDF documents, useful for training or evaluating literature review models.
Description
RefBlock is a multimodal dataset that combines text snippets, positioning blocks in PDF documents, and scientific reference metadata. It is designed to train or evaluate systems that can locate and link citations in complex documents such as academic articles. Each entry includes a text snippet, block coordinates, and data on the cited reference.
What is this dataset for?
- Automatically detect citations in academic PDFs
- Forming models for the link between text and bibliographical references
- Building parsing or indexing systems for scientific archives
Can it be enriched or improved?
Yes, RefBlock can be enriched with other document formats (e.g. Word, LaTeX), additional languages, or citation type annotations (direct, indirect). Adding additional metadata to source documents or standardizing bibliographic styles can also increase its value.
🔎 In summary
🧠 Recommended for
- Academic NLP engineers
- Information extraction researchers
- OCR+LLM projects
🔧 Compatible tools
- Hugging Face Datasets
- PymuPDF
- LayoutLM
- GROBID
- PDFMiner
💡 Tip
Use block coordinates to visualize areas of attention in multimodal models.
Frequently Asked Questions
Does this dataset contain the source PDFs?
No, it only contains the coordinates of the extracted blocks and their textual content, as well as the associated metadata.
Can it be used to train an OCR model?
It is not a raw OCR dataset, but it can complement OCR sets for post-analysis or text and structure alignment.
Is RefBlock suitable for looking for plagiarism or textual similarity?
Yes, by analyzing cross-references and blocks of quotations, it can be used as a basis for projects to detect similarities or plagiarism.




