By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
RefBlock
Multimodal

RefBlock

Multimodal dataset designed to detect and link text blocks containing scientific references in PDF documents, useful for training or evaluating literature review models.

Download dataset
Size

1,250 examples, 18 MB, Parquet format

Licence

CC-BY 4.0

Description

‍

RefBlock is a multimodal dataset that combines text snippets, positioning blocks in PDF documents, and scientific reference metadata. It is designed to train or evaluate systems that can locate and link citations in complex documents such as academic articles. Each entry includes a text snippet, block coordinates, and data on the cited reference.

‍

‍

What is this dataset for?

‍

  • Automatically detect citations in academic PDFs
  • Forming models for the link between text and bibliographical references
  • Building parsing or indexing systems for scientific archives

‍

‍

Can it be enriched or improved?

‍

Yes, RefBlock can be enriched with other document formats (e.g. Word, LaTeX), additional languages, or citation type annotations (direct, indirect). Adding additional metadata to source documents or standardizing bibliographic styles can also increase its value.

‍

‍

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Requires reading block coordinates)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – data well-formatted)
🏷️ Annotation richness⭐⭐⭐⭐✩ (Good – blocks + text + metadata)
📜 Commercial license✅ Yes (CC-BY 4.0)
👨‍💻 Beginner friendly⚠️ Requires basic PDF manipulation knowledge
🔁 Fine-tuning ready🎯 Relevant for training on structured PDF data
🌍 Cultural diversity⚠️ Mainly English – can be expanded

‍

‍

🧠 Recommended for

  • Academic NLP engineers
  • Information extraction researchers
  • OCR+LLM projects

‍

‍

🔧 Compatible tools

  • Hugging Face Datasets
  • PymuPDF
  • LayoutLM
  • GROBID
  • PDFMiner

‍

‍

💡 Tip

Use block coordinates to visualize areas of attention in multimodal models.

Frequently Asked Questions

Does this dataset contain the source PDFs?

No, it only contains the coordinates of the extracted blocks and their textual content, as well as the associated metadata.

Can it be used to train an OCR model?

It is not a raw OCR dataset, but it can complement OCR sets for post-analysis or text and structure alignment.

Is RefBlock suitable for looking for plagiarism or textual similarity?

Yes, by analyzing cross-references and blocks of quotations, it can be used as a basis for projects to detect similarities or plagiarism.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.