By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Summary Screening Dataset
Text

Summary Screening Dataset

Annotated data set to train automatic CV screening systems, based on relevance labels.

Download dataset
Size

10,174 lines, CSV and Parquet (34 MB approximately)

Licence

MIT

Description

Summary Screening Dataset is a structured data set containing more than 10,000 sample resumes annotated according to their relevance to a position. It is intended for training natural language processing models capable of automating the sorting of applications according to their suitability for a given job offer.

What is this dataset for?

  • Develop automated screening systems for recruiters
  • Train binary classification models applied to applications
  • Testing NLP approaches for profile/offer matching

Can it be enriched or improved?

Yes. It is possible to add real job offers, metadata (sector, position, location), or to diversify CVs in language or structure. A multilingual extension or enrichment via syntactic parsing would make the dataset more robust.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Very easy – readable CSV format)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low – columns already structured)
🏷️ Annotation richness⭐⭐✩✩✩ (Basic – relevance present or not)
📜 Commercial license✅ Yes (MIT)
👨‍💻 Beginner friendly🌟 Perfect for starting with NLP classification
🔁 Fine-tuning ready🎯 Useful for fine-tuning BERT-like models on HR data
🌍 Cultural diversity⚠️ Needs strengthening – mostly monolingual and not geolocated

🧠 Recommended for

  • HR AI developers
  • Researchers in NLP applied to recruitment
  • NLP classification students

🔧 Compatible tools

  • Scikit-learn
  • Hugging Face Transformers
  • Pandas
  • FastText

💡 Tip

Complete the dataset with real job offers to model the CV-offer correspondence in a more refined way.

Frequently Asked Questions

Can this dataset be used in a commercial recruitment application?

Yes, the MIT license allows free commercial use with attribution.

Are resumes real or simulated?

The data is anonymized and does not allow real people to be identified.

Is the dataset adapted to languages other than English?

No, the examples are mostly in English. Multilingual adaptation would require translation and lexical readjustment.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.