Summary Screening Dataset
Annotated data set to train automatic CV screening systems, based on relevance labels.
Description
Summary Screening Dataset is a structured data set containing more than 10,000 sample resumes annotated according to their relevance to a position. It is intended for training natural language processing models capable of automating the sorting of applications according to their suitability for a given job offer.
What is this dataset for?
- Develop automated screening systems for recruiters
- Train binary classification models applied to applications
- Testing NLP approaches for profile/offer matching
Can it be enriched or improved?
Yes. It is possible to add real job offers, metadata (sector, position, location), or to diversify CVs in language or structure. A multilingual extension or enrichment via syntactic parsing would make the dataset more robust.
🔎 In summary
🧠 Recommended for
- HR AI developers
- Researchers in NLP applied to recruitment
- NLP classification students
🔧 Compatible tools
- Scikit-learn
- Hugging Face Transformers
- Pandas
- FastText
💡 Tip
Complete the dataset with real job offers to model the CV-offer correspondence in a more refined way.
Frequently Asked Questions
Can this dataset be used in a commercial recruitment application?
Yes, the MIT license allows free commercial use with attribution.
Are resumes real or simulated?
The data is anonymized and does not allow real people to be identified.
Is the dataset adapted to languages other than English?
No, the examples are mostly in English. Multilingual adaptation would require translation and lexical readjustment.




