Synthetic ESCO Skill Sentences
Data set containing synthetic sentences representing job advertisements, associated with 99.5% of the skills in the ESCO v1.1.0 database.
Description
Synthetic ESCO Skill Sentences is a dataset of more than 138,000 sentences automatically generated from the ESCO (European Skills, Competences, Qualifications and Occupations) taxonomy. Each sentence contextualizes an ESCO skill in a typical professional situation, in order to simulate real job advertisements.
What is this dataset for?
- Train competency extraction or recognition models in HR texts
- Test or create systems for recommending profiles or job offers
- Building multi-label classification models on corpora related to employment or training
Can it be enriched or improved?
Yes, it is possible to generate more complex or realistic variants of current sentences, or to annotate sentences with specific business contexts. It can also be adapted to other languages by translating sentences or regenerating them with multilingual templates.
🔎 In summary
🧠 Recommended for
- HR projects
- Skills research
- ATS or guidance tool developers
🔧 Compatible tools
- SpacY
- Scikit-learn
- Hugging Face Transformers
- TARS
- Flair
💡 Tip
Combine this dataset with sentences from real ads to create mixed sets that are more robust to human language variability.
Frequently Asked Questions
Are the sentences in the dataset taken from real job offers?
No, the sentences are generated automatically from the ESCO database for each skill, and do not come from real texts.
Can this dataset be used to train a skill extraction model?
Yes, it is ideal for this because each sentence contains a clearly identified skill, useful for over-learning or testing.
Can we translate this dataset into other languages?
Yes, sentences can be translated automatically or regenerated from ESCO skills in other languages.




