By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
Synthetic ESCO Skill Sentences
Text

Synthetic ESCO Skill Sentences

Data set containing synthetic sentences representing job advertisements, associated with 99.5% of the skills in the ESCO v1.1.0 database.

Download dataset
Size

138,260 phrase-skill pairs, text, JSON format

Licence

CC BY 4.0

Description

Synthetic ESCO Skill Sentences is a dataset of more than 138,000 sentences automatically generated from the ESCO (European Skills, Competences, Qualifications and Occupations) taxonomy. Each sentence contextualizes an ESCO skill in a typical professional situation, in order to simulate real job advertisements.

What is this dataset for?

  • Train competency extraction or recognition models in HR texts
  • Test or create systems for recommending profiles or job offers
  • Building multi-label classification models on corpora related to employment or training

Can it be enriched or improved?

Yes, it is possible to generate more complex or realistic variants of current sentences, or to annotate sentences with specific business contexts. It can also be adapted to other languages by translating sentences or regenerating them with multilingual templates.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐⭐ (Simple structure - sentence + skill)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (None – clean and consistent data)
🏷️ Annotation richness⭐⭐⭐✩✩ (Medium – each sentence linked to an ESCO skill)
📜 Commercial license✅ Yes (CC BY 4.0)
👨‍💻 Beginner friendly✅ Yes – accessible and ready-to-use format
🔁 Fine-tuning ready🗂️ Very good for text classification or tagging models
🌍 Cultural diversity⚠️ Low – sentences only in English, synthetically generated

🧠 Recommended for

  • HR projects
  • Skills research
  • ATS or guidance tool developers

🔧 Compatible tools

  • SpacY
  • Scikit-learn
  • Hugging Face Transformers
  • TARS
  • Flair

💡 Tip

Combine this dataset with sentences from real ads to create mixed sets that are more robust to human language variability.

Frequently Asked Questions

Are the sentences in the dataset taken from real job offers?

No, the sentences are generated automatically from the ESCO database for each skill, and do not come from real texts.

Can this dataset be used to train a skill extraction model?

Yes, it is ideal for this because each sentence contains a clearly identified skill, useful for over-learning or testing.

Can we translate this dataset into other languages?

Yes, sentences can be translated automatically or regenerated from ESCO skills in other languages.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.