RONEC
RONEC is a dataset in French dedicated to the recognition of named entities (NER) with 15 annotation classes, covering more than 80,000 distinct entities over more than 12,000 sentences.
Approximately 12,330 sentences, 0.5 million tokens, JSON/Parquet format, 2.94 MB
MIT
Description
The dataset RONEC contains 12,330 annotated sentences in French with 15 named entity classes. It has more than 80,000 distinct entities, making it a valuable resource for training specialized NER models on the French language.
What is this dataset for?
- Train named entity recognition (NER) models in French
- Improving the extraction of information from French-speaking texts
- Test the accuracy of NER models on manually annotated data
Can it be enriched or improved?
Yes, the dataset can be supplemented with additional annotations, including adding specific classes or correcting existing annotations. It is also possible to expand the corpus to include more fields or contexts.
🔎 In summary
🧠 Recommended for
- French NLP researchers
- AI developers
- Students in automatic language processing
🔧 Compatible tools
- SpacY
- Hugging Face Transformers
- Flair
- Jupyter notebooks
💡 Tip
For best results, combine this dataset with other various French-speaking NER resources.
Frequently Asked Questions
What are the main classes of annotated entities in RONEC?
The dataset contains 15 classes, including people, organizations, places, dates, and other specific entities.
Is the dataset suitable for commercial use?
Yes, the MIT license allows unrestricted commercial use.
What file format is provided for this dataset?
The data is available in JSON and Parquet format, making it easy to integrate it into various NLP pipelines.




