FineFrench v1
FineFrench v1 is a gigantic French-language web corpus automatically filtered using a high-quality AI pipeline, aimed at eliminating noise, spam and useless content for training language models.
66 million documents filtered in plain text, JSONL/parquet format, about several tens of GB
ODC-by 1.0
Description
FineFrench v1 is a French textual corpus resulting from automated sorting on a large scale. From 125 million web documents, around 66 million were retained after rigorous filtering, based on GPT-4-o annotations and a specialized BERT model. The pipeline aims to remove commercial content, spam, low-quality pages, and to keep only rich and informative texts.
What is this dataset for?
- Training French-speaking language models (LLM)
- Experiment with the automated cleaning of web texts
- Serve as a basis for semantic analyses, summaries or content classification
Can it be enriched or improved?
Yes, this corpus can be further refined by thematic classification, semantic enrichment or deduplication by content families. It can also be annotated for specific tasks (NER, QA, summary).
🔎 In summary
🧠 Recommended for
- NLP Laboratories
- French-speaking LLM training
- Language AI startups
🔧 Compatible tools
- Hugging Face Datasets
- PyTorch
- DeepSpeed
- Ray
- Apache Arrow
💡 Tip
Use additional thematic filters if you are looking for specialized corpora (legal, medical, etc.).
Frequently Asked Questions
What is the specificity of FineFrench v1 compared to other web datasets?
It has been automatically filtered using a BERT model and GPT-4-o annotations, ensuring superior quality and massive noise reduction.
Is this dataset multilingual?
No, it's exclusively in French, making it a great resource for French-speaking LLMs.
Is it suitable for a complete pre-training?
Yes, its size and quality allow a complete pre-training or retraining of French language models.




