By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
FineFrench v1
Text

FineFrench v1

FineFrench v1 is a gigantic French-language web corpus automatically filtered using a high-quality AI pipeline, aimed at eliminating noise, spam and useless content for training language models.

Download dataset
Size

66 million documents filtered in plain text, JSONL/parquet format, about several tens of GB

Licence

ODC-by 1.0

Description

FineFrench v1 is a French textual corpus resulting from automated sorting on a large scale. From 125 million web documents, around 66 million were retained after rigorous filtering, based on GPT-4-o annotations and a specialized BERT model. The pipeline aims to remove commercial content, spam, low-quality pages, and to keep only rich and informative texts.

What is this dataset for?

  • Training French-speaking language models (LLM)
  • Experiment with the automated cleaning of web texts
  • Serve as a basis for semantic analyses, summaries or content classification

Can it be enriched or improved?

Yes, this corpus can be further refined by thematic classification, semantic enrichment or deduplication by content families. It can also be annotated for specific tasks (NER, QA, summary).

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐✩✩ (Requires good infrastructure to handle volume)
🧼 Need for cleaning⭐⭐⭐⭐⭐ (Low: already cleaned and filtered automatically)
🏷️ Annotation richness⭐⭐✩✩✩ (No classic labels, but very thorough qualitative filtering)
📜 Commercial license✅ Yes (ODC-By 1.0)
👨‍💻 Beginner friendly❌ No – massive size, requires advanced skills
🔁 Fine-tuning ready✅ Excellent for pretraining or fine-tuning French LLMs
🌍 Cultural diversity🌐 Very large, web content from numerous Francophone sources

🧠 Recommended for

  • NLP Laboratories
  • French-speaking LLM training
  • Language AI startups

🔧 Compatible tools

  • Hugging Face Datasets
  • PyTorch
  • DeepSpeed
  • Ray
  • Apache Arrow

💡 Tip

Use additional thematic filters if you are looking for specialized corpora (legal, medical, etc.).

Frequently Asked Questions

What is the specificity of FineFrench v1 compared to other web datasets?

It has been automatically filtered using a BERT model and GPT-4-o annotations, ensuring superior quality and massive noise reduction.

Is this dataset multilingual?

No, it's exclusively in French, making it a great resource for French-speaking LLMs.

Is it suitable for a complete pre-training?

Yes, its size and quality allow a complete pre-training or retraining of French language models.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.