By clicking "Accept", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. See our Privacy Policy for more information
Open Datasets
English—French Parallel Translation Dataset
Text

English—French Parallel Translation Dataset

Important bilingual English-French corpus aligned sentence by sentence, from web crawl, designed for machine translation and multilingual language models.

Download dataset
Size

~22,500,000 aligned sentences, text format (TSV/CSV) two columns (English/French)

Licence

Open Database License (ODbL) or equivalent

Description

This dataset contains approximately 22,500,000 pairs of parallel English‑French sentences. It was built using web collection and heuristics to align translated documents, and served as the basis for the Workshop on Statistical Machine Translation 2015.

What is this dataset for?

  • Training high performance machine translation models
  • Fine-tuner multilingual LLMs for translation tasks or bilingual generation
  • Study corpus alignments and statistical/hybrid translation methods

Can it be enriched or improved?

Yes, it can be filtered for quality, cleaned for unwanted marks, or enriched with phrase-document alignments or contextual metadata.

🔎 In summary

Criterion Evaluation
🧩 Ease of use⭐⭐⭐⭐✩ (Simple format, but very large file)
🧼 Need for cleaning⭐⭐⭐✩✩ (Moderate to high: some pairs may be noisy)
🏷️ Annotation richness⭐⭐✩✩✩ (Basic: sentence-to-sentence alignment without contextual metadata)
📜 Commercial license✅ Yes – ODbL license (free use with attribution/share)
👨‍💻 Beginner friendly⚠️ Moderate – very heavy file to handle
🔁 Fine-tuning ready✅ Excellent for fine-tuning translation models
🌍 Cultural diversity🌏 Good linguistic balance, but probably biased toward English web

🧠 Recommended for

  • Machine translation researchers
  • NLP engineers
  • Multilingual LLM projects

🔧 Compatible tools

  • OpenNMT
  • MarianMt
  • Fairseq
  • Hugging Face Transformers

💡 Tip

Divide the corpus into sub-selections by domain for more accurate or manageable fine‑tuning.

Frequently Asked Questions

Is this dataset cleaned up and ready to use?

It is aligned, but often requires cleaning up on noisy pairs and encoding errors.

What is the main language used?

Mostly in English and French, from the web, with a probable orientation on English content.

Can it be used for commercial use?

Yes, the ODbL license allows commercial use, provided that derivatives are attributed and shared under the same license.

Similar datasets

See more
Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.

Category

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.