English—French Parallel Translation Dataset
Important bilingual English-French corpus aligned sentence by sentence, from web crawl, designed for machine translation and multilingual language models.
~22,500,000 aligned sentences, text format (TSV/CSV) two columns (English/French)
Open Database License (ODbL) or equivalent
Description
This dataset contains approximately 22,500,000 pairs of parallel English‑French sentences. It was built using web collection and heuristics to align translated documents, and served as the basis for the Workshop on Statistical Machine Translation 2015.
What is this dataset for?
- Training high performance machine translation models
- Fine-tuner multilingual LLMs for translation tasks or bilingual generation
- Study corpus alignments and statistical/hybrid translation methods
Can it be enriched or improved?
Yes, it can be filtered for quality, cleaned for unwanted marks, or enriched with phrase-document alignments or contextual metadata.
🔎 In summary
🧠 Recommended for
- Machine translation researchers
- NLP engineers
- Multilingual LLM projects
🔧 Compatible tools
- OpenNMT
- MarianMt
- Fairseq
- Hugging Face Transformers
💡 Tip
Divide the corpus into sub-selections by domain for more accurate or manageable fine‑tuning.
Frequently Asked Questions
Is this dataset cleaned up and ready to use?
It is aligned, but often requires cleaning up on noisy pairs and encoding errors.
What is the main language used?
Mostly in English and French, from the web, with a probable orientation on English content.
Can it be used for commercial use?
Yes, the ODbL license allows commercial use, provided that derivatives are attributed and shared under the same license.




