Common Voice — Audio corpus for speech recognition
Common Voice is a community audio corpus, comprising thousands of voice recordings with their text transcripts. The dataset is mainly used for training and evaluating automatic speech recognition (ASR) systems.
Several thousand MP3 audio files accompanied by metadata CSV files, multilingual
CC0: Public Domain
Description
The dataset Common Voice contains thousands of MP3 audio files accompanied by transcripts validated by several listeners. It is divided into training, development, and test sets, with detailed metadata about the age, gender, focus, and quality of each recording.
What is this dataset for?
- Train robust ASR models on diverse voice data
- Evaluate speech recognition according to demographic and linguistic criteria
- Perform fine-tuning on specific voices or various languages
Can it be enriched or improved?
Yes, this corpus is open to community contribution via the Common Voice platform. Annotations can be refined by additional validation and languages extended. Additional metadata can also be added to improve use cases.
🔎 In summary
🧠 Recommended for
- Speech recognition researchers
- ASR developers
- Multilingual projects
🔧 Compatible tools
- Kaldi
- ESPnet
- Hugging Face Transformers
- PyTorch
- TensorFlow
💡 Tip
Select validated subassemblies for optimal and reliable training.
Frequently Asked Questions
What languages are covered by Common Voice?
Common Voice covers numerous languages and accents, with a strong presence of English, but also other languages depending on the contributions.
Are all audio files validated?
Only the files from the “valid” sets were validated by several auditors, guaranteeing the audio-text correspondence.
Can we contribute to this dataset?
Yes, Common Voice is an open community project where anyone can record and validate audio clips to enrich the corpus.




