Common Voice segments now available through Mozilla Data Collective
Mozilla Common Voice is a massively multilingual platform for collecting speech data to train automatic speech recognition (ASR). Its mission is simple: to make language technology understand everyone’s mother tongue.
But for datasets to be genuinely useful, they also need to be manageable. Many of the larger Common Voice datasets can be tens of gigabytes in size, which can make processing them difficult for builders and researchers who have limited technical resources.
One of the most consistent requests we’ve heard over the years is that, for many of the world’s larger languages, people want access to only the portion that’s relevant to them. Whether you’re building a model for a specific region or creating resources for a particular community, it shouldn’t require downloading and handling far more data than you actually need.
With that in mind, we’ve started a pilot programme to ship segments of Common Voice, so you can get the data you need for training models or creating resources, without the overhead of working through huge full datasets.
The pilot: targeted language varieties
For the first iteration, we used search queries from Mozilla Data Collective to guide our decisions. We looked at which language varieties were most frequently requested, and created datasets for them.
The initial set includes the most searched-for varieties of:
- Chinese (Mandarin): Beijing, Shanghai
- Dutch: Belgian Dutch (Flemish) and Netherlands Dutch
- English: British English, Australian English, Canadian English, Irish English, Malaysian English, South Asian English, American English and Southern American English
- French: Metropolitan French and Canadian French
- Portuguese: Brazilian Portuguese and Portuguese of Portugal
- Spanish: Mexican Spanish, Caribbean Spanish, and Rioplatense
For Arabic and three of the other varieties (Mexican Spanish, American English and Netherlands Dutch) we have also included two datasets split by gender of the speaker to facilitate work on debiasing.
Where to find the datasets
You can now find these dataset segments live on the site:
- Common Voice Scripted Speech 26.0 - Arabic (Female)
- Common Voice Scripted Speech 26.0 - Arabic (Male)
- Common Voice Scripted Speech 26.0 - Beijing Chinese
- Common Voice Scripted Speech 26.0 - Brazilian Portuguese
- Common Voice Scripted Speech 26.0 - Flemish Dutch
- Common Voice Scripted Speech 26.0 - Netherlands Dutch
- Common Voice Scripted Speech 26.0 - Netherlands Dutch (Female)
- Common Voice Scripted Speech 26.0 - Netherlands Dutch (Male)
- Common Voice Scripted Speech 26.0 - Wu Chinese
- Common Voice Scripted Speech 26.0 - American English (Female)
- Common Voice Scripted Speech 26.0 - American English (Male)
- Common Voice Scripted Speech 26.0 - Australian English
- Common Voice Scripted Speech 26.0 - British English
- Common Voice Scripted Speech 26.0 - Canadian French
- Common Voice Scripted Speech 26.0 - Caribbean Spanish
- Common Voice Scripted Speech 26.0 - French of France
- Common Voice Scripted Speech 26.0 - Irish English
- Common Voice Scripted Speech 26.0 - Malaysian English
- Common Voice Scripted Speech 26.0 - Mexican Spanish (Female)
- Common Voice Scripted Speech 26.0 - Mexican Spanish (Male)
- Common Voice Scripted Speech 26.0 - Rioplatense Spanish
- Common Voice Scripted Speech 26.0 - Scottish English
- Common Voice Scripted Speech 26.0 - South Asian English (India, Pakistan, Sri Lanka)
- Common Voice Scripted Speech 26.0 - Southern American English
Can't find what you're looking for or need a specific segment? Get in touch with us to let us know!