Common Voice segments now available through Mozilla Data Collective

Share

Mozilla Common Voice is a massively multilingual platform for collecting speech data to train automatic speech recognition (ASR). Its mission is simple: to make language technology understand everyone’s mother tongue.

But for datasets to be genuinely useful, they also need to be manageable. Many of the larger Common Voice datasets can be tens of gigabytes in size, which can make processing them difficult for builders and researchers who have limited technical resources.

One of the most consistent requests we’ve heard over the years is that, for many of the world’s larger languages, people want access to only the portion that’s relevant to them. Whether you’re building a model for a specific region or creating resources for a particular community, it shouldn’t require downloading and handling far more data than you actually need.

With that in mind, we’ve started a pilot programme to ship segments of Common Voice, so you can get the data you need for training models or creating resources, without the overhead of working through huge full datasets.

The pilot: targeted language varieties

For the first iteration, we used search queries from Mozilla Data Collective to guide our decisions. We looked at which language varieties were most frequently requested, and created datasets for them.

The initial set includes the most searched-for varieties of:

  • Chinese (Mandarin): Beijing, Shanghai
  • Dutch: Belgian Dutch (Flemish) and Netherlands Dutch
  • English: British English, Australian English, Canadian English, Irish English, Malaysian English, South Asian English, American English and Southern American English
  • French: Metropolitan French and Canadian French
  • Portuguese: Brazilian Portuguese and Portuguese of Portugal
  • Spanish: Mexican Spanish, Caribbean Spanish, and Rioplatense

For Arabic and three of the other varieties (Mexican Spanish, American English and Netherlands Dutch) we have also included two datasets split by gender of the speaker to facilitate work on debiasing. 

Where to find the datasets

You can now find these dataset segments live on the site: 

Can't find what you're looking for or need a specific segment? Get in touch with us to let us know!

Read more

Mozilla Data Collective datasets now discoverable through CLARIN’s Virtual Language Observatory

New collaboration expands visibility for community-governed language datasets and improves exploration of linguistic resources, services and tools. Mozilla Data Collective datasets are now discoverable through CLARIN’s Virtual Language Observatory, making it easier for researchers, developers and language technology practitioners in Europe to find multilingual and community-centered datasets