Common Voice
Pashto becomes third-highest language by volume of data in Common Voice v24
Key highlights from the Common Voice v24 Scripted Speech and v2 Spontaneous Speech release.
Common Voice
Key highlights from the Common Voice v24 Scripted Speech and v2 Spontaneous Speech release.
FAQ
Building with MDC datasets, uploading and sharing data, or joining our community are all ways to get involved.
FAQ
No, you do not need to be a part of an organization to upload a dataset to Mozilla Data Collective.
Common Voice
Firstly, we’d like to thank you for your patience. After introducing Spontaneous Speech early in 2025, we released most locale datasets when the Mozilla Data Collective platform launched in alpha in September of this year. However, upon inspection, the English Spontaneous Speech dataset required some remedial work prior to
Docs
Interested in joining the movement and publishing your dataset on Mozilla Data Collective? This guide will walk you through the steps required, from account creation to submission!
FAQ
When you exclusively host your dataset on MDC, you choose to only make it available on our site.
FAQ
Our priority is technology that is more multilingual, multicultural, and multi-modal.
FAQ
When you use MDC to share or download datasets, you enter into an agreement directly with the data provider/consumer.
News
Nahuatl, Bahasa Indonesia and Bulgarian all feature in our very first community curated datasets to be uploaded to the Mozilla Data Collective platform.
FAQ
We limit access to old versions of Common Voice to respect those speakers who have withdrawn their consent to be included in the dataset.
Common Voice
We’re tightening the circulation of old datasets to protect contributors, while keeping a clear, documented path for researchers who need them.
FAQ
Depending on the nature of the issue, you can contact the dataset uploader, use the report dataset link, or contact us directly.