Common Voice Scripted Speech v27 and Spontaneous Speech v5 are here
Scripted Speech v27 and Spontaneous Speech v5 are now live on the Mozilla Data Collective platform.
Scripted Speech v27 and Spontaneous Speech v5 are now live on the Mozilla Data Collective platform.
Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts This competition evaluates performance on real-world bilingual dialogues in Indonesian-Javanese, Nahuatl-Spanish, and Spanish-English. For speech technologies to truly have an impact on the world, they must adapt to how people actually communicate in digitally-mediated settings.
Helping organisations participate more directly in the AI economy while making it easier for AI builders to discover responsibly sourced datasets. Earlier this month, we shared a preview of Compensated Datasets and our vision for creating more transparent ways for organisations to participate in the AI economy while retaining agency
Mozilla Common Voice is a massively multilingual platform for collecting speech data to train automatic speech recognition (ASR). Its mission is simple: to make language technology understand everyone’s mother tongue. But for datasets to be genuinely useful, they also need to be manageable. Many of the larger Common Voice
New collaboration expands visibility for community-governed language datasets and improves exploration of linguistic resources, services and tools. Mozilla Data Collective datasets are now discoverable through CLARIN’s Virtual Language Observatory, making it easier for researchers, developers and language technology practitioners in Europe to find multilingual and community-centered datasets
Mozilla Data Collective was built to redefine how AI data is created, shared, and governed. As part of our mission to be the data sharing platform for human agency and fair value exchange, we have long-teased what so many partners and community members have requested: a tangible way to
Why fine-tuning Whisper is a dataset problem OpenAI's Whisper changed what's possible in speech recognition: a single multilingual model with strong zero-shot performance across dozens of languages. But "dozens" is the catch. But for the long tail of the world's
For more than 75 years, Radio Free Europe/Radio Liberty (RFE/RL) has promoted democratic values by providing accurate, uncensored news and debate in countries where a free press is threatened. RFE/RL reaches more than 44 million people every week across 18 countries, in 24 languages, including Persian, Russian,
In this post, we walk you through how to create a useful dataset sample as a preview of your dataset, and guide you in uploading it to the MDC platform.
The problem with "low-resource" machine translation Most production machine-translation systems in 2026 are still trained on a fairly narrow set of language pairs: the 50 or so for which the open web supplies enough parallel text to push BLEU scores into useful territory. Below that line,
A curated list of 15 text-to-speech training datasets for teams shipping production voice models in 2026 covering emotional, multi-speaker, audiobook-derived, non-Latin script, indigenous-language datasets and more.
Most voice assistants listen and respond in a handful of languages. Try to build one for your home that speaks your language, though, and you quickly run into a wall: the training data does not exist, or it is locked behind licences that make it unusable for open source projects.