Common Voice Scripted Speech v27 and Spontaneous Speech v5 are here

Scripted Speech v27 and Spontaneous Speech v5 are now live on the Mozilla Data Collective platform.

Share
Common Voice v27 now available

There's a particular kind of satisfaction in watching a language dataset grow, knowing that thousands of individual speakers made the choice to donate their voice, time and time again.

"Tunahisi fahari!" you might say, in Kiswahili, or "Kagum!" in Indonesian.

Scripted Speech v27 and Spontaneous Speech v5 are now live on the Mozilla Data Collective platform, but the numbers only tell part of the story.

The numbers, briefly

Scripted Speech v27.0 now spans 295 languages, comprising 32,349,999 voice clips — a little over 42,593 hours of read speech. Spontaneous Speech v5.0, which is designed to capture conversational, naturalistic speech, now covers 80 languages. It provides 89,754 clips of free-form answers to open questions, of which 302 hours have been transcribed and validated.

It's worth unpacking how Scripted Speech and Spontaneous Speech differ and how they solve different speech technology challenges. Scripted Speech — read from a sentence prompt — provides clean, bounded, phonetically-controlled utterances: excellent for building acoustic models and for language and variety coverage at scale. But we don't always speak that way in real life settings. Spontaneous Speech, by contrast, captures disfluency, hesitation, variable pacing and the general messiness of unscripted speech.

If you've ever wondered why an ASR system trained purely on read speech falls over the moment someone starts speaking naturally — trailing off, self-correcting, restarting a clause — this is why. Spontaneous Speech exists precisely to close that generalisation gap, and is well suited to conversational, real-life speech tasks.

Welcoming new languages to Common Voice

Since the last release, Pa'O (blk) has joined Scripted Speech, and Spontaneous Speech has welcomed six new languages: Chinese (China, zh-CN), Swahili (sw), Palauan (pau), Sundanese (su), Bengali (bn) and Shan (shn).

It's worth noting the work required to onboard a new language to Common Voice. To get to the point of having speech data available in a new language, the Common Voice interface has to be translated through the Pontoon system, and a corpus of copyright-free, linguistically representative sentences must be created for Scripted Speech, and for Spontaneous Speech, a set of prompts calibrated to elicit natural, sustained speech. Six languages joining Spontaneous Speech in a single release cycle is a significant achievement indeed.

In particular we'd like to give a shout out to APNIC Foundation, who facilitated Sundanese on-boarding to Common Voice as part of their broader efforts supporting inclusive and accessible internet in South-East Asia. Thank you, APNIC Foundation!

Community spotlights

While the aggregate numbers are impressive, there are some particular language communities we'd like to highlight for their remarkable achievements.

Laz (lzz) is, by growth rate, the standout of this release — over 67% growth in Scripted Speech since the last release, more than 23,000 new clips yielding 34.3 additional hours of speech. For a language with a comparatively small and geographically concentrated speaker base (Laz is spoken primarily along Turkey's eastern Black Sea coast, with UNESCO classifying it as endangered), that's not incremental progress, it's a step-change driven by concentrated, deliberate community mobilisation.

İsmail Avcı Bucaklişi, President of the Laz Institute, expresses it beautifully:

"For us, Common Voice is not simply a platform for teaching AI to understand Laz. Perhaps for the first time, so many members of the Laz community have come together to write, read and record their own language."

That framing — dataset-as-preservation-infrastructure rather than dataset-as-training-fuel — resonates deeply with us here at Mozilla Data Collective. Language encodes culture and community.

Sindhi (sd) grew 27% — more than 14,000 clips representing 15.6 hours — from a contributor base of just over 30 people, in three months. That's a striking clips-per-contributor ratio and showcases the dedication of the Sindhi language community.

Upper Sorbian (hsb), with around 40 contributors, added just over 1,000 clips (2.2 hours) — a 17% increase. Upper Sorbian is a Slavic minority language spoken by a shrinking population in eastern Germany, and a dataset growing at this rate, from a small contributor pool, is exactly the kind of long-tail, low-resource progress that the whole architecture of Common Voice is designed to make possible in a way that centralised data collection efforts structurally cannot.

And Hausa (hau) got a quieter but no less consequential upgrade: dialectal variant support, with Hausa speakers now able to select which variant they speak in their profile, thanks in part to Community Facilitator Alkasim Yusuf Musa. Thank you Alkasim! This is the sort of work that materially improves downstream model fairness — a model trained without variant labelling will silently average across dialectal variation in ways that degrade performance for whichever variant is under-represented in the mix. Labelling the variant doesn't fix that on its own, but allows model trainers to curate the data used for training, making models perform better across variants.

Visualising Common Voice growth over time

As a little treat, we've visualised the growth of Common Voice Scripted Speech, from its first release in 2019 with only a handful of languages, right up to languages like Sundanese added just last week.

0:00
/0:22

Minutes of scripted speech added to Common Voice per release by language

Why this keeps mattering

The value of Common Voice isn't just that it's a large speech corpus, that's expanded due to the efforts of language communities across the globe across the last seven years.

It's a large speech corpus, with consented, known provenance, and it's released under a permissive licence. Moreover, the language speakers contributing voice samples know exactly what they're contributing to and why. That combination is rarer than it should be in speech data, and it's precisely the property that lets a community like the Laz decide, correctly, that they're not just feeding a model — they're building an artefact of cultural continuity that happens to also be machine-readable.

If your language isn't yet represented, consider contributing your voice at Common Voice. And if you're working with the data itself, both releases are available now on the Mozilla Data Collective.

To every contributor, validator and language community facilitator who made this release possible — thank you. Asante. Terima kasih. Dźakuju so. Didi mardi. Mun gode.

Read more

Mozilla Data Collective datasets now discoverable through CLARIN’s Virtual Language Observatory

New collaboration expands visibility for community-governed language datasets and improves exploration of linguistic resources, services and tools. Mozilla Data Collective datasets are now discoverable through CLARIN’s Virtual Language Observatory, making it easier for researchers, developers and language technology practitioners in Europe to find multilingual and community-centered datasets