> ## Content Index
> Fetch the complete content index at: https://community.mozilladatacollective.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# 15 Datasets for Building a Production TTS Voice in 2026
- URL: https://community.mozilladatacollective.com/15-datasets-for-building-a-production-tts-voice-in-2026/
- Published: 2026-06-02T13:07:24.000Z
- Updated: 2026-06-02T13:20:55.000Z
- Description: A curated list of 15 text-to-speech training datasets for teams shipping production voice models in 2026 covering emotional, multi-speaker, audiobook-derived, non-Latin script, indigenous-language datasets and more.
- Author: Mozilla Data Collective
- Tags: News

## Why dataset selection matters more than scale

There was a time when the entire conversation about text-to-speech could be reduced to a single question: how many hours of clean single-speaker audio do you have? The standard answer was twenty-four hours of [LJSpeech-style](https://mozilladatacollective.com/datasets/cmonjhxee01ako007kohpbg34?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) read audio and from that you got a model that sounded acceptable but flat.

That era is over. The teams shipping production TTS now have moved past "more hours of one speaker." They train on deliberate mixtures: emotional speech for prosody, multi-speaker corpora for speaker generalisation, audiobook-derived data for narrative style, dialect-specific data for regional authenticity, non-Latin script data for under-represented writing systems, indigenous-language data for cultural inclusion. The model's quality is the mixture's quality, and the mixture's quality is a function of which specific datasets you put into it.

In this article is a list of fifteen specific datasets from [Mozilla Data Collectiv](https://mozilladatacollective.com/?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post)e that we think belong in a serious TTS training stack in 2026\. Some are large and provide an acoustic baseline. Some are small and provide a specific signal, an emotion, a script, a voice profile, that nothing else covers. All of them are downloadable today, all of them have clear licensing, and all of them are on Mozilla Data Collective.

## The 15 datasets

- [**Thorsten-Voice Dataset 2021.06 Emotional**](https://mozilladatacollective.com/datasets/cmm4b8f7700mgmh07cha3549n?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Community | Licence: CC0-1.0 | Size: 380.80 MB | Task: TTS | Format: WAV, CSV
- [**Urdu Multi-Speaker TTS Dataset**](https://mozilladatacollective.com/datasets/cmmvykcrs0050ny07vkwww5gi?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Community | Licence: CC-BY-NC-4.0 | Size: 514.54 MB | Task: TTS | Format: WEBM, TSV
- [**LibriVox Italian TTS Female Voice**](https://mozilladatacollective.com/datasets/cmo0qtw7v003knt07u6yupncc?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: MDC Curators | Licence: CC0-1.0 | Size: 61.74 MB | Task: TTS | Format: MP3, TSV
- [**LibriVox Czech TTS Female Voice**](https://mozilladatacollective.com/datasets/cmo0jfvnw00p1nx070preklt5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: MDC Curators | Licence: CC0-1.0 | Size: 178.58 MB | Task: TTS | Format: MP3, TXT, TSV
- [**Kokoro Speech Dataset**](https://mozilladatacollective.com/datasets/cmmknsho4014wmf087kvq5rc6?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Community | Licence: LibriVox Public Domain | Size: 3.98 GB | Task: TTS | Format: FLAC
- [**Yoruba-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmo1nlaah0071mk077mw0qhpv?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 319.05 MB | Task: TTS | Format: MP3, TSV
- [**Hausa-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmnopto3q00t0mf07v2dtc0ej?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post)Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 276.90 MB | Task: TTS | Format: MP3, TSV
- [**isiXhosa-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmo4gtixz00kwny07hayfsk8s?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 276.02 MB | Task: TTS | Format: MP3, TSV
- [**Tiv-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmo4nmfam00nxny07rssox2tj?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 311.58 MB | Task: TTS | Format: MP3, TSV
- [**Duala-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmpmpf0jw021bnu0743hu7763?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 141.26 MB | Task: TTS | Format: MP3, TSV
- [**Bamun-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmnhjbnjp0115mh07kiha0rei?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 219.97 MB | Task: TTS | Format: MP3, TSV
- [**Saraiki 10 Hours TTS Dataset**](https://mozilladatacollective.com/datasets/cmnggqr8z0082mh07vbsbm6t5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: MirasAI | Licence: CC-BY-NC-SA-4.0 | Size: 584.44 MB | Task: TTS | Format: WEBM, TSV
- [**Chuvash TTS**](https://mozilladatacollective.com/datasets/cmnhhi0by00zknn07edrnd82e?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Taruen | Licence: CC-BY-SA-4.0 | Size: 854.02 MB | Task: TTS | Format: PARQUET
- [**Otomí (Hñähñu) TTS Voz Masculina**](https://mozilladatacollective.com/datasets/cmo0cro1g00hlmr07oichasyk?ref=community.mozilladatacollective.com) Contributor: Community | Licence: CC-BY-SA-4.0 | Size: 119.54 MB | Task: TTS | Format: MP3, TXT, TSV
- [**Central Kurdish TTS dataset 1.0**](https://mozilladatacollective.com/datasets/cmj77njd701ljmb07m97pw1p3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: The University of Melbourne | Licence: CC-BY-4.0 | Size: 293.45 MB | Task: TTS | Format: WAV

## Not Scale but Curation

Production TTS in 2026 is no longer a scale problem but a curation problem. The fifteen datasets above won't, on their own, train a model but what they will do is let you build a deliberately diverse training mix that covers prosody, speaker variation, scripts, dialects, and the under-represented languages most commercial TTS still ignores. Mozilla Data Collective exists precisely to provide a platform for that diversity and unlock datasets from around the world that are more multicultural and multilingual.

[Browse all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post)

[Get in touch →](mailto:support@mozilladatacollective.com)

[Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com)