On contributing my Thorsten-Voice voice datasets
Thorsten Müller has created five TTS voice datasets totaling 40 hours of German speech data. In this community-authored post, he speaks to the importance of sharing his voice.
Thorsten Müller has created five TTS voice datasets totaling 40 hours of German speech data. In this community-authored post, he speaks to the importance of sharing his voice.
Thorsten Müller hat fünf TTS-Stimmdatensätze mit insgesamt 40 Stunden an deutschen Sprachdaten erstellt. In diesem von der Community verfassten Beitrag spricht er über die Bedeutung, seine Stimme zu teilen.
By Kathy Reid · Mozilla Data Collective Ask a Queenslander to say "no worries" into a Home Assistant Voice Preview. There's a decent chance it mishears them. Ask someone from Radelaide and things get worse. Ask anyone who grew up calling things "heaps good" and
Panjebar Semangat, a weekly Javanese-language magazine established before Indonesian independence, is collaborating with Mozilla Data Collective to advance community-governed language dataset frameworks.
Mozilla Data Collective is building towards a multicultural, multilingual, and multimodal future that works for all of us. And over the past few months, we’ve listened as people have flagged what kinds of datasets they need, but are struggling to find. So we’re pleased to announce that MDC
This week: 19 new datasets and a few small changes while we're heads down in some exciting new features that will be coming soon...
We get a lot of questions about how to approach licensing your data for AI training. So to help you share your datasets, we’ve compiled some guidance here – it’s intended to be a living document, that we iterate with our partners and communities. Explore Mozilla Data Collective What
Overcoming the complexity of AI Mozilla Data Collective helps communities to offer unique, multilingual, multicultural, and multimodal datasets. From transcribed and translated videos of narrated Ekpeye folktales to complex question-answering text pairs for the Georgian language, the diversity of datasets on our platform is core to our mission. But with
The institutions that safeguard humanity's cultural memory, galleries, libraries, archives, and museums (collectively known as the GLAM sector) are confronting a paradox that defines the current moment in AI development. Years of careful digitization of their archives have transformed physical collections into vast, machine-readable repositories of human knowledge.
In this guide, you will learn how to use the MDC Python SDK Library to download datasets from the Mozilla Data Collective website.
This week: dataset filtering, enhanced uploader request flow, API improvements, and 20 new datasets!
This week: new features for uploaders, updates to dataset search, and new datasets on MDC!