# Mozilla Data Collective > Multilingual, multicultural, multimodal datasets under range of licenses: from open source, to community controlled, to commercial and compensated. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### About Mozilla Data Collective URL: https://community.mozilladatacollective.com/about/ Last updated: 2026-06-23T09:00:09.000Z **What is Mozilla Data Collective’s mission?** We fight for a tech future that is **multilingual, multicultural and multimoda**l. We think we deserve a world where you don’t need to change how you speak or how you look in order to access technology - and we welcome technology’s promise as a connector, enabler; a tool to build and make and shape. We think the right way to get there is by giving everyone, everywhere **the data platform for human agency and fair value exchange**. People should be able to choose where their datasets show up, and they should be able to define what it looks like to benefit; whether that’s swapping data for tool access, or for expertise, donating it to the public, or asking for fair compensation. You can share openly, using existing licenses like Creative Commons, or you can build your own license. You can open up your datasets for everyone, or just for some types of downloaders, you can set custom constraints, ask for exchange, compensation or recognition. You can govern the dataset as an individual, a co-operative, a trust or something else. After all, it’s your dataset. The people who access your datasets are fully authenticated, and held in legally binding contracts, and we have a number of dataset protection features. **What is the history of Mozilla Data Collective?** In 2025, our Founder and CEO E.M. Lewis-Jong, was leading Common Voice (the world’s largest public participation speech dataset) at Mozilla Foundation, and was looking for a release platform that would give Common Voice communities more choice: choice of license, features that undergirded stronger control, and a radically anti-extractivist form of value exchange. The team couldn’t find that platform, so in September we built Mozilla Data Collective. Common Voice was Mozilla Data Collective’s first community user, piloting the platform for its own datasets. Along the road, we had met hundreds of fellow travellers with the same problem; trying to share their data on their own terms, in line with their own values. So in November, we opened up Mozilla Data Collective to close friends, partners and allies. As of April 2026, Mozilla Data Collective has 187 organisations vetted to share datasets on the platform. We work with organisations from libraries, archives and museums, to tech start ups in health, education and language tech. In April 2026, we spun out a UK entity dedicated to housing that work - Mozilla Data Collective. Common Voice continues to be stewarded by Mozilla Foundation. **What is the business model of Mozilla Data Collective?** Mozilla Data Collective is structured as a British company, incubated by Mozilla Foundation, and now a subsidiary of the not-for-profit Mozilla.org. This gives it the flexibility of a company, with the mission lock of a non-profit. In order to facilitate the Mozilla Data Collective community and organisations seeking to set their own value exchange options - including compensation - we needed to be able to transact payments, and provide support services. We chose to establish this organisation in Europe, where data protection is gold standard. Our uploaders keep what they choose to charge (if they choose to charge - most of our datasets are open source). Our business model is to charge the downloader a modest 5% fee, which covers the costs of our storage, infra, maintaining APIs. We will launch a premium subscription later in the year with more sophisticated tools to discover, curate and package datasets, which will also be progressively priced depending on organisation size, with free options for many. That’s it! That’s the model. Being a mission-driven company is not just a necessity, it’s a conscious choice. We are a social enterprise, and proud to be so. In a world where grant funding can disappear in line with political realities, we want to be firmly self-sustaining and independent. We don’t want to cede the multilingual data space (in which we’ve been working since 2017) to dubious marketplace brokers who are gate-keeping data buyers in order to carve out hefty cuts for themselves, or to workforce vendors driving precarious gig work at knockdown prices. We think that people deserve a universe of fair data exchange, of collective bargaining, where they’re in control. Our communities don’t give up their datasets, they invest them as a lever for change. Mozilla Data Collective is structured to give them that platform. **Why did Mozilla decide to invest philanthropic capital in the data space?** Mozilla Foundation's job is to spot the structural fights that will define whether technology serves everyone or just the few. We've been in the data space since 2017 with Common Voice, long before "AI" became a buzzword in every deck. And that’s why we know that for too long, the data economy has looked like a digital land grab: extractive, opaque, and frankly, lazy. We saw a market failure, where the people who actually create the value of AI were being treated as a resource to be mined rather than partners to be respected. We’re in the innovation business, not human fracking, so we wanted to resource an alternative. That became ever more urgent as we saw that whoever controls the data layer controls the AI future. Mozilla Data Collective is our bet that you can build sustainable infrastructure for fair data value exchange, in a way that maximises innovation, no matter your geography or industry proximity. That was a problem worth solving with patient, mission-locked capital – that is, exactly the kind of bet that philanthropy should be making. Mozilla Data Collective is a platform in the truest sense. It’s yours to stand on, and make of it what you will. Mozilla Data Collective works by allowing you to share your data, retain ownership of it, and control who uses it. We imagine and create a better future where AI is built equitably and powered by the people. We do this by providing alternative solutions that challenge extractive data practices by placing the power of how AI data is created and governed in the hands of the people. [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) --- ### Find Us Around the Web: [r/MozillaDataCollective - Reddit](https://www.reddit.com/r/MozillaDataCollective/?ref=community.mozilladatacollective.com) [Mozilla Data Collective | LinkedIn](https://www.linkedin.com/company/mozilla-data-collective?ref=community.mozilladatacollective.com) [Mozilla Data Collective - Discord](https://discord.gg/cs9tJPqQB5?ref=community.mozilladatacollective.com) ### Join Mozilla Data Collective URL: https://community.mozilladatacollective.com/join/ Last updated: 2026-05-19T17:18:22.000Z ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/10/Screenshot-2025-10-07-at-1.08.35---PM.png) Mozilla Data Collective is a platform in the truest sense. It’s yours to stand on, and make of it what you will. We have dual roots in two Mozilla projects - Common Voice, a CC0 public dataset to help tech speak your language - and the Data Futures Lab - an experimental space for instigating new approaches to data stewardship challenges. Mozilla Data Collective works by allowing you to share your data, retain ownership of it, and control who uses it. We partner with organizations and individuals to make their data available through Mozilla Data Collective. You can share openly, using existing licenses like Creative Commons, or you can build your own. You can open up your data for everyone, or just for some types of downloaders, you can set custom constraints, ask for exchange, compensation or recognition. You can govern it as an individual, a co-operative, a trust or something else. After all, it’s your data. [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### Talks URL: https://community.mozilladatacollective.com/dev-talks/ Last updated: 2026-05-19T17:19:04.000Z Learn about Mozilla Data Collective from talks that our team has given and join the movement to reclaim your data. - ➜ [Fine-Tuning a Whisper Model with MDC Datasets](#fine-tuning-a-whisper-model-with-mdc-datasets) - ➜ [Beyond Extraction: Building Community-Centered Speech Data](#beyond-extraction-building-community-centered-speech-data) - ➜ [Your datasets, under your control: Introducing Mozilla Data Collective](#your-datasets-under-your-control-introducing-mozilla-data-collective) [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ## Fine-Tuning a Whisper Model with MDC Datasets Most speech recognition models were built with English, or a handful of well-resourced languages, in mind. If you speak Khmer, Galician, or any of the hundreds of languages underrepresented in mainstream AI, you've probably hit a wall trying to get accurate transcriptions. In [this tutorial](https://community.mozilladatacollective.com/fine-tune-a-speech-to-text-model-for-any-language-including-yours/), we'll walk through how to fine-tune OpenAI's Whisper model on your own language using the Mozilla Data Collective platform's datasets or your own custom audio data. Everything runs locally — even on a laptop — keeping your data private. ## Beyond Extraction: Building Community-Centered Speech Data [![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/10/image.png)](https://opensource.org/ai/webinars/beyond-extraction-building-community-centered-speech-data?ref=community.mozilladatacollective.com) > With the advent of deep learning, speech recognition models like Open AI’s Whisper are now trained on hundreds of thousands of hours of speech data, likely gathered without the consent of the speakers who contributed it. As responsible AI practices grow in prevalence and we continue to advance machine learning-enabled speech technologies, we must define and commit to a set of shared best practices for responsible speech data collection and stewardship. [Watch on opensource.org](https://opensource.org/ai/webinars/beyond-extraction-building-community-centered-speech-data?ref=community.mozilladatacollective.com) --- ## Your datasets, under your control: Introducing Mozilla Data Collective > AI has a data crisis. We're running out of quality training data because the entire web has already been harvested by crawlers to train AI models — leading to the "Token Crisis". What’s left? Synthetic data generated en masse - that’s bland, generic and unrepresentative of the world’s diversity. This data is also problematic for training models, as it can lead to model collapse. Meanwhile, quality datasets from diverse contributors sit unused in silos. Our vision is to encourage the creation of safe, responsible AI that works for *everyone* \- by helping communities to share authentic, ethical and diverse data - a stark contrast to models built by indiscriminately scraping the web and reproducing or synthesising its Anglocentric, white, male biases. [Watch on YouTube](https://www.youtube.com/watch?v=rl7QvFqjXFA&ref=community.mozilladatacollective.com) ### Careers at Mozilla Data Collective URL: https://community.mozilladatacollective.com/careers-at-mozilla-data-collective/ Last updated: 2026-08-20T16:22:13.000Z ## ## About Mozilla Data Collective [**Mozilla Data Collective**](https://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com)is the data sharing platform for human agency and fair value exchange. Our vision is a tech future that is multilingual, multicultural, and multimodal; our mission is to achieve that future by redefining how AI data is created, shared, and governed. Uploaders own their datasets, set terms of use, and decide who benefits from their data. We are an ambitious, growing social enterprise based in the United Kingdom and led by a team of anti-extractivist, pro-people technologists whose expertise spans computational linguistics, AI/ML engineering, and data sovereignty. Fundamentally, we believe AI can be all it promises to be - not all it threatens to be. ## Working at Mozilla Data Collective At Mozilla Data Collective, we’re dedicated to building a vibrant, remote-first environment where you can flourish both professionally and personally. Our benefits are designed to support your long-term well-being, offering competitive health and retirement plans tailored to your location, generous parental leave, and a structured approach to rest with dedicated refresh days and a company-wide December break. We also empower your growth and balance through an annual learning and professional development budget, an annual wellness budget, and a flexible work policy, because we believe that when you are supported to live a flourishing life, our user community also thrives. **OPEN ROLES** [Data Sales Manager](https://drive.google.com/file/d/12wijvt6lzfdwx7k9UiJNDiuLcQim2a8x/view?usp=sharing&ref=community.mozilladatacollective.com) \- in process [Senior Fullstack Software Engineer](https://drive.google.com/file/d/1o4tATMMEzOeroW2bRSYJzYgr8h4BJIyy/view?usp=sharing&ref=community.mozilladatacollective.com) \- in process **COMING SOON** Senior Product Manager (Parental Leave) We're growing fast! Don't see the perfect role, but think you embody our core values of 80/20 thinking, focus and ownership? Send us your CV at [careers@mozilladatacollective.com](mailto:careers@mozilladatacollective.com) ### MDC Data Licence Agreement 1.0 URL: https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0/ Last updated: 2026-07-07T18:03:03.000Z Last Updated: 7 July 2026 This Data Licence Agreement for MDC Datasets (this “**Licence Agreement**”) is entered into by and between Mozilla Data Collective, LTD, a company registered in England and Wales with company number 17054959, whose registered address is located at 167-168 Great Portland Street, London, W1W 5PF (the “**Licensor**” or “**MDC**”), and the individual or entity accepting this License Agreement ("**Licensee**"). By clicking "I Agree" (or other similar assent), downloading, accessing, or using the Licensed Data, Licensee agrees to be bound by the terms of this Licence Agreement and any exhibits hereto . Licensor and Licensee are referred to herein, collectively, as the “**Parties**” and, individually, as each “**Party**.” **WHEREAS**, Licensor has commissioned the creation of certain proprietary datasets and related data compilations; **WHEREAS**, Licensee desires to obtain a licence to access and use certain datasets made available by Licensor; and **WHEREAS**, the parties wish to establish the terms and conditions governing Licensee’s access to and use of such commissioned datasets. **NOW, THEREFORE**, in consideration of the mutual promises, covenants, representations and warranties contained in this License Agreement, and for other good and valuable consideration, the receipt and sufficiency of which are hereby acknowledged, the Parties agree as follows: 1. Definitions.For purposes of this License Agreement, the following terms will have the indicated meanings: 1. “**Applicable Laws**” means all applicable laws, statutes, regulations, rules, ordinances, and other legally-binding requirements of any governmental authority having jurisdiction over a Party or the subject matter of this Licence Agreement. 2. “**Derived Data**” means data created or derived by or on behalf of Licensee as a result of combining, changing, converting, analysing, aggregating, transforming, or otherwise processing Licensed Data with other data, where the resultant data is not reasonably capable of being used, whether directly or indirectly and using any methods, techniques, technologies, or tools now known or later developed, to reconstruct, reproduce, extract, infer, or regenerate Licensed Data. 3. **“Licensed Data”** means all data, datasets, and data compilations licensed pursuant to this Licence Agreement as the same may be further defined in a dataset listing. 4. “**Marks**” means, with respect to a Party, such Party’s trade names, trade dress, trademarks, service marks, logos, brand names and other identifiers, corporate names, meta-tags and universal resource locators, and any applications, registrations and renewals thereof. 5. **“Trained Models and Outputs**” means any machine learning models, artificial intelligence systems, model weights, algorithms, insights, analyses, predictions, reports, results, or other outputs created, developed, generated, trained, fine-tuned, testing, validated, benchmarked, or improved by or on behalf of Licensee through the authorised use of Licensed Data, provided that such Trained Models and Outputs do not disclose, contain, or permit a third party to readily reverse engineer, reconstruct, access, or discern the Licensed Data. 2. Licence Grant. Subject to the terms and conditions of this Licence Agreement, Licensor hereby grants to Licensee a non-exclusive, worldwide, non-transferable, non-sublicensable licence during the Term to access, copy, store, reproduce, use, analyse, and otherwise process the Licensed Data: (a) for Licensee’s internal and commercial business purposes; (b) to develop, train, fine-tune, test, validate, benchmark, and improve artificial intelligence, machine learning, analytics, and related models, systems, and technologies; (c) to create Derived Data; and (d) to create Trained Models and Outputs. Licensee may use, commercialise, distribute, and otherwise exploit any models, insights, analyses, predictions, or outputs generated through Licensee’s authorised use of the Licensed Data, provided that such activities do not result in the disclosure, redistribution, or other unauthorised use of the Licensed Data. For clarity, the Licensed Data shall not include Derived Data or Trained Models and Outputs. 3. Restrictions on Use. Licensee shall not, and shall not permit any third party to: (a) reverse engineer, disassemble, decompile, reconstruct, extract, or otherwise attempt to discover or derive the underlying data elements, source materials, methodology, composition, or contents of the Licensed Data, except as reasonably necessary to access and use the Licensed Data as expressly permitted under this License Agreement; (ii) sell, license, sublicense, distribute, publish, disclose, make available or otherwise provided the Licensed Data, in whole or in part, to any third party except as expressly permitted under this License Agreement; (iii) create, commercialise, distribute, or otherwise make available, any dataset, database, or data product that contains, incorporates, discloses, reproduces, or serves as a substitute for the Licensed Data or any substantial portion thereof; (iv) use the Licensed Data to develop, train, fine-tune, test, validate, benchmark, or improve any artificial intelligence model, machine learning model, product, or service that is designed or intended primarily to reproduce, reconstruct, extract, generate, or otherwise make available the Licensed Data, or any substantial portion thereof; (v) use the Licensed Data to create, develop, market, or offer any dataset, database, data product, or service that competes with, is substantially similar to, or is intended to serve as a substitute for Licensed Data; (vi) remove, alter, obscure, or destroy any copyright notices, Licensor Marks, proprietary legends, attribution requirements, or other notices contained in or accompanying the Licensed Data; (vii) attempt to identify, contact, or otherwise associate any individual with any data contained in the Licensed Data, nor attempt to re-identify any de-identified, anonymised, pseudonymised, or aggregated information; or (viii) use the Licensed Data in any manner that violates, or could reasonably cause Licensor to violate, any Applicable Laws or any third-party rights. 4. **Intellectual Property; Ownership.** 1. Licensed Data. As between the Parties, The Licensed Data, and all right, title, and interest in and to, including all intellectual property rights therein, are and shall remain the sole and exclusive property of Licensor (and its licensors, as applicable). Except for the limited license rights expressly granted under this Licence Agreement, no right, title or interest in or to the Licensed Data is transferred to Licensee. For clarity, the Licensed Data shall not include Derived Data or Trained Models and Outputs. 2. Derived Data; Trained Models and Outputs. As between the Parties, Licensee shall own all right, title and interest in and to the Derived Data and Trained Models and Outputs, subject to Licensor’s ownership of and rights in the Licensed Data and any restrictions set forth in this License Agreement. 5. **Fees and Payment.** In consideration of the licence(s) granted under this Licence Agreement, Licensee shall pay all fees and charges applicable to Licensed Data as specified in the applicable dataset listing. Unless otherwise specified in such documentation, all fees are due upon purchase. All fees are non-cancellable and non-refundable, except as expressly provided in this License Agreement or required by Applicable Law. 6. **Representations and Warranties; Disclaimers.** 1. Mutual Warranties. Each Party hereby represents and warrants to the other Party that such Party’s execution, delivery and performance of this Licence Agreement does not and will not: (i) conflict with or violate any Applicable Laws; or (ii) conflict with or require any consent under any contract, licence, permit, or other agreement to which it is a party. 2. Licensor Warranty. Licensor represents and warrants that it has all rights, licences, permissions and consents necessary to grant Licensee the rights in the Licensed Data as set forth in this Licence Agreement and each dataset listing. 3. Licensee Warranty. Licensee represents and warrants that it shall at all times access, use, process, store, disclose, and otherwise exploit the Licensed Data in compliance with this Agreement and all Applicable Laws. 4. Disclaimer. LICENSEE’S USE OF, OR INABILITY TO USE, THE LICENSED DATA IS AT LICENSEE’S SOLE RISK. THE LICENSED DATA IS PROVIDED ON AN “AS IS” AND “AS AVAILABLE” BASIS, AND LICENSOR HEREBY DISCLAIMS ALL WARRANTIES, WHETHER EXPRESS, IMPLIED, STATUTORY, OR OTHERWISE. LICENSOR SPECIFICALLY DISCLAIMS ALL IMPLIED WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE, AND NON-INFRINGEMENT, AND ALL WARRANTIES ARISING FROM COURSE OF DEALING, USAGE, OR TRADE PRACTICE. LICENSOR MAKES NO WARRANTY OF ANY KIND THAT THE LICENSED DATA, OR RESULTS OF ITS USE, WILL MEET LICENSEE'S OR ANY OTHER PERSON'S REQUIREMENTS, OPERATE WITHOUT INTERRUPTION, ACHIEVE ANY INTENDED RESULT, BE COMPATIBLE OR WORK WITH ANY SOFTWARE, SYSTEM, OR OTHER SERVICES, OR BE SECURE, ACCURATE, COMPLETE, FREE OF HARMFUL CODE, OR ERROR FREE. 7. **Data Privacy.** To the extent that the Licensed Data contains Personal Data (as defined in the DPA), the Parties' processing shall be in accordance with the Data Processing Agreement (the “**DPA**”) attached hereto as Exhibit A, and incorporated by reference herein. 8. **Indemnification.** Licensee shall indemnify and hold harmless, and, at Licensor's option, defend Licensor, its successors and assigns, and each of the respective officers, directors, employees, agents and representatives of the foregoing, from and against any and all claims, damages, assessments, costs, losses and other expenses, including but not limited to reasonable attorneys’ fees and legal costs, in connection with any third party claim, demand, suit, action or other proceeding (each, a “**Claim**”) arising from or relating to Licensee’s (i) use, misuse, or unauthorized use of any Licensed Data, (ii) breach of its representations, warranties or obligations under this License Agreement, (iii) breach of Applicable Laws, and (iv) Derived Data, Trained Models and Outputs, and any products or service developed using the Licensed Data; provided that Licensee may not settle any Claim against Licensor unless such settlement completely and forever releases Licensor from all liability with respect to such Claim or unless Licensor consents to such settlement, and further provided that Licensor shall have the right, at its option, to defend itself against any such Claim or to participate in the defense thereof by counsel of its own choice. 9. **Limitation of Liability.** 1. TO THE MAXIMUM EXTENT PERMITTED BY APPLICABLE LAW, IN NO EVENT WILL LICENSOR, ITS AFFILIATES, AGENTS AND/OR EMPLOYEES BE RESPONSIBLE OR LIABLE FOR ANY INDIRECT, PUNITIVE, INCIDENTAL, SPECIAL, OR CONSEQUENTIAL LOSS, CLAIM, INJURY AND/OR DAMAGE ARISING OUT OF, OR IN ANY WAY CONNECTED WITH, THIS AGREEMENT OR YOUR USE OF THE LICENSED DATA, WHETHER BASED ON CONTRACT, TORT, STRICT LIABILITY, OR OTHERWISE, EVEN IF LICENSOR HAS BEEN ADVISED OF THE POSSIBILITY OF ANY LOSS, CLAIM, INJURY AND/OR DAMAGE. 2. TO THE MAXIMUM EXTENT PERMITTED BY APPLICABLE LAW, IN NO EVENT SHALL LICENSOR, ITS AFFILIATES, AGENTS AND/OR EMPLOYEES BE LIABLE TO YOU FOR ANY CLAIMS, LIABILITIES OR DAMAGES HEREUNDER IN AN AMOUNT EXCEEDING THE GREATER OF THE AMOUNT PAID BY YOU TO LICENSOR DURING THE TWELVE (12) MONTHS PRECEDING THE EVENT GIVING RISE TO THE CLAIM (OR THE DURATION OF YOUR USE OF THE LICENSED DATA, IF LESS THAN TWELVE (12) MONTHS), AND ONE HUNDRED GBP (£100). 10. **Term and Termination** 1. Term. This Licence Agreement shall commence on the Effective Date and shall continue until terminated in accordance with this Section 9. 2. Termination. Licensor may terminate this Licence Agreement immediately upon written notice if Licensee materially breaches this Licence Agreement. In addition, Licensor may suspend or terminate Licensee’s access to the Licensed Data if Licensor reasonably believes that Licensee has used, disclosed or distributed the Licensed Data in violation of the terms of this Licence Agreement. 3. Effect of Termination. Upon expiration or termination of this Licence Agreement, all rights granted to Licensee with respect to the Licensed Data shall immediately cease, and Licensee shall promptly discontinue all access to and use of the Licensed Data and permanently delete or destroy all copies of the Licensed Data in its possession or control. Upon Licensor’s reasonable request, Licensee will certify its compliance with the foregoing obligations in writing. Notwithstanding the foregoing, Licensee may retain and continue to use any Derived Data and Trained Models and Outputs created prior to the effective date of termination, provided that such Derived Data and Trained Models and Outputs do not disclose, contain, or permit access to the Licensed Data and Licensee otherwise remains in compliance with this Licence Agreement. 4. Survival. Sections 3, 4, 5, 6, 7, 8, 9 and 10, and any other provisions that by their nature should survive expiration or termination of this Licence Agreement, shall survive such expiration or termination. 11. **Miscellaneous** 1. Governing Law. This Licence Agreement and any dispute or claim arising out of or in connection with it, its subject matter or formation (including non-contractual disputes or claims) shall be governed by and construed in accordance with the laws of England and Wales. Each Party irrevocably submits to the exclusive jurisdiction of the courts of England and Wales in respect of any dispute, claim or proceeding arising out of or in connection with this Licence Agreement. 2. Entire Agreement. This Licence Agreement, and any applicable dataset listing, constitutes the entire agreement between Parties, and supersedes any prior and contemporaneous agreements between the Parties on the subject matter. 3. Relationship of the Parties. Nothing in this Licence Agreement creates any agency, partnership, joint venture, or employment relationship between Licensor and Licensee. 4. Injunctive Relief. Licensee acknowledges that any unauthorised disclosure or use of Licensed Data may cause irreparable harm for which monetary damages are inadequate, and Licensor may seek injunctive or equitable relief without posting bond. 5. Force Majeure. Under no circumstances will Licensor be liable for any delay or failure in performance resulting directly or indirectly from an event beyond its reasonable control. 6. No Waiver. No waiver of any term of this Licence Agreement shall be deemed a further or continuing waiver of such term or any other term, and Licensor’s failure to assert any right or provision under this Licence Agreement shall not constitute a waiver of such right or provision. 7. Severability. Each of the provisions of this Licence Agreement operates separately. If any court or relevant authority decides that any of them are unlawful or unenforceable, the remaining provisions will remain in full force and effect. If any provision is deemed unlawful or unenforceable, the parties agree that such provision shall be modified or amended by the court or relevant authority to the extent necessary to render it enforceable, in accordance with the intent of the original provision. The modified provision shall be interpreted so as to reflect the original intent of the parties as closely as possible, while remaining compliant with Applicable Law. 8. Assignment. This Licence Agreement, and any rights and licences granted hereunder, may not be transferred or assigned by Licensee without the prior written consent of Licensor. This Licence Agreement may be assigned by Licensor without restriction. Any attempted transfer or assignment in violation hereof shall be null and void. **EXHIBIT A** **Data Processing Agreement** This Data Processing Agreement (“**DPA**”) is incorporated into and forms part of (and if applicable, amends the current version of) the Licence Agreement between Licensee and Mozilla Data Collective LTD (“**Licensor**” or **“MDC”**)**,** each a “**Party**” and collectively the “**Parties**”. This DPA applies to and takes precedence over the Agreement between the Parties, and any associated contractual document between the Parties, such as an order form, statement of work, or data processing agreement thereunder , to the extent of any conflict. Capitalised terms not defined herein or in the Agreement, are defined as in applicable Data Protection Laws. Licensee and MDC agree as follows: 1. **Definitions.** For purposes of this DPA: 1. “**Controller**” is defined as under the GDPR and other Data Protection Laws using that term. 2. “**Data Protection Laws**” means all applicable laws, regulations, and other legal or self-regulatory requirements in any jurisdiction relating to privacy, data protection, data security, breach notification, or the Processing of Personal Data, including without limitation, to the extent applicable, the General Data Protection Regulation, Regulation (EU) 2016/679 (“**GDPR**”), the United Kingdom Data Protection Act of 2018 (“**UK Privacy Act**”), and the Swiss Federal Act on Data Protection (“**FADP**”). 3. “**Licensed Data**” shall have the same meaning as under the Agreement. 4. “**Data Subject**” means an identified or identifiable natural person to whom Personal Data relates and includes “consumer” as defined under Data Protection Laws. 5. “**EU SCCs**” means the Standard Contractual Clauses issued pursuant to Commission Implementing Decision (EU) 2021/914 of 4 June 2021 *on standard contractual clauses for the transfer of personal data to third countries pursuant to Regulation (EU) 2016/679 of the European Parliament and of the Council*, available at [http://data.europa.eu/eli/dec\_impl/2021/914/oj](http://data.europa.eu/eli/dec%5Fimpl/2021/914/oj?ref=community.mozilladatacollective.com). and completed as set forth herein. 6. “**Personal Data**” includes “personal data,” “personal information,” and similar terms, as defined by Data Protection Laws, processed by the parties in connection with the Licensed Data under the Agreement. 7. “**Process**”, “**Processing**”, and their cognates mean any operation or set of operations performed on Personal Data or on sets of Personal Data, whether or not by automated means, such as collection, recording, organisation, creating, structuring, storage, adaptation or alteration, retrieval, consultation, use, disclosure by transmission, dissemination or otherwise making available, alignment or combination, restriction, erasure, or destruction. 8. “**Security Breach**” means any accidental or unlawful acquisition, destruction, loss, alteration, unauthorised disclosure of, or access to, Personal Data. 9. “**UK SCCs**” means the United Kingdom International Data Transfer Addendum to the EU Commission Standard Contractual Clauses (available at [https://ico.org.uk/media/for-organisations/documents/4019539/international-data-transfer-DPA.pdf](https://ico.org.uk/media/for-organisations/documents/4019539/international-data-transfer-addendum.pdf?ref=community.mozilladatacollective.com)) and completed as set forth herein. 2. **Roles of the Parties** 1. This DPA applies to Personal Data processed by the Parties in connection with the Agreement. 2. The Parties are independent Controllers of Personal Data processed under the Agreement. Each Party will comply with the requirements of Data Protection Laws applicable to it as a Controller, and each Party is solely responsible for such compliance. 3. **MDC Obligations** 1. MDC will use commercially reasonable efforts to ensure that any disclosure of Personal Data by MDC to Licensee is supported by an appropriate legal basis under applicable Data Protection Laws. 2. During the term of the Agreement, MDC will maintain a publicly available privacy notice describing its processing activities as required by applicable Data Protection Laws. 4. **Licensee Obligations** 1. Licensee will Process Personal Data for the purposes specified in its Privacy Policy and as specified in this DPA, subject to any other requirements or restrictions under Data Protection Laws and the Agreement. Without limiting the foregoing, Licensee is solely responsible for providing all notices and disclosures, and obtaining any consents, required by applicable Data Protection Laws in connection with its Processing of Personal Data. 2. Licensee acknowledges that it is responsible for responding to requests from or on behalf of Data Subjects with respect to Personal Data that it processed as an independent controller, as required by Data Protection Laws. Licensee will promptly forward to MDC any enquiry or request from or on behalf of a Data Subject, or direct the Data Subject to contact MDC directly, relating to any copy of Personal Data that MDC may have. 3. Licensee will take appropriate technical and organisational measures designed to protect Personal Data against a Security Breach and will lawfully respond to and address potential and confirmed Security Breaches. 5. **Data Transfers** 1. The Parties acknowledge that the Processing contemplated under this Agreement may involve the cross-border transfer of Personal Data from MDC to Licensee. A party may only engage in cross-border transfers or onward cross-border transfers of Personal Data if it has put in place a data transfer mechanism deemed to be valid under Data Protection Laws. 2. To the extent legally required, by executing this DPA, Licensee and MDC are deemed to be signing the EU SCCs, which form part of this DPA and (except as described in Section 5(c) and 5(d) below) will be deemed completed as follows: 1. Module One of the EU SCCs applies to transfers of Personal Data from MDC (as an independent Controller) to Licensee (as an independent Controller). 2. Clause 7 (the optional docking clause) is not included. 3. Clause 11 (Redress): The optional language requiring that data subjects be permitted to lodge a complaint with an independent dispute resolution body is not included. 4. Clause 17 (Governing law): The Parties choose Option 1 and select the law of the Republic of Ireland. 5. Clause 18 (Choice of forum and jurisdiction): The Parties select the courts of the Republic of Ireland. 6. Annex I is completed as set forth in Schedule 1 of this DPA. 7. Annex II is completed as set forth in Schedule 2 of this DPA. 3. To the extent legally required, by executing the Agreement, the Parties are deemed to be signing the UK SCCs, which form part of this DPA and take precedence over the rest of this DPA as set forth in the UK SCCs. The Tables in UK SCCs are deemed completed as follows: 1. Table 1: The Parties’ details shall be the Parties and their affiliates to the extent any of them is involved in such transfer, and the Key Contact shall be the contacts set forth in Schedule 1 of this DPA. 2. Table 2: The Approved EU SCCs referenced in Table 2 shall be the EU SCCs as executed by the Parties and completed in Section 5(b) of this DPA. 3. Table 3: Annexes I and II are set forth in Schedules 1 and 2 below, respectively. 4. Table 4: Either Party may end this DPA as set out in Section 19 of the UK SCCs. 4. For transfers of Personal Data that are subject to the FADP, the EU SCCs form part of this DPA as set forth above, but with the following differences to the extent required by the FADP: (1) references to the GDPR in the EU SCCs are to be understood as references to the FADP insofar as the data transfers are subject exclusively to the FADP and not to the GDPR; (2) references to personal data in the EU SCCs also refer to data about identifiable legal entities until the entry into force of revisions to the FADP that eliminate this broader scope; (3) term “member state” in EU SCCs shall not be interpreted in such a way as to exclude data subjects in Switzerland from the possibility of suing for their rights in their place of habitual residence (Switzerland) in accordance with Clause 18(c) of the EU SCCs; and (4) the relevant supervisory authority is the Swiss Federal Data Protection and Information Commissioner (for transfers subject to the FADP and not the GDPR), or both such Commissioner and the supervisory authority identified in the EU SCCs (where the FADP and GDPR apply, respectively). 6. **Additional Safeguards.** To the extent that Licensee acts as the Importer of Personal Data of Data Subjects located in or subject to the Data Protection Laws of the EEA, Switzerland, or the United Kingdom, Licensee agrees to the following safeguards (“**Additional Safeguards**”) to protect such Personal Data to an equivalent level as such Data Protection Laws: 1. Licensee uses encryption for data in transit. 2. As of the date of this DPA, Licensee has not received any national security orders of the type described in Paragraphs 150-202 of the judgment in the EU Court of Justice Case C-311/18, Data Protection Commissioner v Facebook Ireland Limited and Maximillian Schrems. 3. As of the date of this DPA, no court has found Licensee to be the type of entity eligible to receive process issued under FISA Section 702: (i) an “electronic communication service Exporter” within the meaning of 50 U.S.C § 1881(b)(4) or (ii) a member of any of the categories of entities described within that definition. 4. Licensee will not comply with any request under FISA for bulk surveillance, i.e., a surveillance demand whereby a targeted account identifier is not identified via a specific “targeted selector” (an identifier that is unique to the targeted endpoint of communications subject to the surveillance), or take any action pursuant to U.S. Executive Order 12333. 5. Licensee will use reasonably available legal mechanisms to challenge any demands for data access through national security process that it receives, as well as any non-disclosure provisions attached thereto. 6. Licensee will comply with all requirements of Clauses 14 and 15 of the EU SCCs. In particular, to the extent not prohibited by laws applicable to Importer, Importer will promptly notify Exporter (and where relevant, affected Data Subjects) if Importer (i) receives a legally binding request from a public authority for the disclosure of Personal Data provided to Importer by Exporter (including the Personal Data requested, the requesting authority, the legal basis for the request, and Importer’s response to the request), or (ii) becomes aware that public authorities have directly accessed Personal Data provided to Importer by Exporter. Where Importer is legally prohibited from making such notification, it will use its best efforts to obtain a waiver of this prohibition. 7. Licensee will promptly notify Exporter if it can no longer comply with the EU SCCs, UK SCCs, or these Additional Safeguards, without being required to identify the specific provision with which it can no longer comply. 7. **Indemnification and Limitation of Liability.** To the extent permitted by Data Protection Laws, the Parties will indemnify each other, and their liability will be limited, as provided in the Agreement. 8. **Survival.** The provisions of this DPA survive the termination or expiration of the Agreement for so long as Licensee Processes Personal Data transferred to it under the Agreement. 4,135 words[](https://ghost.org/help/using-the-editor/?ref=community.mozilladatacollective.com) ## Posts ### NOW LIVE: Lost in Transcription Competition URL: https://community.mozilladatacollective.com/no/ Last updated: 2026-08-03T07:38:12.000Z ### Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts This competition evaluates performance on real-world bilingual dialogues in Indonesian-Javanese, Nahuatl-Spanish, and Spanish-English. For speech technologies to truly have an impact on the world, they must adapt to how people actually communicate in digitally-mediated settings. For most of the world, multilingualism is the norm and language boundaries are fluid. Monolingual Automatic Speech Recognition (ASR) systems for high-resource languages - those with abundant training data and mature digital tools like English, German, or Spanish - have reached near-human performance on some benchmark datasets. Still, most of the world’s languages, and even real-world usage of majority languages, remain underrepresented in digital tools. Many existing speech datasets, both training and benchmark data, are monolingual, consist of primarily read rather than spontaneous speech, or fail to capture the domains of use that increasingly define everyday digital communication. Multilingual speech is also rarely just two monolingual streams stitched together. Code-switching (or code-mixing) — moving between two or more languages or varieties within a single conversation, sentence, or phrase — is one of the most common forms of everyday bilingual speech, and it can carry meaning, nuance, and identity that neither language conveys on its own. Because most ASR data assumes a single target language, models trained on monolingual corpora have no consistent way to represent the vocabulary, phonetics, or grammar of a language pair as it’s actually spoken. ## Task **Your task is to build ASR systems that accurately transcribe natural, code-switched speech across three underserved language pairs spanning North and Central America and Southeast Asia:** - North American Spanish-English - Spanish-Nahuatl - Indonesian-Javanese __Example of a mistranscription in Spanish-Nahuatl__ | ![Example of a mistranscribed Spanish-Nahuatl voice note](https://drivendata-public-assets.s3.us-east-1.amazonaws.com/mdc-sp-nh-mistranscription-example.png) | *Ah bueno, entonces ompa timotaj tejua, xamo tikneki tinechtaneutis cien pesos para ika nikoutikisas, no se piyo, lo que pasa tech nin semana fiero nimotak uan nochi niktamik in tomin den niknechikolteua, ajko nechajsi para ijkuak.* Ah ok, so I’ll see you there, by any chance can you lend me 100 pesos so I can buy a chicken as well, just that this week it’s not looking good and I’ve spent all the money that I had saved up for it, I won’t have enough for then. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Example of a mistranscription in Spanish-English__ | ![Example of a mistranscribed Spanish-English voice note](https://drivendata-public-assets.s3.us-east-1.amazonaws.com/mdc-spanish-english-mistranscription.png) | *Pues sí iría si iría but I have to know some of the details like la fecha y el horario you know because like I have to work.* Well yes I would go I would you but I have to know some of the details like the date and the time you know because like I have to work. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | These datasets, along with the held-out test set, consist of real voice-note dialogues between pairs of bilingual speakers, featuring heavy code-switching and lexical borrowing — the kind of natural, multilingual speech that dominates real-world communication but remains largely absent from existing benchmarks. Participants will build ASR systems robust enough to handle the fluidity of real, multilingual speech as it actually occurs in digitally mediated dialogue. Successful models could help transform vast, inaccessible audio-visual collections, such as those held by galleries, archives, libraries, and museums, into searchable, discoverable, and accessible digital assets, and advance speech technology for communities underserved by it. [Join Competition](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/?ref=community.mozilladatacollective.com) ## The languages in this challenge The three pairs span very different linguistic situations, but share a common thread: a widely spoken majority language sits alongside a language that is under-represented in speech technology, and in real life, speakers routinely blend the two. ### North American Spanish-English Tens of millions of bilingual speakers across the United States move fluidly between Spanish and English, often within the same sentence, in a pattern commonly referred to as “Spanglish”. Both languages are individually high-resource, but a model trained separately on monolingual English or Spanish corpora has no way to represent a language boundary that shifts mid-utterance, making this a common yet technically demanding form of natural speech in North America. ### Spanish-Nahuatl Nahuatl, the most widely spoken Indigenous language in Mexico with over 1.5 million speakers, belongs to the Uto-Aztecan family. It is polysynthetic and agglutinative, meaning words are built from many morphemes, the smallest units of meaning in a language, so a single Nahuatl word can carry the meaning of an entire English sentence. This structure, combined with numerous diverse regional varieties (some not mutually intelligible) and centuries of contact-driven borrowing from Spanish, makes Spanish-Nahuatl speech one of the most linguistically complex pairs in this challenge to model with standard ASR architectures. ### Indonesian-Javanese Javanese has tens of millions of native speakers — one of the largest languages in the world by that measure — yet remains under-resourced in digital and speech corpora relative to its size. Javanese and Indonesian are unusually close to begin with: Indonesian absorbed much of its Sanskrit-derived and courtly vocabulary via Old Javanese, and Javanese loanwords are embedded throughout everyday Indonesian, so the lexical boundary between the two is porous rather than clean. As a result, a given word can’t always be cleanly assigned to one language or the other, which complicates the basic assumption behind word-level language tagging. ## Prizes | Place | North American Spanish-English | Spanish-Nahuatl | Indonesian-Javanese | | ---------------------------------------- | ------------------------------ | --------------- | ------------------- | | 1st | $4,000 | $4,000 | $4,000 | | 2nd | $2,000 | $2,000 | $2,000 | | Bonus (best avg. score across all pairs) | $2,000 | | | | **Total** | **$20,000** | | | [Join Competition](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/?ref=community.mozilladatacollective.com) ### Bonus Prize After the competition ends, the organizers will average scores for those who submitted to all three language-pair tracks. A bonus prize will be given to the team that obtains the lowest average WER across tracks. To be eligible, teams must submit to all three tracks with the same team members, but may submit different models. ## How to compete Submissions in this challenge will open at a later date. Join the challenge now to start working with the data and get notified when submissions open. 1. Click the "**Compete**!" button in the sidebar to enroll in the competition. 2. Get familiar with the problem through the [Problem description](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/page/2/?ref=community.mozilladatacollective.com). Additional resources are available on the [About page](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/page/3/?ref=community.mozilladatacollective.com). 3. Explore the highlighted training datasets on the MDC platform, linked on the [Problem description](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/page/2/?ref=community.mozilladatacollective.com) page, or browse the full [MDC catalog](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com). 4. Once submissions are open, package your model files with the code to make predictions according to the runtime repository specification, which will be shared at that time. Then, submit your code as a zip archive for containerized execution from the Submissions page. You’re in! [Join Competition](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/?ref=community.mozilladatacollective.com) ## Competition rules The [competition rules](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/rules/?ref=community.mozilladatacollective.com) are designed to promote fair competition and encourage useful, reproducible solutions. If you are ever unsure whether your solution complies with the rules, contact the organizers by emailing [info@drivendata.org](mailto:info@drivendata.org). ### Note on external data and models **External data and pre-trained models are allowed in this competition.** Participants may use any dataset available on [MDC](https://mozilladatacollective.com/datasets?ref=community.mozilladatacollective.com) for training, including the three datasets highlighted in this competition. Participants may also use external data not currently hosted on MDC, provided that dataset is published to MDC at the conclusion of the competition. For more information on publishing data to MDC, see [MDC’s terms and conditions](https://mozilladatacollective.com/terms/providers?ref=community.mozilladatacollective.com) for data providers, and the [Contribute Datasets](https://mozilladatacollective.com/profile/uploads?ref=community.mozilladatacollective.com) page for instructions on uploading datasets. Open-weight, pre-trained models are also permitted in this challenge and may be run locally. Participants may not send competition data to third-party hosted models, model APIs, or inference services (e.g., OpenAI, Anthropic, or similar). See the [competition rules](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/rules/?ref=community.mozilladatacollective.com) for full details on permitted data and model use. [Join Competition](https://competitions.mozilladatacollective.com/competitions/1/lost-in-transcription/?ref=community.mozilladatacollective.com) If you have questions about data or model use, contact the organizers by emailing [info@drivendata.org](mailto:info@drivendata.org). --- ### Sponsors This challenge is sponsored by the [Mozilla Data Collective](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com) with support from the [Mozilla Foundation](https://www.mozillafoundation.org/en/?ref=community.mozilladatacollective.com). ### Compensated Datasets Is Now Available on Mozilla Data Collective URL: https://community.mozilladatacollective.com/compensated-datasets-is-now-available-on-mozilla-data-collective/ Last updated: 2026-07-29T19:02:59.000Z ## *Helping organisations participate more directly in the AI economy while making it easier for AI builders to discover responsibly sourced datasets.* Earlier this month, we shared a preview of Compensated Datasets and our vision for creating more transparent ways for organisations to participate in the AI economy while retaining agency over how their data is licensed and used. Today, we're excited to make Compensated Datasets publicly available on Mozilla Data Collective. High-quality data is the foundation of better AI. At the same time, many of the organisations and communities creating and stewarding that data have had few transparent ways to participate in the AI economy while retaining agency over how their data is licensed and used. Compensated Datasets helps bridge that gap by enabling organisations to make datasets available through Mozilla Data Collective while retaining control over pricing and licensing terms. In turn, AI builders can discover multilingual, multicultural and multimodal datasets that are responsibly sourced, thoughtfully documented and ready to support the next generation of AI applications. At launch, Compensated Datasets includes contributions from **TAUS, Pangeanic, Karya, Spotlite and ContentX Labs**, among others, representing a growing collection of datasets across languages, domains and modalities. This launch reflects the type of ecosystem we're building: one where AI builders have access to high-quality datasets, and where the organisations and communities creating that data are recognised, supported, and able to participate more directly in the value it creates. Whether you're looking for high-quality datasets for your next AI project or you're interested in making your own datasets available through Mozilla Data Collective, Compensated Datasets is designed to make those connections easier through a marketplace built on transparency, trust and shared value. [**Browse available compensated datasets.**](https://mozilladatacollective.com/datasets?pricing=compensated&ref=community.mozilladatacollective.com) If you're interested in becoming a data provider, we'd love to hear more about your work. [**Tell us about your dataset.**](https://docs.google.com/forms/d/e/1FAIpQLSdFskTVvrZWrw7DfsO5UiEshYsGUAHWhBAkMDNBaXThN4jMHw/viewform?ref=community.mozilladatacollective.com) ## **This is just the beginning** Today's launch marks an important milestone for Mozilla Data Collective, but it's only the beginning. We'll continue expanding the community of organisations contributing datasets, growing the range of languages, domains and modalities available, and making it easier for AI builders and data providers to connect through a marketplace built on transparency, trust and shared value. Whether you're looking for your next dataset or interested in contributing one of your own, we're excited to have you as part of the Mozilla Data Collective community. ### Common Voice segments now available through Mozilla Data Collective URL: https://community.mozilladatacollective.com/common-voice-segments-now-available-through-mozilla-data-collective/ Last updated: 2026-07-28T17:04:36.000Z Mozilla Common Voice is a massively multilingual platform for collecting speech data to train automatic speech recognition (ASR). Its mission is simple: to make language technology understand everyone’s mother tongue. But for datasets to be genuinely useful, they also need to be manageable. Many of the larger Common Voice datasets can be tens of gigabytes in size, which can make processing them difficult for builders and researchers who have limited technical resources. One of the most consistent requests we’ve heard over the years is that, for many of the world’s larger languages, people want access to only the portion that’s relevant to them. Whether you’re building a model for a specific region or creating resources for a particular community, it shouldn’t require downloading and handling far more data than you actually need. With that in mind, we’ve started a pilot programme to ship segments of Common Voice, so you can get the data you need for training models or creating resources, without the overhead of working through huge full datasets. ### **The pilot: targeted language varieties** For the first iteration, we used search queries from Mozilla Data Collective to guide our decisions. We looked at which language varieties were most frequently requested, and created datasets for them. The initial set includes the most searched-for varieties of: - **Chinese (Mandarin):** Beijing, Shanghai - **Dutch**: Belgian Dutch (Flemish) and Netherlands Dutch - **English**: British English, Australian English, Canadian English, Irish English, Malaysian English, South Asian English, American English and Southern American English - **French**: Metropolitan French and Canadian French - **Portuguese**: Brazilian Portuguese and Portuguese of Portugal - **Spanish**: Mexican Spanish, Caribbean Spanish, and Rioplatense For Arabic and three of the other varieties (Mexican Spanish, American English and Netherlands Dutch) we have also included two datasets split by gender of the speaker to facilitate work on debiasing. ### **Where to find the datasets** You can now find these dataset segments live on the site: - [Common Voice Scripted Speech 26.0 - Arabic (Female)](https://mozilladatacollective.com/datasets/cmrv0fgp00022nu077cdkkcny?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Arabic (Male)](https://mozilladatacollective.com/datasets/cmrv0f62m001wnu077iwrjbbo?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Beijing Chinese](https://mozilladatacollective.com/datasets/cmruwsbag00c8md07lk1n7i1e?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Brazilian Portuguese](https://mozilladatacollective.com/datasets/cmruxo9ew00d3md07veethj6k?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Flemish Dutch](https://mozilladatacollective.com/datasets/cmruvk68500awmd07sd3l6lqe?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Netherlands Dutch](https://mozilladatacollective.com/datasets/cmruvj2p300adnx073x7weuyo?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Netherlands Dutch (Female)](https://mozilladatacollective.com/datasets/cmruvkoaj00b0md07fjxzw6x5?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Netherlands Dutch (Male)](https://mozilladatacollective.com/datasets/cmruvkf8p00ajnx07rfn0ecv9?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Wu Chinese](https://mozilladatacollective.com/datasets/cmruwsjti00ccmd07d0jmdzyh?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - American English (Female)](https://mozilladatacollective.com/datasets/cmrt70j4z001qmm07nvfsmgmr?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - American English (Male)](https://mozilladatacollective.com/datasets/cmrt6zbgx000vmm07hfuefigk?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Australian English](https://mozilladatacollective.com/datasets/cmrt710620013mm071t45y6wb?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - British English](https://mozilladatacollective.com/datasets/cmrt6zrob000zmm07yqwjlpwi?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Canadian French](https://mozilladatacollective.com/datasets/cmrumhr9h0005md07ph5dkxpp?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Caribbean Spanish](https://mozilladatacollective.com/datasets/cmr2cf2x202uins07a4pe7m66?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - French of France](https://mozilladatacollective.com/datasets/cmrumhinj0001md07y3huuhs5?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Irish English](https://mozilladatacollective.com/datasets/cmrt71fck001bmm07cj9h312i?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Malaysian English](https://mozilladatacollective.com/datasets/cmrt71n6q001fmm07u64adwuv?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Mexican Spanish (Female)](https://mozilladatacollective.com/datasets/cmr3zunmk00flnt07p7ai9p43?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Mexican Spanish (Male)](https://mozilladatacollective.com/datasets/cmr3jjevj0041nt07vb3sm35d?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Rioplatense Spanish](https://mozilladatacollective.com/datasets/cmr468dhn00lvmm07a22jwgh5?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Scottish English](https://mozilladatacollective.com/datasets/cmrt717ci0017mm075ewqaw6v?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - South Asian English (India, Pakistan, Sri Lanka)](https://mozilladatacollective.com/datasets/cmrt70sar001umm07jwxzhw89?ref=community.mozilladatacollective.com) - [Common Voice Scripted Speech 26.0 - Southern American English](https://mozilladatacollective.com/datasets/cmryvrc6i014vo90786kriix5?ref=community.mozilladatacollective.com) Can't find what you're looking for or need a specific segment? [Get in touch with us](mailto:hello@mozilladatacollective.com) to let us know! ### Mozilla Data Collective datasets now discoverable through CLARIN’s Virtual Language Observatory URL: https://community.mozilladatacollective.com/mozilla-data-collective-datasets-now-discoverable-through-clarins-virtual-language-observatory/ Last updated: 2026-07-15T12:49:06.000Z *New collaboration expands visibility for community-governed language datasets and improves exploration of linguistic resources, services and tools.* [Mozilla Data Collective](https://mozilladatacollective.com/?utm%5Fsource=clarinwebsite&utm%5Fmedium=referral&utm%5Fcampaign=clarin-press-release) datasets are now [discoverable through CLARIN’s Virtual Language Observatory](https://vlo.clarin.eu/search?fq=collection:Mozilla+Data+Collective&ref=community.mozilladatacollective.com), making it easier for researchers, developers and language technology practitioners in Europe to find multilingual and community-centered datasets alongside other linguistic resources, services and tools. Mozilla Data Collective is the data sharing platform for human agency and fair value exchange. It enables communities, organisations, and individuals to share global cultural datasets on their own terms, while helping downloaders build more representative and culturally grounded technologies with data they cannot find anywhere else. The Virtual Language Observatory is CLARIN’s discovery service for language resources. By indexing Mozilla Data Collective dataset metadata, the collaboration creates a new pathway for users to explore the platform’s datasets, while preserving its core approach: data remains governed through Mozilla Data Collective, where uploaders set terms of use, access conditions and documentation for their datasets. The collaboration with [CLARIN](https://www.clarin.eu/?ref=community.mozilladatacollective.com) supports a shared goal: making valuable language resources easier to find, understand and use responsibly. For researchers and practitioners, this means Mozilla Data Collective datasets can now be discovered through a trusted infrastructure already used to explore linguistic resources across languages, modalities and domains. For uploaders, it means greater visibility for datasets that reflect real communities, cultures and linguistic contexts. The availability of Mozilla Data Collective metadata in the Virtual Language Observatory improves discovery without changing where datasets are hosted or how they are governed. Users who identify relevant datasets through CLARIN’s Virtual Language Observatory are directed to Mozilla Data Collective to review the datasheet, licensing information, access conditions and any additional terms set by the uploader. This approach is especially important for multilingual and underrepresented language resources, where context, provenance and community expectations matter. Mozilla Data Collective datasets include detailed documentation designed to help downloaders understand not only what a dataset contains, but how it was created, what it is intended for and what conditions apply to its use. The collaboration between CLARIN and Mozilla Data Collective marks a step towards building a more inclusive AI data ecosystem: one that improves dataset discoverability while respecting how communities choose to share their data. ### Get a Sneak Preview of Mozilla Data Collective’s Compensation Feature! URL: https://community.mozilladatacollective.com/get-a-sneak-preview-of-mozilla-data-collectives-compensation-feature/ Last updated: 2026-07-13T11:42:43.000Z Mozilla Data Collective was built to redefine how AI data is created, shared, and governed. As part of our mission to be the data sharing platform for human agency and fair value exchange, we have long-teased what so many partners and community members have requested: a tangible way to ensure that data owners receive fair compensation from data downloaders in return for access to high-quality, consentful datasets. The wait is now over. We are so excited to announce the release of our compensation feature - starting today, we are making the first datasets available for paid download! Read on for more details on this sneak preview, how you can take advantage of the compensated datasets feature, and what’s next up in this exciting new chapter. ### **Wayfinding in collaboration with Mozilla Data Collective partners** We have worked extensively with community partners to bring online some of the world’s most under-resourced language datasets, spanning Pashto (grassroots community), Catalan (Barcelona Supercomputing Centre, Project Aina), and Kinyarwanda (Digital Umuganda), to name only a few. Now, in response to the clear clarion call for equitable engagement and fair value returning back to these communities, we ran a couple of [Calls for Proposals](https://community.mozilladatacollective.com/call-for-proposals-mdc-is-commissioning-mission-aligned-datasets/) to source multicultural, multilingual, and multimodal datasets, and share them as a path toward a more sustainable model for ethical data sourcing. For each of these datasets, we engaged with trusted, close-to-context partners, and provided all requested compensation upfront in order to reduce the partner’s financial risktaking. This is just one tangible way in which we help to shift the market from exploitative gigwork and consentless scraping to one that centres human expertise and agency. ## **Presenting three of our launch datasets** [**Brazilian Sign Language (LIBRAS) Healthcare Corpus**](https://mozilladatacollective.com/datasets/cmrftle0s01fsnv0755jw602i?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=social&utm%5Fcampaign=payments-launch-2026). This large-scale multimodal corpus was developed by [DataPopAlliance](https://datapopalliance.org/?ref=community.mozilladatacollective.com) to advance inclusive AI and accessibility research for deaf communities in Brazil. It focuses on healthcare communication, including medical terminology, symptoms, preventive care, mental health, and clinical interactions, while maintaining an intersectional and gender-diverse approach. This is an important gap to fill because the health domain remains deeply under-resourced outside English. There are too few high-quality, ethically sourced health datasets in other languages, and even fewer that are multimodal or designed around accessibility needs. [**Medical Domain Mexican Spanish Speech Corpus (aka MEDMEX)**](https://mozilladatacollective.com/datasets/cmrdfjave008uo307ett7d4km?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=social&utm%5Fcampaign=payments-launch-2026). This dataset includes 10.5 hours of medical-domain read speech from two registered nurses in our community, based in Puebla, Mexico. It was designed to support evaluation of Automatic Speech Recognition systems and training of domain-specific Text-to-Speech systems in Mexican Spanish. [**Conversational Speech in Gujarati in the Medical Domain**](https://mozilladatacollective.com/datasets/cmrdfk4ao007po1070z5f4kem?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=social&utm%5Fcampaign=payments-launch-2026). This dataset was developed with our partner [Karya](https://www.karya.in/?ref=community.mozilladatacollective.com) and provides 25 hours of natural, multi-speaker conversational dialogue speech in Gujarati specifically recorded within the high-stakes medical domain. Gujarati is spoken by roughly 60 million people yet remains under-resourced in many AI systems, especially in critical domains across essential services, where language gaps can quickly become access gaps. [Visit Mozilla Data Collective](https://mozilladatacollective.com/?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=payments-launch-2026) and check these datasets out today! ### **How to take advantage of our compensation feature: Get in touch!** Together, these datasets point to the kind of ecosystem we want to help build: one where builders can access data that is useful, contextual, documented, and ethically sourced, and where the people who create those datasets are valued, recognised, and supported to find the right downloaders for their goals. **For those creating, or seeking, high-quality datasets - drop us a note! We would love to hear from you. You can fill out our super-quick question form about your dataset offer or need, and our Partnerships and Data Curation teams will get back to you.** [Get in touch!](https://docs.google.com/forms/d/e/1FAIpQLSdFskTVvrZWrw7DfsO5UiEshYsGUAHWhBAkMDNBaXThN4jMHw/viewform?usp=publish-editor&ref=community.mozilladatacollective.com) ### **What’s next? Get ready for more, directly from data owners around the world** This sneak preview of our curation of multimodal, multicultural datasets is only the beginning. In very short order (we’re talking weeks, not months), the compensation feature will go LIVE for users, and our team will be supporting data owners to share their own compensated datasets that collectively drive toward a more multilingual, multicultural, and multimodal future grounded in human agency. If you want to see what an alternative, sovereign, and human-centric AI data economy can look like, [sign up today](https://mozilladatacollective.com/?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=payments-launch-2026) and explore existing datasets. ## **FAQs** ### **What does this mean for data owners?** For data owners, this is a major unlock: the datasets they have built, curated, cleaned, documented, protected, and stewarded can now be shared with a clearer path to recognition and return. Whether the goal is to support research, sustain community-led data work, fund future collection, or, now, charge for access to high-quality datasets, data providers now have more ways to decide what participation looks like for them. ### **Who sets the pricing for datasets?** As a seller, you set the price for your dataset license. You may want to consider offering your dataset under different sets of terms and conditions for different audiences - for example, offering a dataset available under research terms and conditions for free, while listing it under commercial access terms for a price. You can do this by creating separate dataset listings for each version with the applicable terms for each, in conjunction with the individual access gate feature to review access requests before granting permissions. There is a $USD 100 base floor for datasets on the platform, and Mozilla Data Collective adds a low 5% platform fee to the transaction cost, which is paid by the downloader. ### **What are my primary responsibilities as a seller?** You are responsible for maintaining your datasets, ensuring legal compliance, setting licence terms, managing customer inquiries/refund requests, handling taxes, and maintaining your Stripe account. As a seller, you must provide a customer support phone number to Stripe, a dedicated support email address on your Mozilla Data Collective organisation page, and clearly communicate with buyers and/or MDC in the case of disputes. ### **How do I purchase a dataset?** To purchase a dataset, you will need to sign in to Mozilla Data Collective and navigate to the dataset listing. From there, you can view the terms and conditions for the license, and go through the checkout flow directly from the dataset listing page. The checkout flow is provided by Stripe, and will prompt you to add your payment details before you are able to download the dataset. ### 15 Datasets for Fine-Tuning Whisper on a New Language in 2026 URL: https://community.mozilladatacollective.com/15-datasets-for-fine-tuning-whisper-on-a-new-language-in-2026/ Last updated: 2026-06-09T10:40:51.000Z ## Why fine-tuning Whisper is a dataset problem [OpenAI's Whisper](https://openai.com/index/whisper/?ref=community.mozilladatacollective.com) changed what's possible in speech recognition: a single multilingual model with strong zero-shot performance across dozens of languages. But "dozens" is the catch. But for the long tail of the world's languages, even for some with tens of millions of speakers, Whisper's out-of-the-box output is unusable in terms of word error rate, the most popular metric. Fine-tuning fixes this, and the fine-tuning recipe is well understood by now: take a Whisper checkpoint, feed it paired audio and transcripts in your target language, and the model adapts fast. How fast can be striking: the documentation for the [Khmer ASR Cultural Dataset](https://mozilladatacollective.com/datasets/cmpdy0icy00l9nu07zo5hl1m3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) below reports that adding even a few hundred Khmer speech-text pairs to the training mix measurably lowered Whisper Large V2's character error rate, bringing it down to roughly 8%. What’s clear is that the bottleneck isn't the training code but the data. You need audio that's transcribed accurately, licensed in a way that lets you use it, and ideally varied enough in speaker, dialect, and recording condition that your fine-tuned model generalises beyond the training set. This article lists 15 automatic speech recognition datasets on Mozilla Data Collective that are well suited to Whisper fine-tuning, spanning a range of language families (Dravidian, Bantu, Uralic) and regions (Mesoamerican, South Asian, European). They range from large studio-production corpora to small, phonetically precise field recordings. Several are time-aligned, which is exactly the format Whisper fine-tuning wants. Pick the one in your target language, or combine several related ones to push a whole language family across the usability threshold. ## The 15 datasets ### South Asia - [**Malayalam Time-Aligned Speech Corpus**](https://mozilladatacollective.com/datasets/cmmno795h009hml07dh7uefvp?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: Community | Licence: CC-BY-NC-4.0 | Size: 1.50 GB, 6 hours | Task: ASR | Format: WAV, SRT - [**Tamil Time-Aligned Speech Dataset**](https://mozilladatacollective.com/datasets/cmnmyptri02glo107p5cx5por?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: MirasAI | Licence: CC-BY-NC-SA-4.0 | Size: 37.11 MB, 5 hours | Task: ASR | Format: OGG, SRT - [**Kannada Time-Aligned Speech Corpus**](https://mozilladatacollective.com/datasets/cmnggocgg007ymh079k30st39?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: MirasAI | Licence: CC-BY-NC-SA-4.0 | Size: 355.77 MB, 5 hours | Task: ASR | Format: OGG, SRT ### Southeast Asia - [**Khmer ASR Cultural Dataset (Version 3 - Part 5)**](https://mozilladatacollective.com/datasets/cmpdy19nn00llnu07y94pzeu4?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: DDD-Cambodia | Licence: CC-BY-SA-4.0 | Size: 33.14 GB, 87 hours | Task: ASR | Format: WAV - [**Khmer ASR Cultural Dataset (Version 3 - Part 7)**](https://mozilladatacollective.com/datasets/cmpdy0icy00l9nu07zo5hl1m3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: DDD-Cambodia | Licence: CC-BY-SA-4.0 | Size: 29.93 GB, 81 hours| Task: ASR | Format: WAV - [**Mandar Spontaneous Speech**](https://mozilladatacollective.com/datasets/cml5e30pd00eskr072e6a4rrh?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: Community | Licence: CC-BY-NC-4.0 | Size: 534.45 MB, 10 hours | Task: ASR | Format: MP3, TSV - [**Jember Javanese Spontaneous Speech Corpus**](https://mozilladatacollective.com/datasets/cmlgm5a94008kny07nz2intus?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: Universitas Gadjah Mada | Licence: CC-BY-NC-SA-4.0 | Size: 271.65 MB, 10 hours | Task: ASR | Format: MP3, TSV ### African Region - [**DataTrust Africa: Speech Corpus of Public Radio Recordings from Northern Uganda**](https://mozilladatacollective.com/datasets/cmkfm6xtw00k2nv07oakesnix?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: Community | Licence: NOODL-1.0 | Size: 179.82 MB | Task: ASR | Format: WAV, TSV - [I**siZulu Second Language Learner Speech Corpus**](https://mozilladatacollective.com/datasets/cmo1y8vrv006ol207lc86hc13?ref=community.mozilladatacollective.com) Steward: Community | Licence: CC-BY-SA-4.0 | Size: 5.26 GB | Task: ASR | Format: WAV, SQLite ### Siberia and Northern Eurasia - [**INEL Nganasan Speech Corpus**](https://mozilladatacollective.com/datasets/cmn4kxhp8001tnz07jyfvf2qt?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: University of Hamburg | Licence: CC-BY-NC-SA-4.0 | Size: 1.41 GB, 38.5 hours | Task: ASR | Format: TSV, MP3 - [**INEL Kalmyk Speech Corpus**](https://mozilladatacollective.com/datasets/cmn4kxlaj001xnz07a5yugnew?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: University of Hamburg | Licence: CC-BY-NC-SA-4.0 | Size: 138.31 MB, 3 hours | Task: ASR | Format: TSV, MP3 - [**INEL Evenki Speech Corpus**](https://mozilladatacollective.com/datasets/cmn4kxexu001pnz07kn6wr985?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: University of Hamburg | Licence: CC-BY-NC-SA-4.0 | Size: 103MB, 2.6 hours | Task: ASR | Format: TSV, MP3 - [**INEL Dolgan Speech Corpus**](https://mozilladatacollective.com/datasets/cmn4kqzzt0013nu07caxllg3t?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: University of Hamburg | Licence: CC-BY-NC-SA-4.0 | Size: 583.34 MB, 13 hours | Task: ASR | Format: TSV, MP3 ### Latin America - [**Speech Corpus of English Learners from Mexico**](https://mozilladatacollective.com/datasets/cmo29suhm00gjo2078rhmqn3p?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: Community | Licence: CC-BY-SA-4.0 | Size: 2.45 GB, 8 hours | Task: ASR | Format: MP3, TSV - [**Archivo GELED: Muestra general de audios del cuicateco**](https://mozilladatacollective.com/datasets/cmo29qkrl00g5o207hash6gk2?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) Steward: Community | Licence: CC-BY-NC-SA-4.0 | Size: 1.85 GB, 3 hours | Task: ASR | Format: WAV, TSV ## From scraped to shared: fine-tuning Whisper on community datasets Fine-tuning Whisper on a new language is one of the highest-leverage things you can do in speech AI right now: a few hours of well-transcribed audio can take a language from unusable to production-viable. For a step by step guide on how to fine-tune Whisper models using Mozilla Data Collective datasets you can [check out our step-by-step tutorial](https://community.mozilladatacollective.com/fine-tune-a-speech-to-text-model-for-any-language-including-yours/). The 15 datasets above give you audio across a deliberately wide spread of language families: Dravidian (Tamil, Kannada, Malayalam), Bantu (isiZulu), Uralic and Mongolic (Nganasan, Kalmyk, Evenki, Dolgan); and regions: Mesoamerican (Cuicatec), and Southeast Asian (Khmer, Javanese). Mozilla Data Collective is rebuilding the data ecosystem by keeping communities front and centre and giving the speakers of a language, and the organisations that record and preserve it, a say in how their speech datasets get licensed, credited, and applied. For under-served languages, this is a meaningful shift: away from models trained on whatever material can be scraped and extracted from the internet, and toward systems built on datasets that communities have deliberately shared and contributed on terms they set. [Browse all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=fine-tune-whisper-model) [Get in touch →](mailto:support@mozilladatacollective.com) [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### How Radio Free Europe/Radio Liberty Also Serves Its Communities Through Its Datasets URL: https://community.mozilladatacollective.com/how-radio-free-europe-radio-liberty-also-serves-its-communities-through-its-datasets/ Last updated: 2026-06-08T08:24:20.000Z For more than 75 years, [Radio Free Europe/Radio Liberty](https://www.rferl.org/?ref=community.mozilladatacollective.com) (RFE/RL) has promoted democratic values by providing accurate, uncensored news and debate in countries where a free press is threatened. RFE/RL reaches more than 44 million people every week across 18 countries, in 24 languages, including Persian, Russian, Ukrainian, Belarusian, Kyrgyz, Turkmen, Tatar, Chechen, and Georgian. RFE/RL’s archive is one of the richest collections of high-quality, low-resource language datasets in the world. ## From journalism to datasets Most of today's technologies are trained on data scraped from the open web heavily skewed toward English, toward the commercially valuable, and toward content that was never meant to represent the full diversity of human knowledge. Languages like Dari, Pashto, Kyrgyz, Tatar, Belarusian, and Romanian are radically underrepresented. The communities that speak them often go unheard by the models shaping how most of the world reads, listens, and learns. This is the gap [Mozilla Data Collective](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) was built to close. As the responsible data exchange platform, uploaders always own their datasets, set the terms of use, and decide who benefits. There is no scraping, no extraction, and no unauthorized repurposing of someone else's work. RFE/RL has shared 25 datasets through Mozilla Data Collective, making its multilingual journalism available for natural language processing tasks under terms RFE/RL itself defines. For NLP researchers and developers working on translation, speech recognition, summarization, or content moderation in underrepresented languages, this is a rare opportunity: datasets that are editorially rigorous, ethically sourced, and culturally grounded, contributed by an organization that has spent three quarters of a century earning the trust of its audiences. Some of these datasets include: - [**RFE/RL Chechen News Text Corpus**](https://mozilladatacollective.com/datasets/cmnykm7yi010rny0737lq2fth?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) - [**RFE/RL Kazakh News Text Corpus**](https://mozilladatacollective.com/datasets/cmnylmd5c000ynx07qonxot0v?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) - [**RFE/RL Persian News Text Corpus**](https://mozilladatacollective.com/datasets/cmnhgmkgg00wimh07wc0v3s49?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) - [**RFE/RL Pashto (Pakistani) News Text Corpus**](https://mozilladatacollective.com/datasets/cmnxk6nou00a5nu07kid0mewy?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) - [**RFE/RL Afghan Dari News Text Corpus**](https://mozilladatacollective.com/datasets/cmobui7g400cbmd07kmfe7oj4?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) - [**RFE/RL Tajik News Text Corpus**](https://mozilladatacollective.com/datasets/cmnypmsyf005enx07bc7nr16l?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) - [**RFE/RL Azerbaijani News Text Corpus**](https://mozilladatacollective.com/datasets/cmnz269cd00hcnr075v74km69?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) [Browse all of RFE/RL's datasets](https://mozilladatacollective.com/organization/cmk5l9e9g00a0no07x4bgdylc?utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) ## A model for the rest of the field The significance goes beyond any single dataset. The way institutions like RFE/RL share their datasets matters. When an organization with deep linguistic reach chooses a platform built on community ownership and fair value exchange, it sends a signal to every newsroom, archive, library, and museum watching: there is a third way. The current data economy is extractive, opaque, and dominated by a small number of players. The alternative being built on Mozilla Data Collective is community-driven, transparent, and grounded in real cultures rather than convenient ones. Today, Mozilla Data Collective hosts 600+ datasets across 300 languages, contributed by 190 organizations. Each one is a small refusal of the idea that language diversity is someone else's problem to solve. By participating in Mozilla Data Collective, RFE/RL is sharing its work in service of a more representative, more honest digital future. The languages of Kabul, Yerevan, Tashkent, and Tbilisi deserve to be part of how new technologies are built for the generations to come. Thanks to organizations like RFE/RL, they finally can be. [Explore all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=https&utm%5Fmedium=referral&utm%5Fcampaign=rfeblog) [Get in touch →](mailto:support@mozilladatacollective.com) [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### What makes a good dataset sample — and how to create one URL: https://community.mozilladatacollective.com/what-makes-a-good-dataset-sample/ Last updated: 2026-06-05T11:08:56.000Z Ever downloaded a huge dataset? Watched a Download meter for hours as tens of gigabytes trickle down the internet pipes to your hard drive? Only to then import it into `pandas` and discover it wasn't *quite* what you needed? We share your frustration! That's why providing a high-quality dataset sample is so important, particularly for large datasets. A sample is the "try before you buy" of the data world — a small, representative slice of a larger dataset that lets a potential user assess its suitability before committing to a full download. In this post, we cover two key pieces: what good sampling practice looks like (and how to do it in code), and how to upload a sample file to accompany your dataset on the [Mozilla Data Collective](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com) platform. ## What makes a good dataset sample? ### The case for sampling Providing a sample alongside a dataset is increasingly considered standard practice in responsible data governance. [ISO/IEC 42001:2023](https://www.iso.org/standard/42001?ref=community.mozilladatacollective.com), the international standard for AI management systems, emphasises transparency and accessibility in data documentation as a core requirement for trustworthy AI systems. Good sampling practice is part of that picture: it allows downstream users to make informed decisions about whether a dataset meets their needs, before they invest time or compute in a full download. The FAIR data principles — Findable, Accessible, Interoperable, Reusable — similarly treat dataset transparency as foundational. A sample makes a dataset more assessable, which in turn makes it more reusable. From a practical governance perspective, a sample also gives reviewers, auditors, and ethics boards something concrete to inspect without requiring access to the full dataset. For large datasets, the download cost alone can be a barrier to evaluation. A well-designed sample removes that barrier and increases the likelihood that your dataset reaches the people who could benefit from it. ### Types of sampling and their trade-offs Not all samples are created equal. The choice of sampling method shapes what a potential user can and cannot infer about the full dataset. The table below summarises the most common approaches. | Sampling method | What it does | Pros | Cons | | ---------------------------------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------- | | **Random sampling** | Selects rows uniformly at random, without regard for any column values | Simple to implement; no assumptions about dataset structure required | May under-represent rare categories (e.g. minority accents, low-frequency labels); sample composition is non-deterministic | | **Stratified sampling** | Samples proportionally from within defined subgroups (strata), such as gender, locale, or age | Preserves the distribution of important categorical variables; more representative of the full dataset | Requires knowing which categories are important before sampling; can be complex if many strata intersect | | **Systematic sampling** | Selects every *n*th row | Simple and deterministic; useful for ordered datasets | Can introduce bias if the data has periodic structure (e.g. repeated patterns every *n* rows) | | **Purposive / judgement sampling** | Manually selects rows to illustrate specific properties | Useful for demonstrating edge cases or data quality characteristics | Not statistically representative; may mislead users about dataset composition | | **Cluster sampling** | Randomly selects groups (clusters) and includes all rows from those groups | Efficient for geographically or structurally clustered data | Within-cluster similarity can reduce sample diversity | For most AI and machine learning dataset use cases, **stratified sampling** is the recommended approach. It ensures that categorical variables that matter for model training — such as speaker demographics, accent, or label distribution — appear in the sample in proportions that reflect the full dataset, rather than being left to chance. ### Creating a representative sample with pandas The examples below assume you are working with a Common Voice-style dataset in `.tsv` format, with columns including `client_id`, `sentence_id`, `sentence`, `age`, `gender`, `accents`, and `locale`. The goal is to produce a sample that is representative across the demographic dimensions that matter most for downstream model training and evaluation. #### Step 1: Load the dataset ```pythonimport import pandas as pd df = pd.read_csv('cv-corpus-en.tsv', sep='\t') # Preview the shape and column names print(df.shape)print(df.columns.tolist()) ``` #### Step 2: Inspect the categorical distributions you care about Before sampling, understand what you have. For a speech dataset, `gender` and `accents` are the categories most likely to be unevenly distributed. ``` print(df['gender'].value_counts(dropna=False)) print(df['accents'].value_counts(dropna=False)) ``` You will almost certainly find that some categories are sparse. The `accents` column in Common Voice datasets, for example, typically has a long tail: a few accent categories contain thousands of clips, while many contain fewer than a hundred. This is not a problem to hide — it is important information for a potential user, and a good sample should reflect it honestly. #### Step 3: Stratified sample across gender and accent ``` # Drop rows where both gender and accents are null, # since these cannot be assigned to a stratum df_stratifiable = df.dropna(subset=['gender', 'accents'], how='all') # Fill remaining nulls with a placeholder so they form their own stratum df_stratifiable = df_stratifiable.fillna({ 'gender': 'not_specified', 'accents': 'not_specified' }) # Sample proportionally: ~1000 rows, or 1% of the dataset, whichever is smaller sample_size = min(1000, max(100, int(len(df_stratifiable) * 0.01))) # groupby + apply with a lambda handles unequal stratum sizes gracefully sample = ( df_stratifiable .groupby(['gender', 'accents'], group_keys=False) .apply(lambda x: x.sample( n=min(len(x), max(1, int(sample_size * len(x) / len(df_stratifiable)))), random_state=42 )) .reset_index(drop=True) ) print(f"Sample size: {len(sample)} rows") print(sample[['gender', 'accents']].value_counts()) ``` Setting \`random\_state=42\` (or any fixed integer) makes the sample reproducible — the same code will produce the same sample each time. This is important for auditability. #### Step 4: Include rows with missing demographic data A sample that silently drops rows with null \`gender\` or \`accents\` values may mislead users about how complete the demographic metadata actually is. Consider explicitly including a small number of such rows: ``` df_no_demo = df[df['gender'].isna() & df['accents'].isna()] null_sample = df_no_demo.sample(n=min(20, len(df_no_demo)), random_state=42) sample = pd.concat([sample, null_sample]).reset_index(drop=True) ``` #### Step 5: Save the sample ``` sample.to_csv('cv-corpus-en-sample.tsv', sep='\t', index=False) ``` ### Sample datasets are a reflection of the integrity of your data A sample dataset is a claim about the larger dataset it represents. If your sample over-represents majority categories — because you used random sampling on an imbalanced dataset — users may form incorrect expectations about how useful the dataset is for training models on minority groups. If your sample under-represents them, users may incorrectly conclude the dataset lacks diversity it actually has. The most honest approach is stratified sampling with clear documentation of what strata were used, including a note on categories with very few examples. If a particular accent or demographic group has fewer than 10 clips in the full dataset, say so in your dataset's [Datasheet](https://community.mozilladatacollective.com/datasheets-the-missing-manual-for-your-dataset/). ## How to upload a sample dataset on the MDC platform Once you have exported your sample dataset, compress it as a `.tar.gz` archive. If you're not sure how to do this, our [dataset uploading instructions](https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform/) provide more information. Next, [create a new upload submission](https://mozilladatacollective.com/profile/submissions/create?ref=community.mozilladatacollective.com) (you'll need to be approved to upload datasets first). When creating the upload submission, you'll have the opportunity to upload a sample dataset, as shown below. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/06/image.png) Where to upload a sample dataset when creating a dataset submission After your dataset submission has been submitted and approved, your sample dataset will be available for download from your dataset's page, as shown below: ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/06/image-1.png) How to download a dataset sample You can also add a dataset sample to *existing* datasets by [viewing your uploads](https://mozilladatacollective.com/profile/uploads?ref=community.mozilladatacollective.com), selecting one, then editing it. ## Further reading - **ISO/IEC 42001:2023** — the international standard for AI management systems, which covers data governance and documentation requirements. - [FAIR data principles](https://www.go-fair.org/fair-principles?ref=community.mozilladatacollective.com) - [**Datasheets for Datasets**](https://arxiv.org/abs/1803.09010?ref=community.mozilladatacollective.com) (Gebru et al., 2021) — the foundational paper on structured dataset documentation. - [**pandas documentation on .sample()** ](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.sample.html?ref=community.mozilladatacollective.com) ### 15 Datasets for Building a Low-Resource Translation Model in 2026 URL: https://community.mozilladatacollective.com/15-datasets-for-building-a-low-resource-translation-model-in-2026/ Last updated: 2026-06-04T10:29:28.000Z ## The problem with "low-resource" machine translation Most production machine-translation systems in 2026 are still trained on a fairly narrow set of language pairs: the 50 or so for which the open web supplies enough parallel text to push BLEU scores into useful territory. Below that line, MT quality worsens significantly. For a Hausa-to-English translation system to be worth shipping, you need in the order of millions of aligned sentence pairs. For most of the world's 7,000+ languages, that volume of data doesn't exist anywhere and isn't going to be generated through web crawls alone. What does exist, increasingly, is community-led parallel-data collection. [Translators without Borders](https://translatorswithoutborders.org/?ref=community.mozilladatacollective.com) (now [CLEAR Global](https://clearglobal.org/?ref=community.mozilladatacollective.com)) ran the TWB Gamayun parallel sentence kits specifically to give humanitarian-language translation projects a starting point. Pakistani publishers, Mexican linguists, Sephardic Jewish revitalisation projects, and Nigerian research groups have built smaller but carefully curated parallel corpora in languages mainstream MT simply ignores. Combined, these resources don't replace web-scale parallel data but they're enough to fine-tune an existing multilingual MT base model into something usable for a specific low-resource pair. The 15 datasets below are the working set for that kind of project in 2026\. They range from the TWB Gamayun humanitarian sentence kits (Hausa, Lingala, Tigrinya, Nande, Rohingya, Swahili, Kanuri, Congo Swahili) to South Asian publishing-derived parallel corpora (Saraiki-English, English-Punjabi Shahmukhi), Sephardic Ladino lexical resources for Romance-family transfer, religious-domain multilingual parallel text, and a BOUQuET translation-difficulty evaluation sets that let you benchmark whatever model you end up training. ## The datasets ### African region - [**TWB Parallel Sentence kits - Hausa (30k)**](https://mozilladatacollective.com/datasets/cmom43ixg00hao00731e0o0jg?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 1.68 MB | Task: MT | Format: TSV - [**TWB Parallel Sentence kits - Congo Swahili (25k)**](https://mozilladatacollective.com/datasets/cmosl07v400w9nu07g3puif2t?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 2.18 MB | Task: MT | Format: TSV - [**TWB Parallel Sentence kits - Nande (15k)**](https://mozilladatacollective.com/datasets/cmoskn32a00vhmj07prz6k5ng?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 1.26 MB | Task: MT | Format: TSV - [**TWB Parallel Sentence kits - Swahili (5k)**](https://mozilladatacollective.com/datasets/cmoskxn8k00vtmj07ubxtu0f2?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 347.61 KB | Task: MT | Format: TSV - [**TWB Parallel Sentence kits - Tigrinya (5k)**](https://mozilladatacollective.com/datasets/cmoskmbpj00vxnu07w8lu7rrk?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 404.75 KB | Task: MT | Format: TSV - [**TWB Parallel Sentence kits - Lingala (5k)**](https://mozilladatacollective.com/datasets/cmosknxap00vlmj07kf6mugba?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 494.43 KB | Task: MT | Format: TSV - [**TWB Parallel Sentence kits - Kanuri (5k)**](https://mozilladatacollective.com/datasets/cmoskkop100vjnu0775hclbd1?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 358.46 KB | Task: MT | Format: TSV - [**English Hausa Parallel Corpus**](https://mozilladatacollective.com/datasets/cmn3ht40i00eami07lgydmrgg?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: LocaleNLP | Licence: CC-BY-NC-4.0 | Size: 164.32 KB | Task: MT | Format: CSV ### South and Southeast Asia - [**TWB Parallel Sentence kits - Rohingya (5k)**](https://mozilladatacollective.com/datasets/cmoskwsac00w3nu07b1nydlfb?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 358.88 KB | Task: MT | Format: TSV - [**Saraiki-English Parallel Corpus**](https://mozilladatacollective.com/datasets/cmmaphscg04t2mk07i1f8yc0q?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Kaleem Art Press | Licence: CC-BY-NC-4.0 | Size: 1.92 MB | Task: MT | Format: CSV - [**English–Punjabi (Shahmukhi) Parallel Sentences Corpus (Mediamen Archives)**](https://mozilladatacollective.com/datasets/cmkh9rso90076nv076jwxgjv3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: MEDIAMEN | Licence: CC-BY-NC-4.0 | Size: 1.08 MB | Task: MT | Format: CSV - [**Multilingual Religious Parallel Corpus (Kaleem Art Press)**](https://mozilladatacollective.com/datasets/cmk1bhogs3htwmk07wo7o9p6y?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Kaleem Art Press | Licence: CC-BY-SA-4.0 | Size: 2.27 MB | Task: MT | Format: CSV ### Europe - [**Ladino-Spanish Lexical Resources**](https://mozilladatacollective.com/datasets/cmo1qb62200anmk07ls9feuh5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Community | Licence: CC-BY-4.0 | Size: 39.92 KB | Task: MT | Format: TXT - [**Synthetic Ladino Parallel Corpus**](https://mozilladatacollective.com/datasets/cmpbmhj4i0067nw07tk46v2jp?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Community | Licence: CC-BY-4.0 | Size: 898.32 MB | Task: MT | Format: TSV - [**Sentence translation difficulty in Spanish - BOUQuET**](https://mozilladatacollective.com/datasets/cmngbf1tt0050nn07i49aebnk?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: MDC Curators | Licence: CC-BY-SA-4.0 | Size: 55.83 KB | Task: MT | Format: TSV ## Conclusion A low-resource translation model in 2026 isn't built by waiting for web crawlers to find more text in your target language. It's built by combining the carefully curated parallel corpora that humanitarian organisations, regional publishers, and language communities have produced specifically for this purpose; then fine-tuning a strong multilingual base model on the mix. The 15 datasets above are the working set: not enough alone to train an MT model from scratch, but more than enough to push a multilingual base model into useful territory for the language pair you care about. This is exactly the data ecosystem [Mozilla Data Collective](https://mozilladatacollective.com/?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) was built to enable. Mozilla Data Collective's mission is to put communities at the centre by giving the organisations and contributors who built these corpora real control over how their data is licensed and used, rather than ceding that control to whichever model provider happened to scrape it first. For low-resource translation in particular, the communities who speak these languages are the ones who decided to translate the sentences, validate the alignments, and release the data. At Mozilla Data Collective we’re proud to empower these communities by making that work searchable, downloadable, and properly attributed. [Browse all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) [Get in touch →](mailto:support@mozilladatacollective.com) [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### 15 Datasets for Building a Production TTS Voice in 2026 URL: https://community.mozilladatacollective.com/15-datasets-for-building-a-production-tts-voice-in-2026/ Last updated: 2026-06-02T13:20:55.000Z ## Why dataset selection matters more than scale There was a time when the entire conversation about text-to-speech could be reduced to a single question: how many hours of clean single-speaker audio do you have? The standard answer was twenty-four hours of [LJSpeech-style](https://mozilladatacollective.com/datasets/cmonjhxee01ako007kohpbg34?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) read audio and from that you got a model that sounded acceptable but flat. That era is over. The teams shipping production TTS now have moved past "more hours of one speaker." They train on deliberate mixtures: emotional speech for prosody, multi-speaker corpora for speaker generalisation, audiobook-derived data for narrative style, dialect-specific data for regional authenticity, non-Latin script data for under-represented writing systems, indigenous-language data for cultural inclusion. The model's quality is the mixture's quality, and the mixture's quality is a function of which specific datasets you put into it. In this article is a list of fifteen specific datasets from [Mozilla Data Collectiv](https://mozilladatacollective.com/?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post)e that we think belong in a serious TTS training stack in 2026\. Some are large and provide an acoustic baseline. Some are small and provide a specific signal, an emotion, a script, a voice profile, that nothing else covers. All of them are downloadable today, all of them have clear licensing, and all of them are on Mozilla Data Collective. ## The 15 datasets - [**Thorsten-Voice Dataset 2021.06 Emotional**](https://mozilladatacollective.com/datasets/cmm4b8f7700mgmh07cha3549n?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Community | Licence: CC0-1.0 | Size: 380.80 MB | Task: TTS | Format: WAV, CSV - [**Urdu Multi-Speaker TTS Dataset**](https://mozilladatacollective.com/datasets/cmmvykcrs0050ny07vkwww5gi?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Community | Licence: CC-BY-NC-4.0 | Size: 514.54 MB | Task: TTS | Format: WEBM, TSV - [**LibriVox Italian TTS Female Voice**](https://mozilladatacollective.com/datasets/cmo0qtw7v003knt07u6yupncc?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: MDC Curators | Licence: CC0-1.0 | Size: 61.74 MB | Task: TTS | Format: MP3, TSV - [**LibriVox Czech TTS Female Voice**](https://mozilladatacollective.com/datasets/cmo0jfvnw00p1nx070preklt5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: MDC Curators | Licence: CC0-1.0 | Size: 178.58 MB | Task: TTS | Format: MP3, TXT, TSV - [**Kokoro Speech Dataset**](https://mozilladatacollective.com/datasets/cmmknsho4014wmf087kvq5rc6?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Community | Licence: LibriVox Public Domain | Size: 3.98 GB | Task: TTS | Format: FLAC - [**Yoruba-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmo1nlaah0071mk077mw0qhpv?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 319.05 MB | Task: TTS | Format: MP3, TSV - [**Hausa-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmnopto3q00t0mf07v2dtc0ej?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post)Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 276.90 MB | Task: TTS | Format: MP3, TSV - [**isiXhosa-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmo4gtixz00kwny07hayfsk8s?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 276.02 MB | Task: TTS | Format: MP3, TSV - [**Tiv-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmo4nmfam00nxny07rssox2tj?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 311.58 MB | Task: TTS | Format: MP3, TSV - [**Duala-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmpmpf0jw021bnu0743hu7763?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 141.26 MB | Task: TTS | Format: MP3, TSV - [**Bamun-TTS-Dataset**](https://mozilladatacollective.com/datasets/cmnhjbnjp0115mh07kiha0rei?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Institute of African Digital Humanities | Licence: NOODL-1.0 | Size: 219.97 MB | Task: TTS | Format: MP3, TSV - [**Saraiki 10 Hours TTS Dataset**](https://mozilladatacollective.com/datasets/cmnggqr8z0082mh07vbsbm6t5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: MirasAI | Licence: CC-BY-NC-SA-4.0 | Size: 584.44 MB | Task: TTS | Format: WEBM, TSV - [**Chuvash TTS**](https://mozilladatacollective.com/datasets/cmnhhi0by00zknn07edrnd82e?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: Taruen | Licence: CC-BY-SA-4.0 | Size: 854.02 MB | Task: TTS | Format: PARQUET - [**Otomí (Hñähñu) TTS Voz Masculina**](https://mozilladatacollective.com/datasets/cmo0cro1g00hlmr07oichasyk?ref=community.mozilladatacollective.com) Contributor: Community | Licence: CC-BY-SA-4.0 | Size: 119.54 MB | Task: TTS | Format: MP3, TXT, TSV - [**Central Kurdish TTS dataset 1.0**](https://mozilladatacollective.com/datasets/cmj77njd701ljmb07m97pw1p3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) Contributor: The University of Melbourne | Licence: CC-BY-4.0 | Size: 293.45 MB | Task: TTS | Format: WAV ## Not Scale but Curation Production TTS in 2026 is no longer a scale problem but a curation problem. The fifteen datasets above won't, on their own, train a model but what they will do is let you build a deliberately diverse training mix that covers prosody, speaker variation, scripts, dialects, and the under-represented languages most commercial TTS still ignores. Mozilla Data Collective exists precisely to provide a platform for that diversity and unlock datasets from around the world that are more multicultural and multilingual. [Browse all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=production-tts-post) [Get in touch →](mailto:support@mozilladatacollective.com) [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### Open Home Foundation TTS datasets on Mozilla Data Collective URL: https://community.mozilladatacollective.com/open-home-foundation-tts-datasets-on-mozilla-data-collective/ Last updated: 2026-06-01T08:53:43.000Z Most voice assistants listen and respond in a handful of languages. Try to build one for your home that speaks your language, though, and you quickly run into a wall: the training data does not exist, or it is locked behind licences that make it unusable for open source projects. The [Open Home Foundation](https://www.openhomefoundation.org/about/?ref=community.mozilladatacollective.com) did something about that. Their argument was simple: your home is where you are most yourself, which makes it the worst possible place for a corporation to have a data pipeline. The Open Home Foundation was created in 2024 by the team behind Home Assistant, after years of watching the smart home market tilt decisively toward surveillance capitalism. The foundation now stewards over 250 open source projects, standards, and libraries, including [Piper](https://github.com/OHF-Voice/piper1-gpl?ref=community.mozilladatacollective.com), a lightweight TTS engine designed to run locally. ## **Why this matters** Local voice assistants are only as good as the data used to train them. And voice data raises every question that matters in the current AI economy: who records it, who owns it, who can use it, and under what terms. The Open Home Foundation answered those questions by releasing [23 TTS datasets](https://mozilladatacollective.com/organization/cmhdke5z6001rnp0777akc6i4?ref=community.mozilladatacollective.com) through Mozilla Data Collective, each one a set of scripted recordings from a single speaker, across multiple languages, all published under CC-0 with no attribution required or restrictions on commercial use. This was a specific choice that required a specific infrastructure. You can’t just upload voice recordings to a shared drive and call it ethical data stewardship. You need a platform that can enforce your specific version of open to determine access conditions, authenticate downloaders, and give the people publishing datasets real control over what happens to it. By publishing on Mozilla Data Collective, dataset uploaders enter a values-aligned community that cares just as much as they do about the data shared. Open Home datasets include: - [Lili 1.0 – Slovak](https://mozilladatacollective.com/datasets/cmiup5wqw01hlmf074qy07b80?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog), \~2 hours, female speaker. A West Slavic language with 5 million native speakers and, until recently, almost no open TTS resources. - [Mihai 1.0 – Romanian](https://mozilladatacollective.com/datasets/cmiupa1t801hwmf07xcawi3ve?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog), \~2 hours, male speaker. Romania has 19 million people and a growing tech sector. Now there's open TTS data to match. - [Anna 1.0 – Hungarian](https://mozilladatacollective.com/datasets/cmid4wbrc00hvnv07e448rwv3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog), \~1.6 hours, female speaker. Hungarian is a Ugric language with no close European relatives which makes dedicated TTS data especially valuable for model training. - [Dimitar 1.0 – Bulgarian](https://mozilladatacollective.com/datasets/cmhpahaib00d8mk07ely2m8wh?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog), \~1.4 hours, male speaker. Bulgarian uses the Cyrillic alphabet and sits in a particularly underserved corner of EU language tech. [Browse all the Open Home datasets](https://mozilladatacollective.com/organization/cmhdke5z6001rnp0777akc6i4?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) ## **TTS Datasets in Mozilla Data Collective** The Open Home Foundation's 23 datasets sit alongside a growing curation of TTS and speech data from communities that have historically had no good options for sharing their data on their own terms. For example, Mozilla Data Collective already hosts TTS corpora for under-represented languages from a wide range of regions, including [Otomi](https://mozilladatacollective.com/datasets/cmo0cro1g00hlmr07oichasyk?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) in the Americas; [Punjabi](https://mozilladatacollective.com/datasets/cmnypcx5p004jnr07k6ptxwcq?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) in South Asia; [Javanese](https://mozilladatacollective.com/datasets/cml5bn4k900aame07u0rwidcg?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) and [Betawi](https://mozilladatacollective.com/datasets/cml9hmuis017yo407k0p4i0t4?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) in Southeast Asia; [Ewondo](https://mozilladatacollective.com/datasets/cml16fpkn009lnt07ht6k406o?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog), [isiXhosa](https://mozilladatacollective.com/datasets/cmo4gtixz00kwny07hayfsk8s?ref=community.mozilladatacollective.com), and [Naija](https://mozilladatacollective.com/datasets/cmnykldcz010knu0737o8bgh9?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) in Africa; [Italian](https://mozilladatacollective.com/datasets/cmoiuyem401j5mr07s0jx8rqr?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog), [Czech](https://mozilladatacollective.com/datasets/cmo0jfvnw00p1nx070preklt5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog), and [Croatian](https://mozilladatacollective.com/datasets/cmnyx788r00d3nr07u04smvap?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) in Europe; alongside many other [multilingual TTS datasets](https://mozilladatacollective.com/datasets?task=TTS&utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog). This matters for TTS specifically because voice synthesis in anything other than English, Mandarin, or a handful of European languages remains badly under-served. The datasets exist in scattered archives, institutional servers, and researchers' hard drives. Mozilla Data Collective is one of the few platforms with the governance model to unlock them, giving the people who created those recordings real control over what happens next. ## **Data stewardship is a choice, not a default** The Open Home Foundation's datasets are a great contribution to a large problem. Their work has been inspirational to the many communities working with Mozilla Data Collective. The logic they demonstrate is clear: organisations that care about openness have to be intentional about *how* they share, not just *whether* they share. Mozilla Data Collective is a platform that makes intentional stewardship practical. And it is institutions like the Open Home Foundation that have a real impact on how tech is built by building and sharing datasets that reflect the true diversity of users across the world. If you are working on TTS for a language that is not yet covered, and you have recordings you would like to release, we’d love to hear from you! Email us at [support@mozilladatacollective.com](mailto:support@mozilladatacollective.com). [Explore all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=openhomeblog) [Get in touch →](mailto:support@mozilladatacollective.com) [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### Discover Dataset Insights with the new Data Provider Analytics Portal URL: https://community.mozilladatacollective.com/discover-dataset-insights-with-the-new-data-provider-analytics-portal/ Last updated: 2026-05-27T14:52:13.000Z In the new analytics tab in your user profile, uploaders can now see how their datasets are accessed over time through an interactive dashboard. The analytics view is launching with three activity charts showing completed dataset downloads over time, unique downloaders, and a breakdown of the methods used to access your datasets. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/05/image-6.png) You can filter the chart by dataset or time range, in order to see how your uploads are being accessed on a per-dataset or aggregate basis. The chart automatically adjusts based on the time window that you select (choose between today, last week, last month, last year, and all time) to help you spot patterns and spikes in dataset activity. Alongside the charts, the analytics portal surfaces metrics about how people are engaging with your datasets, showing: - **Total downloads** \- understand overall demand and interest in your datasets over time - **Unique downloaders** \- see how broad your audience is, and whether you have repeat downloaders who are accessing your datasets more than once - **Access breakdown** \- look at how people are downloading your datasets, and whether they're using the API to incorporate your dataset into a programmatic pipeline Together, these insights can help dataset owners better understand who their audience is and how people prefer to work with their data. Whether you're evaluating the impact of a dataset release, identifying growing interest in a language or domain, or planning future dataset updates, the analytics portal can give you a clearer picture of how your work is being used. To get started, data providers can access their analytics portal by visiting [mozilladatacollective.com/profile/analytics](https://mozilladatacollective.com/profile/analytics?ref=community.mozilladatacollective.com). Learn more about [uploading data to Mozilla Data Collective here](https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform/). [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) In the coming months, we'll be expanding the dataset analytics that we make available to data providers on the platform. We'd love to hear what kinds of reporting or insights you'd like to see next to shape this feature - so feel free [to get in touch](mailto:support@mozilladatacollective.com) with your ideas and requests! ### Building an African Voice for AI: Inside the Institute of African Digital Humanities URL: https://community.mozilladatacollective.com/building-an-african-voice-for-ai-inside-the-institute-of-african-digital-humanities/ Last updated: 2026-05-27T14:22:59.000Z When you ask a voice assistant a question in English, French, or Mandarin, the underlying models have been trained on billions of words and millions of hours of speech. Ask the same question in Bafia, Mada, or Suundi, and the technology simply doesn't know how to listen. The [Institute of African Digital Humanities (IADH)](https://inhunumaf.hypotheses.org/?ref=community.mozilladatacollective.com) is helping to change that, one carefully curated dataset at a time. ## Who they are Founded in 2021 as the *Institut des Humanités Numériques d'Afrique francophone* and now known as the Institute of African Digital Humanities, IADH is a research and practice network applying digital and AI methodologies to humanities research with a focus on Africa. It serves as a collaboration platform for Digital Humanities practitioners across the continent, connecting linguists, technologists, and cultural researchers. Since late 2024, the Institute has sharpened its focus on a specific and pressing challenge: designing and publishing NLP and machine-learning datasets for African languages, with priority given to those that are low-resourced and underrepresented in today's AI systems. ## What they've built IADH's contribution to language equity in AI is truly substantial. They have launched 46 under-served African languages on [Mozilla Common Voice](https://commonvoice.mozilla.org/en?ref=community.mozilladatacollective.com), including Adamawa Fulfulde, Bafut, Bakoko, Baoulé, Cameroon Pidgin English, Dagbani, Ewondo, Fang, Ghomala, Ibibio, Mada, Medumba, Mungaka, Tupuri, and many others. The majority had no prior digital speech presence at all. In addition they have contributed an enormous number of African NLP datasets on Mozilla Data Collective. ## The datasets on Mozilla Data Collective IADH has published [38 original datasets](https://mozilladatacollective.com/organization/cmfv3ichk000amd07piai0zoz?utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) onMozilla Data Collective, all released under [the NOODL-1.0 license](https://community.mozilladatacollective.com/how-open-licensing-is-changing-with-ai-the-noodl-license/?utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) and covering more than 20 African languages from Cameroon, Congo, Nigeria, and West Africa. They fall into five broad categories: **Speech datasets** for text-to-speech (TTS) and automatic speech recognition (ASR) including [Bati ASR](https://mozilladatacollective.com/datasets/cmj8fdiyv02egmb07m2wfdltm?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Beembe TTS](https://mozilladatacollective.com/datasets/cmj1gd6j400uvnw07ylungxjz?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Bomitaba TTS](https://mozilladatacollective.com/datasets/cmj2rze7r00j5ny07uhs85go2?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Bulu TTS](https://mozilladatacollective.com/datasets/cml9iik7d01efmn07miuf8yof?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Bamun TTS](https://mozilladatacollective.com/datasets/cmnhjbnjp0115mh07kiha0rei?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Ewondo TTS](https://mozilladatacollective.com/datasets/cml16fpkn009lnt07ht6k406o?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Hausa TTS](https://mozilladatacollective.com/datasets/cmnopto3q00t0mf07v2dtc0ej?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Kituba TTS](https://mozilladatacollective.com/datasets/cmj0az3yz002enw07okqee8jr?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Laari ASR](https://mozilladatacollective.com/datasets/cmj324gbx00p4ny078pi26kfz?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Lingala TTS](https://mozilladatacollective.com/datasets/cmm23jslb003znq07vdska54l?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Mbosi TTS](https://mozilladatacollective.com/datasets/cmj1gdg1s00vrnu07nmlrax7g?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Naija TTS](https://mozilladatacollective.com/datasets/cmnykldcz010knu0737o8bgh9?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Suundi TTS](https://mozilladatacollective.com/datasets/cmj1qhr5n001wnv07hqcatq25?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Teke-Laali TTS](https://mozilladatacollective.com/datasets/cmj8g2kw902fwmb07hub8puq8?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Yaka TTS](https://mozilladatacollective.com/datasets/cmj0c3cpe003enw07zxujgdt1?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Yoruba TTS](https://mozilladatacollective.com/datasets/cmo1nlaah0071mk077mw0qhpv?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog). Sizes range from tens of megabytes to several gigabytes of audio paired with transcripts. **ALCAM multimodal datasets** combining International Phonetic Alphabet transcriptions, audio recordings, and French glosses covering [Akoose](https://mozilladatacollective.com/datasets/cmj0b1ymo002pnu076g7oseat?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Basaa](https://mozilladatacollective.com/datasets/cmind910n0096nx07gme9v3wj?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Bulu](https://mozilladatacollective.com/datasets/cmnopvxxr00t6mf07nm4cp4qs?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Mvele](https://mozilladatacollective.com/datasets/cmmyyd6jq00k5mf07tjn4xh3j?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Yezoum](https://mozilladatacollective.com/datasets/cmndjfw9d000qmj070dece40f?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) and three regional varieties of Ewondo ([Yanda](https://mozilladatacollective.com/datasets/cmiv2rbv901mwmf07iyhv4pva?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), [Fong](https://mozilladatacollective.com/datasets/cmkmoepzx018wnw07cnkupls3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), and [Mbida-Mbani](https://mozilladatacollective.com/datasets/cmkmoepzx018wnw07cnkupls3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog)). These are particularly valuable for phonological research and for building models that handle dialectal variation. In the local context where indigenous language education is hindered by the lack of multimodal, multilingual and/or multidialectal resources to address practical learning needs, these resources are poised to have a significant impact. For example, users can listen to word pronunciations and access word or sentence meanings with the help of French or English translations. **Parallel corpora** for machine translation, pairing African languages with French: [Adamawa Fulfulde–French](https://mozilladatacollective.com/datasets/cml5asbhf009sme079y6sa9hm?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), Bamun–French (versions [1.1](https://mozilladatacollective.com/datasets/cml16bz8v008pnt07scfh66p5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) and [2.0](https://mozilladatacollective.com/datasets/cmn6ay1xg016jnv07drtmw7qo?ref=community.mozilladatacollective.com)), [Ewondo–French](https://mozilladatacollective.com/datasets/cmhqebl6h001pmn071yvbkbyv?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), and [Mada–French](https://mozilladatacollective.com/datasets/cmmamvzrz04qtmk077j1k99vt?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog). **Text corpora** of narratives and speech transcripts, including [FUB Narratives](https://mozilladatacollective.com/datasets/cmhvzlidq0326mn07hk4do3pj?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) (Adamawa Fulfulde), [Mada Narratives](https://mozilladatacollective.com/datasets/cmk19uruw39urmb07alqf493e?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog), and [Spoken Congolese French](https://mozilladatacollective.com/datasets/cmk1bdl0q39v6mb07gee6udp2?ref=community.mozilladatacollective.com), the last of which is over 3 GB and captures a regional variety of French rarely represented in NLP resources. **BOUQuET translation benchmark.** Working with the Mozilla Foundation, IADH produced gold-standard human translations of 1,364 sentences across 324 paragraphs into five Cameroonian languages (Basaa, Eton, Duala, Bafia, and Tupuri) for an international AI translation benchmark. Critically, every translation was produced by human language specialists, with no machine translation or generative AI in the loop. These datasets will be available soon on Mozilla Data Collective. [Browse all the datasets by the Institute of African Digital Humanities](https://mozilladatacollective.com/organization/cmfv3ichk000amd07piai0zoz?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) ## Why it matters There's a quiet asymmetry baked into modern AI: the languages spoken by hundreds of millions of Africans are often invisible to the models that increasingly mediate access to information, services, and culture. IADH's work, methodical, human-led, openly licensed, is one of the most concrete answers to that problem. Every TTS dataset, every parallel corpus, every transcribed narrative is a real step toward AI systems that can hear, read, and respond in languages they currently cannot. For researchers and developers working on multilingual NLP, the [IADH catalogue on Mozilla Data Collective](https://mozilladatacollective.com/organization/cmfv3ichk000amd07piai0zoz?utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) is worth bookmarking. IADH is keen to partner with developers working on tools and projects using these datasets, particularly those with immediate applications to language learning. For everyone else, it's a reminder that the future of AI doesn't have to be monolingual but, instead, aim to include every language possible for a more inclusive AI. [Explore all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=iadhblog) [Get in touch →](mailto:support@mozilladatacollective.com) [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### Never Miss a Dataset with the new Dataset Notification Feature URL: https://community.mozilladatacollective.com/never-miss-a-dataset-with-the-new-dataset-notification-feature/ Last updated: 2026-05-27T11:37:58.000Z With our new recommendation feature, you can now opt-in to a weekly digest that shares similar datasets to ones that you've downloaded on MDC. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/05/image-2.png) After downloading a dataset, you will now be asked whether you'd like to get notified of similar datasets to the datasets that you have downloaded. If you choose to be notified, we'll send you a weekly email digest with datasets tailored specifically to your download history. You can turn this setting on or off at any time in your account profile settings: ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/05/image-3.png) In the coming weeks, we'll be expanding our email notification center out to give you more granularity about the types of emails that you can receive from the Mozilla Data Collective team. In the meantime, if you have feedback about the types of datasets or features you'd like to see, reach out to us at support@mozilladatacollective.com. [Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### CoVoST 2 datasets now available through Mozilla Data Collective URL: https://community.mozilladatacollective.com/covost-2-datasets-now-available-through-mozilla-data-collective/ Last updated: 2026-05-25T08:00:57.000Z In 2020, Meta [introduced a new benchmark dataset](https://ai.meta.com/blog/covost-v2-expanding-the-largest-most-diverse-multilingual-speech-to-text-translation-data-set/?ref=community.mozilladatacollective.com) based on Mozilla Common Voice. CoVoST 2 has 34 translation directions for audio to text machine translation. This is one of the most widely-used benchmark datasets for speech translation, with [nearly 400 citations on Google Scholar](https://www.isca-archive.org/interspeech%5F2021/wang21s%5Finterspeech.pdf?ref=community.mozilladatacollective.com). The dataset is based on version 4.0 of Mozilla Common Voice, and that version is one of the most requested versions that we are asked for. As a result of the [changing expectations about data privacy and the right to be forgotten](https://community.mozilladatacollective.com/were-changing-access-to-older-versions-of-common-voice-datasets/), we put access to old versions of Common Voice behind an email-based process. Researchers who want to access previous versions now have to email us with a description of their project and why they need to use the data. This procedure is less than automatic and still requires researchers to download the whole version 4.0 dataset for each source language in order to extract the subset of clips used in the CoVoST dataset. We recently [launched a request-to-access feature](https://community.mozilladatacollective.com/request-to-access-feature-is-now-available/) on Mozilla Data Collective which allows uploaders to gate their datasets behind individual access requests. We thought that this is an excellent opportunity to make the process of accessing the CoVoST 2 dataset more streamlined. As part of our programme of dataset curation, we are providing experiment-ready versions of CoVoST 2 directly through the MDC platform. Some of the most requested datasets include: - [**CoVoST 2 English - Arabic**](https://mozilladatacollective.com/datasets/cmp5b9e10014mo007qkcmck75?utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=covost2blog) - [**CoVoST 2 Chinese (China) - English**](https://mozilladatacollective.com/datasets/cmp5kv8j101hvo007hjamlqet?utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=covost2blog) - [**CoVoST 2 Spanish - English**](https://mozilladatacollective.com/datasets/cmp704ive02qymp075orb8ok4?utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=covost2blog) [Browse all the CoVoST 2 datasets](https://mozilladatacollective.com/datasets?q=covost+2&utm%5Fsource=https%3A%2F%2Fcommunity.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=covost2blog) If you want to download CoVoST 2, just sign up for an MDC account, search for the dataset through our search bar or through our Data Assistant and click “Request to access” for the translation direction. You will need to give a short description of what your intended use case is, and wait for the access request to be approved – we aim to approve all within one business day. [Explore all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=referral&utm%5Fcampaign=covost2blog) [Get in touch →](mailto:support@mozilladatacollective.com) ### Press Release URL: https://community.mozilladatacollective.com/press-release-new-capabilities-expand-uploader-control-over-access-and-compensation/ Last updated: 2026-05-19T13:33:27.000Z **LONDON, MAY 19, 2026** – [**Mozilla Data Collective**](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com), a data platform redefining how AI data is created, shared, and governed, enters its next phase as a standalone, mission-led entity based in the United Kingdom alongside the introduction of three new platform capabilities: Request to Access, Data Assistant, and an upcoming Payments and Compensation functionality. Together, these capabilities advance Mozilla Data Collective’s vision for human agency and fair value exchange in AI data. ### **A New Chapter for Mozilla Data Collective** Mozilla Data Collective was introduced in November 2025 after the first version of the platform for community datasets went live, grounded in the belief that those who create datasets should control how they are accessed and used. Now operating as a mission-locked British company, Mozilla Data Collective is the first social enterprise incubated by Mozilla Foundation, which has committed up to $10 million to support its development, with $5 million already deployed. Since its soft launch, the platform has grown into a global community of uploaders and developers, with more than 190 vetted organizations sharing over 600 curated datasets across more than 300 languages. These include Hazargi literature from Afghanistan, oral histories in Mada from Cameroon, rescued government FactBooks, and Romansh newspapers from Switzerland. The datasets are already being used by thousands of public labs, journalists, researchers, and technology companies, including AI Unicorns across the UK and Europe. “There’s a false choice in AI right now that you have to choose between respecting data ownership and building high-quality technology,” said EM Lewis-Jong, Founder and CEO of Mozilla Data Collective. “Mozilla Data Collective is built to prove that’s not true. We’re giving communities real control over their data, including how it’s accessed, governed, and what they receive in return, supporting everything from traditional licensing to emerging models like data trusts. That’s what meaningful data sovereignty looks like in practice, and it makes it easier for developers to work with more representative datasets.” ### **New Platform Capabilities** Mozilla Data Collective is introducing three new capabilities that give uploaders more control over how their data is accessed, how value is exchanged, and how easily datasets can be discovered. With [**Request to Access**](https://community.mozilladatacollective.com/request-to-access-feature-is-now-available/), uploaders can require downloaders to submit an access request before any dataset is made available. This allows uploaders to review who is requesting access and confirm alignment with their intended use, whether for research, education, or commercial applications. Downloads are only enabled once the uploader has approved the request. The new [**Data Assistant**](https://community.mozilladatacollective.com/the-mozilla-data-collective-data-assistant-is-now-available-in-alpha/) simplifies dataset discovery by allowing developers to describe their needs in plain language, whether they are searching for existing datasets or looking for help sourcing new ones. In addition to surfacing relevant matches from Mozilla Data Collective’s growing curated collection, the assistant also allows developers to request datasets they may not be able to find elsewhere, particularly in underrepresented languages, regions, and modalities. Mozilla Data Collective will also soon introduce **Payments and Compensation**, allowing uploaders to set their own pricing for dataset access. Downloaders will be able to pay for a license to use the data, and uploaders will receive 100 percent of the license fee directly. Mozilla Data Collective charges downloaders a separate 5 percent platform fee to cover infrastructure and support costs, while uploaders pay nothing to use the platform. Together, these capabilities give uploaders more control over who can access their data, under what terms, and how value is exchanged. These capabilities are rolling out in alpha, with Mozilla Data Collective welcoming feedback, requests, and suggestions from the community at [support@mozilladatacollective.com](mailto:support@mozilladatacollective.com). ### **About Mozilla Data Collective** Mozilla Data Collective is a mission-locked British social enterprise, backed and incubated by Mozilla Foundation, building the data platform for human agency and fair value exchange. Mozilla Data Collective enables communities, organisations, and individuals to share global cultural datasets on their own terms, while helping downloaders build more representative and culturally grounded technologies with data they cannot find anywhere else. Built by the team behind Mozilla’s Common Voice, the world’s largest open, public-participation speech dataset, Mozilla Data Collective already supports more than 190 organisations sharing over 600 datasets across more than 300 languages. Learn more at [mozilladatacollective.com](http://mozilladatacollective.com/?ref=community.mozilladatacollective.com). ### **MEDIA CONTACT** Max Borges Agency for Mozilla Data Collective [mdc@maxborgesagency.com](mailto:mdc@maxborgesagency.com) ### Metadata magic: making datasets discoverable and tractable with Croissant URL: https://community.mozilladatacollective.com/metadata-magic-making-datasets-discoverable-and-tractable-with-croissant/ Last updated: 2026-05-14T12:40:43.000Z If a dataset exists on the internet and nobody can find it, does it actually exist? It's a slightly philosophical question, but it has very practical implications. The Mozilla Data Collective platform hosts hundreds of carefully curated, ethically governed [datasets](https://mozilladatacollective.com/datasets/?ref=community.mozilladatacollective.com) — but hosting alone isn't enough. A dataset that can't be found, understood, or assessed for suitability is a dataset that might not be used, reducing its value. Given the many hours of expertise that goes into collecting, curating and annotating datasets, it's important that we make the "last mile" - making sure they're found - as easy as possible. That's where metadata comes in. ## What is metadata, and why does it matter? *Metadata* is, in the simplest terms, data about data. It's the structured description that tells a potential user — human or machine — what a dataset contains, how it was collected, who created it, what licence applies to it, and how it's intended to be used. It's worth distinguishing metadata from the broader concept of *dataset documentation*. Dataset documentation encompasses everything you might want to know about a dataset: technical specifications, ethical review processes, sample data, usage restrictions, and detailed provenance. Metadata is a structured subset of that documentation — the specific fields and values that can be indexed, searched, and acted upon programmatically. Think of dataset documentation as the full manual, and metadata as the back-cover summary that helps someone decide whether to read the manual in the first place. For people who use datasets such as AI or ML engineers, researchers or analysts, metadata answers the essential first questions: Is this dataset in the right language? Does it cover the right domain? Is it licensed in a way that permits my intended use? Without good metadata, a downloader faces the unappealing prospect of downloading gigabytes of data and importing it into `pandas` only to discover it doesn't suit their needs at all. For AI agents and automated cataloguing systems, metadata is even more critical. A cataloguing agent can only reason about what it can read. Structured, machine-readable metadata enables agents to classify datasets, compare them, assess their fitness for a given task, and route them to the right users — all without human intervention. As agentic AI systems become more prevalent in research and data science workflows, the quality of a dataset's metadata increasingly determines whether that dataset gets discovered at all. ## What is the Croissant metadata standard? The Croissant metadata format is an open, community-built vocabulary for describing machine learning datasets. It was developed under the auspices of [ML Commons](https://mlcommons.org/working-groups/data/croissant/?ref=community.mozilladatacollective.com), the open engineering consortium behind benchmarks such as MLPerf, and is actively maintained by a working group co-chaired by Elena Simperl of King's College London and the Open Data Institute, and Omar Benjelloun of Google. Croissant was originally presented as a [research paper](https://doi.org/10.52202/079017-2610?ref=community.mozilladatacollective.com) in 2024\. Croissant builds on [schema.org](https://schema.org/?ref=community.mozilladatacollective.com), the widely-adopted vocabulary for structured data on the web. Schema.org provides the foundational types — `Dataset`, `Organization`, `URL` — and Croissant extends them with machine learning-specific fields: encoding formats, versioning, content size, language, and task type. Because Croissant is a schema.org extension, Croissant metadata is intelligible to any system that already understands schema.org, which includes most search engines. Over 700,000 datasets on platforms including Hugging Face, Kaggle, and OpenML are already described using Croissant, making it the de facto standard for ML dataset metadata. The format continues to evolve: extensions for geospatial data and life sciences are under active development, and the working group maintains an open-source [Python library for validating and consuming Croissant metadata.](https://pypi.org/project/mlcroissant/?ref=community.mozilladatacollective.com) One particularly important dimension of Croissant is its **RAI extension** — Responsible AI. The[ RAI namespace](http://mlcommons.org/croissant/RAI/?ref=community.mozilladatacollective.com) expands Croissant's descriptive potential to include fields such as `rai:mlTask`, which captures the downstream machine learning task the dataset is intended for. This is significant for us here at Mozilla Data Collective: our mission centres on ethical, responsible data sharing: a metadata standard that can encode not just *what* a dataset contains but *what it's for* aligns naturally with our platform's values. Croissant was the right choice for the Mozilla Data Collective platform for several reasons. It is an open standard under active community development, not a proprietary format. Its schema.org lineage means it integrates seamlessly with web-scale discovery infrastructure. And its RAI extensions allow MDC to encode dataset governance information directly into the metadata record — making responsible data sharing a first-class concern, not an afterthought. ## How we create Croissant metadata automatically from a Datasheet Croissant metadata for every public dataset on the Mozilla Data Collective platform is generated **programmatically and deterministically** from the dataset's Datasheet. This means that as soon as a dataset has a comprehensive Datasheet, its Croissant metadata is created automatically — no manual metadata authoring required. It's a key value-add of hosting your dataset with MDC. The mapping works as follows: | Datasheet field | Croissant field | Notes | | --------------------------------------------------- | --------------------- | ------------------------------------------------------------------------ | | Dataset name | name | Direct mapping | | Long description (or short description as fallback) | description | Prefers long description | | Creation date | dateCreated | Direct mapping | | Publication date | datePublished | Included if present | | URL slug | citeAs | Used as the citation identifier | | Locale / language | inLanguage | ISO language code | | Task type | rai:mlTask | RAI extension field | | Dataset URL | @id and url | JSON-LD identifier and standard URL | | File format(s) | encodingFormat | Parsed and mapped to IANA MIME types | | Task + format tokens | keywords | Combined keyword list for search | | File version number | version | Formatted as semantic version (e.g., 1.0) | | File size in bytes | contentSize | Included as a string value | | Licence URL, abbreviation, or text | license | Three-tier resolution: explicit URL > known abbreviation > verbatim text | | Organisation | creator and publisher | Mapped to sc:Organization nodes | The mapping deliberately excludes contact email addresses and personal names to protect contributor privacy and personally identifying information. This automatic generation approach has an important corollary: **the quality of the Croissant metadata is directly proportional to the quality of the dataset's Datasheet**. A sparse datasheet — one that skips the long description, omits the task type, or leaves the licence as a free-text string rather than a recognised abbreviation — will produce sparse Croissant metadata. A comprehensive datasheet produces rich, discoverable metadata. If you're a dataset producer preparing a submission to the Mozilla Data Collective, two resources are essential reading. The first, [Datasheets: The Missing Manual for your Dataset](https://community.mozilladatacollective.com/datasheets-the-missing-manual-for-your-dataset/), explains what a datasheet is, why it matters, and how to think about the fields. The second, [Uploading your dataset to the Mozilla Data Collective Platform](https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform/), provides field-by-field guidance on completing your submission. The time you invest in a thorough datasheet is time that directly benefits discoverability — for your dataset and for the community that might use it. The Croissant metadata is served as a `JSON-LD` document at a public, cacheable endpoint on each dataset's page. It is embedded directly in the page HTML, which means any crawler that visits a dataset page on the Mozilla Data Collective will find machine-readable metadata waiting for it - automatically. ## How Croissant metadata is used downstream The most immediate downstream consumer of MDC's Croissant metadata is **Google Dataset Search**. Google Dataset Search is a specialised search engine that indexes structured dataset metadata from across the web, enabling researchers and data scientists to discover datasets by topic, language, format, or licence. Because MDC embeds Croissant metadata as schema.org-compatible `JSON-LD`, Google Dataset Search can index MDC datasets automatically. You can see this in action right now: a [search for datasets on the Google Dataset Search platform](https://datasetsearch.research.google.com/search?src=0&query=site:mozilladatacollective.com&ref=community.mozilladatacollective.com) returns MDC datasets with structured previews — names, descriptions, licences, languages — drawn directly from the Croissant metadata our service generates. This is the practical payoff of good metadata. A dataset that is well-described in a standard format, embedded in a crawlable page, is a dataset that gets discovered by researchers who didn't know it existed. It shows up in searches. It gets downloaded. It gets used. For language communities whose data is under-represented in AI training sets — a core audience the Mozilla Data Collective was built to serve — that discoverability is not a nice-to-have feature. It is the point. And as agentic AI systems increasingly mediate how researchers and engineers find and evaluate datasets, the role of structured metadata will only grow. A Croissant record isn't just a search engine artefact. It is a durable, machine-readable description of what a dataset is and what it's for — the kind of description that an autonomous agent can read, reason about, and act on without human intervention. Good metadata helps take your dataset off the shelf of obscurity and into the hands of builders and makers eager to use it. --- For more information on how we generate Croissant metadata, or to learn more about Mozilla Data Collective, drop us a line at [support@mozilladatacollective.com](mailto:support@mozilladatacollective.com?Subject=Enquiry%20from%20the%20Croissant%20blog%20post). ### Best Practices for Sharing Contact Information on Mozilla Data Collective URL: https://community.mozilladatacollective.com/best-practices-for-sharing-contact-information-on-mozilla-data-collective/ Last updated: 2026-05-13T14:38:42.000Z When authoring a datasheet on Mozilla Data Collective, there are several considerations that you may want to make related to how downloaders can contact you about your dataset. In April, we made a change to how contact information was displayed on datasheets. Before the change, all contact information that was entered into the datasheet was displayed publicly on the dataset listing, which meant that data providers did not have control over whether to share their email addresses on the web. To better protect the personally identifiable information of data providers, we removed these fields from the datasheet. This change was intended to be temporary, until we could add a consent opt-in to display this information on the datasheet in order to give data providers more control over how - and when - to share their contact information with the community. We introduced our [conditional gated access feature](https://community.mozilladatacollective.com/request-to-access-feature-is-now-available/), which allows *downloaders* to share their contact information first in order to gain access to a dataset, as one step in a suite of improvements that help facilitate more intentional and specific interactions between data providers and downloaders. Last week, we made an update to the datasheet form. Now, as a data provider, you can choose whether to make the contact details public on your datasheet. If you check this box, your name and email will be visible, like before. By default, this is unchecked, which means the contact information will be hidden unless you explicitly decide to enable it. This can be changed at any time without needing your dataset to be re-approved. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/05/image.png) ## Sharing Contact Information on Mozilla Data Collective When choosing to share contact information on Mozilla Data Collective, consider how you want to engage with the broader community. You may want to selectively enable contact information on individual datasets using the checkbox on the datasheet itself, or put citation/contact information in the 'Description' or 'Other Information' field with additional context about how to get in touch with you about your dataset. You may also want to specify different forms of contact on your Organizational profile page. This is especially helpful if you are representing a team or community, and have different points of contact based on the type of inbound (e.g. support, partnerships). ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/05/image-1.png) If you would prefer to not share your contact information, but would still like downloaders to get in touch with you before accessing your dataset, you can set your dataset to 'Restricted' and use the access gate feature to allow downloaders to get in touch with you first. In the future, we plan on introducing more features that will make on-platform communication easier and more streamlined. If you have feedback on how that might look, we'd [love to hear from you](mailto:support@mozilladatacollective.com)! Join Mozilla Data Collective → ### FAQ: Why Can't I Edit Certain Fields on my Published Dataset? URL: https://community.mozilladatacollective.com/faq-why-cant-i-edit-certain-fields-on-my-published-dataset/ Last updated: 2026-05-14T15:35:33.000Z When you upload a new dataset to Mozilla Data Collective, you will need to fill in a datasheet that accompanies the dataset. After your dataset is approved, you are able to edit some of the fields on the datasheet, but not others. Fields that **cannot** be edited once your dataset is published include: - The license that governs the use of your dataset - Restrictions or prohibited uses for downloading your dataset - Additional terms for your dataset These fields cannot be edited as they collectively make up the legal agreement between you (the uploader), and any downloaders who agree to these terms. Changing these terms can result in a material re-licensing of your dataset, which may (depending on the extent of the changes) need for downloaders to agree to new terms before subsequent downloads. To re-license a dataset by changing these fields, you will need to create a new listing with the new terms and hide the old dataset. While this is not an ideal flow (and we are working on building a re-licensing/versioning for licenses on the platform to enable editing these fields), it ensures that anyone who downloads your dataset at any point will also have an explicit and clear understanding of which license terms applied at the moment of agreement, and a record of that should any conflict arise from the interpretation of the term agreement. ### Engendering Voice: Insights from Common Voice URL: https://community.mozilladatacollective.com/engendering-voice-insights-from-common-voice/ Last updated: 2026-05-07T16:10:29.000Z ## **The origin of Swahili language** Swahili originated from the Indian ocean and has its route from the contacts of Arabian traders with the inhabitants of the east coast of Africa over many centuries. Swahili is the lingua franca of most east african countries spoken largely in countries such as Tanzania,Kenya,Some parts of Congo DRC,Comoros and Uganda among others. In the early 19th century swahili was mostly a trade language and was later adopted by European colonialists, especially the Germans as the language of administration in [Tanganyika](https://www.britannica.com/place/Tanganyika?ref=community.mozilladatacollective.com), thus laying the foundation for its adoption as a national language of independent Tanzania. In other countries within east Africa Swahili is also recognised as one of the national languages although not the formal language of [administration](https://www.britannica.com/topic/Swahili-language?ref=community.mozilladatacollective.com). Over the years several dialects of swahili have emerged to be precise about 15 dialects however the Kiunguja dialect was adopted and made into the standardized swahili which is the modern day swahili spoken in the present day. It has been argued that in the standardization of the Swahili language the voices of the swahili people as a whole was not given a seat hence it was mostly orchestrated by colonial powers and adopted by the communities. ### **Swahili is Gender neutral** Swahili is one of the languages that is largely gender neutral in specific nouns this is also carried forward in pronouns such as “he/she “ which all translate to one swahili word”Yeye” So there is no **masculine or feminine**. A researcher sought to understand the dynamics of gender and the swahili people in one of her writing and this is what she had to say” I seek to counteract widespread tendencies to project assumptions of male dominance onto the past and to uncritically attribute current practices of gender segregation to the presence of Islam. Islam penetrated the Swahili Coast as early as the ninth century, yet gender segregation is a quite recent phenomenon dating back only to the last century. The data presented here cumulatively indicate: 1) that there have been shifts in Swahili society from a time when women occupied positions of greater power, prestige, wealth, and opportunity than is available to them today;⁺ ” ⁺Askew, Kelly M. "Female circles and male lines: gender dynamics along the Swahili coast." **Africa Today*, vol. 46, no. 3-4, summer-autumn 1999, pp. 67+. **Gale Academic OneFile*, link.gale.com/apps/doc/A132906218/AONE?u=anon\~85897988&sid=googleScholar&xid=8ed791e8\. Accessed 11 Dec. 2023. ### **The Swahili Queens** Digging deeper into this we find in history queens along the swahili coast from Mekatili of Mombasa to Mwana wa mema, Queen of Zanzibar and the historical role they played towards fighting colonialisms.A story is told of Mwana Mema of the Zanzibar people in 1650 who joined other Swahili elites in rebellion by forming alliances with the Ya'rubid dynasty of Oman. In 1651, Mwana Mwema invited a Ya'rubid fleet which killed and captured 50-60 Portuguese residents on the island, and she called for further reinforcements by sending two of her ships. However, the reinforcements didn’t arrive, and the elites of Kaole stone-town's rival city on the mainland would ally with the Portuguese to [force the Queen out of Zanzibar by 1652](https://www.africanhistoryextra.com/p/a-history-of-zanzibar-before-the?ref=community.mozilladatacollective.com). This stories prove further that gender imbalances within swahili communities were manifested from a white lens and further given wings by colonialism which inevitably played a crucial role to the spread and adoption of the language.This is visible too in the fact that Swahili drew some of its vocabulary from colonial languages such as portuguese,arabic ,English and germany.Over the years the work around Swahili have not given power to the local native speakers of the language and this has led to marginalization and not enough documentation of the language’s fading or disappearing dialects while standardized swahili takes over. | Mekatilili Wa Menza | Mwana Mwema, Queen of Zanzibar | [Sabani binti Ngumi](https://theconversation.com/ancient-dna-is-restoring-the-origin-story-of-the-swahili-people-of-the-east-african-coast-201154?utm%5Fmedium=email&utm%5Fcampaign=Latest%20from%20The%20Conversation%20for%20March%2030%202023%20-%202585525999&utm%5Fcontent=Latest%20from%20The%20Conversation%20for%20March%2030%202023%20-%202585525999+CID%5Feab407502656e160eec2a036813267c0&utm%5Fsource=campaign%5Fmonitor%5Fglobal&utm%5Fterm=Ancient%20DNA%20is%20restoring%20the%20origin%20story%20of%20the%20Swahili%20people%20of%20the%20East%20African%20coast) | [Mwana Mkisi](https://africaotr.com/powerful-queens-mombasas-history-and-the-legend-mwana-mkisi/?ref=community.mozilladatacollective.com) | | ------------------- | ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | ### **Voice tech and why Swahili** Swahili is now a lingua franca for about 150 million in Africa with the majority of the speakers being from Tanzania,Kenya and Congo DRC.Its one of the most spoken african languages with international recognition such as an official language of the African Union and in 2022 UNESCO horned the language by setting aside a Kiswahili day which is now every 7th of July. The voice tech space currently doesn't support underrepresented languages such as kiswahili hence the need for setting up an infrastructure to ensure the swahili speaking people are incorporated and can access voice enabled services in their language.However in the voice space access to datasets needed is very expensive hence Mozilla’s effort to democratize voice technology by making an open source voice database for kiswahili speakers. > *“Ironically, the story of Swahili origins has been molded almost entirely by non-Swahili peoples, a challenge shared with many other marginalized and colonized peoples who are the modern descendants of cultures of the past with extraordinary achievements.”* [https://www.sapiens.org/biology/ancient-dna-swahili-origins/](https://www.sapiens.org/biology/ancient-dna-swahili-origins/?ref=community.mozilladatacollective.com) ### **Gender disparities in Voice spaces** Despite the majority of Voice assistants being female by default or depicting female names and voices such as Alexa by Amazon, Siri by Apple among others, there is much to be desired in voice tech for gender inclusivity. The major challenges in having more gender integration in voice space include the lack of gender diverse voice datasets which make speech recognition more responsive to female than male users, access gaps due to gender communal disparities, data protection and privacy. This coupled with the fact that majority of builders are male focused hence biases become a challenge. ### **The Gender Action Plan** The Kiswahili voice project was aiming to curate a diverse and inclusive community that will help build an inclusive and trustworthy AI. Hence women are a crucial part of the tech and language community the Kiswahili common voice project hopes to curate. According to World Bank Data, the literacy rate for women in [sub-Saharan Africa is only 58%,](https://data.worldbank.org/indicator/SE.ADT.LITR.ZS?locations=ZG&ref=community.mozilladatacollective.com) Voice technology can help to break down the digital gender divide by reducing the need for traditional literacy in using applications; this is dependent upon the use, and acceptance, of the technology by women. Engagement and trust-building, with partners sensitive to the need for gender equity, will be vital to this outcome. The project seeks to mitigate AI biases by developing strategies for engagement with women, and on equity in data collection and model creation. Hence the need for a Gender Plan, detailing how gender diversity will be ensured during data collection, model creation and use case development. This Gender Action Plan contains specific gender elements to be considered in the project design and implementation,monitoring and evaluating progress. It will help build an effective strategy for gender mainstreaming in the project ensuring diversity and inclusion. --- ## **How to build voice use cases from a gender perspective?** Voice technology has revolutionized how things are done, from the streets workplaces all the way to our homes, people are now using voice as an enabler to access services and navigate menus, roads etc. It's predicted that in the coming years the majority of digital services will be linked to voice, making things more simple. Africa is yet to fully benefit from voice technology as most current applications do not cater for the varying accents and local languages spoken across the continent. For voice technology to be inclusive it needs to cater for the context and languages beyond the western world. In a practical world technology should benefit all genders but the reality is the process building up to building technologies such as voice assistants isn't gender inclusive by design. Let's make it practical: > *“ Company C has built application X for farmers in community Z however majority of the farmers in community Z are women with low-end devices, little to no access to internet yet application X requires its users to have good connection.While some of the women might have smart phones majority speak in their local language Y which is not covered hence this cuts the access by more than half already, adaptation is low and impact not visible.”* While the idea of Company C was great to help rural farmers access markets and implement their design process it did not consider the context and demographics of the community they aimed to serve such as gender.It's imperative to build applications that are user focused and more importantly cater to the needs of those that are usually excluded and this is the case when it comes to gender.In this article I will take you through the journey of how common voice built the kiswahili use cases with gender in mind and draw out lessons and guides to designing voice use cases from a gendered lens. There are key questions to ask in the process of designing voice use cases as it is in any other AI use case that we asked ourselves when we were thinking about the kiswahili common voice use cases that will be fueled by the common voice kiswahili voice dataset. ### **1 - What domains do we want to build in?** We researched the communities that speak Kiswahili to identify what are the most viable domains we could build use cases for that could serve majority of Kiswahili speakers. In most of sub saharan africa majority of economic activities rely on agriculture and hence why we thought it an areas that is still underserved.A report by NEPAD States that “Agriculture forms a significant portion of the economies of all African countries, as a sector it can therefore contribute towards major continental priorities, such as poverty reduction”(*Agriculture in Africa, transformation and outlook(2013)*). A report by McKinsey further highlights this by stating that “More than 60 percent of the population of sub-Saharan Africa is smallholder farmers, and about 23 percent of sub-Saharan Africa’s GDP comes from agriculture.”Voice integration in Agriculture will help meet the needs of more than 60% of Africa that work in the agriculture space (*Winning in Africa’s agricultural market(2019))*. Finance and Agriculture have a proportional relationship,you can not talk of agriculture and leave finance behind. With the adoption of mobile money across Africa, majority are relying on solution such as mpesa to send, receive and save their income even in the remotest places. Voice could play a crucial role in providing access to financial services across backgrounds,education level and even gender. Agriculture and finance are two sides of the same coin, each playing a significant role in enabling the other to build sustainable communities. As voice revolutionizes technology its imperative that it does the same towards agriculture and finance in enabling communities to work together towards achieving SDG’s. ### 2 - **Who is excluded in these domains?** Crucial to meeting the needs within these domains was to ask ourselves who was being excluded, how they were being excluded and how we build voice use cases that won't further exclude them. In our findings majority of farmers are women yet most of them do not have access to the infrastructure needed to facilitate their inclusion in usage of technology. So we deliberately set out to ask questions and engage communities of those excluded. ### 3 - **Who are the communities and stakeholders to engage in defining domain specific use cases?** After understanding that majority of left out communities where women and more precisely women in rural areas we wanted to understand how we can engage them in collectively designing the use cases they needed. In this step we realized we had to first build relationships with this communities including key stakeholders that work in this communities hence we identified Gender focused groups, NGO’s, farmers association, saving and credits associations, academics as well as technical communities building in this domains. This was relevant to get a holistic understanding of the situation,the context ,what has worked, what has not worked and how can we design sustainably with gender in mind. ### **4 - It's important to ask the rights questions,what answers do we need to design with gender in mind?** With the right people to engage that all round represented the communities we hoped to serve we had to think what are the right questions to ask to lay a foundation for the right voice use cases.In this case we divided our questions into understanding five key things: 1. **The demographics of the people we engaged with**: with consent we wanted to capture the gender diversity we captured in our engagement including underlying factors such as age,location,education levels and what field are they engaged in. 2. **The barriers to accessing voice technology**: It was crucial to understand what limits people from embracing voice technology. In this instance our engagement with communities aimed at looking at the social, cultural and economical barriers. 1. **Gender focused Use Cases for agriculture and finance**: We suggested areas of possible intervention of voice technology in agriculture and finance and gave room for the stakeholders to further suggest and pinpoint what they thought was relevant and needed. 3. **Identify groups which are currently benefiting, likely to benefit, groups not benefiting and those at a potential risk of being harmed by voice technology:** Something that is often overlooked when designing use cases is how different demographics of the community might be harmed,benefit or not benefit from it. In this case we identified groups by age,gender,socioeconomic status and explored further with interviews why these groups were like;y to be vulnerable or beneficiaries of said technology. 4. **Inclusion of marginalized communities in voice technology:** Lastly we explored what makes certain members of the community be excluded from such technologies,who are the communities excluded,how are they marginalized and how can we ensure they are included. This thought process of asking this questions led us to the realization of how we wanted to structure our use cases by identifying the barriers that led to gender marginalization in voice technology use cases and technologies. In this instance we set out to center the voices of the marginalized genders by identifying use cases brought out through the interactions with different stakeholders. The factors that limited engagement of diverse genders included: - **Device limitations to technology** : Majority of women and gender diverse groups have limited access to certain technologies such as smart phones and computers hence in designing use cases we needed to think about how this technology would be accessed and used specifically. - **Language barriers** : While voice technology was starting to come up in sub saharan africa the question is what language is this technology in? Majority of this type of technology is available in western languages or key languages leaving indigenous and underrepresented languages left out.Some of this is due to the lack of enough vocabularies of technical terms for example for the case of kiswahili more and more technical terms are being developed however not widely known. - **Affordability** : This was applicable to both access to devices as well as access to internet connection,majority of women in rural areas cannot afford continuous access to the internet due to cost implications.This was also the case when it comes to affordability of smartphones. - **Privacy concerns** : Trust is very important for gender diverse communities since things like trust have to be fostered and privacy addressed.In the context of the communities we were trying to serve this included how information on gender and sexuality was collected,used and the rights they have over their data. Gender is not an add-on when it comes to building AI systems; it needs to be thought through,integrated and weaved into the fabric of every system that is being designed from the very conception stage to the finish line and all the way to evaluation.Most importantly with designing use cases with gender in mind one needs to bear in mind the inclusion of those that are excluded and have their voices being centered.In doing this we learned from guidelines such as the design justice principles which lay a stage for ensuring we design with the voices of those we are building for.It's important to factor in the need for the use cases to function for people with low end device and limited or no connectivity. --- ## **Why does contributor gender matter in voice technology?** Data tells a story; it could be one of inclusion or exclusion depending on which data we chose to amplify over the other hence why contributor gender matters in Voice technology or AI broadly.The common voice platform has been collecting voice datasets since 2017 and in over the years there has been significant improvement in metrics collected and the communities that contribute.One of the key metrics that is collected by this platform is gender data allowing contributors to share information such as gender and age.Below we explore reasons why contribute gender matters in Voice technology. **Increases accuracy of Voice models for different genders**:One of the major reasons for speech recognition being more accurate for males than females is because most of the speech recognition models are trained from datasets that are heavily male.In this case data lays the foundation for the biases of the models that will be built out of it.While platforms such as common voice have made data crowdsourcing easier they have also provided the means to collect data relevant for amplifying inclusion by collecting metrics such as gender. [An article by Seclea simply states](https://www.seclea.com/resources/seclea-blogs/why-does-gender-equality-matter-in-artificial-intelligence/?ref=community.mozilladatacollective.com) “When there are unrepresentative datasets, inadequate models, weak algorithm designs, or historical human biases can result in unfair outcomes”. **Builds representative data**: Collecting contributor gender metrics is critical in ensuring that we are building representative datasets.The benefits of building representative datasets do not only increase accuracy of the models built for different genders but also solves different biases such as selection,algorithm and exclusion biases that are quite common in AI models such as voice assistants.Representation usual goes beyond gender to even considering underlying factors like age,the language variant spoken etc because all this allow for teams building to also have evaluation sets that can be able to measure the model if it's inclusive enough. **It addresses gender Selection bias**: Selection bias happens when the model is created by selecting particular types of instances more than others.In this case when the voice datasets have more male voice datasets than female datasets then selection bias is likely to take place because the model will be able to recognize more male voices than female voices.It's important to consider the diversity of datasets used to build AI products such as gender,age,race etc because they help to ensure that selection bias does not take place. **Breaks down the gender divide in data and AI**: It's common knowledge that there are huge disparities between different genders access to digital tools this ranges from as far as digital skills,access to digital tools,being part of the teams that build digital tools such as voice assistants.When we collect gender data we bring transparency to the table and can spotlight if the dataset weighs heavily on one gender over the other and vice versa which helps to break down some of these barriers.While voice tools are being built if they built off biased data its likely that they will further widen the gender digital gap instead of narrowing it down. Gender contribution matters in any domain. In this case the common voice platform is taking considerable measures to ensure that is captured and represented as part of the metrics of the datasets intentionally to keep gender at the forefront of dataset curation. A [report by the World web foundation](https://webfoundation.org/docs/2018/06/AI-Gender.pdf?ref=community.mozilladatacollective.com) states that “Only the meaningful inclusion of women at all stages will result in policies and technologies that make digital equality a reality.”For the case of emerging technologies such as voice technology, gender should be a base from data collection,model building all the way to how this application's impacts are measured and evaluated. --- ## **Gender and its intersectionality with Age in Voice Technology** Often technology is synonymous with Gen Z generation,majority of the population that actively engage with technology are Millennials and Generation Z, who were born in the 1980s and 1990s onwards. This age group are considered to be digital natives with most of them having grown up with the internet and being highly dependent on technology for various purposes. This often leaves out the older generation who have little to no interactions with most technologies in this case which comprises of ages from 50 upwards. Their is another subsection in this instant that is usually neglected and that is women above 50 and how they engage with different technologies including voice technology. One of the metrics that Mozilla common voice collects is age where contributors have the right to choose which age range they fall upon, studying the Kiswahili voice data set we find close relations between certain age groups and genders with their patterns of contribution to the platform, for example majority of the contributes are in their 20s and in this ages we have a close proximity between women and men contributors and even those who chose not to identify their gender. The bars are almost equally in terms of contributors in their 20’s whether they identify as male, female or other which is not the case from age groups 30’s, forties and this gets even lower almost non existent by the time you reach female contributors between age 50’s and 60’s. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/data-src-image-2fb225d4-63f2-46c2-8898-41d070015d7b.png) **Statistics from Kiswahili common voice data on gender and age as of January 2023* There are proven researches that agree there's a close proximity between age and AI biases,where algorithms are trained off datasets from younger ages hence being trained to be biased against older people.Some have termed this as Digital ageism which refers to ageism reflected design, development, and implementation of AI systems and technologies and its resultant data. [Currently, the prevalence of digital ageism and the sources of AI bias are unknown](https://www.nature.com/articles/s41599-023-01999-y?ref=community.mozilladatacollective.com). For common voice we realized there is a close relationship between digital ageism and gender biases in voice technology since most voice datasets are built out of datasets from men aged thirty and below. More women between forties and fifties have less interaction with technology and are less likely to interact with already biased voice assistants that are trained out of men voice datasets. A report by UNECE states that older persons, particularly older women, are at risk of being left behind in the region’s digital transformation.This is visible also in the embracing of digital skills where majority of older women are less likely to attain access to digital skills, more especially how they can adopt to emerging technologies such as voice assistants.In this Process their a few learnings we gained from our work with building diverse voice datasets on gender and age,which include: 1. Start from the basics: To engage older women in contributing to voice datasets or even embracing AI its goes without saying that you need “build them up” offer the skills they need to get started.This might involve offering teachings that are relevant such as basic digital literacy skills,this can involve introducing them to email creation or even how to use the internet.This will not only motivate their engagement but also built up their exploration into different ways they can use and benefit from technology usage. 2. Incentivise : For most incentives looks like cash and gifts but what we learned is the importance of skills sharing as an incentive. For example in some communities we realized older women wanted to horn or learn new skills such as speaking new languages such as english, family planning or even how to identify digital fraud.One of our grantees in the DRC coupled data collection and user testing of their land rights application with sexual and reproductive health education to older women.This incentivised engagement of many women including some who couldn't read and write but were willing to donate their voices. 3. Partnership with women organizations: One of the ways in which we managed to engage older women was through women only events that we collaborated to plan with organizations that work on women issues.In March 8th on women's day in 2022 one of our community champions partnered with a woman's local organization and brought together over 6o women for voice contributions with over 50% of these participants being in their fifties and sixties. 4. Social cultural contexts: one of the huge learnings in engaging older women was understanding that most of them have grown up in patriarchal societies and hence their cultural tries are very relevant to their participation and involvement in voice technology.In some instances we had to go to this woman to break down barriers such as transport to locations where we normally do our activities but we also had to learn their culture.For instance for women in some communities we had to get consent from their husbands or the head of the household in their homes who were mostly men.Despite that we aim to break down gender stereotypes we had to fit into their context for older women to participate in mozilla common voice. There is undeniably a close relationship between age and gender when it comes to technology uptake whereby the majority of older women are left out in technology.In the case of voice technology when we build with less data from older women we risk building voice assistants that are not responsive to women’s voices.Technology is biased towards certain demographic groups and this includes older women,to build inclusively we need to include women of different age groups along the path. --- ## **6 Ways to conduct Gender assessment of voice technology** There is a proliferation of voice assistants often being designed to replicate female gendered characteristics. However, there is much to be desired in voice tech for gender inclusivity for actual participation and meaningful engagement of women. Recent work has found how gendered biases against women are embedded into these systems.This includes Voice assistants bearing women names and Depicting women voices.In the light of this one of the key questions that kept coming up even after the drafting of the gender action plan was how were we going to assess if gender was really being uplifted in the products that the use case grantees were going to build up. Building from the gender action plan we did wide consultations with different stakeholders ,resources and tools available in the AI space to figure out how best we can guide our awardees through this process.We ended up creating a guide that helps voice technology developers to assess and better plan on how to mitigate gender biases and it draws its findings and assessment checkpoints from various gender resources to streamline them towards voice technology.Lets dive into some ways we can assess for gender in voice tech or generally any AI product. **Identify and frame the problem you are solving**: This is applicable to different stakeholders from product developers,funding institutions and even organizations supporting responsible AI.Let's unpack what it means to define and frame problems,in our case we had two main domains that our use cases where drawn from this is agriculture and finance. For each use case that was accepted we explored with our awardees asking questions such as who are you aiming to serve and in the process of serving them who else might be impacted,how will they be impacted.This questions helps to draw attention to key considerations such as how will you better be able to propose value or tackle the problem with a solution with minimal or no harm.This is important because when it comes to gender often not thinking about impact (positive and negative) is the point where it's easy for marginalized communities such as women to be further marginalized. [Resources - EqualAITools See More AI Literacy Resources See More Issues See More![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/equalAI-favicon-300x300.jpg)EqualAI![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/equal-ai-announcement-.jpg)](https://www.equalai.org/resources/?ref=community.mozilladatacollective.com) **Data quality control** : this is a critical area to address because in this instance you have to consider if the data your building from is biased in terms of gender for the case of the kiswahili use case grantees the dataset from common voice already showed the gender metrics of the dataset.In knowing the gender metrics of the dataset its then easy to tell which gender has less or more data and what would that entail when you build or if you will need additional validation or test datasets to evaluate the biases that your dataset might be carrying.We encourage carrying out tests for error rates for different subgroups to determine is there is a higher false positive rate for a certain gender over the other,this means comparing the character error rate (CER)/word error rate (WER) obtained from the testing set versus the evaluation sets created.This will help in identifying the accuracy of your model for different genders or other subsets.Lastly data quality control helps you to identify if your datasets correctly captures all the demographics of populations likely to be impacted by your product. **Fundamental rights and legal requirements**:It's critical from the onset of designing one's voice/AI solution to consider fundamental rights and ask questions such as are all fundamental rights protected and adhered to? This includes looking at laws in one's jurisdiction including data protection policies,AI regulations etc.One would need to draft privacy policies ,put in place mechanisms to allow for consent of users of the product and most importantly consider if their product does not infringe on the fundamental rights such as privacy,we put an emphasis on privacy because one of the major challenges of uptake of technology by different genders is trust whereby if no proper mechanism are in place to assure different genders that they can trust the product then there won't be incentive to use it.does your product support the right to privacy and to full control over personal data and information? [About | Feminist Principles of the Internet![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Feminist Principles of the Internet![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/logo_full_0.jpg)](https://feministinternet.org/en/page/about?ref=community.mozilladatacollective.com) **User centered design**: the center piece to all this is to center the voices of those you intend to serve in this case consider the women group of farmers you aim to serve and engage them in design thinking.This will not only help to identify what they need or what features would solve their challenges but would help you think through the kind of user interface you need to engage your users particularly considering the different personas of the people you intend to serve.Develop personas of a specific gender and engage them in the process to see how to design an interface that would be easier for them to navigate i.e. Auditing user interface for [gender in-balances](https://designjustice.org/?ref=community.mozilladatacollective.com). **Putting in place safeguards**:one of the worries of different genders in using technologies is their safety and this boils down to the fact that with newer technologies the risk of cyberattacks has also grown.Intelligent systems such as voice assistants collect a lot of personal and private data from users hence its critical to assess if your product is safe from cyberattacks,hackers and even trolls. Testing the vulnerability of your product is not the only place to stop. One has to go further and put mechanisms in place to ensure such doesn't happen and if they do how will you address the harms they pose. **Monitoring and evaluation plan**: Like any good project that is a good product for gender one has to constantly have a monitoring plan to flag down arising signals of bias,attacks and anything that might hinder equal participation of gender.It's imperative to have a plan to evaluate and improve on existing protections in place and user feedback to build better and inclusive to ensure offline harms are not replicated by your products. Assessing gender biases in technology is not a one size fits all depending on the context of the communities you aim to serve its critical to consider who holds power in this communities,how can you engage the power holders in the process and most importantly shifting narratives to allow for gender practices that uplift and not stifle one gender over the other. There is a lot of insights to be drawn from different reports and assessments, our learnings deeply drew some ideas and thoughts from the design justice principles,the feminist principles of the internet and the EqualAI Checklist to Identify Bias in AI. --- ## **Lifting up partnerships from a gender perspective** A strong pillar in building successful projects is “partnerships” not only does it bring different perspectives to the table but it also shapes agendas from a diverse perspective.In this article we explore how we leveraged partnerships from a gender perspective to uplift gender from community building,dataset curation ,model building all the way to dataset uptake and usage. ### **1 - Use Case building** In the process of designing the use cases to be supported by the common voice dataset we conducted a survey study to gain community insights on use cases where voice technology can be applied in agriculture and finance in East Africa. In this end we employed a few methods to capture gender in this design stage by doing the following: - **Wide consultation with feminist and gender based CSO’s working on agriculture and finance from a gender perspective**: we carried out interviews with different societies and groups including chama’s, women farmers associations and groups and the main aim in this was getting insight on how and which use cases will address specific gender concerns.This consultation led to identification of specific gender use cases such as women land rights and support for chama’s. - **Development of gender specific use cases addressing needs identified by gender stakeholders in the field of agriculture and finance.** These use prototypes were also subsequently tested by the respective gender groups to check if they captured the concerns and addressed impact to specific groups. ### **2 - Data collection** Majority of our work was geared towards building a diverse and inclusive dataset hence the efforts were deliberate from the start to ensure that is captured,this was not only an emphasis in the gender action plan but overall in the whole programing of the project. Below we capture ways in which we leveraged partnerships to lift up gender in the data collection stage: - **Engagement of gender groups,in sentence curation as well as voice collection**; we partnered with entities such as universities and gender focused NGOs to build text corpus and in this case we specifically hosted activities that would uplift gender participation by working with partners who were already working on gender. - **Supported women led and women focused organizations/events to host voice drives**:One of the things that came out of this was the realization of gender device gaps hence the partnerships with groups that focused specifically in gender helped to ensure availability of access and devices through partners spaces hubs for gender groups who have no access to contribute.Example of such partnerships included with Pwanitecknogirls in mombasa,the Arusha women school on internet governance in Tanzania and Core23lab in the DRC. ### **3 - Community building** One of the key success factors for the common voice work was that community was at the heart of the project from its inception. In this case we had to think of how we could be able to build on gender from the community level,in this instance both the language and tech community.This is how we leveraged partnerships to build on gender: - **Community champions:** we realized that we wanted to capture diverse and inclusive voice and it was going to be impossible for us to be everywhere hence we worked with communities and stakeholders such as NGOs and hubs to identify community champions who are able to reach communities we couldn't reach.In this effort we were deliberate in picking majority women community champions to helps us spearhead data collection of underrepresented genders.Some of our champions went as far as collecting voices of older rural women to queer communities. - **Partnerships with gender groups**;this goes without saying the most success we had at hosting gender focused drives were those in which we partnered with partners that work specifically on things such as women empowerment programs. ### **4 - Dataset uptake and usage** One of the crucial stages where gender biases can be addressed is the model building stage where after much reflection we utilized partnerships as well to promote uptake of the common voice dataset and its usage from a gender perspective by this i mean how we can build gender responsive models from the dataset.Some ways we did that included: - **Organized hackathons with an emphasis on models that perform well on different genders in partnership with Zindi,Africa’s talking and Swahilipothub;** this hackathons had a strong emphasis on inclusion of women including comparing the overall performance of the model with how it compares on the following evaluation sets representing the following demographics such as Women,[Women under 30 And Women over 30](https://zindi.africa/competitions/mozilla-foundation-mozilla-common-voice-hackathon-i-foundation-mozilla-common-voice-hackathon-i?ref=community.mozilladatacollective.com) among others.These hackathons were conducted across Tanzania and Kenya with support from different partners including community champions. - **Presentation of the common voice work and dataset at partners conferences:** we utilized platforms given by partners such as the africa AI conference to present how one can make use of the data and the diversity of our data.This was a great way to share how our data has training sets on different genders and other demographics critical at evaluating biases. Partnerships level the playing field especially when it comes to gender. In our case partnerships played a crucial role at helping to uplift gender by bringing to the table diverse thoughts,ideologies and perspectives.Different actors can take up tangible actions to deliver on gender equity and equality at different stages of the AI ecosystem. [A report by McKinsey and company](https://www.mckinsey.com/industries/social-sector/our-insights/partnering-for-parity?ref=community.mozilladatacollective.com) States that “ Gender inequality is complex by nature; it cuts across all sectors, interlinks with other areas of development, and requires solutions from many actors including those in government, corporations, and nongovernmental organizations (NGOs).”Hence the need to think through how to establish and uplift partnerships to tackle gender disparities when it comes to emerging technologies such as voice tech. --- ## **Gendered realities: learnings from MCV awardees on Gender** In 2022 mozilla common voice awarded a total of 400,000$ in grants to 8 awardees across three countries of Kenya,Tanzania and the Democratic republic of Congo to build solutions out of the kiswahili common voice datasets.It was deliberate from the beginning to seek projects that are gender inclusive in nature by ensuring the teams of the projects selected comprised of women and had a gender focus in the product they are building.This awardees worked on several projects ranging from the maize value chain,orange flesh potatoes ,climate change,corporate serving groups all the way to land rights.This awardees were supported at different stages on how to build with gender in mind,here are some gendered realities that they experienced and learnt along the way. **Gender device gap**:One of the challenges that the awardees faced both during data collection as well as introducing the applications to the communities was the gender device gaps where majority of women did not own smartphones let alone feature phones.One of the grantees who was working on land rights for women in DRC says they had thought of this gap long before project implementation and hence had budgeted to hand over a few phones to women community leaders to help spearhead the movement.However they still felt that even with providing access through leaders of the women groups there was still a fear that the mean might take it over so they enlisted the help of a woman chief in the community who is ensuring that the phone stay in the hands of the women and for the purpose of sharing knowledge with others. **Communities are patriarchal**:All the awardees agree that the culture in the communities they worked embarrassed patriarchy and hence their practices did not give women the right to engage in big decisions including ownership of mobile devices.One of the awardees SEE Africa who build an orange sweet potato flesh app called Kiazi Bora states that “These cultural practices and beliefs have always placed men above women, this has led to women being excluded from education and income generating opportunities; and as a result majority of women are limited to only doing household chores, majority cannot use smart mobile phones because of lack of education and exposure.”This was the case across the awardees where the communities they worked in did not value the women although they were in most cases the main producers in farms. **Access gaps:**Being able to meet and engage rural women was a big challenges especially because their was not only devices gaps but in rural areas there was also connectivity challenges where by majority had low end devices and in some cases internet was low or unavailable.It was reported by DuniaCom one of the awardees that build a solution for the maize value chain that Poor internet connectivity in rural areas prohibits consistent use of mobile applications and adoption of voice-based agricultural value-added services. “While this has significant implications for gender and inclusion, it also represents a potential market loss for voice-based AgriVas solutions.” **Women had lower literacy levels :** It was noted by most of awardees that the communities they engaged with there was a huge disparity between men and women when it comes to literacy ,more men had education compared to women.Think who developed the chama application says”Most Member that we interviewed with highest level of study as Primary School had a timid look on financial matters and were not as bubbly as more learned persons, The timid persona had little to no control to financial matter Cultural background also plays a role in this. “This was also observed by Core23lab who noticed that most of the women they were trying to engage did not know how to read nor write hence limiting their engagement with digital devices and tools even further. **Trust and data protection challenges**: While laws exist in countries to protect against data breaches there are no gendered laws that address the digital threats women face in digital platforms.DuniaCom states that “For instance, the project came across a significant number of households in which sim-cards used by women were registered under men. This contravenes the existing cybercrimes ACT in 2015\. The development of voice recognition technology in particular needs to be supported by appropriate policy and regulatory frameworks that are understood and adopted by businesses and innovators. Otherwise, it would lead to gaps in transparency, accountability, safety, and ethical standards of the technology that are likely to be detrimental to gender and digital inclusion, and damage the promises of voice recognition technology in development. **Power dynamics**: Women have limited decision-making power at the household level. For example Badili, one of the awardees, states that “we were informed at the workshop that Masai women don’t make any decision regarding selling cattle, goats or sheep. They can only decide on poultry.”Strathmore says” Women have lower decision-making power concerning some farming activities, therefore, climate adaptation practices that are suggested may not be implemented. The design of climate adaptation practices can take a gendered approach to ensure uptake.” **Women are more vulnerable to digital/online risks than men** (e.g., women experience more harassment from strangers, receive unsolicited messages, and face intrusion of privacy). The fact that women are more digitally illiterate than men, makes them more vulnerable to digital/online risks. These risks are also underpinned by gender norms as men are expected to play the role of gatekeepers and always control how women access and use digital technologies due to fears that the technologies are not only unsafe, but also have corrupting influences on women. Some risks are internalized by women themselves as many believe it is acceptable for their devices and communications to be monitored by men who are perceived to be more literate and less prone to digital harm. The gendered realities of technology adaptation prove that there is a lot of work to be done to achieve parity and even adaptation of solutions such as voice AI,it needs a closer look at gender,gaps and finding solutions that allow the community to thrive inclusively beyond glamorous AI solutions.It's advised that involving women in the process gives the project acceptance and legitimacy for other women this includes having women AI experts or even women community champions spearheading the adaptation of the applications.Dismantling patriarchy in a did to do away with digital exclusion requires breaking it down in manner that informs the community that empowering women is not a threat. While these challenges persist the common voice awardees believe there are ways to mitigate them,their learnings on digital exclusion especially where gender is concerned have those two tied together.See Africa believes that “Digital exclusion cannot be arrived at without addressing gender exclusion, gender exclusion precedes digital exclusion; digital inclusion requires a firm background in understanding the values of gender equality which provide room for people of all genders to access education and thus get exposed to other aspects of the developing world including digital technology.” --- ## **A feminist approach to text corpus creation** Language models are [becoming increasingly central to artificial intelligence](https://community.mozilladatacollective.com/how-open-licensing-is-changing-with-ai-the-noodl-license/) through their use in online search, recommendation engines and language genera- tion technologies. It has been urged that concepts of gender can be deeply embedded in textual datasets that are used to train language models, which can have a profound influence on societal conceptions of gender. The representation of gender in large text corpus matters not only from the creation perspective but also from the aspect of how the resulting models perceive gender.Common voice as a platform has provided avenues for text corpus to be created and contributed to by communities through the sentence collector,however there are deliberate efforts on how this texts can be gender inclusive.At the beginning of the work of building the kiswahili text corpus we employed different methodologies and learnt a lot from them,one of the things we realized is that there is a gap in terms of text corpus that uplifts different genders.This coupled with the realization that there is not enough women building AI solutions like language models nor are their stories told especially in the communities we were working with.This learning allowed us to integrate the need to create feminist content and use it to amplify women voices while also creating a feminist text corpus.Our approach included: **Feminist write-a-thons:**It was important to curate content that is relatable to the community hence why this contents were to be developed by the communities themselves.This included conducting feminist write-a-thorns were participants delved into creating content around gender,AI and the people around them.The first session was through MozFEST virtual where the participants who joined the session were introduced to the history of the swahili queens along the east african coast and their role in the growth of the language.Then they engaged in using swahili proverbs to create powerful feminist narratives.This narratives were later cleaned by splitting them into eligible sentences and uploading them on the common voice platform as part of the kiswahili text corpus.The benefit of such narratives was it did not only demystify the concept around the role of gender in development of languages but it also addressed areas of importance such as power dynamics in communities. **Writing competition of feminist stories**:In 2023 a writing competition was launched on women's day with the aim of curating stories of women change makers in the three countries the Kiswahili work was focused on.In this instance a set of guiding principles for participation were put forward including the need for them to focus on women within their communities,writing in their swahili with hopes of bringing to light invisible women in various spaces while also building a feminist text corpus.The writing competition was dubbed “ wanawake mashujaa “ covering stories from Kenya,Tanzania and the Democratic Republic of Congo.The competition was open to people of all genders to submit written content in their preferred swahili dialect under CC0.The[ “Wanawake Mashujaa” Competition — Kiswahili for “Brave Women”](https://foundation.mozilla.org/en/blog/common-voice-announces-nine-winners-in-wanawake-mashujaa-competition/?ref=community.mozilladatacollective.com) — called for written biographies of extraordinary women in local communities across Tanzania, Kenya, and the DRC. There were specifications on how this was to be written in a manner that respected gender diversity but also allowed for the text to go into the common voice Kiswahili corpus.A collection of over 40 stories of 40 powerful women were submitted and out of this 9 winning essays were awarded for being exceptionally written by a panel of linguist and literacy judges of the Kiswahili language.This approach enabled us to not only create a text corpus but ensure it shines a light on impact of women in society. **Validation of Text Corpus**:Common Voice gives power to contributors to decide what they consider an accurate representation of their language.In this instance any text that goes on the platform has to be from verified sources that have approved text to be CC0 but also adheres to the platform principles of inclusivity.The sentence collector accepts sentences but they go through a community drive validation which allows for any text that discriminates against demographics or is delegatory in nature to automatically be dismissed and not approved for public viewing. While this was a way to generate more text from a gendered lens it was also a learning experience about the role of text in fostering gendered biases in a dataset and eventually a language model.Here are excerpts of learning about the dynamics of languages and gender from a feminist perspective: **Gender of words(lexicon)**: Most often words are already gendered when they are written depending on the languages now in the case of swahili languages nouns are what differentiate males from females while the pronouns remain gender neutral.This is not the case for example in English where a chair leader is often referred to as chairman and not chairlady which as people become more gender aware has now changed to chairperson in efforts of being gender neutral.Understanding how text created carry gendered meanings behind them help shape how we approach creating feminist content that can be utilized in language models.Marion and levy(2023) in their paper argue that finding gender-neutral replacements for gendered words, such as *police officer* instead of *policeman* can help to assess gender-inclusive language strategies in a text and thus whether its authors support a binary conceptualization of gender. **Creation of gender inclusive text:**while different languages have different ways of expressing gender one method that can help address this is to strip gender identities to text specifically depending on your intended use.Targeting male-as-norm language replace it with more inclusive forms and, after training AI models on these more gender-inclusive texts, measure the impact of gender-based data curation on gender bias. Fostering the use of more inclusive language and no sexist language, the effect could also be absorbed by AI systems trained on gender-inclusive text. In the case of the writing competitions by Common voice this method was used where participants of the competition were urged to use inclusive language and due to the fact that swahili’s pronouns are neutral hence the text allowed for text data that is neutral. Data has a key role to play in eliminating gender biases in this particular case addressing the challenge from the root included looking at what text is being fed into the systems and how we can better biases at that level.The guiding principles of what kind of text is accepted on the platform also emphasis on this by being keen to ensure text is validated before its officially published on the platform.It is essential to take an approach of curating text data that does not replicate existing gender biases in the communities but rather demystify the concept of gender from a feminist approach of gender and social justice. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/image-3.png) ****Engendering Voice Principles** ### The Mozilla Data Collective Data Assistant is now available in Alpha URL: https://community.mozilladatacollective.com/the-mozilla-data-collective-data-assistant-is-now-available-in-alpha/ Last updated: 2026-05-06T18:31:17.000Z Finding the right dataset for your machine learning project should not take hours. Whether you are building a speech recognition system, training a text-to-speech model, or assembling a machine translation corpus, Mozilla Data Collective holds a substantial and growing library of high-quality, ethically sourced datasets. Until now, navigating that library has required patience and persistence. Today, we are pleased to introduce the Data Assistant: a conversational tool designed to find the right data for you. --- ## Why build the Data Assistant? Mozilla Data Collective exists to make diverse, community-contributed, diverse datasets accessible to researchers, developers, and organisations working on AI and machine learning-based technologies. But finding the right dataset for your project - dataset discovery - can be labourious. You may know the intended *task* you're searching for data to train - such as speech recognition or natural language processing, but not which datasets support it. You may know your target language, but not whether suitable recordings exist. You may refine your search several times before arriving at a meaningful result — and even then, you may be uncertain whether you have considered all the options. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/image-5.png) The Data Assistant aims to reduce that friction. By understanding your requirements in plain language and matching them against Mozilla Data Collective's dataset catalogue, it reduces the gap between "I need data for this project" and "here are the datasets you should look at." --- ## How do I use the Data Assistant? You do not need to know the exact name of a dataset, or the precise terminology used in its metadata. You simply describe what you are looking for, in natural language, and the Data Assistant will interpret your requirements and surface the most relevant datasets. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/image-4.png) The most effective searches include three elements: your **task**, your desired **language**, and any additional requirements around region or recording format. Here are some examples to get you started: - *"I need ASR training data for French, with multiple speakers and a range of recording environments."* - *"Can you help me find high-quality single-speaker audio for a Portuguese TTS system?"* - *"I'm looking for parallel text corpora for English to Swahili machine translation, ideally covering informal or conversational domains."* - *"What datasets are available for training a speech recognition model in a low-resource African language?"* You can also search by *region* if your project requires data from a specific geography, or by *modality* if you have constraints around audio quality, transcription format, or dataset size. The Data Assistant supports queries in multiple languages — if you prefer to search in French, Spanish, or another language, you are welcome to do so. While the data assistant supports queries in multiple languages, we have primarily tested in English and response quality may vary based on the language used. If your initial query is vague or missing key details, the Data Assistant will ask you one or two focused clarifying questions rather than returning a broad or unhelpful set of results. This ensures that the datasets you are recommended are genuinely relevant to your work. --- ## What guardrails does the Data Assistant use to keep me safe? The Data Assistant is built specifically for dataset discovery on the Mozilla Data Collective platform. It does *not* offer general-purpose assistance, and it does *not* draw on information outside the datasets indexed in our catalogue. This narrow focus is intentional: it ensures that every recommendation you receive is grounded in real, available data — and nothing more. No hallucinations here! Several specific safeguards are in place: **If your search falls outside the scope of dataset discovery**, such as a request for general technical advice or content unrelated to Mozilla Data Collective, you will receive a clear explanation of what the Data Assistant is designed to help with, and be guided back towards a relevant query. **If you search for sensitive content** — for example, requests involving personal or biometric data, data about children, or adult content — you will *not* receive a generated response. Instead, you will be redirected to our dataset search page and invited to contact our team directly. Some requests require a human conversation, and the Data Assistant will always route you appropriately. **If no datasets are found matching your search requirements**, you will be told so honestly, and guided on how to broaden or refine your criteria. The Data Assistant will not invent dataset names or fabricate results — if the data does not exist in our catalogue, you will be told that clearly. **If your query lacks sufficient detail** to return meaningful results, the Data Assistant will ask for clarification before proceeding, so that you are not presented with a list of loosely relevant options. ## How do we use the data you enter into the Mozilla Data Assistant? As the premier platform for ethical data and fair value exchange, we're committed to being open and transparent about how we use your data - and the Mozilla Data Assistant is no exception. Mozilla Data Assistant is an AI-powered chatbot, and it uses a large language model (LLM) made available through [together.ai](http://together.ai/?ref=community.mozilladatacollective.com) to help you discover datasets by searching and summarizing dataset listings and datasheet metadata. We change the underlying LLM from time to time depending on what our internal evaluation shows yields the best search results. When you use the Assistant, your messages - the prompts you type into the chatbot - and limited chat context needed to respond - are processed and transmitted to [together.ai](http://together.ai/?ref=community.mozilladatacollective.com) to generate an answer, and may be logged by Mozilla Data Collective for security, troubleshooting, and improving the feature, as described in our [Privacy Policy](https://mozilladatacollective.com/privacy?ref=community.mozilladatacollective.com). We don't sell your data to third parties, such as advertisers. However, as always with chatbots, please do not enter personal data, confidential information, or other sensitive content into the Mozilla Data Assistant. --- The Data Assistant is available now. Log in to your [Mozilla Data Collective](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com) account to get started, accept the usage consent, and ask your first question. We hope you enjoy it as much as we do, and we warmly [welcome your feedback](https://docs.google.com/forms/d/e/1FAIpQLSdEPD2mcuIix6TZaTCJcKY1zX22SK1hG%5FwYKR5VMm1IMr%5FjZg/viewform?ref=community.mozilladatacollective.com). ### How Open Licensing is Changing with AI: The NOODL License URL: https://community.mozilladatacollective.com/how-open-licensing-is-changing-with-ai-the-noodl-license/ Last updated: 2026-05-14T14:32:00.000Z *Author: Alek Tarkowski* For the last twenty five years, standardized open licenses were increasingly seen as a main tool for democratizing access to knowledge. Over less then a decade, a relatively narrow set of canonical choices emerged: [the Creative Commons licensing stack](https://creativecommons.org/licenses/list.en?ref=community.mozilladatacollective.com) [https://creativecommons.org/licenses/list.en](https://creativecommons.org/licenses/list.en?ref=community.mozilladatacollective.com), coupled with a few [Open Data Commons](https://opendatacommons.org/?ref=community.mozilladatacollective.com) licenses. The open movement focused on promoting these standardized licensing tools as means of ensuring access, with as little friction as possible, and with potentially great public value. Since the release of Creative Commons Zero in 2009, no major new licensing options have been designed. Further innovation in licensing did not occur, because it was not needed. And open advocates for at least a decade had a sense that the work of developing sharing frameworks is finished. Things have changed several years ago. The [reuse of 100 million openly licensed photographs for training early AI models](https://openfuture.eu/publication/ai-commons/?ref=community.mozilladatacollective.com) was the first warning sign. Since then, the emergence of AI-related uses of open content has triggered a new wave of licensing innovation. The [Nwulite Obodo Open Data License (NOODL)](https://licensingafricandatasets.com/?ref=community.mozilladatacollective.com) is one of such licensing experiments: a new license that combines open sharing with tiered use conditions. It is a license crafted to address the inequities identified with existing open licenses when used to distribute/share/disseminate African datasets. It is also an effort to address inequalities related to open sharing and content reuse that are increasingly visible and felt, across Africa and the Global South. The NOODL license should be of interest to all open practitioners, data governance experts and stewards of the commons, even though for now it is used to license just a single dataset: [DhoNam](https://mozilladatacollective.com/datasets/cmjepxo6t08nmmk07iauvua6v?ref=community.mozilladatacollective.com), a speech corpus for [Dholuo](https://en.wikipedia.org/wiki/Dholuo?ref=community.mozilladatacollective.com), an indigenous language from Kenya. It is important because it points to a shift in open licensing approaches. Instead of depending on a set of standardized licensing options you can design your own license, community-centered and tailored to local needs. And in doing so, you can aim for ensuring greater equity in a world where open sharing makes you increasingly prone to asymmetries of power. # **Making open licensing equitable** The NOODL license, and other recent licensing experiments - such as the Responsible AI Licenses (RAIL) - can be a source of unease to proponents of open sharing. First, because they put into question the canonical - by now - assumption that standardized open sharing is the “one size fits all” approach for ensuring access to knowledge. And second, because in the name of equity and responsible use, they introduce various conditions and limitations upon open access and reuse. In doing so, they challenge the assumption that more openness is always better. Naturally, open licensing always included several optional limitations (specifically, the Non-Commercial and No Derivatives conditions), but they were seen by many as results of problematic, unnecessary compromises. The new wave of licensing, of which NOODL is part, goes further, by introducing much stricter limitations. In the case of NOODL, free reuse is available only for those people and entities located in developing countries. For those in the developed world (based on OECD criteria), the licence requires meeting additional obligations aimed at “benefit sharing”: various forms of reciprocity, including payments. Other licenses, such as RAIL, introduce strict limits on types of allowed use, for example banning military uses. This is often met with the criticism that this is yet another enclosure of the commons, a shift from open principles towards proprietary approaches to intellectual property. This view might have been correct, if openness was the only principle that we should be caring about, as supporters of the commons. But that should not be the case. A different approach is proposed by the experts and communities that co-designed NOODL. It is an approach that balances openness with the need for equity and social justice. In the words of Dr. Melissa Omino, “It's not that they're saying they don't want to share or they don't want to be open. They're saying that the way open is existing right now is actually harming them.” And the reason for that are power asymmetries and concentrations that are experienced not just by Dholuo speakers, but across the world. There is [a paradox to open sharing](https://paradox.openfuture.eu/?ref=community.mozilladatacollective.com), as mechanisms that help challenge power are also prone to enabling its concentrations. Dr Chijioke Okorie, NOODL's lead author, explains that open licensing appears fair, but enabling commercial reuse without any form of reciprocity reinforces “long-standing power imbalances between well-resourced actors in the Global North and under-resourced researchers on the continent”. The largest global companies are best positioned to benefit from open resources, and to use them to further consolidate power. Sarah Pearson from Creative Commons recently [noted](https://creativecommons.org/2026/02/12/how-to-keep-the-internet-human/?ref=community.mozilladatacollective.com) that “We cannot respond by accepting these risks and harms as inherent and inevitable costs of public sharing knowledge”. When it comes to AI development, democratization must mean not just availability, but also broadening capacity to develop these systems. The aim of NOODL is not to share linguistic data with the world - but to create leverage, with which African communities can build their own AI tools, serving their own, local needs. What does this mean for open sharing? The standardized open licensing stack is a huge achievement of the access to knowledge movement. It is the backbone of many sharing solutions, and remains suitable in many conditions - for example, for sharing data and resources by public institutions. But they should be seen as just one part of a bigger toolbox of commons-based tools. And by paying more attention to the idea of the commons, we can focus on the role of collective decision making, and content of governance - which has been at the heart of the most successful free knowledge projects, like Wikipedia. Recently, Trebor Scholz and Mark Esposito proposed argue that we need [a solidarity ecosystem for AI](https://ssir.org/articles/entry/artificial-intelligence-solidarity-ecosystem?ref=community.mozilladatacollective.com), which combines sharing with cooperative ownership and equity as a foundational principle. The NOODL license fits well within this alternative AI development stack. Both share the assumption that knowledge layer should not just be open, but more importantly collectively governed, and thus community-centered. This is an important reformulation of the purpose of open sharing, which may be controversial to some. It’s important to understand that NOODL is neither a proprietary nor a closed license. The tiered access model introduces limitations and reciprocal mechanisms to ensure that sharing is equitable, and beneficial to the data community. # **The challenges ahead - and how to overcome them** Deployment of NOODL is not without challenges, which are mainly related to enforcement of additional licensing conditions. This is a major issue, and one shared by all other sharing mechanisms that aim to introduce conditionalities, limitations or forms of reciprocity. If a company violates the license terms, what recourse do communities have? This question is not unique to NOODL, but it is particularly acute for licenses centered on benefit-sharing with the community. For now, no obvious solutions have been identified. Instead, there is a sense of a “free for all” when it comes to use of openly shared content - as anything publicly available on the web becomes scraped, often with disregard for any norms or conditions. This is “permissionless innovation” taken to the extreme. Any way forward will require stronger collective norms and enforcement - stewards of the DhoNam dataset will not be able to enforce any rules on their own. Another major challenge relates to license incompatibility and the risk of fragmentation of the open ecosystem. Licensed tailored to needs of specific communities lose the advantage of standardized sharing, which was so important for the scaling of open licensing. One could argue, that NOODL, while focusing on the needs of the data community, paid less attention to the broader digital commons. Here, the solution lies in retaining some level of standardization of licensing options. Hopefully, reciprocal mechanisms proposed by NOODL can be adapted to other communities and contexts. > Both challenges – enforcement and integrity of the sharing ecosystem – point to the role of alternative sharing platforms, like [Mozilla Data Collective](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com). Open licensing frameworks were built ahead of the commons-based platforms that grew in the last two decades. Today, the commons is growing around platforms like Wikimedia, [Mozilla Data Collective](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com), HuggingFace, and various context repositories. It is the infrastructural capacity of these platforms that is just as important as the legal power of licensing tools. Authors of NOODL suggest the need to create an ecosystem that supports diverse licensing approaches—for example, a platform where the various licenses could be shared, explained, and explored by legal experts, data communities, and AI developers alike. Hopefully, the open sharing ecosystem will develop this capacity that, while centralized, would enable broader experimentation with types of data governance that establish a global knowledge commons, while supporting local communities. ### Exciting Updates for Mozilla Data Collective URL: https://community.mozilladatacollective.com/exciting-updates-for-mozilla-data-collective/ Last updated: 2026-04-21T10:00:36.000Z ### **Request to Access feature** We have always wanted to give communities more choice about who accesses their datasets, and the platform supports dozens of licenses, and full downloader authentication. But we also know that some organisations may want to check every downloader themselves, and we want to give them the tools to do that. With the new Request to Access feature, uploaders can require downloaders to request access to their dataset, and ensure that they’re comfortable with sharing with the specific requester. Imagine, your dataset is a multimodal medical corpus intended for research, or a children’s speech corpus intended only for education non-profits or local start-ups; now you are able to double-check that the downloader is aligned with your license intention. Until you, the uploader, approves the request, the platform won’t enable the dataset download. [This feature is available for uploaders to use today](https://community.mozilladatacollective.com/request-to-access-feature-is-now-available/)! ### **Data Assistant** Finding the right dataset for your machine learning project should not take hours. Whether you are building a speech recognition system, training a text-to-speech model, or assembling a machine translation corpus, Mozilla Data Collective holds a substantial and growing library of high-quality, ethically sourced datasets. Until now, searching through the datasets on MDC required you to understand what types of tasks and explicit languages you were searching for. The upcoming Data Assistant feature aims to reduce that friction. By understanding your requirements in plain language and matching them against Mozilla Data Collective's dataset catalogue, it reduces the gap between "I need data for this project" and "here are the datasets you should look at." The MDC Data Assistant will be available starting on May 06, 2026. ### **Payments and Compensation features** Datasets represent hours of labour, care and attention. Curating a multimodal dataset for an underserved community involves linguists, speakers, transcribers, annotators, quality assurance teams, not to mention tooling, hosting and downloader support. Many of our users open source their datasets, in an incredible act of gift and commons-building. But for many, this isn’t possible, or even fair or desirable: we have heard from communities who have lived experience of colonialism and exploitation, who insist that their labour and expertise must be valued and recognised. We’ve spoken to people working in industries whose traditional sustainability models are being cannibalised, and who need new revenue streams to ensure flourishing futures. We are excited that will be releasing a set of compensation and payment features, so that downloaders can pay for a license to use datasets. 100% of the license fee, which is set by the uploader, goes directly to the uploader. We never take a cut of the license fee. We charge the downloader a modest 5% fee to cover our infrastructure, hosting and support costs. We expect to launch this feature in the next few weeks. ### **Differentiated Access** One of our biggest motivations for building Mozilla Data Collective was enabling organisations to set different terms for different contexts. You might be happy to share a sample of your simulated doctor-patient dialogues dataset for free with students and researchers, but expect that a large pharmaceutical company who wants the entire dataset should pay for a license. Now, by stacking our payments and request to access features, you will be able to set up multiple listings and enable differentiated access to your dataset. With these features, you can share your dataset in line with your values. ### **A dedicated entity based in the UK** To power all this for our community, we knew we needed the capabilities and agility of a company, the trustworthiness and purpose of a non-profit, and the data protections of Europe. That’s why we’re excited to announce our new home in the United Kingdom! Mozilla Data Collective is now structured as a mission-locked British company; backed, incubated and governed by Mozilla Foundation, the non-profit that fights for alternative digital futures, and makes good tech the norm. We are a social enterprise, and proud to be so. In a world where grant funding can disappear in line with political realities, we want to be firmly self-sustaining and independent. These features are being released in Alpha, so as always we welcome your feedback, ideas, requests and suggestions for how to make them even better! Reach out at [support@mozilladatacollective.com](mailto:support@mozilladatacollective.com). ### Request to Access Feature is now Available URL: https://community.mozilladatacollective.com/request-to-access-feature-is-now-available/ Last updated: 2026-04-22T11:16:26.000Z We’re excited to release a new way for dataset uploaders on [Mozilla Data Collective ](https://mozilladatacollective.com/?utm%5Fsource=social&utm%5Fmedium=community&utm%5Fcampaign=conditional-access)to control how their data is shared. Until now, dataset access has been governed by a combination of license terms and the addition of additional restrictions as part of the consent-to-download flow. While this provides important legal and usage guardrails, it doesn’t allow for more granular, case-by-case controls. With this release, uploaders can now gate access to their datasets, requiring users to request permission and share their email address before downloading. ### **🔐 For Uploaders** ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/image.png) - **New Access Setting:** When creating a dataset listing, choose between: - **Public Access (default)** – anyone can download (with the existing agreement flow) - **Restricted Access** – users must request permission before downloading and agree to share their email address with the uploader - **Private** \- for temporarily withdrawing your dataset listing from the /datasets page - **Access Requests:** Review incoming requests from potential downloaders through your account. You will receive an email notification when a new access request is submitted. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/data-src-image-90781f16-5c3d-4d40-8a0b-f99b77532b62.png) - **Flexible Controls:** Approve or reject requests, with the ability to update your decision at any time ### **📥 For Downloaders** ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/image-1.png) - **Request Access Flow:** Gated datasets now display a **“Request Access”** button instead of “Download” - **Request Details:** When requesting access, users must provide: - Preferred name - Email address - Organization (from profile) - Optional message explaining intended use ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/04/image-2.png) The uploader of the dataset will be notified of your request and will be able to approve or reject it. - **Status Updates:** - **Pending Access** – request submitted - **Download** – access granted - **Download Unavailable** – access denied This is the first step towards a more granular access control system for organizations hosting their datasets on Mozilla Data Collective. In the future, we will introduce a wider range of available access conditions to better support the range of use cases that dataset owners have set around use of their data. ### Fine-Tune a Speech-to-Text Model for Any Language - Including Yours URL: https://community.mozilladatacollective.com/fine-tune-a-speech-to-text-model-for-any-language-including-yours/ Last updated: 2026-04-15T14:01:29.000Z Most speech recognition models were built with English, or a handful of well-resourced languages, in mind. If you speak Khmer, Galician, or any of the hundreds of languages underrepresented in mainstream AI, you've probably hit a wall trying to get accurate transcriptions. The Mozilla Data Collective (MDC) is working to change that. Our mission is a tech future that is **multilingual, multicultural, and multimodal** \- built on the principle that communities should have genuine data agency and be the primary beneficiaries of sharing their data. The[ speech-to-text-finetune](https://github.com/Mozilla-Data-Collective/speech-to-text-finetune?ref=community.mozilladatacollective.com) blueprint is a practical tool for making that a reality by enabling developers to easily fine-tune Whisper-based or MMS-based STT models. In this tutorial, we'll walk through how to fine-tune OpenAI's Whisper model on your own language using the Mozilla Data Collective platform's datasets or your own custom audio data. Everything runs locally — even on a laptop — keeping your data private. You can also [view a video walk-through of this tutorial here](https://community.mozilladatacollective.com/dev-talks/#fine-tuning-a-whisper-model-with-mdc-datasets). ## **What You'll Build** By the end of this tutorial, you'll have: - A fine-tuned Whisper model optimised for your target language - A local transcription app you can test against your own voice - (Optionally) a model pushed to the Hugging Face Hub for sharing As a concrete example of what's possible: a [Galician Whisper model fine-tuned with this blueprint](https://www.google.com/url?q=https://github.com/Mozilla-Data-Collective/speech-to-text-finetune?tab%3Dreadme-ov-file%23example-result-on-galician&ref=community.mozilladatacollective.com) produces dramatically cleaner transcriptions than the base `whisper-small` model, correctly handling language-specific phonology that the base model garbles. --- ## **Prerequisites** - Python 3.10+ - `ffmpeg` installed on your system - A Mozilla Data Collective API key (free — instructions below) - A GPU is helpful but not required for smaller models ## **Step 1: Clone the Repo and Install Dependencies** ```Shell git clone https://github.com/Mozilla-Data-Collective/speech-to-text-finetune cd speech-to-text-finetune pip install -e . ``` Then install ffmpeg: ```Shell # Ubuntu sudo apt install ffmpeg # macOS brew install ffmpeg ``` ## **Step 2: Get Your MDC API Key** The MDC platform hosts Common Voice datasets across a wide range of languages and provides a Python SDK for accessing them. To use it: 1. Create a free account at[ ](https://datacollective.mozillafoundation.org/api-reference?ref=community.mozilladatacollective.com)https://mozilladatacollective.com/ 2. Generate an API key from your account dashboard 3. Copy the `.env` template and add your key: `cp example_data/.env.example src/speech_to_text_finetune/.env` Then open the .env file and set: `MDC_API_KEY=your_api_key_here` **Note:** Values in this `.env` file will override any environment variables with the same name already set on your system. Join Mozilla Data Collective → ## **Step 3: Choose Your Dataset** You have two options for sourcing training data. ### **Option A: Load via the MDC Python SDK (Recommended)** This is the easiest path. Browse datasets on the[ MDC platform](https://mozilladatacollective.com/datasets?ref=community.mozilladatacollective.com) and find a dataset for your language. Open the dataset page, click "Connect API", and copy either the "Dataset ID" or the "Dataset slug", both can be used interchangeably. `dataset_id_or_slug="khmer-asr-cultural-dataset-4e33cd05" ` ↑ this is your dataset\_id **Important:** You must accept the dataset's terms and conditions on the MDC platform before downloading via the API. ### **Option B: Use a Locally Downloaded Dataset** Download and extract the dataset zip from the MDC platform, then point your config at the local directory. This is useful if you're working offline or want to pre-process the data yourself. ### **Option C: Bring Your Own Data** You can also record and label your own custom dataset using the built-in UI: `python src/speech_to_text_finetune/make_custom_dataset_app.py` The app walks you through recording audio clips and adding transcriptions. Your dataset should produce a `.csv`, `.tsv`, or `.parquet` file with `audio_path` and `transcription` columns. ## **Step 4: Configure Your Fine-tuning Run** Create or edit a config\_whisper.yaml file. Here's a minimal example using the MDC SDK path: ```Python model_id: openai/whisper-tiny dataset_id: cminc35no007no707hql26lzk   # Your MDC dataset ID language: Khmer repo_name: whisper-tiny-km-finetuned download_directory: /path/to/downloads  # Optional path to download the dataset training_hp:   push_to_hub: False   hub_private_repo: True   num_train_epochs: 10   per_device_train_batch_size: 16   learning_rate: 1e-5 ``` **Choosing a model size:** Start with whisper-tiny or whisper-small for faster iteration. Move to whisper-medium or whisper-large-v3-turbo once you're happy with the pipeline and have a GPU available. **Unsupported languages:** If your language wasn't part of Whisper's original training data, don't give up. Use[ Glottolog](https://glottolog.org/?ref=community.mozilladatacollective.com) to find the closest related language that *is* in Whisper's supported list, and set that as your language value. This gives the decoder more appropriate token priors. If no related language appears in Whisper's list, set language: None. ## **Step 5: Fine-tune** `python src/speech_to_text_finetune/finetune_whisper.py` The script handles dataset loading, train/test splitting, tokenization, and training. If your dataset already defines official train and test splits (as Common Voice does), those are preserved automatically and `test_size` in your config is ignored. Training progress will be logged to the console. ## **Step 6: Test Your Model** Once training completes, fire up the transcription app to test your fine-tuned model against your own voice or a sample audio file: `python demo/transcribe_app.py` Enter the local path to your fine-tuned model (or a Hugging Face model ID if you pushed it), record a sample, and see the transcription. To compare your fine-tuned model side-by-side against the base model: `python demo/model_comparison_app.py` This makes it easy to see exactly where fine-tuning improved accuracy on your language's specific phonology and vocabulary. ## **Step 7 (Optional): Speed Up Inference with Faster-Whisper** Once you're happy with your model, you can convert it to the `faster-whisper` format for significantly faster inference — useful if you're deploying this in a production app or on lower-powered hardware. Check out the tutorial by Kathy from the MDC community: [tutorial-whisper-fine-tuning-australian-EO2026](https://github.com/Mozilla-Data-Collective/tutorial-whisper-fine-tuning-australian-EO2026?ref=community.mozilladatacollective.com) ## **Try It Without Installing Anything** Not ready to set up a local environment? You have some other instant options: - **GitHub Codespaces:**[ Launch a ready-to-go environment](https://github.com/codespaces/new?hide%5Frepo%5Fselect=true&ref=main&repo=mozilla-ai/speech-to-text-finetune&skip%5Fquickstart=true&machine=standardLinux32gb) with everything pre-installed to run small scale experiments in a limited resources machine. - **Google Colab:**[ Open the fine-tuning notebook](https://colab.research.google.com/github/mozilla-ai/speech-to-text-finetune/blob/main/demo/notebook.ipynb?ref=community.mozilladatacollective.com) which includes a full MDC flow for Khmer as a worked example (`demo/mdc_khmer.ipynb`) that can be run with free GPU access provided by Google. ## **What's Next?** - Discover and play with new datasets [https://mozilladatacollective.com/datasets](https://mozilladatacollective.com/datasets?ref=community.mozilladatacollective.com) - Share your own fine-tuned model and contribute back to the collection - Join the conversation on[ Discord](https://discord.gg/YuMNeuKStr?ref=community.mozilladatacollective.com) (`#mozilla-data-collective`),[ Reddit](https://reddit.com/r/MozillaDataCollective?ref=community.mozilladatacollective.com), or[ LinkedIn](https://linkedin.com/company/mozilla-data-collective?ref=community.mozilladatacollective.com) *The `speech-to-text-finetune` blueprint is open source under the Apache 2.0 License. It was created by Mozilla.ai and adapted by the Mozilla Data Collective community.* ### Datasheets: The Missing Manual for your Dataset URL: https://community.mozilladatacollective.com/datasheets-the-missing-manual-for-your-dataset/ Last updated: 2026-04-17T13:41:49.000Z 0:00 /3:51 1× When you [upload your dataset to Mozilla Data Collective](https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform/), creating a comprehensive datasheet is a critical part of the process. But what exactly is a datasheet? And how should you think about filling it out? Creating a comprehensive datasheet adds a lot more value to your dataset. It contextualizes the data that you are sharing, so that downloaders understand how it is intended to be used. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/image-2.png) The data card is the first preview that a downloader will have of your dataset. By being specific and clear in these fields, you can help people discover your dataset more easily. Once a potential downloader has found their way to your dataset, your datasheet will help them understand whether it's a good fit for their particular use case. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/image-3.png) The datasheet for the [Jember Javanese Spontaneous Speech Corpus](https://datacollective.mozillafoundation.org/datasets/cmlgm5a94008kny07nz2intus?ref=community.mozilladatacollective.com) on Mozilla Data Collective While only some fields in the dataset onboarding form are required, a thorough datasheet is a concrete tool to help individuals and communities share and govern their data. Join Mozilla Data Collective → ### Additional Resources [Turning Your Data Into a Valuable ML Resource Without Giving Up Control](https://community.mozilladatacollective.com/turning-your-data-into-a-valuable-ml-resource-without-giving-up-control/) \- This guide introduces the principles behind ethical data sharing, ownership, and documentation for community dataset creators [Datasheets for Datasets](https://www.microsoft.com/en-us/research/project/datasheets-for-datasets/?ref=community.mozilladatacollective.com) (Microsoft Research) - Accessible overview of the datasheet concept plus downloadable templates [Data Statements](https://techpolicylab.uw.edu/data-statements/?ref=community.mozilladatacollective.com) (UW Tech Policy Lab) - Practical framework specifically designed for language and speech datasets, including schema elements covering speaker demographics, annotator demographics, recording quality, and more. [Augmented Datasheets for Speech Datasets](https://github.com/SonyResearch/project%5Fethics%5Faugmented%5Fdatasheets%5Ffor%5Fspeech%5Fdatasets?tab=readme-ov-file&ref=community.mozilladatacollective.com) (Sony, GitHub Repo) - Sony Research's companion repo to the FAccT 2023 paper, containing the augmented speech datasheet template and completed example datasheets for five datasets including LibriSpeech and Common Voice [Datasheets for Datasets](http://arxiv.org/abs/1803.09010?ref=community.mozilladatacollective.com) Gebru et al. (2021) - The original paper proposing the datasheet framework, published in Communications of the ACM. Essential reading for understanding the motivation, design decisions, and scope of the datasheet standard [Data Statements for NLP](http://aclanthology.org/Q18-1041?ref=community.mozilladatacollective.com) Bender & Friedman (2018) - Proposes a complementary framework specifically for NLP/speech datasets, including speaker demographics, language variety (BCP-47 tags), recording quality, and speech situation; directly applicable to ASR and Common Voice datasets [The Dataset Nutrition Label](https://arxiv.org/abs/2201.03954?ref=community.mozilladatacollective.com) (2nd gen) Chmielinski et all (2022) - Offers an overview of the 2020 version of the Dataset Nutrition Label The Data Nutrition Project is empowering data practitioners and policymakers with tools to improve AI outcomes. [Learn more about the Data Nutrition Project ](https://datanutrition.org/?ref=community.mozilladatacollective.com) [Augmented Datasheets for Speech Datasets and Ethical Decision-Making](http://arxiv.org/abs/2305.04672?ref=community.mozilladatacollective.com) Papakyriakopoulos et al. (2023) - Extends Gebru et al.'s framework with speech-specific questions covering language diversity, accent, dialect, speech impairment, data subject protection, and speaker compensation. Grounded in a large literature review of SLT datasets. Published at ACM FAccT 2023\. The most directly relevant paper for documenting ASR datasets such as Common Voice [Data Statements: From Technical Concept to Community Practice](http://dl.acm.org/doi/full/10.1145/3594737?ref=community.mozilladatacollective.com) McMillan-Major, Bender & Friedman (2024) - Empirical follow-up on how practitioners actually use data statements, with refined schema (v3) and community-developed best practices [Navigating Dataset Cards on Hugging Face](http://arxiv.org/abs/2401.13822?ref=community.mozilladatacollective.com) (Research Analysis, 2024) - Large-scale empirical analysis of 7,400+ dataset cards on the Hub; reveals what documentation fields are most/least completed and what high-quality cards look like in practice --- ### Animation Credits Music by: Ricky Valadez Written by: Sarah Newman & Jessica Yurkofsky Illustrated by: Jessica Yurkofsky Read by: Sarah Newman Sound editing by: Halsey Burgund Special thanks to Liv Erickson, Katherine Reid, and Francis Tyers Produced by the [Data Nutrition Project](https://datanutrition.org/?ref=community.mozilladatacollective.com) in collaboration with [Mozilla Data Collective](https://mozilladatacollective.com/?utm%5Fsource=ghost&utm%5Fmedium=blog&utm%5Fcampaign=datasheet-video) ### Upcoming Domain Change on 09 April URL: https://community.mozilladatacollective.com/upcoming-domain-change/ Last updated: 2026-04-03T12:20:35.000Z **Hello, Mozilla Data Collective!** 👋 On April 9, we'll be transitioning the platform to a new domain: [mozilladatacollective.com](https://mozilladatacollective.com/?ref=community.mozilladatacollective.com)! While our original domain will redirect after the domain transfer takes place, keep an eye out for the following changes. - Email addresses sent from Mozilla Data Collective will now come from no-reply@mozilladatacollective.com - If you are in contact with members of the Mozilla Data Collective team, individual team emails will come from @mozilladatacollective.com, rather than @mozillafoundation.org. - Our support email address is changing to support@mozilladatacollective.com Dataset listings will now shared under the new domain, but the dataset slug should remain the same: *Before*: [https://datacollective.mozillafoundation.org/datasets/cmndapwry02jnmh07dyo46mot](https://datacollective.mozillafoundation.org/datasets/cmndapwry02jnmh07dyo46mot?ref=community.mozilladatacollective.com) *After*: [https://mozilladatacollective.com/datasets/cmndapwry02jnmh07dyo46mot](https://mozilladatacollective.com/datasets/cmndapwry02jnmh07dyo46mot?ref=community.mozilladatacollective.com) ## Action Required for using the MDC API For developers, you will need to take the following actions once the domain transfer is complete: - Updating API requests to use the new URL: [https://mozilladatacollective.com/api](https://mozilladatacollective.com/api?ref=community.mozilladatacollective.com) - Updating your version of the [Python SDK](https://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com) so that it uses the new URL endpoints This change is coming amidst other exciting features - stay tuned for an action-packed April! 🚀 Join Mozilla Data Collective → ### MDC Release Notes - 30.03.26 URL: https://community.mozilladatacollective.com/mdc-release-notes-30-03-26/ Last updated: 2026-03-30T17:04:25.000Z Hello, Mozilla Data Collective! 👋 It's been an exciting and busy time over here as we've gotten a few ✨major ✨ updates coming your way next month. Last week, we landed a core piece of work into the code base that will unlock our first conditional access feature - individual access gating. This feature will allow uploaders to have more direct control over who is allowed to download their datasets. We've also been iterating on our [Python SDK](https://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com) as we prepare to get uploads and submissions supported in a programmatic pipeline. ### New Features & Changes - Improvements and fixes related to uploading datasets, especially large ones - We fixed a bug where draft datasheets couldn't be saved until all required fields were entered - SO much stuff I want to share now but have to wait until they're publicly available 😉 Join Mozilla Data Collective → ### New Datasets **Anjuman e Katib** [Persian Literature Corpus by Najwai Sukhan | Mozilla Data CollectiveThe Persian Literature Corpus by Najwai Sukhan is a curated collection of Persian (Farsi) literary and educational texts created for research, computational use, and cultural preservation. It contains about 1.26 million tokens across 20 complete works spanning classical literature, poetry, modern prose, educational writing, philosophy, translations, and culturally rooted creative texts. Originally compiled in Microsoft Word format, the corpus was cleaned, normalized, and converted into UTF-8 plain text while preserving original orthography and style. Each file represents a complete work, making the dataset useful for both individual text analysis and broader corpus-level study. The corpus supports corpus linguistics, literary studies, digital humanities, NLP, and Persian language preservation.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-71.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-70.png)](https://datacollective.mozillafoundation.org/datasets/cmn3hs01700e0mb07n3y7brmz?ref=community.mozilladatacollective.com) **Balochistan Educational and Cultural Organization** [BECO Brahui Literature Corpus | Mozilla Data CollectiveThis Brahui literary corpus contains short stories, novels, and other creative literary works, representing a broad range of narrative styles and themes within Brahui literature. The texts reflect both classical and contemporary writing, offering insight into cultural expression and linguistic variation in Brahui. The corpus comprises approximately 355,000 tokens, making it a valuable resource for linguistic research and natural language processing tasks involving an under-resourced language.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-74.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-73.png)](https://datacollective.mozillafoundation.org/datasets/cmmus79780030ms07vyi4dsr5?ref=community.mozilladatacollective.com) **EELLAK - GreekFOSS** [Istorima | Mozilla Data CollectiveDataset Language: Greek Dataset Info: This dataset consists of oral history content collected from the Istorima archive, including transcribed interviews and associated metadata. The material reflects personal narratives and life stories, primarily in Greek, covering a wide range of social, cultural, and historical topics. Metadata Info: This dataset consists of 13,548 oral history interview records, structured as a tabular dataset with mixed data types. Each record includes a unique identifier (id) along with textual fields such as title, summary, transcription, speaker\_name, and researcher\_name. Additional metadata fields capture thematic and categorical information (themes, tags), geographic references (geonames, interview\_place), and temporal attributes (date, published\_at). The dataset also includes numerical and boolean features such as duration\_minutes, is\_age\_restricted, and is\_on\_demand, as well as a language field indicating the interview language. Dataset Statistics: Words: 96,479,186 Tokens: 138,933,365![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-73.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-72.png)](https://datacollective.mozillafoundation.org/datasets/cmmxib8ic00v6nw07bjrci8vj?ref=community.mozilladatacollective.com) **Institute of African Digital Humanities** [Bamun-French Parallel Corpus 2.0 | Mozilla Data CollectiveThis dataset is an extended and updated version of the ‘Bamun-French Parallel Corpus 1.1’ that is published on the Mozilla Data Collective platform. It is a parallel corpus of 4,444 lines in Bamun and French suitable for machine translation tasks. The text was obtained by transcribing raw audio files. Translations were added to enrich the original corpus. Bamun and French text alignment was performed in the process of creating this dataset. This version of the dataset resolves formatting issues flagged in the original and nearly doubles the number of aligned translation units compared to version 1.1.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-58.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-57.png)](https://datacollective.mozillafoundation.org/datasets/cmn6ay1xg016jnv07drtmw7qo?ref=community.mozilladatacollective.com) **Institute of Finno-Ugric/Uralic Studies, University of Hamburg** [INEL Kalmyk Speech Corpus | Mozilla Data CollectiveThis dataset is a machine-learning-ready subset of the INEL Kalmyk Corpus (Version 1.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. The dataset comprises 3 hours and 15 minutes of aligned supervised speech data (1,934 individual clips) across 26 speakers. It strictly populates the primary ‘sentence’ column using the ‘ts’ tier (scientific transcription) to ensure phonetic accuracy, with demographic metadata included where available.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-62.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-61.png)](https://datacollective.mozillafoundation.org/datasets/cmn4kxlaj001xnz07a5yugnew?ref=community.mozilladatacollective.com) [INEL Nganasan Speech Corpus | Mozilla Data CollectiveThis dataset is a machine-learning-ready subset of the INEL Nganasan Corpus (Version 1.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. The dataset comprises 38 hours and 30 minutes of aligned supervised speech data across 42 speakers. It features demographic metadata where available, and prioritizes Cyrillic transcriptions while falling back to Latin or Phonological tiers to ensure complete text coverage for acoustic modeling.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-63.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-62.png)](https://datacollective.mozillafoundation.org/datasets/cmn4kxhp8001tnz07jyfvf2qt?ref=community.mozilladatacollective.com) [INEL Evenki Speech Corpus | Mozilla Data CollectiveThis dataset is a machine-learning-ready subset of the INEL Evenki Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. The dataset comprises 2 hours and 39 minutes of aligned supervised speech data (2,180 individual clips) across 6 speakers. It features demographic metadata where available, and prioritizes Cyrillic transcriptions while falling back to Latin or Phonological tiers to ensure complete text coverage for acoustic modeling.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-64.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-63.png)](https://datacollective.mozillafoundation.org/datasets/cmn4kxexu001pnz07kn6wr985?ref=community.mozilladatacollective.com) [INEL Dolgan Speech Corpus | Mozilla Data CollectiveThis dataset is a machine-learning-ready subset of the INEL Dolgan Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. The dataset comprises 13 hours and 5 minutes of perfectly aligned supervised speech data (10,609 individual clips) across recordings spanning from the 1970s to 2017\. It features demographic metadata where available, and prioritizes Cyrillic transcriptions while falling back to Latin or Phonological tiers to ensure complete text coverage for acoustic modeling.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-65.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-64.png)](https://datacollective.mozillafoundation.org/datasets/cmn4kqzzt0013nu07caxllg3t?ref=community.mozilladatacollective.com) [INEL Kamas Speech Corpus | Mozilla Data CollectiveThis dataset is a machine-learning-ready subset of the INEL Kamas Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. The dataset comprises 13 hours and 51 minutes of aligned supervised speech data (13,197 individual clips) across 6 speakers, representing the entirety of available recorded spoken Kamas data. It features demographic metadata where available, and prioritizes Cyrillic transcriptions while falling back to Latin or Phonological tiers to ensure complete text coverage for acoustic modeling.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-66.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-65.png)](https://datacollective.mozillafoundation.org/datasets/cmn4kq2fl001dnz078d6n7r9a?ref=community.mozilladatacollective.com) [INEL Selkup Speech Corpus | Mozilla Data CollectiveThis dataset is a machine-learning-ready subset of the INEL Selkup Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. The dataset comprises 1 hour and 39 minutes of aligned supervised speech data (1,286 individual clips) across 15 speakers, largely originating from the 1960s and 1970s archive of linguist Angelina Kuzmina. It features demographic metadata where available, and prioritizes Cyrillic transcriptions while falling back to Latin or Phonological tiers to ensure complete text coverage for acoustic modeling.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-67.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-66.png)](https://datacollective.mozillafoundation.org/datasets/cmn4kpt540019nz07sapke78o?ref=community.mozilladatacollective.com) [INEL Nenets Speech Corpus | Mozilla Data CollectiveThis dataset is a machine-learning-ready subset of the INEL Nenets Corpus (Version 1.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. The dataset comprises over 36 minutes of aligned supervised speech data (447 individual clips). It strictly populates the primary \\‘sentence\\’ column using the \\‘st\\’ tier (Cyrillic source transcription) to ensure orthographic accuracy, with demographic metadata included where available.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-68.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-67.png)](https://datacollective.mozillafoundation.org/datasets/cmn4kmw70000nnu07vwcwfdi5?ref=community.mozilladatacollective.com) [Corpus de llenguatge ofensiu en català | Mozilla Data CollectiveThis dataset consists of sentences tagged as offensive-language in the version 25.0 release of Mozilla Common Voice in Catalan. The sentences are provided with the aim that they promote the development of offensive language detection in Catalan.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-69.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-68.png)](https://datacollective.mozillafoundation.org/datasets/cmn4s1j5d0091nu07e1hgzwgn?ref=community.mozilladatacollective.com) **LocaleNLP** [English Hausa Parallel Corpus | Mozilla Data CollectiveThis English–Hausa Parallel Corpus is a curated bilingual dataset of 5,000 aligned sentence pairs, translated from English into Hausa and organized into a clean sentence-level format to ensure reliable alignment. The dataset is designed to support machine translation training and evaluation, bilingual lexicon development, and broader linguistic and natural language processing (NLP) research for Hausa, including data-driven language technology development.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-70.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-69.png)](https://datacollective.mozillafoundation.org/datasets/cmn3ht40i00eami07lgydmrgg?ref=community.mozilladatacollective.com) [![CTA Image](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/discord-logo-discord-icon-transparent-free-png.png)](https://discord.gg/cs9tJPqQB5?ref=community.mozilladatacollective.com) The Mozilla Data Collective is now on the AI @ Mozilla Discord Server. Join us for announcements, community events, and more! [Join us on Discord ](https://discord.gg/cs9tJPqQB5?ref=community.mozilladatacollective.com) **Mozilla Common Voice** ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/image-4.png) The Mozilla Common Voice team has released the Spontaneous Speech 3.0 datasets and Scripted Speech 25.0 datasets on Mozilla Data Collective. You can find all of the Common Voice datasets available on the Common Voice organization page. [Common Voice | Mozilla Data CollectiveCommon Voice is a free, open source platform for community-led data creation. Anyone can preserve, revitalise and elevate their language by sharing, creating and curating text and speech datasets.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-57.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-56.png)](https://datacollective.mozillafoundation.org/organization/cmfh0j9o10006ns07jq45h7xk?ref=community.mozilladatacollective.com) **UP EEEI - Digital Signal Processing Laboratory** [UP - DSP - Philippine Languages Database (UP-DSP-PLD) | Mozilla Data CollectiveThis dataset contains multilingual, text and speech pairs for ten Philippine languages namely Filipino, English, Cebuano, Kapampangan, Hiligaynon, Ilokano, Bikolano, Waray, and Tausug. The dataset contains over 454 hours of recordings, covering multiple domains in news, medical, education, tourism and spontaneous speech. The applicability of the corpus has also been demonstrated in adult and children ASR, phoneme transcriber, voice conversion, and TTS applications.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-76.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-75.png)](https://datacollective.mozillafoundation.org/datasets/cmmxhw46c00tqnw07xyr94zjk?ref=community.mozilladatacollective.com) **MDC Curators** [Corpus de llenguatge ofensiu en català | Mozilla Data CollectiveThis dataset consists of sentences tagged as offensive-language in the version 25.0 release of Mozilla Common Voice in Catalan. The sentences are provided with the aim that they promote the development of offensive language detection in Catalan.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-60.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-59.png)](https://datacollective.mozillafoundation.org/datasets/cmn4s1j5d0091nu07e1hgzwgn?ref=community.mozilladatacollective.com) **Community** [Araina Text Corpus (Occitan Aranese) | Mozilla Data CollectiveThis text corpus includes sentences from three sources. Public domain literary texts translated by Antòni Nogués. Sourced from institutestudisaranesi.cat, Language educational material by Jordi Suïls Subirà, Administrative proceedings from Conselh Generau d’Aran.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-59.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-58.png)](https://datacollective.mozillafoundation.org/datasets/cmn4xrzyx00ednz074sa8dwkp?ref=community.mozilladatacollective.com) [Heroes English-Spanish Dubbed Movie Speech Corpus | Mozilla Data CollectiveHeroes corpus contains mapped bilingual (English and Spanish) speech segments from the TV series Heroes. It contains 7000 single speaker speech segments extracted from the original and Spanish dubbed version of 21 episodes. Audio segments are accompanied with subtitle transcriptions and word-level prosodic/paralinguistic information. Each episode directory contains word-level and segment-level information of the whole episode and also parallel samples extracted under segments\_eng and segments\_spa subdirectories. Each sample is stored as a wave audio file, text file and a csv file containing word timing information and word-level paralinguistic and prosodic features (speaker id, mean f0, mean intensity).![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-72.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-71.png)](https://datacollective.mozillafoundation.org/datasets/cmn3fheeo00cjmb078qc8xr62?ref=community.mozilladatacollective.com) [Oro\_Word | Mozilla Data CollectiveThis dataset contains word-level recordings in Afaan Oromoo collected from native speakers to support the development of open-source speech technologies. The dataset is designed for training and evaluating automatic speech recognition (ASR) and text-to-speech (TTS) systems. Each audio file is paired with its corresponding written word and metadata. Afaan Oromoo is a widely spoken Cushitic language in Ethiopia and neighboring regions, but it remains underrepresented in digital language resources. This contribution aims to expand accessible linguistic data, support research and education, and strengthen the presence of Afaan Oromoo in modern AI technologies.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-61.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-60.png)](https://datacollective.mozillafoundation.org/datasets/cmn4l2u61001nnu07wtrkoqe4?ref=community.mozilladatacollective.com) [Urdu Multi-Speaker TTS Dataset | Mozilla Data CollectiveThis dataset is an Urdu text-to-speech corpus designed for speech technology development and related computational research. It contains approximately 10 hours of speech from 3 speakers, including 2 male and 1 female speaker. The data is distributed across 36 zip files, and each zip file includes a folder of audio files along with a CSV file that maps each audio file to its corresponding transcript. The recordings are drawn from the domains of newspaper, literature, and articles, providing a mix of formal, narrative, and informational language suitable for Urdu TTS, corpus creation, and speaker-based speech modeling.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-75.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-74.png)](https://datacollective.mozillafoundation.org/datasets/cmmvykcrs0050ny07vkwww5gi?ref=community.mozilladatacollective.com) ### On contributing my Thorsten-Voice voice datasets URL: https://community.mozilladatacollective.com/on-contributing-my-thorsten-voice-voice-datasets/ Last updated: 2026-03-26T10:21:24.000Z *Autor: Thorsten Müller* *(*[*Sie können diesen Beitrag auch auf Deutsch lesen*](https://community.mozilladatacollective.com/on-contributing-my-thorsten-voice-voice-datasets-de/)*.)* Was it my original idea to publish my voice under a CC0 license for everyone to use freely? Honestly: no. In fact, it was nearly quite the opposite. When I started diving deeper into speech technology in 2019, I had a very different goal in mind. I wanted to build my own voice assistant. But locally, without any cloud dependency. I didn’t want a microphone sitting in my home that constantly communicates with some server on the internet and streams my voice data somewhere else. Given how sensitive voice data and personal talk is, this dependency on US-based cloud services felt uncomfortable to me. My fascination with voice interaction, however, goes much further back. As a teenager in the 1990s, watching shows like Knight Rider or Star Trek, I was already fascinated by the idea of humans talking to machines and machines responding. Back then, it was pure science fiction and Hollywood stories. Decades later, it suddenly became real. So I started exploring open-source projects that could fit my idea, including Mycroft. In that environment, I also met people who are deeply committed to open technologies and data. Among them [Kathy Reid](https://community.mozilladatacollective.com/author/kathyreid/), who has been advocating for open systems for many years and is now active in [Mozilla Data Collective](https://datacollective.mozillafoundation.org/?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten). Technically, everything was exciting. But the quality of speech or TTS (text to speech) voices, especially in German, was not great. Those classic eSpeak-style voices: fast, efficient, but completely robotic. Fine for testing, but nothing you would actually want to listen to every day. At some point while reading Mycroft documentation, I came across the idea that I could train a synthetic voice using my own recordings. If you read that today, in 2026, you might think. Just record a few seconds and you’re done. Back in 2019, it was a very different story. The recommendation was: record at least 16 hours of clean, neutral audio. So I started recording. Evenings. Weekends. Sentence by sentence. Month after month. And despite all the motivation and excitement I made a lot of mistakes. I used a cheap USB headset instead of a high quality one. I tried to speak as “clean” as possible and ended up losing any natural flow. Simple sentences sounded more like an exaggerated news anchor than a human being. After thousands of recordings (around 10.000) I trained my first model. My computer was running for days. And the result? Well... you could kind of recognize my voice. But it was far from good. Alongside the speech, there was noise, humming, and echo. So I asked the Mycroft community for help and two interesting things happened. First, there was real interest. German-speaking community members asked whether I planned to release the recordings or the trained voice. Interest? In my voice? Even though I know there are many better voices out there. That felt... surprisingly good. And second, something much more important happened. Dominik Kreutz reached out and offered to analyze my recordings. His message was simple: “Send me your recordings, I’ll take a look.”And I thought: wait a second. The whole reason I started this project was to avoid putting my voice on the internet. And now I’m about to send it to a stranger from an online community? I had to think about that for a few days. In the end, I made a conscious decision: I chose to trust. So I sent him the data. And i never had to regret this decision. And his feedback was tough, but honest: the recordings were not good. Some might be salvageable, but most were not usable. That hit hard! I had spent months recording and now I had to accept that much of it was basically useless. Only when I listened at full volume did I notice the real issues: noise, interference, inconsistent setup, unnatural speech patterns. That was the moment I learned one of the most important lessons in this whole journey: Shit in, shit out. Or more politely: the quality of your data defines the quality of your results. If your input data contains noise the machine learning will take that serious and reproduce that noise in TTS output. At that point, I had a choice: stop or start over. The tech enthusiast in me was already hooked, so I kept going. I bought better equipment, built a small recording booth with wood, carpet and acoustic foam panels, and spent the next months of free time recording again. At the same time, I kept sharing progress with the community and realized that the interest in open German speech data was real. And then came the big question. What do I do with all of this? Do I publish nothing? Do I publish only a trained model? Or do I publish the raw data as well? And if I publish it. With restrictions or completely open? I was very aware of what this meant. Speech technology was clearly going to become more important that was obvious even then. And I knew if I release my voice, I give up control. Maybe I even limit future options for myself. For example, voice-based authentication systems e.g. opening doors, accessing sensitive systems all of that becomes questionable if your voice is publicly available. And of course, the obvious concerns came up. What if someone uses my voice for things I strongly disagree with? Political content? Extremism? Fraud? I didn’t feel fear, but I definitely felt respect for these risks. After thinking about it for a few days and discussing it with family and friends, I made my decision. If I do this I do it properly. [CC0](https://creativecommons.org/publicdomain/zero/1.0/legalcode?ref=community.mozilladatacollective.com)! No restrictions. Just like Mozilla [Common Voice](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com). I didn’t want to exclude anyone. Not research, not commercial use. I didn’t want to start adding “yes, but only if...”. If open then truly open. Looking back now, it’s still surprising to me what has grown out of “[Thorsten-Voice](https://datacollective.mozillafoundation.org/datasets/cmm4de9w500ntmh073nx14k7p?ref=community.mozilladatacollective.com)” over the years. Both in very positive ways and in some uncomfortable ones. At one point, someone sent me a video from the so-called “Reichsbürger” scene in Germany. My voice was used in a context that I personally strongly reject. That was the moment when a theoretical risk became reality. Not a great feeling. And still I have never regretted my decision. Because the positive impact clearly outweighs it. I have received so many encouraging messages. A computer science teacher in Berlin told me that his students can now build real speech systems locally, without cloud dependency. The Swiss “Lernstick” uses voices like this for accessibility in education. People use it in smart home setups. Someone told me my voice is speaking from the ceiling of a house on a finca in Mallorca. And there are use cases in screen readers supporting people with visual or reading impairments. These are the moments where you realize: it actually makes a difference. I always include a personal note with my datasets. Not as a restriction but as a personal statement, as I cannot control what is done and said with my voice. But I can communicate what I as a person stand for. > "I believe that all people are equal, regardless of gender, sexual orientation, religion, skin color, or where they were born. I believe in a global world where everyone is welcome. And I believe that knowledge and education should be freely accessible to everyone. And I believe that, as humans, we are capable of achieving amazing things if we trust each other." Today, the situation has evolved even further. Creating synthetic voices has become much easier. You no longer need 16 hours of audio. In many cases, a few seconds are enough. A simple voice message can be sufficient. This creates both opportunities and challenges. Which is exactly why I believe open, transparent, and ethically sourced datasets are so important and why platforms like [Mozilla Data Collectiv](https://datacollective.mozillafoundation.org/?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten)[e](https://datacollective.mozillafoundation.org/?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten) matter. Many modern AI systems have been trained on data where the origin is unclear and consent is questionable. Open data provides a real alternative. Join Mozilla Data Collective → Organizations like Mozilla have been laying the groundwork for this for years with projects like [Common Voice](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com) and now the [Data Collective](https://datacollective.mozillafoundation.org/?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten). Over time, I’ve had the chance to connect with people in this space who are deeply committed to openness and responsible data practices. Maybe this short story encourages others to contribute in their own way. I can say: it doesn’t hurt. Looking back now, after several years, I can say with full conviction: I have never regretted donating my voice. I would do it again. Despite the risks. Because I strongly believe: If we trust each other more, if we share knowledge, if we collaborate openly, then we can achieve a lot — together. And that is exactly why I chose to share my voice. ### On contributing my Thorsten-Voice voice datasets [DE] URL: https://community.mozilladatacollective.com/on-contributing-my-thorsten-voice-voice-datasets-de/ Last updated: 2026-03-26T10:20:34.000Z *Autor: Thorsten Müller* *(*[*You can also read this post in English*](https://community.mozilladatacollective.com/on-contributing-my-thorsten-voice-voice-datasets/)*)* War es meine ursprüngliche Idee, meine Stimme unter der freigiebigen Open-Source CC0-Lizenz zu veröffentlichen, damit sie jeder ohne Einschränkungen nutzen kann? Ehrlich gesagt: nein. Tatsächlich war es fast das Gegenteil. Als ich 2019 begann, mich intensiver mit Sprachtechnologie zu beschäftigen, hatte ich ein ganz simpel klingendes Ziel vor Augen. Ich wollte meinen eigenen Sprachassistenten entwickeln. Aber lokal (ohne Cloud-Abhängigkeit). Ich wollte kein Mikrofon in meiner Wohnung, das ständig mit irgendeinem Server im Internet kommuniziert und meine Sprachdaten austauscht. Angesichts der Sensibilität der Daten und persönlichen Gesprächen im eigenen Zuhause fühlte sich diese intransparente Abhängigkeit von (US-) Cloud-Diensten für mich unangenehm an. Meine Faszination für Sprachinteraktion mit Technologie reicht jedoch viel weiter zurück. Als teenager in den 1990er-Jahren, als ich Serien wie Knight Rider oder Star Trek sah, war ich bereits fasziniert von der Idee, dass Menschen ganz natürlich mit Technologie sprechen und Maschinen auch antworten. Damals war das reine Hollywood Science-Fiction. Jahrzehnte später wurde es Realität. Also begann ich damit Open-Source-Projekte zu erkunden, die zu meiner Vorstellung eines Sprachassistenten der Privatsphäre respektiert passten. Darunter war unter anderem Mycroft. Im Umfeld dieser Community lernte ich viele tolle Menschen kennen, die sich stark für offene (Sprach) Technologien und Daten engagierten. Darunter [Kathy Reid](https://community.mozilladatacollective.com/author/kathyreid/), die sich seit vielen Jahren für offene Systeme einsetzt und heute im [Mozilla Data Collective](https://datacollective.mozillafoundation.org/?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten%5Fde) aktiv ist. Technisch war das alles herausfordernd und spannend. Doch die Qualität der Sprachausgabe bzw. der TTS-Stimmen (Text-to-Speech), insbesondere im Deutschen, ließ stark zu wünschen übrig. Eben die typischen eSpeak-Stimmen: schnell, effizient, aber völlig roboterhaft. Zum Testen okay, aber nichts, was man sich im Alltag längerfristig anhören möchte. Beim Lesen der Mycroft Dokumentation stieß ich auf die Möglichkeit, auf Basis von Audioaufnahmen meine eigene synthetische Stimme zu trainieren. Wenn man das heute, im Jahr 2026, liest dann denkt man vielleicht: „Einfach ein paar Sekunden aufnehmen, fertig.“ 2019 sah die Sache ganz anders aus. Die Empfehlung lautete: Mindestens 16 Stunden sauberes, neutrales Audio aufnehmen. Also fing ich an aufzunehmen. Abends. An Wochenenden. Satz für Satz. Monat für Monat. Und trotz oder vielleicht gerade wegen meiner großen Motivation und Begeisterung habe ich sehr schnell losgelegt und dabei viele Fehler gemacht. Ich benutzte ein billiges USB-Headset statt eines hochwertigen Mikrofons. Ich versuchte, so klar wie möglich zu sprechen und verlor dadurch jeglichen natürlichen Sprachfluss. Die Betonung der Sätze klang eher nach einem überambitionierten Nachrichtensprecher als nach normalem Redefluss. Nach tausenden von Aufnahmen (etwa 10.000) trainierte ich damit mein erstes TTS-Modell. Mein Computer lief tagelang. Und das Ergebnis? Nun ja … man konnte meine Stimme irgendwie erkennen. Aber die Qualität war alles andere als gut. Neben der Stimme gab es Störgeräusche, Brummen und Echo. Also bat ich die Mycroft Community um Hilfe und zwei interessante Dinge passierten. Erstens gab es echtes Interesse. Deutschsprachige Community-Mitglieder fragten, ob ich die Aufnahmen oder die trainierte TTS Stimme veröffentlichen wolle. Interesse? An meiner Stimme? Obwohl ich natürlich weiß, dass es sehr viele attraktivere und bessere Stimmen gibt. Das fühlte sich gut an. Und zusätzlich meldete sich Dominik Kreutz aus der Mycroft Community bei mir und bot an, meine Aufnahmen zu analysieren. Sein Hilfsangebot war simpel: „Schick mir deine Aufnahmen, ich höre sie mir an.“. Da er Audio Expertise zu haben schien, war er auch bereit die Aufnahmen zu optimieren. Aber ich dachte: Moment mal. Der ganze Grund, warum ich dieses Projekt angefangen habe, war doch um meine Stimme nicht ins Internet zu übertragen. Und jetzt soll ich sie einem Fremden aus einer Online-Community schicken?! Darüber musste ich ein paar Tage nachdenken. Schließlich traf ich eine bewusste Entscheidung. Ich wollte ihm vertrauen. Also schickte ich ihm die Daten - und habe diese Entscheidung nie bereut. Sein Feedback zur Qualität meiner Aufnahmen war hart, aber ehrlich: Die Aufnahmen waren schlecht. Einige waren technisch vielleicht noch zu retten, aber die meisten unbrauchbar. Das war ein harter Schlag! Ich hatte monatelang Aufnahmen gemacht, meine Freizeit investiert und musste nun akzeptieren, dass vieles davon im Grunde wertlos war. Beim Anhören auf voller Lautstärke waren dort Rauschen und weitere Störgeräusche deutlich wahrzunehmen. In diesem Moment lernte ich eine der wichtigsten Lektionen auf diesem Weg und im KI Umfeld im Allgemeinen: „shit in, shit out“. Oder etwas höflicher ausgedrückt: Die Qualität deiner Daten bestimmt die Qualität deiner Ergebnisse. Wenn die Eingangsdaten Rauschen enthalten, nimmt das maschinelle Lernen das ernst und erzeugt dieses Rauschen auch in der Sprachausgabe absichtlich wieder. An diesem Punkt stand ich vor der Wahl: aufhören oder von vorne anfangen. Aber man bekommt den Geist ja bekanntermaßen nicht wieder in die Flasche zurück. Der Tech-Enthusiast in mir war weiterhin hoch motiviert. Also machte ich weiter. Ich kaufte besseres Equipment, baute eine kleine Aufnahmekabine aus Holz, Teppich und Akustikschaumstoffplatten und verbrachte die nächsten Monate meiner Freizeit (erneut) mit vielen Aufnahmesitzungen. Gleichzeitig teilte ich meine Fortschritte mit der Community (Mycroft und Mozilla) und erkannte, dass das Interesse an freien deutschen Stimmdaten real war. Und dann kam die große Frage: Was mache ich mit den vielen tausenden Aufnahmen meiner Stimme? Veröffentliche ich gar nichts? Nur Auszüge davon? Nur ein fertig trainiertes TTS-Modell? Oder auch alle Aufnahmen komplett? Und wenn ja, mit Einschränkungen oder komplett offen? Mir war die Tragweite dieser Fragen sehr bewusst. Sprachtechnologie würde eindeutig an Bedeutung gewinnen – da war ich mir damals schon recht sicher. Und ich wusste: Wenn ich meine Stimme veröffentliche, gebe ich die Kontrolle auf. Vielleicht schränke ich damit sogar meine persönlichen zukünftigen Möglichkeiten ein. ielleicht können wir demnächst unsere Wohnungstüren per Stimme öffnen oder auch auf weitere sensible Bereiche wie Bankportale über die Stimme zugreifen. Solche Nutzungsmöglichkeiten kann ich dann nicht nutzen, wenn meine Stimme als Allgemeingut verfügbar ist. Aufgrund der rasend schnellen Entwicklung gelten diese Probleme allerdings mittlerweile auch für Menschen, die ihre Stimme nicht willentlich verschenkt haben. Aber das ist noch eine ganz andere Herausforderung. Und natürlich kamen weiterhin die offensichtlichen Bedenken auf. Was, wenn jemand meine Stimme für Dinge benutzt, mit denen ich absolut nicht einverstanden bin? Politisch sehr fragwürdige Inhalte? Extremismus? Betrug? Ich hatte davor zwar keine Angst, aber ich habe diese möglichen Risiken sehr ernst genommen. Nachdem ich einige Tage darüber nachgedacht und mit Familie und Freunden gesprochen hatte, traf ich meine Entscheidung. Wenn ich das mache, dann richtig. [CC0](https://creativecommons.org/publicdomain/zero/1.0/deed.de?ref=community.mozilladatacollective.com)! Keine Einschränkungen. Genau wie bei [Mozilla Common Voice](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com). Ich wollte niemanden ausschließen. Weder Forschung noch Open Source Projekte oder den Einsatz für kommerzielle Zwecke. Ich wollte nicht anfangen, ständig „Ja, aber nur unter bestimmten Bedingungen“ hinzuzufügen. Wenn offen, dann wirklich offen. Rückblickend überrascht es mich immer noch, was sich im Laufe der Jahre aus „[Thorsten-Voice](https://datacollective.mozillafoundation.org/datasets/cmm4de9w500ntmh073nx14k7p?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten%5Fde)“ entwickelt hat. Glücklicherweise primär, wenn auch nicht ausschließlich, in sehr positiver Weise. Irgendwann schickte mir jemand einen Link zu einem Video aus der sogenannten Reichsbürger Szene, welches mit meiner KI-Stimme vertont wurde. Das war extrem unangenehm, denn diese Sichtweise lehne ich komplett ab. Aber in diesem Moment wurde mir bewusst – aus einem theoretischen Risiko, wurde Realität. Meine Thorsten-Voice Stimme wurde in einem mir sehr unangenehmen Zusammenhang verwendet. Kein schönes Gefühl. Und dennoch habe ich meine Entscheidung nie bereut. Denn die positiven Auswirkungen überwiegen deutlich. Ich habe viele ermutigende Nachrichten erhalten. Ein Informatiklehrer in Berlin erzählte mir, dass seine Klasse nun Sprachprojekte (bspw. ein Telefonmenü mit dynamischer Sprachausgabe) lokal und ohne Cloud-Abhängigkeit entwickeln konnte. Das Schweizer Projekt „Lernstick“ nutzt unter anderem meine Stimme für Barrierefreiheit im Bildungsbereich. Thorsten-Voice wird in Smart Home-Systemen eingesetzt. Jemand erzählte mir, meine Stimme ertöne von der Decke einer Finca auf Mallorca. Und es gibt Anwendungsfälle in Screenreadern, die Menschen mit Seh- oder Leseschwächen unterstützen. Und dies sind nur einige der positiven Rückmeldungen, die ich regelmäßig bekomme. In solchen Momenten wird mir immer wieder bewusst: Meine Stimmspende bewirkt Dinge. Ich füge meinen Datensätzen immer eine persönliche Notiz hinzu. Nicht als Einschränkung, sondern als persönliches Statement, da ich nicht beeinflussen kann, was mit meiner Stimme gesagt wird. Aber ich kann kommunizieren, wofür ich als Person stehe. > **Ich glaube an die Gleichheit aller Menschen, unabhängig von Geschlecht, sexueller Orientierung, Religion, Hautfarbe oder Geburtsort. Ich glaube an eine globale Welt, in der jeder überall willkommen ist. Dass Wissen und Bildung für alle frei zugänglich sein sollte. Und ich glaube, dass wir als Menschheit zu Großartigem fähig sind, wenn wir einander vertrauen und zusammenarbeiten.** Heute hat sich die Situation weiterentwickelt. Die Erstellung synthetischer Stimmen ist deutlich einfacher geworden. Man benötigt keine 16 Stunden Audiomaterial mehr. Oft genügen wenige Sekunden. Eine einfache Sprachnachricht kann ausreichen. Dies birgt sowohl Chancen als auch Herausforderungen – für jeden von uns und als Gesellschaft im Ganzen. Genau deshalb halte ich offene, transparente und ethisch fundierte Datensätze für so wichtig und Plattformen wie [Mozilla Data Collective](https://datacollective.mozillafoundation.org/?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten%5Fde) für so bedeutsam. Viele moderne KI-Systeme wurden mit Daten trainiert, deren Herkunft unklar und deren Einwilligung oft fragwürdig ist. Offene Daten bieten eine echte Alternative. Join Mozilla Data Collective → Organisationen wie Mozilla legen mit Projekten wie [Common Voice](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com) und [Data Collective](https://datacollective.mozillafoundation.org/?utm%5Fsource=ghost&utm%5Fmedium=community-post&utm%5Fcampaign=thorsten%5Fde) seit Jahren den Grundstein dafür. Im Laufe der Zeit hatte ich die Gelegenheit, mit tollen Menschen in diesem Bereich in Kontakt zu treten, die sich sehr engagiert für Offenheit und verantwortungsvollen Umgang mit Daten einsetzen. Vielleicht ermutigt meine kurze Geschichte andere, sich ebenfalls einzubringen. Ich kann nur sagen: Es fühlt sich gut an! Rückblickend, nach einigen Jahren, kann ich mit voller Überzeugung sagen: Ich habe es nie bereut, meine Stimme „verschenkt“ zu haben. Ich würde es jederzeit wieder tun. Trotz möglicher Risiken. Denn ich bin fest davon überzeugt: Wenn wir einander vertrauen, wenn wir Wissen teilen, wenn wir offen zusammenarbeiten, dann können wir gemeinsam viel Gutes erreichen. (und davon kann die Welt aktuell sicherlich einiges vertragen) ### Why Whisper Still Struggles with Australian English - and What We Did About It URL: https://community.mozilladatacollective.com/why-whisper-still-struggles-with-australian-english-and-what-we-did-about-it/ Last updated: 2026-03-23T18:04:25.000Z *By Kathy Reid · Mozilla Data Collective* --- Ask a Queenslander to say "no worries" into a Home Assistant Voice Preview. There's a decent chance it mishears them. Ask someone from Radelaide and things get worse. Ask anyone who grew up calling things "heaps good" and you'll start to understand that the problem isn't the speaker — it's the model. OpenAI's Whisper is genuinely impressive. It supports 99 languages, runs locally, and powers some of the most popular open-source voice assistant stacks in the world — including `faster-whisper`, which sits underneath Home Assistant's local voice pipeline. But Whisper was trained on data that skews hard towards American and British English. Australian English — with its distinct vowel shifts, distinctive prosody, and accent variation across states and subcultures — is underrepresented in that training corpus. The result is a model that routinely mishears the 26 million people who call Australia home. This post explains what we did about it at [Everything Open 2026 in Canberra](https://2026.everythingopen.au/?ref=community.mozilladatacollective.com), and gives you everything you need to replicate the result for Australian English — or adapt the approach for any under-represented accent available in [Common Voice](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com). --- ## The Problem in Plain Terms When a speech recognition model is trained, it learns to map acoustic patterns — the raw shapes of sound — onto text. That mapping is only as good as the data it's trained on. If the training data contains very few examples of a particular accent, vowel pattern, or phonological feature, the model's predictions for speakers with those characteristics will be less accurate. For Australian English specifically, a few features create consistent failure modes in Whisper: - **Vowel shifts**: Australian English has a well-documented system of vowel shifts that diverge significantly from American English. Consider for example the way an Australian-accented person says "vase" - to rhyme with "Mars". This differs from an American-accented person, who is more likely to say "vase" to rhyme with "ways" or "praise". - **High Rising Terminal (HRT)**: the tendency for declarative statements to end with a rising intonation, which models sometimes mis-parse as a question or misread as uncertainty. - **Regional variation**: "General Australian" (the reference accent heard on the ABC) behaves differently to broad Queensland speech, South Australian speech, or the accents common in multicultural urban centres like Western Sydney - such as the emerging [Multicultural Australian English](https://www.tandfonline.com/doi/full/10.1080/07268602.2024.2380680?ref=community.mozilladatacollective.com) (Cox & Penney, 2024) The fix — fine-tuning — is well understood in principle. You take a trained model and continue training it on data that has a distribution that's closer to your use case. In simple terms, you're not starting from scratch; you're nudging the model's weights to pay better attention to the acoustic patterns that match the real-life context it will be deployed in. --- ## The Stack Our approach draws on the [Mozilla.AI Speech-to-Text Finetuning Blueprint](https://blueprints.mozilla.ai/all-blueprints/finetune-an-asr-model-using-common-voice-data?ref=community.mozilladatacollective.com), developed by our colleague [Kostis Saitas Zarkias](https://www.linkedin.com/in/kostissz/?ref=community.mozilladatacollective.com). It's a clean, reproducible pipeline that handles data preparation, training configuration, and evaluation. We wrapped a tutorial around it, customised it for Australian-accented data, and ran it as a hands-on 100-minute workshop. The full stack is: | Component | What it does | | ----------------------------------- | ---------------------------------------------- | | **Mozilla Data Collective** | Data discovery, download, and licensing | | **Common Voice v24 (en-AU subset)** | \~55,600 clips, \~1.92 GB, CC0 licence | | **Mozilla.AI Blueprint** | Fine-tuning pipeline on top of 🤗 Transformers | | **Google Colab (free tier GPU)** | Compute — no local GPU required | | **Hugging Face Hub** | Model storage and sharing | | **faster-whisper** | Deployment target for Home Assistant | Everything is open source. Everything is CC0-licensed on the data side. You don't need an Anthropic API key, an OpenAI account, or a cloud GPU credit card. --- ## The Data: Where It Came From Since 2022, Mozilla Common Voice has allowed contributors to self-report the accent they speak with. That means there's now a growing pool of voice data tagged with granular accent metadata — and we can use that to extract accent-specific subsets for fine-tuning. The dataset we used for the tutorial is the [Common Voice v24 English – en-AU subset](https://datacollective.mozillafoundation.org/datasets/cmko7havo02f5nw07rbwwhowe?ref=community.mozilladatacollective.com), curated and hosted on [Mozilla Data Collective](https://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com). It covers six Australian accent categories: - Australian English - General Australian - South Australia - Educated Australian Accent - Sydney – Middle Eastern seaboard Australian - Queenslandish 55,673 rows. Approximately 4.68 minutes of audio. CSV + MP3\. CC0. If you want to build your own accent subset from Common Voice, the repository includes a dedicated [extract-from-cv-by-accent](https://github.com/Mozilla-Data-Collective/tutorial-whisper-fine-tuning-australian-EO2026/tree/main/extract-from-cv-by-accent?ref=community.mozilladatacollective.com) directory with the subsetting code. This is the part of the work that tends to be invisible in tutorials but is often the most practically useful — the preprocessing and data curation logic. [Discover datasets](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) --- ## Running the Tutorial The tutorial is structured around a single Google Colab notebook: [EO2026\_teach\_whisper\_to\_speak\_Australian.ipynb](https://github.com/Mozilla-Data-Collective/tutorial-whisper-fine-tuning-australian-EO2026/blob/main/EO2026%5Fteach%5Fwhisper%5Fto%5Fspeak%5FAustralian.ipynb?ref=community.mozilladatacollective.com). **What you'll need before you start:** 1. A [Mozilla Data Collective](https://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com) account and the [en-AU dataset download](https://datacollective.mozillafoundation.org/datasets/cmko7havo02f5nw07rbwwhowe?ref=community.mozilladatacollective.com) (1.92 GB) 2. A Hugging Face account and a fine-grained access token (`HF_TOKEN`) with read/write repo permissions and inference access 3. A Google account to run Colab **The pipeline, in order:** 1. **Environmental setup** — install dependencies, configure Hugging Face credentials 2. **Data loading and preparation** — convert audio to the required sample rate, reshape the dataset structure into the format expected by the Blueprint 3. **Fine-tuning on GPU** — Colab's free T4 is enough; training runs while you have a coffee 4. **Evaluation** — measure Word Error Rate (WER) before and after fine-tuning on held-out Australian speech 5. **Export to `faster-whisper` format** — convert the fine-tuned model for deployment in Home Assistant > **Known issue:** There is a currently-unresolved memory exhaustion bug that can surface during the fine-tuning step in the Colab notebook. If you hit it, check your batch size configuration and consider reducing it - or, [PRs welcome](https://github.com/Mozilla-Data-Collective/tutorial-whisper-fine-tuning-australian-EO2026/pulls?ref=community.mozilladatacollective.com)! --- ## Why This Matters Beyond Australian English The pattern here is completely general: 1. Find an accent or dialect under-represented in Whisper's training data 2. Find or build a tagged speech dataset for that accent (Common Voice has tagging for hundreds of accents globally) 3. Use the same Blueprint pipeline to fine-tune 4. Export and deploy Common Voice now carries accent data for accents from South African English to Singapore English to Nigerian English to Scottish English. The methodology in this tutorial scales to all of them. If you're building a voice product for any community whose English sounds different from the American-British norm — or working in a language where the default Whisper model under-performs — this workflow is your starting point! [Discover more datasets](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) --- ## The Bigger Picture: Open Data as Infrastructure The thing that makes this tutorial possible isn't the fine-tuning technique — that's been documented elsewhere. It's the availability of accent-tagged, openly licensed, ethically collected speech data in a form that's easy to discover and use. Mozilla Data Collective is our answer to the problem of AI data being either locked up in proprietary silos or extracted from communities without consent. It's built on a model of community contribution and transparent governance, extended with better tooling for data stewardship, access control, and dataset discovery. Every clip in the `en-AU` dataset was contributed by a person who chose to record it. Every accent tag was self-reported. That's what ethical AI data infrastructure looks like. [Find datasets](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) --- ## Get Involved - **Run the tutorial yourself**: [GitHub repo](https://github.com/Mozilla-Data-Collective/tutorial-whisper-fine-tuning-australian-EO2026?ref=community.mozilladatacollective.com) - **Get the dataset**: [Mozilla Data Collective](https://datacollective.mozillafoundation.org/datasets/cmko7havo02f5nw07rbwwhowe?ref=community.mozilladatacollective.com) The code is Apache 2.0\. The data is CC0\. The methodology is reproducible. If your voice assistant mishears you, you now have the tools to fix that — and the invitation to help others do the same. --- *Kathy Reid is Head of Applied R&D at Mozilla Data Collective.* ### How can a nearly century-old publisher remain relevant and continue to grow? URL: https://community.mozilladatacollective.com/how-can-a-nearly-century-old-publisher-remain-relevant-and-continue-to-grow/ Last updated: 2026-03-19T13:09:35.000Z *By: Yacub Fahmilda & Panjebar Semangat* Most international and national language datasets are well resourced and widely accommodated in technology for AI and other possible forms of tech products. Meanwhile, many local languages remain less visible, low-resource and under-represented in technology advancement. This gap leads to alienating some native speakers of small communities for their own languages and raises demand for linguistic justice. Adapting under-represented language to the technological ecosystem is one of the key points. A community should have their own space to participate and to control over their data. This adaptation idea has been taken by [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com), a weekly Javanese-language magazine established pre-independence of Indonesia. This magazine is recognized as frontier media which remains relevant in today’s and future technological contexts. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/data-src-image-632d8c3d-3d0f-413c-8884-0e0906d0311a.jpeg) *Taken by* [*Yustri Agung Prastiyono*](https://www.linkedin.com/in/yustriagung?ref=community.mozilladatacollective.com) *&* [*Masyhuri Farhan.*](https://www.linkedin.com/in/masyhuri-farhan/?ref=community.mozilladatacollective.com) # The Long Journey of Panjebar Semangat [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com), a leading dataset contributor from Indonesia, may truly “spread the collective spirit”, reflecting the meaning of its name of “Panjebar Semangat”, signifying to initiate, encourage, and inspire others. [Dr. Soetomo](https://id.wikipedia.org/wiki/Soetomo?ref=community.mozilladatacollective.com) was one of the key figures establishing [Panjebar Semangat ](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)on September 2nd 1933 with his colleague[ Imam Supardi](https://panjebarsemangat.id/ngenani/?ref=community.mozilladatacollective.com) who served as the executor of the initiative. Since then, this magazine publisher has continuously contributed to Javanese communities across dynamic social, political and cultural transformations, as outlined in the chronology below. 1. Pre-independence (1933-1942) During the Dutch East Indies colonial period, Javanese language was a means of power to evoke nationalism across Javanese communities as colonizers did not understand it. This spirit grew, contributing them to the idea of integrating an archipelagic nation into Indonesia as one nation, a national identity. 1. Post-independence (after 1949) There was deep silence between 1942 and 1945 as [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com) was shut down during Japan occupation. In 1949 there was a second military aggression of the Dutch, and later the publisher gained power back after the independence declaration. During this post-independence, [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com) played a crucial role in sustaining the spirit of nationalism and fostering nation-building amongst Javanese communities. 1. Industrialisation and transmigration (1990s-2000s) This period marked a critical phase for Javanese community, as the number of the speakers started to decline amid industrialisation and state-led transmigration programs. A large number of Javanese people moved to other regions for new settlements, such as Aceh, Sumatera, Lampung, Jakarta, Kalimantan, Sulawesi, Nusa Tenggara, and other parts of eastern Indonesia. [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com) took an initiative to preserve the language and culture through periodicals, reaching transmigrants of Javanese communities and presenting Javanese perspectives through arts, cultural, and literary works. 1. The digital era (from 2018 onwards) Over the last five years, [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com) has taken a strategic action to embrace technology and mediate their valuable readers by providing the magazines in digital formats. Interestingly, some subscribers who come from overseas, including the Netherlands, seek to learn Javanese. This shift demonstrates that technological adaptation allows a local publisher to gradually expand to a global scale. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/data-src-image-a99d7cc4-1e29-4a53-99dd-cb1efb537341.jpeg) *Taken by* [*Yustri Agung Prastiyono*](https://www.linkedin.com/in/yustriagung?ref=community.mozilladatacollective.com) *&* [*Masyhuri Farhan.*](https://www.linkedin.com/in/masyhuri-farhan/?ref=community.mozilladatacollective.com) More recently, [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com) has extended this adaptation into the field of artificial intelligence through community-led governance. The publisher has shared nearly 2 million words from publications over the last three years through [Mozilla Data Collective](https://datacollective.mozillafoundation.org/datasets/cmk719jvs02z1nt07tjswt52s?ref=community.mozilladatacollective.com). This number represents only a small part of their total archive, indicating even greater potential for future contributions and collaborations. Through this contribution, the publisher encourages those who are working on language and cultural minorities, to take real and measurable action in technology. # Staying Relevant in a Changing World Adapting to dynamic social changes, including technology and people's lifestyle, allows [Panjebar Semangat ](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)to stay relevant for present and future Javanese communities. Rather than allowing unregulated data extraction, the publisher curated and shared texts for a dataset titled “[*Korpus Majalah Bahasa Jawa Panjebar Semangat*](https://datacollective.mozillafoundation.org/datasets/cmk719jvs02z1nt07tjswt52s?ref=community.mozilladatacollective.com)”, intended for data training in NLP tasks and other socially valuable technological products. This contribution represents a small-scale yet globally connected effort. Beyond language communities, language-learning platforms such as [Italki](https://www.italki.com/en/teachers/javanese?ref=community.mozilladatacollective.com), [HelloTalk](https://www.hellotalk.com/id/partners/exchange/javanese?ref=community.mozilladatacollective.com), [Tandem, ](https://tandem.net/learn/language/supported-languages?ref=community.mozilladatacollective.com)and other educational technologies may also benefit from this shared dataset. This reflects [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)’s institutional resilience and courage in adapting across historical periods, from local print culture to global digital ecosystems. [Panjebar Semangat ](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)is one of many language and cultural activists across the Indonesian archipelago who share a similar spirit, commitment, and motivation. [Panjebar Semangat ](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)is a role model which I do believe other communities can follow to take part in ethical, impactful and community-governed data initiatives. # From Archive to Community Governance The rapid expansion of AI systems has exposed a structural weakness in global language datasets: minority and indigenous languages are included without governance, consent, or reciprocity. While open data initiatives have accelerated innovation, they have also unintentionally normalized extraction-based practices that marginalize the very communities they claim to include. In response to this challenge, [Panjebar Semangat ](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)presents a practical counter-model through its collaboration with Mozilla Data Collective through community-centered data governance. The “[*Korpus Majalah Bahasa Jawa Panjebar Semangat*](https://datacollective.mozillafoundation.org/datasets/cmk719jvs02z1nt07tjswt52s?ref=community.mozilladatacollective.com)” consists of curated Javanese-language text, treated not as raw data but as a governed linguistic asset–editorially vetted, context-rich, and incrementally released. This initiative demonstrates that language datasets can be treated as cultural infrastructure, not disposable inputs. [Panjebar Semangat ](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)advances a Community-Governed Language Dataset (CGLD) framework that prioritizes community authority, purpose-bound usage and benefit reciprocity. The model aligns directly with Mozilla’s mission to promote trustworthy, human-centered AI while reducing ethical, reputational, and regulatory risks. For Mozilla Data Collective, CGLD offers a scalable governance layer that strengthens dataset legitimacy without undermining openness. It reframes “open” not as unrestricted extraction, but as accountable collaboration. Institutions that adopt this approach gain more resilient datasets, deeper community trust, and long-term alignment with emerging AI governance norms. The strategic question is no longer how to include more languages in AI, but how to include them without reproducing historical inequities. [Panjebar Semangat](http://panjebarsemangat.id/?ref=community.mozilladatacollective.com)’s collaboration with Mozilla Data Collective provides a real replicable answer, one that positions Mozilla not just as a data aggregator, but as a global standard-setter for ethical language AI. ### Call for proposals: MDC is commissioning mission-aligned datasets! URL: https://community.mozilladatacollective.com/call-for-proposals-mdc-is-commissioning-mission-aligned-datasets/ Last updated: 2026-04-17T13:43:01.000Z # Mozilla Data Collective is building towards a multicultural, multilingual, and multimodal future that works for all of us. And over the past few months, we’ve listened as people have flagged what kinds of datasets they need, but are struggling to find. So we’re pleased to announce that MDC will be commissioning several exciting datasets over the next three months to begin meeting this need (and you can look forward to future rounds as well!) If you represent a small-, medium-sized enterprise (SME), nonprofit, or other mission-aligned organisation that works in dataset creation, we want to hear from you! This is a commissioning programme, the resulting datasets will be owned by Mozilla Data Collective, so we are particularly looking for datasets built from scratch. [Explore Mozilla Data Collective datasets](https://mozilladatacollective.com/datasets?ref=community.mozilladatacollective.com) We welcome all quotes on dataset creation in areas that are un-served, or are under-served, by existing dataset offerings on MDC. **Within our mission of a multilingual, multicultural and multimodal tech future, we seek datasets that speak to common tasks, modalities, and domains:** - Egocentric videos of daily tasks oriented towards safe robotics - Agentic workflows and interactions with AI agents - Computer vision and multimodal datasets relating to physical or online safety - Speech recognition in the healthcare domain - Dialogic interactions in the finance and banking domains - Annotated datasets of performing arts, including singing, dancing This is not an exhaustive list, if you think you have a dataset that you would like to make that meets a real identified need, get in touch! **Particular languages (and varieties) of interest are:** - Arabic (Morocco, Libya, Algeria, Iraq and Syria) - Latin American Spanish - Brazilian Portuguese - Tagalog and other languages of the Philippines - Sign languages (American, British, French, Spanish) - Kannada - Tamil - Burmese and other languages of Burma (Myanmar) - Croatian - Tigrinya - Vietnamese - Thai - Pashto - Uzbek - German - French - Kurdish **Illustrative examples (for inspiration only):** - Dialogic interaction dataset of customer support interactions - Annotated video datasets of public health related lectures - Multimodal dataset of teleoperated robotic arm carrying out tasks like box packing - High value numerical datasets, it could be information relating to pricing, or materials - A single-speaker dataset of studio-quality recordings of Tagalog (male or female speaker) in the health domain aimed at text to speech. - A collection of recordings of alphanumerics in dialectal Arabic (Morocco, Tunisia, Libya, Algeria) in diverse noise environments and spoken by a diverse population. We expect to move quickly, so will start reviewing quotes on a first-come first-served basis starting from the **23th March, 2026\.** You can expecta response to your quote within 1-2 business days. Quotes would ideally include the following information on a per dataset basis: - Description of dataset - Fixed costs associated with the project - Ethical and compliance considerations (we will strongly prefer proposals from close-to-context communities) - Unit price, that could be per hour in the case of ASR, TTS or per thousand tokens or interaction in the case of LLM corpora - Estimated volume deliverable per month - Brief outline of annotation process - Brief explanation of the project team and its qualifications - Earliest delivery date Partners are encouraged to send multiple options. A simple document will suffice, maximum 1-2 pages per dataset. Please do not feel the need to spend time on visual components. *Note: It is important to us that datasets are collected in an ethical and compliant manner. Datasets should be fully anonymised and compliant with relevant privacy regulations.* [Explore Mozilla Data Collective datasets](https://mozilladatacollective.com/datasets?ref=community.mozilladatacollective.com) **Budget:** We expect most quotes to fall between $1000 and $25,000 on a per dataset basis. **Timeline:** - *Announcement of the call:* 17th March - *First quotes received*: 23rd March or before - *Decisions made*: on a rolling basis (first decision by 27th March) - *Delivery window*: 27th March – 27th June (3 months) When evaluating the quotes we will take into account the following factors: delivery speed, volume of data, ethics, price. High quality is assumed 🙂 Please send quotes or requests for clarification directly to [support@mozilladatacollective.com](mailto:support@mozilladatacollective.com?subject=Response%20to%20Dataset%20CFP) with the subject line “Dataset commissioning”. [Join Mozilla Data Collective](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com) ### MDC Release Notes - 13.03.26 URL: https://community.mozilladatacollective.com/mdc-release-notes-13-03-26/ Last updated: 2026-03-13T19:26:11.000Z Hello, Mozilla Data Collective! 👋 These past two weeks, we've been focusing on infrastructure and development of some exciting new features that will be landing in the next couple of weeks 👀 but we've got, as always, a few updates to share and the latest roundup of new datasets on MDC. ### New Features and Changes **Recommended datasets are now visible on individual data listing pages.** These are related datasets to help discover and explore other datasets that might be similar to the ones you're looking at. You can find these on the left side of the page. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/image.png) **It is now possible to report datasets**. While we review each dataset on the platform before it goes live, as we grow, we want to provide trust & safety levers for the community to flag and identify to us if things don't look right. The link emails our team and we'll take a look. Join Mozilla Data Collective → --- ### New Datasets **Aim Foundation** [Dari Literature Corpus by Anjuman e Adabi Nayestan | Mozilla Data CollectiveThe Dari Literature Corpus (Anjuman e Adabi Nayestan) is a curated collection of written Dari (Afghan Persian) literary texts totaling about 1 million tokens. It includes prose, poetry, folklore-inspired narratives, and other culturally significant writings from both contemporary and classical traditions. The texts were collected in Microsoft Word and converted into UTF-8 normalized plain text for computational and linguistic research, including corpus linguistics, digital humanities, and NLP.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-46.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-45.png)](https://datacollective.mozillafoundation.org/datasets/cmmdpikpq003imh077foix53d?ref=community.mozilladatacollective.com) **Collaborative Action for Research &** [IBT Torwali Wordlist | Mozilla Data CollectiveThe IBT Torwali Wordlist contains approximately 20,000 unique entries in Torwali (ISO 639-3: trw), an under-documented Indo-Aryan language spoken in northern Pakistan. The dataset comprises standardized lexical entries covering core vocabulary, function words, and culturally salient terms, with consistent orthography and normalization suitable for linguistic and computational use. Entries are aligned with English and Urdu glosses, and include part-of-speech tag.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-47.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-46.png)](https://datacollective.mozillafoundation.org/datasets/cmmdpbs5z003emh07yvbymzu5?ref=community.mozilladatacollective.com) **Digital Divide Data** [ddd-kenya-somali-68hrs-asr-part1 | Mozilla Data CollectiveThis dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. All transcriptions have undergone a manual verification process to ensure high linguistic accuracy. Recordings feature a balanced mix of genders and various age groups to minimize bias in downstream AI models. This data is specifically designed for training Automatic Speech Recognition (ASR) systems, Text-to-Speech (TTS) synthesis, and general linguistic research for underrepresented African languages.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-34.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-33.png)](https://datacollective.mozillafoundation.org/datasets/cmmng8btl000yl807k8qtx891?ref=community.mozilladatacollective.com) [ddd-kenya-somali-68hrs-asr-part2 | Mozilla Data CollectiveThis dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. All transcriptions have undergone a manual verification process to ensure high linguistic accuracy. Recordings feature a balanced mix of genders and various age groups to minimize bias in downstream AI models. This data is specifically designed for training Automatic Speech Recognition (ASR) systems, Text-to-Speech (TTS) synthesis, and general linguistic research for underrepresented African languages.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-35.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-34.png)](https://datacollective.mozillafoundation.org/datasets/cmmnixx2d0043ml07zbq1i7oi?ref=community.mozilladatacollective.com) [ddd-kenya-somali-68hrs-asr-part3 | Mozilla Data CollectiveThis dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. All transcriptions have undergone a manual verification process to ensure high linguistic accuracy. Recordings feature a balanced mix of genders and various age groups to minimize bias in downstream AI models. This data is specifically designed for training Automatic Speech Recognition (ASR) systems, Text-to-Speech (TTS) synthesis, and general linguistic research for underrepresented African languages.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-38.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-37.png)](https://datacollective.mozillafoundation.org/datasets/cmmniydvi0047ml07nt6z5xud?ref=community.mozilladatacollective.com) **Kaleem Art Press** [Jhoke Publisher Multan’s Saraiki Newspaper Corpus | Mozilla Data CollectiveJhoke Publishers Multan’s Saraiki Newspaper Corpus is a curated text dataset with about 1.25M tokens (1,258K) of Saraiki content collected from Daily Jhoke Saraiki (Multan, Pakistan) and Jhoke Publishers (Multan, Pakistan). Daily Jhoke Multan (ݙین٘ھ وار جھوک ملتان) is a Saraiki newspaper and publishing house based in Multan. It covers regional news and also publishes Saraiki literature, including major literary and religious works (e.g., a Saraiki Quran translation by Professor Dilshad Kalanchvi). The corpus includes three UTF-8 text files (each treated as a separate genre/domain) and a cleaned version with Unicode normalization, standardized whitespace and punctuation, and removal of stray symbols or markup. The dataset reflects contemporary Saraiki usage across journalistic, literary, cultural, and social domains and supports computational and linguistic research.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-52.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-51.png)](https://datacollective.mozillafoundation.org/datasets/cmmao7dc504oamh0710j4wau1?ref=community.mozilladatacollective.com) [Saraiki-English Parallel Corpus | Mozilla Data CollectiveThis English–Saraiki Parallel Corpus is a curated bilingual dataset of 51,447 aligned sentence pairs (about 0.89 million words in total), translated from English into Saraiki by Kaleem Art Press and cleaned into a consistent sentence-level format for reliable alignment; it is designed to support machine translation training and evaluation, bilingual lexicon and terminology work, and broader linguistic and NLP research for Saraiki, including data-driven language technology development.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-51.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-50.png)](https://datacollective.mozillafoundation.org/datasets/cmmaphscg04t2mk07i1f8yc0q?ref=community.mozilladatacollective.com) **Keblagh e Azergi** [Elkhani Hazargi Literature Corpus | Mozilla Data CollectiveThe Hazargi Literature Corpus (Keblagh e Azergi) is a monolingual literary dataset for documenting and supporting computational research on Hazargi (Hazaragi), an eastern Persian (Dari) dialect spoken by Hazara communities in Afghanistan and the diaspora. It contains 12 digitized works (prose, poetry, folklore, drama) converted from Word into UTF-8 normalized plain text while preserving original orthography and dialectal features. Total size: \~0.5M tokens (513,483).![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-45.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-44.png)](https://datacollective.mozillafoundation.org/datasets/cmmdtenxt0050mh0792d10knv?ref=community.mozilladatacollective.com) **Institute of African Digital Humanities** [Mada-French Parallel Corpus 1.0 | Mozilla Data CollectiveThis dataset comprises a parallel corpus of Mada–French literary text translations totalling 2,154 lines. It is designed to support the benchmarking, training and evaluation of machine translation models for Mada, a language spoken in Cameroon. The corpus provides aligned, sentence and paragraph-level translations that capture the stylistic, lexical and syntactic features of literary Mada discourse and how these are rendered in the local variety of French.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-53.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-52.png)](https://datacollective.mozillafoundation.org/datasets/cmmamvzrz04qtmk077j1k99vt?ref=community.mozilladatacollective.com) **Taruen** [Finnish Public Domain 20th Century Literature Text Corpus | Mozilla Data CollectiveThis corpus contains a curated collection of public domain literature from Finland, featuring works by authors who died between 1901 and 1955\. The dataset captures the literary landscape of early 20th-century Finland and includes independent texts in both of the country’s official languages: Finnish (fi) and Swedish (sv). The texts were programmatically extracted from Project Lönnrot, a volunteer-driven digital library. To ensure linguistic relevance for modern NLP tasks, the extraction pipeline strictly filtered for works published in 1901 or later. Language codes for each text were dynamically detected using CLD algorithms. The corpus comprises approximately 69.1 million words across multiple plain text files, with each file prefaced by structured YAML front matter containing relevant metadata (title, author, year, source URL, language), followed by the original project’s boilerplate preamble enclosed in delimiter tags, and finally the literary text proper. All included works are fully in the public domain under Finnish and EU copyright law.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-55.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-54.png)](https://datacollective.mozillafoundation.org/datasets/cmm5078n50168mk07v64792sf?ref=community.mozilladatacollective.com) [![CTA Image](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/03/Redditinc_Thumbnail_Brand_Logo.png)](https://www.reddit.com/r/MozillaDataCollective/?ref=community.mozilladatacollective.com) Mozilla Data Collective is now on reddit! Join us to share your projects, talk data, and contribute your experience and expertise to a growing community of ethical data practitioners. [Visit r/MozillaDataCollective ](https://www.reddit.com/r/MozillaDataCollective/?ref=community.mozilladatacollective.com) **MDC Community Concierge** [Bangor Miami Spanish-English Corpus | Mozilla Data CollectiveThe Bangor Miami Corpus of Spanish-English bilingual speech, containing around 240,000 words over 35 hours of recorded audio conversations. The dataset includes the audios, transcriptions and glosses in CHAT format, and word-level analyses of the transcriptions in .tsv files.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-44.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-43.png)](https://datacollective.mozillafoundation.org/datasets/cmmfulo4r018bnz07py4q9t09?ref=community.mozilladatacollective.com) [Bangor Patagonia Welsh-Spanish Corpus | Mozilla Data CollectiveThe Patagonia Welsh-Spanish corpus contains around 195,000 words: 78% Welsh, 17% Spanish, 5% indeterminate (i.e. the relevant word appears in the dictionaries of both main languages). The dataset includes the audios, transcriptions and glosses in CHAT format, and word-level analyses of the transcriptions in .tsv files.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-50.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-49.png)](https://datacollective.mozillafoundation.org/datasets/cmmccfc5000efmu07ommi3zfr?ref=community.mozilladatacollective.com) [Bangor Siarad Welsh-English Corpus | Mozilla Data CollectiveThe Siarad Welsh-English corpus, containing around 450,000 words, 84% Welsh, 4% English, 13% indeterminate (the relevant word appears in the dictionaries of both main languages). The dataset includes the audios, transcriptions and glosses in CHAT format, and word-level analyses of the transcriptions in .tsv files.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-49.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-48.png)](https://datacollective.mozillafoundation.org/datasets/cmmccg5qi00enmu07w9wjpnrn?ref=community.mozilladatacollective.com) **Community Datasets** [Javanese TTS of Banyumasan Dialect | Mozilla Data CollectiveThis dataset comprises speech data produced by a speaker of the Banyumasan dialect of Javanese (locally known as Ngapak), Central Java Province, Indonesia. All datasets use the informal register (Ngoko) and include various topics.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-54.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-53.png)](https://datacollective.mozillafoundation.org/datasets/cmmamtaf104nemh07xa9e7sdx?ref=community.mozilladatacollective.com) [Kokoro Speech Dataset | Mozilla Data CollectiveKokoro Speech Dataset is a public domain Japanese speech dataset. It contains 43,253 short audio clips of a single speaker reading 14 novel books. The format of the metadata is similar to that of LJ Speech so that the dataset is compatible with modern speech synthesis systems. The texts are from Aozora Bunko, which is in the public domain. The audio clips are from LibriVox project, which is also in the public domain. Readings are estimated by MeCab and UniDic Lite from kanji-kana mixture text. Readings are romanized which are similar to the format used by Julius. The audio clips were split and transcripts were aligned automatically by Kokoro-Align.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-42.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-41.png)](https://datacollective.mozillafoundation.org/datasets/cmmknsho4014wmf087kvq5rc6?ref=community.mozilladatacollective.com) [Malayalam Time-Aligned Speech Corpus | Mozilla Data CollectiveThis dataset is a speaker-organized Malayalam speech corpus consisting of 100 audio recordings and 100 corresponding transcription files in .srt format. The transcriptions are time-aligned and include timestamps matched to the audio. The dataset contains recordings from 5 speakers, including 3 male and 2 female speakers, and the average length of each audio file is approximately 3 minutes. The data is arranged speaker-wise, making it easy to identify and work with each speaker’s recordings and transcriptions separately. This dataset is suitable for automatic speech recognition, forced alignment, speech-text synchronization, subtitle alignment, speech segmentation, and Malayalam speech technology development.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-39.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-38.png)](https://datacollective.mozillafoundation.org/datasets/cmmno795h009hml07dh7uefvp?ref=community.mozilladatacollective.com) [Sundanese TTS | Mozilla Data CollectiveThe Sundanese TTS dataset represents the Sundanese language using the Priangan Sundanese dialect as the standard Sundanese in West Java province, Indonesia, reflecting both traditional forms and modern variations in everyday communication practices. This dataset can be utilized for linguistic research, cultural documentation, sociolinguistic studies, and the development of regional language technologies involving code-mixing with Indonesian.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-43.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-42.png)](https://datacollective.mozillafoundation.org/datasets/cmmj6vyb902ownz07j4k7cunj?ref=community.mozilladatacollective.com) [TODa: Tamazight Open Dataset | Mozilla Data CollectiveWelcome to the Tamazight Open Dataset (TODa), a groundbreaking open-source project dedicated to preserving and advancing the Tamazight language. With its extensive collection of linguistic data, TODa stands as a pioneering collaborative project for Tamazight <=> Englis translation, specifically designed for Natural Language Processing applications. TODa’s unique approach combines both semantic and syntactic categorization methods, offering a rich representation of words in their various contexts and forms. The dataset encompasses a comprehensive collection of linguistic elements, including detailed verb conjugations across different tenses, noun variations, and an extensive compilation of translated expressions that capture the language’s nuances. What sets TODa apart is its inclusive approach to Tamazight’s writing systems. The dataset thoughtfully incorporates Latin alphabets, acknowledging and preserving the diverse writing traditions practiced across Amazigh communities. This dual-script approach ensures broader accessibility and cultural authenticity. Our vision is to establish TODa as the cornerstone resource for Tamazight Natural Language Processing. Through this meticulously curated dataset, we strive to empower developers and researchers to create innovative NLP solutions that authentically serve the Amazigh-speaking community. We take pride in our current progress, yet acknowledge that language documentation is an evolving journey. We actively encourage participation from the Amazigh technology community to contribute their expertise in expanding and refining the dataset. Through collaborative effort, we can create a robust foundation for technological innovations that honor and advance Amazigh linguistic heritage.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-40.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-39.png)](https://datacollective.mozillafoundation.org/datasets/cmmm7rvm200b9md07h3pv8uae?ref=community.mozilladatacollective.com) [TTS Balinese Language | Mozilla Data CollectiveThe Balinese TTS dataset is created and narrated by native Balinese speakers with code-mixing in Indonesian. This dataset is designed to showcase the use of the Balinese language in everyday contexts, covering topics such as family, social interactions, and routine community activities. Each recording reflects natural language use by Balinese speakers, thus representing authentic communication in daily life. This dataset can be utilized for linguistic research, the development of automatic speech recognition systems, and other applications focused on the preservation and advancement of the Balinese language.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-41.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-40.png)](https://datacollective.mozillafoundation.org/datasets/cmmm2ru5r003nmd07p53h9wdw?ref=community.mozilladatacollective.com) ### How to License Your Dataset for AI Training: Some Best Practices URL: https://community.mozilladatacollective.com/how-to-license-your-dataset-for-ai-training-some-best-practices/ Last updated: 2026-03-17T06:52:51.000Z ### We get a lot of questions about how to approach licensing your data for AI training. So to help you share your datasets, we’ve compiled some guidance here – it’s intended to be a living document, that we iterate with our partners and communities. [Explore Mozilla Data Collective](https://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com) ## **What Does It Mean to License a Dataset for AI training?** When you license a dataset, you are not selling it. You are giving a company permission to use it in specific ways, under specific conditions, in exchange for payment or other agreed terms. You keep ownership. The company gets access. A well-written license can therefore help you protect your rights, help with prevention of data being misused, and ensure you are fairly compensated. A poorly written one can leave you with little control over how your data shapes AI systems for years to come. Because of how models are then used in downstream applications, it’s hard to ‘walk back’. ## **Why AI Training Licenses Are Different from other types of license** Licensing data for AI training is different to licensing it for research or publishing. When a company trains a model on your dataset, the data is absorbed into the model's parameters. Aka, it is not stored in a folder somewhere that can simply be deleted. This makes the stakes higher, and it means standard licensing templates may not be enough to protect you or your data. Before you begin this process, please make sure you actually have the right to license this data. Think about: 1. Did you create this data yourself, or did you collect it from other sources? 2. Did the original terms of service or data agreements permit commercial licensing? 3. Does the dataset contain personal data? 4. Do contributors or creators whose work is in the dataset have any rights you need to account for? Once you’ve confirmed you’re within your rights to license your data, and it’s safe to do so: you will need to craft a license that is specific to AI use cases. (Ahead of this, if your dataset has a lot of stakeholders, you should make sure you’ve actually spoken to them too. You could read our article here explaining how to run a [Community Workshop for Dataset Governance](https://community.mozilladatacollective.com/your-data-your-rules-a-community-workshop-for-dataset-governance/).) ## **Some Best Practices for Licensing Your Dataset** ### **1\. Explicitly Define How the Dataset Can Be Used** Be very precise about permitted uses (almost pedantically precise) unless you genuinely want them to be able to use it for anything. That might be fine in your context – only you can decide. Can the company use your data to train a general-purpose model? A commercial product? An internal tool only? Can they use it to fine-tune models they then sell to others? Vague language like "for AI" leaves too much room for interpretation, unless you really mean anything. We see some fall out in the industry around confusing licenses already. Spell out clearly: - Which models or products the data can be used to train - Whether the license covers training, fine-tuning, evaluation, or all three - Whether the resulting AI system can be used commercially or if it’s strictly for research purposes - Whether the company can sublicense the data to third parties (usually, you want to say no) If you already have a possible technology company or non-profit you want to work with, you can get their feedback on the license – how clear do they feel about the language? Where is there ambiguity? Bear in mind that they may have their own incentives, and you should come into conversation clear on your own boundaries. If you don’t have a specific partner in mind – we’re also happy to connect you to people, or to comment, over here at Mozilla Data Collective: the social enterprise for data agency and fair value exchange. [Explore other people's licenses on Mozilla Data Collective](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) ### **2\. Set Clear Restrictions on Redistribution** The data supply chain doesn’t stop with the company. Once your data is in a company's hands, you want to be clear about where it can go next. For example, your license might explicitly prohibit: - Sharing or reselling the raw dataset to other parties - Including your data in open datasets or public releases - Using the data to train models that are then shared as open weights, if that is a real concern for you ### **3\. Address Data Retention and Deletion** One of the hard problems in AI licensing is what happens to your data *after* the training run is complete. Your license needs to document: - How long the user/company can retain copies of your data - Whether they must delete your data after training is finished - How deletion will be verified (an audit right might be useful to consider here) Raw data can be deleted, but the model itself will have already "learned" from your data. Your license should acknowledge this distinction clearly. ### **4\. Negotiate Attribution and Credit** Depending on your situation, you may want to receive credit when your dataset is used. If attribution matters to you, include it as a contractual requirement. Specify the exact form it should take (for example, in model cards, technical reports, or public announcements). Make sure you check what might be required by your own obligations; for example if that data has other stakeholders with their own rights. ### **5\. Build In Spaces for Checks, Audits, and Verification** This depends a little on your own capacity, but if you’re a larger organisation with some technical resources, you may want to have the right to verify that your data is being used in the way the license permits. This does not need to be intrusive, or heavyweight, but it should be real in order to protect you. Maybe think about including: - The right to request written confirmation of how the data was used - The right to commission a third-party audit if you have reasonable concerns - Requirements for the company to keep records of their data usage ### **6\. Get the Fair Value Exchange Part Right (this is the hardest part)** Pricing datasets is difficult. We maintain a repository of pricing data points and the fee structure of different licensing deals that we are happy to share with our data providers and allies upon request. There is no single right way to price a dataset license - it’s dependent on factors like the uniqueness, quality, annotation and potential application of the data. Common models include: - **Flat fee**: a one-time payment for a defined use - **Usage-based pricing**: fees tied to the scale of training runs or the number of models trained - **Revenue sharing**: a percentage of revenue from AI products trained on your data - **Subscription**: ongoing access fees for continued or updated data Deciding this is very custom to your context (how much data, how often it’s going to be updated, who you expect to use it, where in the training cycle etc). Do think expansively about what fair value exchange means to you: you might find it’s not money! Or not just money. It might be that the company lets your community use resulting tools for free for the next 5 years, or that they second an engineer to you for a period of time, or that they agree to an internship system for speakers of your language. Get creative! This type of collaboration agreement won’t necessarily live in the license, but you should think holistically about what would make working with this/any organisation a fair and exciting partnership for you. If they’re not thinking about it as partnership, maybe they aren’t the right fit. ### **7\. Be clear about the Intellectual Property ownership** Your license should leave no ambiguity about who owns what: - You retain full ownership of the underlying dataset - The company/user owns the model they build (this is standard) - Neither party gains IP rights over the other's pre-existing assets Something to consider is making it clear whether any derivative datasets or annotations the company creates from your data belong to them or you or a blended model. For example, say, a company cleans, labels, or augments your dataset as part of their process, the resulting enriched data might be very valuable. Decide upfront who it belongs to. ### **8\. Include Ethical Use Clauses** AI training raises real ethical questions. Consider adding clauses that prohibit the use of your data in projects that you would find problematic. Common examples include: - Surveillance systems or tools for discriminatory profiling - Models intended to generate disinformation - Weapons or systems used in armed conflict These clauses are increasingly common in data licenses and signal that you care about the downstream impact of your work. ### **9\. Agree on a Governing Law and Dispute Resolution Process** Cross-border data deals can get complicated quickly. Be clear about: - Which country's laws govern the contract - How disputes will be resolved (negotiation first, then arbitration or litigation) - Which jurisdiction courts would handle any legal proceedings - Timelines for dispute resolution, expected response times, etc ### **10\. Get Legal Counsel** AI data licensing is a specialist area of law, and the stakes – financial, ethical, and reputational – are significant. If you don’t already have any legal advice or expertise, consider engaging a lawyer with experience in data licensing and AI before you sign anything. **11\. Think about the wider world in which you want to live** Exclusively licensing your data to the single highest bidder may seem appealing. But consider the broader social impact this type of arrangement can have. It generally gate-keeps innovation, locking in monopolies and stifling a more thriving social and economic ecosystem. You might get a larger cheque now, but over the long term, a larger volume of smaller arrangements may in fact be more financially rewarding, and certainly may help to build a tech future that is more biodiverse and thriving. Platforms like Mozilla Data Collective exist for this diversified sharing context. You might also want to consider asking for provisions such as the right to open source the datasets for researchers after a set period, or donating them in whole or part to the public domain in the future. [Join Mozilla Data Collective](https://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com) We’d love to hear your feedback, questions and stories on [mozilladatacollective@mozillafoundation.org](mailto:mozilladatacollective@mozillafoundation.org) ### Behind the scenes: Integrating MDC datasets into your Python project URL: https://community.mozilladatacollective.com/behind-the-scenes-integrating-mdc-datasets-into-your-python-project/ Last updated: 2026-03-05T17:48:21.000Z ### Overcoming the complexity of AI Mozilla Data Collective helps communities to offer unique, multilingual, multicultural, and multimodal datasets. From transcribed and translated [videos of narrated Ekpeye folktales](https://datacollective.mozillafoundation.org/datasets/cmiohwz2t011hnx07urwsx55i ?ref=community.mozilladatacollective.com) to complex [question-answering text pairs for the Georgian language](https://datacollective.mozillafoundation.org/datasets/cmm0n37lm000dnq07vpctdtc9?ref=community.mozilladatacollective.com), the diversity of datasets on our platform is core to our mission. But with the inherent complexity of all these different data formats and representations comes a challenge: how do we make these datasets easily understood by machines? ### A Python and pandas solution As it stands, the most popular programming language for AI development is [Python](https://www.secondtalent.com/resources/top-ai-programming-languages-by-usage-stats/?ref=community.mozilladatacollective.com), and one of Python's most popular data manipulation frameworks is [pandas](https://pandas.pydata.org/?ref=community.mozilladatacollective.com). We therefore made it our priority to support automatic integration of MDC-hosted datasets, making them loadable as a pandas DataFrame with a single function call. From a user and developer perspective, this means that regardless of whether the dataset you want to use is text, video, or audio, in English or in Lingala, for LLM fine-tuning or for Automatic Speech Recognition, the only thing you need to do to integrate it into your Python project is: ```python from datacollective import load_dataset dataframe = load_dataset("my-dataset-id") ``` Single function call for download, extracting and parsing an MDC dataset into a pandas.DataFrame As simple as that 😌 ## Behind the scenes In this blog post we want to give you a glimpse behind the curtain of the inner workings of the [datacollective](https://pypi.org/project/datacollective/?ref=community.mozilladatacollective.com) Python package and how we make this seamless integration possible. There are three main questions that need to be addressed when preparing a dataset for processing: 1. **What kind of task will the dataset be used for?** (ASR, TTS, LM, etc.) 2. **How are the directories and files structured inside the dataset archive?** 3. **How do we map the files and their contents to specific columns in a DataFrame?** We decided the most effective way to answer these questions universally, across all datasets, is by defining a **`schema.yaml`** file for each dataset. ### Why a declarative YAML file? There are a few reasons we went with a schema-first approach rather than, say, shipping custom loader code per dataset. The three most prominent ones are: - **Security:** you never need to download and execute arbitrary code from the internet on your machine. All behaviour is driven by our **open-source** `datacollective` library itself meaning that the schema only describes *what*the data looks like, not *how* to process it. - **Decoupled releases**: because all schemas live in a public registry that the SDK queries at runtime, we can add support for new datasets without releasing a new version of the Python package every time. When a new dataset is released, a new schema file is added in the registry and supported seamlessly. - **Human-readable and community-friendly:** YAML is easy to read, write, and review, lowering the technical entry barrier for maintaining an accurate registry, without requiring any deep Python expertise. ### Answering the three questions 1. **Task type** The first question: **what is this dataset for?** is answered by the *`task`* field, which is set by the data provider when uploading the dataset. Currently supported tasks on the MDC platform include: *`Natural Language Processing, Automatic Speech Recognition, Language Identification, Machine Translation, Language Modelling, Large Language Modelling, Natural Language Understanding, Natural Language Generation, Computer-Aided Language Learning, Retrieval-Augmented Generation, Computer Vision, Machine Learning, Other `* . So far for simplicity, we ask each uploader to choose one task for each dataset to help people who are searching or filtering for a particular dataset for their purposes. 1. **File structure** The second question: ***how are the files laid out?*** is answered by what we call a **loading strategy** which the SDK infers from the fields present in the schema. There are three strategies: - **Index-based** (default): a metadata file (CSV / TSV / pipe-delimited) lists each sample. Key fields: `format`, `index_file`, `columns`. - **Multi-split:** multiple split files (train, dev, test, etc) each contain samples. Key fields: `root_strategy: "multi_split"`, `splits`. - **Paired-glob:** each audio file has a matching `.txt` sidecar and there is no index file at all. Key fields: `root_strategy: "paired_glob"`, `file_pattern`, `audio_extension`. 1. **Column mapping** The third question: ***what goes into each DataFrame column?*** is answered by the `columns` section of the schema. Each key under `columns` becomes a column name in the resulting DataFrame, and its value tells the SDK where to find the data and how to interpret it: ```yaml columns:   audio_path:     source_column: "path"       # column name in the index file     dtype: "file_path"          # resolved to an absolute path on disk   transcription:     source_column: "sentence"     dtype: "string"   speaker_id:     source_column: "client_id"     dtype: "category"     optional: true              # silently skipped if the column is missing ``` Supported dtype values include \`string\`, \`file\_path\` (resolved to an absolute path), \`category\` (pandas Categorical), \`int\`, and \`float\`. ### The `schema.yaml` in practice Here is an example of what a schema looks like for the [Ehugbo TTS: biblical text to speech dataset in Ehugbo Language](https://datacollective.mozillafoundation.org/datasets/cmihqro9h0238o207fgg5cmf6?ref=community.mozilladatacollective.com) dataset: ```python dataset_id: "cmihqro9h0238o207fgg5cmf6" task: "TTS" format: "csv" encoding: "utf-8-sig" checksum: "c29134fe715a9a794f44c94c36022f548e97d6551658bc02f9540e05c1f5f203" index_file: "metadata.csv" base_audio_path: "audios/" columns: audio_path: source_column: "Audio File Name (.wav)" dtype: "string" transcription: source_column: "Transcript" dtype: "string" book: source_column: "Book (audio folder)" dtype: "category" speaker_id: source_column: "Pseudo ID" dtype: "category" gender: source_column: "Gender" dtype: "category" optional: true duration: source_column: "Duration" dtype: "float" optional: true ``` Schema.yaml describing how to parse an MDC dataset This tells our `datacollective` package: *"Read `metadata.csv` as comma-separated (UTF-8 with BOM), map the `Audio File Name (.wav)` column to audio paths (under the `audios/` directory), `Transcript` to transcriptions, `Book (audio folder)` and `Pseudo ID` as categorical metadata columns for book and speaker, with optional `Gender` and `Duration` columns when present."* ### Limitations As our platform and datasets catalogue are ever-evolving, there are a few limitations with our current implementation as of right now (March 2026): - **Not all tasks are supported yet:** the python SDK currently handles ASR and TTS datasets but support for additional tasks is actively being worked on and will be released in the coming weeks! - **A dataset can serve multiple tasks**, but the current schema model supports declaring a single task per schema. Multi-task schemas are on the roadmap. - **Not every MDC dataset has an associated `schema.yaml` yet.** Coverage is growing, and community contributions are very welcome (see the call to contribute below!). ### Next steps The schema-based loading system is already working in production for more than 350+ MDC datasets, and we are excited about where it goes from here. These are the areas we are actively investing in: - **More tasks:** the SDK's task registry is designed to be extended with minimal friction by adding a new loader class - plus a one-line registration is all it takes. Machine Translation (MT), Language Modelling (LM), and other multimodal tasks are on the near-term roadmap. As the MDC dataset catalogue grows, so will the set of supported tasks. - **More schemas:** we are working through the catalogue systematically, and we encourage data providers and community members to contribute schemas for datasets they care about (see the call to action below). - **Broader data format support:** the AI ecosystem is evolving rapidly and we want to keep pace with it. We are planning to expand `load_dataset()` into a universal entry point that can emit data in whichever format best fits your pipeline. ### Want to contribute? We'd love your help to make `datacollective` better for everyone: > *I want to use dataset but its not supported by* `load_dataset()` yet! \-> [Open an Issue](https://github.com/Mozilla-Data-Collective/dataset-schema-registry/issues/new?template=request%5Fschema.yaml&ref=community.mozilladatacollective.com) in our repository to request it. We'll try to add support as quickly as possible! > I have uploaded my dataset at MDC and want to help the community use it by adding support for the `load_dataset()` function. What can I do? → Consider [opening a Pull Request](https://github.com/Mozilla-Data-Collective/dataset-schema-registry?ref=community.mozilladatacollective.com) with the `schema.yaml` for your dataset. We have step-by-step documentation to guide you through the process in this [link](https://mozilla-data-collective.github.io/datacollective-python/add%5Fnew%5Fschema/?ref=community.mozilladatacollective.com). > I have questions or ideas I want to share with you! → Feel free to [open an issue](https://github.com/Mozilla-Data-Collective/datacollective-python/issues?ref=community.mozilladatacollective.com) in our repository if it's something technical. Or you can find us on social media at the bottom of our website. **Feedback from our community is incredibly important to us!** If you found this post useful, don't forget to check out: - 📦 [The source code](https://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com) - 📖 [The full documentation](https://mozilla-data-collective.github.io/datacollective-python/?ref=community.mozilladatacollective.com) ### Cultural Heritage and AI: How Institutions Can Reclaim Control of Their Data URL: https://community.mozilladatacollective.com/cultural-heritage-and-ai-how-institutions-can-reclaim-control-of-their-data/ Last updated: 2026-03-03T18:36:42.000Z The institutions that safeguard humanity's cultural memory, galleries, libraries, archives, and museums (collectively known as the GLAM sector) are confronting a paradox that defines the current moment in AI development. Years of careful digitization of their archives have transformed physical collections into vast, machine-readable repositories of human knowledge. Yet the very openness that would make these collections a gift to the public now exposes them to extraction, commodification, and misuse by the AI industry's insatiable demand for training data. The result is a fundamental question of data governance: How can GLAM institutions maintain sustainable sovereignty over the datasets they are mandated to steward with care, while staying true to their mission of sharing the world's knowledge with the public? ## A Governance Crisis of Extractive Data Harvesting A 2025 report by [GLAM-E Lab](https://www.glamelab.org/what-we-do/?ref=community.mozilladatacollective.com) co-director [Michael Weinberg](https://www.glamelab.org/who-we-are/michael-weinberg/?ref=community.mozilladatacollective.com) titled [*Are AI Bots Knocking Cultural Heritage Offline?*](https://glamelab.org/products/are-ai-bots-knocking-cultural-heritage-offline/?ref=community.mozilladatacollective.com) documented a disturbing trend: automated bots deployed by AI companies are systematically scraping GLAM collections in aggressive swarms that overwhelm server infrastructure and sometimes knock collections entirely offline. The report showed that “Of 43 respondents, 39 had experienced a recent increase in traffic. Twenty-seven of the 39 respondents experiencing an increase in traffic attributed it to AI training data bots, with an additional seven believing that bots could be contributing to the traffic.” Respondents described servers reaching 100% CPU load within minutes and being rendered inoperable until bots moved on to the next target; while others reported sustained denial-of-service-level attacks from AI scrapers. Wikimedia reported that 65% of its most expensive traffic originated from bots, imposing systemic costs on an institution that depends on public support to remain operational. This phenomenon operates at two distinct levels: At the ethical level, it raises foundational questions about the meaning of "open" access in an era of commercial AI development, since rarely do institutions consent, tacitly or otherwise, to having their collections strip-mined by profit-driven AI laboratories without attribution, reciprocity, or institutional dialogue. The governance mechanisms currently available are for all practical purposes inadequate: robots.txt directives are routinely ignored and IP blocking is easily circumvented by bots rotating across hundreds of addresses simultaneously. At the technical level, it imposes increasing infrastructure costs on institutions many times operating under chronic budget constraints, adding further hardship to the [generalised funding cuts](https://news.artnet.com/art-world/museums-politics-2679093?ref=community.mozilladatacollective.com) happening across the world in the cultural and heritage sector. Institutions may thus find themselves helpless in the face of extractive data harvesting, or pressured to enter unfavorable data licensing deals as a way to generate new revenue. Since the legal framework governing AI training data remains deeply ambiguous across jurisdictions, GLAM entities seeking new revenue generation vehicles through their existing datasets may be rightly discouraged to find buyers or fearful of being locked into an exclusivity deal that would prevent them from sharing their data with other institutions. ## Mozilla Data Collective: A New Data Stewardship Paradigm It is precisely within this context that the [Mozilla Data Collective](https://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com) (MDC) emerges as a structurally significant intervention for the GLAM sector. Building on two foundational Mozilla projects, [Common Voice](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com) and the [Data Futures Lab](https://www.mozillafoundation.org/en/data-futures-lab/?ref=community.mozilladatacollective.com), MDC was officially launched at the [15th Mozilla Festival in Barcelona in November 2025](https://www.youtube.com/watch?v=t-8p2HNao%5F8&t=2044s&ref=community.mozilladatacollective.com) and is backed by multi-million-dollar seed funding and operates as the first social enterprise incubated by the [Mozilla Foundation](https://www.mozillafoundation.org/en/?ref=community.mozilladatacollective.com). MDC offers robust, secure, and controlled access to datasets and amplifies their visibility by featuring them alongside other high-value datasets. Its architecture is designed around a principle that stands in direct contrast to the extractive model currently exploited by commercial AI actors: contributors retain full ownership of their datasets and retain full control over the terms of access. Institutions can choose to share openly under existing licenses such as Creative Commons or NOODL, or build custom licensing frameworks tailored to their specific governance requirements. They can open data to all, or restrict access to specific categories of downloaders like academic researchers, non-commercial users, or values-aligned organizations. The datasets are still owned by their rightful owners, MDC is a self-service platform which gives creators full control. If users want to charge for a license to use their data, MDC doesn’t take a cut, they simply charge a modest 5% fee to downloaders, to cover the cost of storage and egress. The principle is anti-extractivist - organisations should continue to own, and be the primary beneficiaries of sharing their data, on their own terms. For GLAM institutions, the implications are transformative. MDC enables institutions to become active, intentional participants in the AI data economy and is designed to accommodate both the organizational and dataset diversity of the GLAM sector. MDC champions the inclusion of multicultural and multilingual data as a foundation for more equitable AI, curating datasets that span from [Indonesian podcast audio](https://datacollective.mozillafoundation.org/datasets?q=podcast&ref=community.mozilladatacollective.com) to [Tatar folklore](https://datacollective.mozillafoundation.org/datasets/cml8gixh60087o407lfgoumgu?ref=community.mozilladatacollective.com), from [public radio from Northern Uganda](https://datacollective.mozillafoundation.org/datasets/cmkfm6xtw00k2nv07oakesnix?ref=community.mozilladatacollective.com) to a speech corpus of testimonials from [Armenian refugees and immigrants](https://datacollective.mozillafoundation.org/datasets/cmi4ma15g09wykx07xbolzpoz?ref=community.mozilladatacollective.com). Institutions retain complete control over how their datasets are used. For example, many opt for the [Creative Commons Attribution Non-Commercial 4.0 International](https://spdx.org/licenses/CC-BY-NC-4.0.html?ref=community.mozilladatacollective.com) (CC-BY-NC-4.0) license but tailor it to suit their particular values. Some choose to forbid the use of their data in [systems intended for surveillance, profiling, or repression of individuals or communities](https://datacollective.mozillafoundation.org/datasets/cmll4tg5200lcl6070hpil2b9?ref=community.mozilladatacollective.com), while others forbid access to their datasets to companies with annual [revenue of more than one million USD](https://datacollective.mozillafoundation.org/datasets/cmhrlxa1100h3mn07l6tpvqwz?ref=community.mozilladatacollective.com). The people who access the datasets are authenticated and held to legally binding contracts to ensure the data is used as intended by the owning institution. ## From Digital Fuel to Cultural Infrastructure The GLAM sector stands at a critical point. The decades of investment that transformed physical collections into digital knowledge are now being leveraged (often without permission, compensation, or acknowledgment) by certain players in the AI industry. The sector's traditional commitment to openness, which has been one of its greatest contributions to public knowledge, has been turned against it by actors for whom cultural heritage is simply another input in the training data supply chain. The response should not be to retreat into restriction. Cultural heritage belongs to the public, and the GLAM sector's mission is to provide access to it. The challenge is to develop governance frameworks sophisticated enough to honor that mission while asserting the institutional agency necessary to ensure such valuable datasets are used by entities and in ways that follow the values and mission of the data providers. Mozilla Data Collective offers such a model by providing a platform where institutions can share, license, and be compensated for their data on their own terms, under legally binding frameworks, transforming the current asymmetry between cultural institutions and commercial AI into a genuine exchange. It empowers GLAM institutions to be active, valued, and sovereign participants in the AI data ecosystem by allowing them to create, curate, and control their datasets as they see fit. It allows GLAM institutions to govern their data in the ways that remain faithful to their mission and sense of purpose. In an age when the cultural record of humanity is being consumed to build the intelligence systems of the future, it is crucial for galleries, libraries, archives, and museums to have a powerful say in the matter. ### Using the MDC Python SDK Library to Download Datasets URL: https://community.mozilladatacollective.com/using-the-mdc-python-sdk-library-to-download-datasets/ Last updated: 2026-03-12T15:26:59.000Z *Recorded by* [*Kostis Saitas - Zarkias*](https://community.mozilladatacollective.com/author/kostis/)*, AI & Data Engineer at Mozilla Data Collective* 0:00 /5:05 1× ### **Prerequisites** 1. Create an account on Mozilla Data Collective and verify your email address Join Mozilla Data Collective → ### **Project Setup** 1. In your profile, create an API credential in [/profile/credentials](https://datacollective.mozillafoundation.org/profile/credentials?ref=community.mozilladatacollective.com). Ensure that you copy the secret key, as you will not be able to view it again once you close the credential creation window. 2. Save your API key in your project .`env` file as an environment variable 3. Install the latest version of the Mozilla Data Collective Python SDK Library - we recommend using a virtual environment ```python uv venv .myenv source .myenv/bin/activate uv pip install datacollective ``` ### Using the package in your project In this example, we prepare a dataset for fine-tuning a speech to text model by downloading, extracting, and bringing a Common Voice dataset into a `pandas` data frame using the following code: ```python from datacollective import load_dataset dataframe = load_dataset("", download_directory="data") ``` You will need replace `` with the dataset ID or slug for the dataset you want to download. To do this, you will need to agree to the terms and conditions for the dataset on the Mozilla Data Collective website. 💡 The interface for agreeing to dataset terms and conditions ensures that each downloader can carefully review the terms for each dataset, as set by the dataset provider, to ensure their use case aligns with the intended use of the data. You can verify that the dataset has been downloaded correctly by printing out the first few elements of the dataframe. ```python print(dataframe.head(5)) ``` ### Downloading datasets to your machine If you want to download a dataset and store it on your local machine, without using it in a specific project, you can do so with `download_dataset("`. 💡 Downloads are automatically resumable, which can be helpful when downloading large datasets ### Getting Dataset Details You can get the metadata associated with a given dataset using `get_dataset_details("`, which can show important details about a dataset before downloading. ### **Additional Links** - [MDC API Documentation](https://datacollective.mozillafoundation.org/api-reference/docs?ref=community.mozilladatacollective.com) - [Mozilla Data Collective Python API Library on PyPi](https://pypi.org/project/datacollective/?ref=community.mozilladatacollective.com) - [Mozilla-Data-Collective/datacollective-python on GitHub](https://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com) ### MDC Release Notes - 27.02.26 URL: https://community.mozilladatacollective.com/mdc-release-notes-27-02-26/ Last updated: 2026-02-27T17:28:25.000Z Hello, Mozilla Data Collective! 👋 This week, we've got a lot of new datasets to share with you, as well as a few on-platform updates and improvements that we've made. ### New Features and Changes - **There is a new filtering option for dataset searches.** Now, you can filter on Task, Language, License, and dataset Format when browsing through the dataset listings page, making it easier to narrow down the available datasets. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/02/Screenshot-2026-02-27-at-11.10.32---AM.png) - **New information requested when becoming an uploader.** When you request to upload datasets to MDC, you will now be asked to fill out a short form with additional information about your datasets. Previously, the process for being approved had a member of our team emailing to collect this information. By making it part of the request flow, our team can more quickly review requests and get uploaders approved sooner! ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/02/image.png) - **The MDC API can now return a dataset via its slug, rather than ID.** The dataset slug can be found in the 'Connect to API page' once terms and conditions have been agreed to for a given dataset. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/02/image-1.png) --- ### Fixes - The Common Voice English Spontaneous Speech 1.0 dataset has been made public to account for the fact that v2.0 was not released Join Mozilla Data Collective → --- ### New Datasets **Institute of African Digital Humanities** [Lingala-TTS-Dataset | Mozilla Data CollectiveThe dataset contains audio and text resources in Lingala, a Bantu language spoken in the Republic of the Congo (also known as ‘Congo Brazzaville’) and the Democratic Republic of the Congo (DRC). These resources are suitable for TTS and ASR tasks and consist of the following: - 8,572 audio clips totalling 4 hours, 25 minutes and 54 seconds; - an audio mapping file containing 8,572 lines.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-17.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-17.png)](https://datacollective.mozillafoundation.org/datasets/cmm23jslb003znq07vdska54l?ref=community.mozilladatacollective.com) **Kaltepetlahtol** [Daily Expressions in Highland Puebla Nahuatl | Mozilla Data CollectiveA corpus of more than 1,000 common expressions in Highland Puebla Nahuatl, annotated for child-directedness and code-switching. 80% of the phrases are accompanied by Spanish translations. All names mentioned in the sentences have been anonymized.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-20.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-20.png)](https://datacollective.mozillafoundation.org/datasets/cmm3vxr6r00bpmk07b7ayxdgx?ref=community.mozilladatacollective.com) [Zacatlán Tepetzintla Nahuatl ASR Dataset | Mozilla Data CollectiveAn ASR dataset of Zacatlán-Ahuacatlán-Tepetzintla (Western Sierra Puebla) Nahuatl, ISO 639-3 nhi. This is a derivative work of the Zacatlán Tepetzintla Nahuatl Audio and Transcriptions datasets. It consists of the subset of larger audio dataset with transcriptions (approximately 14 hours) converted to the Mozilla Common Voice Scripted Speech format. The original stereo audio has been split and aligned with the parsed transcriptions.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-10.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-10.png)](https://datacollective.mozillafoundation.org/datasets/cmls27zfd0043ma07mxvsz8zg?ref=community.mozilladatacollective.com) **MDC Community Concierge** [Finance Sentences - North American Spanish | Mozilla Data CollectiveThis is a public domain corpus of North American Spanish sentences in the finance domain. The corpus was collected in the second half of 2023 in aid of the Mozilla Common Voice project. The dataset contains 79,655 clean sentences (1,325,013 tokens) from nine distinct federal domains. It also contains a total of 209,061 sentences (4,125,637 tokens) of sentences without cleaning.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-21.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-21.png)](https://datacollective.mozillafoundation.org/datasets/cmm3gntdp00jymn07zjetwkyy?ref=community.mozilladatacollective.com) [Cuentos en Kʼicheʼ leídos en voz alta | Mozilla Data CollectiveUna colección de cuentos (audio y texto) en la lengua Kʼicheʼ. 1 hora 51 minutos de audio con 726 oraciones (8,283 palabras) de texto, del Currículo Nacional Base de Guatemala.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-24.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-24.png)](https://datacollective.mozillafoundation.org/datasets/cmm3j5gyb00l9nt073masycl3?ref=community.mozilladatacollective.com) [Cuentos en Mam leídos en voz alta | Mozilla Data CollectiveUna colección de cuentos (audio y texto) en la lengua Mam. 40 cuentos, un total de 1 hora 23 minutos de audio con 958 oraciones (7,441 palabras) de texto, del Currículo Nacional Base de Guatemala.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-25.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-25.png)](https://datacollective.mozillafoundation.org/datasets/cmm3oy19w004umh071icuxao8?ref=community.mozilladatacollective.com) **MDC Curators** [CorCenCC: Corpws Cenedlaethol Cymraeg Cyfoes | Mozilla Data CollectiveThe CorCenCC corpus contains over 11 million words (circa 14.4m tokens) from written, spoken and electronic (online, digital texts) Welsh language sources, taken from a range of genres, language varieties (regional and social) and contexts. The contributors to CorCenCC are representative of the over half a million Welsh speakers in the country. The creation of CorCenCC was a community-driven project, which offered users of Welsh an opportunity to be proactive in contributing to a Welsh language resource that reflects how Welsh is currently used.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-29.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-29.png)](https://datacollective.mozillafoundation.org/datasets/cmm3gotrz00j8nt07bsjz2znh?ref=community.mozilladatacollective.com) **MIT** [ATLAS Cross-Lingual Transfer Matrix | Mozilla Data CollectiveThis matrix is helpful for determining what languages to train a language model with. Given a Target Language, we hope to optimize the performance for (shown as rows), we estimate how beneficial it is to train with each Source Language (shown in the columns). The scores are empirically derived, from 750+ training experiments in the ATLAS multilingual scaling laws paper: https://arxiv.org/pdf/2510.22037\. Higher scores indicate greater synergy, whereas lower scores indicate more interference.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-11.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-11.png)](https://datacollective.mozillafoundation.org/datasets/cmlth9lrp000ams07yjdscjgu?ref=community.mozilladatacollective.com) **OpenCSG** [Finweb-Edu-Chinese-v2.2 | Mozilla Data CollectiveFineweb-Edu-Chinese v2.2 is the updated Fineweb-derived dataset of refined Chinese educational web content. It enhances content quality and expands education-stage coverage, fitting education-focused LLM training & educational AI tools. Get the dataset at www.opencsg.com.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-7.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-7.png)](https://datacollective.mozillafoundation.org/datasets/cmlqmsm0d00pinx072m0mm3aw?ref=community.mozilladatacollective.com) **Taruen** [Finland Public Domain 20th Century Literature Text Corpus | Mozilla Data CollectiveThis corpus contains a curated collection of public domain literature from Finland, featuring works by authors who died between 1901 and 1955\. The dataset captures the literary landscape of early 20th-century Finland and includes independent texts in both of the country’s official languages: Finnish (fi) and Swedish (sv). The texts were programmatically extracted from Project Lönnrot, a volunteer-driven digital library. To ensure linguistic relevance for modern NLP tasks, the extraction pipeline strictly filtered for works published in 1901 or later. Language codes for each text were dynamically detected using CLD algorithms. The corpus comprises approximately 69.1 million words across multiple plain text files, with each file prefaced by structured YAML front matter containing relevant metadata (title, author, year, source URL, language), followed by the original project’s boilerplate preamble enclosed in delimiter tags, and finally the literary text proper. All included works are fully in the public domain under Finnish and EU copyright law.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-30.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-30.png)](https://datacollective.mozillafoundation.org/datasets/cmm5078n50168mk07v64792sf?ref=community.mozilladatacollective.com) [Polish Public Domain 20th Century Literature Text Corpus | Mozilla Data CollectiveThis corpus contains a curated collection of 54 iconic Polish literary works, including major novels, sprawling multi-volume historical epics, and documentary prose from the late 19th and early 20th centuries. The dataset features the complete canonical works of literary titans such as Władysław Reymont, Stefan Żeromski, Henryk Sienkiewicz, Bolesław Prus, Józef Ignacy Kraszewski, Eliza Orzeszkowa, Tadeusz Dołęga-Mostowicz, and Zofia Nałkowska. All texts utilize modern Polish orthography (post-1936 standard) to ensure consistency and utility for training contemporary language models. The corpus comprises approximately 4.2 million words across multiple plain text files, with each file prefaced by structured YAML front matter containing relevant metadata (author, year, source URL). All included works are fully in the public domain under Polish law.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-15.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-15.png)](https://datacollective.mozillafoundation.org/datasets/cmm0nm2ua000eo007qs4r3m8q?ref=community.mozilladatacollective.com) [Dolgan Folklore Text Corpus | Mozilla Data CollectiveThis corpus contains a curated collection of 19 Dolgan fairy tales (15,618 words) digitized from a 2000 academic volume published in Novosibirsk. The Dolgans are the northernmost Turkic-speaking people, and their language is highly endangered. This dataset provides clean, structured data designed to catalyze machine learning and NLP research for low-resource languages, empowering the community to revitalize the Dolgan language. The corpus is structured in plain text files with YAML front matter metadata. It was created with the generous contribution of digitized texts from Karina Sheifer at Dartmouth College, with additional digitization and proofreading by Taruen.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-13.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-13.png)](https://datacollective.mozillafoundation.org/datasets/cmm0n4ro2000hnq079tknw6gv?ref=community.mozilladatacollective.com) [Kyrgyz Folklore Text Corpus | Mozilla Data CollectiveThis corpus contains a curated collection of Kyrgyz folklore texts, including fairy tales, magical tales, tales of everyday life, proverbs, sayings, and aphorisms. The content was digitized from 5 academic volumes published in Bishkek between 2016 and 2017, sourced from the electronic collections of the Central Scientific Library of the National Academy of Sciences of the Kyrgyz Republic. The corpus is structured in plain text files with YAML front matter metadata—notably utilizing the Common Turkic Alphabet for titles—and includes 427,527 words total (338,937 in tales, 71,619 in proverbs, and 16,971 in aphorisms). Text extraction was performed via OCR and LLM processing, followed by strict human proofreading to guarantee accuracy and exclude any hallucinations.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-6.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-6.png)](https://datacollective.mozillafoundation.org/datasets/cmlqoukmi000hnr07cprdmxsc?ref=community.mozilladatacollective.com) **Tbilisi State University** [GeoLogicQA: An LLM Benchmark for Logical Reasoning in Georgian | Mozilla Data CollectiveGeoLogicQA is a manually-curated logical and inferential reasoning dataset for the Georgian language (a Kartvelian language). Designed to evaluate deep language understanding, the dataset bypasses simple pattern recognition in favor of multi-step deduction, reading comprehension, and arithmetic problem-solving. It aims to address the gap in evaluation benchmarks for low-resource languages.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-16.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-16.png)](https://datacollective.mozillafoundation.org/datasets/cmm0n37lm000dnq07vpctdtc9?ref=community.mozilladatacollective.com) **Community** [Thorsten-Voice Dataset 2021.06 Emotional | Mozilla Data CollectiveThorsten-Voice Dataset 2021.06 (emotional) is a German emotional speech dataset recorded by Thorsten Müller and audio-optimized by Dominik Kreutz. It contains 2,400 recordings representing eight distinct emotions. The dataset is released under CC0 to enable unrestricted research and commercial use.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-28.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-28.png)](https://datacollective.mozillafoundation.org/datasets/cmm4b8f7700mgmh07cha3549n?ref=community.mozilladatacollective.com) [Thorsten-Voice Dataset 2022.10 | Mozilla Data CollectiveThorsten-Voice Dataset 2022.10 is a high-quality German neutral speech dataset recorded by Thorsten Müller and audio-optimized by Dominik Kreutz. It contains 12,450 phrases with more than 11 hours of clean speech audio. The dataset is released under CC0 to enable unrestricted research and commercial use.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-27.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-27.png)](https://datacollective.mozillafoundation.org/datasets/cmm4b8y6100nomk07w678zb9d?ref=community.mozilladatacollective.com) [Thorsten-Voice-44kHz-Full | Mozilla Data CollectiveTV-44kHz-Full is a high-quality German speech dataset containing approximately 40 hours of transcribed recordings (38,000+ files) by Thorsten Müller, a single native male speaker. It combines multiple Thorsten-Voice subsets (neutral, emotional, and Hessian dialect) in original 44.1 kHz sampling rate. The dataset is released under CC0 to enable unrestricted research and commercial use.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-26.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-26.png)](https://datacollective.mozillafoundation.org/datasets/cmm4de9w500ntmh073nx14k7p?ref=community.mozilladatacollective.com) [Persian VOA Corpus 2003-2008 | Mozilla Data CollectiveVOA news articles in Persian (Farsi) from 2003 to 2008\. Each entry begins with the original URL, the date of publication, and the headline, in the following format: # File: www... # Date: 2003-01-01 # Headline: ... The text of the article…![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-19.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-19.png)](https://datacollective.mozillafoundation.org/datasets/cmm27p8oy00auo3073kg3f1gj?ref=community.mozilladatacollective.com) [Thorsten-Voice Dataset 2021.02 | Mozilla Data CollectiveThorsten-Voice Dataset 2021.02 is a high-quality German neutral speech dataset recorded by Thorsten Müller and audio-optimized by Dominik Kreutz. It contains 22,668 phrases with more than 23 hours of clean speech audio. The dataset has been publicly available for several years and is released under CC0 to enable unrestricted research and commercial use.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-18.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-18.png)](https://datacollective.mozillafoundation.org/datasets/cmm2fwi490024nt0779vv9o0d?ref=community.mozilladatacollective.com) [Bojonegoro Javanese TTS | Mozilla Data CollectiveThe Bojonegoro Javanese TTS is a speech synthesis dataset containing more than 8 hours of audio recordings, recorded by native speakers of the Javanese language, specifically the Bojonegoro dialect from East Java, Indonesia. This dataset represents the Bojonegoro dialect of Javanese as well as variations of the Aneman dialect. The recordings cover a variety of everyday topics that describe phenomena in daily life in the Bojonegoro area, East Java, Indonesia. The language features in the Bojonegoro Javanese dialect include the use of code-mixing and code-switching between Javanese, Indonesian, and English, reflecting the multilingual nature of the community. This mixing occurs because certain words do not have equivalents in Javanese and are more commonly used in everyday conversation. Therefore, this dataset is suitable for non-commercial linguistic research and can be used to explore phonological and lexical variations in Javanese, particularly the Bojonegoro dialect in East Java, Indonesia.![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/icon/favicon-12.ico)Mozilla Data Collective![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/thumbnail/MDC-Preview-12.png)](https://datacollective.mozillafoundation.org/datasets/cmltnzkug0012mh07k1obic7v?ref=community.mozilladatacollective.com) ### MDC Release Notes - 13.02.26 URL: https://community.mozilladatacollective.com/mdc-release-notes-13-02-26/ Last updated: 2026-02-13T19:14:08.000Z Hello, Mozilla Data Collective! 👋 This week, we're sharing the first of what will become fortnightly release notes, where we'll share a record of major updates from the previous two weeks that have made it onto the platform. We want to hear from you! Fill out our survey to help inform our roadmap for 2026 and we'll send you a free ticket to this year's Mozilla Festival in Barcelona! [Complete the Survey ](https://docs.google.com/forms/d/e/1FAIpQLScSixY3a81-Cvjl0C%5F%5FJpdaub8Io2aKgq3Ti4EWwcjqrkTA9Q/viewform?ref=community.mozilladatacollective.com) ### New Features and Changes - **Dataset download statistics are now available for uploaders.** This feature will allow to you quickly view the status of your submissions, as well as how many people have downloaded your datasets. - **Published datasets can now receive updated .tar.gz files.** This removes the need to create new listings when datasets are updated. Instead, dataset owners can directly upload a new .tar.gz file and make edits to their submissions to published datasets from their upload portal. - **Dataset search was moved to the global navigation bar to make it easier to discover datasets from anywhere on the site.** We also changed the behavior of dataset searches to trigger on enter, instead of searching on incomplete queries. - **Draft datasets can now be deleted.** This should make it easier to clean up dataset upload listings in the upload portal. - **A link to create new uploads was added to the navigation bar and home page.** This should make it quicker to request uploader status, and to create new uploads/view the upload portal. - **Dataset search now operates across more of the datasheet.** This should make it easier to discover datasets that match a query based on a range of possible search terms. ### Fixes - An incomplete dataset migration caused some Common Voice datasets to be unavailable. This has been resolved and the two affected datasets should now be available - API calls made while unauthenticated now return a 401 instead of a 500 empty response - A duplicated paragraph in the terms of service was removed ### New Datasets - Institute of African Digital Humanities - [Adamawa Fulfulde-French Parallel Corpus of Narratives 1.2](https://datacollective.mozillafoundation.org/datasets/cml5asbhf009sme079y6sa9hm?ref=community.mozilladatacollective.com) - Digital Divide Data - [Khmer ASR Cultural Dataset (V2)](https://datacollective.mozillafoundation.org/datasets/cml6ywgg0007xmn07ppq469gt?ref=community.mozilladatacollective.com) - Collaborative Action for Research & Development - [Gawri (گاؤری) Magazine Corpus](https://datacollective.mozillafoundation.org/datasets/cmlgmqqul009lny07rhsa7aey?ref=community.mozilladatacollective.com) - Taruen - [Taruen's Tatar Folklore Text Corpus](https://datacollective.mozillafoundation.org/datasets/cml6ywgg0007xmn07ppq469gt?ref=community.mozilladatacollective.com) - [World Factbook (JSON)](https://datacollective.mozillafoundation.org/datasets/cmll3ucql00mfmn07x1oeiygu?ref=community.mozilladatacollective.com) - Balochistan Educational and Cultural Organization - [NAWA-E-WATAN Balochi Newspaper Corpus](https://datacollective.mozillafoundation.org/datasets/cmlgrqom000jrnx07zywfpblb?ref=community.mozilladatacollective.com) - [Western Balochi Literature Corpus](https://datacollective.mozillafoundation.org/datasets/cmlgv2ucp0005nx07wwocbpux?ref=community.mozilladatacollective.com) - [Talar (تلار) Barahui Magazine Corpus](https://datacollective.mozillafoundation.org/datasets/cmlgx0s7j000jmg07pr4o635t?ref=community.mozilladatacollective.com) - [Brahui Research Work Corpus](https://datacollective.mozillafoundation.org/datasets/cmlgx1kdz000nmg0744z7bbxl?ref=community.mozilladatacollective.com) - Forum for Language Initiatives - [Kohistani Shina Word List](https://datacollective.mozillafoundation.org/datasets/cmlgxhwha0015mg07ufma156h?ref=community.mozilladatacollective.com) - [Khowar Word List](https://datacollective.mozillafoundation.org/datasets/cmlgxqdl80019mg07p0197u76?ref=community.mozilladatacollective.com) - [Khowar Literature Corpus by FLI](https://datacollective.mozillafoundation.org/datasets/cmlgys7fj001knx077yktc0pd?ref=community.mozilladatacollective.com) - [Gojiri Literature Corpus](https://datacollective.mozillafoundation.org/datasets/cmlgz1vdt001omg07rv25gwnu?ref=community.mozilladatacollective.com) - Universidad Autónoma Nacional de México - [Trabajo de Campo - Huave](https://datacollective.mozillafoundation.org/datasets/cmll3ucql00mfmn07x1oeiygu?ref=community.mozilladatacollective.com) - Balochi Academy - [Eastern Balochi Literature Corpus](https://datacollective.mozillafoundation.org/datasets/cmll3ucql00mfmn07x1oeiygu?ref=community.mozilladatacollective.com) - Nick Fox-Gieg - [ABC-Draco](https://datacollective.mozillafoundation.org/datasets/cmll4i2n400l2l607uze0glc2?ref=community.mozilladatacollective.com) - Community Datasets - [TTS Javanese-Lumajang Dialect](https://datacollective.mozillafoundation.org/datasets/cml5bgysg00bhkr07g23kewke?ref=community.mozilladatacollective.com) - [TTS Central Javanese](https://datacollective.mozillafoundation.org/datasets/cml5bgysg00bhkr07g23kewke?ref=community.mozilladatacollective.com) - [Mandar Spontaneous Speech](https://datacollective.mozillafoundation.org/datasets/cml6ywgg0007xmn07ppq469gt?ref=community.mozilladatacollective.com) - [TTS-Tolaki](https://datacollective.mozillafoundation.org/datasets/cml6ywgg0007xmn07ppq469gt?ref=community.mozilladatacollective.com) - [Betawi TTS of Cultural Language (BEKAL)](https://datacollective.mozillafoundation.org/datasets/cml9hmuis017yo407k0p4i0t4?ref=community.mozilladatacollective.com) - [TTS Sasak Language](https://datacollective.mozillafoundation.org/datasets/cml9hmuis017yo407k0p4i0t4?ref=community.mozilladatacollective.com) - [TTS Javanese - Ngapak Dialect](https://datacollective.mozillafoundation.org/datasets/cml9hmuis017yo407k0p4i0t4?ref=community.mozilladatacollective.com) - [Manggarai Language for NLP ](https://datacollective.mozillafoundation.org/datasets/cmll3ucql00mfmn07x1oeiygu?ref=community.mozilladatacollective.com) Join Mozilla Data Collective → ### Your Data, Your Rules: A Community Workshop for Dataset Governance URL: https://community.mozilladatacollective.com/your-data-your-rules-a-community-workshop-for-dataset-governance/ Last updated: 2026-02-04T11:08:13.000Z # *A practical guide for communities creating datasets together—no legal expertise required. Based on our data governance workshop at Mozilla Festival Zambia 2024* --- Your community has created something valuable: a dataset. Maybe it's voice recordings in your language. Maybe it's traditional knowledge, local photographs, or cultural stories. Whatever it is, you built it together. Now comes the hard question: **What happens to it next?** Who gets to use it? For what purposes? Do they need to pay? Credit you? Ask permission first? These aren't just legal questions—they're questions about power, values, and what your community wants to see in the world. But you don't need to be a lawyer to figure this out. You just need to have honest conversations with your community. This guide walks you through a participatory workshop to do exactly that. --- ## Before You Start **Who should be in the room?** Bring together people who contributed to the dataset, community leaders, and anyone who cares about how the data gets used. Aim for 8-20 people—enough for diverse perspectives, small enough for real conversation. **What you'll need:** - A facilitator (can be anyone comfortable guiding discussion) - Colored cards or sticky notes (green, yellow, red) - A large wall or table for sorting - About 2-3 hours - Snacks (always snacks) --- ## Part 1: The Feeling Check (45 minutes) Before diving into legal language, start with gut reactions. This exercise surfaces what your community actually cares about—which might surprise you. ### How it works Read out each scenario below, one at a time. Give people 30 seconds to think silently, then everyone holds up a card: - 🟢 **Green:** This is good / I like it - 🟡 **Yellow:** This is okay / I don't mind - 🔴 **Red:** This is bad / I don't like it After each reveal, spend 2-3 minutes discussing: *Why did people feel differently? What concerns came up?* ### The Scenarios **Scenario 1:** A large tech company downloads your data and uses it to expand free tools like translators that anyone can use. **Scenario 2:** A large tech company downloads your data and uses it to build commercial services, charging people a fee to use them. **Scenario 3:** A local nonprofit uses your data to build tools for social good in your community. **Scenario 4:** Researchers use your dataset extensively and produce important findings, but they never cite, credit, or acknowledge your community publicly. **Scenario 5:** Engineers use your dataset to build somewhat controversial tools—like software that guesses someone's gender or nationality from their name or voice. **Scenario 6:** A local startup uses your data to build commercial products in your community's language—creating jobs locally, but also making profit. **Scenario 7:** People who contribute data get paid when they create it. **Scenario 8:** People who contribute data receive non-cash benefits: training programs, computing resources, or job opportunities. **Scenario 9:** A government uses your data to improve helpful citizen services. Later, after an election, a new government uses the same data to try to identify and monitor certain communities. **Scenario 10:** An outside company hosts a copy of your dataset for free—saving you money, but you lose visibility into who downloads it. **Scenario 11:** Your licensing rules become so complicated that almost no one uses the dataset anymore. ### After all scenarios Take 10 minutes to reflect together: - Where did we mostly agree? - Where were we most divided? - What themes kept coming up? --- ## Part 2: Building Your License (60 minutes) Now that you've surfaced your values, let's get concrete. A data license is just a document that tells people: "Here's what you can and can't do with our stuff." You're going to build one together—not by writing legal text, but by sorting what matters to you. ### The Building Blocks Write each of these on separate cards and spread them on a table: **Who can use it?** - Anyone - Only nonprofits - Only organizations in our region - Only small organizations (under X employees) - Only organizations we approve one-by-one - Only researchers/academics **What can they use it for?** - Anything at all - Only non-commercial purposes - Only research - Only purposes that benefit our community - Only specific applications we approve **What can they NOT use it for?** - Military or weapons - Surveillance - Tools that could discriminate against people - Training AI without explicit permission - Purposes that harm our community **Do they need to give us credit?** - Yes, always—publicly and prominently - Yes, but just in documentation - Not required **Do they need to pay?** - Never—always free - Free for nonprofits, paid for corporations - Everyone pays a licensing fee - Sliding scale based on organization size - Percentage of any revenue they make using it **What happens if someone breaks the rules?** - We can revoke their access - They face legal penalties - We publicly name them - We haven't thought about this yet **Who decides about edge cases?** - Our community makes decisions together - A small committee we elect - The organization hosting the data - Decisions are automatic based on written rules ### How to sort them Create three zones on your table or wall: - **WANT:** This should definitely be in our license - **DON'T WANT:** We don't want this - **UNSURE:** We need to discuss more / it depends Work through the cards together. For each one, discuss briefly and place it in a zone. It's okay to have lots in "unsure"—that's where the real learning happens. --- ## Part 3: The Hard Trade-offs (30 minutes) Here's the uncomfortable truth: you can't have everything. Some choices conflict with others. Discuss these tensions as a group: **Control vs. Impact** The more restrictions you add, the fewer people will use your dataset. A dataset nobody uses doesn't help anyone—but a dataset used in harmful ways doesn't either. Where's your balance point? **Money vs. Mission** Charging fees can sustain your community and feels fair. But it might price out the small local organizations you most want to help. How do you handle this? **Simplicity vs. Precision** A simple license is easy to understand and comply with. A detailed license protects against more misuse scenarios. Which matters more to you? **Trust vs. Verification** Do you trust users to follow the rules, or do you need ways to check? Verification takes resources. Is it worth it? --- ## Part 4: What's Your Starting Point? (15 minutes) Based on everything you've discussed, try to summarize your community's position in plain language. Something like: > "We want our dataset to be used freely by nonprofits and researchers, but commercial companies should pay a fee. Everyone must credit our community. We don't want the data used for surveillance or discrimination. A small committee will review requests from large corporations." This isn't legal text—it's a values statement. You can later work with legal experts to turn it into proper licensing language, but this gives them clear direction. --- ## What Comes Next Your workshop outputs—the scenario reactions, the sorted cards, and your values statement—are gold. Here's what to do with them: 1. **Document everything.** Take photos of your sorted cards. Write up the values statement. 2. **Share with your community.** Not everyone could be in the room. Get feedback from the broader group. 3. **Find legal help.** Organizations like Creative Commons, Mozilla, and various digital rights groups can help translate your values into actual license language. 4. **Revisit regularly.** Your values might change. Technology will definitely change. Plan to have this conversation again in a year or two. --- ## Common Questions **"Can we just use an existing license like Creative Commons?"** You can! Existing licenses are simpler and more widely understood. But they might not capture everything your community cares about. Many communities use a standard license as a base and add specific terms on top. **"What if someone ignores our license?"** This is hard. Enforcement depends on your resources and the legal system in your jurisdiction. Some communities focus on building relationships and trust rather than enforcement. Others partner with larger organizations that can help with legal action if needed. **"What if we can't agree?"** That's normal. Start with what you do agree on. For contested areas, consider: Can you pilot different approaches? Can you revisit the decision in six months with more information? Is there a compromise position? **"Do we need a lawyer for this?"** For the workshop? No. For the final license document? Ideally yes, especially if you want it to be legally enforceable. But going to a lawyer with clear values and priorities is much better than going with nothing. --- ## A Final Thought Your dataset exists because your community came together to create something. The license should reflect that same spirit—collective, intentional, and rooted in what you actually care about. There's no perfect answer. There's just the answer that's right for your community, right now. Good luck! Your data, your rules. --- *This workshop framework is adapted from Mozilla Data Collective work on alternative data licenses. For support developing your license, reach out to us at mozilladatacollective@mozillafoundation.org* ### Help Shape the Future of Mozilla Data Collective URL: https://community.mozilladatacollective.com/help-shape-the-future-of-mozilla-data-collective/ Last updated: 2026-02-03T13:19:57.000Z Mozilla Data Collective has an amazing opportunity for you to get a **free ticket to the 2026 Mozilla Festival** in beautiful Barcelona, Spain this November, 2026\. We are looking for feedback to inform our 2026 roadmap. Help shape the future of ethical data-sharing by filling out the form below ### FAQ: How can I remove published datasets from MDC? URL: https://community.mozilladatacollective.com/faq-how-can-i-remov/ Last updated: 2026-05-14T15:35:25.000Z If you need to remove a dataset from Mozilla Data Collective after it has been published, you can make your dataset private through the following steps: 1. Sign into your account - make sure you are using the account that published the dataset originally 2. Go to your Profile > Uploads and find the dataset that you want to unpublish 3. Select the listing and scroll to the bottom of the page 4. Click the 'Make Dataset Private' button 5. Confirm that you want to remove the dataset listing from MDC in the dialog box that opens ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2026/01/image-1-.png) Once you have unpublished your dataset, it will no longer be accessible to the public on Mozilla Data Collective. You can re-publish your dataset by clicking 'Make Dataset Public' again after it has been made private. ### Turning Your Data Into a Valuable ML Resource Without Giving Up Control URL: https://community.mozilladatacollective.com/turning-your-data-into-a-valuable-ml-resource-without-giving-up-control/ Last updated: 2026-01-28T06:42:30.000Z ## **You might be sitting on something precious** The modern world runs on data. One unfortunate result of this is the fact that many of us are unknowingly producing data for third-party companies, who use our content and actions as data points to make AI models that they then sell back to us in exchange for a subscription fee. On the one hand, making more quality data available for ML and AI training can result in better, more representative models. On the other hand, it should be possible to make our data available without losing control of it outright—and it should also be possible to be compensated for this contribution. At Mozilla Data Collective, we want to make it easier for people and organizations to open up their data to ML and AI model training while giving them control over who accesses their data and how it is used. Many researchers, organizations, and individuals may be in possession of large amounts of data that, with just a little work, could have a big impact on artificial intelligence outcomes. This guide will walk you through some important considerations to help you identify what kind of data you already have, what you can do to maximize its impact for ML research, and how to ensure you benefit from sharing it. ## **Identifying what you have** Sometimes people or organizations tell us that they want to support a better tech ecosystem, but they don't think they have any datasets. Actually, pretty much everybody does. Whether it's B-roll from your video shoots, archives of documents in different languages, those reports you got translated, audio recordings of oral histories, videos of interviews, telemetry data, or workflow data from all those files you categorized—believe us, you can be part of the everyone-a-builder revolution. Here's how to start thinking about what you've got. **Define the subject and scope.** What's the core domain of your data? Language datasets, opinion polls, multimodal content, home assistant logs? Getting clear on this helps you understand who might find it useful. **Consider uniqueness and value.** What problem does this data solve, or what gap does it fill? How is it different from existing public datasets? Data that captures underrepresented languages, specialized domains, or real-world use cases can be especially valuable. **Think through licensing.** This is where you can get creative. Your data can be a lever for change. You can invest it in line with your values—give it to an organization you trust, withhold it from someone you don't. Consider ownership and any existing licensing constraints, and think about what terms make sense for your goals. It is also critical to ensure that any data you share fully complies with all relevant terms of service, ownership rights, and ethical guidelines. This includes strictly avoiding the publication of personal information belonging to others without their explicit, informed consent. For more details, see our [Terms of Service](https://datacollective.mozillafoundation.org/terms/providers?ref=community.mozilladatacollective.com) for data providers. **Explore fair value exchange.** Can or should you charge for dataset use? Is there another form of fair value exchange that would make more sense for you—attribution, reciprocal data sharing, or something else? ## **Investigating quality and scale** To get a sense of the extent to which your data might be useful for machine learning, it's important to assess its size, consistency, representativeness, and metadata. **Scale assessment.** When dealing with datasets, there's no one-size-fits-all solution. Depending on the task, use case, and model, expectations around data volume vary. Here are some rough guidelines for common NLP tasks: - Text-to-speech (TTS): more than 5 hours of clear, deliberately-read speech recordings - Automatic speech recognition (ASR): more than 10 hours of transcribed recordings - LLM fine-tuning: more than 500k tokens - Speech language models: more than 500 hours of untranscribed speech - Text classification: more than 1,000 examples - Morphosyntactic annotation: more than 10k tokens These are broad guidelines. A dataset for any of these tasks could still be quite useful without meeting these volume recommendations. **Data integrity.** Check for missing values, corrupted files, and duplicate entries. Clean data is usable data. **Consistency, representativeness, and bias.** If your data has human annotations, how were these collected? Is there an understanding of inter-annotator agreement? How well does the dataset represent the target phenomena that will be modeled? Given the source of the data and annotators, what are the sources of bias? All human annotations contain biases, but understanding their nature is incredibly important for anyone who wants to use the data responsibly. **Metadata.** An important feature of the datasets published on Mozilla Data Collective is the rich, informative datasheet. Assess what metadata exists for your dataset. Did you create it directly? Good—you should be able to write a great datasheet detailing the process. If your dataset is an extension of existing data (for example, annotating publicly available texts), it's important to have information about the source data as well as the process of extending it. Details like language info, domain, and collection context are valuable in informing potential users about the nature of the dataset. We recommend taking a look at some of our existing datasheets to see what you might want to include in your own. ## **Enhancing usability for ML** Often, data is collected and stored in a way that reflects its original use. A linguist transcribing and annotating language data will store it in a format that works with their preferred annotation program (ELAN, Praat, Transcriber), because that makes the most sense for their workflow. However, the specific machine learning tasks your data could support will likely require—or benefit from—some task-specific considerations. **Segmentation.** Think about how to break down large or complex data into manageable, logical segments. This might mean organizing by date, experiment, category, or some other structure that makes the data easier to work with. **Format and structure.** What state is the data in? CSV, JSON, proprietary binary, raw images or video? Is it structured, semi-structured, or unstructured? Is the format consistent throughout, or varied? Getting clear on this helps you understand what preparation work might be needed. ## **Support you can expect from MDC** You don't have to figure this out alone. Mozilla Data Collective offers support to help you get your data ready for publication. **Technical guidance.** We can assist with data formatting, validation, and storage best practices. If you're unsure how to package your data, we're here to help. **Datasheet guidance.** We'll help you understand what to include in your datasheet and point you to our upload manual for detailed expectations. **Visibility.** MDC increases the exposure of published datasets to the broader ML community. By publishing with us, your data becomes discoverable to researchers and developers who can put it to good use—on terms you've defined. ## **Next steps** Ready to contribute? Create an MDC account and start the upload process. If you're worried about packaging your data or have questions about any of the considerations we've discussed here, get in touch at mozilladatacollective@mozillafoundation.org. We're happy to help. Your data has value. By sharing it thoughtfully—with control over who accesses it and how it's used—you're contributing to more representative ML models and a healthier tech ecosystem. This is collaborative work, and we're glad you're part of it. ### Latest: Asian Language datasets by communities on Mozilla Data Collective URL: https://community.mozilladatacollective.com/latest-asian-language-datasets-by-communities-on-mozilla-data-collective/ Last updated: 2026-01-27T15:34:05.000Z The internet belongs to everyone—but right now, it doesn't work for everyone. From the islands of Borneo to the mountains of Pakistan, hundreds of millions of people speak languages that AI simply can't understand. That's a problem we can solve together. At Mozilla Data Collective, we're proud to host a growing collection of open datasets that are helping researchers, developers, and communities build language technology for Asian languages. These datasets represent countless hours of work by linguists, native speakers, and organizations committed to ensuring no language gets left behind. Here's what's available today. ## Large Language Model Training [**chinese-cosmopedia**](https://datacollective.mozillafoundation.org/datasets/cmkpon5es03j2mg07j2kppzn7?ref=community.mozilladatacollective.com) — OpenCSG's large-scale Chinese text dataset containing approximately 15 million entries (≈60B tokens) covering encyclopedia, education, and multi-domain content. Cleaned and deduplicated for LLM pretraining. *6.09 GB, Apache-2.0* [**smoltalk-chinese**](https://datacollective.mozillafoundation.org/datasets/cmkplyrln03jknw07u3iy35rm?ref=community.mozilladatacollective.com) — A multi-task Chinese conversational dataset covering 19 typical dialogue scenarios, built by OpenCSG. *879.81 MB, Apache-2.0* --- ## Machine Translation [**English–Punjabi (Shahmukhi) Parallel Corpus**](https://datacollective.mozillafoundation.org/datasets/cmkh9rso90076nv076jwxgjv3?ref=community.mozilladatacollective.com) — 30,405 professionally translated sentence pairs from Mediamen archives, supporting machine translation and Punjabi language technology for real-world, contemporary language use. *1.08 MB, CC-BY-NC-4.0* [**Multilingual Religious Parallel Corpus**](https://datacollective.mozillafoundation.org/datasets/cmk1bhogs3htwmk07wo7o9p6y?ref=community.mozilladatacollective.com) — Kaleem Art Press compiled 6,465 aligned sentence units (\~0.98 million words) across Arabic, Urdu, Saraiki, Punjabi (Shahmukhi), and English. *2.27 MB, CC-BY-SA-4.0* --- ## Text Corpora [**Corpus of Panjebar Semangat Javanese-Language Magazine**](https://datacollective.mozillafoundation.org/datasets/cmk719jvs02z1nt07tjswt52s?ref=community.mozilladatacollective.com) — Three years of popular articles from the Javanese weekly magazine Panjebar Semangat, founded in 1933 by national hero Dr. Soetomo. Reflects contemporary themes in the Mataram style of Javanese. *4.31 MB, CC-BY-SA-4.0* [**Sindh Line Publishers Corpus**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — 1.029 million tokens from the Sindhi newspaper Sindh Line (2024-2025), including headlines, editorials, finance news, and advertisements from Karachi, Pakistan. *2.22 MB, CC-BY-SA-4.0* --- ## Speech Recognition [**Khmer ASR Cultural Dataset**](https://datacollective.mozillafoundation.org/datasets/cmkcy8in2004umo0775mye43g?ref=community.mozilladatacollective.com) — Digital Divide Data curated 37.62 hours of speech-text pairs from native Khmer speakers discussing Cambodian cultural topics. Includes speaker metadata for gender, age group, and origin city. *12.59 GB, CC-BY-SA-4.0* [**Common Voice Spontaneous Speech 2.0 - Kenyah**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Kenyah, an Austronesian language of Borneo. *212.06 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Melanau**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Melanau, spoken in Sarawak, Malaysia. *208.47 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Western Penan**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Western Penan, a language of the nomadic and semi-nomadic Penan people of Borneo. *247.12 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Kelabit**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Kelabit, spoken in the highlands of Sarawak. *193.77 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Serian Bidayuh**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Serian Bidayuh from Sarawak, Malaysia. *199.91 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Sabah Malay**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Sabah Malay, a regional variety from Malaysian Borneo. *275.80 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Bahasa Malay**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Bahasa Malay from Malaysia. *125.94 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Gorani**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Gorani, an Iranian language spoken in the border regions of Iran and Iraq. *224.46 MB, CC0-1.0* [**Common Voice Spontaneous Speech 2.0 - Ushojo**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Spontaneous spoken phrases in Ushojo, a Dardic language of northern Pakistan. *102.83 MB, CC0-1.0* --- ## Why This Matters Every dataset here represents a step toward a more inclusive internet. When speech recognition works for Khmer speakers, when translation tools support Punjabi, when AI can process Javanese—technology becomes a bridge instead of a barrier. This collection spans languages spoken by hundreds of millions of people across Southeast Asia, South Asia, and the Pacific. From major languages like Chinese to endangered languages of Borneo, these datasets ensure that communities have the resources they need to build technology that works for them. This work isn't done. We need more contributors, more languages, more voices. If you're working on Asian language data, we'd love to hear from you. [**Explore all datasets →**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) [**Get in touch →**](mailto:mozilladatacollective@mozillafoundation.org) ### Latest: African Language datasets by communities on Mozilla Data Collective URL: https://community.mozilladatacollective.com/latest-african-language-datasets-by-communities-on-mozilla-data-collective/ Last updated: 2026-01-27T15:31:34.000Z The internet belongs to everyone—but right now, it doesn't work for everyone. Millions of people speak languages that AI simply can't understand, and that's a problem we can solve together. At Mozilla Data Collective, we're proud to host a growing collection of open datasets that are helping researchers, developers, and communities build language technology for African languages. These datasets represent countless hours of work by linguists, native speakers, and organizations committed to ensuring no language gets left behind. --- ## Speech Recognition [**Luhya ASR Data (70 hours)**](https://datacollective.mozillafoundation.org/datasets/cmkh2stj6004mmj0751rn8vq6?ref=community.mozilladatacollective.com) — Digital Divide Data collected this substantial corpus of Luhya speech in Kenya. Native speakers recorded sentences to support automatic speech recognition research for this low-resource language. *13.90 GB, CC-BY-4.0* [**DhoNam: Dholuo Speech Dataset**](https://datacollective.mozillafoundation.org/datasets/cmjepxo6t08nmmk07iauvua6v?ref=community.mozilladatacollective.com) — The Maseno Centre for Applied Artificial Intelligence built this 51-hour corpus to supercharge ASR for Dholuo, one of Kenya's major indigenous languages spoken by over 4 million people. Created in partnership with the Dholuo community, who helped determine the licensing framework. *2.49 GB, NOODL-1.0* [**DataTrust Africa: Northern Uganda Radio Corpus**](https://datacollective.mozillafoundation.org/datasets/cmkfm6xtw00k2nv07oakesnix?ref=community.mozilladatacollective.com) — Amara Hub curated over 350 clips from public radio stations including Mega 100 FM, Q FM, Radio Pacis, and Radio Rupiny. Recordings in Acholi, Lango, Lugbara, and Akaramajong are on the way. *179.82 MB, NOODL-1.0* [**Spoken Congolese French Dataset**](https://datacollective.mozillafoundation.org/datasets/cmk1bdl0q39v6mb07gee6udp2?ref=community.mozilladatacollective.com) — Semi-guided interviews from Brazzaville, paired with orthographic transcriptions. A valuable resource for understanding regional French variation in the Republic of the Congo. *3.44 GB, NOODL-1.0* --- ## Machine Translation [**Adamawa Fulfulde–French Parallel Corpus**](https://datacollective.mozillafoundation.org/datasets/cmkmp2jh7019mnw07c2uudmub?ref=community.mozilladatacollective.com) — The Institute of African Digital Humanities compiled 1,977 lines of Fulfulde narratives with French translations, supporting translation research for this Afro-Asiatic language spoken across Central Africa. *112.50 KB, NOODL-1.0* [**Bamun-French Parallel Corpus**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Transcribed audio paired with French translations for Shupamem, spoken in Cameroon. *99.24 KB, NOODL-1.0* --- ## Language Documentation [**Ewondo Fong Multimodal Dataset**](https://datacollective.mozillafoundation.org/datasets/cmkmoepzx018wnw07cnkupls3?ref=community.mozilladatacollective.com) — A multimodal resource for the Fong variety of Ewondo, pairing example sentences with audio recordings and French translations. Built to support speech and language technology for under-resourced African languages. *16.80 MB, NOODL-1.0* [**Ewondo Mbida-Mbani Dataset**](https://datacollective.mozillafoundation.org/datasets/cmk1bbz7o39v2mb07jjupvzud?ref=community.mozilladatacollective.com) — Lexical entries from the Mbida-Mbani speech area, each accompanied by illustrative sentences, word-by-word glosses, French translations, and aligned audio. *19.25 MB, NOODL-1.0* [**Mada Narratives**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) — Seventeen transcribed oral narratives in Mada, an Afro-Asiatic language of Cameroon. These texts capture natural spoken discourse and traditional storytelling. *65.04 KB, NOODL-1.0* --- ## Why we're so excited to work with our communities Every dataset here represents a step toward a more inclusive internet! When speech recognition works for Dholuo speakers, when translation tools support Fulfulde, when AI can process Ewondo—technology becomes a bridge instead of a barrier. This work isn't done. We need more contributors, more languages, more voices. If you're working on African language data, we'd love to hear from you. [**Explore all datasets →**](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com) [**Get in touch →**](mailto:mozilladatacollective@mozillafoundation.org) ### Updates to MDC REST API and Python Library URL: https://community.mozilladatacollective.com/updates-to-mdc-rest-api-and-python-library/ Last updated: 2026-01-15T17:43:41.000Z Since we launched in November 2025, our [Python library](http://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com) and [REST API](https://datacollective.mozillafoundation.org/api-reference/docs?ref=community.mozilladatacollective.com) have seen incredible usage. We're shipping over 65 terabytes and ever increasing worth of data a month to downloaders across the globe, with much of it occurring programmatically through the REST API. But with that growth, we spotted some friction. Namely, the initial download approach was overly complex and slow. Our application servers were streaming datasets to users, which slowed download speeds and bottlenecked traffic. It also increased our costs without adding value for downloaders. So we made it simpler. ### **What's Changed** We've deprecated the `/datasets/{datasetId}/download/{token}` endpoint, it now returns a 410 HTTP status code. While the `/datasets/{datasetId}/download` response remains the same schema-wise, but now includes a *direct* download URL to our storage layer through the `downloadUrl` key. No more token management required. Instead of the two-step process where you'd: 1. POST to get a **token** 2. GET with that **token** to stream the file through our servers You now simply: 1. POST to get a direct, time-limited download URL 2. Download directly from our storage layer This means faster downloads, better reliability, and simpler code. The POST response still includes an `expiresAt` timestamp so you know exactly how long it's valid. ### **What You Need to Do** **If you're using the Python library:** Update to version 0.2.0 or newer. The changes are handled automatically under the hood, no code changes required from you! **If you’re calling the API directly:** Review if the changes above impact on your implementation. ### FAQ: How long does it take to publish my dataset on MDC? URL: https://community.mozilladatacollective.com/faq-how-long-does-it-take-to-publish-my-dataset-on-mdc/ Last updated: 2026-05-14T15:35:14.000Z We review every request to become a data provider on Mozilla Data Collective. **Reviewing Uploader Requests** If you have not already been in contact with our team about uploading data to Mozilla Data Collective, a member of our team will reach out to you via email to discuss your dataset and goals with sharing it. We aim to reach out to everyone who requests to upload to Mozilla Data Collective within a week of receiving the request, though holidays and busy times for the teams may extend this. **Reviewing Datasets** Once you have completed your datasheet, uploaded your dataset, and submitted for review, our team will aim to be in touch with you within a week or two and either a) approve your dataset, at which point it will be live on the platform, or b) contact you with any changes that need to be made before publishing. ### Pashto becomes third-highest language by volume of data in Common Voice v24 URL: https://community.mozilladatacollective.com/pashto-becomes-third-highest-language-by-volume-of-data-in-common-voice-v24/ Last updated: 2026-01-16T21:05:18.000Z We’re delighted to bring you a new dataset release for Common Voice version 24.0 for Scripted Speech, and version 2.0 for Spontaneous Speech. In this blog post, we present key highlights of the release, and show you where you can download the new datasets. ## Huge congratulations to the Pashto community for contributing nearly 3000 hours of speech data Leading our Common Voice v24 Scripted Speech release are the Pashto language community, who have rocketed to the number three spot by volume of data recorded. In just over three months, the Pashto community have recorded nearly 3,000 hours of audio recordings from nearly 5200 additional speakers, and have validated nearly 1000 hours of recordings. Pashto is a language spoken in north-western Pakistan and southern and eastern Afghanistan by over 50 million speakers. It is an official language of Afghanistan, alongside Dari. For comparison, the only two languages on Common Voice with more data are Catalan and English. Both these languages have been active on Common Voice for several years. The Pashto language community overcame several technical barriers to reach this enviable milestone, and we congratulate them for mobilising so effectively to record so much speech data! We’re working to enable Pashto on Spontaneous Speech. [You can join the Pashto language community on Discord here](https://discord.gg/hSgQquMqnj?ref=community.mozilladatacollective.com). [You can download the Pashto dataset here. ](https://datacollective.mozillafoundation.org/datasets/cmj8u3pnb00llnxxbfvxo3b14?ref=community.mozilladatacollective.com) ![Bar graph of total hours versus validated hours by language, showing Pashto as third by volume behind Catalan and English.](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/12/data-src-image-6bf08d99-583c-4978-9c2e-820134b39005.png) Pashto has now surpassed languages like Spanish, German and French by volume of speech data collected in Common Voice. ## Special mentions to Alsatian, Irish, Galician, Kabardian, Adyghe, Igbo, and Kurmanji Kurdish! We would also like to give a special shout out to language communities working on Alsatian, Irish, Galician, Kabardian, Adyghe, French, Igbo, and Kurmanji Kurdish. Targeted campaigns by dedicated volunteers in these communities have achieved laudable outcomes. Alsatian, new to Common Voice this release, has just started gathering data. Thanks to efforts by [*Údarás na Gaeltachta*](https://udaras.ie/en/?ref=community.mozilladatacollective.com)*,* Irish has significantly increased validation efforts in the last three months, with over 80% of all samples now validated thanks. Likewise, Galician, through [*Project Nós*](https://nos.gal/en/proxecto-nos/about-us?ref=community.mozilladatacollective.com) has undertaken significant validation efforts, validating over 50 hours of speech data in the last three months – a commendable achievement! The Kabardian language community has made spectacular progress, recording over 180 hours of speech, and validating almost all of this data. The closely related Adyghe language has not only recorded another 16 hours of data, they have also validated nearly all of their contributions. Campaigns in France have added to the already impressive amount of French speech data - now standing at nearly 1200 hours. Igbo, new to Common Voice early this year, have increased their data by over 50% in just three months, with a special shout out to [Victoria Ofuasia](https://www.linkedin.com/in/soluchi-victoria-ofuasia-39a32b227/?ref=community.mozilladatacollective.com) for all her efforts here. The Kurmanji Kurdish community have also made progress, validating additional data. ## Welcome to Common Voice, Lower Sorbian (dsb), Alsatian (gsw) and Laz (lzz) In this release we welcome to Common Voice three new language communities – Lower Sorbian (`dsb`), Alsatian (`gsw`) and Laz (`lzz`), bringing the total number of languages in Common Voice to 289, from 286 at the last release. A warm welcome to Lower Sorbian, Alsatian and Laz. - **Lower Sorbian (dsb):** Lower Sorbian has around 7,000 speakers in Brandenburg, Germany, particularly around the city of Cottbus. It's a West Slavic language that is heavily endangered, with most native speakers belonging to older generations. [Get the Lower Sorbian dataset here](https://datacollective.mozillafoundation.org/datasets/cmj8u3p0r006dnxxbi7fg8tdh?ref=community.mozilladatacollective.com). - **Alsatian (gsw):** Alsatian is spoken in Alsace, a region in north-eastern France, and is an Alemannic German dialect closely related to Swiss German. Alsatian is spoken by around 500,000 speakers. [Get the Alsatian scripted speech dataset here](https://datacollective.mozillafoundation.org/datasets/cmj8u3p6200a1nxxbf9nbqu7s?ref=community.mozilladatacollective.com), or the [Spontaneous Speech dataset here](https://datacollective.mozillafoundation.org/datasets/cmj8u48az002lnxzp4b47bq1i?ref=community.mozilladatacollective.com). - **Laz (lzz):** Laz has approximately 20,000 native speakers in Turkey and around 1,000 in Georgia. It's spoken along the southeastern shore of the Black Sea, primarily in northeastern Turkey in districts near the Georgian border, and is a Kartvelian language closely related to Mingrelian. [Get the Laz dataset here.](https://datacollective.mozillafoundation.org/datasets/cmj8u3pei00fpnxxb2zv0zgwh?ref=community.mozilladatacollective.com) ## Spontaneous Speech continues to grow We’ve also increased the number of languages represented in Spontaneous Speech in this release from 58 to 62\. You can see all the Spontaneous Speech datasets at: [https://datacollective.mozillafoundation.org/datasets?q=spontaneous+speech](https://datacollective.mozillafoundation.org/datasets?q=spontaneous+speech&ref=community.mozilladatacollective.com) The English (`en`) dataset will not be released at this time due to quality issues, however we are planning to include it in the next release. We are still working on including demographic data such as age, gender and accent of the speaker in the Spontaneous Speech releases, and we expect to include it in the next release. One thing that we do include however are quality tags that are provided for each record in the Spontaneous Speech datasets. [You can read more about them here](https://community.mozilladatacollective.com/improving-the-spontaneous-speech-english-dataset/). ## Datasets for Dholuo There will be two datasets for Dholuo luo released. [The regular Common Voice dataset for Dholuo is released under the CC-0 public domain license and you can find it here.](https://datacollective.mozillafoundation.org/datasets/cmj8u3pe600fhnxxbqol7x8bg?ref=community.mozilladatacollective.com) The DhoNam: Dholuo Speech dataset is a speech corpus designed to supercharge Automatic Speech Recognition (ASR) and other speech technologies for Dholuo, one of Kenya’s major indigenous languages. This dataset is being released under the [Nwulite Obodo Open Data License (NOODL) ](https://licensingafricandatasets.com/?ref=community.mozilladatacollective.com)and has been collected under the supervision of [Dr. Lilian Wanzare](https://www.linkedin.com/in/liliwanzie/?originalSubdomain=ke&ref=community.mozilladatacollective.com) and with funding provided by [GIZ FAIR Forward](https://www.giz.de/en/expertise/digitalisation/artificial-intelligence?ref=community.mozilladatacollective.com). This dataset will be released on MDC before the end of the year. ## Connect with us As always, you can connect with us on [Discord](https://discord.gg/4TjgEdq25Y?ref=community.mozilladatacollective.com) or [Matrix](https://chat.mozilla.org/?ref=community.mozilladatacollective.com#/room/#common-voice:mozilla.org), or contribute speech data at [https://commonvoice.mozilla.org](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com), and find all current Common Voice datasets on the [Mozilla Data Collective](https://datacollective.mozillafoundation.org/datasets?q=common+voice&ref=community.mozilladatacollective.com) platform. ### FAQ: How can I contribute to Mozilla Data Collective? URL: https://community.mozilladatacollective.com/faq-how-can-i-contribute-to-mozilla-data-collective/ Last updated: 2026-05-14T15:11:35.000Z We are a collective of linguists, technologists, activists, researchers and creatives. Whether you’re interested in stewarding data, conducting research, developing new AI and ML technologies, or just want to be part of our community working to make AI all it promises to be - not all it threatens to be - we welcome your involvement! Mozilla Data Collective is a new platform, and we want your help in shaping it. Some ways that you can get involved : - Create an account on[ Mozilla Data Collective](http://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com) to access more than 300+ global language datasets - Learn how to [share your own datasets on Mozilla Data Collective](https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform/) as a data provider - Explore our [public API](https://datacollective.mozillafoundation.org/api-reference?ref=community.mozilladatacollective.com) and [Python package](https://pypi.org/project/datacollective/?ref=community.mozilladatacollective.com) \- or [contribute to its development](https://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com) on GitHub - Join our [Discord server](https://discord.gg/4TjgEdq25Y?ref=community.mozilladatacollective.com) to connect with Mozilla Data Collective and Mozilla Common Voice community members from around the world - [Connect with us on Bluesky](https://bsky.app/profile/mozdatacollective.bsky.social?ref=community.mozilladatacollective.com) and share your AI voice projects with us - Participate in a group effort to improve multilingual AI systems, like our [Shared Task](https://community.mozilladatacollective.com/kick-off-of-the-shared-task-mcv-spontaneous-speech/) or [TidyVoice’s Speaker Verification](https://tidyvoice2026.github.io/?ref=community.mozilladatacollective.com) ### FAQ: Do I need to be part of an organization to upload a dataset to Mozilla Data Collective? URL: https://community.mozilladatacollective.com/faq-do-i-need-to-be-part-of-an-organization-to-upload-a-dataset-to-mozilla-data-collective/ Last updated: 2026-05-13T19:27:44.000Z We recognize that there are many use cases where an individual might want to share datasets that they have created, either on their own or on behalf of a group. If you do not set an organization name that is associated with your account on Mozilla Data Collective, any published datasets that you share will be tagged with the 'Community' tag. An organization on Mozilla Data Collective is meant to be broadly defined. Even as an individual, you can choose to house your data within an 'organization' on MDC - your organization can represent a community group, a project, a company, a school, or even just you! We recommend creating an organization on MDC in the following cases: - you are uploading more than one dataset, especially if the datasets are closely related - the attribution for your dataset is on behalf of a collective group or effort - you are representing an existing organization on MDC - you eventually plan to charge for datasets that you are publishing on MDC Adding or creating an organization helps with discoverability and visibility of your datasets. While it isn't necessary to have an organization in your profile to publish a dataset, we do encourage it if the any of the above criteria apply! ### Improving the Spontaneous Speech English dataset: lifting the lid on speech data quality uplift techniques URL: https://community.mozilladatacollective.com/improving-the-spontaneous-speech-english-dataset/ Last updated: 2026-02-24T15:04:08.000Z Firstly, we’d like to thank you for your patience. After introducing Spontaneous Speech early in 2025, we [released most locale datasets](https://datacollective.mozillafoundation.org/datasets?q=spontaneous&ref=community.mozilladatacollective.com) when the Mozilla Data Collective platform launched in alpha in September of this year. However, upon inspection, the English Spontaneous Speech dataset required some remedial work prior to release. In this blog post, we detail the work undertaken by [Kostis Saitas Zarkias](https://www.linkedin.com/in/kostissz/?ref=community.mozilladatacollective.com), one of our ML Engineers, to improve the data quality of Spontaneous Speech English. ## Elicited speech versus spontaneous speech data: what’s the difference? In November 2024, the [Common Voice platform introduced Spontaneous Speech](https://www.mozillafoundation.org/en/blog/common-voice-navbar-changes-and-spontaneous-speech/?ref=community.mozilladatacollective.com): new functionality which records data contributors answering a range of questions and then allows those recordings to be transcribed. In contrast to *elicited speech*, where a speaker reads a sentence which is then recorded and uploaded to the platform, spontaneous speech records a more natural, conversational type of speech. There are both benefits and drawbacks of this approach. On the plus side, spontaneous speech is of higher value to speech researchers and language technology developers because it reflects how people actually talk in real-world contexts, with disfluencies, false starts, repairs, fillers ("um," "uh"), and overlapping turns. This messy reality is what speech technologies need to handle in practice, whether it's voice assistants, transcription systems, or language models processing conversational input. Additionally, when people speak spontaneously, they display the full range natural variation in prosody (the rhythm and cadence someone speaks with), emotional expression and conversational dynamics (such as the way that Australian-accented speakers finish their sentences like a question, called a “high-rising terminal”). Elicited speech, in contrast, tends to be more controlled by the speaker, meaning that elicited speech data is missing these natural variations. Technologies trained or tested only on elicited speech often perform poorly when confronted with spontaneous speech in real-world situations. Spontaneous speech provides the challenging test cases—incomplete sentences, self-corrections, non-standard constructions, background noise, and simultaneous speakers—that reveal whether systems are truly robust. Having spontaneous speech data available for training can therefore help applications better match their deployment context. However, this variability is one of the downsides of working with spontaneous speech. It presents several challenges for speech data quality, and in turn, for researchers and language technologies applying that data to machine learning models. For example, how do researchers or language technology developers know when there is a disfluency in a recording? This has to be represented in a consistent way in speech data to be of use in training machine learning models. ## How have we improved the quality of Spontaneous Speech English? Several additional processing steps have been performed on Spontaneous Speech English, resulting in a dataset of much higher quality. Below, we provide a methodical technical description of the enhancements performed to support reproducibility and transparency. ### Audio language identification Because the user interface for Spontaneous Speech was slightly different to that used for read speech on Common Voice, there was initially some confusion about language selection, and some speakers contributed to the “English” dataset when they intended to contribute to another language. To remove non-English samples from Spontaneous Speech English, all audio files were processed with the Whisper large-v3-turbo model to provide automatic language identification. A threshold of 0.3 confidence was used to identify audio files where the predicted language did not match the expected language (English). Audio files that had a language identification confidence below the threshold were completely removed from the dataset. Audio files that had a language identification confidence above the threshold but did not match the expected language were moved from the **ss-corpus-en.tsv** to the **ss-report-audios-eng.tsv** and were tagged as "foreign\_language" with a comment indicating the predicted language and the confidence score. ### Disfluency standardization In conversational speech, a disfluency is an interruption or an irregularity in flow of speech. They’re a completely normal and natural part of how people talk in real-life situations. Common types of disfluencies include filler words like “um” or “ah” that indicate someone is thinking about what to say next; repeating sounds or words like “I-I-think so”, false starts, where someone starts to say one thing then says another or prolongations like “Noooo!”. In the [guidelines for transcribing Spontaneous Speech](https://commonvoice.mozilla.org/en/guidelines?tab=spontaneous-speech&ref=community.mozilladatacollective.com#transcribe-the-audio-subheader-1), contributors are asked to mark disfluencies with a disfluency marker, e.g: > *Like \[disfluency\] I dunno, fixing the house, tending to my plants, \[disfluency\] go on a hike, or go somewhere with nice scenery without too many humans, or no humans at all, preferably.* However, not all disfluencies have been tagged this way, because there is often a lot of ambiguity around what is a disfluency and what could be a named entity (a people, a place or a product) or an acronym. For this specific release, if the contributor did **not** mark the disfluency with any disfluency marker (brackets), the disfluency has not been standardized. The wrapping characters for a disfluency marker (\[ \] / ( ) / { }) have all been standardized to square brackets (\[ \]). You can see a worked example of how disfluencies are tagged in the example below: > *Oh, a lot of things! I have a lot of plans in my head. Like \[disfluency\] I dunno, fixing the house, tending to my plants, \[disfluency\] go on a hike, or go somewhere with nice scenery without too many humans, or no humans at all, preferably. And take a great, I dunno, landscape picture with my camera or something like that. But, usually that's the plan, right? Usually I will just stay at home and do nothing.* Unpacking these a little, we see: > “Like, uhmmm, I dunno” => Like \[disfluency\] I dunno > “Uhhh, go on a hike” => \[disfluency\] go on a hike For practitioners who use speech data that is transcribed from spontaneous speech, it’s very helpful to have disfluencies represented in a standard, repeatable way. The transcription guidelines for Spontaneous Speech within Common Voice instruct transcribers to place disfluencies in brackets, e.g. “\[disfluency\] yes, I think so”. Processing has been applied to standardize the representation of disfluencies in transcriptions so they are all bracketed consistently. ### Splitting into train/dev/test splits Splitting a dataset means dividing the dataset into separate subsets—**train, dev, and test**—to build, tune, and evaluate a machine learning model. For Spontaneous Speech datasets, we split based on the Speaker. We “fill” the dev and test split quotas first, to ensure as much variety as possible of Speakers in the dev and test splits, then put the remainder into the train set. We also want to ensure that the test & dev splits have a minimum amount of 1.5 hours of total audio duration. In total, the v1.0 English Spontaneous Speech dataset, after the language identification filtering, had **1.368** audio clips recorded, **1.165** of which were transcribed and **736** of them were validated by other users as correct (received a positive vote from a contributor as a valid audio-transcription pair). We created the splits on the 1.165 audio-transcription pairs which led to the follow distribution: - Train contains 345 audio-transcription pairs from 51 different speakers in a total of 1.66 hours of audio - Dev contains 539 audio-transcription pairs from 38 different speakers in a total of 1.90 hours of audio - Train contains 281 audio-transcription pairs from 42 different speakers in a total of 1.64 hours of audio We try to ensure that no speaker is in more than one of the splits. For example, Alice’s voice can only be heard in the train split, and cannot be found in dev or test. Since we avoid collecting personally identifiable information (PII) from our contributors, we implement this feature through the use of the attribute “client\_id”: a unique identifier automatically assigned to a user session in the Common Voice platform. This means that we cannot guarantee 100% accuracy that a single client\_id is actually a single unique person, as the same person could login from a different account and contribute (so one speaker actually having two different client\_id) or a group of contributors using a single account to record their voices (so multiple speakers under the same client\_id). ### Generating quality tag annotations Processing was also done to identify audio samples without transcriptions, audios that were very short or very long, or where the transcription was in an unexpected orthography (for example where English was transcribed using Japanese characters). We’re happy to report there weren’t any of those in English Spontaneous Speech. As an example, one audio and related transcription in the English Spontaneous Speech dataset was tagged with **short\_transcription**. Upon inspection we can see that the data contributor provided a one-word answer to the prompt: > What do you want to do with your next day off? > “Relax” Depending on the application the data is being used for, this short response may or may not be suitable inclusion. However, having the quality tag allows the researcher or developer to more easily identify short transcriptions, and make this decision. Each audio now has an accompanying quality tag to help you select the best data for your project. ## How is Spontaneous Speech data being used? While it’s still early days for Spontaneous Speech, data from Common Voice Spontaneous Speech is already being used to help improve automatic speech recognition through the [Mozilla Common Voice Spontaneous Speech ASR Shared Task](https://community.mozilladatacollective.com/shared-task-mozilla-common-voice-spontaneous-speech-asr/). In this challenge, researchers and developers are asked to improve the accuracy of automatic speech recognition (ASR) models across 21 under-represented languages from Africa, Asia, Europe and the Americas. [The data for the Shared Task is now available on the Mozilla Data Collective Platform](https://datacollective.mozillafoundation.org/datasets/cmfzu8u8wa555eq8onrk334h4?ref=community.mozilladatacollective.com). Please note that only the splitting and quality tags were applied to the Shared Task data; disfluency tags were not normalised. ## Next steps We’ll continue to make improvements to data quality in the Mozilla Data Collective platform, and we warmly welcome your feedback on any aspect of data quality via email to [mozilladatacollective@mozillafoundation.org](mailto:mozilladatacollective@mozillafoundation.org). You can now download Spontaneous Speech English from the Mozilla Data Collective platform [here](https://datacollective.mozillafoundation.org/datasets?q=Common+Voice+Spontaneous+Speech&locale=en&ref=community.mozilladatacollective.com). ### Uploading your dataset to the Mozilla Data Collective Platform URL: https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform/ Last updated: 2025-12-09T14:32:13.000Z > *Interested in joining the movement and publishing your dataset at Mozilla Data Collective? Here is all you need to know!* ### Making an account First things first, [sign up](https://datacollective.mozillafoundation.org/auth/signup?ref=community.mozilladatacollective.com) for an account on Mozilla Data Collective. We ask that the legal owners of the data are the ones that make an account and upload the dataset through their account in order to ensure proper governance. Make sure that you: - Use your full legal name (e.g. Alice Smith) - Use the legal entity / organization name of the owner of the dataset, if applicable (e.g. Alice Smith Foundation) - Use an email address associated with the organization, if available (e.g. info@alicesmithfoundation.org) - Avoid all capital or lowercase letters, unless part of your branding (e.g. IBM) ### Send a "Request to Upload" In order to ensure that the datasets hosted in our platform are aligned with the Mozilla Data Collective values & manifesto we manually review every request to upload. Simply go to [https://datacollective.mozillafoundation.org/profile/uploads](https://datacollective.mozillafoundation.org/profile/uploads?ref=community.mozilladatacollective.com) and click "Request to Upload" ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/11/image.png) **Note**: Since we are manually reviewing every request, please allow for a few days to get back to you with a reply. At this stage, you should expect an email from us asking further clarifying information about yourself, the organization you are representing, the data you want to upload and what are your goals with sharing this data in our platform. Once we have all the necessary information and have ensured your datasets meet our community standards we will approve your request. ### Guidelines on preparing your dataset for Submission Before starting the processing of submitting a new dataset, take a look at the information that we require you have in hand in order to fill in the Datasheet for your dataset. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/11/image-5.png) - ***Dataset Name***: Use a descriptive, unique name that is not too long or repetitive (e.g. if it is a collection of texts from a publisher, you could mention the publisher name in the title). - ***Description:*** Add a short description of one or two sentences. This will appear in your dataset's "Data Card" on the dataset listings page. - ***Task / Classification:*** If this dataset was curated for a certain downstream application / task, select it from the dropdown. Otherwise, select N/A. - ***Locale:*** An identifier of the language(s) the dataset is comprised of. - Use ISO-639-1 (two letter) or ISO-639-3 (three letter) language codes, if there is an ISO-639-1 code prefer that code; if there are two codes one in English and one from the native language, choose the native language one - e.g.eu **not** eus - If the dataset is in many languages, use the **mul** language code - If the dataset is in two languages and is a parallel corpus use two language codes separated by comma - e.g fr, ewo - If the dataset is only a specific variant, you can use the BCP-47 code for that variant or just use the higher level language code, - e.g. rm-vallader **or** rm - If the dataset contains multiple variants of the same language, use the higher level/macrolanguage code, - e.g. hy **not** hye **or** hyw - ***Format:*** The main file format of your data. - Use uppercase formats without initial full stop, - WAV **not** .wav - Use a comma and a space to separate formats, - e.g. WAV, TSV **not** WAV; TSV - ***License:*** The license attached to your dataset. - If you need to use another licence, choose “custom licence” and fill out: - Long form: e.g. Creative Commons Attribution International 4.0 - Short form: This should be an abbreviation, e.g. CC-BY-4.0 - URL: A url to the full licence text - ***Restrictions/Notes to Coordinators:*** If there are any certain restrictions you want to apply to your dataset you can write them here. For example: - "For research and scientific use only" - "You agree that you will not re-host or re-share this dataset" - ***Forbidden usage:*** If you want to prevent the dataset to be used in certain ways you can define it here. For example: - "You agree not to attempt to determine the identity of speakers in this dataset" - "Any attempt to clone the voice or train models that imitate the speakers in this dataset is forbidden" - "It is forbidden to use this dataset to train chatbots or large language models" - ***Additional Conditions:*** If there are any conditions that do not fit in the two boxes above, you can define them here. For example: - You may use this for evaluation of large language models, but not as part of training or fine-tuning - ***Point of contact:*** This should be the dataset owner / uploader, but it could also be a technical consultant. If the point of contact is not the owner, then you should also fill out the field below "Created by" and "Legal Contact". - ***Funded by:*** If the dataset was funded by a person or organisation who is not the owner/uploader, then that can be listed here. The contact can be an email address or it can be a link. - ***Legal Contact:*** If there is a specific legal contact for the dataset, for example for take-down requests, then that can be listed here. - ***Created by:*** If the point of contact did not create the dataset, put the dataset creator here - ***Intended Usage:*** What is the dataset intended to be used for? This is a free form field, and should consist of a sentence or paragraph describing the intended use. For example: - This dataset is intended for use in creating automatic speech recognition systems. - ***Ethical Review Process***: If the dataset was created in a university or academic/research institution with a review board, you can include that information here. Otherwise, please include information about how informed consent was obtained, for example: - Each participant gave consent to be included in this dataset via a written form and has the ability to revoke their inclusion in the dataset by emailing the authors at any time ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/11/data-src-image-5b1ab68f-cd7a-4903-b4b0-5cb1ddb9a689.png) - ***Technical Datasheet:*** When you are writing or advising on writing the technical datasheet you should put yourself in the shoes of a downloader, who could be a machine learning engineer, or some other person who is interested in using the data in training their machine learning systems. What kind of questions do you have? Think about the following questions: What? Who? Where? When? Why? - Language: What is this language, is it a particular variant, where is it spoken? - Source(s): Where is the data from? What kind of data is it? Who wrote it and when? - Domain(s): What domain is the data from? E.g. Is it general domain, health domain? - Size: if possible, include an approximate size in some unit of measurement that is relevant for the intended task (e.g. hours for ASR, tokens for text) - Structure: What is the structure of the data ? - Sample: What does the data look like? Add a list of examples (randomly sampled or hand-selected, approx 5-10) - If it has text: describe the orthography/writing system, include an alphabet table. - **Note: The technical datasheet is in Markdown, you can click “Preview” to view the result** Certain fields defined above will be used to generate your "Data Card" that will be shown at [https://datacollective.mozillafoundation.org/datasets](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com). For example: ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/11/data-src-image-af3b02f9-f1d3-45f9-adb2-14eb0d3b9ee3.png) The rest of the fields will populate your "Dataset Listing". You can visit any public dataset to get an idea, for example: [https://datacollective.mozillafoundation.org/datasets/cmflnuzz3l9oqw5m1ezpzl062](https://datacollective.mozillafoundation.org/datasets/cmflnuzz3l9oqw5m1ezpzl062?ref=community.mozilladatacollective.com) ### Dataset Exclusivity Clause During the submission process of your dataset you will notice the following checkbox. ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/11/image-4.png) If the same version of your dataset is available on other data-sharing platforms, the **please check the box** before submitting. Note that this clause does **not mean that the *data* in the dataset is not available in some form online.** For more information about this condition please visit [this page](https://community.mozilladatacollective.com/faq-what-does-it-mean-to-exclusively-host-my-dataset-on-moz/). ### Creating a `.tar.gz` file of your dataset All datasets on the Mozilla Data Collective platform need to be uploaded in a single `.tar.gz` file. Here, we show you how to create a `.tar.gz` file on your preferred operating system. #### Linux and Mac Using the command line, use the following command to create a single `.tar.gz` file of your dataset: ```bash tar -czvf dataset-name.tar.gz /path/to/files/or/directories ``` #### Windows If you're using Windows, we recommend using the free, open source software called [7-Zip](https://www.7-zip.org/?ref=community.mozilladatacollective.com) to create a `.tar.gz` file. After downloading and installing 7-Zip, follow these steps: 1. **Right-click** on the folder or files you want to archive. 2. In the context menu, select **`7-Zip`** \> **`Add to archive...`** 3. In the Add to Archive window, under **`Archive format`**, select `**tar**`. Name your file (e.g., `dataset-name`) and click `**OK**`. This creates a `.tar` file. 4. **Right-click** on the newly created `.tar` file. 5. Select **`7-Zip`** \> **`Add to archive...`** again. 6. This time, under **`Archive format`**, select `**gzip**` (which should now be available since there is only one file). Name your file with the final `.tar.gz` extension (e.g., `dataset-name.tar.gz`), and click **`OK`**. ### **Create your first Dataset Submission** After your request has been approved, visit the same "Uploads" page and you will see an option to create a "New Submission". ![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/11/image-1.png) Start the process and fill in the required information following the guidelines defined above. Note that you can start the process, then click "Save Draft" and come back the submission at a later stage to complete it. **When the submission is under the status of "Draft" you are free to make edits to it as you like.** Once you have uploaded the data and filled out all the necessary fields you can click "Submit for Review". **When the submission is under the status of "Submitted for Review" you are no longer able to make edits to the submission.** One of our team members will pick up the submission and review it. It is possible that we will reach back to you for edits. **This will change the status of the submission to "Edits Requested" and you can edit the submission again.** Finally, once all the edits have been finalized, you can click "Submit for Review" again, and, unless there are any other issues that need to be addressed, we will approve your submission. ### View your dataset Once your dataset submission is approved you will be able to see it and share it publicly at [https://mozilladatacollective.org/datasets](https://mozilladatacollective.org/datasets?ref=community.mozilladatacollective.com). ### FAQ: What does it mean to exclusively host my dataset on Mozilla Data Collective? URL: https://community.mozilladatacollective.com/faq-what-does-it-mean-to-exclusively-host-my-dataset-on-moz/ Last updated: 2026-05-14T15:12:00.000Z When you upload a dataset to Mozilla Data Collective, you have the option to make your dataset exclusive to MDC. The default [t](https://datacollective.mozillafoundation.org/terms/providers?ref=community.mozilladatacollective.com#2-data-supplier-rights-and-obligations)[erms of use for data providers](https://datacollective.mozillafoundation.org/terms/providers?ref=community.mozilladatacollective.com#2-data-supplier-rights-and-obligations) on the platform is that datasets are exclusive to MDC. Choosing to host your dataset exclusively with MDC means that you do not plan on hosting the same version of the dataset on other data-sharing platforms, or making it available through other access endpoints. Exclusively hosting your *dataset* on Mozilla Data Collective does not mean that the *data* in the dataset is not available in some form online, but instead refers to the collection of data as a whole. Many exclusively hosted datasets might have components of their data that have been sourced from various places online, but have had additional processing and structured added to them when they were included in the dataset. Mozilla Data Collective provides protections, management controls and visibility for datasets hosted on the Platform. These safeguards and insights apply in full when your dataset is hosted exclusively on the Platform. If your dataset will also be hosted or made accessible in other places, certain of these protections and visibility features may not apply. During the upload process, you should consider whether you intend to make your dataset available through MDC, or if you will also be hosting it on another site. If you are planning to make your dataset available on different platforms in addition to Mozilla Data Collective, you should check this box when you submit your dataset and review [Appendix 1 for data providers in the terms of use](https://datacollective.mozillafoundation.org/terms/providers?ref=community.mozilladatacollective.com#appendix-1). If you are planning to make your dataset available through Mozilla Dataset exclusively, leave this box unchecked. ### FAQ: What kind of datasets can I publish on Mozilla Data Collective? URL: https://community.mozilladatacollective.com/faq-what-kind-of-datasets-can-i-publish-on-mozilla-data-collective/ Last updated: 2026-05-13T19:25:51.000Z We prioritise helping communities unlock content that is not on the web already, and prefer audio, image, and video formats, though we will also accept text documents that advance the above goals. Our expectation is that each dataset is (or can be) prepared in a way that enables its use in machine learning contexts, or is intended to be consumed in such a way for research, evaluation, training, or other similar endeavors. The specific details of how each dataset can be used is up to you, and set via terms on your dataset's corresponding datasheet. Datasets should be organized in a way that makes sense for their contents and intended use. When you upload the dataset to Mozilla Data Collective, you will need to put the contents of your dataset into a .tar.gz format, and upload it as a single file. Datasets must adhere to the Mozilla Data Collective [terms of use](https://datacollective.mozillafoundation.org/terms/providers?ref=community.mozilladatacollective.com). By uploading a dataset to Mozilla Data Collective, you are responsible for ensuring that you have the rights to distribute the dataset and that it does not contain any data in the [Prohibited Data Content](https://datacollective.mozillafoundation.org/terms/providers?ref=community.mozilladatacollective.com#3-data-provider-restrictions) section of the terms. ### FAQ: What are the main points of the MDC Terms of Use? URL: https://community.mozilladatacollective.com/faq-what-are-the-main-points-of-the-mdc-terms-of-use/ Last updated: 2026-05-14T15:12:35.000Z ### About downloading and using datasets **When downloading a dataset, am I getting permission to use it from Mozilla Data Collective or the Data Provider?** When you download a dataset from Mozilla Data Collective, you are entering into an agreement with the Data Provider who published the dataset. Data Providers set their own terms, licenses, and access conditions for their datasets, which Mozilla Data Collective then hosts and distributes on their behalf. **How do I know what terms and conditions apply to my use of a given dataset?** Each dataset hosted on Mozilla Data Collective specifies a specific license that applies to that particular dataset. The collective set of terms and conditions that apply to a dataset include: the license specified for that dataset, the additional restrictions or prohibited uses included on the datasheet, and [the Terms of Use governing your access to and use of the Mozilla Data Collective website itself](https://datacollective.mozillafoundation.org/terms?ref=community.mozilladatacollective.com) that you agreed to when creating an account. The terms and conditions applicable to all downloads from Mozilla Data Collective specifically prohibit rehosting, re-sharing, or otherwise re-distributing the datasets beyond the permissions given by the dataset license and its additional restrictions. **If the license on a dataset permits a certain use, but the additional terms on the datasheet prohibit that use, which do I have to adhere to?** The additional terms on a datasheet and the license on a dataset are considered additive. In places where there is a conflict (e.g. the license permits one thing, but the terms do not allow that use) the more restrictive of the two should be used. **Does Mozilla Data Collective guarantee the accuracy, completeness, or legality of datasets hosted on the platform?** No. Mozilla Data Collective does not guarantee accuracy, completeness, or legality of datasets hosted on the platform. We may, but are not obligated to, review datasets when they are uploaded to the platform and use our discretion to vet data providers, but Mozilla Data Collective does not endorse or guarantee a particular dataset’s quality or suitability for use. You, as the downloader, remain responsible for evaluating whether a given dataset meets your needs. **Can I use automated tools, scripts, and bots to retrieve, download, or index datasets?** You may use the [Mozilla Data Collective public API](https://datacollective.mozillafoundation.org/api-reference?ref=community.mozilladatacollective.com) to access datasets hosted on MDC, provided you have created access credentials through the platform. You may not use any other tools, scripts, or automated scraping technologies to access, copy, scrape, or otherwise access a dataset from Mozilla Data Collective. **Am I allowed to mirror or host a dataset that I downloaded to make it available to my community?** No. Mirroring, re-hosting, re-distributing, or otherwise re-sharing datasets that you download from Mozilla Data Collective is prohibited beyond the permissions given by the dataset license and its additional restrictions. **What obligations do I have about managing and continuing to use data that I have downloaded from Mozilla Data Collective if I terminate my account?** If you terminate your Mozilla Data Collective account, or if Mozilla terminates your account due to violations of the terms of use, you must stop using the datasets that you downloaded. Pursuant to the terms of use applicable to the the individual datasets, you must also delete the copies that you have in your possession. ### About uploading and publishing datasets **Is Mozilla Data Collective free to use?** Our platform offers free hosting of datasets to dataset providers who choose to make them available for free or without economic costs to others. In the future, we plan to charge a commission in cases where the datasets are available for a fee. **If I publish a dataset in Mozilla Data Collective, am I giving it away to Mozilla?** No, by choosing to upload and host a dataset in our platform you’re simply giving us permission to make it available to others under the license and restrictions that you establish. You retain ownership and control over the dataset and may choose to withdraw your dataset at any time. **I already publish my datasets somewhere else, can I publish them in Mozilla Data Collective too?** Dataset creators can decide if they want to host their datasets with us exclusively or not. If you elect to opt out of the exclusivity, you need to indicate that election when completing your datasheet. By doing so, you understand that we won’t be able to offer you the same level of proactive protection we provide to our exclusive datasets. Join Mozilla Data Collective → ### What do the languages Nahuatl, Bahasa Indonesia and Bulgarian have in common? URL: https://community.mozilladatacollective.com/what-do-the-languages-nahuatl-bahasa-indonesia-and-bulgarian-have-in-common/ Last updated: 2025-11-11T09:17:18.000Z At first glance, Western Sierra Puebla Nahuatl - an endangered variety of the indigenous Mexican language [Nahuatl](https://en.wikipedia.org/wiki/Nahuatl?ref=community.mozilladatacollective.com), spoken in the state of Puebla in Mexico, [Bahasa Indonesia](https://en.wikipedia.org/wiki/Indonesian%5Flanguage?ref=community.mozilladatacollective.com) \- the official national language of the Republic of Indonesia, spoken by over 250,000,000 people, and [Bulgarian](https://en.wikipedia.org/wiki/Bulgarian%5Flanguage?ref=community.mozilladatacollective.com) \- a Slavic language spoken by nearly 8 million people in south-eastern Europe - might seem to have little in common. Not so fast! They're all languages featured in the very first community curated datasets to be uploaded to the Mozilla Data Collective platform. With huge thanks to the organizations and individuals who created them, they're now available for you to explore. ## [Dimitar - a 1.4 hour corpus of Bulgarian from a single speaker](https://datacollective.mozillafoundation.org/datasets/cmhpahaib00d8mk07ely2m8wh?ref=community.mozilladatacollective.com) Different tasks in machine learning require different sorts of data. For example, in [automatic speech recognition](https://en.wikipedia.org/wiki/Speech%5Frecognition?ref=community.mozilladatacollective.com) (ASR), the task is to accurately recognize speech from a wide variety of speakers - from varying genders, ages and accents. So, data from speakers of varying genders, ages and accents is needed to create a robust ASR model. ASR takes spoken audio and predicts written words. The complement to ASR is [speech synthesis](https://en.wikipedia.org/wiki/Speech%5Fsynthesis?ref=community.mozilladatacollective.com) \- also called "text to speech" or TTS. TTS models take written words and generate spoken audio. TTS models, in contrast to ASR models, need high quality data from a single speaker. The Dimitar corpus provides just that data in Bulgarian. Curated by the [Open Home Foundation](https://www.openhomefoundation.org/?ref=community.mozilladatacollective.com) \- the not-for-profit foundation behind the [Home Assistant](https://www.home-assistant.io/?ref=community.mozilladatacollective.com) home automation platform, the Dimitar corpus is suitable for training TTS models that create synthesized speech in Bulgarian. For example, it could be used with the [Piper TTS system](https://github.com/OHF-Voice/piper1-gpl/?ref=community.mozilladatacollective.com), which is the TTS system the Home Assistant team use to generate synthetic voices for [Home Assistant Voice](https://www.home-assistant.io/voice-pe/?ref=community.mozilladatacollective.com). ### Datasheets help data practitioners understand how to use a dataset Each dataset uploaded to the Mozilla Data Collective must be accompanied by a *datasheet*. A datasheet is a form of dataset documentation: transparent descriptions of a dataset's contents that helps a data practitioner understand the type of data in the dataset, what it's useful for - and - equally - what it *shouldn*'t be used for. Looking at the datasheet for the Dimitar corpus, we can see that it provides specifics useful for building TTS, such as the median characters per sentence and median words per sentence, which are useful metrics for speech technologists building synthetic speech models. It also helpfully provides the full Bulgarian alphabet. This information helps speech technologists match phonemes - the building blocks of speech - to individual characters. Particular thanks go to [Dr. Michael Hansen](https://www.linkedin.com/in/michael-hansen-9885b2105/?ref=community.mozilladatacollective.com), Voice Engineering Lead at Nabu Casa, for all his work on this corpus. [*Download Dimitar*](https://datacollective.mozillafoundation.org/datasets/cmhpahaib00d8mk07ely2m8wh?ref=community.mozilladatacollective.com) ## [Podcast Hari Minggoean - a ten-hour corpus of Javanese-accented Bahasa Indonesia featuring code-switching and contemporary Indonesian speech](https://datacollective.mozillafoundation.org/datasets/cmhlvswz200fio007c7szypw8?ref=community.mozilladatacollective.com) Derived from the "Hari Minggoean" podcast, this dataset has several features that are attractive to practitioners building ASR models for Indonesian and other Malay languages. Firstly, the podcast hosts content tailored for young Indonesian audiences, and uses contemporary language. Data of contemporary language use is particularly important: languages are changing all the time, and ML models need to keep up. Just five years ago, there was no need for ASR to recognize phrases like "Skibidi Ohio rizz"! But if ASR is to recognize the speech of young people accurately, we need data *from* young people. Secondly, the dataset features [code-switching](https://en.wikipedia.org/wiki/Code-switching?ref=community.mozilladatacollective.com). In linguistics, code-switching is where a speaker uses two languages or two varieties in the same utterance. For young people in Indonesia, it's very common to alternate between Bahasa Indonesia and English. However, most speech recognition models are trained on only a single language, and [speech recognition of code-switched speech is still an emerging research area](https://asru2019.signalprocessingsociety.org/asru2019.org/wp/indexdbb7.html?page%5Fid=1881&ref=community.mozilladatacollective.com) \- one this dataset might be very suited for! Terima kasih banyak kepada [Yacub Fahmilda](https://id.linkedin.com/in/yacub-fahmilda-520489163?ref=community.mozilladatacollective.com) atas hasil kerja yang bagus! [*Download podcast Hari Minggoean*](https://datacollective.mozillafoundation.org/datasets/cmhlvswz200fio007c7szypw8?ref=community.mozilladatacollective.com) [*Listen to the Hari Minggoean podcast /* ](https://open.spotify.com/show/3dkeZ25Zr5KHorvzzMu8kh?si=ab58ff80a0e5464e&ref=community.mozilladatacollective.com) [*Mendengarkan podcast Hari Minggoean*](https://open.spotify.com/show/3dkeZ25Zr5KHorvzzMu8kh?si=ab58ff80a0e5464e&ref=community.mozilladatacollective.com) ## [Tetelancingo Nahuatl Corpus: A corpus of audio and annotated transcriptions of Western Sierra Puebla Nahuatl](https://datacollective.mozillafoundation.org/datasets/cmhkl8z2a007rnr07p9bm5kmz?ref=community.mozilladatacollective.com) From Indonesia we go across the Pacific to central Mexico, which is where PhD Candidate and Mozilla Data Collective linguist, [Robert Pugh](https://www.linkedin.com/in/robert-pugh-ab50b45a/?ref=community.mozilladatacollective.com), and a team of collaborators collected monologues and dialogues from 5 individual speakers from a community in Zacatlán de las Manzanas. What's important for data practitioners in this corpus is that each recording has both spontaneous and standard transcriptions, a Spanish translation of the recording and word-level language tags. These multiple forms of annotation help create a richer dataset which lends itself for use in many ML tasks. For example, because the recordings contain both a Nahuatl and a Spanish transcription, the corpus could be used to build a [machine translation](https://en.wikipedia.org/wiki/Machine%5Ftranslation?ref=community.mozilladatacollective.com) model between these languages. That is, it's a "parallel corpus" - a corpus of two languages that have the same semantic meaning. Machine translation is notoriously difficult in endangered languages, so this dataset is one more step in the direction of speech technologies that work better for more of the world's 7000 spoken languages. Additionally, the word-level tags make this an excellent dataset for identifying when speakers switch between languages or variants within the same utterance - again, a useful application for endangered languages. [*Download Tetelancingo Nahuatl Corpus*](https://datacollective.mozillafoundation.org/datasets/cmhkl8z2a007rnr07p9bm5kmz?ref=community.mozilladatacollective.com) > For more information about the dataset and some preliminary experiments, see the paper [Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla Nahuatl](https://aclanthology.org/2025.naacl-long.181.pdf?ref=community.mozilladatacollective.com). ## How do I get started curating and uploading my own dataset to the Mozilla Data Collective platform? Have you been inspired by these datasets to curate your own? [Mozilla Data Collective](https://datacollective.mozillafoundation.org/?ref=community.mozilladatacollective.com) provides you with the ability to host and share your datasets, on your own terms. To upload your own dataset, first [create an account](https://datacollective.mozillafoundation.org/auth/signup?ref=community.mozilladatacollective.com), then [request to be able to upload data](https://datacollective.mozillafoundation.org/profile/uploads?ref=community.mozilladatacollective.com). We verify each dataset creator as a quality assurance measure. What community datasets will we see next? Stay tuned! #### #### ### FAQ: Why can’t I download old versions of Common Voice datasets? URL: https://community.mozilladatacollective.com/faq-why-cant-i-download-old-versions-of-common-voice-datasets/ Last updated: 2026-05-14T15:13:09.000Z In order to better serve our community and to keep up with current changes in best practices for data stewardship, we are changing how previous releases of the Common Voice datasets are accessed. They will now be accessible to interested researchers and others via an email-based request process. You will send us an email using a valid return address and agree to the terms of use, and we will send you a link to download the dataset. This means that we will be able to more comprehensively honour deletion requests. We are committed to supporting researchers and teachers in carrying out reproducible research, but we are adding a step so that we can confirm purpose and track usage. **Read the full post for details and how to request an older version:** [We’re changing access to older versions of Common Voice datasets](https://community.mozilladatacollective.com/were-changing-access-to-older-versions-of-common-voice-datasets/) ### We're Changing Access to Older Versions of Common Voice datasets URL: https://community.mozilladatacollective.com/were-changing-access-to-older-versions-of-common-voice-datasets/ Last updated: 2026-06-23T07:48:01.000Z Common Voice has grown to more than 33,000 hours of speech across 137 languages, contributed by hundreds of thousands of volunteers worldwide. And each one of those people donating their voices to Common Voice is helping improve technology to reflect how real communities speak. With those contributions comes a responsibility to do two things well: respect contributors’ rights—including when they ask us to delete their data—and support researchers who need past versions to reproduce results for work. At times, those goals pull in different directions, creating tension for us in how we store the data we collect. ## The right to be forgotten Roughly every three months, we publish a new Common Voice dataset version (e.g., 23.0 from September 2025) of all the voice clips plus matching text for each language. Sometimes, contributors ask us to remove their data for privacy or other reasons. We remove it from the database and leave it out of future dataset versions. We feel strongly that this is the right thing to do. However, older releases may still contain earlier contributions, which researchers sometimes need to replicate past results (e.g., a paper that used the 17.0 dataset). They want to be able to make sure that the results are the same, to know that their implementation was correct. ## What’s changing To better balance privacy and reproducibility, we’re updating access to old datasets: 1. All dataset downloaders will need to use a *validated* email to access datasets. 2. You need to agree to basic terms when downloading datasets: please don’t share the files further. 3. Current dataset releases will remain self-serve. 4. For older dataset versions, you’ll need to get in touch to tell us which version you need and why (for example, to reproduce a paper). We will help you get the right version for your use case. ## What it means for the community We’ll keep honoring deletion requests in future releases and reduce the casual spread of older datasets. Researchers can still get historical snapshots, but this short new ‘request’ step adds accountability, aligns with contributor privacy, fits today’s risk landscape (where voice cloning needs much less data than before), and creates a single point of access to find the right version with the right terms. ## How to request an older dataset version For the moment, to request an old dataset, just [send us an email](mailto:commonvoice@mozilla.com?subject=Request%20for%20Historical%20Common%20Voice%20Access) providing: - The version number you need - A short reason (e.g., replicating a paper that used v17) - Confirm your consent to the no‑reshare and research‑only terms We’ll reply with the next steps. ## Get in touch We’re happy to take feedback and hear how these changes affect your workflow. If centralized access makes your work harder, get in contact with us to explain how, and the context of your setup (classroom, workshop, lab, solo research), the constraints, and the versions you rely on. This kind of input will shape how we design and develop the platform so we can keep responsibly stewarding the datasets whilst continuing to make improvements for ease of access. ### FAQ: I've noticed an issue with a data listing - what do I do? URL: https://community.mozilladatacollective.com/faq-ive-noticed-an-issue-with-a-data-listing-what-do-i-do/ Last updated: 2026-05-14T15:14:12.000Z If you discover a minor issue with the accuracy of a data listing on Mozilla Data Collective (a typo, missing attribute in datasheet, etc.) and want to bring it to the attention of the data provider, reach out to the point of contact as specified in the data listing if available. If you believe that your work has been used in a way that constitutes copyright infringement related to a dataset hosted on Mozilla Data Collective, please review [Section 11 of the Terms of Use](https://datacollective.mozillafoundation.org/terms?ref=community.mozilladatacollective.com) to contact Mozilla's Designated DMCA Agent. To report other issues with a data listing, or if you believe that a dataset violates our terms of use, [please contact us](mailto:mozilladatacollective@mozillafoundation.org?subject=Report%20an%20Issue%20with%20a%20Data%20Listing) or use the report dataset link on the datasheet page. ### Kick-off of the Shared Task: MCV Spontaneous Speech URL: https://community.mozilladatacollective.com/kick-off-of-the-shared-task-mcv-spontaneous-speech/ Last updated: 2025-10-13T15:50:09.000Z Join us for the official kick‑off of the **Shared Task: Mozilla Common Voice Spontaneous Speech**! ​This live, online gathering will bring together researchers, engineers, and language‑technology enthusiasts from around the world to launch the challenge focused on building robust, multilingual ASR systems for under‑represented languages. In a collaborative Zoom session, you’ll hear an overview of the challenge timeline, [data resources](http:// https://datacollective.mozillafoundation.org/), evaluation metrics, and prize structure, followed by an open Q&A where participants can discuss ideas, share insights, and explore cross‑lingual strategies. ​Whether you’re already prototyping a model or just curious about the problem space, this event is the perfect place to connect, ask questions, and start shaping the next generation of inclusive speech technology. ​**For more on the Shared Task Challenge, follow this link:** There will be two sessions in order to accommodate multiple time zones: - Thursday, October 16 - 6pm UTC+3 ([https://luma.com/cr9fehms](https://luma.com/cr9fehms?ref=community.mozilladatacollective.com)) for European afternoon/Americas morning - 6pm UTC-6 ([https://luma.com/0lwc2oyk](https://luma.com/0lwc2oyk?ref=community.mozilladatacollective.com)) for Americas afternoon/Oceania morning ### FAQ: Why can't I re-host or share Common Voice datasets that I download from MDC? URL: https://community.mozilladatacollective.com/faq-why-cant-i-re-host-or-share-common-voice-datasets-that-i-download-from-mdc/ Last updated: 2026-05-14T15:15:08.000Z Mozilla Data Collective wants to provide effective, ethical stewardship support for datasets. This is challenging or impossible when datasets are mirrored or split across a range of forks. Our principles include trying to enable good governance, e.g. the right to be forgotten, as much as possible, which means we need to be able to maintain robust versioning and communication channels. CC0 remains the license *for computational use*, whilst not allowing mirroring the datasets is a *platform term.* Common Voice communities are now starting to collect data under a range of difference licenses. Some of these licenses will include conditions that require stronger transparency measures on downloads and use. This is only possible with safeguards around redistribution. If you have an edge case that might need support, we're always happy to connect and discuss. You can reach us at [commonvoice@mozilla.com](mailto:commonvoice@mozilla.org) or [mozilladatacollective@mozillafoundation.org](mailto:mozilladatacollective@mozillafoundation.org). ### Shared Task: Mozilla Common Voice Spontaneous Speech ASR URL: https://community.mozilladatacollective.com/shared-task-mozilla-common-voice-spontaneous-speech-asr/ Last updated: 2025-10-24T23:17:10.000Z --- **Quick links** - [Registration Form](https://forms.gle/DUj5fVMD62aqsY7R9?ref=community.mozilladatacollective.com) - [Link to datasets](https://datacollective.mozillafoundation.org/datasets/cmfzu8u8wa555eq8onrk334h4?ref=community.mozilladatacollective.com) - [Codabench](https://www.codabench.org/competitions/10820/?ref=community.mozilladatacollective.com) page (to submit results during the testing period) - Contact: sharedtask@mozillafoundation.org --- **Overview** Automatic speech recognition (ASR) has come a long way – but most systems are still trained on polished, read-aloud speech. So we set out to build a model that can handle the messy, beautiful reality of spontaneous responses and languages long ignored by mainstream tech. We’re raising the standards for accuracy scores and building speech technology that works for everyone, not just the few. So along with Mozilla Data Collective’s new Spontaneous Speech datasets, we’re launching a shared task that challenges researchers and developers to push ASR further, across 21 underrepresented languages from Africa, Asia, Europe and the Americas. Join Mozilla Data Collective → The goal of this shared task is to promote the development of robust automatic speech recognition (ASR) systems for spontaneous speech in a number of lower-resource languages that have historically been underrepresented in speech technology research. Many of the large, widely-used ASR datasets are either read speech (previous releases from Mozilla Common Voice), predominantly English (Switchboard datasets), or both (LibriSpeech, WSJ). This shared task is based on the recently released spontaneous speech datasets from Mozilla Common Voice. In these datasets, participants freely respond to prompts, and the responses are transcribed and validated. The available spontaneous speech datasets represent a wide range of under-served language communities. The task will evaluate systems based on overall performance, the best improvement over the baseline for any single language, and resource-constrained system performance, encouraging innovative approaches to handle the nuances of spontaneous speech recognition. **Tasks** The shared task includes one main task: - *Multilingual ASR Performance* **(Task 1)** : The average Word Error Rate (WER) on all languages (excluding unseen languages). and 3 subtasks: - *Best improvement on a single language* **(Task 2)**: The largest improvement on WER for any single language as compared to our baseline system's performance. Submissions to this task can be for any of the languages in Task 1 or the unseen languages from Task 4 (see below). - *Model-size constrained improvement on single language* **(Task3)**: The best WER improvement over baseline for any language with a model that is less than 500 MB in size. Submissions to this task can be for any of the languages in Task 1 or the unseen languages from Task 4 (see below). - *Unseen language ASR* **(Task 4)**: In addition to the set of languages that have training data (see the Data section below for details), we also include 5 languages for which no training data will be provided. We will provide the language names, but it is up to the participating teams to find additional data or leverage cross-lingual techniques. The best average WER on the set of unseen languages will win this task. Any additional data used must be able to be shared openly. **Data** The dataset (available on Mozilla Data Collective [here](https://datacollective.mozillafoundation.org/datasets/cmfzu8u8wa555eq8onrk334h4?ref=community.mozilladatacollective.com)) includes approximately 9 hours each of 21 total languages from Africa, the Americas, Europe, and Asia. Each language's dataset is available on the Mozilla Data Collective website. The following table displays all 21 languages: | | No. | Language | ISO 639 | | ---------- | --- | ---------------------------- | ------- | | *Africa* | | | | | | 1 | Bukusu | bxk | | | 2 | Chiga | cgg | | | 3 | Nubi | kcn | | | 4 | Konzo | koo | | | 5 | Lendu | led | | | 6 | Kenyi | lke | | | 7 | Thur | lth | | | 8 | Ruuli | ruc | | | 9 | Amba | rwm | | | 10 | Rutoro | ttj | | | 11 | Kuku | ukv | | *Americas* | | | | | | 12 | Wixárika | hch | | | 13 | Southwestern Tlaxiaco Mixtec | meh | | | 14 | Michoacán Mazahua | mmc | | | 15 | Papantla Totonac | top | | | 16 | Toba Qom | tob | | *Europe* | | | | | | 17 | Gheg Albanian | aln | | | 18 | Cypriot Greek | el-CY | | | 19 | Scots | sco | | *Asia* | | | | | | 20 | Betawi | bew | | | 21 | Western Penan | pne | Additionally, for Task 4, we include 5 languages for which only test data will be released: *Adyghe (ady), Kabardian (kbd), Basaa (bas), Puno Quechua (qxp), Ushojo (ush)* For these languages, teams are encouraged to consult potential useful data and/or leverage cross-lingual approaches. Any data used must be openly licensed to facilitate reproducibility. **Prizes** - Task 1: $5,000 USD - Tasks 2-4: $2,000 USD each **Note**: *Contestants are not eligible to receive prizes if they are on the US Specifically Designated Nationals (SDN) list or if there are sanctions against the contestant’s country such that Mozilla is prohibited from paying them.* **Registration** Please register for the competition through the following [form](https://forms.gle/DUj5fVMD62aqsY7R9?ref=community.mozilladatacollective.com). **Important dates** - 26th September, 2025: [Train/Dev data](https://datacollective.mozillafoundation.org/datasets/cmfzu8u8wa555eq8onrk334h4?ref=community.mozilladatacollective.com) released (via *Mozilla Data Collective*) - 1st October, 2025: Shared task announced - 1st December, 2025: Test data released - 8th December, 2025: Deadline for submitting **final results** and **system description paper** - 12th December, 2025: Winners announced **Submission** Once we release the test data on the 1st December (audio only), teams will have 1 week to submit their system's predicted transcriptions for the relevant tasks on the shared task [CodaBench page](https://www.codabench.org/competitions/10820/?ref=community.mozilladatacollective.com). Submissions should take the form of a zip file containing three subdirectories, each with a set of one or more tsv files (1 per language being attempted): - multilingual-general - aln.tsv - bew.tsv - … - small-model - aln.tsv - bew.tsv - … - unseen-langs - ady.tsv - bas.tsv - … The tsv files should have two columns, the first is the name of the audio file, and the second is the predicted transcription. The scores for each task are the **average** over all of the languages in the respective task (21 for general and small model, 5 for unseen). **If you do not submit transcriptions for a given language, we will treat it as a blank transcription, resulting in a WER of 1.0.** For the “Biggest improvement over baseline” task, we will automatically select the language from the *multilingual-general* and *multilingual-small-model* that improves the most over our baseline. The team with the best performance in each task will be asked to submit their model and a script to perform inference, so that we can reproduce the results (to avoid the possibility of, e.g., post-editing predicted transcriptions to improve performance). If we are unable to reproduce the system performance, the team will be disqualified and we will request the model and inference script from the next-highest-scoring team. **System description papers** Each team must submit, in addition to their system's predicted transcriptions on the test data, a system description paper, between 4-8 pages (excluding acknowledgments and references). Please use the [ACL Template](https://github.com/acl-org/acl-style-files?ref=community.mozilladatacollective.com). Submissions should omit the author names for review. **Organisers** *Programme chairs*: - Francis M. Tyers, Indiana University - Robert Pugh, Mozilla Data Collective - Anastasia Kuznetsova, Rev.com - Jean Maillard, Meta *Programme committee*: - Antonios Anastasopoulos, George Mason University - Kathy Reid, Australian National University - Miguel del Rio, Rev.com - Pooneh Mousavi, MILA - Abteen Ebrahimi, University of Colorado, Boulder - Ximena Gutierrez Vasquez, UNAM *Advisory committee*: - Emmanuel Ngué Um, University of Yaounde 1 - Belu Ticona, George Mason University - Jennifer Smith, University of Glasgow - Joyce Nabende, Makerere University - Jonathan Mukiibi, Makerere University - Elwin Huaman, Innsbruck University - Yacub Fahmilda, Universitas Gadjah Mada - Riska Legistari Febri, Universitas Gadjah Mada - Murat Topçu, Okan University - Rosario de Fátima Alvarez García, Universidad Autónoma Metropolitana - Athziri Madeleine Vega Martínez, Universidad Nacional Autónoma de México - Marlon Vargas Méndez, Escuela Nacional de Antropología e Historia - Antonio Hayuaneme García Mijarez, Nación Wixárika - Vivian Stamou, Archimedes Athena Research Centre - Meesum Alam, Indiana University - Jonathan Lewis-Jong, University of Oxford ### Common Voice 23.0 Live On Mozilla Data Collective URL: https://community.mozilladatacollective.com/common-voice-23-0-live-on-mozilla-data-collective/ Last updated: 2025-09-30T14:10:27.000Z Common Voice 23.0 is now live – and available for [download via Mozilla Data Collective](https://datacollective.mozillafoundation.org/datasets?ref=community.mozilladatacollective.com). Mozilla Data Collective is a sister platform from the team behind Common Voice, designed to let dataset owners and creators offer their data on their own terms. Mozilla Data Collective was built in response to community feedback, as many of you asked for more flexibility in licensing and ways to host and manage datasets collected outside the Common Voice platform. **Common Voice 23.0 features:** This release adds 2,100 new hours of open speech data, bringing the total to 35,921 hours. Almost 2,000 newly validated hours have been added to Common Voice (24,600 validated in total to date). With 149 new languages added to the dataset, this more than doubles the total to 286 languages represented in Common Voice.. This unprecedented expansion of new contributions spans the globe, including Lassi, Scots, Tupuri and Puno Quechua. **Spontaneous Speech data is now available:** Spontaneous Speech mode launched this year on Common Voice, allowing contributors to record natural responses to prompts in their own words for the first time. Common Voice 23.0 features 357 hours of Spontaneous Speech data across 51 languages, many of which have been previously excluded from open speech datasets (e.g., Betawi, Western Penan and Ligurian). **Introducing datasheets:** Each dataset in Common Voice 23.0 is now accompanied by [detailed datasheets](https://github.com/common-voice/cv-datasheets?ref=community.mozilladatacollective.com) to help users to understand more about dataset terms, the language the dataset includes and more key features of the data.. **Mozilla Data Collective offers more download support:** Common Voice 23.0 – and all future releases – can now be downloaded through Mozilla Data Collective. We’ve also added a [download via API](https://datacollective.mozillafoundation.org/api-reference?ref=community.mozilladatacollective.com) option, as well as an API-supported [open-source Python library](https://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com), allowing for programmatic downloads of datasets. You can learn more about the Mozilla Data Collective [Alpha launch here](https://community.mozilladatacollective.com/mozilla-data-collective-alpha-goes-live/). **Common Voice is a community effort:** We want to extend our thanks to the language communities, activists, developers and community members who made this release possible. A special thanks to Zakia Mustafa for running a Kiswahili community event that validated over 2400 clips. Zakia is a long-time advocate for open technology and language inclusion, with a passion for building resources that strengthen the representation of low-resource languages in digital spaces. She has previously worked as a Mozilla Community Champion and continues to support community-driven initiatives through Common Voice. We’re taking ongoing applications for a volunteer program to help enable community members like Zakiya to do outreach in your language communities with funding and logistics support. You can learn more about [the program here](https://discourse.mozilla.org/t/applications-open-common-voice-contributor-support-program/144118?ref=community.mozilladatacollective.com) or apply [directly via this form](https://docs.google.com/forms/d/e/1FAIpQLScPjIsEcZhAksNEnFgD-bdztSwGg%5FjpbsdU75DwRlVQJZDmAg/viewform?ref=community.mozilladatacollective.com). **Tell us what you think:** We’re always excited for your feedback. You can reach us on [Discord](https://discord.gg/A4sVvDYdwG?ref=community.mozilladatacollective.com), Matrix, GitHub or email us directly at [commonvoice@mozilla.com](mailto:commonvoice@mozilla.com). We also run regular drop-in calls with the community and will be running office hours calls on [October (23-10-2025 7am GMT)](https://mozilla.zoom.us/meeting/register/hQMe8v-7TuiQtRXJY5XZFA?ref=community.mozilladatacollective.com) and [November: (20-11-2025 5pm GMT)](https://mozilla.zoom.us/meeting/register/vyVop0JvTBienKdXLWSreQ?ref=community.mozilladatacollective.com). We’d love to hear what you think about this release and what you’d like to see next. ### Mozilla Data Collective Alpha Goes Live URL: https://community.mozilladatacollective.com/mozilla-data-collective-alpha-goes-live/ Last updated: 2025-09-29T15:27:42.000Z Mozilla Data Collective is live in Alpha – the new platform from Mozilla Foundation that puts communities in control of how datasets are shared! Starting today, the latest Common Voice datasets (23.0) are available to download through Mozilla Data Collective. We’re so excited to introduce 149 new languages in this release, alongside an important first: Spontaneous Speech datasets, which include transcribed, spontaneous responses to prompts that help train models on more realistic speech patterns. [Create a Mozilla Data Collective account ](https://datacollective.mozillafoundation.org/auth/signup?ref=community.mozilladatacollective.com)to download all the datasets. From there, you can use the new datasheets pages to explore and access resources, access. You can also get access [via API or integrate them easily with our new ](https://datacollective.mozillafoundation.org/api-reference?ref=community.mozilladatacollective.com)[open-source Python library](https://github.com/Mozilla-Data-Collective/datacollective-python?ref=community.mozilladatacollective.com), which allows easy access to datasets programmatically. This means increased global access, more options for dataset downloads and new ways to work with data at scale. Common Voice datasets are just the beginning! Mozilla Data Collective is a platform to let all dataset communities and creators share their data under their own terms. We’ll be adding more partner datasets soon. If you have a dataset you would like to see on Mozilla Data Collective, tell us about it at [mozilladatacollective@mozillafoundation.org](mailto:mozilladatacollective@mozillafoundation.org). Mozilla Data Collective wouldn’t exist without the language communities, dataset users and amazing community that makes up Common Voice. Thank you for building with us. We’re excited to hear your feedback, wishlists and ideas for what you want us to build, so get in touch at [mozilladatacollective@mozillafoundation.org](mailto:mozilladatacollective@mozillafoundation.org). ### Your datasets, under your control: Mozilla Data Collective at PyConAU in Melbourne, Australia URL: https://community.mozilladatacollective.com/your-datasets-under-your-control-mozilla-data-collective-at-pyconau-in-melbourne-australia/ Last updated: 2025-09-15T02:55:57.000Z This weekend, Kathy Reid will present at [PyConAU](https://2025.pycon.org.au/?ref=community.mozilladatacollective.com) in Melbourne, Australia about the upcoming Mozilla Data Collective (MDC) platform - a sister platform to [Common Voice](https://commonvoice.mozillafoundation.org/?ref=community.mozilladatacollective.com). *In this blog post, we preview Kathy’s talking points, and show how you can get involved in MDC.* ## Tokenomics: why human-generated data is so valuable Kathy will first give an overview of tokens - building blocks of data - and how large language models work by creating relationships between tokens. AI models need ever-increasing volumes of tokens – billions and trillions. In fact, estimates for the the volume of training data required by GPT-5 suggest that it was trained on 114 billion tokens of data - orders of magnitude more data then earlier models like GPT-2 and GPT-3\. This increase in demand for training data is shown in the image below ![LLM Models: number of training tokens plotted over release year, showing steep increase to 114 billion tokens for GPT-5](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/09/data-src-image-64d6d9d1-e1bf-4caa-84c1-095a2e8270af.png) Source: [https://github.com/KathyReid/token-wars-dataviz](https://github.com/KathyReid/token-wars-dataviz?ref=community.mozilladatacollective.com) And we’re reaching the limits of just how much data is available on the open web. And while synthetic data - data produced by generative models - can help fill the gaps in some cases, training those same models with synthetic data leads to a condition called [model collapse](https://en.wikipedia.org/wiki/Model%5Fcollapse?ref=community.mozilladatacollective.com). So, just like fossil fuel, tokens are becoming scarce, rare, and contested ## Token harvesting ### Extractive practices And harvesting tokens to train AI is now big business. But tokens are scraped in ways that are exploitative. Many system administrators now spend a lot of time and energy preventing bots and scrapers from overloading their infrastructure. And harvesting tokens is not just exploitative of hosting infrastructure - token harvesting also extracts value from data creators - whether that’s writers, or artists, or corporations who put content on the web. And it’s also exploitative of the workers who are paid pennies to label, or moderate the content that is scraped. So, it’s exploitative all around - technically, from an infrastructure perspective, creatively and from a labour perspective. ![Newspaper headline showing three authors, unhappy that their text was scraped to train AI](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/09/data-src-image-a079b51e-5a42-4808-b784-c2101532c309.png) Nicola Heath, “[Australian authors challenge Productivity Commission’s proposed copyright law exemption for AI](https://www.abc.net.au/news/2025-08-13/productivity-commission-ai-report-copyright-law-authors-respond/105646086?ref=community.mozilladatacollective.com)”, ABC News, 13th August 2025 ### Concentration of power Token harvesting also works to concentrate power - meaning that it’s a small group of players shaping AI innovation. If you have a lot of data - like Reddit or Stack Overflow - then you can block certain scrapers, and only allow some companies to scrape your data - in exchange for wads of cash - which is exactly what Reddit and Stack Overflow have done, inking exclusive agreements with OpenAI. ### Not everyone is represented in harvested data And importantly, all those billions of tokens that have been scraped from the web don’t represent *everyone*. The language, the words, the sentences that are harvested represent a particular culture - often a dominant culture - but not *everyone*. Compare the writing on Reddit with the Washington Post or with a Tumblr blog - they’re vastly different styles, word choices and forms of expression. Moreover, the internet is predominantly English - so token harvesting for AI training also reduces the opportunities for AI to develop in the other 7000 languages still spoken today The internet is also full of biases - sexism, racism - and if we scrape the internet blindly and feed it to AI models, they’re going to inherit those biases. ### Bulldozing rights And current approaches to token harvesting bulldoze rights. They don’t respect values like data sovereignty and [CARE principles for the treatment of Indigenous data](https://www.gida-global.org/care?ref=community.mozilladatacollective.com). Language encodes culture. Language encodes history. Language encodes stories, and songs and dreams. And if we scrape language data without permission, we’re scraping culture and scraping history and scraping stories and scraping songs and dreams. ## The Mozilla Data Collective: A better way In an industry that relies on extractive practices, Mozilla Data Collective is rebuilding the AI data ecosystem - with communities at the centre. So, what are the two sides of the platform? ### Data contributors Data contributors are those people who make datasets available through the Mozilla Data Collective platform. **Researchers and non-profits** often have valuable data, and want to help that data be discovered to increase its impact - especially because it’s often expensive to produce **Creatives, media organisations and SMEs** want generate revenue and benefits for the people who’ve created the data - the words, the images, the creative works. And **governments, funders and public knowledge organisations** also have datasets they want to make public - unlocking their benefits for AI innovation. *But at the moment, there are limited options for these data contributors to share their data on their terms.* ### Data consumers And on the other side of the platform we have Data Consumers - people who need data. **Model trainers and data engineers** need high quality datasets - human, authentic, curated with care, ready for training. **Technical organisations** \- those deploying models into production - want to be able to connect with the communities behind datasets - so they can ask questions like “where did this data come from” or “what decisions are behind this data” and “how does this data reflect community values?”. And **compliance professionals** want to be able to de-risk their use of data by understanding its context - where and who and what it came from. *Again, there are limited options for data consumers to obtain high quality, human data in a compliant and ethical way.* ### Mozilla Data Collective - a better way The Mozilla Data Collective brings together these two groups - **data contributors and data consumers** \- unlocking data abundance by giving people and communities control over their data. **Data contributors** can promote their datasets, see how they’re being shared, and unlock new value by combining them with other datasets. **Data consumers** can find the data they’re looking for, connect with communities who created that data, and know that they have supply chain transparency. ## Join us! And we’d love for you to join us! [Come chat with us, send us feedback - or let us know about a dataset you think should be on the Mozilla Data Collective](mailto:mozilladatacollective@mozillafoundation.org?Subject=Query%20from%20the%20PyConAU%20blog%20post). And you can connect with us on social media. [![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/09/data-src-image-58246829-2617-41d0-887f-e14f0fef232a.png "Discord icon.png")](https://discord.gg/4TjgEdq25Y?ref=community.mozilladatacollective.com) [![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/09/data-src-image-eeb221fa-559b-4477-819a-c3a48c53ff4d-1-1-1-1.png)](https://mozdatacollective.bsky.social/?ref=community.mozilladatacollective.com) [![](https://storage.ghost.io/c/ff/ca/ffcaf12e-8ea8-4d92-a7c4-6446e2c332cd/content/images/2025/09/data-src-image-3ffe9229-abe6-4a45-9c0b-60d7f697e02d.png)](https://https//fosstodon.org/@mozilladatacollective?ref=community.mozilladatacollective.com) ## Slides [YourDataSetsYourControlPyConAU2025Kathy Reid's slide deck from PyConAU 2025YourDataSetsYourControlPyConAU2025.pdf4 MBdownload-circle](https://community.mozilladatacollective.com/content/files/2025/09/YourDataSetsYourControlPyConAU2025.pdf "Download") You can see the slides from the talk above (PDF, 4.2MB) ## Further reading - Associated Press. A[I startup Anthropic agrees to pay $1.5bn to settle book piracy lawsuit](https://www.theguardian.com/technology/2025/sep/05/anthropic-settlement-ai-book-lawsuit?ref=community.mozilladatacollective.com). The Guardian. September 5, 2025. - Cummins M. [How much LLM training data is there, in the limit? ](https://www.educatingsilicon.com/2024/05/09/how-much-llm-training-data-is-there-in-the-limit/?ref=community.mozilladatacollective.com)Educating Silicon. May 9, 2024.[ ](https://www.educatingsilicon.com/2024/05/09/how-much-llm-training-data-is-there-in-the-limit/?ref=community.mozilladatacollective.com) - Global Indigenous Data Alliance. [CARE Principles for Indigenous Data Governance](https://www.gida-global.org/care?ref=community.mozilladatacollective.com).[ ](https://www.gida-global.org/care?ref=community.mozilladatacollective.com) - Hao K. [Artificial intelligence is creating a new colonial world order.](https://www.technologyreview.com/2022/04/19/1049592/artificial-intelligence-colonialism/?ref=community.mozilladatacollective.com) MIT Technology Review. April 19, 2022. - Heath N. [Authors warn AI copyright exception a “free pass” for Big Tech to steal work](https://www.abc.net.au/news/2025-08-13/productivity-commission-ai-report-copyright-law-authors-respond/105646086?ref=community.mozilladatacollective.com). August 13, 2025\. - Jones PL, Mahelona K, Duncan S, Leoni G. Kaitiaki: closing the door on open Indigenous data. *International Journal on Digital Libraries*. 2025;26(1):1\. doi:[10.1007/s00799-025-00410-2](https://doi.org/10.1007/s00799-025-00410-2?ref=community.mozilladatacollective.com) - OpenAI. [OpenAI and Reddit Partnership](https://openai.com/index/openai-and-reddit-partnership/?ref=community.mozilladatacollective.com). May 16, 2024\. ### FAQ: What is the MDC Public API? URL: https://community.mozilladatacollective.com/faq-what-is-the-mdc-public-api/ Last updated: 2026-05-14T15:15:26.000Z Users must agree to dataset terms through the web interface before downloading. Each download token can only be used for one complete download session, and downloads are proxied through our API server. [API Reference Documentation](https://dev.mozilladatacollective.com/api-reference/docs?ref=community.mozilladatacollective.com) ### FAQ: Can I get the Common Voice or other MDC datasets from other platforms like GitHub or Hugging Face? URL: https://community.mozilladatacollective.com/faq-can-i-get-the-common-voice-or-other-mdc-datasets-from-other-platforms-like-github-or-hugging-face/ Last updated: 2026-05-13T19:25:30.000Z We have no plans to host Mozilla community datasets through third parties at this time, as it makes governance and stewardship extremely challenging. For example, when someone chooses to revoke their consent to be included in a dataset, we need a way to remove them from the dataset and update the data listings. When a given dataset is mirrored and hosted in multiple places, it becomes difficult to respect these requests and ensure that available versions of the dataset exclude those individuals. Mozilla community datasets, including Mozilla Common Voice datasets are exclusively available through MDC for this reason. Our new terms reflect this. Some of our contributors’ open datasets are available in other places. We want to make sure that those of you who enjoy Hugging Face’s great model and training features can still use them easily, so we’ve published an [API reference page](https://datacollective.mozillafoundation.org/api-reference?ref=community.mozilladatacollective.com) with instructions on how to create access credentials and download datasets programmatically. ### Roadmap URL: https://community.mozilladatacollective.com/roadmap/ Last updated: 2025-09-03T19:12:40.000Z We're excited to share a high-level roadmap for the Mozilla Data Collective platform, leading up to our 1.0 launch in early Q1 2026: **September: Mozilla Data Collective Alpha Launch** **October: New Datasets Available** **November: Mozilla Data Collective Beta Launch** - Dataset and datasheet on-boarding and upload flow - Feature support for conditional dataset access **Q1 2026: Mozilla Data Collective 1.0 Release** - Support for financial or ecosystem contributions as specified by dataset owner - Organizational accounts - Extended dataset augmentation and monitoring ### FAQ: What is the long-term sustainability model for Mozilla Data Collective? URL: https://community.mozilladatacollective.com/faq-what-is-the-long-term-sustainability-model-for-mozilla-data-collective/ Last updated: 2026-05-14T15:17:18.000Z We are a mission-driven, community-centred tech organisation that was incubated at Mozilla Foundation. As of April 2026, Mozilla Data Collective is an independent for-profit entity based in the United Kingdom. We anticipate that we will have both social enterprise and non-profit components eventually, as we think that in today’s volatile, politicised grant funding environment, it’s never been more important to be independent. Our business model is simple, ethical and transparent. In the future, if you choose to ask for financial contributions to make use of your dataset, we will charge the downloader a 5% platform fee. That’s it! This will help cover the platform development, storage and hosting costs. One day, we may ask those using datasets at scale - such as major corporations - to pay a bit more for an enterprise-grade API. ### FAQ: Who is behind Mozilla Data Collective? URL: https://community.mozilladatacollective.com/faq-who-is-behind-mozilla-data-collective/ Last updated: 2025-08-26T14:36:24.000Z We are backed and stewarded by Mozilla Foundation - the non-profit, movement-building, and philanthropy arm of Mozilla. ### FAQ: How does Mozilla Data Collective work? URL: https://community.mozilladatacollective.com/faq-how-does-mozilla-data-collective-work/ Last updated: 2026-05-13T19:24:03.000Z We partner with organizations and individuals to make their data available through Mozilla Data Collective. You can share openly, using existing licenses like Creative Commons, or you can build your own. You can open up your data for everyone, or just for some types of downloaders, you can set custom constraints, ask for exchange, compensation or recognition. You can govern it as an individual, a co-operative, a trust or something else. After all, it’s your data. The people who access your datasets are authenticated, and held in legally binding contracts, and we have a number of dataset protection features. If you are interested in hosting data on Mozilla Data Collective, please reach out to us at mozilladatacollective@mozillafoundation.org. ### FAQ: What is Mozilla Data Collective? URL: https://community.mozilladatacollective.com/faq-why-mozilla-data-collective/ Last updated: 2026-05-13T19:23:51.000Z Mozilla Data Collective is a platform in the truest sense. It’s yours to stand on, and make of it what you will. We have dual roots in two Mozilla projects - Common Voice, a CC0 public dataset to help tech speak your language - and the Data Futures Lab - an experimental space for instigating new approaches to data stewardship challenges. Mozilla Data Collective works by allowing you to share your data, retain ownership of it, and control who uses it. ### Exciting News! Mozilla Data Collective URL: https://community.mozilladatacollective.com/coming-soon/ Last updated: 2025-08-07T13:10:36.000Z Over the last eight years, the Common Voice community has shared wishlists with us for ways to create, curate, and control their data that extend beyond our current platform capabilities. For example, supporting the collection and release of datasets under different licences to CC-0, and the ability to contribute datasets collected externally to Common Voice. We’re so excited to announce that we have grown the team that developed Common Voice, and are expanding to build **Mozilla Data Collective,** a sister platform to enable you and your communities to share datasets in new ways. We’ll be building and releasing this platform throughout this year to help dataset owners and data creators to share their data with developers, researchers and others on their own terms. We wanted to share our plans with our Common Voice community early on to make space for you to participate in shaping Mozilla Data Collective. We want the Collective to meet your needs, and support Common Voice thriving alongside it. **What does this mean for Common Voice?** More support and optionality for Common Voice community members! Mozilla Data Collective is designed to help bring the community new options for control over more types of their own data. The [Common Voice platform](https://commonvoice.mozilla.org/?ref=community.mozilladatacollective.com) will continue to have full-time engineering support focused on improving what exists, and will remain accessible under the same MPL 2.0 licence. Common Voice will continue to be supported by Mozilla Foundation team members, our contributors and the wider community. The existing Common Voice datasets will continue to be accessible under the CC-0 licence they were released with. Historical datasets will continue to be available through the Common Voice website. Future versions of the datasets will be released through Mozilla Data Collective. **What are the benefits for my language community?** The experience will be significantly improved. We will be: - integrating robust, detailed datasheets about what is contained in the datasets - adding programmatic access through a developer API By keeping the services somewhat separate, we will be optimising for a high speed, and scalable performance, which are all issues on which we have received helpful feedback- and we thank the Common Voice community for your valuable inputs. As part of the work with Mozilla Data Collective, Common Voice will also be offering more granularity around licensing and access options to our data communities. For instance, the Creative Commons licences CC-BY, CC-BY-SA or customised versions like the Nwulite Obodo (NOODL) licence. **Next steps:** Common Voice dataset users, contributors or community members don’t have to do anything right now. We’ll share more information about Mozilla Data Collective to this channel soon. **Questions and participation:** We would love to hear your thoughts. You can also chat to us about datasets you’d like to explore making available through MDC - the MCV community will of course be getting early access! You can email the team at [commonvoice@mozilla.com](mailto:commonvoice@mozilla.com) any time, in the language you’re most comfortable using. In the next open office hours session on the [28th August, 2025 ](https://mozilla.zoom.us/meeting/register/qoxMhoXtRQuTbOPAWQSYhw?ref=community.mozilladatacollective.com)if you would like to discuss Mozilla Data Collective or any other Common Voice topics with the team and community. See you there!