# Mozilla Data Collective > Multilingual, multicultural, multimodal datasets under range of licenses: from open source, to community controlled, to commercial and compensated. Public Ghost content for AI and LLM tooling. Use `/llms-full.txt` for consolidated page and post context. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages - [About Mozilla Data Collective](https://community.mozilladatacollective.com/about.md) - What is Mozilla Data Collective’s mission? We fight for a tech future that is multilingual, multicultural and multimodal. We think we deserve a world where you don’t need to change how you speak or how you look in order to access technology - and we welcome technology’s promise as a connector, enab… - [Careers at Mozilla Data Collective](https://community.mozilladatacollective.com/careers-at-mozilla-data-collective.md) - About Mozilla Data Collective Mozilla Data Collective is the data sharing platform for human agency and fair value exchange. Our vision is a tech future that is multilingual, multicultural, and multimodal; our mission is to achieve that future by redefining how AI data is created, shared, and gover… - [Talks](https://community.mozilladatacollective.com/dev-talks.md) - Learn about Mozilla Data Collective from talks that our team has given and join the movement to reclaim your data. * ➜ Fine-Tuning a Whisper Model with MDC Datasets * ➜ Beyond Extraction: Building Community-Centered Speech Data * ➜ Your datasets, under your control: Introducing Mozilla Data Collect… - [Join Mozilla Data Collective](https://community.mozilladatacollective.com/join.md) - Mozilla Data Collective is a platform in the truest sense. It’s yours to stand on, and make of it what you will. We have dual roots in two Mozilla projects - Common Voice, a CC0 public dataset to help tech speak your language - and the Data Futures Lab - an experimental space for instigating new ap… - [MDC Data Licence Agreement 1.0](https://community.mozilladatacollective.com/mdc-data-licence-agreement-1-0.md) - Last Updated: 7 July 2026 This Data Licence Agreement for MDC Datasets (this “Licence Agreement”) is entered into by and between Mozilla Data Collective, LTD, a company registered in England and Wales with company number 17054959, whose registered address is located at 167-168 Great Portland Street… ## Posts - [NOW LIVE: Lost in Transcription Competition](https://community.mozilladatacollective.com/no.md) - Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts This competition evaluates performance on real-world bilingual dialogues in Indonesian-Javanese, Nahuatl-Spanish, and Spanish-English. For speech technologies to truly have an impact on the world, they must adap… - [Compensated Datasets Is Now Available on Mozilla Data Collective](https://community.mozilladatacollective.com/compensated-datasets-is-now-available-on-mozilla-data-collective.md) - Helping organisations participate more directly in the AI economy while making it easier for AI builders to discover responsibly sourced datasets. Earlier this month, we shared a preview of Compensated Datasets and our vision for creating more transparent ways for organisations to participate in th… - [Common Voice segments now available through Mozilla Data Collective](https://community.mozilladatacollective.com/common-voice-segments-now-available-through-mozilla-data-collective.md) - Mozilla Common Voice is a massively multilingual platform for collecting speech data to train automatic speech recognition (ASR). Its mission is simple: to make language technology understand everyone’s mother tongue. But for datasets to be genuinely useful, they also need to be manageable. Many of… - [Mozilla Data Collective datasets now discoverable through CLARIN’s Virtual Language Observatory](https://community.mozilladatacollective.com/mozilla-data-collective-datasets-now-discoverable-through-clarins-virtual-language-observatory.md) - New collaboration expands visibility for community-governed language datasets and improves exploration of linguistic resources, services and tools. Mozilla Data Collective datasets are now discoverable through CLARIN’s Virtual Language Observatory, making it easier for researchers, developers and l… - [Get a Sneak Preview of Mozilla Data Collective’s Compensation Feature!](https://community.mozilladatacollective.com/get-a-sneak-preview-of-mozilla-data-collectives-compensation-feature.md) - Mozilla Data Collective was built to redefine how AI data is created, shared, and governed. As part of our mission to be the data sharing platform for human agency and fair value exchange, we have long-teased what so many partners and community members have requested: a tangible way to ensure that… - [15 Datasets for Fine-Tuning Whisper on a New Language in 2026](https://community.mozilladatacollective.com/15-datasets-for-fine-tuning-whisper-on-a-new-language-in-2026.md) - Why fine-tuning Whisper is a dataset problem OpenAI's Whisper changed what's possible in speech recognition: a single multilingual model with strong zero-shot performance across dozens of languages. But "dozens" is the catch. But for the long tail of the world's languages, even for some with tens o… - [How Radio Free Europe/Radio Liberty Also Serves Its Communities Through Its Datasets](https://community.mozilladatacollective.com/how-radio-free-europe-radio-liberty-also-serves-its-communities-through-its-datasets.md) - For more than 75 years, Radio Free Europe/Radio Liberty (RFE/RL) has promoted democratic values by providing accurate, uncensored news and debate in countries where a free press is threatened. RFE/RL reaches more than 44 million people every week across 18 countries, in 24 languages, including Pers… - [What makes a good dataset sample — and how to create one](https://community.mozilladatacollective.com/what-makes-a-good-dataset-sample.md) - In this post, we walk you through how to create a useful dataset sample as a preview of your dataset, and guide you in uploading it to the MDC platform. - [15 Datasets for Building a Low-Resource Translation Model in 2026](https://community.mozilladatacollective.com/15-datasets-for-building-a-low-resource-translation-model-in-2026.md) - The problem with "low-resource" machine translation Most production machine-translation systems in 2026 are still trained on a fairly narrow set of language pairs: the 50 or so for which the open web supplies enough parallel text to push BLEU scores into useful territory. Below that line, MT qualit… - [15 Datasets for Building a Production TTS Voice in 2026](https://community.mozilladatacollective.com/15-datasets-for-building-a-production-tts-voice-in-2026.md) - A curated list of 15 text-to-speech training datasets for teams shipping production voice models in 2026 covering emotional, multi-speaker, audiobook-derived, non-Latin script, indigenous-language datasets and more. - [Open Home Foundation TTS datasets on Mozilla Data Collective](https://community.mozilladatacollective.com/open-home-foundation-tts-datasets-on-mozilla-data-collective.md) - Most voice assistants listen and respond in a handful of languages. Try to build one for your home that speaks your language, though, and you quickly run into a wall: the training data does not exist, or it is locked behind licences that make it unusable for open source projects. The Open Home Foun… - [Discover Dataset Insights with the new Data Provider Analytics Portal](https://community.mozilladatacollective.com/discover-dataset-insights-with-the-new-data-provider-analytics-portal.md) - Today, we're excited to share a new way for dataset providers to better understand how their datasets are being used on Mozilla Data Collective with a new data provider analytics portal. - [Building an African Voice for AI: Inside the Institute of African Digital Humanities](https://community.mozilladatacollective.com/building-an-african-voice-for-ai-inside-the-institute-of-african-digital-humanities.md) - When you ask a voice assistant a question in English, French, or Mandarin, the underlying models have been trained on billions of words and millions of hours of speech. Ask the same question in Bafia, Mada, or Suundi, and the technology simply doesn't know how to listen. The Institute of African Di… - [Never Miss a Dataset with the new Dataset Notification Feature](https://community.mozilladatacollective.com/never-miss-a-dataset-with-the-new-dataset-notification-feature.md) - Today, we're excited to share a new way to stay informed about the latest datasets on Mozilla Data Collective - the ability to subscribe to get updated about similar datasets to your previous downloads. - [CoVoST 2 datasets now available through Mozilla Data Collective](https://community.mozilladatacollective.com/covost-2-datasets-now-available-through-mozilla-data-collective.md) - In 2020, Meta introduced a new benchmark dataset based on Mozilla Common Voice. CoVoST 2 has 34 translation directions for audio to text machine translation. This is one of the most widely-used benchmark datasets for speech translation, with nearly 400 citations on Google Scholar. The dataset is ba… - [Press Release](https://community.mozilladatacollective.com/press-release-new-capabilities-expand-uploader-control-over-access-and-compensation.md) - New capabilities expand uploader control over access and compensation, while helping developers discover more representative datasets - [Metadata magic: making datasets discoverable and tractable with Croissant](https://community.mozilladatacollective.com/metadata-magic-making-datasets-discoverable-and-tractable-with-croissant.md) - Learn how we automatically generate Croissant metadata to describe datasets on the Mozilla Data Collective platform, making them more discoverable. - [Best Practices for Sharing Contact Information on Mozilla Data Collective](https://community.mozilladatacollective.com/best-practices-for-sharing-contact-information-on-mozilla-data-collective.md) - In this guide, we'll walk through the different options available on the platform for sharing your contact information with downloaders and setting expectations about how downloaders or other community members can reach out to you. - [FAQ: Why Can't I Edit Certain Fields on my Published Dataset?](https://community.mozilladatacollective.com/faq-why-cant-i-edit-certain-fields-on-my-published-dataset.md) - Fields that make up the terms of use of your dataset cannot be edited after publishing. - [Engendering Voice: Insights from Common Voice](https://community.mozilladatacollective.com/engendering-voice-insights-from-common-voice.md) - The origin of Swahili language Swahili originated from the Indian ocean and has its route from the contacts of Arabian traders with the inhabitants of the east coast of Africa over many centuries. Swahili is the lingua franca of most east african countries spoken largely in countries such as Tanzan… - [The Mozilla Data Collective Data Assistant is now available in Alpha](https://community.mozilladatacollective.com/the-mozilla-data-collective-data-assistant-is-now-available-in-alpha.md) - We're excited to share that the Mozilla Data Collective Data Assistant is now available in Alpha. Visit https://mozilladatacollective.com/chat to get started. - [How Open Licensing is Changing with AI: The NOODL License](https://community.mozilladatacollective.com/how-open-licensing-is-changing-with-ai-the-noodl-license.md) - Author: Alek Tarkowski For the last twenty five years, standardized open licenses were increasingly seen as a main tool for democratizing access to knowledge. Over less then a decade, a relatively narrow set of canonical choices emerged: the Creative Commons licensing stack https://creativecommons.… - [Exciting Updates for Mozilla Data Collective](https://community.mozilladatacollective.com/exciting-updates-for-mozilla-data-collective.md) - Today we’re so excited to announce exciting platform changes to Mozilla Data Collective that bring us closer to fulfilling our promise to give everyone the Data Platform for Human Agency and Fair Value Exchange. - [Request to Access Feature is now Available](https://community.mozilladatacollective.com/request-to-access-feature-is-now-available.md) - With this release, uploaders can now gate access to their datasets, requiring users to request permission and share their email address before downloading. - [Fine-Tune a Speech-to-Text Model for Any Language - Including Yours](https://community.mozilladatacollective.com/fine-tune-a-speech-to-text-model-for-any-language-including-yours.md) - A step-by-step developer tutorial from Kostis at Mozilla Data Collective - [Datasheets: The Missing Manual for your Dataset](https://community.mozilladatacollective.com/datasheets-the-missing-manual-for-your-dataset.md) - In this video, produced by the Data Nutrition Project and illustrated by Jessica Yurkofsky, you'll learn more about the role of the datasheet and how you can use it to give clear guidance to potential downloaders about how your data can (and can't!) be used. - [Upcoming Domain Change on 09 April](https://community.mozilladatacollective.com/upcoming-domain-change.md) - In mid-April, Mozilla Data Collective's primary domain will change to mozilladatacollective.com - [MDC Release Notes - 30.03.26](https://community.mozilladatacollective.com/mdc-release-notes-30-03-26.md) - 379 new datasets with a Mozilla Common Voice update, improvements to the Python SDK (make sure you update to the latest!) and a preview of an upcoming feature. 👀 - [On contributing my Thorsten-Voice voice datasets](https://community.mozilladatacollective.com/on-contributing-my-thorsten-voice-voice-datasets.md) - Thorsten Müller has created five TTS voice datasets totaling 40 hours of German speech data. In this community-authored post, he speaks to the importance of sharing his voice. - [On contributing my Thorsten-Voice voice datasets [DE]](https://community.mozilladatacollective.com/on-contributing-my-thorsten-voice-voice-datasets-de.md) - Thorsten Müller hat fünf TTS-Stimmdatensätze mit insgesamt 40 Stunden an deutschen Sprachdaten erstellt. In diesem von der Community verfassten Beitrag spricht er über die Bedeutung, seine Stimme zu teilen. - [Why Whisper Still Struggles with Australian English - and What We Did About It](https://community.mozilladatacollective.com/why-whisper-still-struggles-with-australian-english-and-what-we-did-about-it.md) - By Kathy Reid · Mozilla Data Collective Ask a Queenslander to say "no worries" into a Home Assistant Voice Preview. There's a decent chance it mishears them. Ask someone from Radelaide and things get worse. Ask anyone who grew up calling things "heaps good" and you'll start to understand that the p… - [How can a nearly century-old publisher remain relevant and continue to grow?](https://community.mozilladatacollective.com/how-can-a-nearly-century-old-publisher-remain-relevant-and-continue-to-grow.md) - Panjebar Semangat, a weekly Javanese-language magazine established before Indonesian independence, is collaborating with Mozilla Data Collective to advance community-governed language dataset frameworks. - [Call for proposals: MDC is commissioning mission-aligned datasets!](https://community.mozilladatacollective.com/call-for-proposals-mdc-is-commissioning-mission-aligned-datasets.md) - Mozilla Data Collective is building towards a multicultural, multilingual, and multimodal future that works for all of us. And over the past few months, we’ve listened as people have flagged what kinds of datasets they need, but are struggling to find. So we’re pleased to announce that MDC will be… - [MDC Release Notes - 13.03.26](https://community.mozilladatacollective.com/mdc-release-notes-13-03-26.md) - This week: 19 new datasets and a few small changes while we're heads down in some exciting new features that will be coming soon... - [How to License Your Dataset for AI Training: Some Best Practices](https://community.mozilladatacollective.com/how-to-license-your-dataset-for-ai-training-some-best-practices.md) - We get a lot of questions about how to approach licensing your data for AI training. So to help you share your datasets, we’ve compiled some guidance here – it’s intended to be a living document, that we iterate with our partners and communities. Explore Mozilla Data Collective What Does It Mean to… - [Behind the scenes: Integrating MDC datasets into your Python project](https://community.mozilladatacollective.com/behind-the-scenes-integrating-mdc-datasets-into-your-python-project.md) - Overcoming the complexity of AI Mozilla Data Collective helps communities to offer unique, multilingual, multicultural, and multimodal datasets. From transcribed and translated videos of narrated Ekpeye folktales to complex question-answering text pairs for the Georgian language, the diversity of d… - [Cultural Heritage and AI: How Institutions Can Reclaim Control of Their Data](https://community.mozilladatacollective.com/cultural-heritage-and-ai-how-institutions-can-reclaim-control-of-their-data.md) - The institutions that safeguard humanity's cultural memory, galleries, libraries, archives, and museums (collectively known as the GLAM sector) are confronting a paradox that defines the current moment in AI development. Years of careful digitization of their archives have transformed physical coll… - [Using the MDC Python SDK Library to Download Datasets](https://community.mozilladatacollective.com/using-the-mdc-python-sdk-library-to-download-datasets.md) - In this guide, you will learn how to use the MDC Python SDK Library to download datasets from the Mozilla Data Collective website. - [MDC Release Notes - 27.02.26](https://community.mozilladatacollective.com/mdc-release-notes-27-02-26.md) - This week: dataset filtering, enhanced uploader request flow, API improvements, and 20 new datasets! - [MDC Release Notes - 13.02.26](https://community.mozilladatacollective.com/mdc-release-notes-13-02-26.md) - This week: new features for uploaders, updates to dataset search, and new datasets on MDC! - [Your Data, Your Rules: A Community Workshop for Dataset Governance](https://community.mozilladatacollective.com/your-data-your-rules-a-community-workshop-for-dataset-governance.md) - A practical guide for communities creating datasets together—no legal expertise required. Based on our data governance workshop at Mozilla Festival Zambia 2024 Your community has created something valuable: a dataset. Maybe it's voice recordings in your language. Maybe it's traditional knowledge, l… - [Help Shape the Future of Mozilla Data Collective](https://community.mozilladatacollective.com/help-shape-the-future-of-mozilla-data-collective.md) - Mozilla Data Collective has an amazing opportunity for you to get a free ticket to the 2026 Mozilla Festival in beautiful Barcelona, Spain this November, 2026. We are looking for feedback to inform our 2026 roadmap. Help shape the future of ethical data-sharing by filling out the form below - [FAQ: How can I remove published datasets from MDC?](https://community.mozilladatacollective.com/faq-how-can-i-remov.md) - In your datasheet settings, there is a button to make your dataset private. - [Turning Your Data Into a Valuable ML Resource Without Giving Up Control](https://community.mozilladatacollective.com/turning-your-data-into-a-valuable-ml-resource-without-giving-up-control.md) - You might be sitting on something precious The modern world runs on data. One unfortunate result of this is the fact that many of us are unknowingly producing data for third-party companies, who use our content and actions as data points to make AI models that they then sell back to us in exchange… - [Latest: Asian Language datasets by communities on Mozilla Data Collective](https://community.mozilladatacollective.com/latest-asian-language-datasets-by-communities-on-mozilla-data-collective.md) - The internet belongs to everyone—but right now, it doesn't work for everyone. From the islands of Borneo to the mountains of Pakistan, hundreds of millions of people speak languages that AI simply can't understand. That's a problem we can solve together. At Mozilla Data Collective, we're proud to h… - [Latest: African Language datasets by communities on Mozilla Data Collective](https://community.mozilladatacollective.com/latest-african-language-datasets-by-communities-on-mozilla-data-collective.md) - The internet belongs to everyone—but right now, it doesn't work for everyone. Millions of people speak languages that AI simply can't understand, and that's a problem we can solve together. At Mozilla Data Collective, we're proud to host a growing collection of open datasets that are helping resear… - [Updates to MDC REST API and Python Library](https://community.mozilladatacollective.com/updates-to-mdc-rest-api-and-python-library.md) - Tl;dr: Update to newest version of the MDC Python library (0.2.0 or newer) to continue downloading datasets - [FAQ: How long does it take to publish my dataset on MDC?](https://community.mozilladatacollective.com/faq-how-long-does-it-take-to-publish-my-dataset-on-mdc.md) - We aim to review all dataset submissions within 7 days. - [Pashto becomes third-highest language by volume of data in Common Voice v24](https://community.mozilladatacollective.com/pashto-becomes-third-highest-language-by-volume-of-data-in-common-voice-v24.md) - Key highlights from the Common Voice v24 Scripted Speech and v2 Spontaneous Speech release. - [FAQ: How can I contribute to Mozilla Data Collective?](https://community.mozilladatacollective.com/faq-how-can-i-contribute-to-mozilla-data-collective.md) - Building with MDC datasets, uploading and sharing data, or joining our community are all ways to get involved. - [FAQ: Do I need to be part of an organization to upload a dataset to Mozilla Data Collective?](https://community.mozilladatacollective.com/faq-do-i-need-to-be-part-of-an-organization-to-upload-a-dataset-to-mozilla-data-collective.md) - No, you do not need to be a part of an organization to upload a dataset to Mozilla Data Collective. - [Improving the Spontaneous Speech English dataset: lifting the lid on speech data quality uplift techniques](https://community.mozilladatacollective.com/improving-the-spontaneous-speech-english-dataset.md) - Firstly, we’d like to thank you for your patience. After introducing Spontaneous Speech early in 2025, we released most locale datasets when the Mozilla Data Collective platform launched in alpha in September of this year. However, upon inspection, the English Spontaneous Speech dataset required so… - [Uploading your dataset to the Mozilla Data Collective Platform](https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform.md) - Interested in joining the movement and publishing your dataset on Mozilla Data Collective? This guide will walk you through the steps required, from account creation to submission! - [FAQ: What does it mean to exclusively host my dataset on Mozilla Data Collective?](https://community.mozilladatacollective.com/faq-what-does-it-mean-to-exclusively-host-my-dataset-on-moz.md) - When you exclusively host your dataset on MDC, you choose to only make it available on our site. - [FAQ: What kind of datasets can I publish on Mozilla Data Collective?](https://community.mozilladatacollective.com/faq-what-kind-of-datasets-can-i-publish-on-mozilla-data-collective.md) - Our priority is technology that is more multilingual, multicultural, and multi-modal. - [FAQ: What are the main points of the MDC Terms of Use?](https://community.mozilladatacollective.com/faq-what-are-the-main-points-of-the-mdc-terms-of-use.md) - When you use MDC to share or download datasets, you enter into an agreement directly with the data provider/consumer. - [What do the languages Nahuatl, Bahasa Indonesia and Bulgarian have in common?](https://community.mozilladatacollective.com/what-do-the-languages-nahuatl-bahasa-indonesia-and-bulgarian-have-in-common.md) - Nahuatl, Bahasa Indonesia and Bulgarian all feature in our very first community curated datasets to be uploaded to the Mozilla Data Collective platform. - [FAQ: Why can’t I download old versions of Common Voice datasets?](https://community.mozilladatacollective.com/faq-why-cant-i-download-old-versions-of-common-voice-datasets.md) - We limit access to old versions of Common Voice to respect those speakers who have withdrawn their consent to be included in the dataset. - [We're Changing Access to Older Versions of Common Voice datasets](https://community.mozilladatacollective.com/were-changing-access-to-older-versions-of-common-voice-datasets.md) - We’re tightening the circulation of old datasets to protect contributors, while keeping a clear, documented path for researchers who need them. - [FAQ: I've noticed an issue with a data listing - what do I do?](https://community.mozilladatacollective.com/faq-ive-noticed-an-issue-with-a-data-listing-what-do-i-do.md) - Depending on the nature of the issue, you can contact the dataset uploader, use the report dataset link, or contact us directly. - [Kick-off of the Shared Task: MCV Spontaneous Speech](https://community.mozilladatacollective.com/kick-off-of-the-shared-task-mcv-spontaneous-speech.md) - Join us for the official kick‑off of the Shared Task: Mozilla Common Voice Spontaneous Speech! This live, online gathering will bring together researchers, engineers, and language‑technology enthusiasts from around the world to launch the challenge focused on building robust, multilingual ASR syste… - [FAQ: Why can't I re-host or share Common Voice datasets that I download from MDC?](https://community.mozilladatacollective.com/faq-why-cant-i-re-host-or-share-common-voice-datasets-that-i-download-from-mdc.md) - By exclusively hosting Common Voice datasets on MDC, we are best able to govern and respect the wishes of contributors. - [Shared Task: Mozilla Common Voice Spontaneous Speech ASR](https://community.mozilladatacollective.com/shared-task-mozilla-common-voice-spontaneous-speech-asr.md) - Quick links * Registration Form * Link to datasets * Codabench page (to submit results during the testing period) * Contact: sharedtask@mozillafoundation.org Overview Automatic speech recognition (ASR) has come a long way – but most systems are still trained on polished, read-aloud speech. So we se… - [Common Voice 23.0 Live On Mozilla Data Collective](https://community.mozilladatacollective.com/common-voice-23-0-live-on-mozilla-data-collective.md) - Common Voice 23.0 is now live – and available for download via Mozilla Data Collective. Mozilla Data Collective is a sister platform from the team behind Common Voice, designed to let dataset owners and creators offer their data on their own terms. Mozilla Data Collective was built in response to c… - [Mozilla Data Collective Alpha Goes Live](https://community.mozilladatacollective.com/mozilla-data-collective-alpha-goes-live.md) - Mozilla Data Collective is now in live alpha, offering the Common Voice 23.0 datasets. - [Your datasets, under your control: Mozilla Data Collective at PyConAU in Melbourne, Australia](https://community.mozilladatacollective.com/your-datasets-under-your-control-mozilla-data-collective-at-pyconau-in-melbourne-australia.md) - Kathy Reid's presentation to PyConAU in Melbourne, Australia covers tokenomics, harvesting tokens, and how Mozilla Data Collective offers a better way forward. - [FAQ: What is the MDC Public API?](https://community.mozilladatacollective.com/faq-what-is-the-mdc-public-api.md) - The Mozilla Data Collective REST API provides a way for developers to access datasets from their own applications, using the programming language of their choice. - [FAQ: Can I get the Common Voice or other MDC datasets from other platforms like GitHub or Hugging Face?](https://community.mozilladatacollective.com/faq-can-i-get-the-common-voice-or-other-mdc-datasets-from-other-platforms-like-github-or-hugging-face.md) - Common Voice datasets are exclusively available through Mozilla Data Collective. - [Roadmap](https://community.mozilladatacollective.com/roadmap.md) - We're excited to share a high-level roadmap for the Mozilla Data Collective platform, leading up to our 1.0 launch in early Q1 2026: September: Mozilla Data Collective Alpha Launch October: New Datasets Available November: Mozilla Data Collective Beta Launch * Dataset and datasheet on-boarding and… - [FAQ: What is the long-term sustainability model for Mozilla Data Collective?](https://community.mozilladatacollective.com/faq-what-is-the-long-term-sustainability-model-for-mozilla-data-collective.md) - We ask for a 5% dataset access fee from downloaders of paid datasets and may offer additional other subscription capabilities in the future. - [FAQ: Who is behind Mozilla Data Collective?](https://community.mozilladatacollective.com/faq-who-is-behind-mozilla-data-collective.md) - We are backed and stewarded by Mozilla Foundation - the non-profit, movement-building, and philanthropy arm of Mozilla. - [FAQ: How does Mozilla Data Collective work?](https://community.mozilladatacollective.com/faq-how-does-mozilla-data-collective-work.md) - We partner with organizations and individuals to make their data available through Mozilla Data Collective. - [FAQ: What is Mozilla Data Collective?](https://community.mozilladatacollective.com/faq-why-mozilla-data-collective.md) - Mozilla Data Collective is a data platform for human agency and fair value exchange. - [Exciting News! Mozilla Data Collective](https://community.mozilladatacollective.com/coming-soon.md) - Over the last eight years, the Common Voice community has shared wishlists with us for ways to create, curate, and control their data that extend beyond our current platform capabilities. For example, supporting the collection and release of datasets under different licences to CC-0, and the abilit… ## Optional - [RSS Feed](https://community.mozilladatacollective.com/rss/) - [Sitemap](https://community.mozilladatacollective.com/sitemap.xml) - [Full content of pages and posts](https://community.mozilladatacollective.com/llms-full.txt)