Compensated Datasets Is Now Available on Mozilla Data Collective

Share
Compensated Datasets Is Now Available on Mozilla Data Collective

Helping organisations participate more directly in the AI economy while making it easier for AI builders to discover responsibly sourced datasets.

Earlier this month, we shared a preview of Compensated Datasets and our vision for creating more transparent ways for organisations to participate in the AI economy while retaining agency over how their data is licensed and used.

Today, we're excited to make Compensated Datasets publicly available on Mozilla Data Collective.

High-quality data is the foundation of better AI. At the same time, many of the organisations and communities creating and stewarding that data have had few transparent ways to participate in the AI economy while retaining agency over how their data is licensed and used.

Compensated Datasets helps bridge that gap by enabling organisations to make datasets available through Mozilla Data Collective while retaining control over pricing and licensing terms. In turn, AI builders can discover multilingual, multicultural and multimodal datasets that are responsibly sourced, thoughtfully documented and ready to support the next generation of AI applications.

At launch, Compensated Datasets includes contributions from TAUS, Pangeanic, Karya, Spotlite and ContentX Labs, among others, representing a growing collection of datasets across languages, domains and modalities.

This launch reflects the type of ecosystem we're building: one where AI builders have access to high-quality datasets, and where the organisations and communities creating that data are recognised, supported, and able to participate more directly in the value it creates.

Whether you're looking for high-quality datasets for your next AI project or you're interested in making your own datasets available through Mozilla Data Collective, Compensated Datasets is designed to make those connections easier through a marketplace built on transparency, trust and shared value.

Browse available compensated datasets.

If you're interested in becoming a data provider, we'd love to hear more about your work.

Tell us about your dataset.

This is just the beginning

Today's launch marks an important milestone for Mozilla Data Collective, but it's only the beginning.

We'll continue expanding the community of organisations contributing datasets, growing the range of languages, domains and modalities available, and making it easier for AI builders and data providers to connect through a marketplace built on transparency, trust and shared value.

Whether you're looking for your next dataset or interested in contributing one of your own, we're excited to have you as part of the Mozilla Data Collective community.

Read more

Mozilla Data Collective datasets now discoverable through CLARIN’s Virtual Language Observatory

New collaboration expands visibility for community-governed language datasets and improves exploration of linguistic resources, services and tools. Mozilla Data Collective datasets are now discoverable through CLARIN’s Virtual Language Observatory, making it easier for researchers, developers and language technology practitioners in Europe to find multilingual and community-centered datasets