> ## Content Index
> Fetch the complete content index at: https://community.mozilladatacollective.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# 15 Datasets for Building a Low-Resource Translation Model in 2026
- URL: https://community.mozilladatacollective.com/15-datasets-for-building-a-low-resource-translation-model-in-2026/
- Published: 2026-06-04T10:29:28.000Z
- Updated: 2026-06-04T10:29:28.000Z
- Author: Mozilla Data Collective
- Tags: News

## The problem with "low-resource" machine translation

Most production machine-translation systems in 2026 are still trained on a fairly narrow set of language pairs: the 50 or so for which the open web supplies enough parallel text to push BLEU scores into useful territory. Below that line, MT quality worsens significantly. For a Hausa-to-English translation system to be worth shipping, you need in the order of millions of aligned sentence pairs. For most of the world's 7,000+ languages, that volume of data doesn't exist anywhere and isn't going to be generated through web crawls alone.

What does exist, increasingly, is community-led parallel-data collection. [Translators without Borders](https://translatorswithoutborders.org/?ref=community.mozilladatacollective.com) (now [CLEAR Global](https://clearglobal.org/?ref=community.mozilladatacollective.com)) ran the TWB Gamayun parallel sentence kits specifically to give humanitarian-language translation projects a starting point. Pakistani publishers, Mexican linguists, Sephardic Jewish revitalisation projects, and Nigerian research groups have built smaller but carefully curated parallel corpora in languages mainstream MT simply ignores. Combined, these resources don't replace web-scale parallel data but they're enough to fine-tune an existing multilingual MT base model into something usable for a specific low-resource pair.

The 15 datasets below are the working set for that kind of project in 2026\. They range from the TWB Gamayun humanitarian sentence kits (Hausa, Lingala, Tigrinya, Nande, Rohingya, Swahili, Kanuri, Congo Swahili) to South Asian publishing-derived parallel corpora (Saraiki-English, English-Punjabi Shahmukhi), Sephardic Ladino lexical resources for Romance-family transfer, religious-domain multilingual parallel text, and a BOUQuET translation-difficulty evaluation sets that let you benchmark whatever model you end up training.

## The datasets

### African region

- [**TWB Parallel Sentence kits - Hausa (30k)**](https://mozilladatacollective.com/datasets/cmom43ixg00hao00731e0o0jg?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 1.68 MB | Task: MT | Format: TSV
- [**TWB Parallel Sentence kits - Congo Swahili (25k)**](https://mozilladatacollective.com/datasets/cmosl07v400w9nu07g3puif2t?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 2.18 MB | Task: MT | Format: TSV
- [**TWB Parallel Sentence kits - Nande (15k)**](https://mozilladatacollective.com/datasets/cmoskn32a00vhmj07prz6k5ng?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 1.26 MB | Task: MT | Format: TSV
- [**TWB Parallel Sentence kits - Swahili (5k)**](https://mozilladatacollective.com/datasets/cmoskxn8k00vtmj07ubxtu0f2?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 347.61 KB | Task: MT | Format: TSV
- [**TWB Parallel Sentence kits - Tigrinya (5k)**](https://mozilladatacollective.com/datasets/cmoskmbpj00vxnu07w8lu7rrk?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 404.75 KB | Task: MT | Format: TSV
- [**TWB Parallel Sentence kits - Lingala (5k)**](https://mozilladatacollective.com/datasets/cmosknxap00vlmj07kf6mugba?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 494.43 KB | Task: MT | Format: TSV
- [**TWB Parallel Sentence kits - Kanuri (5k)**](https://mozilladatacollective.com/datasets/cmoskkop100vjnu0775hclbd1?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 358.46 KB | Task: MT | Format: TSV
- [**English Hausa Parallel Corpus**](https://mozilladatacollective.com/datasets/cmn3ht40i00eami07lgydmrgg?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: LocaleNLP | Licence: CC-BY-NC-4.0 | Size: 164.32 KB | Task: MT | Format: CSV

### South and Southeast Asia

- [**TWB Parallel Sentence kits - Rohingya (5k)**](https://mozilladatacollective.com/datasets/cmoskwsac00w3nu07b1nydlfb?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: CLEAR Global | Licence: CC-BY-4.0 | Size: 358.88 KB | Task: MT | Format: TSV
- [**Saraiki-English Parallel Corpus**](https://mozilladatacollective.com/datasets/cmmaphscg04t2mk07i1f8yc0q?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Kaleem Art Press | Licence: CC-BY-NC-4.0 | Size: 1.92 MB | Task: MT | Format: CSV
- [**English–Punjabi (Shahmukhi) Parallel Sentences Corpus (Mediamen Archives)**](https://mozilladatacollective.com/datasets/cmkh9rso90076nv076jwxgjv3?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: MEDIAMEN | Licence: CC-BY-NC-4.0 | Size: 1.08 MB | Task: MT | Format: CSV
- [**Multilingual Religious Parallel Corpus (Kaleem Art Press)**](https://mozilladatacollective.com/datasets/cmk1bhogs3htwmk07wo7o9p6y?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Kaleem Art Press | Licence: CC-BY-SA-4.0 | Size: 2.27 MB | Task: MT | Format: CSV

### Europe

- [**Ladino-Spanish Lexical Resources**](https://mozilladatacollective.com/datasets/cmo1qb62200anmk07ls9feuh5?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Community | Licence: CC-BY-4.0 | Size: 39.92 KB | Task: MT | Format: TXT
- [**Synthetic Ladino Parallel Corpus**](https://mozilladatacollective.com/datasets/cmpbmhj4i0067nw07tk46v2jp?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: Community | Licence: CC-BY-4.0 | Size: 898.32 MB | Task: MT | Format: TSV
- [**Sentence translation difficulty in Spanish - BOUQuET**](https://mozilladatacollective.com/datasets/cmngbf1tt0050nn07i49aebnk?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) Contributor: MDC Curators | Licence: CC-BY-SA-4.0 | Size: 55.83 KB | Task: MT | Format: TSV

## Conclusion

A low-resource translation model in 2026 isn't built by waiting for web crawlers to find more text in your target language. It's built by combining the carefully curated parallel corpora that humanitarian organisations, regional publishers, and language communities have produced specifically for this purpose; then fine-tuning a strong multilingual base model on the mix. The 15 datasets above are the working set: not enough alone to train an MT model from scratch, but more than enough to push a multilingual base model into useful territory for the language pair you care about.

This is exactly the data ecosystem [Mozilla Data Collective](https://mozilladatacollective.com/?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post) was built to enable. Mozilla Data Collective's mission is to put communities at the centre by giving the organisations and contributors who built these corpora real control over how their data is licensed and used, rather than ceding that control to whichever model provider happened to scrape it first. For low-resource translation in particular, the communities who speak these languages are the ones who decided to translate the sentences, validate the alignments, and release the data. At Mozilla Data Collective we’re proud to empower these communities by making that work searchable, downloadable, and properly attributed.

[Browse all Mozilla Data Collective datasets →](https://mozilladatacollective.com/datasets?utm%5Fsource=community.mozilladatacollective.com&utm%5Fmedium=blog&utm%5Fcampaign=mt-post)

[Get in touch →](mailto:support@mozilladatacollective.com)

[Join Mozilla Data Collective → ](https://mozilladatacollective.com/auth/signup?ref=community.mozilladatacollective.com)