NOW LIVE: Lost in Transcription Competition
Lost In Transcription: Advancing Speech Recognition for Underserved Linguistic Contexts
This competition evaluates performance on real-world bilingual dialogues in Indonesian-Javanese, Nahuatl-Spanish, and Spanish-English.
For speech technologies to truly have an impact on the world, they must adapt to how people actually communicate in digitally-mediated settings. For most of the world, multilingualism is the norm and language boundaries are fluid. Monolingual Automatic Speech Recognition (ASR) systems for high-resource languages - those with abundant training data and mature digital tools like English, German, or Spanish - have reached near-human performance on some benchmark datasets. Still, most of the world’s languages, and even real-world usage of majority languages, remain underrepresented in digital tools. Many existing speech datasets, both training and benchmark data, are monolingual, consist of primarily read rather than spontaneous speech, or fail to capture the domains of use that increasingly define everyday digital communication.
Multilingual speech is also rarely just two monolingual streams stitched together. Code-switching (or code-mixing) — moving between two or more languages or varieties within a single conversation, sentence, or phrase — is one of the most common forms of everyday bilingual speech, and it can carry meaning, nuance, and identity that neither language conveys on its own. Because most ASR data assumes a single target language, models trained on monolingual corpora have no consistent way to represent the vocabulary, phonetics, or grammar of a language pair as it’s actually spoken.
Task
Your task is to build ASR systems that accurately transcribe natural, code-switched speech across three underserved language pairs spanning North and Central America and Southeast Asia:
- North American Spanish-English
- Spanish-Nahuatl
- Indonesian-Javanese
|
Ah bueno, entonces ompa timotaj tejua, xamo tikneki tinechtaneutis cien pesos para ika nikoutikisas, no se piyo, lo que pasa tech nin semana fiero nimotak uan nochi niktamik in tomin den niknechikolteua, ajko nechajsi para ijkuak. Ah ok, so I’ll see you there, by any chance can you lend me 100 pesos so I can buy a chicken as well, just that this week it’s not looking good and I’ve spent all the money that I had saved up for it, I won’t have enough for then. |
|
Pues sí iría si iría but I have to know some of the details like la fecha y el horario you know because like I have to work. Well yes I would go I would you but I have to know some of the details like the date and the time you know because like I have to work. |
These datasets, along with the held-out test set, consist of real voice-note dialogues between pairs of bilingual speakers, featuring heavy code-switching and lexical borrowing — the kind of natural, multilingual speech that dominates real-world communication but remains largely absent from existing benchmarks.
Participants will build ASR systems robust enough to handle the fluidity of real, multilingual speech as it actually occurs in digitally mediated dialogue. Successful models could help transform vast, inaccessible audio-visual collections, such as those held by galleries, archives, libraries, and museums, into searchable, discoverable, and accessible digital assets, and advance speech technology for communities underserved by it.
The languages in this challenge
The three pairs span very different linguistic situations, but share a common thread: a widely spoken majority language sits alongside a language that is under-represented in speech technology, and in real life, speakers routinely blend the two.
North American Spanish-English
Tens of millions of bilingual speakers across the United States move fluidly between Spanish and English, often within the same sentence, in a pattern commonly referred to as “Spanglish”. Both languages are individually high-resource, but a model trained separately on monolingual English or Spanish corpora has no way to represent a language boundary that shifts mid-utterance, making this a common yet technically demanding form of natural speech in North America.
Spanish-Nahuatl
Nahuatl, the most widely spoken Indigenous language in Mexico with over 1.5 million speakers, belongs to the Uto-Aztecan family. It is polysynthetic and agglutinative, meaning words are built from many morphemes, the smallest units of meaning in a language, so a single Nahuatl word can carry the meaning of an entire English sentence. This structure, combined with numerous diverse regional varieties (some not mutually intelligible) and centuries of contact-driven borrowing from Spanish, makes Spanish-Nahuatl speech one of the most linguistically complex pairs in this challenge to model with standard ASR architectures.
Indonesian-Javanese
Javanese has tens of millions of native speakers — one of the largest languages in the world by that measure — yet remains under-resourced in digital and speech corpora relative to its size. Javanese and Indonesian are unusually close to begin with: Indonesian absorbed much of its Sanskrit-derived and courtly vocabulary via Old Javanese, and Javanese loanwords are embedded throughout everyday Indonesian, so the lexical boundary between the two is porous rather than clean. As a result, a given word can’t always be cleanly assigned to one language or the other, which complicates the basic assumption behind word-level language tagging.
Prizes
| Place | North American Spanish-English | Spanish-Nahuatl | Indonesian-Javanese |
|---|---|---|---|
| 1st | $4,000 | $4,000 | $4,000 |
| 2nd | $2,000 | $2,000 | $2,000 |
| Bonus (best avg. score across all pairs) | $2,000 | ||
| Total | $20,000 | ||
Bonus Prize
After the competition ends, the organizers will average scores for those who submitted to all three language-pair tracks. A bonus prize will be given to the team that obtains the lowest average WER across tracks. To be eligible, teams must submit to all three tracks with the same team members, but may submit different models.
How to compete
Submissions in this challenge will open at a later date. Join the challenge now to start working with the data and get notified when submissions open.
- Click the "Compete!" button in the sidebar to enroll in the competition.
- Get familiar with the problem through the Problem description. Additional resources are available on the About page.
- Explore the highlighted training datasets on the MDC platform, linked on the Problem description page, or browse the full MDC catalog.
- Once submissions are open, package your model files with the code to make predictions according to the runtime repository specification, which will be shared at that time. Then, submit your code as a zip archive for containerized execution from the Submissions page. You’re in!
Competition rules
The competition rules are designed to promote fair competition and encourage useful, reproducible solutions. If you are ever unsure whether your solution complies with the rules, contact the organizers by emailing info@drivendata.org.
Note on external data and models
External data and pre-trained models are allowed in this competition.
Participants may use any dataset available on MDC for training, including the three datasets highlighted in this competition. Participants may also use external data not currently hosted on MDC, provided that dataset is published to MDC at the conclusion of the competition. For more information on publishing data to MDC, see MDC’s terms and conditions for data providers, and the Contribute Datasets page for instructions on uploading datasets.
Open-weight, pre-trained models are also permitted in this challenge and may be run locally. Participants may not send competition data to third-party hosted models, model APIs, or inference services (e.g., OpenAI, Anthropic, or similar). See the competition rules for full details on permitted data and model use.
If you have questions about data or model use, contact the organizers by emailing info@drivendata.org.
Sponsors
This challenge is sponsored by the Mozilla Data Collective with support from the Mozilla Foundation.