Accès ouvert

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

Article scientifique 2026 Autre

Résumé

Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyond matched evaluation. Results revealed clear interactions between architecture and language, with W2V-BERT outperforming Whisper on Nguni languages by 3 to 4 WER points whilst Whisper performed better on Sotho-Tswana languages. W2V-BERT with tone conditioning reached 28.41% average WER across datasets and 23.79% on Xitsonga transfer. No single model suited all 6 languages, so deployment should pair model selection per language with validation across corpora.

Citer ce document

Mokgosi, K., Marivate, V., Mundia, S., Netshifhefhe, U., Mogale, T., Sindane, T. (2026). Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition. https://doi.org/10.48550/arxiv.2606.31642

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0