Accès ouvert

Fine-tuning Whisper Tiny for Swahili ASR: Challenges and Recommendations for Low-Resource Speech Recognition

Article scientifique 2025 Anglais

Résumé

Automatic Speech Recognition (ASR) technologies have seen significant advancements, yet many widely spoken languages remain underrepresented.This paper explores the finetuning of OpenAI's Whisper Tiny model (39M parameters) for Swahili, a lingua franca for over 100 million people across East Africa.Using a dataset of 5,520 Swahili audio samples, we analyze the model's performance, error patterns, and limitations after fine-tuning.Our results demonstrate the potential of fine-tuning for improving transcription accuracy, while also highlighting persistent challenges such as phonetic misinterpretations, named entity recognition failures, and difficulties with morphologically complex words.We provide recommendations for improving Swahili ASR, including scaling to larger model variants, architectural adaptations for agglutinative languages, and data enhancement strategies.This work contributes to the growing body of research on adapting pre-trained multilingual ASR systems to low-resource languages, emphasizing the need for approaches that account for the unique linguistic features of Bantu languages.

Citer ce document

Sharma, A., Pandya, M., Shukla, A. (2025). Fine-tuning Whisper Tiny for Swahili ASR: Challenges and Recommendations for Low-Resource Speech Recognition. https://doi.org/10.18653/v1/2025.africanlp-1.11

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0