Accès ouvert

Evaluating Robustness of LLMs to Typographical Noise in Yorùbá QA

Article scientifique 2025 Anglais

Résumé

Generative AI models are primarily accessed through chat interfaces, where user queries often contain typographical errors.While these models perform well in English, their robustness to noisy inputs in low-resource languages like Yorùbá remains underexplored.This work investigates a Yorùbá question-answering (QA) task by introducing synthetic typographical noise into clean inputs.We design a probabilistic noise injection strategy that simulates realistic human typos.In our experiments, each character in a clean sentence is independently altered, with noise levels ranging from 10% to 40%.We evaluate performance across three strong multilingual models using two complementary metrics: (1) a multilingual BERTScore to assess semantic similarity between outputs on clean and noisy inputs, and (2) an LLM-asjudge approach, where the best Yorùbá-capable model rates fluency, comprehension, and accuracy on a 1-5 scale.Results show that while English QA performance degrades gradually, Yorùbá QA suffers a sharper decline.At 40% noise, GPT-4o experiences over a 50% drop in comprehension ability, with similar declines for Gemini 2.0 Flash and Claude 3.7 Sonnet.We conclude with recommendations for noiseaware training and dedicated noisy Yorùbá benchmarks to enhance LLM robustness in lowresource settings.

Citer ce document

Okewunmi, P., James, F., Fajemila, O. (2025). Evaluating Robustness of LLMs to Typographical Noise in Yorùbá QA. https://doi.org/10.18653/v1/2025.africanlp-1.29

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0