Accès ouvert

ArGAN: Arabic Gender, Ability, and Nationality Dataset for Evaluating Biases in Large Language Models

Article scientifique 2025 Anglais

Résumé

Large language models (LLMs) are pretrained on substantial, unfiltered corpora, assembled from a variety of sources.This risks inheriting the deep-rooted biases that exist within them, both implicit and explicit.This is even more apparent in low-resource languages, where corpora may be prioritized by quantity over quality, potentially leading to more unchecked biases, particularly in low-resource languages, where all available data is leveraged solely to expand volume due to inherent scarcity.More specifically, we address the biases present in the Arabic language in both general-purpose and Arabic-specialized architectures in three dimensions of demographics: gender, ability, and nationality.We introduce ArGAN, a dataset for evaluating the fairness of these models across three demographic axes: gender, ability and nationality.Where we experiment with bias-revealing, template-based prompts and measure performance and bias using existing and evaluation metrics, and propose adaptations to others.

Citer ce document

Aly, R., Allam, Y., Gaber, R., Basta, C. (2025). ArGAN: Arabic Gender, Ability, and Nationality Dataset for Evaluating Biases in Large Language Models. https://doi.org/10.18653/v1/2025.gebnlp-1.23

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0