Accès ouvert

MBZUAI at AMIYA Shared Task 2026: Adapting Open-Source LLMs for Dialectal Arabic

Article scientifique 2026 Autre

Résumé

This paper presents our contribution to the closed data track of the AMIYA Shared Task on Dialectal Arabic text generation.In this track, we train fully open-source Large Language Models (LLMs) on five Arabic dialects: Egyptian, Moroccan, Palestinian, Saudi, and Syrian, using the provided training datasets.We experiment with different base and instruct models using several pretraining and instruction tuning approaches.In total, five models were submitted, with three variants per dialect.Our best-performing models for the five dialects are ALLaM for Egyptian, LLaMa for Moroccan, and Palestinian, and Aya for Saudi and Syrian.

Citer ce document

Gaber, R., Allam, Y., Amin, S., Aly, R., Alhafni, B. (2026). MBZUAI at AMIYA Shared Task 2026: Adapting Open-Source LLMs for Dialectal Arabic. https://doi.org/10.18653/v1/2026.vardial-1.31

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0