Phoenix at Palmx: Exploring Data Augmentation for Arabic Cultural Question Answering
Résumé
Large Language Models (LLMs) have become central to natural language processing, but their performance in low-resource cultural domains remains limited, mainly due to the dominance of English data in training.This limitation is especially evident in open small models.Evaluating and improving LLMs' performance in Arabic culture is therefore necessary.This paper presents Phoenix and PhoenixIs, two models fine-tuned for the Palmx 2025 general culture and Islamic culture subtasks.Phoenix uses the Palmx-GC and Palmx-IC datasets as seed data and applies diverse data augmentation strategies to construct an enriched fine-tuning dataset.Phoenix achieves an accuracy of 71.35% on the general culture subtask, while PhoenixIs reaches 83.82% on the Islamic culture subtask.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueAuteur(s)
Statistiques
Consultations : 1
Téléchargements : 0