Accès ouvert

LMSA at AraGenEval Shared Task: Ensemble-Based Detection of AI-Generated Arabic Text Using Multilingual and Arabic-Specific Models

Article scientifique 2025 Autre

Résumé

We address the problem of distinguishing between human-authored and AI-generated text in low-resource languages, particularly Arabic.We present the LMSA 1 team's participation in the ARATECT (Arabic AI-Generated Text Detection) subtask of the AraGenEval 2 shared task, which targets the detection of AI-generated Arabic texts.We propose an ensemble-based classification framework that integrates multilingual and Arabic-specific pre-trained language models, namely Fanar, AraBERT, and XLM-R, optimized through a dedicated fine-tuning pipeline.The approach is evaluated on the balanced Arabic text dataset provided by the shared task organizers.Our system achieved an F1-score of 0.864 and ranked first among all participating teams.

Citer ce document

Zita, K., Nehar, A., Khelil, A., Bellaouar, S., Cherroun, H. (2025). LMSA at AraGenEval Shared Task: Ensemble-Based Detection of AI-Generated Arabic Text Using Multilingual and Arabic-Specific Models. https://doi.org/10.18653/v1/2025.arabicnlp-sharedtasks.4

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0