LMSA at AraGenEval Shared Task: Ensemble-Based Detection of AI-Generated Arabic Text Using Multilingual and Arabic-Specific Models
Résumé
We address the problem of distinguishing between human-authored and AI-generated text in low-resource languages, particularly Arabic.We present the LMSA 1 team's participation in the ARATECT (Arabic AI-Generated Text Detection) subtask of the AraGenEval 2 shared task, which targets the detection of AI-generated Arabic texts.We propose an ensemble-based classification framework that integrates multilingual and Arabic-specific pre-trained language models, namely Fanar, AraBERT, and XLM-R, optimized through a dedicated fine-tuning pipeline.The approach is evaluated on the balanced Arabic text dataset provided by the shared task organizers.Our system achieved an F1-score of 0.864 and ranked first among all participating teams.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueStatistiques
Consultations : 1
Téléchargements : 0