ANLP-UniSo at MAHED Shared Task: Detection of Hate and Hope Speech in Arabic Social Media based on XLM-RoBERTa and Logistic Regression
Résumé
In this paper, we present our system for Subtask 1 of the MAHED 2025 shared task, which involves classifying Arabic text into three categories: Hate, Hope, and not_applicable.Our methodology integrates XLM-RoBERTa embeddings with supervised ML and deep learning techniques.After applying Arabicspecific preprocessing, we extract contextual embeddings and mitigate class imbalance using SMOTE .We then train LR and LSTM classifiers on the augmented features space, supplemented by a similarity calculation with Zero-Shot for prediction validation.The system was evaluated in two phases: using the initial validation set, and the official updated datasets.Results show competitive performance, particularly in boosting recall for minority classes with a macro score of 0.60.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueStatistiques
Consultations : 1
Téléchargements : 0