USTHB at NADI 2023 shared task: Exploring Preprocessing and Feature Engineering Strategies for Arabic Dialect Identification
Résumé
In this research paper, we undertake a comprehensive examination of several pivotal factors that impact the performance of Arabic Disinformation Detection in the ArAIEval'2023 shared task.Our exploration encompasses the influence of surface preprocessing, morphological preprocessing, the FastText vector model, and the weighted fusion of TF-IDF features.To carry out classification tasks, we employ the Linear Support Vector Classification (LSVC) model.In the evaluation phase, our system showcases significant results, achieving an F 1 micro score of 76.70% and 50.46% for binary and multiclass classification scenarios, respectively.These accomplishments closely correspond to the average F 1 micro scores achieved by other systems submitted for the second subtask, standing at 77.96% and 64.85% for binary and multiclass classification scenarios, respectively.
Citer ce document
Accès au document
Ce lien n'est plus accessible actuellement. Contactez l'institution d'origine.
Statistiques
Consultations : 1
Téléchargements : 0