Offensive Language Detection in Arabizi
Résumé
Detecting offensive language in underresourced languages presents a significant real-world challenge for social media platforms.This paper is the first work focused on the issue of offensive language detection in Arabizi, an under-explored topic in an under-resourced form of Arabic.For the first time, a comprehensive and critical overview of the existing work on the topic is presented.In addition, we carry out experiments using different BERT-like models and show the feasibility of detecting offensive language in Arabizi with high accuracy.Throughout a thorough analysis of results, we emphasize the complexities introduced by dialect variations and out-ofdomain generalization.We use in our experiments a dataset that we have constructed by leveraging existing, albeit limited, resources.To facilitate further research, we make this dataset publicly accessible to the research community.
Citer ce document
Accès au document
Voir sur le dépôt sourceCe document est hébergé sur son dépôt institutionnel d'origine.
Auteur(s)
Statistiques
Consultations : 1
Téléchargements : 0