Accès ouvert

A Tunisian benchmark social media data set for COVID-19 sentiment analysis and sarcasm detection

Article scientifique 2022 Anglais

Résumé

Abstract The aim of this research is to better understand public perceptions of COVID-19 pandemic patterns and to identify key themes of concern expressed by Tunisian dialect social media users throughout the epidemic. We collected around 23K comments written in Tunisian dialect in both Arabic and Latin letters. These comments were manually annotated by native language experts for sentiment analysis (optimist, pessimist and neutral) and sarcasm detection (sarcastic and non-sarcastic). In addition to health, our data set includes comments relating to additional COVID-19-influenced thematic areas, such as entertainment, social, sports, religion and politics. This paper deals with an extensive analysis of the sentiments and sarcasm expressed in Tunisian social media comments about the novel COVID-19 since its release at the beginning of 2020. On the data set, we also report benchmarking results for sentiment analysis and sarcasm detection using machine learning and deep learning techniques. The best models achieved an accuracy of above 70% on both sentiment analysis and sarcasm detection.

Citer ce document

Mekki, A., Zribi, I., Ellouze, M., Belguith, L. (2022). A Tunisian benchmark social media data set for COVID-19 sentiment analysis and sarcasm detection. https://doi.org/10.21203/rs.3.rs-2321298/v1

Accès au document

Voir sur le dépôt source

Ce document est hébergé sur son dépôt institutionnel d'origine.

Statistiques

Consultations : 1

Téléchargements : 0