Accès ouvert

AraMinds at MAHED 2025: Leveraging Vision-Language Models and Contrastive Multi-task Learning for Multimodal Hate Speech Detection

Article scientifique 2025 Autre

Résumé

Detecting hate speech in social media content is essential to provide a safe space for people to connect.Memes have been used lately to sarcastically express one's opinion, and they can be used to hide harmful intentions and spread hateful speech.In this work, we build our system that detects hateful speech in memes by combining visual and textual features and merging them using different techniques to detect the inherent meaning and overcome the challenge of vast dialectal differences and the variety of topics discussed.To improve our system's robustness, we combine different techniques, such as multi-tasking, contrastive learning, and vision language modeling in a final ensemble model that secured us the third place in the MAHED 2025 shared-task leaderboard with a macro-f1 score of 0.74, showing strong performance on the evaluation set.

Citer ce document

Zaytoon, M., Salem, A., Sakr, A., Elkordi, H. (2025). AraMinds at MAHED 2025: Leveraging Vision-Language Models and Contrastive Multi-task Learning for Multimodal Hate Speech Detection. https://doi.org/10.18653/v1/2025.arabicnlp-sharedtasks.85

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0