Accès ouvert

Learning Word Embeddings from Glosses: A Multi-Loss Framework for Arabic Reverse Dictionary Tasks

Article scientifique 2025 Autre

Résumé

We address the task of reverse dictionary modeling in Arabic, where the goal is to retrieve a target word given its definition.The task comprises two subtasks: (1) generating embeddings for Arabic words based on Arabic glosses, and (2) a cross-lingual setting where the gloss is in English and the target embedding is for the corresponding Arabic word.Prior approaches have largely relied on BERT models such as CAMeLBERT or MARBERT trained with mean squared error loss.In contrast, we propose a novel ensemble architecture that combines MARBERTv2 with the encoder of AraBART, and we demonstrate that the choice of loss function has a significant impact on performance.We apply contrastive loss to improve representational alignment, and introduce structural and center losses to better capture the semantic distribution of the dataset.This multi-loss framework enhances the quality of the learned embeddings and leads to consistent improvements in both monolingual and cross-lingual settings.Our system achieved the best rank metric in both subtasks compared to the previous approaches.These results highlight the effectiveness of combining architectural diversity with task-specific loss functions in representational tasks for morphologically rich languages like Arabic.

Citer ce document

Ibrahim, E., Adel, F., Torki, M., El-Makky, N. (2025). Learning Word Embeddings from Glosses: A Multi-Loss Framework for Arabic Reverse Dictionary Tasks. https://doi.org/10.18653/v1/2025.arabicnlp-main.31

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0