Accès ouvert

Codezone Research Group at ImageEval Shared-Task 2: Arabic Image Captioning Using BLIP and M2M100: A Two-Stage Translation Approach for ImageEval 2025

Article scientifique 2025 Autre

Résumé

This paper details the ImageEval 2025 Shared Task on Arabic image captioning.We designed a two-step, zero-shot framework that utilises the BLIP multimodal vision-language model to first generate English captions.These captions are then converted to Arabic via the M2M100 multilingual translation model.We tested the full pipeline on the official ImageEval 2025 benchmarking set, obtaining a cosine similarity of 0.383 and an LLM Judge score of 15.14.The corroborating numerical and qualitative findings confirm the viability of a translation-driven methodology for cross-lingual image captioning in Arabic, a language often classified as low-resource.Nonetheless, the experiments also uncovered weaknesses: subtle semantic layers and culturally specific references are inadequately conveyed in the output and merit focused attention in subsequent iterations.

Citer ce document

Bichi, A. (2025). Codezone Research Group at ImageEval Shared-Task 2: Arabic Image Captioning Using BLIP and M2M100: A Two-Stage Translation Approach for ImageEval 2025. https://doi.org/10.18653/v1/2025.arabicnlp-sharedtasks.53

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0