VISH-GUARD: a multi-agent and LLM-powered framework for multilingual voice phishing detection
Résumé
Abstract Voice phishing (vishing) has emerged as a major cybersecurity threat, leveraging persuasive speech and psychological manipulation to deceive victims in real time. Existing detection approaches remain limited by monolingual assumptions, unimodal analysis, and poor interpretability, reducing their effectiveness against adaptive and multilingual attacks. This paper presents VISH-GUARD, a multi-agent vishing detection framework that combines multimodal analysis with large language model (LLM)-based reasoning to achieve robust, multilingual, and explainable threat detection. The system coordinates specialized agents responsible for acoustic feature analysis, semantic intent recognition, emotion detection, and behavioral profiling. Their outputs are integrated into a unified risk assessment, while an LLM-based reasoning layer generates structured explanations to support human analysts. To enable systematic evaluation, we introduce a new annotated dataset of synthetic yet realistic vishing calls in English, French, and Arabic, incorporating background noise, emotional cues, and diverse persuasion strategies. Experimental results show that VISH-GUARD achieves an F1-score of 0.96, outperforming unimodal agents, multimodal deep learning baselines, and an end-to-end multimodal transformer baseline, while reducing false negatives in scenarios involving emotional manipulation and behavioral irregularities. Overall, VISH-GUARD demonstrates how multi-agent collaboration combined with LLM-driven reasoning can provide an effective and transparent defense against modern voice phishing attacks.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueAuteur(s)
Statistiques
Consultations : 1
Téléchargements : 0