EmotionNet: A Novel Hybrid Deep Learning Model for Arabic Speech Emotion Recognition
Résumé
This study presents EmotionNet, a novel hybrid deep learning model designed for Arabic speech emotion recognition. EmotionNet integrates a Variational Auto-Encoder (VAE) for latent representation learning with a lightweight classification branch enhanced by latent-space refinement. Evaluated on the KEDAS dataset, which includes five emotionally acted categories, the model achieved a test accuracy of 93.99% and outperformed conventional classifiers such as SVM, MLP, and Random Forest. The proposed approach employs a compound loss function and KL annealing to jointly optimize reconstruction and classification. Although the results are promising, the acted nature of KEDAS may overstate real-world performance, highlighting the need for evaluation on spontaneous, multimodal datasets, an effort currently underway in an ongoing interdisciplinary project.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueAuteur(s)
Statistiques
Consultations : 1
Téléchargements : 0