Evaluation of Supervised Learning Models for Automatic Spam Email Detection
Résumé
Abstract This paper compares the performance of different supervised learning algorithms for email spam detection. The comparison considered performance measures such as the area under the curve (AUC), F-score, precision, and confusion matrix. The paper evaluated the performance of eight supervised learning algorithms for email spam detection. The first stage collected the dataset of emails from the Kaggle repository. The second stage involves pre-processing, duplicate removal, and the dataset features scaling. After the pre-processing stage, the study employed a synthetic minority technique for balancing samples representing the spam and no spam emails in the dataset. In the final stage, the supervised learning algorithms are trained on the pre-processed dataset and then the test result is analyzed. The comparison shows the random forest (RF) model performing at higher accuracy than the other models. The RG model achieved 96.6% accuracy in email spam detection. Thus, the result demonstrates that different models tend to perform differently in email spam detection.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueAuteur(s)
Statistiques
Consultations : 1
Téléchargements : 0