An XGBoost-based predictive framework for diabetes mellitus multi-classification
Résumé
Diabetes mellitus is a persistent metabolic condition that requires accurate and early diagnosis to prevent severe complications. This paper proposes an Extreme Gradient Boosting (XGBoost)-based predictive framework for multi-classification of diabetes mellitus into non-diabetic, pre-diabetic, and diabetic classes. After standardization and exclusion of non-clinical identifiers, Duplicate clinical records were removed from the original dataset, leaving 826 unique records. Comprehensive preprocessing pipeline used a stratified 70:30 train-test split and five-fold cross-validation; scaling and resampling were performed only within training partitions. Experimental results on the original dataset XGBoost achieved an accuracy of 99.60%. Both Random Over Sampling (ROS) and Syntenic minority over sampling technique (SMOTE) have also achieved 99.60% accuracy but provided improved generalization at the expense of higher computational cost. In contrast, Random Under Sampling (RUS) and Cluster Centroids (CC) reduced accuracy to 92.74% and 90.32%, respectively due to information loss. Across five folds, the original XGBoost model achieved 98.79 ± 1.13% accuracy. Benchmarking against Logistic Regression, Random Forest, Support Vector Machine, Decision Tree, and K-Nearest Neighbors showed that XGBoost provided the strongest performance. These findings highlight the effectiveness of XGBoost while emphasizing classification accuracy in multiclass diabetes prediction systems.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueAuteur(s)
Statistiques
Consultations : 1
Téléchargements : 0