A Predictive Risk Framework for P2P Micro-Lending Default Prediction with Anomaly Detection
Résumé
Machine-learning credit assessment is often discussed as a substitute for manual underwriting, yet high-stakes lending decisions require a more careful design in which predictive models augment human judgement. This paper presents a supervised learning framework for peer-to-peer micro-lending default prediction using a reproducible pipeline built around LightGBM classification, engineered affordability features, class-imbalance handling, Isolation Forest anomaly detection, and SHAP explainability. The dataset is processed through exploratory analysis, feature construction, stratified train-test splitting, median and modal imputation, robust scaling, categorical encoding, SMOTE rebalancing, model training, diagnostic evaluation, business-impact estimation, and global and local explanation. The experimental run reports an AUC-ROC of 0.7459, average precision of 0.3042, macro-F1 of 0.5614, and default-class F1 of 0.3332. At the operating threshold used in the pipeline, the model identifies 4,081 defaults while missing 1,850 defaults, representing $236,020,901.20 in potential principal at risk. SHAP analysis ranks age, interest rate, months employed, credit-to-income, dependents, and co-signer status among the strongest drivers, revealing both useful affordability signals and governance-sensitive proxy variables. The findings support an augmentation thesis: AI can improve default-risk triage, anomaly surfacing, and explanation consistency, but false positives, missed defaults, proxy-feature risks, and highvalue edge cases require human-in-the-loop underwriting, fairness review, and accountable lending policy.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueStatistiques
Consultations : 1
Téléchargements : 0