Towards a fair and transparent early warning system for identifying students at risk of academic failure
Résumé
Early warning systems (EWS) help schools identify students at risk of academic failure and intervene before the school year ends. Machine learning models for this task are commonly evaluated using accuracy and AUC-ROC, but both can be misleading when failure is the minority outcome and student characteristics are unevenly distributed; fairness and interpretability are usually treated as separate concerns, if at all. We address all three together. We frame deployment as a constrained problem: maximize precision subject to a regulatory recall threshold of 0.85 on Fail. Using first-semester data on 451,852 middle school students to predict end-of-year Pass/Fail, we trained nine models, applied isotonic calibration, and tuned per-model thresholds. We then verified the recall constraint under sampling uncertainty with a stratified bootstrap 95% confidence interval. Four models met it: Multilayer Perceptron (MLP), Logistic Regression, Random Forest, and an SGD Classifier. A Borda count over Recall, Precision, AUC-PR, and Matthews Correlation Coefficient (MCC) selected the MLP, which correctly flags 86.7% of students who will fail (Precision 0.751, AUC-PR 0.888, MCC 0.730), five months before the academic year ends. To make predictions usable by teachers and school leaders, we add explanations at three levels: SAGE for global feature importance, Accumulated Local Effects (ALE) for marginal effects, and Diverse Counterfactual Explanations (DiCE) for the changes a student would need to flip Fail to Pass. A fairness audit across Gender×Location intersections shows subgroup gaps in true-positive rates; Equal Opportunity post-processing closes most of these gaps and raises recall to 0.910 with a measurable but acceptable precision–recall trade-off given the policy mandate (precision 0.751 → 0.702, false-positive rate 0.106 → 0.142), while Equalized Odds reduces overall performance. The result is an EWS that meets the recall mandate, can justify its predictions to non-technical users, and exposes inequities that aggregate metrics would hide. However, the model was developed and evaluated on a single cohort from one Moroccan region, and future work should assess temporal and external generalization across cohorts, academic years, and regions.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueStatistiques
Consultations : 1
Téléchargements : 0