A leakage aware machine learning approach to classify the technological proficiency of Moroccan university students in artificial intelligence
Résumé
Abstract This study investigates AI-related technological proficiency among undergraduate students at the University of Casablanca and identifies, using a machine learning-based classification framework, the questionnaire items most strongly associated with this construct. Data were collected from 600 students across scientific and humanities disciplines using a validated 30-item questionnaire covering AI applications, AI-related skills, and improvement strategies. In addition to descriptive analysis, a classification approach was adopted to estimate construct-level proficiency categories from a reduced subset of questionnaire items to capture complex relationships between variables. To address a critical methodological challenge in questionnaire-based modeling, namely circularity and target leakage, the target variable was defined at the construct level, ensuring a clear separation from individual predictive items. Feature selection was performed exclusively on the training set, resulting in a subset of 17 informative features. Two classification models were developed and evaluated: SVC and GNB, using a 75/25 train-test split and repeated stratified tenfold cross-validation. In terms of construct-recovery accuracy, the degree to which the reduced-item model reproduces the construct-level proficiency category derived from the full 30-item instrument, the SVC reached 96.7% (AUROC = 0.9666), while GNB reached 94.7% (AUROC = 0.9560). Baseline models, including Logistic Regression (85%) and Decision Tree (88%), were also implemented to contextualize performance. The results confirm that the proposed models capture non-trivial patterns within the reduced item set and show stable performance across cross-validation folds (mean accuracy = 96.6%, SD = 0.9%), indicating that the selected model specification behaves consistently within this dataset rather than providing evidence of external predictive generalization. The findings indicate that a reduced subset of questionnaire items can closely approximate the construct-level proficiency category derived from the full 30-item instrument. Beyond empirical results, this approach contributes a reproducible, leakage-aware methodology for construct-level classification that can inform rather than independently drive measurement practices in higher education.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueStatistiques
Consultations : 1
Téléchargements : 0