MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors
Résumé
We present MSA-MATHEVAL, our submission to the BEA 2025 Shared Task on evaluating AI tutor responses across four instructional dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability.Our approach uses a unified training pipeline to fine-tune a single instructiontuned language model across all tracks, without any task-specific architecture modifications.To improve prediction reliability, we introduce a disagreement-aware ensemble inference strategy that enhances coverage of minority labels.Our system achieves strong performance across all tracks, ranking 1 st in Providing Guidance, 3 rd in Actionability, and 4 th in both Mistake Identification and Mistake Location.These results demonstrate the effectiveness of scalable instruction tuning and disagreementdriven modeling for robust, multi-dimensional evaluation of LLMs as educational tutors.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueAuteur(s)
Statistiques
Consultations : 1
Téléchargements : 0