Accès ouvert

MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors

Article scientifique 2025 Anglais

Résumé

We present MSA-MATHEVAL, our submission to the BEA 2025 Shared Task on evaluating AI tutor responses across four instructional dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability.Our approach uses a unified training pipeline to fine-tune a single instructiontuned language model across all tracks, without any task-specific architecture modifications.To improve prediction reliability, we introduce a disagreement-aware ensemble inference strategy that enhances coverage of minority labels.Our system achieves strong performance across all tracks, ranking 1 st in Providing Guidance, 3 rd in Actionability, and 4 th in both Mistake Identification and Mistake Location.These results demonstrate the effectiveness of scalable instruction tuning and disagreementdriven modeling for robust, multi-dimensional evaluation of LLMs as educational tutors.

Citer ce document

Hikal, B., Basem, M., Oshallah, I., Hamdi, A. (2025). MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors. https://doi.org/10.18653/v1/2025.bea-1.95

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0