Accès ouvert

Isnad AI at IslamicEval 2025: A Rule-Based System for Identifying Religious Texts in LLM Outputs

Article scientifique 2025 Autre

Résumé

This paper presents the Isnad AI system developed for the IslamicEval 2025 Shared Task 1A, which focuses on identifying characterlevel spans of Quranic verses (Ayahs) and Prophetic sayings (Hadiths) within Large Language Model (LLM) outputs.This task is formulated as a token classification problem using a fine-tuned AraBERTv2 model.The primary contribution is a novel rule-based data preprocessing and augmentation pipeline, through which a large-scale, high-quality training corpus is systematically generated from raw religious texts.Through comprehensive ablation studies, it is demonstrated that the controlled synthetic data generation approach significantly outperforms traditional database lookup methods and basic fine-tuning approaches.The system achieved an F1 score of 66.97% in the official test set, demonstrating the effectiveness of principled synthetic data generation for specialized religious text verification tasks.To support reproducibility and future research in Islamic citation detection, all code, generated datasets, and experimental resources are made publicly available on GitHub and Hugging Face.

Citer ce document

Elden, F. (2025). Isnad AI at IslamicEval 2025: A Rule-Based System for Identifying Religious Texts in LLM Outputs. https://doi.org/10.18653/v1/2025.arabicnlp-sharedtasks.74

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0