Accès ouvert

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study

Article scientifique 2026 Autre

Résumé

Ancient and endangered languages pose a unique challenge for NLP: their datasets are inherently scarce, difficult to expand, and built from formulaic corpora-making data-quality issues especially consequential yet rarely audited.Motivated by the need to understand what current NMT can realistically achieve for such languages, we investigate hieroglyphicto-German translation, where a recent study reported 61.5 BLEU using fine-tuned M2M-100.Our reproduction yields only 37.0 BLEU with the released model.Investigating this gap, we find 32% of test targets appear identically in training (16/50; 50% under 8gram overlap at 70% threshold).This contamination inflates scores dramatically: contaminated samples achieve up to 83.8 BLEU / 0.924 COMET-22 versus 30.9-39.2BLEU / 0.622-0.676COMET-22 on clean samples across five model configurations spanning two architectures.Document-level decontamination reduces contaminated BLEU by only 4.6 points because 8/16 targets persist via other source documents-target-level deduplication is required.We release a decontaminated 34sample test set and establish corrected baselines (30.9-39.2BLEU), providing a realistic assessment of NMT capability for this endangered writing system.

Citer ce document

Toutou, A., Harb, A., Basta, C. (2026). Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study. https://doi.org/10.18653/v1/2026.nlp4dh-1.6

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0