Building a Named Entity Recognition model for Ethiopian Languages: a comparative analysis of composite feature embedding
Résumé
Abstract Named Entity Recognition (NER) has become a critical and essential step in information extraction, machine translation, and question-and-answering systems in a different language. One of the most important factors which directly and significantly affects the quality of the NER is the selection and encoding of the input features to generate rich semantic and grammatical representation vectors. However, the existing NER models are insufficient to handle new and unseen entity types from the growing Amharic digital data, and the development of more effective and accurate NER models is being widely researched. In this regard, herein, we propose a deep learning NER model that effectively represents word tokens through the design of a combinatorial feature embedding; and performed a comparative analysis with the existing models for Ethiopian languages. The word vectors built for all tokens using an unsupervised learning algorithm was merged with a set of specifically developed language-independent features and together fed to the neural network model to predict the classes of the words. Empirical results over the Ethiopian language dataset show that the use of character-level word embeddings in conjunction with other features in BiLSTM-CRF models leads to comparable state-of-the-art performance. Besides just showing the ability of our model to generalize to different languages, we evaluated the model and obtained state-of- the-art performances: 92.88%, and 82.35% of accuracy on AM_NER and Oro_ NER datasets, respectively.
Citer ce document
Accès au document
Voir sur le dépôt sourceCe document est hébergé sur son dépôt institutionnel d'origine.
Auteur(s)
Statistiques
Consultations : 1
Téléchargements : 0