Text mining based on analogical proportions: Application to Arabic text summarization and classification
Résumé
A huge number of textual documents regularly enrich the World Wide Web. This big amount of data, produced by various sources such as news websites, is among the main sources of documents on the Internet and Web 2.0. However, the extraction of useful information from these unstructured data becomes a hard task requiring a complex step of data processing and management. Given the huge amount of textual data in the Web, there is a serious need for text mining applications such as text summarization and classification. This thesis studies the ability of Analogical proportions as a promising tool for text summarization and classification. Analogical proportions (AP) are statements of the form “a is to b as c is to d” usually denoted a : b :: c : d and expressing that “a differs from b as c differs from d”, as well as “b differs from a as d differs from c”. Analogical inference is based on the assumption that if four items a, b, c, d are making a valid analogical proportion and if d is unknown, this enables us to predict the value of d on the basis of the values of the triplet (a, b, c). Based on the above principle, analogical proportions have recently proven their efficiency to classify “structured” datasets. Our main interest in this work is to validate their efficiency to deal with “unstructured” datasets. We first propose in this thesis a new free structured Arabic dataset in a standard XML TREC format, named ANT corpus (https://antcorpus.github.io/). This latter is collected using RSS feeds from various news websites. It is used for many text-mining tasks, such as Arabic text summarization and classification. The ANT collection contains documents from five various news websites that are classified into six categories. This resource is available online in different versions. We exploit the ANT corpus v2.1 to treat two different problems: i) extractive text summarization and ii) text classification. The goal of automatic text summarization is to represent the main text by an extracted set of sentences. Various types of summarizers that are either extractive, abstractive or hybrid algorithms are available in the literature, such as TextRank, LexRank, LSA and Luhn. In this current work, we focus on unsupervised single- document extractive summarization. For this purpose, we first investigate the ability of analogical proportions to represent the relationship between documents and their corresponding summaries. Then, we suggest two various procedures to quantify the relevance/irrelevance of a given term (keyword), issued from the original text, to build its summary. In the first model, the analogical proportion, which represents this relationship, considers only the existence/non- existence of the keyword in the main text or summary in a binary way without any interest in the keyword frequency of the input text. In the second model, the keyword frequency is quantified in the analogical proportion. Then, we develop two algorithms respectively called AATS1 and AATS2 that implement respectively the first and second models. To assess the efficiency of the suggested algorithms, we exploit two datasets: ANT corpus v2.1 and EASC (https://sourceforge.net/projects/easc-corpus/). In both collections, the input document and its reference summary are available. To evaluate the proposed summarizers, we compute the ROUGE and BLEU metrics that compare the extracted summaries, from each category to the reference summaries. The achieved results show the interest of analogical proportions for text summarization if compared to language-independent summarizers or Arabic summarizers, using large (ANT corpus) and small (EASC) corpora. The proposed summarizers outperform the language-independent summarizers in terms of BLEU-1 and ROUGE-1 scores for the "Education" category from the EASC dataset. On the other hand, our AATS algorithms perform better than the three Arabic single document extractive summarizers (the score-based summarizer (2021), ASDKGA (2018) and LCEAS (2015)), tested on the EASC dataset, in terms of ROUGE-1 score. In particular, the AATS algorithms significantly outperform three among four language- independent summarizers (LexRank, TextRank, Luhn and LSA) in case of BLEU-1 for the ANT corpus, and they are not significantly outperformed by any other summarizer in the case of the EASC corpus. We also use the ANT corpus v2.1 for automatic text classification. The process of classifying or categorizing texts consists to assign a predefined class label to an unlabeled text based on its content. The classification model efficiency is sensitive to the high dimensionality of data that can affect the classifier accuracy. Many fields of application can benefit from text classification achievements, in particular recommendation systems, search engines and NLP tools. Great research effort has been devoted to studying the classification problems of English texts. On the contrary, only a limited number of research studies have been devoted to an in-depth review and improvement of the classification of Arabic texts, which is still limited by the challenges of Arabic language processing. Despite the existence of some Arabic text classifiers in the literature, most of them are based on existing ML classifiers. A few Arabic text classifiers have also used deep learning techniques to enhance the accuracy of the used classifiers. Moreover, Arabic text classification is rarely addressed through new classification models. For this purpose, we investigate the effectiveness of analogical proportions as a promising tool to represent the relationship between texts and their appropriate categories. We suggest two different procedures to quantify the relevance/irrelevance of a given term (keyword) used to predict the document class. The suggested language-independent algorithms, named AATC1 and AATC2, are based on the following concept: we first examine the analogical relationship between a new document to be classified and each triplet of pairs from the training set (each pair is represented as a document and its corresponding class). If such a triplet builds a valid analogical proportion with the new document, this states the basis for predicting the class for the new document. For the first algorithm, the set of keywords is extracted from the document being classified, whereas for the second algorithm it is extracted from the set of classes. In the latter, we assume that relevant keywords extracted from the whole class are generally useful for classifying any document belonging to this class. In addition, some keywords are not necessarily relevant to the document to be classified but are still useful for its classification. The experiments prove the effectiveness of analogical proportions in the process of classifying Arabic texts. Compared to some existing text classifiers, the suggested analogical classifiers provide significant statistical improvements, in terms of Precision, Recall, F1 and Accuracy Rate, using large (ANT v2.1) and small (BBC-Arabic) datasets. As perspectives, we plan to use the proposed ANT corpus in other NLP and text mining tasks such as Arabic text clustering, abstractive single and multi-document Arabic text summarization, question answering, etc. Indeed, an analogical abstractive single and multi- document Arabic text summarizer can be developed and then compared to the existing abstractive Arabic summarizers. Besides, given the generic nature of the language-independent analogical-based Arabic text summarizers and classifiers, it is relevant to investigate the efficiency of AATS and AATC using benchmark textual datasets in other languages. Moreover, given the encouraging results obtained by analogical proportions in the field of classification and summarization of Arabic texts, as well as in the field of IR, we wonder if these tools could be exploited in other related fields such as Arabic NLP or Arabic Mono, Multi- and cross-language IR. The ANT corpus can be promising if we transform it into a collection of standard texts for Arabic extraction or IR/CLIR tasks. The IR/CLIR tasks require a set of queries along with their corresponding sets of relevant documents using a recognized matching pattern. It can also be useful in the area of identifying and assessing the reliability of news.
Citer ce document
Accès au document
Voir sur le dépôt sourceCe document est hébergé sur son dépôt institutionnel d'origine.
Auteur(s)
Statistiques
Consultations : 2
Téléchargements : 0