Learning to Retrieve Relevant Passages and Questions in Open Domain and Community Question Answering
Résumé
Question Answering (QA) aims to directly return succinct and accurate answers to natural language questions. Passage Retrieval (PR) is deemed to be the kernel of a typical QA system where the goal is to reduce the search space from a huge set of documents to a few number of relevant passages, from which the required answer can be found. Although there has been an abundance of work on this task, it still requires non-trivial endeavor. Recently, community Question Answering (cQA) services have evolved into a popular way of online information seeking, where users can interact and exchange knowledge in the form of questions and answers. The Question Retrieval (QR) problem in cQA is to certain extent analogue to the PR task in traditional QA. While passage retrieval matches the user question with the document passages to search for correct excerpts in response to the user, question retrieval matches the user’s question with the archived questions to find out those that are semantically similar to the queried one. By the time, with the sharp increase of community archives and the accumulation of duplicated questions, the QR problem has become increasingly alarming and it remains more challenging than PR due to the shortness of the community questions as well as the lexical gap problem. In this thesis, we tackle both tasks: PR in open domain QA and QR in cQA. We propose different approaches to improve these critical problems in different languages. For PR, we were mainly based on SVM and n-grams while for QR, we were opted for neural networks mainly word embeddings and Long Short-Term Memory (LSTM). We run our experiments on large scale data sets from CLEF and Yahoo! Answers in different languages to show the efficiency and generality of our proposed approaches. Interestingly, the obtained results transcend that of other previously proposed ones.
Citer ce document
Accès au document
Voir sur le dépôt sourceCe document est hébergé sur son dépôt institutionnel d'origine.
Auteur(s)
Statistiques
Consultations : 2
Téléchargements : 0