Some contributions to deep learning for metagenomics
Résumé
Metagenomic data from human microbiome is a novel source of data for improving diagnosis and prognosis in human diseases. However, to do a prediction based on individual bacteria abundance is a challenge, since the number of features is much bigger than the number of samples. Therefore, we face the difficulties related to high dimensional data processing, as well as to the high complexity of heterogeneous data. Machine Learning (ML) in general, and Deep Learning (DL) in particular, has obtained great achievements on important metagenomics problems linked to OTU-clustering, binning, taxonomic assignment, comparative metagenomics, and gene prediction. ML offers powerful frameworks to integrate a vast amount of data from heterogeneous sources, to design new models, and to test multiple hypotheses and therapeutic products. The contribution of this PhD thesis is multi-fold: 1) we introduce a feature selection framework for efficient heterogeneous biomedical signature extraction, and 2) a novel DL approach for predicting diseases using artificial image representations. The first contribution is an efficient feature selection approach based on visualization capabilities of Self-Organising Maps (SOM) for heterogeneous data fusion. We reported that the framework is efficient on a real and heterogeneous dataset called MicrObese, containing metadata, genes of adipose tissue, and gut flora metagenomic data with a reasonable classification accuracy compared to the state-of-the-art methods. The second approach developed in the context of this PhD project, is a method to visualize metagenomic data using a simple fill-up method, and also various state-of-the-art dimensional reduction learning approaches. The new metagenomic data representation can be considered as synthetic images, and used as a novel data set for an efficient deep learning method such as Convolutional Neural Networks. We also explore applying Local Interpretable Model-agnostic explanations (LIME), Saliency Maps and Gradient-weighted Class Activation (Grad-CAM) to identify important regions in the newly constructed artificial images which might help to explain the predictive models. We show by our experimental results that the proposed methods either achieve the state-of-the-art predictive performance, or outperform it on public rich metagenomic benchmarks.
Citer ce document
Accès au document
Voir sur le dépôt sourceCe document est hébergé sur son dépôt institutionnel d'origine.
Auteur(s)
Statistiques
Consultations : 3
Téléchargements : 0