Accès ouvert

An end-to-end transformer based approach for disease prediction from metagenomic data

Thèse 2025 Anglais

Résumé

The human gut microbiome is a complex and diverse ecosystem composed of millions of microorganisms spanning thousands of species. It plays a crucial role in maintaining human health, and changes in its composition are frequently associated with various diseases. As such, the state of the gut microbiome serves as a valuable indicator of health, and understanding its internal dynamics is key to both diagnosing and preventing disease. Recent advances in sequencing technologies have enabled the collection of microbial samples from both healthy individuals and patients, providing access to the DNA content of these communities. These metagenomic samples consist of millions to tens of millions of short DNA sequences, making them highly complex and challenging to analyze. Current bioinformatics approaches typically rely on aligning these sequences to reference genomes to identify the species present in a sample, generating abundance profiles that can be linked to health phenotypes. While effective, these methods have notable limitations: they depend on reference databases, are computationally intensive, and reduce rich sequence-level data to species-level summaries. In this work, we propose an alternative approach that leverages recent advances in Artificial Intelligence—particularly Transformer architectures from Natural Language Processing—to analyze metagenomic DNA sequences by treating DNA as a language. Our method, MetagenBERT, consists of two main steps. First, we use Transformer models to produce high-dimensional mathematical representations (embeddings) of individual DNA reads. Second, we aggregate these embeddings to generate a sample-level representation of the metagenome, which can then be used for disease classification. We explore different clustering strategies to achieve this aggregation, and demonstrate that our method performs on par with or outperforms state-of-the-art approaches. These clustering techniques also enable novel representations of metagenomes—either as sets of well-chosen vectors or as new abundance profiles derived from embedding clusters rather than species identity. In doing so, we develop an end-to-end pipeline for metagenome classification that operates directly on raw reads, is independent of reference genomes, and preserves information beyond traditional taxonomic frameworks. Our findings show that this approach can serve as a powerful alternative or complement to conventional bioinformatics tools, opening new perspectives for the application of AI in metagenomics and precision medicine.

Citer ce document

Roy, G. (2025). An end-to-end transformer based approach for disease prediction from metagenomic data.

Accès au document

Voir sur le dépôt source

Ce document est hébergé sur son dépôt institutionnel d'origine.

Auteur(s)

Statistiques

Consultations : 4

Téléchargements : 0