Deep Learning for Metagenomic Data: using 2D Embeddings and Convolutional Neural Networks

Nguyen, Thanh Hai; Chevaleyre, Yann; Prifti, Edi; Sokolovska, Nataliya; Zucker, Jean-Daniel

Computer Science > Computer Vision and Pattern Recognition

arXiv:1712.00244 (cs)

[Submitted on 1 Dec 2017]

Title:Deep Learning for Metagenomic Data: using 2D Embeddings and Convolutional Neural Networks

Authors:Thanh Hai Nguyen, Yann Chevaleyre, Edi Prifti, Nataliya Sokolovska, Jean-Daniel Zucker

View PDF

Abstract:Deep learning (DL) techniques have had unprecedented success when applied to images, waveforms, and texts to cite a few. In general, when the sample size (N) is much greater than the number of features (d), DL outperforms previous machine learning (ML) techniques, often through the use of convolution neural networks (CNNs). However, in many bioinformatics ML tasks, we encounter the opposite situation where d is greater than N. In these situations, applying DL techniques (such as feed-forward networks) would lead to severe overfitting. Thus, sparse ML techniques (such as LASSO e.g.) usually yield the best results on these tasks. In this paper, we show how to apply CNNs on data which do not have originally an image structure (in particular on metagenomic data). Our first contribution is to show how to map metagenomic data in a meaningful way to 1D or 2D images. Based on this representation, we then apply a CNN, with the aim of predicting various diseases. The proposed approach is applied on six different datasets including in total over 1000 samples from various diseases. This approach could be a promising one for prediction tasks in the bioinformatics field.

Comments:	Accepted at NIPS 2017 Workshop on Machine Learning for Health (this https URL In Proceedings of the NIPS ML4H 2017 Workshop in Long Beach, CA, USA;
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:1712.00244 [cs.CV]
	(or arXiv:1712.00244v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1712.00244

Submission history

From: Thanh Hai Nguyen [view email]
[v1] Fri, 1 Dec 2017 09:18:04 UTC (867 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Deep Learning for Metagenomic Data: using 2D Embeddings and Convolutional Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Deep Learning for Metagenomic Data: using 2D Embeddings and Convolutional Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators