Zhao, Haiyong and Wang, Shuang and Yuan, Xiguo (2020) Detection of Pathogenic Microbe Composition Using Next-Generation Sequencing Data. Frontiers in Genetics, 11. ISSN 1664-8021
pubmed-zip/versions/2/package-entries/fgene-11-603093-r1/fgene-11-603093.pdf - Published Version
Download (2MB)
Abstract
Next-generation sequencing (NGS) technologies have provided great opportunities to analyze pathogenic microbes with high-resolution data. The main goal is to accurately detect microbial composition and abundances in a sample. However, high similarity among sequences from different species and the existence of sequencing errors pose various challenges. Numerous methods have been developed for quantifying microbial composition and abundance, but they are not versatile enough for the analysis of samples with mixtures of noise. In this paper, we propose a new computational method, PGMicroD, for the detection of pathogenic microbial composition in a sample using NGS data. The method first filters the potentially mistakenly mapped reads and extracts multiple species-related features from the sequencing reads of 16S rRNA. Then it trains an Support Vector Machine classifier to predict the microbial composition. Finally, it groups all multiple-mapped sequencing reads into the references of the predicted species to estimate the abundance for each kind of species. The performance of PGMicroD is evaluated based on both simulation and real sequencing data and is compared with several existing methods. The results demonstrate that our proposed method achieves superior performance. The software package of PGMicroD is available at https://github.com/BDanalysis/PGMicroD.
Item Type: | Article |
---|---|
Subjects: | Lib Research Guardians > Medical Science |
Depositing User: | Unnamed user with email support@lib.researchguardians.com |
Date Deposited: | 30 Jan 2023 10:44 |
Last Modified: | 13 Jul 2024 13:25 |
URI: | http://eprints.classicrepository.com/id/eprint/53 |