Phymm and PhymmBL: metagenomic phylogenetic classification with interpolated Markov models

Arthur Brady; Steven L Salzberg

doi:10.1038/nmeth.1358

Phymm and PhymmBL: metagenomic phylogenetic classification with interpolated Markov models

Nat Methods. 2009 Sep;6(9):673-6. doi: 10.1038/nmeth.1358. Epub 2009 Aug 2.

Authors

Arthur Brady¹, Steven L Salzberg

Affiliation

¹ Center for Bioinformatics and Computational Biology, University of Maryland, College Park, Maryland, USA. abrady@umiacs.umd.edu

Abstract

Metagenomics projects collect DNA from uncharacterized environments that may contain thousands of species per sample. One main challenge facing metagenomic analysis is phylogenetic classification of raw sequence reads into groups representing the same or similar taxa, a prerequisite for genome assembly and for analyzing the biological diversity of a sample. New sequencing technologies have made metagenomics easier, by making sequencing faster, and more difficult, by producing shorter reads than previous technologies. Classifying sequences from reads as short as 100 base pairs has until now been relatively inaccurate, requiring researchers to use older, long-read technologies. We present Phymm, a classifier for metagenomic data, that has been trained on 539 complete, curated genomes and can accurately classify reads as short as 100 base pairs, a substantial improvement over previous composition-based classification methods. We also describe how combining Phymm with sequence alignment algorithms improves accuracy.

Publication types

Research Support, N.I.H., Extramural

MeSH terms

Artificial Intelligence*
Bacteria / classification*
Bacteria / genetics
Base Sequence
DNA / classification*
DNA / genetics
Genomics / methods*
Hydrogen-Ion Concentration
Markov Chains*
Mining
Models, Genetic*
Phylogeny
Sequence Alignment
Soil Microbiology

Substances

DNA

Abstract

Publication types

MeSH terms

Substances

Grants and funding