MUSCLE: multiple sequence alignment with high accuracy and high throughput
Nucleic Acids Research · 2004 · 47,425 citationsOpen access
Abstract
We describe MUSCLE, a new computer program for creating multiple alignments of protein sequences. Elements of the algorithm include fast distance estimation using kmer counting, progressive alignment using a new profile function we call the log-expectation score, and refinement using tree-dependent restricted partitioning. The speed and accuracy of MUSCLE are compared with T-Coffee, MAFFT and CLUSTALW on four test sets of reference alignments: BAliBASE, SABmark, SMART and a new benchmark, PREFAB. MUSCLE achieves the highest, or joint highest, rank in accuracy on each of these sets. Without refinement, MUSCLE achieves average accuracy statistically indistinguishable from T-Coffee and MAFFT, and is the fastest of the tested methods for large numbers of sequences, aligning 5000 sequences of average length 350 in 7 min on a current desktop computer. The MUSCLE program, source code and PREFAB test data are freely available at http://www.drive5. com/muscle.
Cite this paper
Edgar, R. C. (2004). MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research, 32(5), 1792–1797. https://doi.org/10.1093/nar/gkh340
Read it with every claim anchored
Add this paper to a project, ask questions of it, and get answers that point to the exact passage.
Start freeRelated papers
- Gapped BLAST and PSI-BLAST: a new generation of protein database search programs1997
- CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice1994
- STAR: ultrafast universal RNA-seq aligner2012
- MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability2013
- Cutadapt removes adapter sequences from high-throughput sequencing reads2011
Metadata from OpenAlex (CC0). Citations are generated from the published record.