Skip to content

CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice

Julie Thompson, Desmond G. Higgins, Toby J. Gibson

Nucleic Acids Research · 1994 · 65,374 citationsOpen access

Abstract

The sensitivity of the commonly used progressive multiple sequence alignment method has been greatly improved for the alignment of divergent protein sequences. Firstly, individual weights are assigned to each sequence in a partial alignment in order to down-weight near-duplicate sequences and up-weight the most divergent ones. Secondly, amino acid substitution matrices are varied at different alignment stages according to the divergence of the sequences to be aligned. Thirdly, residue-specific gap penalties and locally reduced gap penalties in hydrophilic regions encourage new gaps in potential loop regions rather than regular secondary structure. Fourthly, positions in early alignments where gaps have been opened receive locally reduced gap penalties to encourage the opening up of new gaps at these positions. These modifications are incorporated into a new program, CLUSTAL W which is freely available.

Cite this paper

Thompson, J., Higgins, D. G., & Gibson, T. J. (1994). CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic Acids Research, 22(22), 4673–4680. https://doi.org/10.1093/nar/22.22.4673

Read it with every claim anchored

Add this paper to a project, ask questions of it, and get answers that point to the exact passage.

Start free
  1. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs1997
  2. STAR: ultrafast universal RNA-seq aligner2012
  3. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability2013
  4. MUSCLE: multiple sequence alignment with high accuracy and high throughput2004
  5. Cutadapt removes adapter sequences from high-throughput sequencing reads2011

Metadata from OpenAlex (CC0). Citations are generated from the published record.