Skip to content

The Pfam Protein Families Database

Alex G. Bateman

Nucleic Acids Research · 2002 · 14,196 citationsOpen access

Abstract

Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the World Wide Web in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgb.ki.se/Pfam/, in France at http://pfam.jouy.inra.fr/ and in the US at http://pfam.wustl.edu/. The latest version (6.6) of Pfam contains 3071 families, which match 69% of proteins in SWISS-PROT 39 and TrEMBL 14. Structural data, where available, have been utilised to ensure that Pfam families correspond with structural domains, and to improve domain-based annotation. Predictions of non-domain regions are now also included. In addition to secondary structure, Pfam multiple sequence alignments now contain active site residue mark-up. New search tools, including taxonomy search and domain query, greatly add to the functionality and usability of the Pfam resource.

Cite this paper

Bateman, A. G. (2002). The pfam protein families database. Nucleic Acids Research, 30(1), 276–280. https://doi.org/10.1093/nar/30.1.276

Read it with every claim anchored

Add this paper to a project, ask questions of it, and get answers that point to the exact passage.

Start free
  1. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs1997
  2. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice1994
  3. STAR: ultrafast universal RNA-seq aligner2012
  4. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability2013
  5. MUSCLE: multiple sequence alignment with high accuracy and high throughput2004

Metadata from OpenAlex (CC0). Citations are generated from the published record.