Search the NCBI Hidden Markov models collection against your favorite prokaryotic proteins

The NCBI Hidden Markov models (HMM) 6.0 release, available on our FTP site, has 15,247 models supported at NCBI. We created 80 more new HMMs and consolidated the collection by removing 2,151 HMMs that were nearly identical to another. Release 6.0 also incorporates 12,656 PFAM from release 34 that apply to prokaryotic proteins. You can use the HMMER sequence analysis package to search the collection against your favorite prokaryotic proteins to identify their function. We have also added more specific names or associated EC number, gene symbols and publication to over 500 HMMs.

Gene Ontology (GO) term attributes are now available for 20% of HMM models (see Figure 1 below). We added most of these based on existing mappings, but our experts are working on creating more associations. Starting in the fall, we’ll start propagating GO terms from HMMs to annotated genomes and proteins!

Example Protein Family Model, TIGR03697.1 for the global nitrogen regulator NtcA protein family, with newly shown GO terms (framed in red).
Figure 1. Example Protein Family Model, TIGR03697.1 for the global nitrogen regulator NtcA protein family, with newly shown GO terms (framed in red).

You can search HMMs in the Entrez Protein Family Model collection, which also includes conserved domain architectures and BlastRules, and find all RefSeq proteins they name, using a variety of terms such as gene symbol, product name and now, GO identifier or description. See this earlier post for more details.

One thought on “Search the NCBI Hidden Markov models collection against your favorite prokaryotic proteins

Leave a Reply