Tag: RefSeq

Evidence for naming the protein now on non-redundant refseq records (WP_ accessions)

We are now showing the curated evidence used for assigning names and, if possible, gene symbols, publications, and Enzyme Commission numbers on nearly 70% (83 million) microbial RefSeq proteins. This evidence includes a hierarchical collection of curated Hidden Markov Model (HMM)-based and BLAST-based protein families, and conserved domain architectures.

Continue reading “Evidence for naming the protein now on non-redundant refseq records (WP_ accessions)” →

RefSeq release 95: naming evidence added to all relevant WP proteins

RefSeq release 95 is accessible online, via FTP and through NCBI’s Entrez programming utilities, E-utilities.

This full release incorporates genomic, transcript, and protein data available, as of July 8, 2019 and contains 206,416,381 records, including 146,381,777 proteins, 27,212,750 RNAs, and sequences from 93,618 organisms.

Continue reading “RefSeq release 95: naming evidence added to all relevant WP proteins” →

March-April 2019 RefSeq eukaryotic annotations

In March and April, the NCBI Eukaryotic Genome Annotation Pipeline released new annotations in RefSeq for the following organisms:

Continue reading “March-April 2019 RefSeq eukaryotic annotations” →

RefSeq release 94 with MANE and RefSeq Select markup, protein name evidence, and improved [Candida] auris assembly

RefSeq release 94 is now available through NCBI web services, FTP and through NCBI’s Entrez programming utilities, E-utilities.

This full release incorporates genomic, transcript, and protein data available, as of May 13, 2019 and contains 200,311,267 records, including 141,839,334 proteins, 26,534,602 RNAs, and sequences from 91,873 organisms. The release is provided in several directories as a complete dataset and also as divided by logical groupings.

Continue reading “RefSeq release 94 with MANE and RefSeq Select markup, protein name evidence, and improved [Candida] auris assembly” →

Human genome annotation will be updated every 2 months

NCBI will be updating the human genome RefSeq annotation more frequently to incorporate improvements made to genes and transcripts by RefSeq curation experts. Faster updates will allow us to include the latest datasets.

In the past, we’ve produced a full re-annotation of the human genome about once a year. The last full annotation, Homo sapiens Annotation Release 109, was in March 2018. A full annotation is produced by two main processes:

Continue reading “Human genome annotation will be updated every 2 months” →

Expanded accession formats appear in RefSeq release 93

RefSeq release 93 is accessible online, via FTP and through NCBI’s Entrez programming utilities, E-utilities.

This full release incorporates genomic, transcript, and protein data available as of March 13, 2019. It contains 192,722,653 records, including 135,670,032 proteins, 25,840,272 RNAs, and sequences from 88,816 organisms.

Continue reading “Expanded accession formats appear in RefSeq release 93” →

New RefSeq annotations for big brown bat, peregrine falcon and more

Hibernating brown bat

In January and February, the NCBI Eukaryotic Genome Annotation Pipeline released new annotations in RefSeq for the following organisms:

Aphis gossypii (cotton aphid)
Balaenoptera acutorostrata scammoni (minke whale)
Bombyx mandarina (wild silkworm)
Chelonia mydas (green sea turtle)
Corapipo altera (white-ruffed manakin)
Empidonax traillii (willow flycatcher)
Eptesicus fuscus (big brown bat)
Eumetopias jubatus (Steller sea lion)
Falco cherrug (Saker falcon)
Falco peregrinus (peregrine falcon)
Marmota flaviventris (yellow-bellied marmot)
Monomorium pharaonis (pharaoh ant)
Neopelma chrysocephalum (saffron-crested tyrant-manakin)
Ovis aries (sheep)
Pipra filicauda (wire-tailed manakin)
Rhopalosiphum maidis (corn leaf aphid)
Solanum pennellii (eudicot)
Tupaia chinensis (Chinese tree shrew)
Vigna unguiculata (cowpea)
Vombatus ursinus (common wombat)
Xiphophorus couchianus (Monterrey platyfish)

See more details on the Eukaryotic RefSeq Genome Annotation Status page.

RefSeq release 92 updates 10,000 human transcripts

RefSeq release 92 is accessible online, via FTP and through NCBI’s Entrez programming utilities, E-utilities.

This full release incorporates genomic, transcript, and protein data available, as of January 4, 2019 and contains 185,738,687 records, including 130,366,644 proteins, 25,088,890 RNAs, and sequences from 86,867 organisms. The release is provided in several directories as a complete dataset and as divided by logical groupings.

Continue reading “RefSeq release 92 updates 10,000 human transcripts” →

November and December 2018 RefSeq annotations – hybrid cattle, cheetah, fly & more

two cheetahs sit next to each other in natural habitat

In November and December, the NCBI Eukaryotic Genome Annotation Pipeline released new annotations in RefSeq for the following organisms:

Continue reading “November and December 2018 RefSeq annotations – hybrid cattle, cheetah, fly & more” →

Join NCBI at PAG in San Diego, January 12–16, 2019

Next week, NCBI staff will attend the Plant and Animal Genome (PAG) Conference. We have several activities planned, including 1 booth (#223), 4 workshops, 1 talk and 2 posters.

Read on to learn more about what you can look forward to if you’re attending PAG this year. (Note: The listed times are Pacific time.)

Continue reading “Join NCBI at PAG in San Diego, January 12–16, 2019” →