A Fresh Look for Exploring 3D Protein Structures and Finding Related Proteins

A Fresh Look for Exploring 3D Protein Structures and Finding Related Proteins

NLM’s NCBI is excited to share the modernized look and functionality of two tools used for protein structure research: Structure Summary Pages and the Vector Alignment Search Tool (VAST). The updated versions are available to try now and will fully replace the current versions in October. 

Structure Summary Pages show basic information about macromolecular structures and offer detailed 3D views to interactively study their shapes, functional sites, and how they bind to other molecules.  

Our Vector Alignment Search Tool (VAST) helps you find and compare protein structures with similar 3D shapes. It uses simplified geometric representations to detect matching folding patterns and identify “structural neighbors,” which can reveal evolutionary relationships even when the protein sequences are very different.   Continue reading “A Fresh Look for Exploring 3D Protein Structures and Finding Related Proteins”

NCBI Taxonomy Has a New Home

NCBI Taxonomy Has a New Home

We are excited to introduce the new NCBI Taxonomy, with updated tools and resources to improve how you browse taxonomic information and access NCBI data, including sequences, genomes, and other resources. Highlights of the redesigned site include the Taxonomy BrowserCommon Tree, and Taxonomy name/ID check status tool. 

What’s new? 

Redesigned Taxonomy Browser 

As previously announced, the new Taxonomy Browser provides search autosuggestions, more informed navigation, flexible tree views, and enhanced access to sequences, genomes, and other NCBI resources.    Continue reading “NCBI Taxonomy Has a New Home”

GenBank Release 273.0 is Available!

GenBank Release 273.0 is Available!

GenBank release 273.0 (8/15/2026) is now available on the NCBI FTP site. This release has 60.07 trillion bases and 6.68 billion records. 

The current release has:  

  • 267,383,895 traditional records containing 8,236,878,868,450 base pairs of sequence data
  • 5,132,699,813 WGS records containing 50,829,714,144,609 base pairs of sequence data
  • 1,082,802,170 bulk-oriented TSA records containing 923,937,913,366 base pairs of sequence data
  • 193,606,949 bulk-oriented TLS records containing 80,213,567,020 base pairs of sequence data

Continue reading “GenBank Release 273.0 is Available!”

Now Available: Updated Bacterial and Archaeal Reference Genome Collection

Now Available: Updated Bacterial and Archaeal Reference Genome Collection

Download the updatedbacterial and archaeal reference genome collection! We built this collection of 23,519 genomes by selecting the “best” genome assembly for each species among the 500,000+ prokaryotic genomes inRefSeq.  

What’s new?  
  • Twenty-one species are represented in this collection for the first time  
  • 114 species are represented by a better assembly 
  • Eight species were removed because of changes in NCBI Taxonomy or uncertainty in their species assignment  

Continue reading “Now Available: Updated Bacterial and Archaeal Reference Genome Collection”

BioCollections: Connecting Sequence Data to Physical Specimens

BioCollections: Connecting Sequence Data to Physical Specimens

NCBI’s BioCollections is a curated database of metadata for institutions that house specimens, cultures, and other biological materials. These institutions include museums, herbaria, stock centers, and culture collection centers. Each BioCollections record describes an institution and its collection, and includes standardized details such as the institution name, institution code, collection code, collection type, country, and available link out URL formula. When researchers submit sequences to NCBI, the BioCollections institution code and specimen identifier help link the sequence record to the corresponding physical specimen held in an institutional repository. 

What’s new? 

BioCollections is moving to a new home with an updated design where you can browse, search, and find a collection record.   Continue reading “BioCollections: Connecting Sequence Data to Physical Specimens”

Upcoming Transition of the Public Access Compliance Monitor (PACM) in Support of the 2024 NIH Public Access Policy

As shared in a previous National Center for Biotechnology Information (NCBI) Insights post, NCBI at the National Library of Medicine (NLM) is making ongoing updates to several services to support implementation of the 2024 NIH Public Access Policy, which went into effect on July 1, 2025. As part of the effort, NCBI has been exploring alternative strategies for helping institutions identify their funded articles and track manuscript submissions.

After September 30, 2026, the services that NCBI offers to assist institutions with public access compliance monitoring will move from the Public Access Compliance Monitor (PACM) to a new service within the NIH Manuscript Submission System (NIHMS). After September 30 PACM will no longer be available.

The new service, the NIHMS Institutional Monitor, is now available to assist institutions who have received NIH awards to track the status of Author Accepted Manuscripts (AAMs) submitted to the NIH Manuscript Submission (NIHMS) system. Continue reading “Upcoming Transition of the Public Access Compliance Monitor (PACM) in Support of the 2024 NIH Public Access Policy”

Now Available: RefSeq Release 236

Now Available: RefSeq Release 236

RefSeq release 236 is now available online and from the FTP site! You can access RefSeq data through NCBI Datasets. The release is provided in several directories as a complete dataset and also as divided by logical groupings.    

What’s included in this release? 

As of July 6, 2026, this full release incorporates genomic, transcript, and protein data containing:   

  • 629,953,391 records  
  • 482,864,455 proteins  
  • 83,806,434 RNAs  
  • Sequences from 182,465 organisms  

Continue reading “Now Available: RefSeq Release 236”

Now Available! NCBI Hidden Markov Models (HMM) Release 20.0

Now Available! NCBI Hidden Markov Models (HMM) Release 20.0

Download release 20.0 of the NCBI protein profile Hidden Markov models (HMMs) used by the Prokaryotic Genome Annotation Pipeline (PGAP). You can search this collection against your favorite prokaryotic proteins to identify their function using the HMMER sequence analysis package. 

What’s new? 

Release 20.0 contains: 

  • 18,950 HMMs maintained by NCBI 
  • 497 new HMMs since release 19.0 

Continue reading “Now Available! NCBI Hidden Markov Models (HMM) Release 20.0”

Standards for Sequence Read Archive (SRA) Data Submission

Standards for Sequence Read Archive (SRA) Data Submission

What you need to know for 2026 and beyond! 

The volume of genomic sequencing data is growing rapidly, and we want to ensure that publicly shared datasets remain useful, findable, and scientifically trustworthy. To support this goal, the Sequence Read Archive (SRA) is implementing a set of data submission standards that improve data consistency, quality, and long-term value—benefiting both the research community and the data submitters. 

Why new standards? 

High-quality data and metadata ensure that researchers can effectively reuse sequencing data and that SRA can process submissions quickly and accurately. These standards align SRA submissions with evolving INSDC expectations and best practices in genomic data stewardship by: 

    • Preventing common formatting and metadata errors 
    • Ensuring that sequencing runs remain useful for future scientific research 
    • Speeding up SRA data processing and public release timelines 

These standards are expected to go into effect at the end of 2026. Here’s what to expect:  Continue reading “Standards for Sequence Read Archive (SRA) Data Submission”

New Minimum Requirements for BioProject and BioSample Fields

New Minimum Requirements for BioProject and BioSample Fields

BioProject and BioSample support the discovery, access, and reuse of nucleotide sequence data at NCBI. BioProject records provide a centralized place for accessing diverse data generated as part of a single research initiative. BioSample records describe the biological source materials used to generate those data. To ensure data quality, NCBI, in conjunction with the International Nucleotide Sequence Database Collaboration (INSDC), has established minimum criteria for submission acceptance. These changes are part of NCBI’s broader effort to implement INSDC minimum specifications across sequence data resources and submission workflows. 

What’s new? 

Starting early 2027, new validation requirements will be applied during BioProject and BioSample submissions.  Continue reading “New Minimum Requirements for BioProject and BioSample Fields”