dbGaP Modernization Update

dbGaP Modernization Update

NCBI is modernizing the database of Genotype and Phenotype (dbGaP), a repository that collects human research study data from NIH-funded investigators and facilitates controlled-access to individual-level data. As previously announced, we released a modernized homepage in 2025. Now, we’ve extended that modernization to the public search and study record pages.

What’s new?

The modernized public study record pages provide a summary of each study that has been registered in dbGaP (see an example in the image below). Use these pages to understand the research design, affiliations and references, associated files and file access, and other metadata for a study. These record pages can help you understand if the individual-level data for a study will be helpful for your research before you initiate a request for access to the controlled data. Continue reading “dbGaP Modernization Update”

Discover the Enhanced SRA Run Browser User Interface

Discover the Enhanced SRA Run Browser User Interface

If you use NCBI’s Sequence Read Archive’s (SRA) Run Browser, you will be excited to explore the updated user interface (UI). The latest enhancements are designed to make your data exploration smoother, more intuitive, and more informative than ever before. Let’s dive into what’s new and how these features can help you get the most out of your STAT (SRA Taxonomy Analysis Tool) output. 

Expanded taxonomy tree for deeper insights 

The highlight of the new Run Browser UI is the expanded taxonomy tree with additional metadata and easy-to-use expand and collapse buttons. This interactive feature allows you to drill down into taxonomic hierarchies, revealing the taxonomic tree of an SRA Run with just a click. With expandable taxonomic levels and additional metadata, the taxonomy tree makes it easier to navigate and interpret the taxa detected in the Run.   Continue reading “Discover the Enhanced SRA Run Browser User Interface”

Now Available: RefSeq Release 237

Now Available: RefSeq Release 237

RefSeq release 237 is now available online and from the FTP site! You can access RefSeq data through NCBI Datasets. The release is provided in several directories as a complete dataset and also as divided by logical groupings.    

What’s included in this release? 

As of August 31, 2026, this full release incorporates genomic, transcript, and protein data containing:

  • 645,505,406 records 
  • 495,290,394 proteins 
  • 86,067,871 RNAs 
  • Sequences from 184,752 organisms 

Continue reading “Now Available: RefSeq Release 237”

Register for NCBI’s Pre-Conference Workshop at ASM BIG 2026

Register for NCBI’s Pre-Conference Workshop at ASM BIG 2026

Are you heading to ASM BIG (Bioinformatics, Genomics, and Big Data Conference) in Washington, DC this fall? If so, make sure to register for the workshop hosted by the National Center for Biotechnology Information (NCBI) as part of the pre-conference activities. This workshop, Empowering Your Microbial Research with Expert Tools and Resources: Unlock the Power of NCBI Data, will showcase a suite of services designed specifically for researchers working with pathogenic bacteria. Whether you are new to genomic data analysis or an experienced scientist looking for advanced analysis tools, this hands-on workshop will introduce NCBI resources and workflows for working with pathogenic bacterial data. 

Workshop goals 

  • Gain practical skills for searching, retrieving, and analyzing genomic and metagenomic datasets with powerful NCBI resources like Datasets, Sequence Read Archive (SRA), and Pathogen Detection  
  • Participate in hands-on exercises using SRA Toolkit, STAT, HRRT and cutting-edge technologies for outbreak investigation and antimicrobial resistance research 
  • Network with fellow scientists and NCBI experts to stay ahead in a rapidly evolving field 

Continue reading “Register for NCBI’s Pre-Conference Workshop at ASM BIG 2026”

Interactive Annotation for Viruses Using VADR

Interactive Annotation for Viruses Using VADR

NCBI Virus is excited to announce the launch of a web interface for Viral Annotation DefineR (VADR). Web VADR can be used to validate and annotate virus sequence data using a web interface and does not require any command line setup like the standalone version. You can upload sequences through the web interface, start a VADR run, and download the results when processing is complete. 

What is Web VADR? 

With just a few clicks in Web VADR you can: 

  • Annotate genes, proteins, and other key features on one or more virus nucleotide sequences 
  • Quickly identify potential issues in input sequences 
  • Review results based on curated reference models 
  • Download an annotation file (feature file) formatted and ready to use in submissions to GenBank  
  • Download accession lists indicating passing and failing sequences, and other files useful for reviewing your results 

Continue reading “Interactive Annotation for Viruses Using VADR”

A Fresh Look for Exploring 3D Protein Structures and Finding Related Proteins

A Fresh Look for Exploring 3D Protein Structures and Finding Related Proteins

NLM’s NCBI is excited to share the modernized look and functionality of two tools used for protein structure research: Structure Summary Pages and the Vector Alignment Search Tool (VAST). The updated versions are available to try now and will fully replace the current versions in October. 

Structure Summary Pages show basic information about macromolecular structures and offer detailed 3D views to interactively study their shapes, functional sites, and how they bind to other molecules.  

Our Vector Alignment Search Tool (VAST) helps you find and compare protein structures with similar 3D shapes. It uses simplified geometric representations to detect matching folding patterns and identify “structural neighbors,” which can reveal evolutionary relationships even when the protein sequences are very different.   Continue reading “A Fresh Look for Exploring 3D Protein Structures and Finding Related Proteins”

NCBI Taxonomy Has a New Home

NCBI Taxonomy Has a New Home

We are excited to introduce the new NCBI Taxonomy, with updated tools and resources to improve how you browse taxonomic information and access NCBI data, including sequences, genomes, and other resources. Highlights of the redesigned site include the Taxonomy BrowserCommon Tree, and Taxonomy name/ID check status tool. 

What’s new? 

Redesigned Taxonomy Browser 

As previously announced, the new Taxonomy Browser provides search autosuggestions, more informed navigation, flexible tree views, and enhanced access to sequences, genomes, and other NCBI resources.    Continue reading “NCBI Taxonomy Has a New Home”

GenBank Release 273.0 is Available!

GenBank Release 273.0 is Available!

GenBank release 273.0 (8/15/2026) is now available on the NCBI FTP site. This release has 60.07 trillion bases and 6.68 billion records. 

The current release has:  

  • 267,383,895 traditional records containing 8,236,878,868,450 base pairs of sequence data
  • 5,132,699,813 WGS records containing 50,829,714,144,609 base pairs of sequence data
  • 1,082,802,170 bulk-oriented TSA records containing 923,937,913,366 base pairs of sequence data
  • 193,606,949 bulk-oriented TLS records containing 80,213,567,020 base pairs of sequence data

Continue reading “GenBank Release 273.0 is Available!”

Now Available: Updated Bacterial and Archaeal Reference Genome Collection

Now Available: Updated Bacterial and Archaeal Reference Genome Collection

Download the updatedbacterial and archaeal reference genome collection! We built this collection of 23,519 genomes by selecting the “best” genome assembly for each species among the 500,000+ prokaryotic genomes inRefSeq.  

What’s new?  
  • Twenty-one species are represented in this collection for the first time  
  • 114 species are represented by a better assembly 
  • Eight species were removed because of changes in NCBI Taxonomy or uncertainty in their species assignment  

Continue reading “Now Available: Updated Bacterial and Archaeal Reference Genome Collection”

BioCollections: Connecting Sequence Data to Physical Specimens

BioCollections: Connecting Sequence Data to Physical Specimens

NCBI’s BioCollections is a curated database of metadata for institutions that house specimens, cultures, and other biological materials. These institutions include museums, herbaria, stock centers, and culture collection centers. Each BioCollections record describes an institution and its collection, and includes standardized details such as the institution name, institution code, collection code, collection type, country, and available link out URL formula. When researchers submit sequences to NCBI, the BioCollections institution code and specimen identifier help link the sequence record to the corresponding physical specimen held in an institutional repository. 

What’s new? 

BioCollections is moving to a new home with an updated design where you can browse, search, and find a collection record.   Continue reading “BioCollections: Connecting Sequence Data to Physical Specimens”