
This post accompanies my short Janecka Genomics video tutorial, How to Rename GenBank Sequences for Cleaner Phylogenetic Trees.
When you download sequences from GenBank, FASTA headers can contain accession numbers, species names, gene descriptions, and other source information. This information is valuable, but carrying the entire header forward as the sequence identifier can quickly make alignments and phylogenetic trees difficult to read.
A simple solution is to give each sequence a short, unique, and informative name while preserving the original information in a separate metadata table.
Keep Sequence Names Short, Unique, and Informative
There is no single correct way to rename sequences.
For a dataset containing one RAG1 sequence from each species, for example, a short identifier based on species and gene may work well. For datasets containing multiple individuals from the same species, however, you may need to include sample IDs, populations, locations, accession numbers, or another identifier.
The goal is to choose names that communicate the information you need for that particular analysis.
Whatever naming convention you choose, make sure that:
-
Every sequence has a unique name.
-
Names remain consistent across your alignment, trees, and other analyses.
-
Names are short enough to remain easy to read.
-
Spaces are avoided when possible. Although many programs tolerate spaces, different software may handle them differently.
Consistency is especially important when moving between programs. Some analyses require the sequence labels in an alignment and phylogenetic tree to match exactly. You can rename tips later—for example, by editing a Newick tree—but doing this manually makes it easy to introduce typos or mismatched identifiers.
Why This Matters for Phylogenetic Trees
Long GenBank identifiers can produce very wide and cluttered phylogenetic figures because the full sequence identifier becomes the tip label.
If you maintain a fixed figure width, long labels leave less space for displaying the tree itself. Alternatively, maintaining the desired visual scale of the tree may require a much wider figure.
Neither changes the underlying phylogeny or estimated branch lengths—it causes a figure readability problem.
Short, biologically informative labels make it much easier for a reader to recognize taxa or samples and concentrate on the topology, branch lengths, and support values.
Don’t Throw Away the Original Information
Shortening sequence names should never lead to losing the original sequence information or provenance.
I keep a CSV metadata table in which each short sequence name is linked to its original GenBank accession number and source information. For example:
|
short_name |
accession |
species |
population |
location |
|---|---|---|---|---|
|
Tiger_RAG1 |
GenBank # |
Panthera tigris |
— |
— |
The exact columns will depend on the project.
This approach becomes especially useful later because the CSV file can be imported directly into R, Python, and other analysis software. Additional columns can be added for populations, sampling locations, experimental groups, clades, phenotypes, geographic coordinates, or other variables needed for downstream analyses and figures.
It also gives you a reliable mapping between your short identifiers and the original sequence information. If you later need to rename tree tips or reorganize identifiers, that mapping can be used to automate the process in R rather than manually editing every label.
Finally, always retain your original downloaded sequence file unchanged. Make identifier changes, trimming, or other modifications to a working copy and document what you changed.
A few minutes spent organizing sequence identifiers and metadata at the beginning of a project can save considerable time when you begin analyzing the data, creating figures, or preparing a manuscript.
Tutorial resources: The sequence files and other materials used in the tutorial are available through the Janecka Genomics GitHub repository.
By Jan E. Janecka, PhD - August 30, 2026