Report on Annotation: Recommended Tools

Version 4.0—May 2026

To accompany the Recommendations, the EBP provides A Report on Annotation Standards.

The annotation subcommittee has assembled a list of tools that have proven useful for given steps of the annotation process to one or more of its members. This list was last reviewed in May 2026 and is not meant to be exhaustive. Comments provided for each tool on scope, parameters or performance are based on the experience of this committee only. It is recommended to always use the latest version of a tool.



Repeat masking

RepeatModeler2

For building a de novo library of repeats that can be used for repeat masking. Options for additional LTR identification.

URL: https://www.repeatmasker.org/RepeatModeler/

Publication: https://doi.org/10.1073/pnas.1921046117

TE-Trimmer

Automates manual curation of TE libraries: TE boundary definition and classification to improve library construction. 

URL: https://github.com/qjiangzhao/TEtrimmer

Publication: https://doi.org/10.1101/2024.06.27.600963

RepeatMasker

For annotating and masking repetitive elements in genomic sequences.

URL: https://www.repeatmasker.org/

Tandem Repeat Finder

Integrated in RepeatMasker but can additionally be run with different parameters to mask more repeats in large genomes. 

URL: https://github.com/Benson-Genomics-Lab/TRF

Publication: https://doi.org/10.1093/nar/27.2.573

Ultra

For the identification and classification of tandem repeats (including satellites). 

URL: https://github.com/TravisWheelerLab/ULTRA

Publication: https://doi.org/10.1101/2024.06.03.597269

RepeatDetector

For detecting repeats in genomic sequences. K-mer based. Works well for vertebrates, insects, and plants.

URL: https://github.com/BioinformaticsToolsmith/Red

Publication: https://doi.org/10.1093/nargab/lqac089

Windowmasker

For masking low-complexity regions in genomic sequences. K-mer based. Tends to mask less than other tools. 

URL: https://www.ncbi.nlm.nih.gov/tools/windowmasker/

Publication: https://doi.org/10.1093/bioinformatics/bti774

The Extensive de novo TE Annotator (EDTA)

For automated de novo TE annotation and species-specific TE library construction.

URL: https://github.com/oushujun/EDTA

Publication: https://doi.org/10.1186/s13059-019-1905-y


Short read alignment

STAR

RNA-Seq aligner that handles spliced alignments, with an option for diploid genomes. Two-pass mapping increases intron discovery. 

URL: https://github.com/alexdobin/STAR

Publication: https://doi.org/10.1093/bioinformatics/bts635

HISAT2

Fast, RAM-efficient and sensitive RNA-Seq aligner. 

URL: https://github.com/DaehwanKimLab/hisat2

Publication: https://doi.org/10.1038/s41587-019-0201-4


Long read/cDNA alignment

MINIMAP2

IsoSeq consensus or ONT reads aligner. Options for ONT cDNA and direct RNA.

URL: https://github.com/lh3/minimap2

Publication: https://doi.org/10.1093/bioinformatics/btab705


Transcript reconstruction

Stringtie2

For RNA-Seq and IsoSeq or ONT long reads.

URL: https://github.com/gpertea/stringtie

Publication: https://doi.org/10.1186/s13059-019-1910-1

Psiclass

For RNA-Seq reads. Can operate with a variable number of libraries, optimised for rare isoform construction.  

URL: https://github.com/splicebox/PsiCLASS

Publication: https://doi.org/10.1038/s41467-019-12990-0

Scallop

For RNA-Seq reads. Sensitive.

URL: https://github.com/Kingsford-Group/scallop

Publication: https://doi.org/10.1186/s13059-019-1883-0

Mikado

For combining and filtering transcript assemblies originating from multiple methods, samples, or sequencing technologies. 

URL: https://github.com/EI-CoreBioinformatics/mikado

Publication: https://doi.org/10.1093/gigascience/giy093

Isoquant

For genome-guided long-read transcript reconstruction and quantification from PacBio/ONT data.

URL: https://github.com/ablab/IsoQuant

Publication: https://doi.org/10.1038/s41587-022-01565-y


Protein-to-genome alignment

Miniprot

Extremely fast. Reasonably accurate if the distance is not very large (e.g., within mammals). 

URL: https://github.com/lh3/miniprot

Publication: https://doi.org/10.1093/bioinformatics/btad014

Spaln

Useful for larger evolutionary distances.  

URL: https://github.com/ogotoh/spaln

Publications: https://doi.org/10.1093/nar/gks708 and https://doi.org/10.1093/bioinformatics/btae517


Protein-coding gene predictors

BRAKER4

Use when RNA-Seq data is available, or with proteins only (BRAKER2 mode). 

URL: https://github.com/Gaius-Augustus/BRAKER4

Publication: https://doi.org/10.1101/gr.278090.123

GALBA2

For larger genomes. Quick, useful when annotating closely related species where no RNA-Seq is available. 

URL: https://github.com/Gaius-Augustus/GALBA2

Publication: https://doi.org/10.1186/s12859-023-05449-z

GEMOMA

Requires proteins and reference annotations from closely related genome/species, optionally uses short read RNA-Seq alignments. 

URL: https://github.com/Jstacs/Jstacs/tree/master/projects/gemoma 

Publication: https://doi.org/10.1007/978-1-4939-9173-0_9

EVIANN

Requires proteins and reference annotations from closely related genome/species, optionally uses short read RNA-Seq and/or IsoSeq data. 

URL: https://github.com/alekseyzimin/EviAnn_release/

Publication: https://doi.org/10.1101/2025.05.07.652745

EASEL

Combined structural/functional annotation; requires short read (or IsoSeq) RNA-Seq alignments; pre-trained models for plants, invertebrates, and vertebrates, designed for complex genomes.

URL: https://gitlab.com/PlantGenomicsLab/easel

Webinar link: https://www.youtube.com/watch?v=01ZCy5f72F8

EGAPX

Supports chordates, arthropods, echinoderms, cnidaria, monocots, dicots; also adds functional annotation (product names) to gene structures.

URL: https://github.com/ncbi/egapx


Protein-coding gene set combiners 

TSEBRA

Combines gene sets from BRAKER, Tiberius, and GALBA. 

URL: https://github.com/Gaius-Augustus/TSEBRA

Publication: https://doi.org/10.1186/s12859-021-04482-0


Protein-coding gene predictors (ab initio

HELIXER

Models are available for: land plants, fungi, vertebrates, and invertebrates; suggest to use in combination with other tools.

URL: https://github.com/weberlab-hhu/Helixer 

Publication: https://doi.org/10.1038/s41592-025-02939-1

Web interface: https://www.plabipd.de/helixer_main.html

ANNEVO

Models are available for: embryophyta, fungi, insecta, invertebrates, mammalia, and other vertebrates.

URL: https://github.com/xjtu-omics/ANNEVO

Publication: https://doi.org/10.21203/rs.3.rs-6402260/v1

ORIONGENO

Models are available for vertebrates (mammals, fish, birds, and others), invertebrates (arthropods and others), plants (angiosperms and bryophytes), fungi, and some algae and protists.

URL: https://github.com/BGIResearch/OrionGeno

Publication: https://doi.org/10.64898/2026.04.26.720859

TIBERIUS

Models are available for: diatoms, eudicotyledons, fungi, insecta, mammalia, monocotyledons, vertebrates, and chlorophyta.

URL: https://github.com/Gaius-Augustus/Tiberius 

Publication: https://doi.org/10.1093/bioinformatics/btae685 


Annotation by projection 

LIFTOFF

Annotation projection from “closely” related species. 

URL: https://github.com/agshumate/Liftoff  

Publication: https://doi.org/10.1093/bioinformatics/btaa1016 

Lifton

Annotation projection from the same or different species. 

URL: https://github.com/Kuanhao-Chao/LiftOn 

Publication: https://doi.org/10.1101/2024.05.16.593026  

CESAR2

Requires the regions to be pre-aligned. Can work at large distances. 

URL: https://github.com/hillerlab/CESAR2.0

Publication: https://doi.org/10.1093/bioinformatics/btx527 

CAT

Good for smaller distances (e.g., primates). 

URL: https://github.com/ComparativeGenomicsToolkit/Comparative-Annotation-Toolkit 

Publication: https://doi.org/10.1101/gr.233460.117


Functional annotation 

Interproscan

For predicting domain architectures. 

URL: https://github.com/ebi-pf-team/interproscan 

Publication: https://doi.org/10.1093/bioinformatics/btu031 

entap

Sequence similarity, gene family, ontology, protein domain, contaminant and HGT detection. 

URL: https://gitlab.com/PlantGenomicsLab/EnTAP

Publication: https://doi.org/10.1111/1755-0998.13106 

Eggnog-mapper

Fast orthology-based functional annotation that complements InterProScan.

URL: https://github.com/eggnogdb/eggnog-mapper

Publication: https://doi.org/10.1093/molbev/msab293

Fantasia

For predicting GO terms, when used in combination with topGO. Very fast on GPU (e.g. a few minutes on proteome of one species).  

URL: https://github.com/MetazoaPhylogenomicsLab/FANTASIA

Publications: https://doi.org/10.1093/nargab/lqae078 and https://doi.org/10.1101/2024.02.28.582465


Non-coding gene predictors 

RFAM + Infernal cmsearch

For predicting short non-coding RNAs.

RFAM URL: https://rfam.org/ and publication: https://doi.org/10.1093/nar/gkaa1047

INFERNAL URL: http://eddylab.org/infernal/ and publication: https://doi.org/10.1093/bioinformatics/btt509 

Tranascan-SE

For predicting tRNAs. Has a eukaryotic high confidence filter that can be applied to significantly reduce false positives. 

URL: https://github.com/UCSC-LoweLab/tRNAscan-SE 

Publication: https://doi.org/10.1007/978-1-4939-9173-0_1  


Annotation QC

Omark

For assessing the annotation quality of predicted proteomes, including completeness, overprediction, and contamination checks.

URL: https://github.com/DessimozLab/OMArk

Publication: https://doi.org/10.1038/s41587-024-02147-w

BUSCO

For assessing the completeness of annotated gene sets using conserved single-copy orthologs.

URL: https://busco.ezlab.org

Publication: https://doi.org/10.1093/nar/gkae987

GFFcompare

For comparing multiple structural annotations.

URL: https://github.com/gpertea/gffcompare

Publication: https://doi.org/10.12688/f1000research.23297.2

AGAT

For checking, fixing, and standardising GFF/GTF annotation files before downstream use.

URL: https://github.com/NBISweden/AGAT


About the Subcommittee

This Report on Annotation Tools Recommendations was developed by EBP’s Scientific Subcommittee for Annotation.