Report on Annotation: Recommended Tools
Version 4.0—May 2026
To accompany the Recommendations, the EBP provides A Report on Annotation Standards.
The annotation subcommittee has assembled a list of tools that have proven useful for given steps of the annotation process to one or more of its members. This list was last reviewed in May 2026 and is not meant to be exhaustive. Comments provided for each tool on scope, parameters or performance are based on the experience of this committee only. It is recommended to always use the latest version of a tool.
Repeat masking
RepeatModeler2
For building a de novo library of repeats that can be used for repeat masking. Options for additional LTR identification.
URL: https://www.repeatmasker.org/RepeatModeler/
Publication: https://doi.org/10.1073/pnas.1921046117
TE-Trimmer
Automates manual curation of TE libraries: TE boundary definition and classification to improve library construction.
URL: https://github.com/qjiangzhao/TEtrimmer
Publication: https://doi.org/10.1101/2024.06.27.600963
RepeatMasker
For annotating and masking repetitive elements in genomic sequences.
URL: https://www.repeatmasker.org/
Tandem Repeat Finder
Integrated in RepeatMasker but can additionally be run with different parameters to mask more repeats in large genomes.
URL: https://github.com/Benson-Genomics-Lab/TRF
Publication: https://doi.org/10.1093/nar/27.2.573
Ultra
For the identification and classification of tandem repeats (including satellites).
URL: https://github.com/TravisWheelerLab/ULTRA
Publication: https://doi.org/10.1101/2024.06.03.597269
RepeatDetector
For detecting repeats in genomic sequences. K-mer based. Works well for vertebrates, insects, and plants.
URL: https://github.com/BioinformaticsToolsmith/Red
Publication: https://doi.org/10.1093/nargab/lqac089
Windowmasker
For masking low-complexity regions in genomic sequences. K-mer based. Tends to mask less than other tools.
URL: https://www.ncbi.nlm.nih.gov/tools/windowmasker/
Publication: https://doi.org/10.1093/bioinformatics/bti774
The Extensive de novo TE Annotator (EDTA)
For automated de novo TE annotation and species-specific TE library construction.
URL: https://github.com/oushujun/EDTA
Publication: https://doi.org/10.1186/s13059-019-1905-y
Short read alignment
STAR
RNA-Seq aligner that handles spliced alignments, with an option for diploid genomes. Two-pass mapping increases intron discovery.
URL: https://github.com/alexdobin/STAR
Publication: https://doi.org/10.1093/bioinformatics/bts635
HISAT2
Fast, RAM-efficient and sensitive RNA-Seq aligner.
URL: https://github.com/DaehwanKimLab/hisat2
Publication: https://doi.org/10.1038/s41587-019-0201-4
Long read/cDNA alignment
MINIMAP2
IsoSeq consensus or ONT reads aligner. Options for ONT cDNA and direct RNA.
URL: https://github.com/lh3/minimap2
Publication: https://doi.org/10.1093/bioinformatics/btab705
Transcript reconstruction
Stringtie2
For RNA-Seq and IsoSeq or ONT long reads.
URL: https://github.com/gpertea/stringtie
Publication: https://doi.org/10.1186/s13059-019-1910-1
Psiclass
For RNA-Seq reads. Can operate with a variable number of libraries, optimised for rare isoform construction.
URL: https://github.com/splicebox/PsiCLASS
Publication: https://doi.org/10.1038/s41467-019-12990-0
Scallop
For RNA-Seq reads. Sensitive.
URL: https://github.com/Kingsford-Group/scallop
Publication: https://doi.org/10.1186/s13059-019-1883-0
Mikado
For combining and filtering transcript assemblies originating from multiple methods, samples, or sequencing technologies.
URL: https://github.com/EI-CoreBioinformatics/mikado
Publication: https://doi.org/10.1093/gigascience/giy093
Isoquant
For genome-guided long-read transcript reconstruction and quantification from PacBio/ONT data.
URL: https://github.com/ablab/IsoQuant
Publication: https://doi.org/10.1038/s41587-022-01565-y
Protein-to-genome alignment
Miniprot
Extremely fast. Reasonably accurate if the distance is not very large (e.g., within mammals).
URL: https://github.com/lh3/miniprot
Publication: https://doi.org/10.1093/bioinformatics/btad014
Spaln
Useful for larger evolutionary distances.
URL: https://github.com/ogotoh/spaln
Publications: https://doi.org/10.1093/nar/gks708 and https://doi.org/10.1093/bioinformatics/btae517
Protein-coding gene predictors
BRAKER4
Use when RNA-Seq data is available, or with proteins only (BRAKER2 mode).
URL: https://github.com/Gaius-Augustus/BRAKER4
Publication: https://doi.org/10.1101/gr.278090.123
GALBA2
For larger genomes. Quick, useful when annotating closely related species where no RNA-Seq is available.
URL: https://github.com/Gaius-Augustus/GALBA2
Publication: https://doi.org/10.1186/s12859-023-05449-z
GEMOMA
Requires proteins and reference annotations from closely related genome/species, optionally uses short read RNA-Seq alignments.
URL: https://github.com/Jstacs/Jstacs/tree/master/projects/gemoma
Publication: https://doi.org/10.1007/978-1-4939-9173-0_9
EVIANN
Requires proteins and reference annotations from closely related genome/species, optionally uses short read RNA-Seq and/or IsoSeq data.
URL: https://github.com/alekseyzimin/EviAnn_release/
Publication: https://doi.org/10.1101/2025.05.07.652745
EASEL
Combined structural/functional annotation; requires short read (or IsoSeq) RNA-Seq alignments; pre-trained models for plants, invertebrates, and vertebrates, designed for complex genomes.
URL: https://gitlab.com/PlantGenomicsLab/easel
Webinar link: https://www.youtube.com/watch?v=01ZCy5f72F8
EGAPX
Supports chordates, arthropods, echinoderms, cnidaria, monocots, dicots; also adds functional annotation (product names) to gene structures.
Protein-coding gene set combiners
TSEBRA
Combines gene sets from BRAKER, Tiberius, and GALBA.
URL: https://github.com/Gaius-Augustus/TSEBRA
Publication: https://doi.org/10.1186/s12859-021-04482-0
Protein-coding gene predictors (ab initio)
HELIXER
Models are available for: land plants, fungi, vertebrates, and invertebrates; suggest to use in combination with other tools.
URL: https://github.com/weberlab-hhu/Helixer
Publication: https://doi.org/10.1038/s41592-025-02939-1
Web interface: https://www.plabipd.de/helixer_main.html
ANNEVO
Models are available for: embryophyta, fungi, insecta, invertebrates, mammalia, and other vertebrates.
URL: https://github.com/xjtu-omics/ANNEVO
Publication: https://doi.org/10.21203/rs.3.rs-6402260/v1
ORIONGENO
Models are available for vertebrates (mammals, fish, birds, and others), invertebrates (arthropods and others), plants (angiosperms and bryophytes), fungi, and some algae and protists.
URL: https://github.com/BGIResearch/OrionGeno
Publication: https://doi.org/10.64898/2026.04.26.720859
TIBERIUS
Models are available for: diatoms, eudicotyledons, fungi, insecta, mammalia, monocotyledons, vertebrates, and chlorophyta.
URL: https://github.com/Gaius-Augustus/Tiberius
Publication: https://doi.org/10.1093/bioinformatics/btae685
Annotation by projection
LIFTOFF
Annotation projection from “closely” related species.
URL: https://github.com/agshumate/Liftoff
Publication: https://doi.org/10.1093/bioinformatics/btaa1016
Lifton
Annotation projection from the same or different species.
URL: https://github.com/Kuanhao-Chao/LiftOn
Publication: https://doi.org/10.1101/2024.05.16.593026
CESAR2
Requires the regions to be pre-aligned. Can work at large distances.
URL: https://github.com/hillerlab/CESAR2.0
Publication: https://doi.org/10.1093/bioinformatics/btx527
CAT
Good for smaller distances (e.g., primates).
URL: https://github.com/ComparativeGenomicsToolkit/Comparative-Annotation-Toolkit
Publication: https://doi.org/10.1101/gr.233460.117
Functional annotation
Interproscan
For predicting domain architectures.
URL: https://github.com/ebi-pf-team/interproscan
Publication: https://doi.org/10.1093/bioinformatics/btu031
entap
Sequence similarity, gene family, ontology, protein domain, contaminant and HGT detection.
URL: https://gitlab.com/PlantGenomicsLab/EnTAP
Publication: https://doi.org/10.1111/1755-0998.13106
Eggnog-mapper
Fast orthology-based functional annotation that complements InterProScan.
URL: https://github.com/eggnogdb/eggnog-mapper
Publication: https://doi.org/10.1093/molbev/msab293
Fantasia
For predicting GO terms, when used in combination with topGO. Very fast on GPU (e.g. a few minutes on proteome of one species).
URL: https://github.com/MetazoaPhylogenomicsLab/FANTASIA
Publications: https://doi.org/10.1093/nargab/lqae078 and https://doi.org/10.1101/2024.02.28.582465
Non-coding gene predictors
RFAM + Infernal cmsearch
For predicting short non-coding RNAs.
RFAM URL: https://rfam.org/ and publication: https://doi.org/10.1093/nar/gkaa1047
INFERNAL URL: http://eddylab.org/infernal/ and publication: https://doi.org/10.1093/bioinformatics/btt509
Tranascan-SE
For predicting tRNAs. Has a eukaryotic high confidence filter that can be applied to significantly reduce false positives.
URL: https://github.com/UCSC-LoweLab/tRNAscan-SE
Publication: https://doi.org/10.1007/978-1-4939-9173-0_1
Annotation QC
Omark
For assessing the annotation quality of predicted proteomes, including completeness, overprediction, and contamination checks.
URL: https://github.com/DessimozLab/OMArk
Publication: https://doi.org/10.1038/s41587-024-02147-w
BUSCO
For assessing the completeness of annotated gene sets using conserved single-copy orthologs.
URL: https://busco.ezlab.org
Publication: https://doi.org/10.1093/nar/gkae987
GFFcompare
For comparing multiple structural annotations.
URL: https://github.com/gpertea/gffcompare
Publication: https://doi.org/10.12688/f1000research.23297.2
AGAT
For checking, fixing, and standardising GFF/GTF annotation files before downstream use.
URL: https://github.com/NBISweden/AGAT
About the Subcommittee
This Report on Annotation Tools Recommendations was developed by EBP’s Scientific Subcommittee for Annotation.