IT and Informatics Standards
Version 1.0—March 2021
Authors: Vincent (Xiaofeng) Wei, Monica Poelchau, Rob Davey, Kim Pruitt, Nick Salmon, Keith A. Crandall, Juan Carlos Castilla Rubio, Stephen Richards
A. Overview:
Coordinating the collection, integration, standardization, analysis, archiving and sharing of genetic data and related metadata is the core mission of Earth Biogenome Project (EBP). The EBP aims to undertake systematic sequencing and analysis of genomes from all known species, through collaboration with different institutes or groups around the world. Thus, EBP serves as a platform to connect biologists/genomicists within the EBP community and other stakeholders to foster collaboration and data sharing, in line with access and benefit sharing of the Convention of Biological Diversity Nagoya Protocol and related international agreements and regulations (see Access and Sharing Policies section). In this regard, the IT/informatics subcommittee needs to define minimum requirements, standard operating procedures (SOPs), recommended practices, and infrastructure requirements to support data and metadata handling, whether it is the collection of voucher specimen information, documenting the processes of data analysis, or providing access to data outputs. These informatics standards facilitate the EBP and its members to contextualise and reuse the genomic and associated meta-data collected as part of this international effort. By following standards, we expect to develop a model and build a platform which covers the infrastructures, production systems, archive repositories and analysis pipelines to manage the overall progress and data of the EBP more efficiently, to comply with the data sharing regulations and/or laws of different countries and political entities, and to promote sharing, mining and application of the data.
Figure 1. Establishing an IT/informatics standard framework with best practices and recommended resources to support the project networks of EBP and other subcommittees.
A1. Access and Sharing Policies
The recommendation is to support the FAIR (Findability, Accessibility, Interoperability, and Reuse) principles) and place no restrictions on data access by submitting consensus reference genomes and corresponding sample metadata to INSDC (International Nucleotide Sequence Database Collaboration) managed data repositories. However, considering the Nagoya Protocol or related regulations in different countries and political entities, under certain circumstances the submitted genomics data may include a comment that links back to the submitters website for more information about restrictions/licenses but the comment may not include the license information itself. INSDC data repositories do not enforce licenses. In-country data generators are encouraged to work with an INSDC member to establish a workflow to broker submissions. The EBP recognizes that some researchers need to abide by nation-specific data policies and may need to establish local plans for how to meet the goals of data sharing.
A2. Data and Metadata Standards
In the archiving and sharing of omics data, most of the international public archiving systems and databases, such as those managed by the INSDC, already reflect community data format standards through the archiving process, data structure and data management practices. Genome sequence, annotation, and sample information standards should be compatible with those already established by the INSDC. Considering that the field of biodiversity research involves different species groups, these existing data schemas can be extended by working with INSDC members to supplement the existing standards if needed.
A3. Analysis and Pipelines
Considering different infrastructure capabilities and data analysis running environments across EBP partners, analysis pipelines will inevitably be varied, which may directly affect reproducibility. We should establish strategies for standardization and sharing of pipelines using descriptive and controlled workflow languages, versioning systems, and code repositories.
A4. Portal and User Services
EBP includes different research communities who may be using different standards or specifications or tools for their data analysis pipelines, data schemas, data archiving, and data sharing. A EBP central user services portal will provide useful services to these research communities in tracking and sharing this information, thereby making it easier for researchers to manage and use data files and pipelines.
B. Best Practices and Recommended Resources:
B1. Access and Sharing Policies
Compared to human genetic resources, there are fewer restrictions on data sharing of biodiversity genomic data, but some countries and regions have specific biodiversity protection or management laws and regulations. Therefore, under the premise of aligning to the Nagoya Protocol where appropriate, EBP can provide a more inclusive design to promote the data sharing of the entire project.
B.1.1. Repositories
The EBP involves many types of metadata associated with the targeted genomic data, including information on the biological sample, sample vouchers, and even images. Although most of this can be archived by using existing international public biological repositories and the management systems of biobanks and museums, it is still necessary to clarify what types of data we need to store at every stage. Data (such as sequences and images) and metadata that are submitted to a repository should result in a persistent identifier for public citation (e.g., accession, digital object identifier (DOI)).
B.1.2. Sharing Policies and Licenses
Figure 2. If in future EBP has its own data service portal, researchers can easily choose the appropriate license for their own project based on suggestions made within the portal.
Unrestricted access to data and analysis tools and software is the main recommendation of the EBP. Much of the data and metadata generated by the EBP will be housed in INSDC repositories, where data sharing is subject to the INSDC data policy. In general, records in INSDC repositories are available for free and unrestricted re-use, while citing the original submission based on scientific best practices. Based on this policy, a specific separate license is not recommended for submissions to the INSDC. For submissions to other public repositories, we recommend the use of CC-0 or CC-BY. We acknowledge that this may not meet all of the situations and sometimes there are other factors that need to be considered. For example, the biodiversity law of Brazil allows conducting research with Digital Sequence Information (DSI) derived from genetic resources, but registration is required for publication or patent application for DSI users (*Silva MD, Oliveira DR. The new Brazilian legislation on access to the biodiversity (Law 13,123/15 and Decree 8772/16). Braz J Microbiol. 2018;49(1):1-4. doi:10.1016/j.bjm.2017.12.001). For these cases, other common licences are an option, such as the Creative Commons suite of licences (see Appendix 1). If in future EBP has its own data service portal, researchers can easily choose the appropriate license for their own project (Fig. 2) based on suggestions made within the portal.
C. DATA AND METADATA STANDARDS:
The Earth BioGenome Project will generate massive amounts of data across the tree of life. The EBP’s overarching goals require that these data are accessible and re-usable by both humans and machines (FAIR Guiding Principles for scientific data management and stewardship). With smaller datasets and analyses, researchers can pull the information that they need that is not described in the metadata record from additional text (usually associated publications, which are not always open-access). The large scale of the Earth BioGenome Project data will require public, preserved, structured, relevant, machine-readable metadata for effective re-use by the scientific community. The recommended data repositories in Table 1 provide mechanisms for metadata entry and storage. These general-use repositories necessarily require minimal metadata for the data types generated by a standard genome project; here, we encourage EBP members and affiliated project networks to provide a high standard of metadata beyond the minimum requirements.
We also recognize that collecting and reporting metadata can be onerous for the data submitter, in particular if it is dubious how and whether other scientists will (or even should) use it. How much metadata is enough? Will the required metadata sufficiently capture the complexity of the underlying methodology or experiment? In some projects, where the data are generated by only a few providers, rich structured metadata can be readily harvested; however, this becomes a challenge in a distributed project such as the EBP. Recognizing these challenges, we provide concrete and practical recommendations that we hope will allow a consistent and reasonable approach to metadata across large swaths of the project. Each affiliated project network should strive to implement a consistent approach to metadata within their network. See here for an example implementation of genome metadata.
While we anticipate many different use cases for the EBP data, we expect the primary data generators and users to be in the realms of bioinformatics and biodiversity informatics. This expectation influences our recommendations.
C1. Recommendations
Incorporate metadata collection and entry into your project planning.
Why? Experimental details are more likely to be remembered and recorded accurately when the experiment is performed, rather than after the fact.
How? Use a platform that brokers submissions to an INSDC repository. For example, the COPO project and CyVerse platform provide mechanisms for metadata entry. Sample- and project-level metadata can be collected (and published) early in the project.
Collect the metadata for your project with future scientists in mind.
Why?
The data generated by the Earth BioGenome Project should be re-used by future scientists. If relevant experimental details are only reported in published articles behind paywalls, then the data can’t be effectively re-used by all.
How?
Sample metadata: We are in the process of generating a EBP checklist for sample metadata, based on The Darwin Tree of Life Project’s metadata entry forms. These are designed to capture the most relevant metadata for this large-scale project, and following their templates will help ensure that your metadata is reasonably well scoped. However, if your metadata needs for your project differ substantially from the EBP recommendations, other INSDC sample checklists are perfectly acceptable, as long as as many metadata fields are completed as is possible and reasonable.
Sequence: Submission of sequence data to any INSDC repository should provide sufficient access to the data. Use of data brokers for submission to INSDC repositories is also acceptable. The sequence submission should be linked to the Earth BioGenome umbrella BioProject number, PRJNA533106 and, as relevant, to an additional BioProject representing an affiliated network. Sequence, assembly, and sample data should be submitted to the same INSDC center; e.g., do not submit assembly and sample metadata to one INSDC center and the assembled genome to a different center..
Assembly: Submission of genome assembly data to any INSDC repository should provide sufficient access to the data. The assembly submission should be linked to the Earth BioGenome umbrella BioProject number, PRJNA533106. Some projects may choose to provide additional information, structured or unstructured, about the assembly analysis details in other repositories, such as Zenodo. When choosing a repository, consider the longevity of the repository, and whether the metadata can be bidirectionally linked to the assembled genome sequence in INSDC and accessed by humans and machines.
Variants: Genetic variant data can be submitted to a center that brokers submissions to INSDC or directly to EMBL-EBI or DDBJ as NCBI no longer accepts submissions of non-human variant data. The preferred format for variant data is VCF (variant call format).
Project: Submission of project or study data to any INSDC repository should provide sufficient access to the data. We recommend generating an umbrella project, under the EBG umbrella project, for each EBP affiliated network.
Voucher: We recommend that EBP projects submit voucher data and metadata to an appropriate national repository (e.g., GGBN member repositories) and include voucher IDs when submitting sample metadata. We will supply more specific recommendations in phase II.
Image: We recommend the Audubon core standard for multimedia resources and Zenodo for archiving the images and receiving DOIs.
Given the existing fields in the metadata entry form, be as complete as possible.
Why? Attaching rich, structured metadata to a dataset is one of the best ways to make your data re-usable by others.
How?
Decide on your metadata forms in advance, and collect the metadata as you retrieve it. If you are working with a commercial sequence provider, ask in advance whether they can provide you with the necessary standard of metadata.
Use ontology or controlled vocabulary terms for your entries to the extent possible. EBI’s ontology lookup service is a great resource for this. It is often helpful for EBP networks to discuss and recommend ontologies that can be used for a particular domain or taxonomic group, given the large number of options available.
If you cannot supply sufficient information in the designated metadata fields, use a description field. In some cases, the existing repositories do not allow for sufficient granularity of your metadata for others to re-use it, i.e. the “bare minimum”. In this case, repositories will often provide additional fields where free-text is allowed. Providing detailed information there will allow future scientists to better understand your data, without having to access a manuscript. Again, it is helpful to include ontology terms or descriptions in these fields, for example the DTOL metadata schema suggests the use of the Environment Ontology (ENVO) for describing the habitat of collected organisms.
D. ANALYSIS AND PIPELINE:
D.1. IT Infrastructure
The EBP infrastructure requirements are large, not only for the final genome assembly data that needs to be archived, but also the storage and computing resources needed for intermediate analyses. We recommend the development of a mechanism to share the existing infrastructure capabilities for each affiliated project (Fig. 3).
D.1.1. Local Resources
When considering the sharing of IT hardware resources in the various EBP affiliated projects, we recommend aggregating links to the different infrastructure resources which can be shared into a “hardware resources list”, and keep it regularly updated. If any groups encounter the lack of hardware resources which cannot be carried out, then they can find the resources from the “list” and use the internet or expressway to transfer data to other centers for analysis.
Figure 3. If in future EBP has its own data service portal, researchers can easily apply for the hardware resources in a “List” which are shared by the EBP affiliated project networks .
D.1.2. Cloud Resources
Commercial Cloud
If local and shared computing resources cannot meet the demand of EBP computing needs, elastic commercial cloud resources can be a suitable choice to complement local computing resources and provide on-the-fly expanded resources for targeted use. Cloud resources can be purchased on demand and expanded elastically, though the overall cost can be higher than local compute and storage. There are low-cost services such as offline storage, e.g., Amazon Glacier, but transmission of data into and out of these services can incur a significant cost. If your requirement for resources is temporary and your analysis workflows are well understood, cloud services are suitable and much cheaper than building new hardware facilities due to their pay-for-what-you-use pricing structure.
With respect to security, cloud services, such as Amazon Web Services (AWS),Microsoft Azure, and Google Cloud Platform (GCP) are designed with high security in mind and are kept iteratively updated. However, users of cloud computing service provider’s should confirm that such use is in line with the laws and regulations of the countries or regions from which the data are generated. Similarly, security of cloud resources supplied by these vendors is often left to individual users, so care should be taken to ensure data integrity and system security.
Self-Built Cloud
For institutions or groups with an IT team and suitable underlying hardware in managed data centers, an open source cloud framework like Openstack can be used to build a cloud service that can then be supplied to fulfill the same requirements listed above.
D.2. Security (Backup)
Large international omics data repositories, such as INSDC and CNGB, already have mechanisms for data backup and security. Unless there are extenuating factors, there will be no data loss. But many times, when data are produced, they may not be archived into the data submission repositories in a timely fashion, and often there are no data backups locally. Considering the huge cost of genome sequencing projects in the EBP, it is required to archive data as soon as possible (Fig. 4).
At present, if a scientific research project, e.g., ICGC, can be added to the scientific data lists of some cloud service providers such as the Open Data project of AWS, the space for storing data can be free of charge.
If not, this will need to be discussed within the EBP partner project meetings to set out processes for data security and backups.
Figure 4. The large international omics data repositories, like INSDC, are the first choices for data backup.
D.3. Code Repository
Although open access data are the basis of ‘omics research, the repeatability of data analysis pipelines is also critical for the accuracy and reliability of final research conclusions. In addition to the differences in experimental methods and sequencing platforms, the differences between analysis pipelines, algorithms and parameters directly affect the accuracy and quality of the data, and downstream interpretation.
We recommend that each EBP project establish a publicly accessible code repository, which can be Zenodo, or Github/Gitlab, to version, archive and manage the code of the analysis pipelines (Fig. 5).
Figure 5. The management and sharing of the pipelines. Note: The standardized and dockerized method by WDL and Docker are not mandatory, but recommended options.
D.4. Workflows
Data processing workflow languages, like Workflow Description Language (WDL), can be used to standardize pipelines. By storing this descriptive specification about the tools used to complete an analysis in a suitable repository, we can ensure that the pipelines are easier to be used and maintained. For example, protocols.io can house descriptions of laboratory or analytical processes to aid reproducibility. Actual workflow description files can be stored and versioned in code repositories such as Github or GitLab. Therefore, we recommend that assembly and variant calling pipeline tools, parameters, and any called scripts are standardised and shared accordingly.
D.5. Containerisation
To aid reproducibility and reuse of data, EBP may provide online analysis tools and infrastructure to ensure the consistency of redeployment of analysis pipelines. There are suitable container technologies available that we can use to provide consistent pipelines in the later phases. We recommend Docker or Singularity images to contain any scripts, software and workflows. Bear in mind that DockerHub has recently stated they will delete Docker images that have not been used or updated within a 6 month time frame, so other mechanisms of persisting and sharing containers may be required.
E. PORTAL AND USER SERVICES:
E.1 Design - The Data Services Portal
There are many biodiversity networks or groups under the EBP. Each group has its own strategy, which involves the design of scientific research topics, application of funding, selection of samples, selection of sequencing platforms, data analysis and data sharing. If we want to follow the progress of the different affiliated projects, we can not get the details from different data repositories, coordinate different IT infrastructure resources, common data/sample processing issues easily.
So a unified portal is necessary, though it needs to be designed and developed. Below are some basic functions and ideas for such a portal (Fig. 6.):
Figure 6. The basic design of the EBP data portal (which is in need of development) and its basic services.
E.1.1. Progress
Registers can update the progress of their projects, and every register can check the progress of all projects, including active querying of particular species and target species overviews for the entire project.
E.1.2. Exploration
Since different groups in EBP deposit data into different repositories, like NCBI SRA, EBI ENA or local storage, an EBP portal should allow for linkages across storage locations for quick and unified access to genomic data across the EBP. Such access and linkage would require a minimal metadata set that can be established based on the pre-defined vocabularies of different data repositories and indexed by the query engine.
E.1.3. Tools
While it would be impractical to store tools for genomic research in the EBP portal, we can provide links to effective tools and registered pipelines for analysis consistency and repeatability. Users can directly click and jump to the tool home pages through the outlink in this portal.
E.1.4. Brokering
In addition to the basic functions mentioned above, if we can design a component based on an integrated and comprehensive EBP data standard, at the same time collect the required information of each process in the production system (sample management system or lab management system), and also can be an agent or gateway to broker the data to EBI or NCBI, which would be very useful for the EBP and other genome research projects. The DTOL project has started such an integrated systematic approach involving open source data and metadata management tools as well as internal LIMS integration, and this model could provide insights into future EBP partner projects and their implementation of these recommendations.
E.2 Best Practices - Sample Tracking Systems
E.2.1. Introduction
EBP partners will need to consider how they track the progress of samples as they move through the genomics pipeline. This has two main purposes:
It allows the EBP partner to manage their own sample management, sequencing and assembly processes efficiently;
It generates standard metrics and progress updates that can be escalated to EBP and combined with similar data from other partners to provide an overall assessment of EBP progress.
E.2.2. EBP Progress Metrics
To comply with EBP standards, all partners should be capable of reporting monthly on core metrics for EBP genomics (Table 2).
Table 2. Core metrics for reporting across the EBP for monthly progress reports.
The EBP portal should include a front-end Dashboard with live, updated summary statistics across the metrics summarized in Table 2. This allows for accurate updating across the project and helps avoid duplicative genome sequencing efforts. Duplicate genomes should be allowed where appropriate metadata warrant such efforts.
E.2.3. Case Study - Wellcome Sanger Institute’s Tree of Life Programme
The Wellcome Sanger Institute’s (Sanger) Tree of Life Programme (ToL) is leading two EBP-affiliated projects: 1) the Darwin Tree of Life (DToL) project which aims to sequencing all 60k identified eukaryotic species in Britain and Ireland; and, 2) the Aquatic Symbiosis Genomics (ASG) project, which aims to sequence 500 symbiotic systems (>1000 genomes) from across the world’s freshwater and marine ecosystems. The collection of samples is through biodiversity partners, and DNA extractions are carried out within ToL. ToL works with the Sanger Institute’s Scientific Operations core to make sequencing libraries and generate raw sequence data. These data are assessed, assembled and curated by ToL teams before assemblies are submitted to the European Nucleotide Archive (ENA) and the International Nucleotide Sequence Database Consortium (INSDC). To support the receipt, storage and management of the tens of thousands of samples for these projects ToL has launched an IT systems project called Sample Tracking Systemisation (STS). The goals of STS are:
Enable easy entry and storage of sample metadata against defined schema;
Support sequence data submission and metadata brokering with genome repositories;
Track the status of every ToL sample at any stage of the pipeline;
Manage multiple identifiers for each sample and track relationships between samples;
Link samples and metadata to quality control information (e.g., DNA extraction yield);
Automate decision making and processing based on metadata wherever possible;
Enable clear reporting of pipeline and project metrics;
Provide easy access to ToL data and pipeline metrics for partners.
A discovery phase for this project was run between May-July 2020 and the project is now in build phase with an estimated initial go-live of the end of October 2020.
STS is being built using common open source tools including:
Web applications using the React Javascript framework
Python/Django for server-side code
RESTful APIs
Databases using PostgreSQL
Message/event queues using RabbitMQ/Kafka
Openstack
STS also uses COPO (originally Collaborative Open Plant Omics portal, built by Rob Davey’s group at the Earlham Institute) to provide sample metadata collation, validation and data brokering support.
High level roadmap for STS:
Phase 1 (end Oct 2020)
Go live with support for sample metadata ingestion via COPO and storage in the STS database, plus early sample management, compliance management and work request generation. Go live of DToL public portal including project metrics and sample look-up.
Phase 2 (end Jan 2021)
Go live with support for production and R&D including laboratory information management systems (LIMS), data warehouse and electronic lab notebook integration. Will deliver sample tracking capability throughout sample management and production. Phase 2 will also deliver continuous improvements to the sample management capability provided in Phase 1.
Phase 3 (end Apr 2021)
Go live with support for assembly and curation including integration with informatics management tools and databases. Will deliver sample tracking capability throughout the entire Sanger Institute pipeline. Phase 3 will also deliver continuous improvements to the capabilities provided in Phase 1 and 2.
F. Appendix
Reference: Licenses
In EBP, we choose INSDC as the default repositories for researchers to share their datasets. But in order to ensure that the researchers get credit for the data or softwares, and follow Nagoya Protocol or related regulations in different countries and political entities, publishing without restriction is not always ideal. We have also enabled the option to publish under other licenses.
The following table can bring you some suggestions.
ABOUT THE SUBCOMMITTEE
This Report on IT and Informatics Standards was developed by EBP’s Scientific Subcommittee for IT and Informatics.