Introduction
BIOTA is a unified and structured collection of omics data from official open European EMBL-EBI and NCBI biological databases.
It is dedicated to Gencovery Web Services (GWS) for the conception and use of digital twins of cell metabolism, and more broadly for omics data integration and analysis.
All data distributed in BIOTA are provided in their original versions without alteration by the Gencovery team. Users can verify the reliability of the data through the cited data providers and related scientific publications.
📚 BIOTA at a Glance
BIOTA is organized into two main sections:
🧩 Ontology Database
A collection of structured biological vocabularies used to describe and classify biological information.
🧪 Molecular Database
A collection of compounds, enzymes, proteins, genes, and metabolic reactions classified using ontology data.
📊 Ontology Database
🔬 Molecular Database
🧩 Ontology Databases
BIOTA's ontology database contains controlled biological terms used to describe biological information.
Ontologies define:
- hierarchical relationships between terms;
- biological semantics;
- standardized descriptions for biological data;
- shared vocabularies that support data analysis and interpretation.
🧬 Gene Ontology
The Gene Ontology (GO) is a major bioinformatics initiative designed to unify the representation of gene and gene-product attributes across all species.
🌐 Source: geneontology.org📄 License: Creative Commons CC BY 4.0
⚙️ Systems Biology Ontology
The Systems Biology Ontology (SBO) provides controlled and relational vocabularies commonly used in systems biology and computational modelling.
It helps standardize the description of models and biochemical experiments, facilitating their interpretation and reuse.
🌐 Source: EMBL-EBI SBO📄 License: Artistic License 2.0
🔎 Evidence and Conclusion Ontology
The Evidence and Conclusion Ontology (ECO) contains terms describing types of evidence and assertion methods.
These terms are commonly used in biological curation to describe the evidence supporting biological assertions.
🌐 Source: Evidence Ontology📄 License: Creative Commons CC0 1.0 Universal
🧫 BRENDA Tissue Ontology
The BRENDA Tissue Ontology (BTO) provides a structured vocabulary describing biological sources of enzymes.
It includes terms for:
- tissues;
- organs;
- anatomical structures;
- cell types;
- cell lines;
- cell cultures;
- plant structures.
It covers organisms from all taxonomic groups, including animals, plants, and fungi.
🌐 Source: EMBL-EBI Ontology Lookup Service📄 License: Creative Commons CC BY 4.0
🌳 NCBI Taxonomy
The NCBI Taxonomy Database provides curated classification and nomenclature for organisms represented in public sequence databases.
NCBI molecular databases include resources covering:
- nucleotide sequences;
- protein sequences;
- macromolecular structures;
- molecular variation;
- gene expression;
- mapping data.
🌐 Source: NCBI Taxonomy
NCBI places no restrictions on the use or distribution of the data contained in these resources.
🗺️ Pathways
Pathway information in BIOTA is collected from:
Reactome
Reactome is a free, open-source, curated, and peer-reviewed pathway database.
It supports the visualization, interpretation, and analysis of pathway knowledge for:
- basic research;
- genome analysis;
- modelling;
- systems biology;
- education.
BKMS
BKMS-react is an integrated and non-redundant biochemical reaction database.
It combines biochemical reactions collected from:
- BRENDA;
- KEGG;
- MetaCyc;
- SABIO-RK.
These reactions are integrated by matching their substrates and products.
⚗️ Enzyme Classification
Enzyme classification data are collected from the Expasy-ENZYME database.
The classification is based on the Enzyme Commission number (EC number), a numerical system that classifies enzymes according to the chemical reactions they catalyze.
BIOTA contains the enzyme functional-class hierarchies to help users interpret enzyme-related data even when the exact enzyme classification is not known.
🌐 Source: Expasy-ENZYME📄 License: Creative Commons CC BY 4.0
🧪 Molecular Database
BIOTA's molecular database contains biological entities found in living organisms, including:
- compounds and metabolites;
- enzymes;
- proteins;
- genes;
- metabolic reactions.
These data are collected and structured from several open databases to support the conception of digital twins of cell metabolism.
🧪 Metabolic Compounds
Metabolic compound data are collected from ChEBI — Chemical Entities of Biological Interest.
ChEBI is a freely available dictionary focused on small molecular entities.
These entities include:
- naturally occurring compounds;
- synthetic compounds used to intervene in biological processes.
Genome-encoded macromolecules such as nucleic acids and proteins are generally not included.
🌐 Source: ChEBI📄 License: Creative Commons CC BY 4.0
⚙️ Enzymes
Enzyme data are based primarily on the European reference databases BRENDA and Expasy.
They provide information on known enzymes across living organisms, including:
- classifications;
- gene sequences;
- kinetic parameters;
- environmental conditions;
- biological origin.
🔗 Enzyme vs. Enzyme-Ortholog
BIOTA distinguishes between two concepts:
Enzyme Orthologs are useful when working with enzyme functional classes.
Enzymes provide more detailed organism-specific information, including taxonomy and protein sequence information.
BIOTA currently contains approximately:
- 6,900 Enzyme-Orthologs
- 112,000 Enzymes
🧪 BRENDA Database
BRENDA is a major collection of enzyme functional data available to the scientific community.
For characterized enzymes associated with an EC number, BRENDA can provide information including:
- tissue and cellular localization;
- kinetic parameters;
- environmental conditions such as pH;
- cofactors;
- genetic sequences in FASTA format.
Because EC numbers describe enzyme-catalyzed reactions rather than individual enzymes, a single reaction can be related to several EC numbers.
🌐 Source: BRENDA📄 License: Creative Commons CC BY 4.0
🌐 Expasy Database
Expasy is the bioinformatics resource portal of the SIB Swiss Institute of Bioinformatics.
It provides access to more than 160 databases and software tools supporting areas such as:
- genomics;
- proteomics;
- structural biology;
- evolution and phylogeny;
- systems biology;
- medicinal chemistry.
Expasy-ENZYME provides information related to enzyme nomenclature and EC-number classification.
🌐 Source: Expasy📄 License: Creative Commons CC BY 4.0
🔗 Enzyme Orthologs — Enzo
The Enzyme-Ortholog (Enzo) concept is not a standard concept in biology or bioinformatics.
It was introduced in BIOTA as a concept analogous to KEGG Orthologs, allowing enzymes to be uniquely referenced according to their EC-number characteristics independently of the organism.
Enzos are particularly suited to:
- characterize metabolic pathways;
- work with functional enzyme information;
- support fast reconstruction of metabolic models.
This concept is used, for example, in the work of Tabei and co-authors, Bioinformatics, 2016.
🔄 Metabolic Reactions
Metabolic reaction data are provided by Rhea.
Rhea is an expert-curated knowledgebase of chemical and transport reactions of biological interest.
It uses chemical and biological information from resources including:
- ChEBI;
- UniProtKB;
- InChIKey;
- Gene Ontology.
Rhea is connected to BRENDA and Expasy through enzyme EC numbers.
🌐 Source: Rhea📄 License: Creative Commons CC BY 4.0
📊 Current Reaction Data
The substrate, product, and enzyme figures represent relationship links, not unique biological entities.
The same compound or enzyme may therefore be counted several times when linked to multiple reactions.
🧬 Proteins
Protein data are collected from UniProtKB.
The UniProt Knowledgebase is a central resource for functional protein information and provides rich and consistent biological annotation.
Its records include information such as:
- amino acid sequences;
- protein names and descriptions;
- taxonomic information;
- citation information.
BIOTA contains the manually reviewed and annotated UniProtKB records from Swiss-Prot.
UniProt is updated approximately every eight weeks, although BIOTA is currently not updated at the same frequency.
🌐 Source: UniProtKB / Swiss-Prot📄 License: Creative Commons CC BY 4.0
🌐 Main Data Sources
BIOTA brings together data from several major biological resources:

⚖️ Notice
Gencovery Numerical Resources (GNR) refer to the software, libraries, and data provided through Gencovery web services.
GNR may be covered by third-party licenses.
Gencovery guarantees that GNR are accessible for commercial and non-commercial use through Gencovery web services.
For ad-hoc use of GNR outside Gencovery web services, users should verify applicable third-party licenses to ensure that they are legally authorized to use the corresponding resources.
Gencovery does not warrant or assume legal liability or responsibility for the accuracy or completeness of information disclosed through Gencovery web services.
This section is not a legal notice. Please refer to the Gencovery terms of use for applicable legal information.