UNIT 17 · Case Study
mDeepFRI
Metagenomic-DeepFRI is a high-performance pipeline for annotating protein sequences with Gene Ontology (GO) terms using DeepFRI.
A pipeline for annotation of genes with DeepFRI, a deep learning model for functional protein annotation with Gene Ontology (GO) terms and mapping to Cluster of Orthologous Groups (COG) categories. It incorporates FoldComp databases of predicted protein structures for fast annotation of metagenomic gene catalogues.
🔍 Overview
Metagenomic-DeepFRI is a high-performance pipeline for annotating protein sequences with Gene Ontology (GO) terms using DeepFRI, a deep learning model for functional protein annotation.
Protein function prediction is increasingly important as sequencing technologies generate vast numbers of novel sequences. Metagenomic-DeepFRI combines:
- Structure information from FoldComp databases (AlphaFold, ESMFold, PDB, etc.)
- Sequence-based predictions using DeepFRI’s neural networks
- Fast searches with MMseqs2 for database alignment
- Significant speedup of 2-12× compared to standard DeepFRI implementation.
📋 Pipeline stages
- Search proteins similar to query in PDB and supply
FoldCompdatabases withMMseqs2. - Find the best alignment among
MMseqs2hits usingPyOpal. - Align target protein contact map to query protein with unknown structure.
- Run
DeepFRIwith the structure if found in the database, otherwise runDeepFRIwith sequence only.