Published January 1, 2018 | Version v1
Journal article Open

Large-scale automated function prediction of protein sequences and an experimental case study validation on PTEN transcript variants

  • 1. Istanbul Tech Univ, Dept Comp Engn, TR-34467 Istanbul, Turkey
  • 2. Middle East Tech Univ, Grad Sch Informat, CanSyL, TR-06800 Ankara, Turkey
  • 3. European Bioinformat Inst, European Mol Biol Lab, Prot Funct Dev Team, Cambridge CB10 1SD, England
  • 4. Middle East Tech Univ, Dept Comp Engn, TR-06800 Ankara, Turkey

Description

Recent advances in computing power and machine learning empower functional annotation of protein sequences and their transcript variations. Here, we present an automated prediction system UniGOPred, for GO annotations and a database of GO term predictions for proteomes of several organisms in UniProt Knowledgebase (UniProtKB). UniGOPred provides function predictions for 514 molecular function (MF), 2909 biological process (BP), and 438 cellular component (CC) GO terms for each protein sequence. UniGOPred covers nearly the whole functionality spectrum in Gene Ontology system and it can predict both generic and specific GO terms. UniGOPred was run on CAFA2 challenge target protein sequences and it is categorized within the top 10 best performing methods for the molecular function category. In addition, the performance of UniGOPred is higher compared to the baseline BLAST classifier in all categories of GO. UniGOPred predictions are compared with UniProtKB/TrEMBL database annotations as well. Furthermore, the proposed tool's ability to predict negatively associated GO terms that defines the functions that a protein does not possess, is discussed. UniGOPred annotations were also validated by case studies on PTEN protein variants experimentally and on CHD8 protein variants with literature. UniGOPred protein functional annotation system is available as an open access tool at .

Files

bib-7406bef7-9f79-491e-8de7-7ea8ca748cb5.txt

Files (308 Bytes)

Name Size Download all
md5:9ae7c8a129fe61ba4a7dd73988c67601
308 Bytes Preview Download