ABSTRACT
Background DisProt is the primary repository of Intrinsically Disordered Proteins. This database is manually curated and the annotations there have strong experimental support. Currently DisProt contains a relatively small number of proteins highlighting the importance of transferring verified disorder and other annotations, in such a way as to increase the number of proteins that could benefit from this valuable information. While the principles and practicalities of homology transfer are well-established for globular proteins, these are largely lacking for disordered proteins.
Methods We used DisProt to evaluate the transferability of the annotation terms to orthologous proteins. For each protein, we looked for their orthologs, with the assumption that they will have a similar function. Then, for each protein and their orthologs we made multiple sequence alignments (MSAs). Global and regional quality of the MSAs was evaluated with the NorMD score.
Results We have designed a pipeline to obtain good quality MSAs and to transfer annotations from any protein to their orthologs. Applying the pipeline to DisProt proteins, from the 1931 entries with 5,623 annotations we can reach 97,555 orthologs and transfer a total of 301,190 terms by homology. We also provide a web server for consulting the results of DisProt proteins and execute the pipeline for any other protein. The server Homology Transfer IDP (HoTIDP) is accessible at http://hotidp.leloir.org.ar.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
Elizabeth Martínez Pérez emartinez{at}leloir.org.ar
Mátyás Pajkos matyas.pajkos{at}ttk.elte.hu
Silvio C.E. Tosatto silvio.tosatto{at}unipd.it
Toby J. Gibson toby.gibson{at}embl.de
Zsuzsanna Dosztányi zsuzsanna.dosztanyi{at}ttk.elte.hu
Cristina Marino Buslje cmb{at}leloir.org.ar