Prediction of mono- and di-nucleotide-specific DNA-binding sites in proteins using neural networks-Reference-Cited by-同舟云学术

Prediction of mono- and di-nucleotide-specific DNA-binding sites in proteins using neural networks

Published:2009-05-13 Issue:1 Volume:9 Page:
ISSN:1472-6807
Container-title:BMC Structural Biology
language:en
Short-container-title:BMC Struct Biol

Author:

Andrabi Munazah,Mizuguchi Kenji,Sarai Akinori,Ahmad Shandar

Abstract

Abstract Background DNA recognition by proteins is one of the most important processes in living systems. Therefore, understanding the recognition process in general, and identifying mutual recognition sites in proteins and DNA in particular, carries great significance. The sequence and structural dependence of DNA-binding sites in proteins has led to the development of successful machine learning methods for their prediction. However, all existing machine learning methods predict DNA-binding sites, irrespective of their target sequence and hence, none of them is helpful in identifying specific protein-DNA contacts. In this work, we formulate the problem of predicting specific DNA-binding sites in terms of contacts between the residue environments of proteins and the identity of a mononucleotide or a dinucleotide step in DNA. The aim of this work is to take a protein sequence or structural features as inputs and predict for each amino acid residue if it binds to DNA at locations identified by one of the four possible mononucleotides or one of the 10 unique dinucleotide steps. Contact predictions are made at various levels of resolution viz. in terms of side chain, backbone and major or minor groove atoms of DNA. Results Significant differences in residue preferences for specific contacts are observed, which combined with other features, lead to promising levels of prediction. In general, PSSM-based predictions, supported by secondary structure and solvent accessibility, achieve a good predictability of ~70–80%, measured by the area under the curve (AUC) of ROC graphs. The major and minor groove contact predictions stood out in terms of their poor predictability from sequences or PSSM, which was very strongly (>20 percentage points) compensated by the addition of secondary structure and solvent accessibility information, revealing a predominant role of local protein structure in the major/minor groove DNA-recognition. Following a detailed analysis of results, a web server to predict mononucleotide and dinucleotide-step contacts using PSSM was developed and made available at http://sdcpred.netasa.org/ or http://tardis.nibio.go.jp/netasa/sdcpred/. Conclusion Most residue-nucleotide contacts can be predicted with high accuracy using only sequence and evolutionary information. Major and minor groove contacts, however, depend profoundly on the local structure. Overall, this study takes us a step closer to the ultimate goal of predicting mutual recognition sites in protein and DNA sequences.

Publisher

Springer Science and Business Media LLC

Subject

Structural Biology

Link

https://link.springer.com/content/pdf/10.1186/1472-6807-9-30.pdf

Reference42 articles.

1. Nadassy K, Wodak SJ, Janin J: Structural features of protein-nucleic acid recognition sites. Biochemistry 1999, 38: 1999–2017. 10.1021/bi982362d

2. Jones S, van Heyningen P, Berman HM, Thornton JM: Protein-DNA interactions: A structural analysis. J Mol Biol 1999, 287: 877–896. 10.1006/jmbi.1999.2659

3. Pabo CO, Nekludova L: Geometric analysis and comparison of protein-DNA interfaces: why is there no simple code for recognition? J Mol Biol 2000, 301: 597–624. 10.1006/jmbi.2000.3918

4. Kinney JB, Tkacik G, Callan CG Jr: Precise physical models of protein-DNA interaction from high-throughput data. Proc Natl Acad Sci 2007, 104(2):501–506. 10.1073/pnas.0609908104

5. Morozov AV, Havranek JJ, Baker D, Siggia ED: Protein-DNA binding specificity predictions with structural models. Nucl Acids Res 2005, 33(18):5781–5798. 10.1093/nar/gki875

Cited by 35 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Predictive modeling of moonlighting DNA-binding proteins;NAR Genomics and Bioinformatics;2022-10-06

2. Inadequacy of Evolutionary Profiles Vis-a-vis Single Sequences in Predicting Transient DNA-Binding Sites in Proteins;Journal of Molecular Biology;2022-07

3. Detailed profiling with MaChIAto reveals various genomic and epigenomic features affecting the efficacy of knock-out, short homology-based knock-in and Prime Editing;2022-06-30

4. DNAPred_Prot: Identification of DNA-Binding Proteins Using Composition- and Position-Based Features;Applied Bionics and Biomechanics;2022-04-13

5. Dissecting and predicting different types of binding sites in nucleic acids based on structural information;Briefings in Bioinformatics;2021-10-09