Assessing the limits of genomic data integration for predicting protein networks-Reference-Cited by-同舟云学术

Assessing the limits of genomic data integration for predicting protein networks

Published:2005-07 Issue:7 Volume:15 Page:945-953
ISSN:1088-9051
Container-title:Genome Research
language:en
Short-container-title:Genome Res.

Author:

Lu Long J.,Xia Yu,Paccanaro Alberto,Yu Haiyuan,Gerstein Mark

Abstract

Genomic data integration—the process of statistically combining diverse sources of information from functional genomics experiments to make large-scale predictions—is becoming increasingly prevalent. One might expect that this process should become progressively more powerful with the integration of more evidence. Here, we explore the limits of genomic data integration, assessing the degree to which predictive power increases with the addition of more features. We focus on a predictive context that has been extensively investigated and benchmarked in the past—the prediction of protein–protein interactions in yeast. We start by using a simple Naive Bayes classifier for integrating diverse sources of genomic evidence, ranging from coexpression relationships to similar phylogenetic profiles. We expand the number of features considered for prediction to 16, significantly more than previous studies. Overall, we observe a small, but measurable improvement in prediction performance over previous benchmarks, based on four strong features. This allows us to identify new yeast interactions with high confidence. It also allows us to quantitatively assess the inter-relations amongst different genomic features. It is known that subtle correlations and dependencies between features can confound the strength of interaction predictions. We investigate this issue in detail through calculating mutual information. To our surprise, we find no appreciable statistical dependence between the many possible pairs of features. We further explore feature dependencies by comparing the performance of our simple Naive Bayes classifier with a boosted version of the same classifier, which is fairly resistant to feature dependence. We find that boosting does not improve performance, indicating that, at least for prediction purposes, our genomic features are essentially independent. In summary, by integrating a few (i.e., four) good features, we approach the maximal predictive power of current genomic data integration; moreover, this limitation does not reflect (potentially removable) inter-relationships between the features.

Publisher

Cold Spring Harbor Laboratory

Subject

Genetics(clinical),Genetics

Reference64 articles.

1. Alberts, B. 2002. Molecular biology of the cell. Garland Science, New York.

2. Gene Ontology: tool for the unification of biology

3. Protein Structure Prediction and Structural Genomics

4. Structure and mechanism of DNA topoisomerase II

5. Bishop, C.M. 1995. Neural networks for pattern recognition. Clarendon Press, Oxford University Press, Oxford, UK.

Cited by 162 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Integration of data‐independent acquisition (DIA) with co‐fractionation mass spectrometry (CF‐MS) to enhance interactome mapping capabilities;PROTEOMICS;2023-05-05

2. Computational approaches for the design of modulators targeting protein-protein interactions;Expert Opinion on Drug Discovery;2023-02-23

3. Identification of copper-related biomarkers and potential molecule mechanism in diabetic nephropathy;Frontiers in Endocrinology;2022-10-18

4. Expanding interactome analyses beyond model eukaryotes;Briefings in Functional Genomics;2022-05-12

5. Identification of significant protein in protein-protein interaction of Alzheimer disease using top-k representative skyline query;Jurnal Teknologi dan Sistem Komputer;2021-04-24