Quantitative modeling of transcription factor binding specificities using DNA shape-Reference-Cited by-同舟云学术

Quantitative modeling of transcription factor binding specificities using DNA shape

Published:2015-03-09 Issue:15 Volume:112 Page:4654-4659
ISSN:0027-8424
Container-title:Proceedings of the National Academy of Sciences
language:en
Short-container-title:Proc Natl Acad Sci USA

Author:

Zhou Tianyin,Shen Ning,Yang Lin,Abe Namiko,Horton John,Mann Richard S.,Bussemaker Harmen J.,Gordân Raluca,Rohs Remo

Abstract

DNA binding specificities of transcription factors (TFs) are a key component of gene regulatory processes. Underlying mechanisms that explain the highly specific binding of TFs to their genomic target sites are poorly understood. A better understanding of TF−DNA binding requires the ability to quantitatively model TF binding to accessible DNA as its basic step, before additional in vivo components can be considered. Traditionally, these models were built based on nucleotide sequence. Here, we integrated 3D DNA shape information derived with a high-throughput approach into the modeling of TF binding specificities. Using support vector regression, we trained quantitative models of TF binding specificity based on protein binding microarray (PBM) data for 68 mammalian TFs. The evaluation of our models included cross-validation on specific PBM array designs, testing across different PBM array designs, and using PBM-trained models to predict relative binding affinities derived from in vitro selection combined with deep sequencing (SELEX-seq). Our results showed that shape-augmented models compared favorably to sequence-based models. Although both k-mer and DNA shape features can encode interdependencies between nucleotide positions of the binding site, using DNA shape features reduced the dimensionality of the feature space. In addition, analyzing the feature weights of DNA shape-augmented models uncovered TF family-specific structural readout mechanisms that were not revealed by the DNA sequence. As such, this work combines knowledge from structural biology and genomics, and suggests a new path toward understanding TF binding and genome function.

Funder

HHS | National Institutes of Health

National Science Foundation

Publisher

Proceedings of the National Academy of Sciences

Subject

Multidisciplinary

Reference41 articles.

1. Transcriptional enhancers: from properties to genome-wide predictions

2. In pursuit of design principles of regulatory sequences

3. Absence of a simple code: how transcription factors read the genome

4. From the editors

5. Modeling the specificity of protein-DNA interactions

Cited by 217 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. DNA breathing integration with deep learning foundational model advances genome-wide binding prediction of human transcription factors;Nucleic Acids Research;2024-09-13

2. Geometric deep learning of protein–DNA binding specificity;Nature Methods;2024-08-05

3. DNA shape features improve prediction of CRISPR/Cas9 activity;Methods;2024-06

4. Construction of transcript regulation mechanism prediction models based on binding motif environment of transcription factor AoXlnR in Aspergillus oryzae;Journal of Bioinformatics and Computational Biology;2024-06

5. Interpretable Protein-DNA Interactions Captured by Structure-based Optimization;2024-05-27