Optimizing data integration improves Gene Regulatory Network inference in Arabidopsis thaliana-Reference-Cited by-同舟云学术

Optimizing data integration improves Gene Regulatory Network inference in Arabidopsis thaliana

Published:2023-09-29 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Cassan Océane^ORCID,Lecellier Charles-Henri^ORCID,Martin Antoine^ORCID,Bréhélin Laurent^ORCID,Lèbre Sophie^ORCID

Abstract

AbstractMotivationsGene Regulatory Networks (GRN) are traditionnally inferred from gene expression profiles monitoring a specific condition or treatment. In the last decade, integrative strategies have successfully emerged to guide GRN inference from gene expression with complementary prior data. However, datasets used as prior information and validation gold standards are often related and limited to a subset of genes. This lack of complete and independent evaluation calls for new criteria to robustly estimate the optimal intensity of prior data integration in the inference process.ResultsWe address this issue for two common regression-based GRN inference models, an integrative Random Forest (weigthedRF) and a generalized linear model with stability selection estimated under a weighted LASSO penalty (weightedLASSO). These approaches are applied to data from the root response to nitrate induction inArabidopsis thaliana. For each gene, we measure how the integration of transcription factor binding motifs influences model prediction. We propose a new approach, DIOgene, that uses model prediction error and a simulated null hypothesis for optimizing data integration strength in a hypothesis-driven, gene-specific manner. The resulting integration scheme reveals a strong diversity of optimal integration intensities between genes. In addition, it provides a good trade-off between prediction error minimization and validation on experimental interactions, while master regulators of nitrate induction can be accurately retrieved.Availability and implementationThe R code and notebooks demonstrating the use of the proposed approaches are available in the repositoryhttps://github.com/OceaneCsn/integrative_GRN_N_induction.

Publisher

Cold Spring Harbor Laboratory

Reference63 articles.

1. SCENIC: single-cell regulatory network inference and clustering;Nature Methods,2017

2. Systems approach identifies TGA1 and TGA4 transcription factors as important regulatory components of the nitrate response ofArabidopsis thalianaroots

3. José M. Alvarez , Anna-Lena Schinke, Matthew D . Brooks Angelo Pasquino , Lauriebeth Leonelli , Kranthi Varala , Alaeddine Safi , Gabriel Krouk , Anne Krapp , and Gloria M. Coruzzi . Transient genome-wide interactions of the master transcription factor NLP7 initiate a rapid nitrogen-response cascade. Nature Communications, 11(1), March 2020.

4. An atlas of active enhancers across human cell types and tissues

5. An experimentally supported model of the Bacillus subtilis global transcriptional regulatory network