Predicting transcriptional responses to heat and drought stress from genomic features using a machine learning approach in rice-Reference-Cited by-同舟云学术

Predicting transcriptional responses to heat and drought stress from genomic features using a machine learning approach in rice

Published:2023-07-17 Issue: Volume:14 Page:
ISSN:1664-462X
Container-title:Frontiers in Plant Science
language:
Short-container-title:Front. Plant Sci.

Author:

Smet Dajo,Opdebeeck Helder,Vandepoele Klaas

Abstract

Plants have evolved various mechanisms to adapt to adverse environmental stresses, such as the modulation of gene expression. Expression of stress-responsive genes is controlled by specific regulators, including transcription factors (TFs), that bind to sequence-specific binding sites, representing key components of cis-regulatory elements and regulatory networks. Our understanding of the underlying regulatory code remains, however, incomplete. Recent studies have shown that, by training machine learning (ML) algorithms on genomic sequence features, it is possible to predict which genes will transcriptionally respond to a specific stress. By identifying the most important features for gene expression prediction, these trained ML models allow, in theory, to further elucidate the regulatory code underlying the transcriptional response to abiotic stress. Here, we trained random forest ML models to predict gene expression in rice (Oryza sativa) in response to heat or drought stress. Apart from thoroughly assessing model performance and robustness across various input training data, the importance of promoter and gene body sequence features to train ML models was evaluated. The use of enriched promoter oligomers, complementing known TF binding sites, allowed us to gain novel insights in DNA motifs contributing to the stress regulatory code. By comparing genomic feature importance scores for drought and heat stress over time, general and stress-specific genomic features contributing to the performance of the learned models and their temporal variation were identified. This study provides a solid foundation to build and interpret ML models accurately predicting transcriptional responses and enables novel insights in biological sequence features that are important for abiotic stress responses.

Publisher

Frontiers Media SA

Subject

Plant Science

Reference88 articles.

1. Permutation importance: a corrected feature importance measure;Altmann;Bioinformatics,2010

2. ArrowK. J. BarankinE. W. BlackwellD. BottR. DalkeyN. DresherM. Princeton University PressContributions to the theory of games (AM-28)1953

3. Recent insights into signaling responses to cope drought stress in rice;Aslam;Rice Sci.,2022

4. The cis-regulatory codes of response to combined heat and drought stress in arabidopsis thaliana;Azodi;NAR Genom. Bioinform.,2020

5. Trimmomatic: a flexible trimmer for illumina sequence data;Bolger;Bioinformatics,2014

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ASPTF: A computational tool to predict abiotic stress-responsive transcription factors in plants by employing machine learning algorithms;Biochimica et Biophysica Acta (BBA) - General Subjects;2024-06

2. Predicting gene expression responses to environment inArabidopsis thalianausing natural variation in DNA sequence;2024-04-28

3. Deep learning the cis-regulatory code for gene expression in selected model plants;Nature Communications;2024-04-25