Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision-Reference-Cited by-同舟云学术

Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Published:2024 Issue:1 Volume:4 Page:
ISSN:2635-0041
Container-title:Bioinformatics Advances
language:en
Short-container-title:

Author:

Malusare Aditya¹²^ORCID,Kothandaraman Harish²,Tamboli Dipesh³,Lanman Nadia A²⁴,Aggarwal Vaneet¹²³

Affiliation:

1. School of Industrial Engineering, Purdue University , West Lafayette, IN 47907, United States

2. Institute for Cancer Research, Purdue University , West Lafayette, IN 47907, United States

3. Elmore Family School of Electrical and Computer Engineering, Purdue University , West Lafayette, IN 47907, United States

4. Department of Comparative Pathobiology, Purdue University , West Lafayette, IN 47907, United States

Abstract

Abstract Summary This article presents the Ensemble Nucleotide Byte-level Encoder-Decoder (ENBED) foundation model, analyzing DNA sequences at byte-level precision with an encoder–decoder Transformer architecture. ENBED uses a subquadratic implementation of attention to develop an efficient model capable of sequence-to-sequence transformations, generalizing previous genomic models with encoder-only or decoder-only architectures. We use Masked Language Modeling to pretrain the foundation model using reference genome sequences and apply it in the following downstream tasks: (i) identification of enhancers, promotors, and splice sites, (ii) recognition of sequences containing base call mismatches and insertion/deletion errors, an advantage over tokenization schemes involving multiple base pairs, which lose the ability to analyze with byte-level precision, (iii) identification of biological function annotations of genomic sequences, and (iv) generating mutations of the Influenza virus using the encoder–decoder architecture and validating them against real-world observations. In each of these tasks, we demonstrate significant improvement as compared to the existing state-of-the-art results. Availability and implementation The source code used to develop and fine-tune the foundation model has been released on Github (https://github.itap.purdue.edu/Clan-labs/ENBED).

Funder

National Science Foundation

Publisher

Oxford University Press (OUP)

Link

https://academic.oup.com/bioinformaticsadvances/advance-article-pdf/doi/10.1093/bioadv/vbae117/58884622/vbae117.pdf

Reference41 articles.

1. A global reference for human genetic variation;1000 Genomes Project Consortium;Nature,2015

2. Effective gene expression prediction from sequence by integrating long-range interactions;Avsec;Nat Methods,2021

3. The influenza virus resource at the national center for biotechnology information;Bao;J Virol,2008

4. MutaGAN: a sequence-to-sequence GAN framework to predict mutations of evolving protein populations;Berman;Virus Evol,2023