Exploring Genomic Large Language Models: Bridging the Gap between Natural Language and Gene Sequences-Reference-Cited by-同舟云学术

Exploring Genomic Large Language Models: Bridging the Gap between Natural Language and Gene Sequences

Published:2024-02-29 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Liu Huaqing,Zhou Shuxian,Chen Peiyi,Liu Jiahui,Huo Ku-Geng,Han Lanqing

Abstract

AbstractMotivationWith the rapid development of genomic sequencing technologies and accumulation of sequencing data, there is an increasing demand for analysis tools that are more user-friendly for non-programmer users. In support of this initiative, we developed an all-in-one tool called GenomicLLM that can understand simple grammar in the question input and perform different types of analyses and tasks accordingly.ReaultsWe trained the GenomicLLM model using three large open-access datasets, namely GenomicLLM_GRCh38, Genome Understanding Evaluation and GenomicBenchmarks, and developed a hybrid tokenization approach to allow better comprehension from mixed corpora that include sequence and non-sequence inputs. GenomicLLM can carry out a wider range of tasks. In the classification tasks that are also available in the state-of-the-art DNABERT-2 and HyenaDNA, GenomicLLM has comparable performance. Moreover, GenomicLLM can also carry out other regression and generation tasks that are not accomplishable by these tools. In summary, we demonstrated here a successful large language model with a mixture of gene sequences and natural language corpus that enables a wider range of applications.Availability and implementationCodes and data can be accessed athttps://github.com/Huatsing-Lau/GenomicLLMandhttps://zenodo.org/records/10695802

Publisher

Cold Spring Harbor Laboratory

Reference6 articles.

1. Genomic benchmarks: a collection of datasets for genomic sequence classification

2. Lin, C.-Y. ROUGE: A Package for Automatic Evaluation of Summaries. In, Annual Meeting of the Association for Computational Linguistics. 2004.

3. Papineni, K. , et al. BLEU: a method for automatic evaluation of machine translation. In, Proceedings of the 40th Annual Meeting on Association for Computational Linguistics. Philadelphia, Pennsylvania: Association for Computational Linguistics; 2002. p. 311–318.