Affiliation:
1. Imperial College London, UK
Abstract
The yearly global production of data is growing exponentially, outpacing the capacity of existing storage media, such as tape and disk, and surpassing our ability to store it. DNA storage—the representation of arbitrary information as sequences of nucleotides—offers a promising storage medium. DNA is nature’s information-storage molecule of choice and has a number of key properties: It is extremely dense, offering the theoretical possibility of storing 455 EB/g; it is durable, with a half-life of approximately 520 years that can be increased to thousands of years when DNA is chilled and stored dry; and it is amenable to automated synthesis and sequencing. Furthermore, biochemical processes that act on DNA potentially enable highly parallel data manipulation.
While biological information is encoded in DNA via a specific mapping from triplet sequences of nucleotides to amino acids, DNA storage is not limited to a single encoding scheme, and there are many possible ways to map data to chemical sequences of nucleotides for synthesis, storage, retrieval, and data manipulation. However, there are several biological, error-tolerance, and information-retrieval considerations that an encoding scheme needs to address to be viable.
This comprehensive review focuses on comparing existing work done in encoding arbitrary data within DNA in terms of their encoding schemes, methods to address biological constraints, and measures to provide error correction. We compare encoding approaches on the overall information density and coverage they achieve, as well as the data-retrieval method they use (i.e., sequential or random access). We also discuss the background and evolution of the encoding schemes.
Publisher
Association for Computing Machinery (ACM)
Subject
General Computer Science,Theoretical Computer Science
Reference37 articles.
1. 2014. Keyboard scan codes. Retrieved June 20 2019 from https://www.marjorie.de/ps2/scancode-set2.htm
2. 2021. TPC-H Decision Support Benchmark. Retrieved April 14 2021 from http://www.tpc.org/tpch/
3. An improved Huffman coding method for archiving text, images, and music characters in DNA
4. The half-life of DNA in bone: measuring decay kinetics in 158 dated fossils
5. Basic local alignment search tool
Cited by
3 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献
1. Multi-file dynamic compression method based on classification algorithm in DNA storage;Medical & Biological Engineering & Computing;2024-06-26
2. Survey for a Decade of Coding for DNA Storage;IEEE Transactions on Molecular, Biological, and Multi-Scale Communications;2024-06
3. Non-Iterative MSNN Decoder for Image Transmission over AWGN Channel using LDPC;2024 26th International Conference on Digital Signal Processing and its Applications (DSPA);2024-03-27