Archiving scientific data-Reference-Cited by-同舟云学术

Archiving scientific data

Published:2004-03 Issue:1 Volume:29 Page:2-42
ISSN:0362-5915
Container-title:ACM Transactions on Database Systems
language:en
Short-container-title:ACM Trans. Database Syst.

Author:

Buneman Peter¹,Khanna Sanjeev²,Tajima Keishi³,Tan Wang-Chiew⁴

Affiliation:

1. University of Edinburgh, Edinburgh, Scotland

2. University of Pennsylvania, Philadelphia, Pennsylvania, PA

3. Japan Advanced Institute of Science and Technology, Ishikawa, Japan

4. University of California, Santa Cruz, Santa Cruz, California

Abstract

Archiving is important for scientific data, where it is necessary to record all past versions of a database in order to verify findings based upon a specific version. Much scientific data is held in a hierachical format and has a key structure that provides a canonical identification for each element of the hierarchy. In this article, we exploit these properties to develop an archiving technique that is both efficient in its use of space and preserves the continuity of elements through versions of the database, something that is not provided by traditional minimum-edit-distance diff approaches. The approach also uses timestamps. All versions of the data are merged into one hierarchy where an element appearing in multiple versions is stored only once along with a timestamp. By identifying the semantic continuity of elements and merging them into one data structure, our technique is capable of providing meaningful change descriptions, the archive allows us to easily answer certain temporal queries such as retrieval of any specific version from the archive and finding the history of an element. This is in contrast with approaches that store a sequence of deltas where such operations may require undoing a large number of changes or significant reasoning with the deltas. A suite of experiments also demonstrates that our archive does not incur any significant space overhead when contrasted with diff approaches. Another useful property of our approach is that we use XML format to represent hierarchical data and the resulting archive is also in XML. Hence, XML tools can be directly applied on our archive. In particular, we apply an XML compressor on our archive, and our experiments show that our compressed archive outperforms compressed diff-based repositories in space efficiency. We also show how we can extend our archiving tool to an external memory archiver for higher scalability and describe various index structures that can further improve the efficiency of some temporal queries on our archive.

Publisher

Association for Computing Machinery (ACM)

Subject

Information Systems

Link

https://dl.acm.org/doi/pdf/10.1145/974750.974752

Reference39 articles.

1. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000

2. Keys for XML

Cited by 61 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Modelling for Efficient Scientific Data Storage Using Simple Graphs in DNA;SN Computer Science;2024-04-01

2. Provenance Framework for Multi-Depth Querying Using Zero-Information Loss Database;International Journal of Information Technology & Decision Making;2022-11-30

3. Modelling of Efficient Graph-aware Data Storage using DNA;Proceedings of the 11th International Conference on Data Science, Technology and Applications;2022

4. From papers to practice;Proceedings of the VLDB Endowment;2021-07

5. Fine-grained lineage for safer notebook interactions;Proceedings of the VLDB Endowment;2021-02