Discovering longest-lasting correlation in sequence databases-Reference-Cited by-同舟云学术

Discovering longest-lasting correlation in sequence databases

Published:2013-09 Issue:14 Volume:6 Page:1666-1677
ISSN:2150-8097
Container-title:Proceedings of the VLDB Endowment
language:en
Short-container-title:Proc. VLDB Endow.

Author:

Li Yuhong¹,U Leong Hou¹,Yiu Man Lung²,Gong Zhiguo¹

Affiliation:

1. Department of Computer and Information Science, University of Macau, Macau

2. Department of Computing, Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong

Abstract

Most existing work on sequence databases use correlation (e.g., Euclidean distance and Pearson correlation) as a core function for various analytical tasks. Typically, it requires users to set a length for the similarity queries. However, there is no steady way to define the proper length on different application needs. In this work we focus on discovering longest-lasting highly correlated subsequences in sequence databases, which is particularly useful in helping those analyses without prior knowledge about the query length. Surprisingly, there has been limited work on this problem. A baseline solution is to calculate the correlations for every possible subsequence combination. Obviously, the brute force solution is not scalable for large datasets. In this work we study a space-constrained index that gives a tight correlation bound for subsequences of similar length and offset by intra-object grouping and inter-object grouping techniques. To the best of our knowledge, this is the first index to support normalized distance metric of arbitrary length subsequences. Extensive experimental evaluation on both real and synthetic sequence datasets verifies the efficiency and effectiveness of our proposed methods.

Publisher

VLDB Endowment

Subject

General Earth and Planetary Sciences,Water Science and Technology,Geography, Planning and Development

Link

https://dl.acm.org/doi/pdf/10.14778/2556549.2556552

Cited by 23 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A Compact and Efficient Neural Data Structure for Mutual Information Estimation in Large Timeseries;Proceedings of the 36th International Conference on Scientific and Statistical Database Management;2024-07-10

2. Static and Streaming Discovery of Maximal Linear Representation Between Time Series;IEEE Transactions on Knowledge and Data Engineering;2024-01

3. Correlation Joins over Time Series Data Streams Utilizing Complementary Dimension Reduction and Transformation;Proceedings of the ACM on Management of Data;2023-12-08

4. TSM-Bench: Benchmarking Time Series Database Systems for Monitoring Applications;Proceedings of the VLDB Endowment;2023-07

5. Constructing Compact Time Series Index for Efficient Window Query Processing;2022 IEEE 38th International Conference on Data Engineering (ICDE);2022-05