The unified logging infrastructure for data analytics at Twitter-Reference-Cited by-同舟云学术

The unified logging infrastructure for data analytics at Twitter

Published:2012-08 Issue:12 Volume:5 Page:1771-1780
ISSN:2150-8097
Container-title:Proceedings of the VLDB Endowment
language:en
Short-container-title:Proc. VLDB Endow.

Author:

Lee George¹,Lin Jimmy¹,Liu Chuang¹,Lorek Andrew¹,Ryaboy Dmitriy¹

Affiliation:

1. Twitter, Inc.

Abstract

In recent years, there has been a substantial amount of work on large-scale data analytics using Hadoop-based platforms running on large clusters of commodity machines. A less-explored topic is how those data, dominated by application logs, are collected and structured to begin with. In this paper, we present Twitter's production logging infrastructure and its evolution from application-specific logging to a unified "client events" log format, where messages are captured in common, well-formatted, flexible Thrift messages. Since most analytics tasks consider the user session as the basic unit of analysis, we pre-materialize "session sequences", which are compact summaries that can answer a large class of common queries quickly. The development of this infrastructure has streamlined log collection and data analysis, thereby improving our ability to rapidly experiment and iterate on various aspects of the service.

Publisher

VLDB Endowment

Subject

General Earth and Planetary Sciences,Water Science and Technology,Geography, Planning and Development

Link

https://dl.acm.org/doi/pdf/10.14778/2367502.2367516

Cited by 56 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. LogFlux: A Software Suite for Replicating Results in Automated Log Parsing;Proceedings of the 2nd ACM Conference on Reproducibility and Replicability;2024-06-18

2. PreLog: A Pre-trained Model for Log Analytics;Proceedings of the ACM on Management of Data;2024-05-29

3. Exploiting Data-pattern-aware Vertical Partitioning to Achieve Fast and Low-cost Cloud Log Storage;ACM Transactions on Storage;2024-02-19

4. Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics;2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE);2023-10-09

5. LogGrep: Fast and Cheap Cloud Log Storage by Exploiting both Static and Runtime Patterns;Proceedings of the Eighteenth European Conference on Computer Systems;2023-05-08