Methods of Data Popularity Evaluation in the ATLAS Experiment at the LHC-Reference-Cited by-同舟云学术

Methods of Data Popularity Evaluation in the ATLAS Experiment at the LHC

Published:2021 Issue: Volume:251 Page:02013
ISSN:2100-014X
Container-title:EPJ Web of Conferences
language:
Short-container-title:EPJ Web Conf.

Author:

Beermann Thomas,Chuchuk Olga,Di Girolamo Alessandro,Grigorieva Maria,Klimentov Alexei,Lassnig Mario,Schulz Markus,Sciaba Andrea,Tretyakov Eugeny

Abstract

The ATLAS Experiment at the LHC generates petabytes of data that is distributed among 160 computing sites all over the world and is processed continuously by various central production and user analysis tasks. The popularity of data is typically measured as the number of accesses and plays an important role in resolving data management issues: deleting, replicating, moving between tapes, disks and caches. These data management procedures were still carried out in a semi-manual mode and now we have focused our efforts on automating it, making use of the historical knowledge about existing data management strategies. In this study we describe sources of information about data popularity and demonstrate their consistency. Based on the calculated popularity measurements, various distributions were obtained. Auxiliary information about replication and task processing allowed us to evaluate the correspondence between the number of tasks with popular data executed per site and the number of replicas per site. We also examine the popularity of user analysis data that is much less predictable than in the central production and requires more indicators than just the number of accesses.

Publisher

EDP Sciences

Link

https://www.epj-conferences.org/10.1051/epjconf/202125102013/pdf

Reference11 articles.

1. ATLAS collaboration, JINST 3, S08003 (2008)

2. Molfetas A. et al., Journal of Physics: Conference Series 331, 062018 (2011)

3. Maeno T., De K., Panitkin S., Journal of Physics: Conference Series 396, 32070 (2012)

4. Elmsheuser J., Di Girolamo A. (ATLAS), EPJ Web Conf. 214, 03010 (2019)

5. Barisits M.S., Ph.D. thesis, Vienna, Tech. U. (2017)

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Operational Analytics Studies for ATLAS Distributed Computing: Data Popularity Forecast and Utilization of the WLCG Centers;EPJ Web of Conferences;2024

2. Data Quality Computation For Obsolescence Detection Within Connected Environments;2023 International Conference on Innovations in Intelligent Systems and Applications (INISTA);2023-09-20

3. Exploring Hierarchical Forecasting of Data Popularity in High-Energy Physics Experiments;Lobachevskii Journal of Mathematics;2023-08

4. Leveraging History to Predict Infrequent Abnormal Transfers in Distributed Workflows;Sensors;2023-06-10

5. Predicting Slow Network Transfers in Scientific Computing;Fifth International Workshop on Systems and Network Telemetry and Analytics;2022-06-27