Automated Trace Clustering Pipeline Synthesis in Process Mining-Reference-Cited by-同舟云学术

Automated Trace Clustering Pipeline Synthesis in Process Mining

Published:2024-04-20 Issue:4 Volume:15 Page:241
ISSN:2078-2489
Container-title:Information
language:en
Short-container-title:Information

Author:

Grigore Iuliana Malina¹^ORCID,Tavares Gabriel Marques²³^ORCID,Silva Matheus Camilo da¹^ORCID,Ceravolo Paolo⁴^ORCID,Barbon Junior Sylvio¹^ORCID

Affiliation:

1. Dipartimento di Ingegneria e Architettura, Università Degli Studi di Trieste, 34127 Trieste, Italy

2. Chair of Database Systems and Data Mining, Ludwig-Maximilians-Universität München, 80538 Munich, Germany

3. Munich Center for Machine Learning (MCML), 80539 Munich, Germany

4. Dipartimento di Informatica, Università Degli Studi di Milano Statale, 20122 Milano, Italy

Abstract

Business processes have undergone a significant transformation with the advent of the process-oriented view in organizations. The increasing complexity of business processes and the abundance of event data have driven the development and widespread adoption of process mining techniques. However, the size and noise of event logs pose challenges that require careful analysis. The inclusion of different sets of behaviors within the same business process further complicates data representation, highlighting the continued need for innovative solutions in the evolving field of process mining. Trace clustering is emerging as a solution to improve the interpretation of underlying business processes. Trace clustering offers benefits such as mitigating the impact of outliers, providing valuable insights, reducing data dimensionality, and serving as a preprocessing step in robust pipelines. However, designing an appropriate clustering pipeline can be challenging for non-experts due to the complexity of the process and the number of steps involved. For experts, it can be time-consuming and costly, requiring careful consideration of trade-offs. To address the challenge of pipeline creation, the paper proposes a genetic programming solution for trace clustering pipeline synthesis that optimizes a multi-objective function matching clustering and process quality metrics. The solution is applied to real event logs, and the results demonstrate improved performance in downstream tasks through the identification of sub-logs.

Publisher

MDPI AG

Link

https://www.mdpi.com/2078-2489/15/4/241/pdf

Reference39 articles.

1. Business Process Management: A Comprehensive Survey;ISRN Softw. Eng.,2013

2. Opportunities and Challenges for Process Mining in Organizations: Results of a Delphi Study;Martin;Bus. Inf. Syst. Eng.,2021

3. van der Aalst, W.M.P., and Carmona, J. (2022). Process Mining Handbook, Springer.

4. Xavier-Junior, J.C., and Rios, R.A. (2022). Proceedings of the Intelligent Systems, Springer International Publishing.

5. Neubauer, T.R., Pamponet Sobrinho, G., Fantinato, M., and Peres, S.M. (2021, January 15–18). Visualization for enabling human-in-the-loop in trace clustering-based process mining tasks. Proceedings of the 2021 IEEE International Conference on Big Data (Big Data), Orlando, FL, USA.