Automated Trace Clustering Pipeline Synthesis in Process Mining

Iuliana Malina Grigore; Gabriel Marques Tavares; Matheus Camilo da Silva; Paolo Ceravolo; Sylvio Barbon Junior

doi:10.3390/info15040241

Information (Apr 2024)

Automated Trace Clustering Pipeline Synthesis in Process Mining

Iuliana Malina Grigore,
Gabriel Marques Tavares,
Matheus Camilo da Silva,
Paolo Ceravolo,
Sylvio Barbon Junior

Affiliations

Iuliana Malina Grigore: Dipartimento di Ingegneria e Architettura, Università Degli Studi di Trieste, 34127 Trieste, Italy
Gabriel Marques Tavares: Chair of Database Systems and Data Mining, Ludwig-Maximilians-Universität München, 80538 Munich, Germany
Matheus Camilo da Silva: Dipartimento di Ingegneria e Architettura, Università Degli Studi di Trieste, 34127 Trieste, Italy
Paolo Ceravolo: Dipartimento di Informatica, Università Degli Studi di Milano Statale, 20122 Milano, Italy
Sylvio Barbon Junior: Dipartimento di Ingegneria e Architettura, Università Degli Studi di Trieste, 34127 Trieste, Italy

DOI: https://doi.org/10.3390/info15040241
Journal volume & issue: Vol. 15, no. 4
p. 241

Abstract

Read online

Business processes have undergone a significant transformation with the advent of the process-oriented view in organizations. The increasing complexity of business processes and the abundance of event data have driven the development and widespread adoption of process mining techniques. However, the size and noise of event logs pose challenges that require careful analysis. The inclusion of different sets of behaviors within the same business process further complicates data representation, highlighting the continued need for innovative solutions in the evolving field of process mining. Trace clustering is emerging as a solution to improve the interpretation of underlying business processes. Trace clustering offers benefits such as mitigating the impact of outliers, providing valuable insights, reducing data dimensionality, and serving as a preprocessing step in robust pipelines. However, designing an appropriate clustering pipeline can be challenging for non-experts due to the complexity of the process and the number of steps involved. For experts, it can be time-consuming and costly, requiring careful consideration of trade-offs. To address the challenge of pipeline creation, the paper proposes a genetic programming solution for trace clustering pipeline synthesis that optimizes a multi-objective function matching clustering and process quality metrics. The solution is applied to real event logs, and the results demonstrate improved performance in downstream tasks through the identification of sub-logs.

Published in Information

ISSN: 2078-2489 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: http://www.mdpi.com/journal/information/

About the journal

Abstract

Keywords