It is shown that strong overall accuracy can coexist with poor performance on rare but operationally critical log templates, and the need for evaluation protocols that explicitly account for template frequency when assessing log parsers in real-world settings is highlighted.
Abstract
Log parsing is a critical step in automated log analysis, enabling tasks such as debugging, monitoring, root cause analysis, and anomaly detection by transforming unstructured log messages into structured representations. Despite extensive research on log parsing, real-world log data exhibits a pronounced long-tail structure at the template level: a small number of frequent log templates dominate the data, while a large fraction of templates occur only a few times. These rare templates often correspond to failures, abnormal behaviors, or rare system events, yet they are severely underrepresented in commonly used benchmarks such as LogHub-2.0.In this paper, we conduct a comprehensive empirical study across multiple real-world log datasets to quantify the prevalence of rare log templates and to examine their impact on log parsing evaluation. Our analysis shows that, on average, approximately 12% of log templates are rare, while collectively accounting for less than 0.03% of log messages. In our research, rare log events defined as groups that having fewer than five associated log instances. We further demonstrate that commonly used evaluation metrics, such as Parsing Accuracy (PA) and Grouping Accuracy (GA), are dominated by frequent templates and can obscure substantial performance degradation on rare ones. Using template-level metrics, including F1_score of Template accuracy(FTA) and F1_score of grouping accuracy(FGA), we show that strong overall accuracy can coexist with poor performance on rare but operationally critical log templates. These findings highlight the need for evaluation protocols that explicitly account for template frequency when assessing log parsers in real-world settings.
An extensive empirical study of state-of-the-art DGAD models is conducted, revealing that detection difficulty increases consistently with anomaly complexity, from simple localized irregularities to coordinated and temporally persistent structures.
Mohamed Nazim Mezhoudi, Guillaume Lachaud, Yan-Lei Diao et al.· Proceedings of the 32nd ACM...· 0 citations
: Rapid growth in volume and complexity of system logs in modern computing environments requires efficient anomaly detection to identify potential system failures, security breaches, or performance issues. This paper introduces a novel approach to detecting anomalies in system logs using Small Language Models (SLMs). W...
Pulkit Gupta, Jayani Shah, Param Chheda et al.· Proceedings of the 1st Inter...· 0 citations
In large-scale systems, fault localization remains expensive because bug reports are often ambiguous and incomplete. In practice, developers rely heavily on runtime logs and coverage data as critical clues for reasoning about how faults propagate through systems. However, these rich diagnostic signals are rarely integr...
Zheyuan Lin, Yang Feng, Jian-Jun Chen et al.· Fall Joint Computer Conferen...· 0 citations
Log parsing, which transforms unstructured log messages into structured formats, is a critical step in automated log analysis and directly impacts the effectiveness of downstream tasks. However, existing log parsers struggle to balance effectiveness with efficiency and show limited capability in handling log inconsiste...
Shu-Ting Lai, Haiyu Huang, Peng-Fei Chen et al.· ACM Transactions on Software...· 0 citations
A scalable framework that leverages LLMs as data translators to bridge the gap between unstructured textual resources and structured event data is proposed, and finetuning LLMs on a newly created text-to-log dataset shows that the resulting models can extract high-fidelity event logs from unstructured resources.
Maximilian Seeth, G. Tavares, Daniel Schuster· 0 citations
Understanding binary programs is challenging due to the loss of high-level abstractions during compilation. Type inference plays a key role in recovering information such as variable types, data structures, and class hierarchies, which is crucial for reverse engineering (RE), decompilation, and security analysis. This...
Raisul Arefin, Ryan Vrecenar, Samuel Mulder· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.