REVIEW 2 cited by
LogRCA: Log-based Root Cause Analysis for Distributed Services
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
To assist IT service developers and operators in managing their increasingly complex service landscapes, there is a growing effort to leverage artificial intelligence in operations. To speed up troubleshooting, log anomaly detection has received much attention in particular, dealing with the identification of log events that indicate the reasons for a system failure. However, faults often propagate extensively within systems, which can result in a large number of anomalies being detected by existing approaches. In this case, it can remain very challenging for users to quickly identify the actual root cause of a failure. We propose LogRCA, a novel method for identifying a minimal set of log lines that together describe a root cause. LogRCA uses a semi-supervised learning approach to deal with rare and unknown errors and is designed to handle noisy data. We evaluated our approach on a large-scale production log data set of 44.3 million log lines, which contains 80 failures, whose root causes were labeled by experts. LogRCA consistently outperforms baselines based on deep learning and statistical analysis in terms of precision and recall to detect candidate root causes. In addition, we investigated the impact of our deployed data balancing approach, demonstrating that it considerably improves performance on rare failures.
Forward citations
Cited by 2 Pith papers
-
DDB: Source-Level Interactive Debugging for Distributed Applications
DDB extends interactive source-level debugging to distributed applications via cross-RPC backtrace reconstruction, intent-preserving breakpoint propagation, and pause-erased time virtualization, achieving 100% fault l...
-
AdaptiveLog: An Adaptive Log Analysis Framework with the Collaboration of Large and Small Language Model
An adaptive log analysis framework that routes uncertain SLM predictions to an LLM with error-case prompts, improving accuracy and cutting LLM cost by about 73%.
Discussion (0). Continue with ORCID to comment.