Pith. sign in

REVIEW 1 cited by

Sage: Using Unsupervised Learning for Scalable Performance Debugging in Microservices

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.00267 v1 pith:7TQA27NM submitted 2021-01-01 cs.DC cs.PF

classification cs.DCcs.PF
keywords microservicesperformancesagecausecloudrootclustersdebugging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cloud applications are increasingly shifting from large monolithic services to complex graphs of loosely-coupled microservices. Despite the advantages of modularity and elasticity microservices offer, they also complicate cluster management and performance debugging, as dependencies between tiers introduce backpressure and cascading QoS violations. We present Sage, a machine learning-driven root cause analysis system for interactive cloud microservices. Sage leverages unsupervised ML models to circumvent the overhead of trace labeling, captures the impact of dependencies between microservices to determine the root cause of unpredictable performance online, and applies corrective actions to recover a cloud service's QoS. In experiments on both dedicated local clusters and large clusters on Google Compute Engine we show that Sage consistently achieves over 93% accuracy in correctly identifying the root cause of QoS violations, and improves performance predictability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Multi-Agent Fault Localization System Based on Monte Carlo Tree Search Approach

    cs.SE 2025-07 conditional novelty 6.0 of 10

    An LLM multi-agent system using Monte Carlo Tree Search over a Fault Mining Tree reports 49-128% higher root cause localization accuracy and much lower token use than prior LLM-based RCA methods.

Pith tools