Pith. sign in

REVIEW 2 major objections 9 minor 94 references

DDB: Source-Level Interactive Debugging for Distributed Applications

T0 review · 2 major / 9 minor · reviewed 2026-07-08 · glm-5.2

Pith's one-line read Distributed debugger pauses a cluster without triggering timeouts

desk verdict DDB is a working distributed interactive debugger with three genuinely novel mechanisms; the main soft spot is PET's coverage gap for kernel-level timer mechanisms, which is real but narrower than it first appears. read the letter →

arxiv 2607.06107 v1 pith:BLTT5MGY submitted 2026-07-07 cs.DC

classification cs.DC
keywords distributeddebugginginteractivetimevirtualizationbacktracetimeoutcascadepreventionRPCpause-the-worldclock
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

DDB brings source-level interactive debugging—setting breakpoints, inspecting call stacks, examining live variables—to distributed applications by solving three problems that previously made it impractical. First, Distributed Backtrace (DBT) embeds compact caller-context metadata in every RPC payload, allowing a unified call stack to be reconstructed across process boundaries regardless of user-level thread rescheduling. Second, an intent-preserving control plane propagates breakpoints and debug commands across dynamic process sets, so breakpoints automatically follow auto-scaling, restarts, and computation migrations without manual per-process management. Third, Pause-Erased Time (PET) virtualizes each process's clock by intercepting POSIX time APIs and subtracting accumulated pause durations, preventing the timeout cascades (leader elections, transaction aborts, crash recovery) that would otherwise destroy the state being debugged. The paper claims this works with 20-60 lines of integration code per RPC framework, achieves 30ms median cross-RPC backtrace latency at 122-process scale, keeps perceived time drift under 5ms, adds 1-5% throughput overhead, and enabled 100% fault localization in a controlled user study versus 38.5% for baseline tools.

What carries the argument

Pause-Erased Time (PET) with Virtual Deadline Enforcement; Distributed Backtrace (DBT) with caller-context metadata embedding; Intent-preserving control plane with logical groups and scope resolution

What would settle it

If a distributed application uses raw rdtsc instructions or NIC hardware timestamps for its timeout logic, PET cannot intercept these, and debugger-induced pauses would still trigger timeout cascades—making interactive debugging unsafe for such applications.

Watch

Extended reading notes

Core claim

The central mechanism is Pause-Erased Time (PET), specifically its Virtual Deadline Enforcement component. The key insight is that a debugger can safely pause an entire distributed cluster if every process's perception of elapsed time is adjusted to subtract out the pause duration. PET maintains a cumulative offset of all pause durations and intercepts time API calls (get_time, sleep_until, sleep) to return adjusted values. The subtlety that makes this non-trivial is that processes sleeping during a pause will be prematurely woken by the kernel when real time has advanced past their deadline. PET's Virtual Deadline Enforcement traps these premature wakeups, recalculates the adjusted deadline

Load-bearing premise

PET assumes all time-sensitive application behavior routes through POSIX time APIs that an LD_PRELOAD shim can intercept; applications using raw rdtsc, bypassing libc, or interacting with external services enforcing strict physical-time leases would still experience timeout cascades.

Editorial extensions

If this is right

  • If PET's approach is sound, distributed systems developers could replace iterative log-and-redeploy cycles with live interactive debugging sessions, reducing fault localization time from days to minutes.
  • The 20-60 LoC integration cost per framework suggests that distributed interactive debugging could become a standard feature of RPC frameworks with minimal engineering effort.
  • The user study's 100% vs 38.5% fault localization gap implies that a significant fraction of distributed system debugging failures are caused by lack of cross-service runtime state visibility rather than developer skill or tooling familiarity.
  • PET's temporal virtualization could potentially be applied beyond debugging—to controlled testing scenarios where pausing a cluster is needed for inspection without disrupting protocol invariants.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • PET's correctness is limited to applications that route time through interceptable POSIX APIs; systems using raw rdtsc or hardware timestamps would need deeper interception (eBPF, hypervisor trapping) that the paper describes as transferable but does not implement.
  • The pause-the-world model assumes a coordinated global pause is feasible, which the paper scopes to staging and test environments; production debugging would require a different approach.
  • The 5ms time drift bound provides an order-of-magnitude safety margin against the most aggressive real-world timeouts (50-100ms), but this margin could erode for systems with sub-10ms failure detection thresholds.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 9 minor

Summary. The paper presents DDB, a source-level interactive debugger for distributed applications built on three mechanisms: (1) Distributed Backtrace (DBT), which embeds caller-context metadata in RPC payloads to reconstruct unified call stacks across process boundaries; (2) an intent-preserving control plane that propagates breakpoints across dynamic process sets; and (3) Pause-Erased Time (PET), which virtualizes each process's clock via LD_PRELOAD interception of POSIX time APIs to prevent debugger-induced timeout cascades. The system integrates with four RPC frameworks (gRPC, Nu, Quicksand, ServiceWeaver) in 10–60 LoC each. The evaluation reports 30 ms median backtrace latency at 122-process scale, sub-5 ms time drift under repeated pauses, 1–5% throughput overhead, and a controlled user study showing 100% fault localization vs. 38.5% for baseline tools. PET invariants (monotonicity, timer correctness) are formally proven in Appendix A.

Significance. The paper addresses a well-motivated problem: interactive debugging has been considered impractical for distributed systems due to call-stack termination at process boundaries, dynamic topology, and timeout cascades. The three-pillar design is clean and each pillar is evaluated independently. Strengths include: (1) formal correctness proofs for PET invariants (Appendix A.1–A.3) with clean mathematical notation; (2) a user study (§6.1) providing empirical evidence of diagnostic efficacy gains over GDB and OpenTelemetry baselines; (3) reproducible integration effort quantified in LoC (Table 2); (4) falsifiable performance claims benchmarked against real-world timeout thresholds from LogCabin, RAMCloud, and Nu (Table 5). The 20–60 LoC integration claim is a concrete, verifiable metric. The evaluation covers responsiveness, PET effectiveness, and overhead across multiple frameworks and scales.

major comments (2)
  1. §4.3 and Table 1: PET's LD_PRELOAD shim intercepts libc wrappers for POSIX time APIs (clock_gettime, gettimeofday, pthread_cond_timedwait, sleep, nanosleep, sem_timedwait). However, several common kernel-level timer mechanisms do not route through libc in a way that the offset-subtraction or Virtual Deadline Enforcement mechanism can handle. Specifically: (a) timerfd_create/timerfd_settime arms kernel timers based on real time; when the fd becomes readable in an epoll loop, the application reads the timer event and believes the interval has elapsed, even though PET should have erased the pause duration. Unlike sleep/pthread_cond_timedwait, there is no blocking primitive for Virtual Deadline Enforcement to intercept and re-arm. (b) epoll_wait with a relative timeout counts real elapsed time against the timeout during a pause; the shim could intercept the libc epoll_wait wrapper and adjust
  2. §6.3: PET effectiveness is evaluated using a synthetic application that calls POSIX time APIs in a tight loop. This does not test the timerfd or epoll_wait timeout patterns raised above. Since the paper's central claim is that PET makes pausing safe, and since timerfd/epoll-based event loops are a common pattern in the evaluated frameworks (e.g., gRPC C++ uses epoll), the evaluation should include at least one test case that exercises these kernel-level timer mechanisms. If PET does not handle them, the limitation should be stated more explicitly in §5.1 alongside the rdtsc gap, rather than only acknowledging raw rdtsc and NIC timestamps.
minor comments (9)
  1. Abstract: states '20-60 LoC' but §1 and §5 state '10-60 LoC' (ServiceWeaver is ~10 LoC per Table 2). The abstract should match the body.
  2. Table 2: The 'Framework Support' rows list gRPC≈20, Nu≈30, Quicksand≈60, ServiceWeaver≈10, but the text in §5 says '≈20 LoC changes to gRPC, ≈30 LoC changes to Nu, ≈60 LoC changes to Quicksand, and ≈10 LoC changes to ServiceWeaver.' These are consistent, but the abstract's '20-60' range should be '10-60' to include ServiceWeaver.
  3. §6.1, Study 2: The sample size is 9 participants with 3 tools in a within-subjects design. While the Latin Square counterbalancing is appropriate, the paper should report whether the differences in localization success rate (100% vs. 44% vs. 33%) are statistically significant (e.g., Fisher's exact test or chi-square), given the small sample.
  4. Figure 10: The y-axis label and units are unclear. The caption mentions 'time jumps' but the axis values and what 'PET' vs. 'physical time' represent in the plot could be labeled more explicitly.
  5. §4.1: The term 'thread context' is used to describe the caller-context metadata embedded in RPC payloads, but the exact contents (which registers, how many bytes) are not specified until §6.4 mentions 52 bytes. Clarifying this earlier would help readers understand the metadata overhead.
  6. Table 5: The 'Transaction Timeout' row for RAMCloud lists 'N×50ms' with a dagger footnote, but N is not defined in the table caption. It is mentioned in the text as 'the number of participants involved in the transaction' but should be noted in the table itself (or the footnote should be expanded).
  7. Appendix A.1, Invariant 1 proof: The partition into running intervals R and paused intervals P assumes that t1 and t2 both fall in running intervals. The proof states this correctly, but the case where a get_time call occurs during a pause is not discussed. Since the application is suspended during pauses, this is impossible by construction, but stating this assumption explicitly would make the proof self-contained.
  8. §7: The related work mentions TotalView and Linaro DDT for HPC but does not discuss how DDB's pause-the-world model compares to HPC debuggers' parallel pause mechanisms in terms of scalability. A brief comparison would strengthen the positioning.
  9. The paper uses both '10-60 LoC' (§1, §5) and '20-60 LoC' (abstract). Standardize throughout.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the careful reading and for identifying a genuine gap in PET's coverage and its evaluation. The referee is correct that kernel-level timer mechanisms (timerfd, epoll_wait timeouts) are not handled by the current LD_PRELOAD shim and that this limitation is not adequately disclosed in §5.1. We will revise the manuscript accordingly.

read point-by-point responses
  1. Referee: §4.3 and Table 1: PET's LD_PRELOAD shim intercepts libc wrappers for POSIX time APIs, but kernel-level timer mechanisms such as timerfd_create/timerfd_settime and epoll_wait with relative timeouts do not route through libc in a way that the offset-subtraction or Virtual Deadline Enforcement mechanism can handle. (a) timerfd arms kernel timers based on real time; when the fd becomes readable after a pause, the application believes the interval elapsed, and there is no blocking primitive for VDE to intercept and re-arm. (b) epoll_wait with a relative timeout counts real elapsed time during a pause.

    Authors: The referee is correct on both points. We will revise the manuscript to explicitly acknowledge these limitations in §5.1, alongside the existing rdtsc and NIC timestamp caveats. We address each sub-point below. (a) timerfd: This is a genuine gap that cannot be resolved within the current LD_PRELOAD architecture. When timerfd_settime arms a kernel timer, the kernel independently tracks real time against the fd. During a debugger pause, the kernel timer expires and marks the fd as readable. On resume, the application's epoll loop observes the readable fd and reads the timer expiration event, concluding that the interval has elapsed. Unlike sleep or pthread_cond_timedwait, there is no blocking libc primitive where Virtual Deadline Enforcement can intercept a premature wakeup and re-arm the wait. The shim could in principle intercept timerfd_settime and read() on timer fds to track and suppress premature expirations, but this would require maintaining a mapping of timer fds to their PET-adjusted deadlines and intercepting read() selectively—a substantially more complex mechanism that we have not implemented or evaluated. We will state this as a limitation. (b) epoll_wait: The relative timeout parameter to epoll_wait is, in principle, interceptable: the shim could intercept the libc epoll_wait wrapper, adjust the timeout by the cumulative pause offset, and apply Virtual Deadline Enforcement logic analogous to the sleep case (re-arming with a corrected timeout on premature return). However, our current implementation does not intercept epoll_wait, and Table 1 does not list it. We will acknowledge this in the revision. We note that for the frameworks evaluated in this paper (gRPC C++, Nu, Quicksand, ServiceWeaver), the applications we tested do not rely on epoll_wait relative超 revision: no

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: DDB's three mechanisms are independently designed and evaluated against external baselines.

full rationale

The paper presents DDB, a distributed interactive debugger with three pillars: Distributed Backtrace (DBT), an intent-preserving control plane, and Pause-Erased Time (PET). Each pillar is defined by its own mechanism and evaluated against external benchmarks, not by circular derivation. DBT embeds caller-context metadata in RPC payloads and reconstructs cross-process call stacks via iterative DWARF unwinding (§4.1, Algorithm 1); this is a constructive mechanism, not a definition that presupposes its output. The control plane propagates debug-intents across dynamic process sets (§4.2); its guarantees are enforced by invariant maintenance, not by fitting to the result. PET virtualizes clocks via pause-offset accumulation and Virtual Deadline Enforcement (§4.3); its correctness is established by two formal invariants (Monotonicity and Timer Correctness, Appendix A.1) that are proven from the mathematical definition of the offset function O(t), not from the conclusion. The evaluation measures latency, overhead, and PET effectiveness against external systems (LogCabin, RAMCloud, Nu) and baselines (GDB, OpenTelemetry). The user study compares DDB against GDB and OTel as independent baselines. No step in the derivation chain reduces to its inputs by construction. The formal proofs in Appendix A.1 define T_v(t) = T(t) - O(t) and prove monotonicity and timer correctness from the properties of O(t) and the re-arming loop; these are genuine proofs from definitions, not tautologies. The paper's limitations (POSIX-only interception, rdtsc gap, intra-cluster scope) are honestly acknowledged and do not introduce circularity. Self-citations are absent from the load-bearing arguments. The derivation is self-contained.

Assumptions & free parameters 2 free parameters · 6 assumptions · 3 invented entities

The axiom ledger captures the design parameters and assumptions underlying DDB's three mechanisms. The free parameters are design choices (metadata size, grouping policy) rather than fitted constants. The domain assumptions define the scope of applicability (POSIX time APIs, intra-cluster isolation, staging/test environments, GDB compatibility, symmetric address spaces for migration). The invented entities are the three core mechanisms, each with independent empirical or formal evidence.

free parameters (2)
  • Caller-context metadata size = 52 bytes per RPC payload
    The 52-byte metadata size is a design choice for the caller-context embedded in each RPC payload, chosen to be compact while carrying thread context registers, IP address, and PID.
  • Logical group default policy = binary-based grouping
    The default grouping policy (processes running the same binary belong to the same group) is a design choice that determines how debug-intents are scoped and resolved.
assumptions (6)
  • domain assumption All time-sensitive application behavior in the target distributed systems routes through POSIX time APIs (gettimeofday, clock_gettime, sleep, nanosleep, pthread_cond_timedwait, sem_timedwait).
    PET's LD_PRELOAD shim layer intercepts POSIX time APIs. If applications use raw rdtsc, NIC hardware timestamps, or bypass libc, PET cannot virtualize time. Stated in §5.1: 'DDB only provides PET on top of POSIX time APIs. Thus, it does not work in applications that directly use raw rdtsc or NIC hardware timestamp.'
  • domain assumption The debugged cluster is isolated from external true-time observers that enforce strict physical-time deadlines.
    PET's temporal guarantees are strictly intra-cluster. Stated in §4.3: 'PET cannot virtualize the physical flow of time for external, true-time observers. If the cluster interacts with external systems that rely on physical time... those external services will perceive the developer's pause as a genuine timeout.'
  • domain assumption DDB targets staging and test environments, not production, where coordinated global pauses are feasible.
    The pause-the-world model requires all attached processes to be concurrently paused. Stated in §1: 'Because interactive debugging naturally targets staging and test environments rather than production, the overhead of these mechanisms is acceptable.'
  • standard math The kernel's monotonic clock is non-decreasing, and NTP steps on the wall clock are an application-level responsibility.
    Used in the proof of Invariant 1 (Monotonicity) in Appendix A.1 and in Appendix A.3 for wall clock semantics. The kernel clock's monotonicity is a standard OS property.
  • domain assumption GDB can interpret and manage the target binaries and programming languages.
    DDB's DKnot agent is an extended version of GDB. Stated in §5.1: 'Our prototype currently assumes GDB as the underlying debugger. Therefore, our prototype does not work for binaries and programming languages that GDB cannot interpret and manage.'
  • domain assumption For frameworks with computation migration (Nu, Quicksand), symmetric address-space layouts exist with ASLR disabled, and a logical heap locator API is available.
    DDB's heap restoration during migration relies on these properties. Stated in Appendix E: 'The underlying frameworks deploy an identical, universally linked binary across all cluster nodes with Address Space Layout Randomization (ASLR) strictly disabled.'
invented entities (3)
  • Distributed Backtrace (DBT) independent evidence
    purpose: Reconstructs a unified call stack across RPC boundaries by embedding caller-context metadata in RPC payloads and iteratively contacting upstream processes.
    DBT is evaluated empirically: Table 3 shows 30ms median latency at 122-process scale, Table 4 shows latency scaling with call depth. The algorithm is specified in Algorithm 1 and Appendix B.1. Falsifiable: if the metadata embedding fails to capture correct caller context under uthread rescheduling, backtraces would be wrong.
  • Intent-Preserving Control Plane independent evidence
    purpose: Manages debug-intents (source location, scope, command) as specifications continuously enforced against dynamic process topologies.
    Evaluated through the user study (breakpoints auto-propagate to new replicas) and the socialnet benchmark (36 services, 5 replicas each). The invariant enforcement property is stated and the mechanism is described in §4.2.
  • Pause-Erased Time (PET) with Virtual Deadline Enforcement independent evidence
    purpose: Virtualizes each process's clock to make debugger pauses invisible to application-level timeouts and timers.
    PET is formally proven correct (Appendix A.1-A.3: monotonicity and timer correctness invariants) and empirically evaluated (Figure 10: sub-5ms time drift, Table 6: nanosecond-scale interception overhead). Virtual Deadline Enforcement is distinguished from existing virtual clock mechanisms by its handling of dynamic offset growth during static application state.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DDB: Source-Level Interactive Debugging for Distributed Applications." pith.science (2026). https://pith.science/paper/BLTT5MGY

@misc{pith2026260706107,
  author       = {Pith},
  title        = {Pith review of: DDB: Source-Level Interactive Debugging for Distributed Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BLTT5MGY}},
  note         = {Machine review of arXiv:2607.06107}
}
read the original abstract

Interactive debugging is an effective tool for understanding program behavior at the source level, allowing developers to pause execution, navigate the call stack, and inspect runtime state. However, interactive debuggers are designed for single-process execution, and interactive debugging has been widely considered impractical for distributed systems. Call stacks stop at process boundaries, debugging state fails to survive infrastructure dynamics, and, most critically, debugger-induced execution pauses trigger catastrophic timeout cascades that destroy the intended debug flow. Consequently, developers are forced to abandon live hypothesis testing in favor of unwieldy and iterative log-and-redeploy cycles. We present DDB, a source-level interactive debugger that extends interactive debugging capabilities to distributed applications. We show that each of these challenges admits a targeted solution. To bridge disjoint processes, Distributed Backtrace (DBT) embeds compact causality metadata in every RPC and reconstructs a unified call stack across RPC boundaries. To manage the lifecycle of a distributed session, an intent-preserving control plane automatically coordinates and propagates breakpoints across dynamic process sets. To make pausing safe, Pause-Erased Time (PET) virtualizes each process's clock, decoupling logical time from physical pauses and preventing timeout cascades. DDB integrates with an RPC framework in 20-60 lines of code. Evaluated on gRPC, ServiceWeaver, Nu, and Quicksand across up to 122 processes, DDB achieves 30ms median cross-RPC backtrace latency, sub-5 ms time jump under repeated execution pauses, and adds 1-5% throughput overhead, comparable to attaching a single-process debugger. In a controlled user study, DDB achieves a 100% fault localization success rate (compared to 38.5% for baseline tools) with a median localization time of ~8 minutes.

Figures

Figures reproduced from arXiv: 2607.06107 by the authors.

Figure 1
Figure 1. Distributed programming paradigms and infrastructure have become much more accessible to ordinary application devel￾opers. The adoption of distributed programming has increased and the required expertise barriers have dropped substantially. execution across processes [42, 71]. The result is that appli￾cation developers—who may have no profound expertise in distributed systems—are now routinely writing code whose exe… view at source ↗
Figure 2
Figure 2. Three challenges of employing interactive debugging in distributed applications. conclusion that a single interactive inspection of the cross￾service call chain could have reached in minutes. Monitoring, tracing, logging, and record-and-replay each address a distinct need, but none provides the capability this workflow demands: live, interactive inspection across service boundaries. The following subsections detail … view at source ↗
Figure 3
Figure 3. The DDB VSCode graphical frontend debugging a three-node Raft cluster (with one leader and two followers in the screenshot). Both followers have paused at the local AppendEntries RPC handler (middle). The Call Stack (top-right) displays a Distributed Backtrace extending to the leader’s remote calling frame. The DDB Sessions pane (bottom-right) and debugger control pane (top) allow granular per-process execution cont… view at source ↗
Figures from the paper (7 more)
Figure 5
Figure 5. Figure 5: Illustration of the frame injection. Blue numbers on the left indicates the frame number, where 0 is the outermost frame and 4 is the innermost frame. Frame #2 is the injected frame. The local variable meta is inside the extraction frame. injected extraction frame (DDB…
Figure 4
Figure 4. Figure 4: Overview of DDB’s three pillars. Distributed Backtrace (DBT) reconstructs cross-RPC call stacks for live inspection. The intent-preserving control plane automatically propagates debug operations across dynamic process sets and enables granular per￾process control. Paus…
Figure 6
Figure 6. Figure 6: The design overview of DDB. Components in gray are DDB components. kernel wakeup from a debugger resume, causes the shim to re-arm. Application post-sleep code executes only when the intended PET duration has genuinely passed. These invariants ensure that application l…
Figure 7
Figure 7. Figure 7: Left: Integration success rates for integrating DDB vs. OpenTelemetry by satisfying the requirement to view linear (G1) and concurrent (G2) caller-callee traces. Right: Completion time for successful DDB integrations (complete both G1 and G2). OpenTelemetry is not incl…
Figure 9
Figure 9. Figure 9: Case-by-case comparison of time-to-localize and time￾to-fix across different bug cases. Case 1: Random Segmentation Fault; Case 2: Logic Error; Case 3: Distributed Deadlock. The GDB outlier in Case 3 corresponds to a participant who resolved the fault based on prior in…
Figure 10
Figure 10. Figure 10: The application perceived time (PET) with multiple pauses (P) and resumes (R). The gray area presents time jumps that the app perceives between the execution pause and resume. time jumps, which is clearly smaller than the 100 ms inter￾val. Because PET accumulates and …
Figure 11
Figure 11. Figure 11: The performance overhead when DDB metadata is embedded, and DDB is attached across Nu’s socialnet, Ser￾viceWeaver (SW)’s socialnet, and gRPC-based Raft consensus implementation. DDB introduces minimal throughput degradation on socialnet. Specifically, socialnet runnin…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

94 extracted references · 94 canonical work pages

  1. [1]

    https://akka.io

    Akka: Build and Run Apps That React to Change. https://akka.io. Accessed: 2025-02-25

  2. [2]

    https://github.com/AlDanial/cloc

    AlDanial/Cloc. https://github.com/AlDanial/cloc. Accessed: 2025-05-06

  3. [3]

    https://learn.microsoft

    Azure Monitor Logs - Azure Monitor. https://learn.microsoft. com/en-us/azure/azure-monitor/logs/data-platform-logs . Accessed: 2024-10-20

  4. [4]

    https://cloud.google.com/ logging

    Cloud Logging on Google Cloud. https://cloud.google.com/ logging. Accessed: 2024-10-20

  5. [5]

    https://kubernetes.io/docs/concepts/ workloads/pods/ephemeral-containers/

    Ephemeral Containers. https://kubernetes.io/docs/concepts/ workloads/pods/ephemeral-containers/. Accessed: 2024-12-10

  6. [6]

    https: //grafana.com/

    Grafana: The Open and Composable Observability Platform. https: //grafana.com/. Accessed: 2025-02-27

  7. [7]

    Accessed: 2024-10-20

    gRPC.https://grpc.io/. Accessed: 2024-10-20

  8. [8]

    https://www

    Jaeger: Open Source, Distributed Tracing Platform. https://www. jaegertracing.io/. Accessed: 2024-11-14

Show all 94 references
  1. [9]

    https://www.linaroforge.com/linaro-ddt/

    Linaro DDT. https://www.linaroforge.com/linaro-ddt/. Ac- cessed: 2025-05-14

  2. [10]

    Accessed: 2025-09-15

    LLDB.https://lldb.llvm.org/. Accessed: 2025-09-15

  3. [11]

    https: //engineering.fb.com/2017/08/31/core-infra/logdevice- a-distributed-data-store-for-logs/

    LogDevice: A Distributed Data Store for Logs. https: //engineering.fb.com/2017/08/31/core-infra/logdevice- a-distributed-data-store-for-logs/. Accessed: 2024-10-20

  4. [12]

    https://opentelemetry.io/

    OpenTelemetry. https://opentelemetry.io/. Accessed: 2024-10- 20

  5. [13]

    https://zipkin.io/

    OpenZipkin: A Distributed Tracing System. https://zipkin.io/. Accessed: 2024-11-14

  6. [14]

    https://docs.ray.io/en/latest/ray- observability/user-guides/debug-apps/ray-debugging.html

    Ray Debugger. https://docs.ray.io/en/latest/ray- observability/user-guides/debug-apps/ray-debugging.html . Accessed: 2024-11-04

  7. [15]

    https://rr- project.org/

    RR: Lightweight Recording & Deterministic Debugging. https://rr- project.org/. Accessed: 2025-09-15

  8. [16]

    https://totalview.io/

    TotalView: Debugger for HPC Computing. https://totalview.io/. Accessed: 2025-05-14

  9. [17]

    https://github.com/logcabin/logcabin

    Logcabin/Logcabin. https://github.com/logcabin/logcabin. Ac- cessed: 2024-11-22

  10. [18]

    https://dapr.io

    Dapr: Distributed Application Runtime. https://dapr.io. CNCF Graduated Project, 2024

  11. [19]

    https://www.sourceware.org/ gdb/

    GDB: The GNU Project Debugger. https://www.sourceware.org/ gdb/

  12. [20]

    Rohan Achar, Pritha Dawn, and Cristina V. Lopes. GoTcha: An Inter- active Debugger for GoT-based Distributed Systems. InProceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software (Onward! 2019). Associat...

  13. [21]

    Cloud-Scale Runtime Verification of Serverless Applications

    Kalev Alpernas, Aurojit Panda, Leonid Ryzhyk, and Mooly Sagiv. Cloud-Scale Runtime Verification of Serverless Applications. InPro- ceedings of the ACM Symposium on Cloud Computing. ACM, Seattle WA USA, 92–107. doi:10.1145/3472883.3486977

  14. [22]

    TraceWeaver: Distributed Request Tracing for Microservices Without Application Modification

    Sachin Ashok, Vipul Harsh, Brighten Godfrey, Radhika Mittal, Srini- vasan Parthasarathy, and Larisa Shwartz. TraceWeaver: Distributed Request Tracing for Microservices Without Application Modification. InProceedings of the ACM SIGCOMM 2024 Conference. ACM, Sydney NSW Australia...

  15. [23]

    Barr and Mark Marron

    Earl T. Barr and Mark Marron. Tardis: Affordable Time-Travel Debug- ging in Managed Runtimes.ACM SIGPLAN Notices49, 10 (Oct. 2014), 67–82. doi:10.1145/2714064.2660209

  16. [24]

    Orleans: Distributed Virtual Actors for Programmability and Scalability

    Philip A Bernstein, Sergey Bykov, Alan Geller, Gabriel Kliot, and Jorgen Thelin. Orleans: Distributed Virtual Actors for Programmability and Scalability

  17. [25]

    Ivan Beschastnikh, Patty Wang, Yuriy Brun, and Michael D. Ernst. Debugging Distributed Systems.Commun. ACM59, 8 (July 2016), 32–37. doi:10.1145/2909480

  18. [26]

    Blumofe, Christopher F

    Robert D. Blumofe, Christopher F. Joerg, Bradley C. Kuszmaul, Charles E. Leiserson, Keith H. Randall, and Yuli Zhou. Cilk: An Effi- cient Multithreaded Runtime System.SIGPLAN Not.30, 8 (Aug. 1995), 207–216. doi:10.1145/209937.209958

  19. [27]

    Transparent Checkpoints of Closed Distributed Systems in Emulab

    Anton Burtsev, Prashanth Radhakrishnan, Mike Hibler, and Jay Lep- reau. Transparent Checkpoints of Closed Distributed Systems in Emulab. InProceedings of the 4th ACM European Conference on Com- puter Systems (EuroSys ’09). Association for Computing Machinery, New York, NY, USA...

  20. [28]

    Mani Chandy and Leslie Lamport

    K. Mani Chandy and Leslie Lamport. Distributed Snapshots: Deter- mining Global States of Distributed Systems.ACM Trans. Comput. Syst.3, 1 (Feb. 1985), 63–75. doi:10.1145/214451.214456

  21. [29]

    Spanner: Google’s globally distributed database.ACM Transactions on Computer Systems (TOCS) 31, 3 (2013), 1–22

    James C Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christo- pher Frost, Jeffrey John Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, et al. Spanner: Google’s globally distributed database.ACM Transactions on Computer Systems (TOCS) 31,...

  22. [30]

    The Tail at Scale.Commun

    Jeffrey Dean and Luiz André Barroso. The Tail at Scale.Commun. ACM56, 2 (Feb. 2013), 74–80. doi:10.1145/2408776.2408794

  23. [31]

    Revelio: ML-generated Debugging Queries for Distributed Systems

    Pradeep Dogga, Karthik Narasimhan, Anirudh Sivaraman, Shiv Kumar Saini, George Varghese, and Ravi Netravali. Revelio: ML-generated Debugging Queries for Distributed Systems. doi: 10.48550/arXiv. 2106.14347arXiv:2106.14347 [cs]

  24. [32]

    The Design and Operation of CloudLab

    Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David John- son, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and P...

  25. [33]

    Fowler and W

    J. Fowler and W. Zwaenepoel. Causal Distributed Breakpoints. In Proceedings.,10th International Conference on Distributed Computing Systems. 134–141. doi:10.1109/ICDCS.1990.89277

  26. [34]

    Caladan: Mitigating Interference at Microsecond Timescales

    Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, and Adam Belay. Caladan: Mitigating Interference at Microsecond Timescales. InPro- ceedings of the 14th USENIX Conference on Operating Systems Design and Implementation (OSDI’20). USENIX Association, USA, 281–297. doi:10.5555/348876...

  27. [35]

    Fujimoto

    Richard M. Fujimoto. Parallel Discrete Event Simulation.Commun. ACM33, 10 (Oct. 1990), 30–53. doi:10.1145/84537.84545

  28. [36]

    Sage: Practical and Scalable ML-driven Performance Debugging in Microservices

    Yu Gan, Mingyu Liang, Sundar Dev, David Lo, and Christina Delim- itrou. Sage: Practical and Scalable ML-driven Performance Debugging in Microservices. InProceedings of the 26th ACM International Confer- ence on Architectural Support for Programming Languages and Oper- ating Sy...

  29. [37]

    An Open-Source Benchmark Suite for Microservices and Their Hardware-Software Implications for Cloud & Edge Systems

    Yu Gan, Yanqi Zhang, Dailun Cheng, Ankitha Shetty, Priyal Rathi, Nayan Katarki, Ariana Bruno, Justin Hu, Brian Ritchken, Brendon Jackson, Kelvin Hu, Meghna Pancholi, Yuan He, Brett Clancy, Chris Colen, Fukang Wen, Catherine Leung, Siyuan Wang, Leon Zaruvinsky, Mateo Espinosa, ...

  30. [38]

    Cloud-Native Applications.IEEE Cloud Computing4, 5 (Sept

    Dennis Gannon, Roger Barga, and Neel Sundaresan. Cloud-Native Applications.IEEE Cloud Computing4, 5 (Sept. 2017), 16–21. doi: 10. 1109/MCC.2017.4250939

  31. [39]

    Friday: Global Comprehension for Distributed Replay

    Dennis Geels, Gautam Altekar, Petros Maniatis, Timothy Roscoe, and Ion Stoica. Friday: Global Comprehension for Distributed Replay. In Proceedings of the 4th USENIX Conference on Networked Systems Design & Implementation (NSDI’07). USENIX Association, USA, 21. 13

  32. [40]

    Replay Debugging for Distributed Applications

    Dennis Geels, Gautam Altekar, Scott Shenker, and Ion Stoica. Replay Debugging for Distributed Applications. In2006 USENIX Annual Technical Conference (USENIX ATC 06). USENIX Associa- tion, Boston, MA. https://www.usenix.org/conference/2006- usenix-annual-technical-conference/r...

  33. [41]

    The Google file system

    Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung. The Google file system. InProceedings of the nineteenth ACM symposium on Oper- ating systems principles. 29–43

  34. [42]

    Towards Mod- ern Development of Cloud Applications

    Sanjay Ghemawat, Robert Grandl, Srdjan Petrovic, Michael Whit- taker, Parveen Patel, Ivan Posva, and Amin Vahdat. Towards Mod- ern Development of Cloud Applications. InProceedings of the 19th Workshop on Hot Topics in Operating Systems (HotOS ’23). Asso- ciation for Computing ...

  35. [43]

    Lorch, Bryan Parno, Michael L

    Chris Hawblitzel, Jon Howell, Manos Kapritsos, Jacob R. Lorch, Bryan Parno, Michael L. Roberts, Srinath Setty, and Brian Zill. IronFleet: Proving Practical Distributed Systems Correct. InProceedings of the 25th Symposium on Operating Systems Principles (SOSP ’15). Association ...

  36. [44]

    Wolfcw/Libfaketime

    Wolfgang Hommel. Wolfcw/Libfaketime. https://github.com/ wolfcw/libfaketime

  37. [45]

    Tprof: Performance Profiling via Structural Aggregation and Automated Analysis of Distributed Sys- tems Traces

    Lexiang Huang and Timothy Zhu. Tprof: Performance Profiling via Structural Aggregation and Automated Analysis of Distributed Sys- tems Traces. InProceedings of the ACM Symposium on Cloud Computing. ACM, Seattle WA USA, 76–91. doi:10.1145/3472883.3486994

  38. [46]

    Sambasivan

    Darby Huye, Yuri Shkuro, and Raja R. Sambasivan. Lifting the Veil on Meta’s Microservice Architecture: Analyses of Topology and Request Workflows. In2023 USENIX Annual Technical Conference (USENIX ATC 23). 419–432. https://www.usenix.org/conference/atc23/ presentation/huye

  39. [47]

    Virtual Time.ACM Transactions on Programming Languages and Systems (TOPLAS)(1985)

    David Jefferson. Virtual Time.ACM Transactions on Programming Languages and Systems (TOPLAS)(1985). https://dl.acm.org/doi/ 10.1145/3916.3988

  40. [48]

    Gonzalez, Raluca Ada Popa, Ion Stoica, and David A

    Eric Jonas, Johann Schleier-Smith, Vikram Sreekanti, Chia-Che Tsai, Anurag Khandelwal, Qifan Pu, Vaishaal Shankar, Joao Carreira, Karl Krauth, Neeraja Yadwadkar, Joseph E. Gonzalez, Raluca Ada Popa, Ion Stoica, and David A. Patterson. Cloud Programming Simplified: A Berkeley V...

  41. [49]

    Canopy: An End-to-End Performance Tracing And Analysis System

    Jonathan Kaldor, Jonathan Mace, Michał Bejda, Edison Gao, Wiktor Kuropatwa, Joe O’Neill, Kian Win Ong, Bill Schaller, Pingjia Shan, Brendan Viscomi, Vinod Venkataraman, Kaushik Veeraraghavan, and Yee Jiun Song. Canopy: An End-to-End Performance Tracing And Analysis System. InP...

  42. [50]

    Gunawi, Cody Hammock, Joe Mambretti, Alexander Barnes, François Halbach, Alex Rocha, and Joe Stubbs

    Kate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau, Paul Ruth, Dan Stanzione, Mert Cevik, Jacob Colleran, Haryadi S. Gunawi, Cody Hammock, Joe Mambretti, Alexander Barnes, François Halbach, Alex Rocha, and Joe Stubbs. Lessons Learned from the Chameleon Testbed. InProceeding...

  43. [51]

    Always-On Recording Frame- work for Serverless Computations: Opportunities and Challenges

    Shreyas Kharbanda and Pedro Fonseca. Always-On Recording Frame- work for Serverless Computations: Opportunities and Challenges. In Proceedings of the 1st Workshop on SErverless Systems, Applications and MEthodologies. ACM, Rome Italy, 41–49. doi: 10.1145/3592533. 3592810

  44. [52]

    King, George W

    Samuel T. King, George W. Dunlap, and Peter M. Chen. Debug- ging Operating Systems with Time-Traveling Virtual Machines. InProceedings of the Annual Conference on USENIX Annual Technical Conference (ATEC ’05). USENIX Association, USA, 1. https://www.usenix.org/conference/2005-...

  45. [53]

    Jepsen-Io/Jepsen

    K Kingsbury. Jepsen-Io/Jepsen. https://github.com/jepsen-io/ jepsenAccessed: 2025-09-15

  46. [54]

    Cassandra: a decentralized structured storage system.ACM SIGOPS operating systems review44, 2 (2010), 35–40

    Avinash Lakshman and Prashant Malik. Cassandra: a decentralized structured storage system.ACM SIGOPS operating systems review44, 2 (2010), 35–40

  47. [55]

    Frans Kaashoek, and Zheng Zhang

    Xuezheng Liu, Zhenyu Guo, Xi Wang, Feibo Chen, Xiaochen Lian, Jian Tang, Ming Wu, M. Frans Kaashoek, and Zheng Zhang. D3S: Debugging Deployed Distributed Systems. InProceedings of the 5th USENIX Symposium on Networked Systems Design and Implementation (NSDI’08). USENIX Associa...

  48. [56]

    Pivot Tracing: Dynamic Causal Monitoring for Distributed Systems

    Jonathan Mace, Ryan Roelke, and Rodrigo Fonseca. Pivot Tracing: Dynamic Causal Monitoring for Distributed Systems. InProceedings of the 25th Symposium on Operating Systems Principles (SOSP ’15). Association for Computing Machinery, New York, NY, USA, 378–393. doi:10.1145/28154...

  49. [57]

    Gabime/Spdlog

    Gabi Melman. Gabime/Spdlog. https://github.com/gabime/ spdlog

  50. [58]

    Inc. Meta. Folly Fiber. https://github.com/facebook/folly/ blob/main/folly/fibers/README.mdAccessed: 2026-03-10

  51. [59]

    Miller and J.-D

    B.P. Miller and J.-D. Choi. Breakpoints and Halting in Distributed Programs. In[1988] Proceedings. The 8th International Conference on Distributed. 316–323. doi:10.1109/DCS.1988.12532

  52. [60]

    Jordan, and Ion Stoica

    Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. Ray: A Distributed Framework for Emerging AI Applications. In13th USENIX Symposium on Operating Systems Des...

  53. [61]

    How Amazon Web Services Uses Formal Methods.Commun

    Chris Newcombe, Tim Rath, Fan Zhang, Bogdan Munteanu, Marc Brooker, and Michael Deardeuff. How Amazon Web Services Uses Formal Methods.Commun. ACM58, 4 (March 2015), 66–73. doi: 10. 1145/2699417

  54. [62]

    In Search of an Understand- able Consensus Algorithm

    Diego Ongaro and John Ousterhout. In Search of an Understand- able Consensus Algorithm. In2014 USENIX Annual Technical Confer- ence (USENIX ATC 14). USENIX Association, Philadelphia, PA, 305–

  55. [63]

    https://www.usenix.org/conference/atc14/technical- sessions/presentation/ongaro

  56. [64]

    In Search of an Understandable Consensus Algorithm

    Diego Ongaro and John Ousterhout. In Search of an Understandable Consensus Algorithm. InProceedings of the 2014 USENIX Conference on USENIX Annual Technical Conference (USENIX ATC’14). USENIX Association, USA, 305–320

  57. [65]

    Rumble, Ryan Stutsman, John Ousterhout, and Mendel Rosenblum

    Diego Ongaro, Stephen M. Rumble, Ryan Stutsman, John Ousterhout, and Mendel Rosenblum. Fast Crash Recovery in RAMCloud. InProceed- ings of the Twenty-Third ACM Symposium on Operating Systems Princi- ples. ACM, Cascais Portugal, 29–41. doi: 10.1145/2043556.2043560

  58. [66]

    Shenango: Achieving High CPU Efficiency for Latency-Sensitive Datacenter Workloads

    Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, and Hari Balakrishnan. Shenango: Achieving High CPU Efficiency for Latency-Sensitive Datacenter Workloads. In16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). 361–378. doi:10.5555/3323234.3323265

  59. [67]

    The RAMCloud Storage System.ACM Transactions on Computer Systems33, 3 (Sept

    John Ousterhout, Arjun Gopalan, Ashish Gupta, Ankita Kejriwal, Collin Lee, Behnam Montazeri, Diego Ongaro, Seo Jin Park, Henry Qin, Mendel Rosenblum, Stephen Rumble, Ryan Stutsman, and Stephen Yang. The RAMCloud Storage System.ACM Transactions on Computer Systems33, 3 (Sept. 2...

  60. [68]

    Prometheus: A Next-Generation Monitoring System

    Björn Rabenstein and Julius Volz. Prometheus: A Next-Generation Monitoring System. (2015). https://www.usenix.org/conference/ srecon15europe/program/presentation/rabenstein

  61. [69]

    RecPlay: A Fully Integrated Practical Record/Replay System.ACM Trans

    Michiel Ronsse and Koen De Bosschere. RecPlay: A Fully Integrated Practical Record/Replay System.ACM Trans. Comput. Syst.17, 2 (May 1999), 133–152. doi:10.1145/312203.312214 14

  62. [70]

    Aguilera, Adam Belay, Seo Jin Park, and Malte Schwarzkopf

    Zhenyuan Ruan, Shihang Li, Kaiyan Fan, Marcos K. Aguilera, Adam Belay, Seo Jin Park, and Malte Schwarzkopf. Unleashing True Utility Computing with Quicksand. InProceedings of the 19th Workshop on Hot Topics in Operating Systems. ACM, Providence RI USA, 196–205. doi:10.1145/359...

  63. [71]

    Aguilera, Adam Belay, and Malte Schwarzkopf

    Zhenyuan Ruan, Shihang Li, Kaiyan Fan, Seo Jin Park, Marcos K. Aguilera, Adam Belay, and Malte Schwarzkopf. Quicksand: Harness- ing Stranded Datacenter Resources with Granular Computing. In22nd USENIX Symposium on Networked Systems Design and Implementa- tion (NSDI 25). 147–16...

  64. [72]

    Aguilera, Adam Belay, and Malte Schwarzkopf

    Zhenyuan Ruan, Seo Jin Park, Marcos K. Aguilera, Adam Belay, and Malte Schwarzkopf. Nu: Achieving Microsecond-Scale Resource Fungi- bility with Logical Processes. In20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 1409–1427. https: //www.usenix.o...

  65. [73]

    AIFM: High-Performance, Application-Integrated Far Memory

    Zhenyuan Ruan, Malte Schwarzkopf, Marcos K Aguilera, and Adam Belay. AIFM: High-Performance, Application-Integrated Far Memory. InNSDI

  66. [74]

    XFaaS: Hyperscale and Low Cost Serverless Functions at Meta

    Alireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang, Abhigna Nagaraja, Neeraj Pathak, Girish Joshi, Carla Souza, Bo Huang, Wyatt Cook, Andrii Golovei, Pradeep Venkat, Andrew Mcfague, Dim- itrios Skarlatos, Vipul Patel, Ravinder Thind, Ernesto Gonzalez, Yun Jin, a...

  67. [75]

    Network-Centric Distributed Tracing with DeepFlow: Trou- bleshooting Your Microservices in Zero Code

    Junxian Shen, Han Zhang, Yang Xiang, Xingang Shi, Xinrui Li, Yunxi Shen, Zijian Zhang, Yongxiang Wu, Xia Yin, Jilong Wang, Mingwei Xu, Yahui Li, Jiping Yin, Jianchang Song, Zhuofeng Li, and Runjie Nie. Network-Centric Distributed Tracing with DeepFlow: Trou- bleshooting Your M...

  68. [76]

    Dapper, a Large-Scale Distributed Systems Tracing Infrastructure

    Benjamin H Sigelman, Luiz Andre Barroso, Mike Burrows, Pat Stephen- son, Manoj Plakal, Donald Beaver, Saul Jaspan, and Chandan Shanbhag. Dapper, a Large-Scale Distributed Systems Tracing Infrastructure. ([n. d.])

  69. [77]

    Modular Monolith: Is This the Trend in Software Architecture?

    Ruoyu Su and Xiaozhou Li. Modular Monolith: Is This the Trend in Software Architecture?. InProceedings of the 1st International Workshop on New Trends in Software Architecture. ACM, Lisbon Portugal, 10–13. doi:10.1145/3643657.3643911

  70. [78]

    Vearne/Grpcreplay

    Vearne. Vearne/Grpcreplay. https://github.com/vearne/ grpcreplay

  71. [79]

    Bond, Ravi Netravali, Miryung Kim, and Guo- qing Harry Xu

    Chenxi Wang, Haoran Ma, Shi Liu, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen, Michael D. Bond, Ravi Netravali, Miryung Kim, and Guo- qing Harry Xu. Semeru: A {Memory-Disaggregated} Managed Runtime. In14th USENIX Symposium on Operating Systems Design and Implemen- tation (OSDI 20). ...

  72. [80]

    Model Checking Guided Testing for Distributed Systems

    Dong Wang, Wensheng Dou, Yu Gao, Chenao Wu, Jun Wei, and Tao Huang. Model Checking Guided Testing for Distributed Systems. In Proceedings of the Eighteenth European Conference on Computer Systems. ACM, Rome Italy, 127–143. doi:10.1145/3552326.3587442

  73. [81]

    Wilcox, Doug Woos, Pavel Panchekha, Zachary Tatlock, Xi Wang, Michael D

    James R. Wilcox, Doug Woos, Pavel Panchekha, Zachary Tatlock, Xi Wang, Michael D. Ernst, and Thomas Anderson. Verdi: A Framework for Implementing and Formally Verifying Distributed Systems. In Proceedings of the 36th ACM SIGPLAN Conference on Programming Language Design and Im...

  74. [82]

    doi:10.1145/2737924.2737958

  75. [83]

    LogRCA: Log- Based Root Cause Analysis for Distributed Services

    Thorsten Wittkopp, Philipp Wiesner, and Odej Kao. LogRCA: Log- Based Root Cause Analysis for Distributed Services. doi: 10.48550/ arXiv.2405.13599arXiv:2405.13599 [cs]

  76. [84]

    MODIST: Transparent Model Checking of Unmodified Distributed Systems

    Junfeng Yang, Tisheng Chen, Ming Wu, Zhilei Xu, Xuezheng Liu, Haoxiang Lin, Mao Yang, Fan Long, Lintao Zhang, and Lidong Zhou. MODIST: Transparent Model Checking of Unmodified Distributed Systems. 213–228

  77. [85]

    Jain, and Michael Stumm

    Ding Yuan, Yu Luo, Xin Zhuang, Guilherme Renna Rodrigues, Xu Zhao, Yongle Zhang, Pranay U. Jain, and Michael Stumm. Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems. In11th USENIX Sympo- sium on Operati...

  78. [86]

    https://www.usenix.org/conference/osdi14/technical- sessions/presentation/yuan

  79. [87]

    Characterizing Log- ging Practices in Open-Source Software

    Ding Yuan, Soyeon Park, and Yuanyuan Zhou. Characterizing Log- ging Practices in Open-Source Software. InProceedings of the 34th International Conference on Software Engineering (ICSE ’12). IEEE Press, Zurich, Switzerland, 102–112. https://dl.acm.org/doi/10.5555/ 2337223.2337236

  80. [88]

    The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems

    Lei Zhang, Zhiqiang Xie, Vaastav Anand, Ymir Vigfusson, and Jonathan Mace. The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems. In20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 321–339. https://www.usenix.org/ conference/nsdi23/...

  81. [89]

    Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study.IEEE Transactions on Software Engineering47, 2 (Feb

    Xiang Zhou, Xin Peng, Tao Xie, Jun Sun, Chao Ji, Wenhai Li, and Dan Ding. Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study.IEEE Transactions on Software Engineering47, 2 (Feb. 2021), 243–260. doi: 10. 1109/TSE.2018.2887384

  82. [90]

    Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study.IEEE Trans

    Xiang Zhou, Xin Peng, Tao Xie, Jun Sun, Chao Ji, Wenhai Li, and Dan Ding. Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study.IEEE Trans. Softw. Eng.47, 2 (Feb. 2021), 243–260. doi: 10.1109/TSE.2018. 2887384

  83. [91]

    caller_id

    Gefei Zuo, Jiacheng Ma, Andrew Quinn, Pramod Bhatotia, Pedro Fonseca, and Baris Kasikci. Execution Reconstruction: Harnessing Failure Reoccurrences for Failure Reproduction. InProceedings of the 42nd ACM SIGPLAN International Conference on Programming Lan- guage Design and Imp...

  84. [92]

    Presence Verification.For every reconstructed caller frame,DDBqueries the local runtime to verify whether the caller’s target heap remains physically present in the local address space

  85. [93]

    Surgical Heap Restoration.If the heap is absent,DDB queries the framework’s locator API to pinpoint the cur- rent physical owner. It then coordinates a targeted remote dump of the latest synchronized heap state.DDBphysi- cally writes these bytes into the reconstructed caller’s...

  86. [94]

    The process memory state rigidly reverts to its pristine original configuration before execution resumes

    Ephemeral Cleanup.Upon intercepting a continue command,DDBaggressively scrubs any temporary restorations. The process memory state rigidly reverts to its pristine original configuration before execution resumes. 21

Pith tools

Reviewed July 8, 2026 · model on record in the stance chip above.