Pith. sign in

REVIEW 3 major objections 6 minor 148 references

Bodega: Serving Linearizable Reads Locally from Anywhere at Anytime via Roster Leases

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Bodega claims to be the first consensus protocol that serves linearizable reads locally from any designated replica even while interfering writes are in progress, via all-to-all roster leases.

desk verdict The roster-lease idea is genuinely new and the evaluation is strong, but the transition logic as written has a safety bug that breaks the proof; a major revision is needed. read the letter →

arxiv 2509.07158 v1 pith:N5XHQJTW submitted 2025-09-08 cs.DC

classification cs.DC
keywords consensuslinearizablereadsleasesrosterlocalreplicationPaxoswide-areasystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Linearizable reads — reads that observe every acknowledged write in real time — usually force a client to contact a leader or a quorum, which is slow across wide-area networks. Bodega claims to break this constraint: any replica designated as a responder can answer reads locally, at any time, even while writes to the same key are actively committing. The protocol protects a generalization of leadership called the roster with all-to-all roster leases, and requires writes to reach all active responders before committing. The paper reports 5.6x–13.1x faster average read latency than prior approaches under moderate write interference, with comparable write performance, in a replicated key-value store implementation.

What carries the argument

The central mechanism is the roster lease, an all-to-all generalization of the one-to-one leader lease. Each replica grants timed, renewable leases for a specific roster (a ballot number together with the leader and per-key responder assignments) to every peer; piggybacked on heartbeats, these leases establish a stable roster when a node holds a majority of grants. A stable-roster check also requires the node to have committed every slot below a safety threshold learned from peers, and local reads are served only after that check passes. This off-critical-path leasing is what makes local reads safe: a write cannot commit before all active responders have acknowledged it, so a responder with a stable roster never misses a committed value.

What would settle it

Crash or suspend a node long enough for clock drift or lease timers to violate the bounded-drift assumption, then check whether two nodes can simultaneously hold majority lease sets for different roster ballots; if the protocol's own stable condition ever allows that, a local read can miss a committed write and the central claim is false. A cheaper version: instrument the implementation to log grantor-side and grantee-side expiration times under skewed clocks and test whether the guard-then-renew invariant from the lease background is ever reversed.

Watch

Extended reading notes

Core claim

On its own terms, Bodega's discovery is that a consensus protocol can keep reads local without quiet periods or leader round-trips if the cluster agrees not just on who leads but on who may answer reads. That agreement is the roster, and it is protected by roster leases: every node grants timed leases to every other node for the current roster ballot, so holding leases from a majority of nodes proves that no other stable roster exists. Once that condition holds, a responder can serve the latest committed value directly, and when an interfering write is still in flight it can optimistically hold the read until the write commits or enough early accept notifications arrive. Bodega therefore locates reads at arbitrary replicas at all times, at the cost of requiring writes to be acknowledged by all responders.

Load-bearing premise

The protocol's safety rests on the assumption that node clocks never drift apart by more than a fixed small bound, so that lease expiration times cannot overlap in a way that lets two competing rosters both appear stable.

Editorial extensions

If this is right

  • A client near any responder gets read latency close to a local memory access even while writes to the same key are continuous; the only degraded period after an interfering write lasts about half a majority round trip.
  • Writes carry a modest tax: they must collect replies from all responders of the key, not just a majority, but the paper measures write throughput as comparable to classic consensus.
  • Roster changes can be applied proactively in two message rounds, while failure-triggered changes wait for lease expiration, preserving availability under minority faults.
  • The protocol retains the fault tolerance of classic consensus and needs no external membership or metadata service.
  • In YCSB workloads, Bodega matches the performance of sequentially consistent etcd and ZooKeeper, which do not offer linearizable local reads.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Roster leases may protect other rare metadata beyond read responders — membership, asymmetric quorums, node-specific hints — extending local-read safety to those mutable settings.
  • The read-locality benefit is largest when the roster is tuned to read-heavy, location-skewed keys; an adaptive controller that changes responders as workload skew shifts could realize the paper's anytime promise in production, but that policy is not developed here.
  • The safety of the scheme depends on bounded clock drift; deployments could add clock-drift monitoring and refuse local reads when drift approaches the bound, a testable safeguard the paper does not specify.
  • The same all-to-all lease idea might extend to Byzantine fault tolerance by raising the majority threshold to 2f+1 and reusing early accept notifications, but Bodega itself targets crash faults.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents Bodega, a consensus protocol that aims to serve linearizable reads locally at designated responder replicas even while writes are in flight. It introduces the roster, a ballot-tagged cluster metadata that assigns a leader and per-key responder sets, and protects roster agreement with all-to-all 'roster leases' that generalize leader leases. The authors argue that a node holding a majority of lease grants can serve reads locally, that writes commit only after reaching all active responders, and that optimistic holding and early accept notifications keep reads local during write interference. They implement Bodega and prior work in Vineyard, an async-Rust KV store, and report 5.6–13.1x read-latency improvements over prior linearizable-read protocols on WAN clusters, comparable write throughput, fast roster changes, and YCSB performance matching sequentially consistent etcd and ZooKeeper. Correctness is argued via a short proof sketch (§4.2) and a PlusCal/TLA+ model checked for a 3-node configuration (Appendix B).

Significance. If the protocol as written were sound, this would be a meaningful advance: it is the first self-contained consensus-based scheme we are aware of that extends local linearizable reads to arbitrary responders under active write interference, and the roster abstraction is a clean generalization of leadership. The implementation breadth and comparison against several protocols and production systems are valuable, and the reported speedups are substantial and measured rather than fitted to the claims. The paper also ships machine-checked TLA+ (though, as discussed below, an idealized lease model). The central correctness claim, however, rests on a stale-lease transition invariant that the detailed pseudocode does not preserve, so the contribution is currently not established as presented.

major comments (3)
  1. [§3.3.2, Figure 17; §4.2 Case #1] The revoke procedure does not implement the prose. §3.3.2 says a node S 'invokes the revoke_leases(bal) procedure synchronously to ensure that all the leases it is granting or holding with the older ballot bal are safely revoked and removed,' but do_revoke_leases (Figure 17 lines 15–19) clears only {}guarding (S's grantor-side guard set) and waits only on {}endowing; it never purges S's grantee-side {}endowed or {}guarded. Moreover, Recv leaseRevoke (line 18) removes the sender only if bal >= current ballot, so a Revoke carrying an old ballot is ignored by a node that has already adopted a higher ballot, and the stale grant remains until timer expiry. Consequently, when S adopts a new ballot bal' while peers are still on bal, S's {}endowed can still contain a majority of old-ballot grants; Condition (1) in §3.3.1 can then hold for bal' even though no majority has adopted bal'. In that window an old-ballot leader can commit a write without S, and S can serve a local read on bal' that misses that committed write, violating Case #1 of the proof in §4.2 and the claimed 'at most one stable roster' invariant. The pseudocode must clear all grantee-side entries whose roster ballot is not current at the moment of adopting a new ballot, and the Recv leaseRevoke predicate must also remove stale grants when the received ballot is older than the current ballot.
  2. [Appendix B (TLA+ model)] The model-checked TLA+ specification does not exercise the revocation/transition logic that is the basis of the bug above. The global `grants` set and the `Lease` macro replace each grantor's grant with a single grant to the new roster, and the comment explicitly says messages are 'removed' to model expiry 'probably making way for switching to a different roster.' There are no per-grantee {}endowed sets, no Guard/Revoke messages, and no ordering constraint on revocation before adopting a higher ballot. The check of 3 nodes, 3 ballots, 2 writes, and 2 reads therefore cannot detect stale-grant violations of Condition (1), and the statement in Contribution 5 that the TLA+ specification verifies the protocol should be scoped to the idealized lease model.
  3. [§4.2 Case #3] The proof of Case #3 assumes that for any size-m subset E of S's {}endowed, at least one grantor P accepted the committed write W 'before granting to S,' implying thresh_p >= x. But thresh_p is recorded only in the initial Guard (Figure 17 line 24), and renewals (lines 25–28) do not update it. If the grant counted in E is an old-ballot grant that was established before W committed, thresh_p can be below x even though P is in a majority; the argument requires either that grants carry the ballot and are refreshed with fresh thresholds, or that old grants are provably purged at roster transitions. Neither is guaranteed by the present pseudocode.
minor comments (6)
  1. [Figure 17, line 18] The guard 'bal >= current ballot' for Recv leaseRevoke is not explained and appears to have the wrong sign for the stale-grant case; the intended semantics should be stated in the text.
  2. [Appendix B] The phrase 'cheated' model in the inlined comments should be reflected in the main text's description of the formal verification, so readers know the TLA+ result does not cover the lease revocation ordering.
  3. [§6, Figures 8–12, 14–15] The evaluation figures do not report error bars, confidence intervals, or the number of runs; please state the run-to-run variability.
  4. [Table 1] Table 1 cells contain '#' and 'G' markers that are not defined in the caption; the caption only defines the filled circle symbols.
  5. [§6.2] The sentence beginning 'Bodega 1 shortens' should read 'Bodega (1) shortens'.
  6. [Figure 5] Figure 5's legend ('#20 x4', '#32 x1', '#11 x1') is cryptic without a pointer to §3.3; consider explaining the notation in the caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the protocol behavior follows from its stated lease and commit conditions and is validated by external measurements.

full rationale

The paper's derivation chain is self-contained: the roster and roster-lease mechanism are defined in §3.1 and §3.3 from standard lease primitives, the write commit condition ('must also have received replies from all the responders') and the stable condition (|endowed| >= m plus threshold catch-up) are protocol definitions rather than fitted outputs, and the §4.2 linearizability proof reduces to majority intersection and lease-expiration properties. The performance claims are supported by measured CloudLab/GEO experiments against external baselines (etcd, ZooKeeper, Leader Leases, EPaxos, PQR, Quorum Leases) rather than by fitting the model to the results. The only self-citation (ref. [63], a survey by the same authors) appears alongside the standard Herlihy-Wing reference for the definition of linearizability and is not load-bearing. The appendix TLA+ model is a model-checking sanity check; the skeptic's concern that the Revoke/endowed transition is not fully modeled is a potential soundness gap, not circular reasoning. Accordingly, no circular step is identified.

Assumptions & free parameters 5 free parameters · 5 assumptions · 1 invented entities

The protocol rests on standard lease assumptions (bounded clock drift, lease safety), standard majority-quorum math, and a new 'roster' metadata abstraction. The timing parameters (heartbeats, lease durations, batching) are hand-chosen default values used in the evaluation, not derived from the underlying theory.

free parameters (5)
  • t_hb_send (heartbeat interval) = 120 ms
    Chosen default for wide-area clusters; affects lease renewal frequency and failure detection speed.
  • t_hb_fail (failure timeout) = 1200 ms (±300 ms randomized)
    Used to detect failed peers; contributes to roster change latency.
  • t_guard and t_lease (guard/lease durations) = 2500 ms (±100 ms randomized)
    Determines how long a lease is valid before renewal; parameter for the leasing mechanism.
  • request batching interval = 1 ms
    Used in the Vineyard implementation to batch non-local-read commands; directly impacts measured latencies.
  • Smart roster coverage thresholds = >95% read ratio; >20% of reads from a site
    Policy for when to automatically add responders; not central to correctness but used in YCSB experiments.
assumptions (5)
  • domain assumption Bounded clock drift between nodes by t_Delta
    Stated in Section 2.2 and Section 3.3; required for lease safety. No hard bound is derived; it is assumed to hold in cloud environments.
  • domain assumption Asynchronous network with fail-stop/fail-slow nodes
    Standard consensus model; assumed from Section 2.1.
  • standard math Majority quorum intersection (2m > n)
    Used in the linearizability proof in Section 4.2 to argue that any two majority sets intersect.
  • standard math Safety and liveness of the standard lease primitive
    The proof in Section 4.2 relies on 'well-established results of the safety and liveness of leases [51]'.
  • domain assumption No blind writes
    Assumed in Section 2.1 to restrict scope to non-transactional commands; excludes blind writes.
invented entities (1)
  • Roster
    purpose: Cluster metadata that records the leader and, for each key, the set of responder nodes allowed to serve local reads.
    A new conceptual abstraction introduced by the paper. It is not an external entity; its correctness is established by the protocol's proof and evaluation, not by independent evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bodega: Serving Linearizable Reads Locally from Anywhere at Anytime via Roster Leases." pith.science (2026). https://pith.science/paper/N5XHQJTW

@misc{pith2026250907158,
  author       = {Pith},
  title        = {Pith review of: Bodega: Serving Linearizable Reads Locally from Anywhere at Anytime via Roster Leases},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N5XHQJTW}},
  note         = {Machine review of arXiv:2509.07158}
}
read the original abstract

We present Bodega, the first consensus protocol that serves linearizable reads locally from any desired node, regardless of interfering writes. Bodega achieves this via a novel roster leases algorithm that safeguards the roster, a new notion of cluster metadata. The roster is a generalization of leadership; it tracks arbitrary subsets of replicas as responder nodes for local reads. A consistent agreement on the roster is established through roster leases, an all-to-all leasing mechanism that generalizes existing all-to-one leasing approaches (Leader Leases, Quorum Leases), unlocking a new point in the protocol design space. Bodega further employs optimistic holding and early accept notifications to minimize interruption from interfering writes, and incorporates smart roster coverage and lightweight heartbeats to maximize practicality. Bodega is a non-intrusive extension to classic consensus; it imposes no special requirements on writes other than a responder-covering quorum. We implement Bodega and related works in Vineyard, a protocol-generic replicated key-value store written in async Rust. We compare it to previous protocols (Leader Leases, EPaxos, PQR, and Quorum Leases) and two production coordination services (etcd and ZooKeeper). Bodega speeds up average client read requests by 5.6x-13.1x on real WAN clusters versus previous approaches under moderate write interference, delivers comparable write performance, supports fast proactive roster changes as well as fault tolerance via leases, and closely matches the performance of sequentially-consistent etcd and ZooKeeper deployments across all YCSB workloads. We will open-source Vineyard upon publication.

Figures

Figures reproduced from arXiv: 2509.07158 by the authors.

Figure 1
Figure 1. Frequency of touching a node on the critical [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Demonstration of standard leasing. Left: the guard phase establishes the first iteration of promise coverage; grantee welcomes the first Renew only if it is received within the guarded period (C < A’). This allows the grantor to derive a safe D’ = B’+𝑡lease+𝑡Δ even if the RenewReply is lost, such that C’ < D’. Right: the grantor attempts to extend the promise with a Renew (or to actively revoke it with a Revoke), bu… view at source ↗
Figure 4
Figure 4. Normal case operations of Bodega. Assume all nodes agree on the same example roster: S0 is the leader and S0,2-4 are responders for a key while S1 is not. See §3.2. 3.2.1 Writes Writes follow the same leader-based process as in MultiPaxos ( [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (9 more)
Figure 5
Figure 5. Figure 5: All-to-all roster leases demonstrated. S0,3,4 are each holding ≥ majority grants of roster #20; S4 has not seen all slots up to #20’s safety threshold. S1 is stuck with an older roster of #11. S2 is initiating a new roster of #32. See §3.3. have observed the latest com…
Figure 6
Figure 6. Figure 6: Timeline comparison across protocols of linearizable reads [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Evaluation settings. Orange denotes designated leader node and Red denotes other responders, if relevant. The edges mark per-pair RTT values in milliseconds. See §6. Similarly, clients may request a server to send the roster along with a command reply, and then cache t…
Figure 8
Figure 8. Figure 8: Normalized throughput and latency at different client locations [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Latency CDFs of requests in the WAN setting across different write intensities, focusing on one specific key. See §6.2. 0 25 50 75 100 125 150 175 Time (ms) 0 20 40 Read Lat (ms) write arrives Quorum Leases Qrm Ls (passive) Bodega [PITH_FULL_IMAGE:figures/full_fig_p01…
Figure 10
Figure 10. Figure 10: We see that the write introduces an interruption [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]
Figure 13
Figure 13. Figure 13: Failure-triggered vs. regular roster changes in Bodega. The x-axis is time at which reqs finish. See §6.3. WI +MA +SC +APT +UT Responders Set 0 20 40 60 80 Latency (ms) Write Read [PITH_FULL_IMAGE:figures/full_fig_p011_13.png]
Figure 16
Figure 16. Figure 16: YCSB workloads on Vineyard, etcd, & ZooKeeper. [PITH_FULL_IMAGE:figures/full_fig_p011_16.png]
Figure 17
Figure 17. Figure 17: Complete Summary of the Bodega algorithm. Lists all actions a node 𝑟 would take upon certain conditions, grouped by purposes for clarity: triggers for a new roster, granting procedure of new roster leases, heartbeats and lease renewals, handling client write requests,…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

148 extracted references · 17 canonical work pages

  1. [1]

    Aguilera, Naama Ben-David, Rachid Guerraoui, Virendra J

    Marcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, Virendra J. Marathe, Athanasios Xygkis, and Igor Zablotchi. 2020. Microsecond Consensus for Microsecond Applications. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 599–616.https://www.usenix.org/conference/osdi20/p resentation/aguilera

  2. [2]

    Aguilera, Naama Ben-David, Rachid Guerraoui, An- toine Murat, Athanasios Xygkis, and Igor Zablotchi

    Marcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, An- toine Murat, Athanasios Xygkis, and Igor Zablotchi. 2023. UBFT: Microsecond-Scale BFT Using Disaggregated Memory. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Vancouver, BC, Canada)(ASPLOS 2023). Associati...

  3. [3]

    Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, and Tevfik Kosar. 2020. WPaxos: Wide Area Network Flexible Consensus.IEEE Trans. Parallel Distrib. Syst.31, 1 (Jan. 2020), 211–223. doi:10.1109/ TPDS.2019.2929793

  4. [4]

    Mohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, and Samer Al-Kiswany. 2020. Toward a Generic Fault Tolerance Technique for Partial Network Partitioning. In14th USENIX Symposium on Operat- ing Systems Design and Implementation (OSDI 20). USENIX Associa- tion, 351–368.https://www.usenix.org/conference/osdi20/presentat ion/alfatafta

  5. [5]

    Sérgio Almeida, João Leitão, and Luís Rodrigues. 2013. ChainReac- tion: a causal+ consistent datastore based on chain replication. In Proceedings of the 8th ACM European Conference on Computer Systems (Prague, Czech Republic)(EuroSys ’13). Association for Computing Machinery, New York, NY, USA, 85–98. doi:10.1145/2465351.2465361

  6. [6]

    Artem on StackOverflow. 2017. Is ZooKeeper always consistent in terms of CAP theorem?https://stackoverflow.com/questions/353877

  7. [7]

    Balaji Arun, Sebastiano Peluso, Roberto Palmieri, Giuliano Losa, and Binoy Ravindran. 2017. Speeding up Consensus by Chasing Fast Decisions. In47th IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). 49–60. doi:10.1109/DSN.2017.35

  8. [8]

    Hagit Attiya and Jennifer L. Welch. 1994. Sequential Consistency versus Linearizability.ACM Trans. Comput. Syst.12, 2 (may 1994), 91–122. doi:10.1145/176575.176576

Show all 148 references
  1. [9]

    AWS. 2024. AWS Global Infrastructure.https://aws.amazon.com/a bout-aws/global-infrastructure/, Last accessed on 2024-04-28

  2. [10]

    AWS. 2024. Workload Characteristics.https://docs.aws.amazon.com/ prescriptive-guidance/latest/oracle-exadata-blueprint/workload- characteristics.html. Accessed: 2024-12-01

  3. [11]

    Corbett, JJ Furman, Andrey Khor- lin, James Larson, Jean-Michel Leon, Yawei Li, Alexander Lloyd, and Vadim Yushprakh

    Jason Baker, Chris Bond, James C. Corbett, JJ Furman, Andrey Khor- lin, James Larson, Jean-Michel Leon, Yawei Li, Alexander Lloyd, and Vadim Yushprakh. 2011. Megastore: Providing Scalable, Highly Available Storage for Interactive Services. InProceedings of the Conference on In...

  4. [12]

    Mahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi, Ahmed Jafri, Xiao Shi, Santosh Ghosh, Hazem Hassan, Aaryaman Sagar, Rhed Shi, Jingming Liu, Filip Gruszczynski, Xianan Zhang, Huy Hoang, Ahmed Yossef, Francois Richard, and Yee Jiun Song

  5. [13]

    Davis, Vijayan Prab- hakaran, Michael Wei, and Ted Wobber

    Mahesh Balakrishnan, Dahlia Malkhi, John D. Davis, Vijayan Prab- hakaran, Michael Wei, and Ted Wobber. 2013. CORFU: A distributed shared log.ACM Trans. Comput. Syst.31, 4, Article 10 (Dec. 2013), 24 pages. doi:10.1145/2535930

  6. [14]

    Davis, Sriram Rao, Tao Zou, and Aviad Zuck

    Mahesh Balakrishnan, Dahlia Malkhi, Ted Wobber, Ming Wu, Vijayan Prabhakaran, Michael Wei, John D. Davis, Sriram Rao, Tao Zou, and Aviad Zuck. 2013. Tango: Distributed Data Structures over a Shared Log. InProceedings of the Twenty-Fourth ACM Symposium on Operating Systems Prin...

  7. [16]

    Jeff Barr. 2023. Amazon S3 Update – Strong Read-After-Write Con- sistency.https://aws.amazon.com/blogs/aws/amazon-s3-update- strong-read-after-write-consistency/, Last accessed on 2023-11-19

  8. [17]

    Michael Ben-Or. 1983. Another advantage of free choice (Extended Abstract): Completely asynchronous agreement protocols. InPro- ceedings of the Second Annual ACM Symposium on Principles of Dis- tributed Computing(Montreal, Quebec, Canada)(PODC ’83). As- sociation for Computing...

  9. [18]

    Hal Berenson, Phil Bernstein, Jim Gray, Jim Melton, Elizabeth O’Neil, and Patrick O’Neil. 1995. A critique of ANSI SQL isolation levels. InProceedings of the 1995 ACM SIGMOD International Conference on Management of Data(San Jose, California, USA)(SIGMOD ’95). Association for ...

  10. [19]

    Reiser, and Alysson Bessani

    Christian Berger, Hans P. Reiser, and Alysson Bessani. 2021. Making Reads in BFT State Machine Replication Fast, Linearizable, and Live. In2021 40th International Symposium on Reliable Distributed Systems (SRDS). 1–12. doi:10.1109/SRDS53918.2021.00010

  11. [20]

    Carlos Eduardo Bezerra, Fernando Pedone, and Robbert Van Renesse

  12. [21]

    Changyu Bi, Vassos Hadzilacos, and Sam Toueg. 2022. Pa- rameterized algorithm for replicated objects with local reads. arXiv:2204.01228 [cs.DC]https://arxiv.org/abs/2204.01228

  13. [22]

    Marc Brooker, Tao Chen, and Fan Ping. 2020. Millions of tiny databases. InNSDI 2020.https://www.amazon.science/publica tions/millions-of-tiny-databases

  14. [23]

    Matthew Burke, Audrey Cheng, and Wyatt Lloyd. 2020. Gryff: Uni- fying Consensus and Shared Registers. In17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20). USENIX Association, Santa Clara, CA, 591–617.https://www.usenix.org/con ference/nsdi20/presen...

  15. [24]

    Mike Burrows. 2006. The Chubby Lock Service for Loosely-Coupled Distributed Systems. InProceedings of the 7th Symposium on Operat- ing Systems Design and Implementation(Seattle, Washington)(OSDI ’06). USENIX Association, USA, 335–350

  16. [25]

    Georgia Butler. 2024. Google Cloud accidentally deleted UniSuper’s Private Cloud Subscription.Data Center Dynamics(2024).https: //www.datacenterdynamics.com/en/news/google-cloud-accidental ly-deleted-unisupers-private-cloud-subscription/

  17. [26]

    Miguel Castro and Barbara Liskov. 1999. Practical Byzantine Fault Tolerance. InThird Symposium on Operating Systems Design and Implementation (OSDI). USENIX Association, Co-sponsored by IEEE TCOS and ACM SIGOPS, New Orleans, Louisiana

  18. [27]

    Miguel Castro and Barbara Liskov. 2002. Practical byzantine fault tolerance and proactive recovery.ACM Trans. Comput. Syst.20, 4 (Nov. 2002), 398–461. doi:10.1145/571637.571640

  19. [28]

    Data Centre. 2024. Alibaba Cloud hit by Digital Realty fire in Singa- pore.Frontier Enterprise(2024).https://www.frontier-enterprise.c 13 om/alibaba-cloud-hit-by-digital-realty-fire-in-singapore/

  20. [29]

    Chandra, Robert Griesemer, and Joshua Redstone

    Tushar D. Chandra, Robert Griesemer, and Joshua Redstone. 2007. Paxos Made Live: An Engineering Perspective. InProceedings of the 26th ACM Symposium on Principles of Distributed Computing(Port- land, OR, USA)(PODC ’07). Association for Computing Machinery, New York, NY, USA, 3...

  21. [30]

    Chandra, Vassos Hadzilacos, and Sam Toueg

    Tushar D. Chandra, Vassos Hadzilacos, and Sam Toueg. 2016. An Algorithm for Replicated Objects with Efficient Reads. InProceedings of the 2016 ACM Symposium on Principles of Distributed Computing (Chicago, Illinois, USA)(PODC ’16). Association for Computing Ma- chinery, New Yo...

  22. [31]

    Aleksey Charapko, Ailidani Ailijiang, and Murat Demirbas. 2019. Linearizable Quorum Reads in Paxos. In11th USENIX Workshop on Hot Topics in Storage and File Systems (HotStorage 19). USENIX Asso- ciation, Renton, WA.https://www.usenix.org/conference/hotstora ge19/presentation/charapko

  23. [32]

    Aleksey Charapko, Ailidani Ailijiang, and Murat Demirbas. 2021. PigPaxos: Devouring the Communication Bottlenecks in Distributed Consensus. InProceedings of the 2021 International Conference on Management of Data(Virtual Event, China)(SIGMOD ’21). Asso- ciation for Computing M...

  24. [33]

    Inho Choi, Ellis Michael, Yunfan Li, Dan R. K. Ports, and Jialin Li. 2023. Hydra: Serialization-Free Network Ordering for Strongly Consistent Distributed Applications. In20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Boston,...

  25. [34]

    Google Cloud. 2024. Google Cloud locations.https://cloud.google.c om/about/locations/, Last accessed on 2024-11-30

  26. [35]

    Cooper, Adam Silberstein, Erwin Tam, Raghu Ramakrishnan, and Russell Sears

    Brian F. Cooper, Adam Silberstein, Erwin Tam, Raghu Ramakrishnan, and Russell Sears. 2010. Benchmarking Cloud Serving Systems with YCSB. InProceedings of the 1st ACM Symposium on Cloud Computing (Indianapolis, IN, USA)(SoCC ’10). Association for Computing Ma- chinery, New York...

  27. [36]

    Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, J

    James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson Hsieh, Sebastian Kan- thak, Eugene Kogan, Hongyi Li, Alexander Lloyd, Sergey Melnik, David Mwaura, Davi...

  28. [37]

    Heming Cui, Rui Gu, Cheng Liu, Tianyu Chen, and Junfeng Yang

  29. [38]

    Huynh Tu Dang, Pietro Bressana, Han Wang, Ki Suh Lee, Noa Zil- berman, Hakim Weatherspoon, Marco Canini, Fernando Pedone, and Robert Soulé. 2020. P4xos: Consensus as a Network Service. IEEE/ACM Trans. Netw.28, 4 (Aug. 2020), 1726–1738. doi:10.1109/TN ET.2020.2992106

  30. [39]

    Huynh Tu Dang, Daniele Sciascia, Marco Canini, Fernando Pedone, and Robert Soulé. 2015. NetPaxos: consensus at network speed. In Proceedings of the 1st ACM SIGCOMM Symposium on Software Defined Networking Research(Santa Clara, California)(SOSR ’15). Association for Computing M...

  31. [40]

    Cong Ding, David Chu, Evan Zhao, Xiang Li, Lorenzo Alvisi, and Robbert Van Renesse. 2020. Scalog: Seamless Reconfiguration and Total Order in a Scalable Shared Log. In17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20). USENIX Association, Santa Clar...

  32. [41]

    Jiaqing Du, Daniele Sciascia, Sameh Elnikety, Willy Zwaenepoel, and Fernando Pedone. 2014. Clock-RSM: Low-Latency Inter-datacenter State Machine Replication Using Loosely Synchronized Physical Clocks. In2014 44th Annual IEEE/IFIP International Conference on Dependable Systems ...

  33. [42]

    Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David John- son, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and P...

  34. [43]

    Vitor Enes, Carlos Baquero, Tuanir França Rezende, Alexey Gotsman, Matthieu Perrin, and Pierre Sutra. 2020. State-Machine Replication for Planet-Scale Systems. InProceedings of the Fifteenth European Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20). Association f...

  35. [44]

    etcd. 2023. etcd: A distributed, reliable key-value store for the most critical data.https://etcd.io/, Last accessed on 2023-11-13

  36. [45]

    FireScroll. 2023. FireScroll: The config database to deploy everywhere. https://github.com/FireScroll/FireScroll, Last accessed on 2024-09-05

  37. [46]

    Pedro Fouto, Nuno Preguiça, and Joao Leitão. 2022. High Through- put Replication with Integrated Membership Management. In2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX As- sociation, Carlsbad, CA, 575–592.https://www.usenix.org/confere nce/atc22/presentation/fouto

  38. [47]

    Aishwarya Ganesan, Ramnatthan Alagappan, Andrea Arpaci- Dusseau, and Remzi Arpaci-Dusseau. 2020. Strong and Efficient Consistency with Consistency-Aware Durability. In18th USENIX Conference on File and Storage Technologies (FAST 20). USENIX Asso- ciation, Santa Clara, CA, 323–...

  39. [49]

    Jinkun Geng, Anirudh Sivaraman, Balaji Prabhakar, and Mendel Rosenblum. 2022. Nezha: Deployable and High-Performance Con- sensus Using Synchronized Clocks.Proc. VLDB Endow.16, 4 (dec 2022), 629–642. doi:10.14778/3574245.3574250

  40. [51]

    Gray and D

    C. Gray and D. Cheriton. 1989. Leases: an efficient fault-tolerant mechanism for distributed file cache consistency. InProceedings of the Twelfth ACM Symposium on Operating Systems Principles (SOSP ’89). Association for Computing Machinery, New York, NY, USA, 202–210. doi:10.1...

  41. [52]

    Joshua Guarnieri and Aleksey Charapko. 2023. Linearizable Low- latency Reads at the Edge. InProceedings of the 10th Workshop on Principles and Practice of Consistency for Distributed Data(Rome, Italy)(PaPoC ’23). Association for Computing Machinery, New York, NY, USA, 77–83. d...

  42. [53]

    Rachid Guerraoui, Antoine Murat, Javier Picorel, Athanasios Xygkis, Huabing Yan, and Pengfei Zuo. 2022. uKharon: A Membership Service for Microsecond Applications. In2022 USENIX Annual Technical Conference (ATC 22). USENIX Association, Carlsbad, CA, 101–120. 14 https://www.use...

  43. [54]

    Gunawi, Mingzhe Hao, Tanakorn Leesatapornwongsa, Tiratat Patana-anake, Thanh Do, Jeffry Adityatama, Kurnia J

    Haryadi S. Gunawi, Mingzhe Hao, Tanakorn Leesatapornwongsa, Tiratat Patana-anake, Thanh Do, Jeffry Adityatama, Kurnia J. Eliazar, Agung Laksono, Jeffrey F. Lukman, Vincentius Martin, and Anang D. Satria. 2014. What Bugs Live in the Cloud? A Study of 3000+ Issues in Cloud Syste...

  44. [55]

    Rachael Harding, Dana Van Aken, Andrew Pavlo, and Michael Stone- braker. 2017. An evaluation of distributed concurrency control.Proc. VLDB Endow.10, 5 (Jan 2017), 553–564. doi:10.14778/3055540.3055548

  45. [56]

    HashiCorp. 2020. Consul.https://consul.io. Accessed: 2024-10-16

  46. [59]

    Maurice Herlihy. 1987. Dynamic quorum adjustment for partitioned data.ACM Trans. Database Syst.12, 2 (Jun 1987), 170–194. doi:10.1 145/22952.22953

  47. [60]

    Herlihy and Jeannette M

    Maurice P. Herlihy and Jeannette M. Wing. 1990. Linearizability: A Correctness Condition for Concurrent Objects.ACM Trans. Program. Lang. Syst.12, 3 (jul 1990), 463–492. doi:10.1145/78969.78972

  48. [61]

    Long Hoang Le, Enrique Fynn, Mojtaba Eslahi-Kelorazi, Robert Soulé, and Fernando Pedone. 2019. DynaStar: Optimized Dynamic Parti- tioning for Scalable State Machine Replication. In2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS). 1453–1465. doi:...

  49. [62]

    Heidi Howard, Dahlia Malkhi, and Alexander Spiegelman. 2016. Flex- ible Paxos: Quorum intersection revisited. arXiv:1608.06696 [cs.DC]

  50. [63]

    Guanzhou Hu, Andrea Arpaci-Dusseau, and Remzi Arpaci-Dusseau

  51. [65]

    Dongxu Huang, Qi Liu, Qiu Cui, Zhuhe Fang, Xiaoyu Ma, Fei Xu, Li Shen, Liu Tang, Yuxing Zhou, Menglong Huang, Wan Wei, Cong Liu, Jian Zhang, Jianjun Li, Xuelian Wu, Lingyu Song, Ruoxi Sun, Shuaipeng Yu, Lei Zhao, Nicholas Cameron, Liquan Pei, and Xin Tang. 2020. TiDB: A Raft-B...

  52. [66]

    Lorch, Yingnong Dang, Murali Chintalapati, and Randolph Yao

    Peng Huang, Chuanxiong Guo, Lidong Zhou, Jacob R. Lorch, Yingnong Dang, Murali Chintalapati, and Randolph Yao. 2017. Gray Failure: The Achilles’ Heel of Cloud-Scale Systems. InProceedings of the 16th Workshop on Hot Topics in Operating Systems(Whistler, BC, Canada)(HotOS ’17)....

  53. [67]

    Junqueira, and Benjamin Reed

    Patrick Hunt, Mahadev Konar, Flavio P. Junqueira, and Benjamin Reed. 2010. ZooKeeper: Wait-Free Coordination for Internet-Scale Systems. InProceedings of the 2010 USENIX Conference on USENIX Annual Technical Conference(Boston, MA)(USENIXATC’10). USENIX Association, USA, 11

  54. [68]

    Randall Hunt. 2017. Keeping Time With Amazon Time Sync Service. https://aws.amazon.com/blogs/aws/keeping-time-with-amazon- time-sync-service/. Accessed: 2017-11-29

  55. [69]

    Sagar Jha, Jonathan Behrens, Theo Gkountouvas, Mae Milano, Wei- jia Song, Edward Tremel, Robbert Van Renesse, Sydney Zink, and Kenneth P. Birman. 2019. Derecho: Fast State Machine Replication for Cloud Services.ACM Trans. Comput. Syst.36, 2, Article 4 (apr 2019), 49 pages. doi...

  56. [70]

    Jonathan Kaldor, Jonathan Mace, Michał Bejda, Edison Gao, Wik- tor Kuropatwa, Joe O’Neill, Kian Win Ong, Bill Schaller, Pingjia Shan, Brendan Viscomi, Vinod Venkataraman, Kaushik Veeraragha- van, and Yee Jiun Song. 2017. Canopy: An End-to-End Performance Tracing And Analysis S...

  57. [71]

    Siavash Katebzadeh, Arpit Joshi, Aleksandar Dragojevic, Boris Grot, and Vijay Nagara- jan

    Antonios Katsarakis, Vasilis Gavrielatos, M.R. Siavash Katebzadeh, Arpit Joshi, Aleksandar Dragojevic, Boris Grot, and Vijay Nagara- jan. 2020. Hermes: A Fast, Fault-Tolerant and Linearizable Replica- tion Protocol. InProceedings of the Twenty-Fifth International Con- ference ...

  58. [72]

    KRaft. 2025. KRaft: Apache Kafka Without ZooKeeper.https: //developer.confluent.io/learn/kraft/, Last accessed on 2025-04-12

  59. [73]

    H. T. Kung and John T. Robinson. 1981. On optimistic methods for concurrency control.ACM Trans. Database Syst.6, 2 (jun 1981), 213–226. doi:10.1145/319566.319567

  60. [74]

    Accessed: 2024-12-01

  61. [75]

    Leslie Lamport. 1998. The Part-Time Parliament.ACM Trans. Comput. Syst.16, 2 (may 1998), 133–169. doi:10.1145/279227.279229

  62. [76]

    Leslie Lamport. 2001. Paxos Made Simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001)(December 2001), 51–58.https://www.microsoft.com/en- us/research/publication/paxos-made-simple/

  63. [77]

    Leslie Lamport. 2005. Generalized consensus and Paxos.Microsoft Research Technical Report(2005)

  64. [78]

    Leslie Lamport. 2006. Fast Paxos.Distrib. Comput.19, 2 (oct 2006), 79–103. doi:10.1007/s00446-006-0005-x

  65. [79]

    Leslie Lamport. 1979. How to Make a Multiprocessor Computer That Correctly Executes Multiprocess Programs.IEEE Transactions on Computers C-289 (September 1979), 690–691.https://www.microsof t.com/en-us/research/publication/make-multiprocessor-computer- correctly-executes-multi...

  66. [80]

    Butler Lampson. 2001. The ABCD’s of Paxos. InProceedings of the Twentieth Annual ACM Symposium on Principles of Distributed Com- puting(Newport, RI, USA)(PODC ’01). Association for Computing Machinery, New York, NY, USA, 13. doi:10.1145/383962.383969

  67. [82]

    Collin Lee, Seo Jin Park, Ankita Kejriwal, Satoshi Matsushita, and John Ousterhout. 2015. Implementing Linearizability at Large Scale 15 and Low Latency. InProceedings of the 25th Symposium on Operat- ing Systems Principles(CA)(SOSP ’15). Association for Computing Machinery, N...

  68. [83]

    Edward K. F. Lee and Chandramohan A. Thekkath. 1996. Petal: distributed virtual disks. InASPLOS VII.https://api.semanticscholar. org/CorpusID:314852

  69. [84]

    Leslie Lamport, Dahlia Malkhi, and Lidong Zhou. 2009. Vertical paxos and primary-backup replication. InProceedings of the 28th ACM Symposium on Principles of Distributed Computing(Calgary, AB, Canada)(PODC ’09). Association for Computing Machinery, New York, NY, USA, 312–313. ...

  70. [85]

    Sharma, Adriana Szekeres, and Dan R

    Jialin Li, Ellis Michael, Naveen Kr. Sharma, Adriana Szekeres, and Dan R. K. Ports. 2016. Just Say NO to Paxos Overhead: Replacing Consensus with Network Ordering. In12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). USENIX Association, Savannah, G...

  71. [86]

    Yuliang Li, Gautam Kumar, Hema Hariharan, Hassan Wassel, Peter Hochschild, Dave Platt, Simon Sabato, Minlan Yu, Nandita Dukkipati, Prashant Chandra, and Amin Vahdat. 2020. Sundial: Fault-tolerant Clock Synchronization for Datacenters. In14th USENIX Symposium on Operating Syste...

  72. [87]

    Wei Lin, Mao Yang, Lintao Zhang, and Lidong Zhou. 2008. PacificA: Replication in Log-Based Distributed Storage Systems.https://api. semanticscholar.org/CorpusID:18304090

  73. [88]

    Linux man pages. 2011. tc-netem(8) — Linux manual page.https: //man7.org/linux/man-pages/man8/tc-netem.8.html. [Online; accessed 29-November-2023]

  74. [89]

    Ki Suh Lee, Han Wang, Vishal Shrivastav, and Hakim Weatherspoon

  75. [90]

    Faleiro, Juno Kim, Soham Sankaran, Daniel J

    Joshua Lockerman, Jose M. Faleiro, Juno Kim, Soham Sankaran, Daniel J. Abadi, James Aspnes, Siddhartha Sen, and Mahesh Bal- akrishnan. 2018. The FuzzyLog: A Partially Ordered Shared Log. In 13th USENIX Symposium on Operating Systems Design and Implemen- tation (OSDI 18). USENI...

  76. [92]

    Xuhao Luo, Weihai Shen, Shuai Mu, and Tianyin Xu. 2022. DepFast: Orchestrating Code of Quorum Systems. In2022 USENIX Annual Technical Conference (ATC 22). USENIX Association, 557–574.https: //www.usenix.org/conference/atc22/presentation/luo

  77. [93]

    Haojun Ma, Hammad Ahmad, Aman Goel, Eli Goldweber, Jean- Baptiste Jeannin, Manos Kapritsos, and Baris Kasikci. 2022. Sift: Using Refinement-guided Automation to Verify Complex Distributed Systems. In2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, C...

  78. [94]

    Sakallah

    Haojun Ma, Aman Goel, Jean-Baptiste Jeannin, Manos Kapritsos, Baris Kasikci, and Karem A. Sakallah. 2019. I4: Incremental In- ference of Inductive Invariants for Verification of Distributed Pro- tocols. InProceedings of the 27th ACM Symposium on Operating Systems Principles(Hu...

  79. [95]

    Freedman, Michael Kaminsky, and David G

    Wyatt Lloyd, Michael J. Freedman, Michael Kaminsky, and David G. Andersen. 2011. Don’t settle for eventual: scalable causal consistency for wide-area storage with COPS. InProceedings of the Twenty-Third ACM Symposium on Operating Systems Principles(Cascais, Portugal) (SOSP ’11...

  80. [96]

    Junqueira, and Keith Marzullo

    Yanhua Mao, Flavio P. Junqueira, and Keith Marzullo. 2008. Mencius: Building Efficient Replicated State Machines for WANs. InProceed- ings of the 8th USENIX Conference on Operating Systems Design and Implementation(San Diego, California)(OSDI’08). USENIX Associa- tion, USA, 369–384

  81. [97]

    Jim Martin, Jack Burbank, William Kasch, and Professor David L. Mills. 2010. Network Time Protocol Version 4: Protocol and Algo- rithms Specification. RFC 5905. doi:10.17487/RFC5905

  82. [98]

    Syed Akbar Mehdi, Cody Littley, Natacha Crooks, Lorenzo Alvisi, Nathan Bronson, and Wyatt Lloyd. 2017. I Can’t Believe It’s Not Causal! Scalable Causal Consistency with No Slowdown Cascades. In14th USENIX Symposium on Networked Systems Design and Im- plementation (NSDI 17). US...

  83. [99]

    Microsoft. 2024. Azure Global Infrastructure.https://datacenters.mi crosoft.com/globe/explore/, Last accessed on 2024-11-30

  84. [100]

    Andersen, and Michael Kaminsky

    Iulian Moraru, David G. Andersen, and Michael Kaminsky. 2013. There is More Consensus in Egalitarian Parliaments. InProceedings of the 24th ACM Symposium on Operating Systems Principles(Farminton, Pennsylvania)(SOSP ’13). Association for Computing Machinery, New York, NY, USA,...

  85. [101]

    Kai Ma, Cheng Li, Enzuo Zhu, Ruichuan Chen, Feng Yan, and Kang Chen. 2024. Noctua: Towards Automated and Practical Fine-grained Consistency Analysis. InProceedings of the Nineteenth European Con- ference on Computer Systems(Athens, Greece)(EuroSys ’24). Asso- ciation for Compu...

  86. [102]

    Andersen, and Michael Kaminsky

    Iulian Moraru, David G. Andersen, and Michael Kaminsky. 2014. Paxos Quorum Leases: Fast Reads Without Sacrificing Writes. Carnegie Mellon University PDL Technical Report(2014)

  87. [103]

    Achour Mostefaoui, Matthieu Perrin, and Julien Weibel. 2024. Brief Announcement: Randomized Consensus: Common Coins Are not the Holy Grail!. InProceedings of the 43rd ACM Symposium on Principles of Distributed Computing(Nantes, France)(PODC ’24). Association for Computing Mach...

  88. [104]

    Antoine Murat, Clément Burgelin, Athanasios Xygkis, Igor Zablotchi, Marcos Kawazoe Aguilera, and Rachid Guerraoui. 2024. SWARM: Replicating Shared Disaggregated-Memory Data in No Time. InPro- ceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP ’24)....

  89. [105]

    Faisal Nawab, Divyakant Agrawal, and Amr El Abbadi. 2018. DPaxos: Managing Data Closer to Users for Low-Latency and Mobile Ap- plications. InProceedings of the 2018 International Conference on Management of Data(Houston, TX, USA)(SIGMOD ’18). Associ- ation for Computing Machin...

  90. [106]

    Khiem Ngo, Siddhartha Sen, and Wyatt Lloyd. 2020. Tolerating Slowdowns in Replicated State Machines using Copilots. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 583–598.https://www.usenix.org/c onference/osdi20/presentation/ngo

  91. [107]

    Andersen, and Michael Kaminsky

    Iulian Moraru, David G. Andersen, and Michael Kaminsky. 2014. Paxos Quorum Leases: Fast Reads Without Sacrificing Writes. In Proceedings of the ACM Symposium on Cloud Computing(Seattle, WA, USA)(SOCC ’14). Association for Computing Machinery, New York, NY, USA, 1–13. doi:10.11...

  92. [108]

    2014.Consensus: Bridging Theory and Practice

    Diego Ongaro. 2014.Consensus: Bridging Theory and Practice. Ph. D. Dissertation. Stanford, CA, USA. Advisor(s) K., Ousterhout, John and David, Mazières, and Mendel, Rosenblum,. AAI28121474

  93. [109]

    Diego Ongaro and John Ousterhout. 2014. In Search of an Under- standable Consensus Algorithm. InProceedings of the 2014 USENIX Conference on USENIX Annual Technical Conference(Philadelphia, PA)(USENIX ATC’14). USENIX Association, USA, 305–320

  94. [110]

    Team Live Optics. 2021. Live Optics Basics: Read / Write Ratio. https://support.liveoptics.com/hc/en-us/articles/229590547-Live- Optics-Basics-Read-Write-Ratio. Accessed: 2024-12-01

  95. [111]

    Haochen Pan, Jesse Tuglu, Neo Zhou, Tianshu Wang, Yicheng Shen, Xiong Zheng, Joseph Tassarotti, Lewis Tseng, and Roberto Palmieri

  96. [112]

    Seo Jin Park and John Ousterhout. 2019. Exploiting Commutativ- ity For Practical Fast Replication. In16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). USENIX Association, Boston, MA, 47–64.https://www.usenix.org/conferenc e/nsdi19/presentation/park

  97. [113]

    Oki and Barbara H

    Brian M. Oki and Barbara H. Liskov. 1988. Viewstamped Replica- tion: A New Primary Copy Method to Support Highly-Available Distributed Systems. InProceedings of the Seventh Annual ACM Sym- posium on Principles of Distributed Computing(Toronto, Ontario, Canada)(PODC ’88). Assoc...

  98. [114]

    Spreitzer, Douglas B

    Karin Petersen, Mike J. Spreitzer, Douglas B. Terry, Marvin M. Theimer, and Alan J. Demers. 1997. Flexible update propagation for weakly consistent replication. InProceedings of the Sixteenth ACM Symposium on Operating Systems Principles(Saint Malo, France) (SOSP ’97). Associa...

  99. [115]

    Dan R. K. Ports, Jialin Li, Vincent Liu, Naveen Kr. Sharma, and Arvind Krishnamurthy. 2015. Designing Distributed Systems Using Approx- imate Synchrony in Data Center Networks. In12th USENIX Sym- posium on Networked Systems Design and Implementation (NSDI 15). USENIX Associati...

  100. [117]

    RabbitMQ. 2025. RabbitMQ: One broker to queue them all.https: //www.rabbitmq.com/, Last accessed on 2025-04-08

  101. [118]

    Redpanda. 2024. Redpanda: The Unified Streaming Data Platform. https://www.redpanda.com/, Last accessed on 2024-09-05

  102. [119]

    Schneider

    Robbert Van Renesse and Fred B. Schneider. 2004. Chain Replication for Supporting High Throughput and Availability. In6th Symposium on Operating Systems Design & Implementation (OSDI 04). USENIX Association, CA.https://www.usenix.org/conference/osdi-04/chain- replication-suppo...

  103. [120]

    Suraj Pasuparthy and Lokesh Agarwal. 2023. Benchmarking Span- ner’s price-performance for key-value workloads.https://cloud.goog le.com/blog/products/databases/benchmarking-spanner-for-key- value-workloads/. Accessed: 2024-12-01

  104. [121]

    Sajjad Rizvi, Bernard Wong, and Srinivasan Keshav. 2017. Canopus: A Scalable and Massively Parallel Consensus Protocol. InProceedings of the 13th International Conference on Emerging Networking EXperi- ments and Technologies(Incheon, Republic of Korea)(CoNEXT ’17). Association...

  105. [122]

    Fedor Ryabinin, Alexey Gotsman, and Pierre Sutra. 2024. SwiftPaxos: Fast Geo-Replicated State Machines. In21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). USENIX Association, Santa Clara, CA, 345–369.https://www.usenix.org/con ference/nsdi24/pres...

  106. [123]

    Schneider

    Fred B. Schneider. 1990. Implementing Fault-Tolerant Services Using the State Machine Approach: A Tutorial.ACM Comput. Surv.22, 4 (dec 1990), 299–319. doi:10.1145/98163.98167

  107. [124]

    ScyllaDB. 2023. Beyond Legacy NoSQL: 7 Design Principles Behind ScyllaDB.https://lp.scylladb.com/real-time-big-data-database- principles-thanks.html, Last accessed on 2023-11-13

  108. [125]

    Chrysoula Stathakopoulou, Matej Pavlovic, and Marko Vukolić. 2022. State Machine Replication Scalability Made Simple. InProceedings of the Seventeenth European Conference on Computer Systems(Rennes, France)(EuroSys ’22). Association for Computing Machinery, New York, NY, USA, ...

  109. [126]

    Xudong Sun, Wenjie Ma, Jiawei Tyler Gu, Zicheng Ma, Tej Chajed, Jon Howell, Andrea Lattuada, Oded Padon, Lalith Suresh, Adriana Szekeres, and Tianyin Xu. 2024. Anvil: Verifying Liveness of Cluster Management Controllers. In18th USENIX Symposium on Operating Systems Design and ...

  110. [127]

    David K. Rensin. 2015.Kubernetes - Scheduling the Future at Cloud Scale. O’Reilly and Associates, 1005 Gravenstein Highway North Sebastopol, CA 95472. All pages.http://www.oreilly.com/webops- perf/free/kubernetes.csp

  111. [128]

    Hatem Takruri, Ibrahim Kettaneh, Ahmed Alquraan, and Samer Al- Kiswany. 2020. FLAIR: Accelerating Reads with Consistency-Aware Network Routing. In17th USENIX Symposium on Networked Systems Design & Implementation (NSDI 20). USENIX Association, CA, 723– 737.https://usenix.org/c...

  112. [129]

    Freedman

    Jeff Terrace and Michael J. Freedman. 2009. Object Storage on CRAQ: High-Throughput Chain Replication for Read-Mostly Workloads. In 2009 USENIX Annual Technical Conference (USENIX ATC 09). USENIX Association, San Diego, CA.https://www.usenix.org/conference/us enix-09/object-st...

  113. [130]

    Terry, Alan J

    Douglas B. Terry, Alan J. Demers, Karin Petersen, Mike J. Spreitzer, Marvin M. Theimer, and Brent B. Welch. 1994. Session Guarantees for Weakly Consistent Replicated Data. InProceedings of the Third International Conference on on Parallel and Distributed Information Systems(Au...

  114. [131]

    Myles Thiessen, Aleksey Panas, Guy Khazma, and Eyal de Lara. 2024. Towards Reconfigurable Linearizable Reads. arXiv:2404.05470 [cs.DC]https://arxiv.org/abs/2404.05470

  115. [132]

    TigerBeetle. 2024. TigerBeetle: The Financial Transactions Database. https://tigerbeetle.com/, Last accessed on 2024-11-12

  116. [133]

    Sarah Tollman, Seo Jin Park, and John Ousterhout. 2021. EPaxos Revisited. In18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21). USENIX Association, 613–632.https: //www.usenix.org/conference/nsdi21/presentation/tollman

  117. [134]

    Florian Suri-Payer, Matthew Burke, Zheng Wang, Yunhao Zhang, Lorenzo Alvisi, and Natacha Crooks. 2021. Basil: Breaking up BFT with ACID (Transactions). InProceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles(Virtual Event, Germany) (SOSP ’21). Associatio...

  118. [135]

    Madhyastha

    Muhammed Uluyol, Anthony Huang, Ayush Goel, Mosharaf Chowd- hury, and Harsha V. Madhyastha. 2020. Near-Optimal Latency Versus 17 Cost Tradeoffs in Geo-Distributed Storage. In17th USENIX Sympo- sium on Networked Systems Design and Implementation (NSDI 20). USENIX Association, S...

  119. [136]

    Nathan VanBenschoten, Arul Ajmani, Marcus Gartner, Andrei Matei, Aayush Shah, Irfan Sharif, Alexander Shraer, Adam Storm, Rebecca Taft, Oliver Tan, Andy Woods, and Peyton Walters. 2022. En- abling the Next Generation of Multi-Region Applications with Cock- roachDB. InProceedin...

  120. [137]

    Kaushik Veeraraghavan, Justin Meza, Scott Michelson, Sankar- alingam Panneerselvam, Alex Gyori, David Chou, Sonia Margulis, Daniel Obenshain, Shruti Padmanabha, Ashish Shah, Yee Jiun Song, and Tianyin Xu. 2018. Maelstrom: Mitigating Datacenter-level Dis- asters by Draining Int...

  121. [138]

    Werner Vogels. 2008. Eventually Consistent: Building Reliable Dis- tributed Systems at a Worldwide Scale Demands Trade-Offs Be- tween Consistency and Availability.Queue6, 6 (oct 2008), 14–19. doi:10.1145/1466443.1466448

  122. [139]

    Cheng Wang, Jianyu Jiang, Xusheng Chen, Ning Yi, and Heming Cui

  123. [141]

    Bohdan Trach, Rasha Faqeh, Oleksii Oleksenko, Wojciech Ozga, Pramod Bhatotia, and Christof Fetzer. 2020. T-Lease: a trusted lease primitive for distributed systems. InProceedings of the 11th ACM Symposium on Cloud Computing(Virtual Event, USA)(SoCC ’20). Association for Comput...

  124. [142]

    Hellerstein, Heidi Howard, Ion Stoica, and Adriana Szekeres

    Michael Whittaker, Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Neil Giridharan, Joseph M. Hellerstein, Heidi Howard, Ion Stoica, and Adriana Szekeres. 2021. Scaling Replicated State Machines with Compartmentalization.Proc. VLDB Endow.14, 11 (jul 2021), 2203–2215. doi...

  125. [143]

    Hellerstein, Heidi Howard, and Ion Stoica

    Michael Whittaker, Aleksey Charapko, Joseph M. Hellerstein, Heidi Howard, and Ion Stoica. 2021. Read-Write Quorum Systems Made Practical. InProceedings of the 8th Workshop on Principles and Practice of Consistency for Distributed Data(Online, United Kingdom)(PaPoC ’21). Associ...

  126. [144]

    Jian Yi, Qing Li, Bin Zhang, Yong Jiang, Dan Zhao, Yuan Yang, and Zhenhui Yuan. 2023. Gleaning the Consensus for Linearizable and Conflict-Free Per-Replica Local Reads. InProceedings of the 7th Asia- Pacific Workshop on Networking(Hong Kong, China)(APNet ’23). Association for ...

  127. [145]

    Reiter, Guy Golan Gueta, and Ittai Abraham

    Maofan Yin, Dahlia Malkhi, Michael K. Reiter, Guy Golan Gueta, and Ittai Abraham. 2019. HotStuff: BFT Consensus with Linearity and Responsiveness. InProceedings of the 2019 ACM Symposium on Principles of Distributed Computing(Toronto ON, Canada)(PODC ’19). Association for Comp...

  128. [146]

    Hanze Zhang, Ke Cheng, Rong Chen, and Haibo Chen. 2024. Fast and Scalable In-network Lock Management Using Lock Fission. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). USENIX Association, Santa Clara, CA, 251–268.https: //www.usenix.org/confe...

  129. [147]

    Yunhao Zhang, Srinath Setty, Qi Chen, Lidong Zhou, and Lorenzo Alvisi. 2020. Byzantine Ordered Consensus without Byzantine Oli- garchy. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 633–649.https: //www.usenix.org/confere...

  130. [148]

    Hanyu Zhao, Quanlu Zhang, Zhi Yang, Ming Wu, and Yafei Dai

  131. [149]

    Xingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen, and Haibo Chen. 2023. Characterizing Off-path SmartNIC for Accelerating Dis- tributed Systems. In17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). USENIX Association, Boston, MA, 987–1004.https://w...

  132. [150]

    Yang Zhou, Zezhou Wang, Sowmya Dharanipragada, and Minlan Yu

  133. [158]

    Siyuan Zhou and Shuai Mu. 2021. Fault-Tolerant Replication with Pull-Based Consensus in MongoDB. In18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21). USENIX Association, 687–703.https://www.usenix.org/conference/nsdi21/p resentation/zhou

  134. [2014]

    In2014 44th Annual IEEE/I- FIP International Conference on Dependable Systems and Networks

    Scalable State-Machine Replication. In2014 44th Annual IEEE/I- FIP International Conference on Dependable Systems and Networks. 331–342. doi:10.1109/DSN.2014.41

  135. [2016]

    In Proceedings of the 2016 ACM SIGCOMM Conference(Florianopolis, Brazil)(SIGCOMM ’16)

    Globally Synchronized Time via Datacenter Networks. In Proceedings of the 2016 ACM SIGCOMM Conference(Florianopolis, Brazil)(SIGCOMM ’16). Association for Computing Machinery, New York, NY, USA, 454–467. doi:10.1145/2934872.2934885

  136. [2017]

    InProceedings of the 2017 Symposium on Cloud Computing(Santa Clara, California) (SoCC ’17)

    APUS: fast and scalable paxos on RDMA. InProceedings of the 2017 Symposium on Cloud Computing(Santa Clara, California) (SoCC ’17). Association for Computing Machinery, New York, NY, USA, 94–107. doi:10.1145/3127479.3128609

  137. [2018]

    InProceedings of the ACM Symposium on Cloud Com- puting(Carlsbad, CA, USA)(SoCC ’18)

    SDPaxos: Building Efficient Semi-Decentralized Geo-replicated State Machines. InProceedings of the ACM Symposium on Cloud Com- puting(Carlsbad, CA, USA)(SoCC ’18). Association for Computing Machinery, New York, NY, USA, 68–81. doi:10.1145/3267809.3267837

  138. [2020]

    In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)

    Virtual Consensus in Delos. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 617–632.https://www.usenix.org/conference/osdi20/p resentation/balakrishnan

  139. [2023]

    none” /∈ Replicas Population ∆ = Cardinality(Replicas) MajorityNum ∆ = ( Population÷ 2) +1 WritesAssumption ∆ = ∧ IsFiniteSet(Writes) ∧ Cardinality(Writes)≥ 1 ∧ “nil

    Electrode: Accelerating Distributed Protocols with eBPF. In20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Boston, MA, 1391–1407.https: //www.usenix.org/conference/nsdi23/presentation/zhou A Complete Presentation of theBodega...

  140. [2024]

    arXiv:2409.01576 [cs.DC]https://arxiv.org/abs/2409.01576

    A Unified, Practical, and Understandable Summary of Non-transactional Consistency Levels in Distributed Replication. arXiv:2409.01576 [cs.DC]https://arxiv.org/abs/2409.01576

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.