Pith. sign in

REVIEW 1 major objections 61 references

Existing benchmarks for geo-distributed OLTP databases miss network instability, data locality, and cross-region transfer costs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-28 23:56 UTC pith:TYVYBCRL

load-bearing objection Gaia framework adds network variability, locality, and cost tracking to geo-distributed DB evaluation and reports three practical findings, but the abstract leaves the experimental details thin. the 1 major comments →

arxiv 2605.30156 v1 pith:TYVYBCRL submitted 2026-05-28 cs.DB

The Missing Dimensions in Geo-Distributed Database Evaluation

classification cs.DB
keywords geo-distributed databasesOLTP evaluationnetwork costsfault tolerancecloud benchmarksdata localityperformance evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper establishes that standard evaluation methods for geo-distributed databases assume stable networks, omit explicit control over where data and clients sit, and ignore the expense of moving data between cloud regions. It introduces Gaia as a framework that adds variable network conditions, multiple geo-distribution patterns, locality settings, and cost tracking to the test process. When the framework is applied to current systems, it shows that performance degrades under network fluctuations, that data transfer dominates total cloud expenses, and that multi-region fault-tolerance features add measurable delays to the critical path. A reader would care because these omissions mean published results may not predict behavior in actual global deployments and may hide important design trade-offs among speed, reliability, and cost.

Core claim

By deploying geo-distributed OLTP systems across cloud regions under controlled variations in network stability, data placement, client locations, and geo-distribution patterns while measuring transfer costs, the evaluation finds that most systems are sensitive to network instabilities, that network costs dominate cloud deployment expenses, and that multi-region fault-tolerance mechanisms incur measurable critical-path overhead that prior evaluations overlooked. The authors therefore conclude that future designs must rethink the trade-offs among performance, fault tolerance, and cost.

What carries the argument

Gaia, the evaluation framework that extends benchmarks with explicit settings for variable cross-region network conditions, data and client locality, multiple geo-distribution patterns, and measurement of data transfer costs.

Load-bearing premise

The assumption that stable networks, ignored locality settings, and unmeasured data transfer costs are sufficient to represent real geo-distributed database behavior.

What would settle it

Running the same systems under Gaia's variable network conditions and observing neither performance sensitivity to instabilities nor network costs as the dominant expense.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Database systems must incorporate handling for network instabilities into their core design.
  • Cost models for cloud deployments must treat inter-region data transfer as the primary expense driver.
  • Multi-region fault-tolerance mechanisms require explicit optimization to reduce their overhead on the critical path.
  • Benchmark practices must expand to include diverse geo-distribution patterns and unstable network models.
  • Design of future geo-distributed databases must explicitly balance performance against fault-tolerance overhead and network costs.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Published performance numbers obtained under stable-network benchmarks may not generalize to production geo-deployments.
  • Standard benchmark suites would need updates that treat network variability and transfer costs as first-class parameters.
  • The observed overheads might be reduced by redesigning systems to minimize cross-region traffic during normal operation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 0 minor

Summary. The paper proposes Gaia, a comprehensive evaluation framework for geo-distributed OLTP databases. It addresses gaps in prior benchmarks (assumed stable networks, missing explicit data/client locality settings, ignored cross-region data transfer costs, and limited geo-distribution patterns) by deploying systems across multiple cloud regions under variable network conditions and patterns. The evaluation yields three main findings: most systems are sensitive to network instabilities, network costs dominate cloud deployment expenses, and multi-region fault-tolerance mechanisms incur measurable critical-path overhead often overlooked previously. The authors conclude that future designs must rethink trade-offs among performance, fault-tolerance, and cost.

Significance. If the experimental results hold under rigorous verification, the work would be significant for the database community by exposing practical gaps in how geo-distributed systems are benchmarked and by incorporating cost as a first-class dimension. The proposal of a reusable framework that varies network conditions and locality is a constructive contribution; credit is due for explicitly including data-transfer costs, which prior evaluations have largely omitted.

major comments (1)
  1. [Abstract] Abstract: the three findings (sensitivity to instabilities, cost dominance, and overlooked fault-tolerance overhead) are stated without any accompanying description of the experimental methodology, the specific systems evaluated, the network models or instability injection methods, the cost-calculation procedure, the number of runs, or error bars. Because these details are load-bearing for the central empirical claims, the findings cannot be assessed for support from the provided text.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for their review. The single major comment concerns the level of detail in the abstract. We respond point-by-point below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the three findings (sensitivity to instabilities, cost dominance, and overlooked fault-tolerance overhead) are stated without any accompanying description of the experimental methodology, the specific systems evaluated, the network models or instability injection methods, the cost-calculation procedure, the number of runs, or error bars. Because these details are load-bearing for the central empirical claims, the findings cannot be assessed for support from the provided text.

    Authors: Abstracts are intentionally concise summaries (typically <250 words) and do not enumerate methodology; this is standard practice across the database literature. The manuscript body provides the requested details: Section 3 describes the Gaia framework, the three evaluated systems, network modeling via latency/bandwidth variation and instability injection, and the cost model based on public cloud pricing tables; Section 4 reports all results with 5+ runs per configuration and error bars. The three findings are therefore directly supported by the full experimental evidence. We do not believe expanding the abstract with these specifics would improve readability or conform to venue norms. revision: no

Circularity Check

0 steps flagged

Empirical framework proposal with no derivation chain

full rationale

The paper proposes the Gaia evaluation framework and reports empirical observations from deploying geo-distributed OLTP systems under varied network conditions, locality patterns, and cost models. No equations, fitted parameters, predictions, or mathematical derivations appear in the abstract or described content. Claims rest on experimental measurements rather than any self-referential reduction or self-citation chain. This is a standard empirical benchmark paper whose central contribution is the framework itself and the reported measurements, with no load-bearing step that reduces to its own inputs by construction.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

The contribution is an empirical evaluation framework and set of observations rather than a theoretical derivation; no free parameters, axioms, or invented entities are referenced in the abstract.

pith-pipeline@v0.9.1-grok · 5715 in / 1031 out tokens · 31456 ms · 2026-06-28T23:56:31.858336+00:00 · methodology

0 comments
read the original abstract

Geo-distributed OLTP databases are widely deployed across cloud regions, yet current evaluation practices do not cover the challenges of this aspect. Existing benchmarks assume stable network conditions; they lack explicit settings for data and client locality, and they largely ignore data transfer costs across regions. In addition, most evaluations rely on a limited set of geo-distribution patterns. In this paper, we propose Gaia, a comprehensive evaluation framework that addresses these gaps. We use Gaia to perform a comprehensive evaluation of existing geo-distributed OLTP systems. We deploy them across multiple cloud regions, using different geo-distribution patterns and variable cross-region network conditions. Among other interesting findings, our framework reveals that: i) most systems are sensitive to network instabilities, ii) network costs dominate cloud deployment expenses iii) multi-region fault-tolerance mechanisms incur measurable critical-path overhead that is often overlooked in prior evaluations. We argue that for the design of future geo-distributed databases, we must rethink the trade-offs between performance, fault-tolerance, and cost.

Figures

Figures reproduced from arXiv: 2605.30156 by Asterios Katsifodimos, George Christodoulou, Kyriakos Psarakis, Oto Mraz, Paris Carbone.

Figure 1
Figure 1. Figure 1: Gaia framework design. Gaia generates transactions based on user [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Three transaction types distinguished by Gaia (replicas omitted). [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Average round-trip times (ms) for each region pair used across 3 runs. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The YCSB and TPC-C throughput, latency, data transfers, and cost per transaction with increasing input throughput. Opaque lines in the latency [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The YCSB and TPC-C throughput, p50 (bold bars) and p99 (opaque bars) latency, data transfers, and VM+storage cost (bold bars) and data transfer [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The YCSB and TPC-C throughput, p50 (bold bars) and p99 (opaque bars) latency, data transfers, and VM+storage cost (bold bars) and data transfer [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: The YCSB and TPC-C throughput, latency, data transfers, and cost per transaction with varying geo-distribution percentages. Opaque lines in the [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The relative YCSB and TPC-C throughput, latency, data transfers, and cost per transaction under extra network latency and jitter. Opaque lines in the [PITH_FULL_IMAGE:figures/full_fig_p010_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: The relative YCSB and TPC-C throughput, latency, data transfers, and cost per transaction under extra packet loss. Opaque lines in the latency figure [PITH_FULL_IMAGE:figures/full_fig_p011_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Trace of CRDB’s performance under a single node and full region [PITH_FULL_IMAGE:figures/full_fig_p011_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

61 extracted references · 1 canonical work pages

  1. [1]

    Detock: High perfor- mance multi-region transactions at scale,

    C. D. Nguyen, J. K. Miller, and D. J. Abadi, “Detock: High perfor- mance multi-region transactions at scale,”Proceedings of the ACM on Management of Data, vol. 1, no. 2, pp. 1–27, 2023

  2. [2]

    An evaluation of distributed concurrency control,

    R. Harding, D. Van Aken, A. Pavlo, and M. Stonebraker, “An evaluation of distributed concurrency control,”Proceedings of the VLDB Endow- ment, vol. 10, no. 5, pp. 553–564, 2017

  3. [3]

    Slog: Serializable, low-latency, geo- replicated transactions,

    K. Ren, D. Li, and D. J. Abadi, “Slog: Serializable, low-latency, geo- replicated transactions,”Proceedings of the VLDB Endowment, vol. 12, no. 11, 2019

  4. [4]

    Calvin: fast distributed transactions for partitioned database systems,

    A. Thomson, T. Diamond, S.-C. Weng, K. Ren, P. Shao, and D. J. Abadi, “Calvin: fast distributed transactions for partitioned database systems,” inProceedings of the 2012 ACM SIGMOD international conference on management of data, 2012, pp. 1–12

  5. [5]

    Caerus: Low-latency distributed transactions for geo-replicated systems,

    J. Hildred, M. Abebe, and K. Daudjee, “Caerus: Low-latency distributed transactions for geo-replicated systems,”Proceedings of the VLDB Endowment, vol. 17, no. 3, pp. 469–482, 2023

  6. [6]

    CockroachDB: The resilient geo-distributed SQL database,

    R. Taft, I. Sharif, A. Matei, N. VanBenschoten, J. Lewis, T. Grieger, K. Niemi, A. Woods, A. Birzin, R. Posset al., “CockroachDB: The resilient geo-distributed SQL database,” inProceedings of the 2020 ACM SIGMOD international conference on management of data, 2020, pp. 1493–1509

  7. [7]

    Faunadb,

    Fauna, “Faunadb,” https://fauna.com, 2020

  8. [8]

    Oltp through the looking glass 16 years later: Communication is the new bottleneck,

    X. Zhou, V . Leis, X. Yu, and M. Stonebraker, “Oltp through the looking glass 16 years later: Communication is the new bottleneck,” in15th Annual Conference on Innovative Data Systems Research (CIDR’25), 2025

  9. [9]

    Sunstorm: Geographically distributed transactions over aurora-style systems,

    C. D. Nguyen, P. Nilangekar, H. Linnakangas, and D. J. Abadi, “Sunstorm: Geographically distributed transactions over aurora-style systems,”Proceedings of the VLDB Endowment, vol. 18, no. 13, pp. 5555–5568, 2025

  10. [10]

    Concurrency control in distributed database systems,

    P. A. Bernstein and N. Goodman, “Concurrency control in distributed database systems,”ACM Computing Surveys (CSUR), vol. 13, no. 2, pp. 185–221, 1981

  11. [11]

    Aria: A fast and practical deterministic oltp database,

    Y . Lu, X. Yu, L. Cao, and S. Madden, “Aria: A fast and practical deterministic oltp database,”Proceedings of the VLDB Endowment, vol. 13, no. 11, 2020

  12. [12]

    Spanner: Google’s globally distributed database,

    J. C. Corbett, J. Dean, M. Epstein, A. Fikes, C. Frost, J. J. Furman, S. Ghemawat, A. Gubarev, C. Heiser, P. Hochschildet al., “Spanner: Google’s globally distributed database,”ACM Transactions on Computer Systems (TOCS), vol. 31, no. 3, pp. 1–22, 2013

  13. [13]

    Amazon aurora: Design considerations for high throughput cloud-native relational databases,

    A. Verbitski, A. Gupta, D. Saha, M. Brahmadesam, K. Gupta, R. Mittal, S. Krishnamurthy, S. Maurice, T. Kharatishvili, and X. Bao, “Amazon aurora: Design considerations for high throughput cloud-native relational databases,” inProceedings of the 2017 ACM International Conference on Management of Data, 2017, pp. 1041–1052

  14. [14]

    Building consistent transactions with inconsistent replication,

    I. Zhang, N. K. Sharma, A. Szekeres, A. Krishnamurthy, and D. R. Ports, “Building consistent transactions with inconsistent replication,” ACM Transactions on Computer Systems (TOCS), vol. 35, no. 4, pp. 1–37, 2018

  15. [15]

    Consolidating concurrency control and consensus for commits under conflicts,

    S. Mu, L. Nelson, W. Lloyd, and J. Li, “Consolidating concurrency control and consensus for commits under conflicts,” in12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), 2016, pp. 517–532

  16. [16]

    Mako: Speculative dis- tributed transactions with Geo-Replication,

    W. Shen, Y . Cui, S. Sen, S. Angel, and S. Mu, “Mako: Speculative dis- tributed transactions with Geo-Replication,” in19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25), 2025, pp. 129–152

  17. [17]

    Maat: Effective and scalable coordination of distributed transactions in the cloud,

    H. A. Mahmoud, V . Arora, F. Nawab, D. Agrawal, and A. El Abbadi, “Maat: Effective and scalable coordination of distributed transactions in the cloud,”Proceedings of the VLDB Endowment, vol. 7, no. 5, pp. 329–340, 2014

  18. [18]

    Easycommit: A non-blocking two-phase commit protocol

    S. Gupta and M. Sadoghi, “Easycommit: A non-blocking two-phase commit protocol.” inEDBT, 2018, pp. 157–168

  19. [19]

    Q-store: Distributed, multi- partition transactions via queue-oriented execution and communication

    T. Qadah, S. Gupta, and M. Sadoghi, “Q-store: Distributed, multi- partition transactions via queue-oriented execution and communication.” inEDBT, 2020, pp. 73–84

  20. [20]

    Enabling the next generation of multi-region applications with CockroachDB,

    N. VanBenschoten, A. Ajmani, M. Gartner, A. Matei, A. Shah, I. Sharif, A. Shraer, A. Storm, R. Taft, O. Tanet al., “Enabling the next generation of multi-region applications with CockroachDB,” inProceedings of the 2022 International Conference on Management of Data, 2022, pp. 2312– 2325

  21. [21]

    A study of database performance sensitivity to experiment settings

    Y . Wang, M. Yu, Y . Hui, F. Zhou, Y . Huang, R. Zhu, X. Ren, T. Li, and X. Lu, “A study of database performance sensitivity to experiment settings.”Proceedings of the VLDB Endowment, vol. 15, no. 7, 2022

  22. [22]

    TPC benchmark C (re- vision 5.11),

    Transaction Processing Performance Council, “TPC benchmark C (re- vision 5.11),” https://www.tpc.org/tpc documents current versions/pdf/ tpc-c v5.11.0.pdf, 2010

  23. [23]

    Benchmarking cloud serving systems with ycsb,

    B. F. Cooper, A. Silberstein, E. Tam, R. Ramakrishnan, and R. Sears, “Benchmarking cloud serving systems with ycsb,” inProceedings of the 1st ACM symposium on Cloud computing, 2010, pp. 143–154

  24. [24]

    The cost of serial- izability on platforms that use snapshot isolation,

    M. Alomari, M. Cahill, A. Fekete, and U. Rohm, “The cost of serial- izability on platforms that use snapshot isolation,” in2008 IEEE 24th international conference on data engineering. IEEE, 2008, pp. 576– 585

  25. [25]

    Gaia: Extended evaluation report,

    O. Mraz, K. Psarakis, G. Christodoulou, P. Carbone, and A. Katsifodi- mos, “Gaia: Extended evaluation report,” https://github.com/delftdata/ gaia/blob/main/appendix.pdf, 2026

  26. [26]

    Bonspiel: Low tail latency transactions in geo-distributed databases,

    F. Cui, E. Lo, S. Srivastava, and Z. Lai, “Bonspiel: Low tail latency transactions in geo-distributed databases,”Proceedings of the VLDB Endowment, vol. 18, no. 11, pp. 3840–3853, 2025

  27. [27]

    Dynamast: Adaptive dynamic mastering for replicated systems,

    M. Abebe, B. Glasbergen, and K. Daudjee, “Dynamast: Adaptive dynamic mastering for replicated systems,” in2020 IEEE 36th inter- national conference on data engineering (ICDE). IEEE, 2020, pp. 1381–1392

  28. [28]

    Zeus: locality-aware distributed transactions,

    A. Katsarakis, Y . Ma, Z. Tan, A. Bainbridge, M. Balkwill, A. Dragojevic, B. Grot, B. Radunovic, and Y . Zhang, “Zeus: locality-aware distributed transactions,” inProceedings of the Sixteenth European Conference on Computer Systems, 2021, pp. 145–161

  29. [29]

    Transaction management in the r* distributed database management system,

    C. Mohan, B. Lindsay, and R. Obermarck, “Transaction management in the r* distributed database management system,”ACM Transactions on Database Systems (TODS), vol. 11, no. 4, pp. 378–396, 1986

  30. [30]

    Amazon EC2 on-demand pricing,

    Amazon AWS, “Amazon EC2 on-demand pricing,” https://aws.amazon. com/ec2/pricing/on-demand/, Feb. 2025

  31. [31]

    Why tpc is not enough: An analysis of the amazon redshift fleet,

    A. van Renen, D. Horn, P. Pfeil, K. Vaidya, W. Dong, M. Narayanaswamy, Z. Liu, G. Saxena, A. Kipf, and T. Kraska, “Why tpc is not enough: An analysis of the amazon redshift fleet,” Proceedings of the VLDB Endowment, vol. 17, no. 11, pp. 3694–3706, 2024

  32. [32]

    Fast commit- ment for geo-distributed transactions via decentralized co-coordinators,

    Z. Zhang, H. Hu, X. Zhou, Y . Tu, W. Qian, and A. Zhou, “Fast commit- ment for geo-distributed transactions via decentralized co-coordinators,” Proceedings of the VLDB Endowment, vol. 17, no. 10, pp. 2555–2567, 2024

  33. [33]

    Mea- suring latency variation in the internet,

    T. Høiland-Jørgensen, B. Ahlgren, P. Hurtig, and A. Brunstrom, “Mea- suring latency variation in the internet,” inProceedings of the 12th International on Conference on emerging Networking EXperiments and Technologies, 2016, pp. 473–480

  34. [34]

    AWS latency monitoring,

    Amazon AWS, “AWS latency monitoring,” https://www.cloudping.co/, May 2026

  35. [35]

    Measuring network latency from a wireless isp: Variations within and across subnets,

    S. Sundberg, A. Brunstrom, S. Ferlin-Reiter, T. Høiland-Jørgensen, and R. Chac ´on, “Measuring network latency from a wireless isp: Variations within and across subnets,” inProceedings of the 2024 ACM on Internet Measurement Conference, 2024, pp. 29–43

  36. [36]

    Measuring network la- tency variation impacts to high performance computing application performance,

    R. Underwood, J. Anderson, and A. Apon, “Measuring network la- tency variation impacts to high performance computing application performance,” inProceedings of the 2018 ACM/SPEC International Conference on Performance Engineering, 2018, pp. 68–79

  37. [37]

    Measuring and improving the relia- bility of wide-area cloud paths,

    O. Haq, M. Raja, and F. R. Dogar, “Measuring and improving the relia- bility of wide-area cloud paths,” inProceedings of the 26th International Conference on World Wide Web, 2017, pp. 253–262

  38. [38]

    Lossradar: Fast detection of lost packets in data center networks,

    Y . Li, R. Miao, C. Kim, and M. Yu, “Lossradar: Fast detection of lost packets in data center networks,” inProceedings of the 12th International on Conference on emerging Networking EXperiments and Technologies, 2016, pp. 481–495

  39. [39]

    All networking pricing,

    Google GCP, “All networking pricing,” https://cloud.google.com/vpc/ network-pricing, Feb. 2025

  40. [40]

    Bandwidth pricing,

    Azure, “Bandwidth pricing,” https://azure.microsoft.com/en-us/pricing/ details/bandwidth/, Feb. 2025

  41. [41]

    iftop: display bandwidth usage on an interface,

    P. Warren and C. Lightfoot, “iftop: display bandwidth usage on an interface,” https://pdw.ex-parrot.com/iftop/, Feb. 2025

  42. [42]

    How AWS pricing works,

    Amazon AWS, “How AWS pricing works,” https://docs.aws.amazon. com/whitepapers/latest/how-aws-pricing-works/key-principles.html, Apr. 2026

  43. [43]

    Cloud storage cost: a taxonomy and survey,

    A. Q. Khan, M. Matskin, R. Prodan, C. Bussler, D. Roman, and A. Soylu, “Cloud storage cost: a taxonomy and survey,”World Wide Web, vol. 27, no. 4, p. 36, 2024

  44. [44]

    An efficient model to estimate and optimise the cloud migration costs from on- premises web apps,

    V . Prakash, A. Kumar, M. Shahid, L. Garg, and S. Bawa, “An efficient model to estimate and optimise the cloud migration costs from on- premises web apps,”Discover Computing, vol. 28, no. 1, p. 151, 2025

  45. [45]

    Cost modelling and optimisation for cloud: a graph-based approach,

    A. Q. Khan, M. Matskin, R. Prodan, C. Bussler, D. Roman, and A. Soylu, “Cost modelling and optimisation for cloud: a graph-based approach,”Journal of Cloud Computing, vol. 13, no. 1, p. 147, 2024

  46. [46]

    Introduction to network transformation on AWS,

    Amazon AWS, “Introduction to network transformation on AWS,” https://aws.amazon.com/blogs/networking-and-content-delivery/ introduction-to-network-transformation-on-aws-part-1/, May 2026

  47. [47]

    Network emulation with netem,

    S. Hemmingeret al., “Network emulation with netem,” inLinux conf au, vol. 5, 2005, p. 8

  48. [48]

    Cloud analytics benchmark,

    A. Van Renen and V . Leis, “Cloud analytics benchmark,”Proceedings of the VLDB Endowment, vol. 16, no. 6, pp. 1413–1425, 2023

  49. [49]

    Cdsben: Benchmarking the performance of storage services in cloud-native database system at ByteDance,

    J. Zhang, W. Jiang, B. Tang, H. Ma, L. Cao, Z. Jiang, Y . Nie, F. Wang, L. Zhang, and Y . Liang, “Cdsben: Benchmarking the performance of storage services in cloud-native database system at ByteDance,” Proceedings of the VLDB Endowment, vol. 16, no. 12, pp. 3584–3596, 2023

  50. [50]

    Lauca: Generating application-oriented synthetic workloads,

    Y . Li, R. Zhang, Y . Li, K. Shu, S. Zhang, and A. Zhou, “Lauca: Generating application-oriented synthetic workloads,”arXiv preprint arXiv:1912.07172, 2019

  51. [51]

    Cockroach Labs, “Morr,” https://www.cockroachlabs.com/docs/stable/ movr, 2025

  52. [52]

    Clay: Fine-grained adaptive partitioning for general database schemas,

    M. Serafini, R. Taft, A. J. Elmore, A. Pavlo, A. Aboulnaga, and M. Stonebraker, “Clay: Fine-grained adaptive partitioning for general database schemas,”Proceedings of the VLDB Endowment, vol. 10, no. 4, pp. 445–456, 2016

  53. [53]

    Why tpc is not enough: An analysis of the amazon redshift fleet,

    A. Van Renen, D. Horn, P. Pfeil, K. E. Vaidya, W. Dong, M. Narayanaswamy, Z. Liu, G. Saxena, A. Kipf, and T. Kraska, “Why tpc is not enough: An analysis of the amazon redshift fleet,”Proceedings of the VLDB Endowment, 2024

  54. [54]

    Workload insights from the snowflake data cloud: What do production analytic queries really look like?

    J. V . Szlang, S. Bress, S. Cattes, J. Dees, F. Funke, M. Heimel, M. Oleynik, I. Oukid, and T. Maltenberger, “Workload insights from the snowflake data cloud: What do production analytic queries really look like?”Proceedings of the VLDB Endowment, vol. 18, no. 12, pp. 5126–5138, 2025

  55. [55]

    Staring into the abyss: An evaluation of concurrency control with one thousand cores,

    X. Yu, G. Bezerra, A. Pavlo, S. Devadas, and M. Stonebraker, “Staring into the abyss: An evaluation of concurrency control with one thousand cores,”Proceedings of the VLDB Endowment, vol. 8, no. 3, 2014

  56. [56]

    OLTP through the looking glass, and what we found there,

    S. Harizopoulos, D. J. Abadi, S. Madden, and M. Stonebraker, “OLTP through the looking glass, and what we found there,”Proceedings of the 2008 ACM SIGMOD international conference on management of data, pp. 981–992, 2008

  57. [57]

    Towards cost-optimal query processing in the cloud,

    V . Leis and M. Kuschewski, “Towards cost-optimal query processing in the cloud,”Proceedings of the VLDB Endowment, vol. 14, no. 9, pp. 1606–1612, 2021

  58. [58]

    Eva: Cost-efficient cloud-based clus- ter scheduling,

    T.-T. Chang and S. Venkataraman, “Eva: Cost-efficient cloud-based clus- ter scheduling,” inProceedings of the Twentieth European Conference on Computer Systems. ACM New York, NY , USA, 2025, pp. 1399– 1416

  59. [59]

    Are database system researchers making correct assumptions about transaction work- loads?

    C. D. Nguyen, K. Chen, C. DeCarolis, and D. J. Abadi, “Are database system researchers making correct assumptions about transaction work- loads?”Proceedings of the ACM on Management of Data, vol. 3, no. 3, pp. 1–26, 2025

  60. [60]

    TPC-DI: The first industry benchmark for data integration,

    M. Poess, T. Rabl, H.-A. Jacobsen, and B. Caufield, “TPC-DI: The first industry benchmark for data integration,”Proceedings of the VLDB Endowment, vol. 7, no. 13, pp. 1367–1378, 2014

  61. [61]

    TPC benchmark E,

    Transaction Processing Performance Council, “TPC benchmark E,” https://www.tpc.org/tpc documents current versions/pdf/tpc-e v1. 12.0.pdf, 2010