REVIEW 1 major objections 5 minor 71 references
Enhancing OLAP Resilience at LinkedIn
T0 review · 1 major / 5 minor · reviewed 2026-07-15 · grok-4.5
Pith's one-line read A layered defense of workload isolation, zone-aware placement, and adaptive routing keeps petabyte OLAP SLAs stable under continuous churn.
desk verdict Solid production systems paper: three Pinot-specific mechanisms with real algorithms, inductive arguments, and fleet-scale evidence that other large OLAP operators will actually use. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The layered resiliency framework: Query Workload Isolation (sampling-based in-JVM accounting plus per-window budgets), Maintenance Zone–aware mirrored assignment with impact-free rebalance steps, and Adaptive Server Selection scoring servers inside each Mirror Server Set.
What would settle it
Run the zone-aware swap algorithm under a sustained, highly skewed zone inventory (one or two zones chronically under-provisioned) and check whether replica diversity, data-movement volume, and query SLAs still meet the claimed safety bounds during rebalance and zone drain.
Extended reading notes
Core claim
The paper claims that OLAP resilience at cloud scale is achieved by coordinating three layers inside the engine: measured per-workload CPU/memory budgets with host-local enforcement, topology-aware mirrored placement repaired by greedy swaps and impact-free drained rebalancing, and broker-side adaptive selection within mirror server sets. Together these absorb continuous operational entropy without treating resource management, placement, and routing as separate problems.
Load-bearing premise
The design assumes maintenance zones stay roughly evenly stocked with nodes and that replacements or scale-outs arrive from effectively random zones so the global balance still holds.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a production-deployed layered resiliency framework for Apache Pinot at LinkedIn that unifies three failure vectors: (1) Query Workload Isolation (QWI), an in-JVM sampling-based CPU/memory accounting and budget-enforcement system across brokers and servers with sub-1% overhead; (2) Maintenance Zone–aware replica placement via a greedy swap algorithm (Algorithm 3) plus an impact-free rebalance procedure (Algorithm 4) that drains query traffic before substantial segment movement; and (3) Adaptive Server Selection (ADSS), which scores servers within each Mirror Server Set using broker-local in-flight and latency signals (hybrid score with optional softmax) to divert traffic from fail-slow nodes. The mechanisms are evaluated with controlled trace replay on three clusters, simulation of ADSS edge cases, and production evidence including rollout to 1,500+ hosts, a 10 PB migration with zero SLA violations, 99.9% availability under up to 7% daily node churn, and ~90% reduction in fail-slow latency degradation.
Significance. If the reported production outcomes hold, this is a substantial systems contribution for real-time OLAP: it shows how workload isolation, topology-aware placement under mirrored replica groups, and adaptive routing can be co-designed inside a single engine rather than bolted on as external schedulers or cgroups. Strengths include an inductive correctness argument for Algorithm 3 under stated MZ assumptions, explicit safety invariants for Algorithm 4, open discussion of limitations (no bursting, heterogeneous-hardware calibration, ADSS inability to add capacity under systemic failure), and large-scale empirical evidence (overhead tables, isolation P95 recovery, 10 PB migration, 12-week degradation-prevention rate). The work is directly transferable to other multi-tenant OLAP systems that share a JVM and use replica-group scatter-gather, and the production metrics are external observations rather than quantities defined by the same equations used to claim success.
major comments (1)
- The central claim is empirically supported and the algorithms are carefully specified; I do not find a load-bearing technical error that would require major revision. The weakest assumption (roughly balanced MZ node counts under random replacement, §3.2 Assumptions 1–4) is already stated as a platform property, conditions the inductive progress argument for Algorithm 3, and is consistent with the reported 99.9% availability under real churn. Remaining gaps (QWI budget enforcement still rolling out incrementally after cost collection; no bursting/quota stealing; ADSS cannot add capacity under cluster-wide failure) are openly listed in the conclusion and do not falsify the production outcomes claimed.
minor comments (5)
- Table 2 and the production-scale paragraph (§2.6) would be clearer if the authors stated explicitly which of the 10 rolled-out clusters currently have full budget enforcement versus cost collection only, so readers can separate accounting overhead from enforcement effects.
- Figure 10 caption and surrounding text should state the exact simulation parameters (QPS, broker count, replica count, α, N, τ) used for the hybrid vs. softmax comparison so the oscillation-control claim is reproducible from the text alone.
- In §2.7, the GC discussion (G1 → Shenandoah → ZGC path) is operationally valuable but slightly informal; a short table of false-positive rates or heap-pressure kill counts under each GC would strengthen the lesson.
- Notation for Mirror Server Set (MSS) is introduced in §1.2 and reused in §4; a single forward reference or glossary entry would help readers who jump to the ADSS section.
- Related Work §5.1 correctly contrasts RU-based pre-classification (TiDB) with measured-cost charging; a one-sentence note on whether QWI budgets can be exported as signals to higher-level schedulers (YARN/K8s) would round out the positioning.
Circularity Check
No significant circularity: production-measured systems results, not inputs renamed as predictions.
full rationale
This is an empirical systems paper whose load-bearing claims are production observations (QWI overhead and isolation P95 recovery, ADSS diversion/prevention rates, 99.9% availability under churn, 10 PB zero-SLA-violation migration), not quantities derived from the same equations used to define success. QWI charges measured ThreadMXBean CPU and getThreadAllocatedBytes during execution and reports external before/after metrics (Tables 2, Figures 6–7); those outcomes are not forced by the budget configuration. The MZ-aware greedy swap (Algorithm 3) has an inductive progress argument conditioned on stated platform assumptions (balanced MZ node counts, random replacement); the proof is a conditional correctness argument, not a self-definitional loop, and production availability is an independent observation under real churn. ADSS hybrid scoring adapts external C3/ELS/Peak-EWMA ideas; softmax τ is set from an explicit design heuristic (1.5× healthy score → P_slow < 0.1%), then diversion effectiveness is measured in simulation and production rather than fitted to the final 90% figure. Self-citations (original Pinot architecture, Helix placement) supply background, not uniqueness theorems that force the central claims. No fitted-input-called-prediction, self-definitional reduction, or load-bearing self-citation chain is present.
Assumptions & free parameters
free parameters (6)
- QWI sampling interval =
1 ms
- QWI enforcement window W =
5–10 s
- ADSS hybrid exponent N =
~3
- ADSS softmax temperature τ =
≈0.07 × average score
- ADSS EMA α =
~2/3
- Heap-pressure kill thresholds =
96% / 99%
assumptions (5)
- domain assumption Pods/nodes are provisioned from a fixed number of Maintenance Zones and an even distribution across MZs is maintained (Assumption 1, §3.2).
- domain assumption Replacement or scale-out nodes arrive from effectively random MZs while the global balance assumption continues to hold.
- domain assumption Pinot’s static mirrored replica-group assignment and scatter-gather routing model (Section 1.2).
- domain assumption Broker-local in-flight and latency signals are sufficiently informative to detect fail-slow servers without server-side feedback.
- standard math Standard inductive reasoning and finite-state termination arguments for Algorithms 3 and 4.
invented entities (2)
-
Query Workload Isolation (QWI) framework (Budget Manager + Resource Accountant + Budget Enforcer)
independent evidence
-
Impact-Free Rebalance state model (Initial/Desired/Added/Removed/Down sets)
independent evidence
Cite this review
Pith. "Pith review of Enhancing OLAP Resilience at LinkedIn." pith.science (2026). https://pith.science/paper/IO5AXIAM
@misc{pith2026260307382,
author = {Pith},
title = {Pith review of: Enhancing OLAP Resilience at LinkedIn},
year = {2026},
howpublished = {\url{https://pith.science/paper/IO5AXIAM}},
note = {Machine review of arXiv:2603.07382}
}
read the original abstract
Real-time OLAP datastores are critical infrastructure for modern enterprises, powering interactive analytics on petabyte-scale datasets with subsecond latency requirements. As these systems become integral to service architectures, maintaining strict SLAs under failures, load spikes, and cluster changes is as important as raw performance. We present a set of resiliency mechanisms developed for Apache Pinot at LinkedIn, applicable to modern OLAP systems broadly. We introduce Query Workload Isolation (QWI), which provides workload-level CPU and memory budgeting across Pinot's broker and server tiers via fine-grained resource accounting and sub-millisecond enforcement, delivering predictable tail latency and fairness with under 1% overhead. We present Impact-Free Rebalancing for SLA-safe data movement during routine operations (e.g., upgrades, scale-out, and recovery), and Maintenance Zone Awareness to place replicas across fault domains and mitigate correlated failures. We also describe Adaptive Server Selection, which routes queries using real-time load and performance signals to avoid slow or failing nodes while preserving balanced utilization. Together, these mechanisms form a holistic resiliency framework deployed in production at LinkedIn, enabling stable query latency and high availability at scale.
Reference graph
Works this paper leans on
-
[1]
Kubernetes Docu- mentation
Kubernetes Documentation 2025.Local Ephemeral Storage. Kubernetes Docu- mentation. https://kubernetes.io/docs/concepts/storage/ephemeral-storage/
2025
-
[2]
Kubernetes Documentation
Kubernetes Documentation 2025.Pod Lifecycle. Kubernetes Documentation. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/
2025
-
[3]
Amazon Web Services. 2024. Amazon Kinesis. https://aws.amazon.com/kinesis/ Accessed: 2025-01-18
2024
-
[4]
Amazon Web Services. 2024. Placement Groups. https://docs.aws.amazon.com/ AWSEC2/latest/UserGuide/placement-groups.html Accessed: 2025-01-18
2024
-
[5]
Amazon Web Services. 2025. Amazon Redshift Workload Manage- ment. https://docs.aws.amazon.com/redshift/latest/dg/cm-c-implementing- workload-management.html Accessed: 2025-01-18
2025
-
[6]
Apache Cassandra. 2024. Data Replication. https://cassandra.apache.org/doc/ latest/cassandra/architecture/dynamo.html Accessed: 2025-01-18
2024
-
[7]
Apache Cassandra. 2024. Snitches. https://cassandra.apache.org/doc/latest/ cassandra/architecture/snitch.html Accessed: 2025-01-18
2024
-
[8]
Apache Doris. 2025. Workload Group — Apache Doris Documenta- tion. https://doris.apache.org/docs/admin-manual/workload-management/ workload-group/. Accessed: 2026-01-28
2025
Show all 71 references
-
[9]
Apache Druid. 2024. Basic Cluster Tuning. https://druid.apache.org/docs/latest/ operations/basic-cluster-tuning Accessed: 2025-01-18
2024
-
[10]
Apache Hadoop. 2019. Rack Awareness. https://hadoop.apache.org/docs/r3.1.2/ hadoop-project-dist/hadoop-common/RackAwareness.html Accessed: 2026-01- 28
2019
-
[11]
Apache Hadoop. 2024. Capacity Scheduler. https://hadoop.apache.org/docs/ current/hadoop-yarn/hadoop-yarn-site/CapacityScheduler.html Accessed: 2025- 01-18
2024
-
[12]
Apache Helix. 2025. CRUSHED Rebalancer in Apache Helix. https://helix. apache.org/1.3.1-docs/design_crushed.html Accessed: 2025-03-10
2025
-
[13]
Apache Helix. 2025. Weight-Aware Globally-Even Distribute Rebal- ancer. https://github.com/apache/helix/wiki/Weight-Aware-Globally-Even- Distribute-Rebalancer. Accessed: 2025-02-21
2025
-
[14]
Apache Software Foundation. 2024. Apache Hive Documentation. https: //hive.apache.org/ Accessed: 2025-01-18
2024
-
[15]
Apache Software Foundation. 2024. Apache Kafka. https://kafka.apache.org/ Accessed: 2025-01-18
2024
-
[16]
Apache Software Foundation. 2024. Apache Spark. https://spark.apache.org/ Accessed: 2025-01-18
2024
-
[17]
Apache Software Foundation. 2024. Hadoop MapReduce Tutorial. https://hadoop.apache.org/docs/current/hadoop-mapreduce-client/hadoop- mapreduce-client-core/MapReduceTutorial.html Accessed: 2025-01-18
2024
-
[18]
Nathan Bronson, Zach Amsden, George Cabrera, Prasad Chakka, Peter Dimov, Hui Ding, Jack Ferris, Anthony Giardullo, Sachin Kulkarni, Harry C Li, et al. 2013. TAO: Facebook’s Distributed Data Store for the Social Graph. InProceedings of the 2013 USENIX Annual Technical Conferenc...
2013
-
[19]
ClickHouse. 2025. Parallel Replicas. https://clickhouse.com/docs/deployment- guides/parallel-replicas Accessed: 2026-05-02
2025
-
[20]
Cockroach Labs. 2024. Configure Replication Zones. https://www.cockroachlabs. com/docs/stable/configure-replication-zones Accessed: 2025-01-18
2024
-
[21]
James C Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, JJ Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, et al. 2013. Spanner: Google’s Globally Distributed Database.ACM Transactions on Computer Systems31, 3 (2013), 1–22
2013
-
[22]
Benoit Dageville, Thierry Cruanes, Marcin Zukowski, Vadim Antonov, Artin Avanes, Jon Bock, Jonathan Claybaugh, Daniel Engovatov, Martin Hentschel, Jiansheng Huang, et al. 2016. The Snowflake Elastic Data Warehouse. InPro- ceedings of the 2016 International Conference on Manage...
2016
-
[23]
Jeffrey Dean and Luiz André Barroso. 2013. The Tail at Scale.Commun. ACM56, 2 (2013), 74–80
2013
-
[24]
Envoy Proxy. 2024. Load Balancing. https://www.envoyproxy.io/docs/envoy/ latest/intro/arch_overview/upstream/load_balancing/load_balancing Accessed: 2025-01-18
2024
-
[25]
Google Cloud. 2024. Google Cloud Pub/Sub. https://cloud.google.com/pubsub Accessed: 2025-01-18
2024
-
[26]
Google Cloud. 2025. BigQuery: Introduction to Slots and Workload Management. https://cloud.google.com/bigquery/docs/slots Accessed: 2025-01-18
2025
-
[27]
Rihan Hai, Christoph Quix, and Matthias Jarke. 2021. Data Lake Concept and Systems: A Survey.CoRRabs/2106.09592 (2021). https://arxiv.org/abs/2106.09592
2021 arXiv
-
[28]
Dongxu Huang, Qi Liu, Qiu Cui, Zhuhe Fang, Xiaoyu Ma, Fei Xu, Li Shen, Liu Tang, Yuxing Zhou, Menglong Huang, et al. 2020. TiDB: A Raft-based HTAP Database. InProceedings of the VLDB Endowment, Vol. 13. 3072–3084
2020
-
[29]
IBM. 2024. DB2 Workload Manager Guide and Reference. https://www.ibm. com/docs/en/db2 Accessed: 2025-01-18
2024
-
[30]
Jean-François Im, Kishore Gopalakrishna, Subbu Subramaniam, Mayank Shrivas- tava, Adwait Tumbde, Xiaotian Jiang, Jennifer Dai, Seunghyun Lee, Neha Pawar, Jialiang Li, and Ravi Aringunram. 2018. Pinot: Realtime OLAP for 530 Million Users. InProceedings of the 2018 International...
2018 doi
-
[31]
Kubernetes. 2024. Resource Quotas. https://kubernetes.io/docs/concepts/policy/ resource-quotas/ Accessed: 2025-01-18
2024
-
[32]
Kubernetes. 2024. Specifying a Disruption Budget for your Application. https: //kubernetes.io/docs/tasks/run-application/configure-pdb/ Accessed: 2025-01- 18
2024
-
[33]
Kubernetes. 2025. Pod Topology Spread Constraints. https://kubernetes.io/docs/ concepts/scheduling-eviction/topology-spread-constraints/ Accessed: 2025-01- 18
2025
-
[34]
X. Lei. 2017. Powering Helix’s Auto Rebalancer with Topology- Aware Partition Placement. LinkedIn Engineering Blog. https: //www.linkedin.com/blog/engineering/archive/powering-helix_s-auto- rebalancer-with-topology-aware-partition-p Accessed: 2026-05-08
2017
-
[35]
Viktor Leis, Peter Boncz, Alfons Kemper, and Thomas Neumann. 2014. Morsel- Driven Parallelism: A NUMA-Aware Query Evaluation Framework for the Many- Core Age. InProceedings of the 2014 ACM SIGMOD International Conference on Management of Data. 743–754
2014
-
[36]
LinkedIn. 2024. Who Viewed Your Profile. https://www.linkedin.com/help/ linkedin/answer/a540651 Accessed: 2025-01-18
2024
-
[37]
LinkedIn Engineering. 2024. Recommended Content Feed. https://engineering. linkedin.com/blog/topic/feed Accessed: 2025-01-18. 13
2024
-
[38]
LinkedIn Marketing Solutions. 2024. Reporting and Analytics. https://business. linkedin.com/marketing-solutions/reporting-analytics Accessed: 2025-01-18
2024
-
[39]
LinkedIn Talent Solutions. 2024. LinkedIn Talent Insights. https://business. linkedin.com/talent-solutions/talent-insights Accessed: 2025-01-18
2024
-
[40]
Ruiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue, and Peng Huang. 2025. One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed Systems. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, Ph...
2025
-
[41]
Ryan Marcus and Olga Papaemmanouil. 2016. WiSeDB: A Learning-based Work- load Management Advisor for Cloud Databases. InProceedings of the VLDB Endowment, Vol. 9. 780–791
2016
-
[42]
Olivier Michaelis, Stefan Schmid, and Habib Mostafaei. 2024. L3: Latency-aware Load Balancing in Multi-Cluster Service Mesh. InProceedings of the 25th Interna- tional Middleware Conference
2024
-
[43]
Microsoft. 2024. Use Scale-in Policies with Azure Virtual Machine Scale Sets. https://learn.microsoft.com/en-us/azure/virtual-machine-scale-sets/ virtual-machine-scale-sets-scale-in-policy Accessed: 2025-01-18
2024
-
[44]
Microsoft. 2025. Automatic Instance Repairs with Azure Virtual Machine Scale Sets. https://learn.microsoft.com/en-us/azure/virtual-machine-scale- sets/virtual-machine-scale-sets-automatic-instance-repairs Accessed: 2025-01- 18
2025
-
[45]
Microsoft. 2025. Availability Zones Overview. https://learn.microsoft.com/en- us/azure/reliability/availability-zones-overview Accessed: 2025-01-18
2025
-
[46]
MongoDB. 2024. Read Preference. https://www.mongodb.com/docs/manual/ core/read-preference/ Accessed: 2025-01-18
2024
-
[47]
Vivek Narasayya, Surajit Das, Manoj Syamala, Badrish Chandramouli, and Surajit Chaudhuri. 2013. SQLVM: Performance Isolation in Multi-Tenant Relational Database-as-a-Service. InProceedings of the 6th Biennial Conference on Innovative Data Systems Research (CIDR)
2013
-
[48]
2011.Workload Adaptation in Autonomic Database Management Systems
Baoning Niu. 2011.Workload Adaptation in Autonomic Database Management Systems. Ph.D. Dissertation. Queen’s University
2011
-
[49]
Oracle. 2024. Database Resource Manager. https://docs.oracle.com/en/database/ oracle/oracle-database/19/admin/managing-resources-with-oracle-database- resource-manager.html Accessed: 2025-01-18
2024
-
[50]
PingCAP. 2025. Use Resource Control to Achieve Resource Group Limitation and Flow Control — TiDB Documentation. https://docs.pingcap.com/tidb/stable/tidb- resource-control-ru-groups/. Accessed: 2026-01-28
2025
-
[51]
Waleed Reda, Marco Canini, Lalith Suresh, Dejan Kostić, and Sean Braithwaite
-
[52]
InProceedings of the ACM Symposium on Cloud Computing (SoCC)
Héron: Replica Selection for Heterogeneous Latencies. InProceedings of the ACM Symposium on Cloud Computing (SoCC). 380–393
-
[53]
Bart Samwel, John Cieslewicz, Ben Handy, Jason Gober, Petros Venetis, Chanjun Yang, Keith Peters, Jeff Shute, Daniel Tenedorio, Harnesh Apte, et al. 2018. F1 Query: Declarative Querying at Scale. InProceedings of the VLDB Endowment, Vol. 11. 1835–1848
2018
-
[54]
SAP. 2024. SAP HANA Workload Management. https://help.sap.com/docs/ SAP_HANA_PLATFORM Accessed: 2025-01-18
2024
-
[55]
Shopify Engineering. 2023. Shard Balancing: Moving Shops Confidently with Zero-Downtime at Terabyte-scale. https://shopify.engineering/mysql-database- shard-balancing-terabyte-scale Accessed: 2025-01-18
2023
-
[56]
2023.Liminal: Pre- dictable Resource Scheduling for LinkedIn’s Query Workloads
Avinash Shukla, Shivaram Venkataraman, and Ankit Soni. 2023.Liminal: Pre- dictable Resource Scheduling for LinkedIn’s Query Workloads. Technical Report. LinkedIn Engineering
2023
-
[57]
Spotify Engineering. 2015. ELS, Part 2 – The Trees. https://engineering.atspotify. com/2015/12/els-part-2/ Accessed: 2025-02-13
2015
-
[58]
StarRocks. 2025. Resource Group — StarRocks Documentation. https://docs.starrocks.io/docs/administration/management/resource_ management/resource_group/. Accessed: 2026-01-28
2025
-
[59]
Lalith Suresh, Marco Canini, Stefan Schmid, and Anja Feldmann. 2015. C3: Cutting Tail Latency in Cloud Data Stores via Adaptive Replica Selection. In12th USENIX Symposium on Networked Systems Design and Implementation (NSDI). 513–527
2015
-
[60]
Rebecca Taft, Irfan Sharber, Andrei Matei, Nathan VanBenschoten, Jordan Lewis, Tobias Grieger, Kai Niber, Andy Woods, Anne Biber, Raphael Poss, et al. 2020. CockroachDB: The Resilient Geo-Distributed SQL Database. InProceedings of the 2020 ACM SIGMOD International Conference o...
2020
-
[61]
The Linux Kernel. 2025. Control Group v2 — The Linux Kernel Documentation. https://www.kernel.org/doc/html/latest/admin-guide/cgroup-v2.html. Accessed: 2026-01-28
2025
-
[62]
Steven Tozer, Tim Brecht, and Ashraf Aboulnaga. 2010. Q-Cop: Avoiding Bad Query Mixes to Minimize Client Timeouts under Heavy Loads. InProceedings of the 26th International Conference on Data Engineering (ICDE). 397–408
2010
-
[63]
Twitter Finagle. 2017. PeakEwma.scala. https://github.com/ twitter/finagle/blob/9cc08d15216497bb03a1cafda96b7266cfbbcff1/finagle- core/src/main/scala/com/twitter/finagle/loadbalancer/PeakEwma.scala Ac- cessed: 2026-05-02
2017
-
[64]
Abhishek Verma, Luis Pedrosa, Madhukar Korupolu, David Oppenheimer, Eric Tune, and John Wilkes. 2015. Large-scale Cluster Management at Google with Borg. InProceedings of the 10th European Conference on Computer Systems (Eu- roSys). 1–17
2015
-
[65]
Vitess. 2024. Vitess: A Database Clustering System for Horizontal Scaling of MySQL. https://vitess.io/ Accessed: 2025-01-18
2024
-
[66]
Sage A Weil, Scott A Brandt, Ethan L Miller, and Carlos Maltzahn. 2006. CRUSH: Controlled, Scalable, Decentralized Placement of Replicated Data. InProceedings of the 2006 ACM/IEEE Conference on Supercomputing. 122
2006
-
[67]
Wikipedia contributors. 2024. Moving Average – Exponential Moving Average. https://en.wikipedia.org/wiki/Moving_average Accessed: 2025-02-13
2024
-
[68]
Fangjin Yang, Eric Tschetter, Xavier Léauté, Nelson Ray, Gian Merlino, and Deep Ganguli. 2014. Druid: A Real-time Analytical Data Store. InProceedings of the 2014 ACM SIGMOD International Conference on Management of Data. 157–168
2014
-
[69]
Yugabyte. 2024. Multi-Zone and Multi-Region Deployments. https://docs. yugabyte.com/preview/deploy/multi-dc/ Accessed: 2025-01-18
2024
-
[70]
Xiang Zhou, Ji Sun, Ryan Marcus, and Olga Papaemmanouil. 2023. IconqSched: Query Scheduling with Learned Concurrency Models. InProceedings of the 2023 International Conference on Management of Data (SIGMOD). 1–15
2023
-
[71]
Yuhang Zhou, Zhibin Wang, Peng Jiang, Haoran Xia, Junhe Lu, Qianyu Jiang, Rong Gu, Hengxi Xu, Xinjing Huang, Guanghuan Fang, Zhiheng Hu, Jingyi Zhang, Yongjin Cai, Jian He, and Chen Tian. 2025. Odyssey: Adaptive Policy Selection for Resilient Distributed Training. arXiv:2508.2...
2025 arXiv
Reviewed July 15, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.