REVIEW 3 major objections 4 minor 65 references
Skyrise: Exploiting Serverless Cloud Infrastructure for Elastic Data Processing
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Skyrise shows a complete SQL engine can run on serverless functions alone and match cloud systems on TPC-H.
desk verdict Skyrise delivers the first fully serverless SQL engine, but the abstract's TPC-H competitiveness claim is broader than the three scan-heavy queries it actually evaluates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a per-query coordinator running as a serverless function. It compiles SQL through a rule-based logical optimizer and a physical optimizer, breaks the plan into data-parallel pipelines, and for each pipeline chooses a worker count from the input size and the network burst capacity of a function. Workers are invoked via a two-level fan-out, fetch only the relevant columnar slices from object storage in parallel, and push vectorized batches to the final operator, which writes a single output object. Because every worker is stateless, the coordinator can re-trigger stragglers mid-query without invalidating results — an idempotent worker recomputes and overwrites the same deterministic file. The second mechanism is an intermediate-result registry: the coordinator checks a hash of the logically optimized plan before scheduling a pipeline and skips it if a matching result already exists.
What would settle it
Run the join- and shuffle-heavy TPC-H queries, for example Q5, Q8, or Q9, at scale factor 1,000 on Skyrise and compare latency and cost against Snowflake or Athena. If those queries are several times slower or more expensive than the scan-heavy set, the paper's general claim that Skyrise is competitive on the analytical TPC-H benchmark would be refuted.
Extended reading notes
Core claim
Skyrise is the first end-to-end SQL query processor built entirely on serverless infrastructure: the query coordinator and all workers run as cloud functions, data is read from and written to object storage, and no server is involved at any point. To make this viable, Skyrise compiles SQL into staged pipelines, sizes the number of workers per pipeline from input size and network capacity, invokes workers through a two-level fan-out to speed up cluster startup, and aggressively re-triggers straggling workers or storage requests. Workers are stateless and idempotent, writing a single deterministic output file, so a re-triggered worker can overwrite a racing copy safely and aborted queries can resume from stage checkpoints. A cache of intermediate results, keyed by a hash of the logically optimized plan, lets recurring queries skip pipelines. In the evaluation, Skyrise's runtime and cost on TPC-H Q1, Q6, and Q12 at scale factor 1,000 are on par with the research baseline and competitive with the commercial systems, and query latency grows by less than an order of magnitude as input size grows from about 1 GB to about 10 TB.
Load-bearing premise
The evaluation rests on three scan-heavy TPC-H queries (Q1, Q6, Q12) that the authors chose because they are not dominated by shuffling; the paper's broad claim of competitive performance and cost for terabyte-scale TPC-H would fail if join- and shuffle-heavy queries perform much worse.
Editorial extensions
If this is right
- A full SQL engine can operate with zero idle infrastructure: no coordinator, no worker pool, and no shuffle service are kept running between queries.
- For scan-heavy analytical workloads, serverless execution is cost-competitive with provisioned warehouses and with query-as-a-service offerings.
- Elasticity becomes automatic: worker count and query cost track input size, with latency staying within one order of magnitude across five orders of magnitude of data.
- Recurring or overlapping queries become cheaper because intermediate results are reused via the serverless storage cache.
- Stragglers, a known weakness of function-as-a-service platforms, can be handled by re-triggering idempotent workers without compromising correctness.
Reading between the lines
- The paper leaves join- and shuffle-heavy TPC-H queries untested; if those also stay competitive, the result would extend serverless processing from scan-dominated analytics to general decision-support workloads, but that remains an open bet.
- The same architecture could plausibly be ported to other function-as-a-service providers, since the design relies only on functions, object storage, a queue, and a metadata store, though the paper only evaluates one cloud environment.
- The result cache keyed on a logical-plan hash implies that semantically equivalent queries with different physical plans can share intermediate results, a form of cross-query reuse the paper does not explore.
- Because the bottleneck at 10 TB is object-store request limits and function-invocation stragglers, the next limit to test is whether hotter storage tiers remove the remaining latency gap.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Skyrise, a query processor built entirely on serverless AWS infrastructure: the coordinator and workers run as Lambda functions, all persistent and intermediate data resides on serverless object storage, and messaging uses SQS. The system compiles SQL through a rule-based logical optimizer and a physical optimizer that introduces pipeline breakers and chooses shuffle strategies, then executes stage-wise with adaptive retriggering of straggling workers and an intermediate-result cache. The evaluation compares Skyrise against Lambada, Athena, and several Snowflake configurations on TPC-H at scale factor 1000, measuring runtime and cost for the three scan-heavy queries Q1, Q6, and Q12, and reports an elasticity experiment from SF 1 to SF 10000. The central claims are that Skyrise is the first fully serverless query processor with no server-based component and that its performance and cost are competitive to other cloud data systems for terabyte-scale TPC-H queries.
Significance. If the architecture and measurements hold, the paper makes a notable contribution: it is one of the first open-source systems to demonstrate that a complete SQL query processor, including coordinator, workers, and shuffle path, can run end-to-end on FaaS plus object storage while scaling from zero workers to thousands. The adaptive worker retriggering and the serverless intermediate-result cache are clearly described and are reasonable engineering contributions. The elasticity experiment, showing query latency within one order of magnitude across five orders of magnitude of data size, is an informative and rare measurement for this class of systems. The main weakness is that the empirical support for the headline TPC-H competitiveness claim is limited to three deliberately scan-heavy queries, and one of the comparison baselines (Lambada) is not measured on the same dataset.
major comments (3)
- [Abstract; §4.2] The abstract and Section 1 claim that 'both Skyrise's performance and cost are competitive to other cloud data systems for terabyte-scale queries of the analytical TPC-H benchmark.' The evaluation in §4.2 supports this only for the three scan-heavy queries Q1, Q6, and Q12, which the authors themselves describe as 'well-suited for serverless execution as they are not dominated by shuffling operators.' No TPC-H query with multi-way joins, substantial grouping and sorting, or large data movement is tested, even though the tiered shuffle mechanism and the physical optimizer's repartition/broadcast join decisions are central components of the architecture. The observed scan performance therefore does not establish the broad TPC-H competitiveness claim; the claim should be narrowed to scan-heavy analytical workloads unless join- or shuffle-heavy TPC-H queries are evaluated.
- [§4.2.1, Fig. 5] The Lambada comparison uses runtimes and costs published in [22] rather than measurements on the same dataset. The paper itself notes in §4.2.1 that the Lambada dataset replaces string columns with integers and sorts lineitem on l_shipdate, which enables partition pruning; both differences directly reduce I/O and parsing costs for the selected queries. Consequently, the statement in §4.2.2 that Skyrise achieves latencies and costs 'comparable' to Lambada is not supported by a like-for-like comparison. Please re-run Lambada on the identical dataset and Parquet files, or clearly present the comparison as indicative and quantify the expected effect of the dataset differences.
- [§4.2, Figs. 5–6] All comparative latency and cost figures report only the median of five executions, with no measure of spread. Given that the paper's own elasticity experiment (Fig. 7) shows considerable variance at larger scale factors and that serverless functions are described as straggler-prone, the median alone is insufficient to determine whether observed differences between Skyrise and the comparison systems are meaningful. Please report per-execution values or at least min/max ranges for each configuration, or justify why the median is sufficient.
minor comments (4)
- [§4.3] The abbreviation 'OOM' is used for 'order of magnitude' without being defined; the phrase 'less than an order of magnitude (OOM) difference' is also redundant and should be simplified.
- [Abstract] The phrase 'competitive to other cloud data systems' should be 'competitive with other cloud data systems'.
- [§4.2.2, Fig. 6] The sentence 'Athena being 7 × cheaper on average' is ambiguous because it does not state which configuration it is compared with; clarify the comparison and whether the ratio is per query or based on a different normalization.
- [Table 3] The header 'T ail Latency' contains a typo ('T ail'), and the entry '>1k' for S3 Standard read tail latency should specify units and the percentile measured.
Circularity Check
No circularity: Skyrise is an experimental systems paper whose claims are evaluated against external TPC-H benchmarks and commercial/academic systems, with no fitted parameter renamed as a prediction.
full rationale
The paper contains no derivation chain that reduces to its own inputs. Skyrise's headline claims—first fully serverless query processor and competitive TPC-H performance and cost—are supported by an empirical evaluation in Section 4 against Athena, Snowflake, and Lambada, using the external TPC-H benchmark. The one parameter that comes from prior work is the worker-count sizing rule: 'The number of workers per pipeline is based on the total input size and the network burst capacity per function, as determined in prior work [42]' (Section 3.2). That is a configuration heuristic taken from the authors' empirical infrastructure study [42], not a fitted parameter that is later reported as a prediction of Skyrise's performance; the measured latencies and costs are actual experimental results. The paper also explicitly states 'we deactivate the query result cache' (Section 4.1), so the measured performance is not produced by re-reading cached outputs. The abstract's generality is broader than the tested three scan-heavy queries (Q1, Q6, Q12), but that is an external-validity or overclaim concern, not circularity: the evidence is independent of the conclusion. The self-citations [36, 42, 44] are not load-bearing in the sense of substituting for evidence; they point to the system's prior design and to an empirical infrastructure study, while the central experimental comparison stands on external systems and benchmark data. No circular step can be exhibited with a specific equation or fitted-input reduction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption AWS Lambda, S3, SQS, DynamoDB, and Glue pricing and performance characteristics (Tables 1-3) are accurate and stable for the evaluation period.
- domain assumption TPC-H benchmark at scale factor 1,000 is representative of analytical cloud workloads.
- domain assumption Published Lambada results [22] provide a valid baseline even though the datasets differ (integer-encoded strings, sorted lineitem).
Cite this review
Pith. "Pith review of Skyrise: Exploiting Serverless Cloud Infrastructure for Elastic Data Processing." pith.science (2026). https://pith.science/paper/2VNIIT2U
@misc{pith2026250108479,
author = {Pith},
title = {Pith review of: Skyrise: Exploiting Serverless Cloud Infrastructure for Elastic Data Processing},
year = {2026},
howpublished = {\url{https://pith.science/paper/2VNIIT2U}},
note = {Machine review of arXiv:2501.08479}
}
read the original abstract
Serverless computing offers elasticity unmatched by conventional server-based cloud infrastructure. Although modern data processing systems embrace serverless storage, such as Amazon S3, they continue to manage their compute resources as servers. This is challenging for unpredictable workloads, leaving clusters often underutilized. Recent research shows the potential of serverless compute resources, such as cloud functions, for elastic data processing, but also sees limitations in performance robustness and cost efficiency for long running workloads. These challenges require holistic approaches across the system stack. However, to the best of our knowledge, there is no end-to-end data processing system built entirely on serverless infrastructure. In this paper, we present Skyrise, our effort towards building the first fully serverless SQL query processor. Skyrise exploits the elasticity of its underlying infrastructure, while alleviating the inherent limitations with a number of adaptive and cost-aware techniques. We show that both Skyrise's performance and cost are competitive to other cloud data systems for terabyte-scale queries of the analytical TPC-H benchmark.
Reference graph
Works this paper leans on
-
[22]
M¨ uller, I., Marroqu ´ ın, R., Alonso, G.: Lam- bada: Interactive Data Analytics on Cold Data Using Serverless Cloud Infrastructure. In: ACM SIGMOD, pp. 115–130 (2020)
work page 2020
-
[1]
Gupta, A., Agarwal, D., Tan, D., Kulesza, J., Pathak, R., Stefani, S., Srinivasan, V.: Amazon Redshift and the Case for Simpler Data Warehouses. In: ACM SIGMOD, pp. 1917–1923 (2015)
work page 2015
-
[2]
Melnik, S., Gubarev, A., Long, J.J., Romer, G., Shivakumar, S., Tolton, M., Vassilakis, T.: Dremel: Interactive Analysis of Web-Scale Datasets. PVLDB 3(1), 330–339 (2010)
work page 2010
-
[4]
https: //itmarketstrategy.com/2022/12/09/dbms- market-transformation-2021-the-big-picture
IT Market Strategy: DBMS Market Trans- formation 2021: The Big Picture. https: //itmarketstrategy.com/2022/12/09/dbms- market-transformation-2021-the-big-picture. Accessed: 2024-10-07 (2022)
work page 2022
-
[5]
https://itmarketstrategy.com/2024/05/31/ dbms-market-2023-more-momentum-shifts
IT Market Strategy: DBMS Mar- ket 2023 – More Momentum Shifts. https://itmarketstrategy.com/2024/05/31/ dbms-market-2023-more-momentum-shifts. Accessed: 2024-10-07 (2024)
work page 2024
-
[6]
Singh, A., Ong, J., Agarwal, A., Anderson, G., Armistead, A., Bannon, R., Boving, S., Desai, G., Felderman, B., Germano, P., Kana- gala, A., Provost, J., Simmons, J., Tanda, E., Wanderer, J., H¨ olzle, U., Stuart, S., Vahdat, A.: Jupiter Rising: A Decade of Clos Topolo- gies and Centralized Control in Google’s Dat- acenter Network. In: ACM SIGCOMM, pp. 18...
work page 2015
-
[7]
Firestone, D., Putnam, A., Mundkur, S., Chiou, D., Dabagh, A., Andrewartha, M., Angepat, H., Bhanu, V., Caulfield, A.M., Chung, E.S., Chandrappa, H.K., Chaturmo- hta, S., Humphrey, M., Lavier, J., Lam, N., Liu, F., Ovtcharov, K., Padhye, J., Popuri, G., Raindel, S., Sapre, T., Shaw, M., Silva, G., Sivakumar, M., Srivastava, N., Verma, A., Zuhair, Q., Bans...
work page 2018
-
[8]
https:// aws.amazon.com/ec2/nitro/
Amazon: A WS Nitro System. https:// aws.amazon.com/ec2/nitro/. Accessed: 2024- 10-07 (2024)
work page 2024
Show all 65 references
-
[9]
https: //aws.amazon.com/blogs/aws/new-per- second-billing-for-ec2-instances-and-ebs- volumes
Amazon: Per-Second Billing for EC2 Instances and EBS Volumes. https: //aws.amazon.com/blogs/aws/new-per- second-billing-for-ec2-instances-and-ebs- volumes. Accessed: 2024-10-07 (2017)
2017
-
[10]
https:// cloud.google.com/pricing/
Google: Google Cloud Pricing. https:// cloud.google.com/pricing/. Accessed: 2024- 10-07 (2024)
2024
-
[11]
In: ACM SIGMOD, pp
Armenatzoglou, N., Basu, S., Bhanoori, N., Cai, M., Chainani, N., Chinta, K., Govin- daraju, V., Green, T.J., Gupta, M., Hillig, S., Hotinger, E., Leshinksy, Y., Liang, J., McCreedy, M., Nagel, F., Pandis, I., Parchas, P., Pathak, R., Polychroniou, O., Rahman, F., Saxena, G., ...
2022
-
[12]
PVLDB 13(12), 3461–3472 (2020)
Melnik, S., Gubarev, A., Long, J.J., Romer, G., Shivakumar, S., Tolton, M., Vassi- lakis, T., Ahmadi, H., Delorey, D., Min, S., Pasumansky, M., Shute, J.: Dremel: A Decade of Interactive SQL Analysis at Web Scale. PVLDB 13(12), 3461–3472 (2020)
2020
-
[13]
In: USENIX NSDI, pp
Vuppalapati, M., Miron, J., Agarwal, R., 10 Truong, D., Motivala, A., Cruanes, T.: Build- ing An Elastic Query Engine on Disaggre- gated Storage. In: USENIX NSDI, pp. 449– 462 (2020)
2020
-
[14]
https: //aws.amazon.com/lambda/
Amazon: A WS Lambda. https: //aws.amazon.com/lambda/. Accessed: 2024-10-07 (2023)
2023
-
[15]
https: //azure.microsoft.com/services/functions/
Microsoft: Azure Functions: Execute Event-driven Serverless Code with an End- to-end Development Experience. https: //azure.microsoft.com/services/functions/. Accessed: 2024-03-27 (2024)
2024
-
[16]
In: ACM SoCC, pp
Jonas, E., Pu, Q., Venkataraman, S., Stoica, I., Recht, B.: Occupy the Cloud: Distributed Computing for the 99%. In: ACM SoCC, pp. 445–451 (2017)
2017
-
[17]
In: IEEE IISWC, pp
Ustiugov, D., Amariucai, T., Grot, B.: Ana- lyzing Tail Latency in Serverless Clouds with STeLLAR. In: IEEE IISWC, pp. 51–62 (2021)
2021
-
[18]
In: ACM SOSP, pp
Seemakhupt, K., Stephens, B.E., Khan, S.M., Liu, S., Wassel, H.M.G., Yeganeh, S.H., Sno- eren, A.C., Krishnamurthy, A., Culler, D.E., Levy, H.M.: A Cloud-Scale Characterization of Remote Procedure Calls. In: ACM SOSP, pp. 498–514 (2023)
2023
-
[19]
https: //parquet.apache.org/
Parquet contributors: Apache Parquet. https: //parquet.apache.org/. Accessed: 2024-10-07 (2024)
2024
-
[20]
https: //aws.amazon.com/serverless/
Amazon: Serverless on A WS. https: //aws.amazon.com/serverless/. Accessed: 2024-03-27 (2024)
2024
-
[21]
https:// cloud.google.com/serverless
Google: Serverless. https:// cloud.google.com/serverless. Accessed: 2024-03-27 (2024)
2024
-
[23]
In: ACM SIG- MOD, pp
Perron, M., Fernandez, R.C., DeWitt, D.J., Madden, S.: Starling: A Scalable Query Engine on Cloud Functions. In: ACM SIG- MOD, pp. 131–141 (2020)
2020
-
[24]
https://aws .amazon.com/ athena/
Amazon: Amazon Athena: Analyze Petabyte- scale Data Where it Lives with Ease and Flexibility. https://aws .amazon.com/ athena/. Accessed: 2024-10-07 (2024)
2024
-
[25]
https://aws.amazon.com/redshift/redshift- serverless/
Amazon: Amazon Redshift Serverless. https://aws.amazon.com/redshift/redshift- serverless/. Accessed: 2024-10-07 (2024)
2024
-
[26]
https: //docs.snowflake.com/en/user-guide/tasks- intro#serverless-tasks
Snowflake: Serverless Tasks. https: //docs.snowflake.com/en/user-guide/tasks- intro#serverless-tasks. Accessed: 2024-10-07 (2024)
2024
-
[27]
https:// aws.amazon.com/s3/
Amazon: Amazon S3. https:// aws.amazon.com/s3/. Accessed: 2024-10-07 (2023)
2023
-
[28]
https:// cloud.google.com/storage/
Google: Cloud Storage. https:// cloud.google.com/storage/. Accessed: 2024-03-27 (2024)
2024
-
[29]
https: //docs.aws.amazon.com/lambda/latest/ dg/gettingstarted-limits.html
Amazon: Lambda Quotas. https: //docs.aws.amazon.com/lambda/latest/ dg/gettingstarted-limits.html. Accessed: 2024-03-28 (2024)
2024
-
[30]
https://aws.amazon.com/ec2/instance- types/c6g/
Amazon: Amazon EC2 C6g Instances. https://aws.amazon.com/ec2/instance- types/c6g/. Accessed: 2024-03-27 (2023)
2023
-
[31]
https://azure.microsoft.com/products/ storage/blobs/
Microsoft: Azure Blob Storage: Mas- sively Scalable and Secure Object Storage for Cloud-native Workloads, Archives, Data Lakes, High-performance Computing, and Machine Learning. https://azure.microsoft.com/products/ storage/blobs/. Accessed: 2024-03-27 (2024)
2024
-
[32]
https://aws .amazon.com/ about-aws/whats-new/2018/11/announcing- amazon-dynamodb-on-demand/
Amazon: Announcing Amazon DynamoDB On-Demand. https://aws .amazon.com/ about-aws/whats-new/2018/11/announcing- amazon-dynamodb-on-demand/. Accessed: 2024-03-26 (2018)
2018
-
[33]
https: //cloud.google.com/firestore/
Google: Cloud Firestore. https: //cloud.google.com/firestore/. Accessed: 2024-03-27 (2024) 11
2024
-
[34]
https://aws .amazon.com/ blogs/aws/new-announcing-amazon-efs- elastic-throughput/
Amazon: Announcing Amazon EFS Elas- tic Throughput. https://aws .amazon.com/ blogs/aws/new-announcing-amazon-efs- elastic-throughput/. Accessed: 2024-03-26 (2022)
2022
-
[35]
https://azure .microsoft.com/ products/storage/files/
Microsoft: Azure Files: Simple, Secure, and Serverless Enterprise-grade Cloud File Shares. https://azure .microsoft.com/ products/storage/files/. Accessed: 2024-03-27 (2024)
2024
-
[36]
In: VLDB PhD Workshop (2020)
Bodner, T.: Elastic Query Processing on Function as a Service Platforms. In: VLDB PhD Workshop (2020)
2020
-
[37]
In: EDBT, pp
Dreseler, M., Kossmann, J., Boissier, M., Klauck, S., Uflacker, M., Plattner, H.: Hyrise Re-engineered: An Extensible Database Sys- tem for Research in Relational In-Memory Data Management. In: EDBT, pp. 313–324 (2019)
2019
-
[38]
https: //www.datadoghq.com/state-of-serverless/
Datadog: The State of Serverless 2023. https: //www.datadoghq.com/state-of-serverless/. Accessed: 2024-03-28 (2023)
2023
-
[39]
Journal of Systems and Software 170, 110708 (2020)
Scheuner, J., Leitner, P.: Function-as-a- Service Performance Evaluation - A Multi- vocal Literature Review. Journal of Systems and Software 170, 110708 (2020)
2020
-
[40]
https://docs.aws.amazon.com/lambda/ latest/dg/urls-invocation.html
Amazon: Invoking Lambda Function URLs. https://docs.aws.amazon.com/lambda/ latest/dg/urls-invocation.html. Accessed: 2024-03-25 (2024)
2024
-
[41]
PVLDB 13(8), 1206– 1220 (2020)
Dreseler, M., Boissier, M., Rabl, T., Uflacker, M.: Quantifying TPC-H Choke Points and Their Optimizations. PVLDB 13(8), 1206– 1220 (2020)
2020
-
[42]
arXiv preprint arXiv:2501.07771 (2025) arXiv:2501.07771 [cs.DB]
Bodner, T., Radig, T., Justen, D., Rit- ter, D., Rabl, T.: An Empirical Evalua- tion of Serverless Cloud Infrastructure for Large-Scale Data Processing. arXiv preprint arXiv:2501.07771 (2025) arXiv:2501.07771 [cs.DB]
2025 arXiv
-
[43]
https://aws .amazon.com/s3/ storage-classes/express-one-zone/
Amazon: Amazon S3 Express One Zone Storage Class. https://aws .amazon.com/s3/ storage-classes/express-one-zone/. Accessed: 2024-03-25 (2024)
2024
-
[44]
In: Joint Proceedings of Workshops at the 49th International Con- ference on Very Large Data Bases
Mahling, F., R¨ oßler, P., Bodner, T., Rabl, T.: BabelMR: A Polyglot Framework for Server- less MapReduce. In: Joint Proceedings of Workshops at the 49th International Con- ference on Very Large Data Bases. CEUR Workshop Proceedings, vol. 3462 (2023). https://ceur-ws.org/Vol-3...
2023
-
[45]
PVLDB 16(11), 2769–2782 (2023)
Durner, D., Leis, V., Neumann, T.: Exploiting Cloud Object Storage for High- Performance Analytics. PVLDB 16(11), 2769–2782 (2023)
2023
-
[46]
Ailamaki, A., DeWitt, D.J., Hill, M.D.: Data Page Layouts for Relational Databases on Deep Memory Hierarchies. VLDB J. 11(3), 198–215 (2002)
2002
-
[47]
https:// orc.apache.org/
ORC contributors: Apache ORC. https:// orc.apache.org/. Accessed: 2024-03-26 (2024)
2024
-
[48]
https: //docs.aws.amazon.com/A WSEC2/latest/ UserGuide/using-regions-availability- zones.html#concepts-availability-zones
Amazon: Regions and Zones. https: //docs.aws.amazon.com/A WSEC2/latest/ UserGuide/using-regions-availability- zones.html#concepts-availability-zones. Accessed: 2024-03-28 (2024)
2024
-
[49]
https: //docs.aws.amazon.com/lambda/latest/dg/ foundation-arch.html
Amazon: Lambda Instruction Set Architectures (ARM/x86). https: //docs.aws.amazon.com/lambda/latest/dg/ foundation-arch.html. Accessed: 2024-03-27 (2024)
2024
-
[50]
https://www.snowflake.com/de/data-cloud/ platform/
Snowflake: Snowflake Cloud Data Platform. https://www.snowflake.com/de/data-cloud/ platform/. Accessed: 2024-10-07 (2024)
2024
-
[51]
https://www.tpc.org/tpch/
Transaction Processing Performance Coun- cil: Specification of the TPC-H Benchmark. https://www.tpc.org/tpch/. Accessed: 2024- 03-27 (2024)
2024
-
[52]
In: IEEE ICDE, pp
Sethi, R., Traverso, M., Sundstrom, D., Phillips, D., Xie, W., Sun, Y., Yegitbasi, N., Jin, H., Hwang, E., Shingte, N., Berner, C.: Presto: SQL on Everything. In: IEEE ICDE, pp. 1802–1813 (2019)
2019
-
[53]
In: IEEE CLOUD, pp
Kim, Y., Lin, J.: Serverless Data Analytics with Flint. In: IEEE CLOUD, pp. 451–455 12 (2018)
2018
-
[54]
In: USENIX NSDI, pp
Pu, Q., Venkataraman, S., Stoica, I.: Shuf- fling, Fast and Slow: Scalable Analytics on Serverless Infrastructure. In: USENIX NSDI, pp. 193–206 (2019)
2019
-
[55]
In: USENIX NSDI, pp
Zhang, H., Tang, Y., Khandelwal, A., Chen, J., Stoica, I.: Caerus: NIMBLE Task Schedul- ing for Serverless Analytics. In: USENIX NSDI, pp. 653–669 (2021)
2021
-
[56]
https://github.com/bcongdon/corral/
Congdon, B.: Corral: A Serverless MapRe- duce Framework Written for A WS Lambda. https://github.com/bcongdon/corral/. Accessed: 2024-03-26 (2018)
2018
-
[57]
Proceedings of the ACM on Management of Data 1(4), 233–123325 (2023)
Perron, M., Fernandez, R.C., DeWitt, D.J., Cafarella, M.J., Madden, S.: Cackle: Analyti- cal Workload Cost and Performance Stability With Elastic Pools. Proceedings of the ACM on Management of Data 1(4), 233–123325 (2023)
2023
-
[58]
Proceedings of the ACM on Man- agement of Data 1(2), 161–116127 (2023)
Bian, H., Sha, T., Ailamaki, A.: Using Cloud Functions as Accelerator for Elastic Data Analytics. Proceedings of the ACM on Man- agement of Data 1(2), 161–116127 (2023)
2023
-
[59]
In: ACM EuroSys, pp
Khandelwal, A., Tang, Y., Agarwal, R., Akella, A., Stoica, I.: Jiffy: Elastic far- memory for stateful serverless analytics. In: ACM EuroSys, pp. 697–713 (2022)
2022
-
[60]
In: USENIX OSDI, pp
Klimovic, A., Wang, Y., Stuedi, P., Trivedi, A., Pfefferle, J., Kozyrakis, C.: Pocket: Elas- tic ephemeral storage for serverless analytics. In: USENIX OSDI, pp. 427–444 (2018)
2018
-
[61]
In: CIDR (2021)
Wawrzoniak, M., M¨ uller, I., Alonso, G., Bruno, R.: Boxer: Data Analytics on Network-enabled Serverless Platforms. In: CIDR (2021)
2021
- [62]
-
[63]
In: ACM EuroSys, pp
Sharma, P., Guo, T., He, X., Irwin, D.E., Shenoy, P.J.: Flint: Batch-Interactive Data- Intensive Processing on Transient Servers. In: ACM EuroSys, pp. 6–1615 (2016). https://doi.org/10.1145/2901318.2901319 . https://doi.org/10.1145/2901318.2901319
2016
-
[64]
Foundations and Trends in Databases 10(1), 1–107 (2021)
Narasayya, V.R., Chaudhuri, S.: Cloud Data Services: Workloads, Architectures and Multi-Tenancy. Foundations and Trends in Databases 10(1), 1–107 (2021)
2021
-
[65]
In: ACM SIGMOD, pp
Dageville, B., Cruanes, T., Zukowski, M., Antonov, V., Avanes, A., Bock, J., Clay- baugh, J., Engovatov, D., Hentschel, M., Huang, J., Lee, A.W., Motivala, A., Munir, A.Q., Pelley, S., Povinec, P., Rahn, G., Triantafyllis, S., Unterbrunner, P.: The Snowflake Elastic Data Wareh...
2016
-
[66]
In: CIDR (2021) 13
Zaharia, M., Ghodsi, A., Xin, R., Armbrust, M.: Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics. In: CIDR (2021) 13
2021
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.