REVIEW 3 major objections 4 minor 112 references
Approximation-First Timeseries Monitoring Query At Scale
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read PromSketch caches overlapping window aggregations in compact sketches, cutting rule-query latency by up to two orders of magnitude while keeping mean errors at 5% or below.
desk verdict Real system, real latency wins, but the 400x cost claim prices the measured ingestion overhead at zero. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Exponential Histogram used as a cache index: a sequence of non-overlapping buckets whose sizes grow exponentially as they age, each bucket holding a sketch. Because EH buckets are mergeable (the target statistics are weakly additive), a query over any sub-window $(t_1,t_2)$ is answered by merging the buckets strictly between the two boundary buckets, with the discarded boundary buckets contributing controlled rank or norm error. The paper pairs EH with KLL sketches for quantiles (EHKLL), with universal sketches for GSum statistics (EHUniv), and with uniform sampling for sum/avg/stddev, and adds a hybrid optimization in which small buckets keep exact item-frequency maps before converting to sketches. This machinery is what lets the cache answer many overlapping and drill-down windows from one compact in-memory structure.
What would settle it
Run the Table 1 1000-node, 268-billion-sample workload on the released PromSketch code and bill the extra ingestion CPU (the measured 1.3x–3x insertion slowdown) at the same per-vCPU-hour rate the paper applies to query processing; if the resulting total monthly cost is not roughly two orders of magnitude below Prometheus's, the central cost claim is falsified.
Extended reading notes
Core claim
The central discovery is that repeated data scans and repeated query computations over overlapping windows are the two bottlenecks in rule-query engines, and both can be bypassed by caching intermediate results rather than raw samples or final answers. PromSketch maintains, per timeseries, an Exponential Histogram whose buckets hold composable sketches—KLL sketches for quantiles (and min/max), universal sketches for entropy, distinct count, L2 norm, and top-k, and a sliding-window uniform sample for average, sum, and variance. Any sub-window of the recent window is answered by merging the buckets that cover it, so a 5-, 10-, or 15-minute drill-down can be served from the same cached structure without rescanning storage. The paper proves rank-error bounds for the EH+KLL combination (final normalized rank error at most $2\epsilon_{EH}+\epsilon_{KLL}$) and memory bounds for the EH+universal-sketch combination, and it reports end-to-end latency and cost results on synthetic and real traces.
Load-bearing premise
The cost-reduction numbers assume that the ingestion-time sketching work, which the paper itself measures as a 1.3x–3x insertion-throughput slowdown, is operationally free: Table 1 charges PromSketch only for query processing while keeping storage and data-ingestion costs at the uncached baseline.
Editorial extensions
If this is right
- Rule queries with overlapping windows, such as 10-minute windows evaluated every minute, can be served from precomputed bucket merges, so alerting and recording rules no longer rescan storage on every evaluation.
- Drill-down queries that zoom from a 1M-sample window into 100K and 10K sub-windows keep mean errors under 5% at a fraction of the memory of exact storage, which supports anomaly localization on the same cache.
- Operators can trade memory for accuracy explicitly through knobs such as $k_{EH}$ and $k_{KLL}$, since error decreases as memory grows for all evaluated statistics.
- Because the cache is a standalone module integrated through a small query-parser patch, the same design transfers to any PromQL-like engine, including distributed deployments where nodes shard timeseries by consistent hashing.
Reading between the lines
- Beyond the paper, the cost comparison depends heavily on the query-to-ingestion ratio: PromSketch's advantage should grow with more frequent rule evaluations and shrink for workloads that mostly ingest data and rarely query, since the ingestion-path overhead remains in either case.
- Beyond the paper, the same EH-as-cache design could be applied to label-dimensional aggregation or to metrics with string values beyond IP addresses, since universal sketches operate on any item universe, though the paper leaves that extension unproven.
- Beyond the paper, because small EH buckets use exact maps, a natural adaptive policy would set bucket sizes so that typical drill-down queries touch only exact-map buckets, potentially pushing mean error far below 5% for common workloads.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PromSketch, an approximate in-memory intermediate-result cache for Prometheus-compatible timeseries monitoring systems. The system maintains Exponential Histogram (EH) buckets over recent time windows, using KLL sketches per bucket for quantile queries, universal sketches with an exact-map hybrid for GSum statistics (entropy, distinct count, L2 norm, top-K), and uniform sampling for average/stddev queries. At query time it answers overlapping sub-window aggregations by merging bucket sketches instead of rescanning raw samples. The authors provide error-bound analyses for EHKLL and EHUniv, describe single-machine and distributed integrations into Prometheus and VictoriaMetrics, and report experiments on synthetic and real traces claiming mean errors at or below 5%, latency reductions of up to about 200x, and query-processing cost reductions of 400x versus Prometheus and at least 4x versus VictoriaMetrics. They also report ingestion-throughput slowdowns of 1.3-3.1x over Prometheus and 1.4-7.2x over VictoriaMetrics in §6.1.3.
Significance. If the cost accounting is corrected and the sub-window error bound is fixed, this is a substantial contribution: it is, to my knowledge, the first end-to-end approximate intermediate-cache design for overlapping time-window rule queries in Prometheus-like systems; it combines a general window framework with mergeable sketches and includes provable memory-accuracy statements; and it ships a public Go implementation and uses public datasets. The latency and error measurements in Table 4 and Figures 6-13 are systematic, repeated, and reported against two real baselines, which is a genuine strength. The two issues below, however, affect the two headline claims (cost reduction and guaranteed sub-window accuracy), so the paper needs revision before the claims can be accepted as stated.
major comments (3)
- [§6.1.1, Table 1; §6.1.3, Fig. 7] The headline cost reductions of 400x and at least 4x are computed with an incomplete cost model. Table 1 keeps PromSketch's ingestion costs at the baseline ($9,186 for PS-PM and 'incl.' for PS-VM) while moving query processing to $28.6/month, and §6.1.1 states that PromSketch reduces query-processing cost 'while not increasing the storage and data ingestion costs.' But §6.1.3 and Fig. 7 measure a 1.3-3.1x insertion-throughput slowdown over Prometheus and a 1.4-7.2x slowdown over VictoriaMetrics, so the precomputation consumes real CPU on the ingestion path. Under the resource-based billing model used for VictoriaMetrics in the same table, those cycles are billable; under a per-sample model, reprocessing the 268B samples at the paper's own $0.1/B-sample query rate would add roughly $26.8/month to the PS-PM row. Please recompute the cost comparison with the ingestion-path precomputation priced (e.g., vCPU-hour for PS-VM and per-sample for PS-PM), report both query-path and total operational costs, and state explicitly whether the 'two orders of magnitude' claim refers only to query-path spend.
- [§4.2.1, EHKLL Error Guarantee] The sub-window rank-error analysis for EHKLL appears incorrect. The text says the total rank difference from the EH sub-window query is at most C_i - C_j, where B_i is the discarded bucket containing t1 and B_j is the included bucket containing t2. Each of these buckets contributes a nonnegative error: excluding part of B_i misses up to C_i items, and including all of B_j adds up to C_j items after t2. The combined rank error should therefore be bounded by a sum of the two contributions (or by max(C_i, C_j) in the best case), not by their difference. As written, C_i - C_j is non-positive under Invariant 2 (C_1 <= ... <= C_l), so the subsequent upper bound of 2*epsilon_EH * N_t1/(N_t1 - N_t2) + epsilon_KLL does not follow. Please rederive this bound with absolute values or explicit worst-case analysis, and state the assumptions needed for the claimed sub-window guarantee.
- [§4.2.1, Algorithm 1 and query semantics] The sub-window query procedure in Algorithm 1 includes the whole bucket B_j when t2 falls strictly inside B_j, which adds the suffix of B_j after t2 to the approximate window. The error analysis should explicitly account for this suffix, since the stated guarantee is for the window (t1, t2). The current discussion mentions that B_j contributes at most C_j, but the formula then uses C_i - C_j, and the relationship between the approximate window [end(B_i), end(B_j)] and the true window [t1, t2] is never made precise. A precise statement of which samples are included and excluded would also resolve the sign issue in the preceding comment.
minor comments (4)
- [§2.3 and §6.1.1] The AWS Prometheus cost figures are internally inconsistent: the Introduction gives $11,520 for query processing and $9,256 for ingestion, while §6.1.1 gives $11,560 and $9,186, and Table 1 gives $11,520 and $9,186. Please harmonize these numbers and show the arithmetic, since 10 queries/minute each processing 8B samples per month does not obviously produce $11,520 at the stated $0.1/B-samples rate without additional assumptions.
- [Abstract and Table 3] The claim that PromSketch covers '70% of Prometheus' aggregation over time queries' is not supported by any query-corpus analysis. Please state the query corpus, the weighting of queries, and the counting methodology that yields the 70% figure.
- [§6.1.2 and Table 4] The body text says the drill-down queries cover '10^5-, 10^4-, and 10^3-second time windows' while the Table 4 caption says '10K-, 100K-, and 1M-sample windows.' Please clarify whether the axes are time durations or sample counts, and use consistent units throughout.
- [Throughout] There are several typographical errors and inconsistent spellings, e.g., 'remoevd' in §4.3, 'Kubernates' in the introduction and §4.4, and 'Google Promethues' in reference [101]. A thorough proofreading pass is needed.
Circularity Check
No significant circularity: the latency/accuracy claims come from independent experiments, and the cost-model caveat about free precomputation is an accounting limitation rather than a self-referential derivation.
full rationale
The paper's central latency, throughput, and accuracy results are measured end-to-end against Prometheus and VictoriaMetrics on public and synthetic datasets, so they are not generated by the paper's own prior claims. EHKLL's error bound is derived from Exponential Histogram invariants [60] and KLL's rank-error guarantee [72], both external; EHUniv's construction and guarantees are cited to [69], which shares an author with the present paper, but [69] is a published theorem with assumptions that do not include the present evaluation targets, so under the stated rules it is independent support rather than a circular self-citation. No fitted constants are renamed as predictions: the 5% accuracy target is an explicit configuration objective (k_EH, k_KLL), and the reported errors are measured, not assumed. The one substantive caveat is the operational-cost model: §6.1.1 claims PromSketch reduces query-processing cost 'while not increasing the storage and data ingestion costs,' yet §6.1.3/Fig. 7 measures a 1.3–3.1x insertion-throughput slowdown for the Prometheus integration and 4.1–4.5x for EHKLL/EHUniv on VictoriaMetrics, meaning sketch precomputation consumes real CPU that Table 1 does not price. That is a limitation in the cost accounting (the 400x and 4x ratios are arithmetic consequences of the chosen billing assumptions), but it is not circular: the cost estimate is an explicit model-based calculation, not a prediction that reduces to its own inputs, and it does not contaminate the independent latency/accuracy measurements.
Assumptions & free parameters
free parameters (4)
- EH bucket parameter k_EH (EHKLL) =
50
- KLL space parameter k_KLL =
256
- EH bucket parameter k_EH (EHUniv) =
20
- Uniform sampling rate p =
10%
assumptions (6)
- standard math Target functions are weakly additive or admit composable sketches (Eq. 1)
- standard math Universal sketching GSum estimation theorem (Theorem 2 in [52])
- standard math KLL sketches are mergeable with rank error guarantees
- domain assumption Monitoring rule queries tolerate approximate results
- domain assumption Out-of-order and duplicate samples are rejected upstream
- domain assumption Rule query workload is dominated by periodic overlapping windows
Cite this review
Pith. "Pith review of Approximation-First Timeseries Monitoring Query At Scale." pith.science (2026). https://pith.science/paper/PIO2RF3V
@misc{pith2026250510560,
author = {Pith},
title = {Pith review of: Approximation-First Timeseries Monitoring Query At Scale},
year = {2026},
howpublished = {\url{https://pith.science/paper/PIO2RF3V}},
note = {Machine review of arXiv:2505.10560}
}
read the original abstract
Timeseries monitoring systems such as Prometheus play a crucial role in gaining observability of the underlying system components. These systems collect timeseries metrics from various system components and perform monitoring queries over periodic window-based aggregations (i.e., rule queries). However, despite wide adoption, the operational costs and query latency of rule queries remain high. In this paper, we identify major bottlenecks associated with repeated data scans and query computations concerning window overlaps in rule queries, and present PromSketch, an approximation-first query framework as intermediate caches for monitoring systems. It enables low operational costs and query latency, by combining approximate window-based query frameworks and sketch-based precomputation. PromSketch is implemented as a standalone module that can be integrated into Prometheus and VictoriaMetrics, covering 70% of Prometheus' aggregation over time queries. Our evaluation shows that PromSketch achieves up to a two orders of magnitude reduction in query latency over Prometheus and VictoriaMetrics, while lowering operational dollar costs of query processing by two orders of magnitude compared to Prometheus and by at least 4x compared to VictoriaMetrics with at most 5% average errors across statistics. The source code has been made available at https://github.com/Froot-NetSys/promsketch.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[69]
Nikita Ivkin, Ran Ben Basat, Zaoxing Liu, Gil Einziger, Roy Friedman, and Vladimir Braverman. 2019. I know what you did last summer: Network mon- itoring using interval queries. Proceedings of the ACM on Measurement and Analysis of Computing Systems 3, 3 (2019), 1–28
work page 2019
-
[1]
Prometheus: Monitoring at SoundCloud
2015. Prometheus: Monitoring at SoundCloud. https://github.com/pingcap/ docs/blob/master/tidb-monitoring-framework.md
2015
-
[2]
Packets-per-second limits in EC2
2019. Packets-per-second limits in EC2. Retrieved 2024 from https://stressgrid. com/blog/pps_limits_in_ec2/
2019
-
[3]
Google ClusterData 2019 traces
2020. Google ClusterData 2019 traces. Retrieved 2024 from https://github.com/ google/cluster-data/blob/master/ClusterData2019.md
2020
-
[4]
Amazon EC2 On-Demand Pricing
2022. Amazon EC2 On-Demand Pricing. Retrieved 2024 from https://aws. amazon.com/ec2/pricing/on-demand/
2022
-
[5]
Kubernetes
2022. Kubernetes. https://kubernetes.io/
2022
-
[6]
Prometheus handling out-of-order samples
2022. Prometheus handling out-of-order samples. Retrieved 2024 from https://promlabs.com/blog/2022/12/15/understanding-duplicate-samples- and-out-of-order-timestamp-errors-in-prometheus/
2022
-
[7]
VictoriaMetrics
2022. VictoriaMetrics. https://victoriametrics.com
2022
Show all 112 references
-
[8]
PromCon 2023 - Yet Another Streaming PromQL Engine
2023. PromCon 2023 - Yet Another Streaming PromQL Engine. Retrieved 2024 from https://www.youtube.com/watch?v=3kM2Asj6hcg
2023
-
[9]
Prometheus Metrics based autoscaling in Kubernetes
2023. Prometheus Metrics based autoscaling in Kubernetes. Retrieved 2024 from https://gcore.com/learning/prometheus-metrics-based-autoscaling-in- kubernetes/
2023
-
[10]
Amazon Managed Service for Prometheus pricing
2024. Amazon Managed Service for Prometheus pricing. Retrieved 2024 from https://aws.amazon.com/prometheus/pricing/
2024
-
[11]
Awesome Prometheus alerts
2024. Awesome Prometheus alerts. https://samber.github.io/awesome- prometheus-alerts/
2024
-
[12]
Cloudflare Blog - Monitoring our monitoring: how we validate our Prometheus alert rules
2024. Cloudflare Blog - Monitoring our monitoring: how we validate our Prometheus alert rules. https://blog.cloudflare.com/monitoring-our- monitoring/
2024
-
[13]
Cluster VictoriaMetrics
2024. Cluster VictoriaMetrics. Retrieved 2024 from https://docs.victoriametrics. com/cluster-victoriametrics/
2024
-
[14]
DataDog Anomaly Detection
2024. DataDog Anomaly Detection. https://docs.datadoghq.com/monitors/ types/anomaly/
2024
-
[15]
Fastcache used by VictoriaMetrics
2024. Fastcache used by VictoriaMetrics. https://github.com/VictoriaMetrics/ fastcache
2024
-
[16]
Google Cloud scales based on Monitoring metrics
2024. Google Cloud scales based on Monitoring metrics. https://cloud.google. com/compute/docs/autoscaler/scaling-cloud-monitoring-metrics
2024
-
[17]
Google Kubernetes Engine Quotas and Limits
2024. Google Kubernetes Engine Quotas and Limits. https://cloud.google.com/ kubernetes-engine/quotas
2024
-
[18]
Grafana Dashboards
2024. Grafana Dashboards. https://grafana.com/grafana/dashboards/
2024
-
[19]
Grafana Mimir
2024. Grafana Mimir. https://grafana.com/docs/mimir/latest/
2024
-
[20]
Grafana Mimir uses Redis or Memcached as chunks-cache, index-cache, results-cache and metadata-cache
2024. Grafana Mimir uses Redis or Memcached as chunks-cache, index-cache, results-cache and metadata-cache. https://grafana.com/docs/helm-charts/ mimir-distributed/latest/configure/configure-redis-cache/
2024
-
[21]
InfluxDB
2024. InfluxDB. https://www.influxdata.com/
2024
-
[22]
Kubernetes monitoring with Prometheus
2024. Kubernetes monitoring with Prometheus. Retrieved 2024 from https://prometheus.io/docs/prometheus/latest/configuration/configuration/ #kubernetes_sd_config
2024
-
[23]
Memcached
2024. Memcached. https://memcached.org
2024
-
[24]
Monitoring Juniper Networks with Prometheus
2024. Monitoring Juniper Networks with Prometheus. Retrieved 2024 from https://github.com/czerwonk/junos_exporter
2024
-
[25]
Open5GS Metrics with Prometheus
2024. Open5GS Metrics with Prometheus. Retrieved 2024 from https://open5gs. org/open5gs/docs/tutorial/04-metrics-prometheus/
2024
-
[26]
Prometheus Configurations
2024. Prometheus Configurations. Retrieved 2024 from https://prometheus.io/ docs/prometheus/latest/configuration/configuration/
2024
-
[27]
Prometheus functions
2024. Prometheus functions. Retrieved 2024 from https://prometheus.io/docs/ prometheus/latest/querying/functions/
2024
-
[28]
Prometheus Query Language
2024. Prometheus Query Language. Retrieved 2024 from https://prometheus. io/docs/prometheus/latest/querying/basics/
2024
-
[29]
Prometheus SNMP exporter
2024. Prometheus SNMP exporter. Retrieved 2024 from https://github.com/ prometheus/snmp_exporter
2024
-
[30]
2024. Redis. https://redis.io
2024
-
[31]
2024. Thanos. https://thanos.io/
2024
-
[32]
Thanos Downsampling, Resolution and Retention
2024. Thanos Downsampling, Resolution and Retention. Retrieved 2024 from https://thanos.io/v0.8/components/compact/
2024
-
[33]
The CAIDA UCSD Anonymized Internet Traces
2024. The CAIDA UCSD Anonymized Internet Traces. https://www.caida.org/ catalog/datasets/passive_dataset/. Online
2024
-
[34]
VictoriaMetrics Anomaly Detection
2024. VictoriaMetrics Anomaly Detection. Retrieved 2024 from https://victoriametrics.com/blog/victoriametrics-anomaly-detection- handbook-chapter-2/index.html
2024
-
[35]
VictoriaMetrics backfilling support for out-of-order samples
2024. VictoriaMetrics backfilling support for out-of-order samples. Retrieved 2025 from https://docs.victoriametrics.com/#backfilling
2024
-
[36]
VictoriaMetrics Deduplication
2024. VictoriaMetrics Deduplication. Retrieved 2024 from https://docs. victoriametrics.com/#deduplication
2024
-
[37]
VictoriaMetrics parallel query in vm-select
2024. VictoriaMetrics parallel query in vm-select. https://github.com/ VictoriaMetrics/VictoriaMetrics/issues/2886
2024
-
[38]
VictoriaMetrics Pricing compared to Prometheus
2024. VictoriaMetrics Pricing compared to Prometheus. Retrieved 2024 from https://victoriametrics.com/blog/managed-prometheus-pricing/
2024
-
[39]
VictoriaMetrics rollup functions
2024. VictoriaMetrics rollup functions. Retrieved 2024 from https://docs. victoriametrics.com/metricsql/#rollup-functions
2024
-
[40]
VictoriaMetrics Single Version
2024. VictoriaMetrics Single Version. Retrieved 2024 from https://docs. victoriametrics.com/single-server-victoriametrics/
2024
-
[41]
Lior Abraham, John Allen, Oleksandr Barykin, Vinayak Borkar, Bhuwan Chopra, Ciprian Gerea, Daniel Merl, Josh Metzler, David Reiss, Subbu Subramanian, et al
-
[42]
Swarup Acharya, Phillip B Gibbons, Viswanath Poosala, and Sridhar Ra- maswamy. 1999. The aqua approximate query answering system. InProceedings of the 1999 ACM SIGMOD international conference on Management of data . 574– 576
1999
-
[43]
Sameer Agarwal, Barzan Mozafari, Aurojit Panda, Henry Milner, Samuel Mad- den, and Ion Stoica. 2013. BlinkDB: queries with bounded errors and bounded response times on very large data. In Proceedings of the 8th ACM European conference on computer systems . 29–42
2013
-
[44]
Nitin Agrawal and Ashish Vulimiri. 2017. Low-latency analytics on colossal data streams with summarystore. In Proceedings of the 26th Symposium on Operating Systems Principles. 647–664
2017
-
[45]
Manos Antonakakis, Tim April, Michael Bailey, Matt Bernhard, Elie Bursztein, Jaime Cochran, Zakir Durumeric, J Alex Halderman, Luca Invernizzi, Michalis Kallitsis, et al. 2017. Understanding the mirai botnet. In 26th USENIX security symposium (USENIX Security 17) . 1093–1110
2017
-
[46]
Arvind Arasu and Gurmeet Singh Manku. 2004. Approximate counts and quantiles over sliding windows. InProceedings of the twenty-third ACM SIGMOD- SIGACT-SIGART symposium on Principles of database systems . 286–296
2004
-
[47]
Ran Ben Basat, Gil Einziger, Isaac Keslassy, Ariel Orda, Shay Vargaftik, and Erez Waisbard. 2018. Memento: Making sliding windows efficient for heavy hitters. In Proceedings of the 14th International Conference on Emerging Networking EXperiments and Technologies. 254–266
2018
-
[48]
Ran Ben-Basat, Gil Einziger, Roy Friedman, and Yaron Kassner. 2016. Heavy hitters in streams and sliding windows. InIEEE INFOCOM 2016-The 35th Annual IEEE International Conference on Computer Communications . IEEE, 1–9
2016
-
[49]
Ran Ben Basat, Gil Einziger, Roy Friedman, and Yaron Kassner. 2019. Succinct summing over sliding windows. Algorithmica 81 (2019), 2072–2091
2019
-
[50]
Vance W Berger and YanYan Zhou. 2014. Kolmogorov–smirnov test: Overview. Wiley statsref: Statistics reference online (2014). Zeying Zhu, Jonathan Chamberlain, Kenny Wu, David Starobinski, and Zaoxing Liu
2014
-
[51]
Burton H Bloom. 1970. Space/time trade-offs in hash coding with allowable errors. Commun. ACM 13, 7 (1970), 422–426
1970
-
[52]
Vladimir Braverman, Stephen R Chestnut, David P Woodruff, and Lin F Yang
-
[53]
Vladimir Braverman and Rafail Ostrovsky. 2007. Smooth histograms for sliding windows. In 48th Annual IEEE Symposium on Foundations of Computer Science (FOCS’07). IEEE, 283–293
2007
-
[54]
Vladimir Braverman and Rafail Ostrovsky. 2010. Zero-one frequency laws. In Proceedings of the forty-second ACM symposium on Theory of computing . 281–290
2010
-
[55]
Yousra Chabchoub and Georges Heébrail. 2010. Sliding hyperloglog: Estimating cardinality in a data stream over a sliding window. In 2010 IEEE International Conference on Data Mining Workshops. IEEE, 1297–1303
2010
-
[56]
Shiyu Chang, Yang Zhang, Jiliang Tang, Dawei Yin, Yi Chang, Mark A Hasegawa- Johnson, and Thomas S Huang. 2017. Streaming recommender systems. In Proceedings of the 26th international conference on world wide web . 381–389
2017
-
[57]
Moses Charikar, Kevin Chen, and Martin Farach-Colton. 2002. Finding frequent items in data streams. In International Colloquium on Automata, Languages, and Programming. Springer, 693–703
2002
-
[58]
Graham Cormode and Shan Muthukrishnan. 2005. An improved data stream summary: the count-min sketch and its applications. Journal of Algorithms 55, 1 (2005), 58–75
2005
-
[59]
Chuck Cranor, Theodore Johnson, Oliver Spataschek, and Vladislav Shkapenyuk
-
[60]
Mayur Datar, Aristides Gionis, Piotr Indyk, and Rajeev Motwani. 2002. Main- taining stream statistics over sliding windows. SIAM journal on computing 31, 6 (2002), 1794–1813
2002
-
[61]
Ronen Ben David and Anat Bremler Barr. 2021. Kubernetes autoscaling: Yoyo attack vulnerability and mitigation. arXiv preprint arXiv:2105.00542 (2021)
2021 arXiv
-
[62]
Cristian Estan and George Varghese. 2002. New directions in traffic measure- ment and accounting. In Proceedings of the 2002 conference on Applications, technologies, architectures, and protocols for computer communications . 323–336
2002
-
[63]
Edward Gan, Peter Bailis, and Moses Charikar. 2020. Coopstore: Optimizing precomputed summaries for aggregation. Proceedings of the VLDB Endowment 13, 12 (2020), 2174–2187
2020
-
[64]
Xiangyang Gou, Long He, Yinda Zhang, Ke Wang, Xilai Liu, Tong Yang, Yi Wang, and Bin Cui. 2020. Sliding sketches: A framework using time zones for data stream processing in sliding windows. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery &...
2020
-
[65]
Michael Greenwald and Sanjeev Khanna. 2001. Space-efficient online computa- tion of quantile summaries. ACM SIGMOD Record 30, 2 (2001), 58–66
2001
-
[66]
Georges Hebrail and Alice Berard. 2012. Individual Household Elec- tric Power Consumption. UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C58K54
2012 doi
-
[67]
Ching-Tien Ho, Rakesh Agrawal, Nimrod Megiddo, and Ramakrishnan Srikant
-
[68]
Peng Huang, Chuanxiong Guo, Lidong Zhou, Jacob R Lorch, Yingnong Dang, Murali Chintalapati, and Randolph Yao. 2017. Gray failure: The achilles’ heel of cloud-scale systems. In Proceedings of the 16th Workshop on Hot Topics in Operating Systems. 150–155
2017
-
[70]
Uwe Jugel, Zbigniew Jerzak, Gregor Hackenbroich, and Volker Markl. 2014. M4: a visualization-oriented time series data aggregation. Proceedings of the VLDB Endowment 7, 10 (2014), 797–808
2014
-
[71]
David Karger, Eric Lehman, Tom Leighton, Rina Panigrahy, Matthew Levine, and Daniel Lewin. 1997. Consistent hashing and random trees: Distributed caching protocols for relieving hot spots on the world wide web. In Proceedings of the twenty-ninth annual ACM symposium on Theory ...
1997
-
[72]
Zohar Karnin, Kevin Lang, and Edo Liberty. 2016. Optimal quantile approxima- tion in streams. In 2016 ieee 57th annual symposium on foundations of computer science (focs). IEEE, 71–78
2016
-
[73]
Yehuda Koren. 2009. Collaborative filtering with temporal dynamics. InProceed- ings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining. 447–456
2009
-
[74]
Ashwin Lall, Vyas Sekar, Mitsunori Ogihara, Jun Xu, and Hui Zhang. 2006. Data streaming algorithms for estimating entropy of network traffic. ACM SIGMETRICS Performance Evaluation Review 34, 1 (2006), 145–156
2006
-
[75]
Xi Liang, Stavros Sintos, Zechao Shang, and Sanjay Krishnan. 2021. Combining aggregation and sampling (nearly) optimally for approximate query processing. In Proceedings of the 2021 International Conference on Management of Data . 1129–1141
2021
- [76]
-
[77]
Gangmuk Lim, Mohamed S Hassan, Ze Jin, Stavros Volos, and Myeongjae Jeon
-
[78]
Zaoxing Liu, Antonis Manousis, Gregory Vorsanger, Vyas Sekar, and Vladimir Braverman. 2016. One sketch to rule them all: Rethinking network flow mon- itoring with univmon. In Proceedings of the 2016 ACM SIGCOMM Conference . 101–114
2016
-
[79]
Zaoxing Liu, Hun Namkung, Georgios Nikolaidis, Jeongkeun Lee, Changhoon Kim, Xin Jin, Vladimir Braverman, Minlan Yu, and Vyas Sekar. 2021. Jaqen: A High-Performance Switch-Native Approach for Detecting and Mitigating Volumetric DDoS Attacks with Programmable Switches. In30th U...
2021
-
[80]
Xuesong Lu, Wee Hyong Tok, Chedy Raissi, and Stéphane Bressan. 2010. A simple, yet effective and efficient, sliding window sampling algorithm. In Data- base Systems for Advanced Applications: 15th International Conference, DASFAA 2010, Tsukuba, Japan, April 1-4, 2010, Proceedi...
2010
-
[81]
Samuel Madden, Michael J Franklin, Joseph M Hellerstein, and Wei Hong
-
[82]
Antonis Manousis, Zhuo Cheng, Ran Ben Basat, Zaoxing Liu, and Vyas Sekar
-
[83]
Stavros Maroulis, Vassilis Stamatopoulos, George Papastefanatos, and Manolis Terrovitis. 2024. Visualization-aware Time Series Min-Max Caching with Error Bound Guarantees. Proceedings of the VLDB Endowment 17, 8 (2024), 2091–2103
2024
-
[84]
Octavian Mart, Catalin Negru, Florin Pop, and Aniello Castiglione. 2020. Observ- ability in kubernetes cluster: Automatic anomalies detection using prometheus. In 2020 IEEE 22nd International Conference on High Performance Computing and Communications; IEEE 18th International ...
2020
-
[85]
Artur Marzano, David Alexander, Osvaldo Fonseca, Elverton Fazzion, Cristine Hoepers, Klaus Steding-Jessen, Marcelo HPC Chaves, Ítalo Cunha, Dorgival Guedes, and Wagner Meira. 2018. The evolution of bashlite and mirai iot botnets. In 2018 IEEE Symposium on Computers and Communi...
2018
-
[86]
Charles Masson, Jee E Rim, and Homin K Lee. 2019. DDSketch: A fast and fully-mergeable quantile sketch with relative-error guarantees. arXiv preprint arXiv:1908.10693 (2019)
2019 arXiv
-
[87]
In Proceedings of the 2003 ACM SIGMOD international conference on Management of data
The design of an acquisitional query processor for sensor networks. In Proceedings of the 2003 ACM SIGMOD international conference on Management of data. 491–502
2003
-
[88]
Muhammad Aashiq Moosa, Apurva K Vangujar, and Dnyanesh Pramod Ma- hajan. 2023. Detection and Analysis of DDoS Attack Using a Collaborative Network Monitoring Stack. In 2023 16th International Conference on Security of Information and Networks (SIN) . IEEE, 1–9
2023
-
[89]
Ohsita, S
Y. Ohsita, S. Ata, and M. Murata. 2004. Detecting distributed denial-of- service attacks by analyzing TCP SYN packets statistically. In IEEE Global Telecommunications Conference, 2004. GLOBECOM ’04. , Vol. 4. 2043–2049 Vol.4. https://doi.org/10.1109/GLOCOM.2004.1378371
2004 arXiv
-
[90]
Alessandro Vittorio Papadopoulos, Ahmed Ali-Eldin, Karl-Erik Årzén, Johan Tordsson, and Erik Elmroth. 2016. PEAS: A performance evaluation framework for auto-scaling strategies in cloud applications. ACM Transactions on Modeling and Performance Evaluation of Computing Systems ...
2016
-
[91]
Odysseas Papapetrou, Minos Garofalakis, and Antonios Deligiannakis. 2012. Sketch-based querying of distributed sliding-window data streams. arXiv preprint arXiv:1207.0139 (2012)
2012 arXiv
-
[92]
Yongjoo Park, Barzan Mozafari, Joseph Sorenson, and Junhao Wang. 2018. Verdictdb: Universalizing approximate query processing. In Proceedings of the 2018 International Conference on Management of Data . 1461–1476
2018
-
[93]
Tuomas Pelkonen, Scott Franklin, Justin Teller, Paul Cavallaro, Qi Huang, Justin Meza, and Kaushik Veeraraghavan. 2015. Gorilla: A fast, scalable, in-memory time series database. Proceedings of the VLDB Endowment 8, 12 (2015), 1816– 1827
2015
-
[94]
Michael Mitzenmacher, Thomas Steinke, and Justin Thaler. 2012. Hierarchical heavy hitters with the space saving algorithm. In 2012 Proceedings of the Four- teenth Workshop on Algorithm Engineering and Experiments (ALENEX) . SIAM, 160–174
2012
-
[95]
James Pope, Francesco Raimondo, Vijay Kumar, Ryan McConville, Rob Piechocki, George Oikonomou, Thomas Pasquier, Bo Luo, Dan Howarth, Ioan- nis Mavromatis, et al. 2021. Container escape detection for edge devices. In Approximation-First Timeseries Monitoring Query At Scale Proc...
2021
-
[96]
Athanasios Priovolos, Dimitris Lioprasitis, Georgios Gardikis, and Socrates Costicoglou. 2021. Using anomaly detection techniques for securing 5G infras- tructure and applications. In 2021 IEEE International Mediterranean Conference on Communications and Networking (MeditCom) ...
2021
-
[97]
Chunhui Shen, Qianyu Ouyang, Feibo Li, Zhipeng Liu, Longcheng Zhu, Yujie Zou, Qing Su, Tianhuan Yu, Yi Yi, Jianhong Hu, et al . 2023. Lindorm TSDB: A Cloud-Native Time-Series Database for Large-Scale Monitoring Systems. Proceedings of the VLDB Endowment 16, 12 (2023), 3715–3727
2023
-
[98]
Mor Sides, Anat Bremler-Barr, and Elisha Rosensweig. 2015. Yo-Yo Attack: vul- nerability in auto-scaling mechanism. InProceedings of the 2015 ACM Conference on Special Interest Group on Data Communication . 103–104
2015
-
[99]
James Turnbull. 2018. Monitoring with Prometheus. Turnbull Press
2018
-
[100]
Zhiqi Wang, Jin Xue, and Zili Shao. 2021. Heracles: an efficient storage model and data flushing for performance monitoring timeseries. Proceedings of the VLDB Endowment 14, 6 (2021), 1080–1092
2021
-
[101]
Jinglin Peng, Dongxiang Zhang, Jiannan Wang, and Jian Pei. 2018. Aqp++ connecting approximate query processing with aggregate precomputation for interactive analytics. In Proceedings of the 2018 International Conference on Management of Data. 1477–1492
2018
-
[102]
Yuhan Wu, Shiqi Jiang, Siyuan Dong, Zheng Zhong, Jiale Chen, Yutong Hu, Tong Yang, Steve Uhlig, and Bin Cui. 2023. MicroscopeSketch: Accurate Sliding Estimation Using Adaptive Zooming. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 2660–2671
2023
-
[103]
future-proof
Mingran Yang, Junbo Zhang, Akshay Gadre, Zaoxing Liu, Swarun Kumar, and Vyas Sekar. 2020. Joltik: enabling energy-efficient" future-proof" analytics on low-power wide-area networks. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking . 1–14
2020
-
[104]
Tong Yang, Junzhi Gong, Haowei Zhang, Lei Zou, Lei Shi, and Xiaoming Li. 2018. Heavyguardian: Separate and guard hot items in data streams. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2584–2593
2018
-
[105]
Bohan Zhao, Xiang Li, Boyu Tian, Zhiyu Mei, and Wenfei Wu. 2021. Dhs: Adaptive memory layout organization of sketch slots for fast and accurate data stream processing. In Proceedings of ACM SIGKDD . 2285–2293
2021
-
[108]
Jamie Wilkinson. 2016. Google Promethues: A practical guide to alerting at scale. Retrieved 2024 from https://docs.google.com/presentation/d/ 1X1rKozAUuF2MVc1YXElFWq9wkcWv3Axdldl8LOH9Vik/edit#slide=id. g598ef96a6_0_341
2016
-
[1997]
ACM SIGMOD Record 26, 2 (1997), 73–88
Range queries in OLAP data cubes. ACM SIGMOD Record 26, 2 (1997), 73–88
1997
-
[2003]
In Proceedings of the 2003 ACM SIGMOD international conference on Management of data
Gigascope: A stream database for network applications. In Proceedings of the 2003 ACM SIGMOD international conference on Management of data . 647– 651
2003
-
[2013]
Proceedings of the VLDB Endowment 6, 11 (2013), 1057–1067
Scuba: Diving into data at facebook. Proceedings of the VLDB Endowment 6, 11 (2013), 1057–1067
2013
-
[2016]
In Proceedings of the 35th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems
Streaming space complexity of nearly all functions of one variable on frequency vectors. In Proceedings of the 35th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems . 261–276
-
[2020]
In 2020 IEEE 36th International Conference on Data Engineering (ICDE)
Approximate quantiles for datacenter telemetry monitoring. In 2020 IEEE 36th International Conference on Data Engineering (ICDE) . IEEE, 1914–1917
2020
-
[2022]
arXiv preprint arXiv:2208.04927 (2022)
Enabling efficient and general subpopulation analytics in multidimen- sional data streams. arXiv preprint arXiv:2208.04927 (2022)
2022 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.