REVIEW 5 major objections 5 minor 59 references
QPET: A Versatile and Portable Quantity-of-Interest-Preservation Framework for Error-Bounded Lossy Compression
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read QPET claims that any sufficiently differentiable quantity of interest can be preserved under error-bounded lossy compression using per-point Taylor bounds plus lossless outlier correction, yielding 2x-10x speedups over prior approaches.
desk verdict QPET is a genuine generalization of QoI-preserving compression, but its reported gains rest on an unmeasured outlier-correction overhead and a probabilistic bound with unverified assumptions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the point-wise error-bound estimator built on a second-order Taylor expansion with remainder: for univariate QoIs, $f(x') \approx f(x)+f'(x)(x'-x)+\frac{f''(x)}{2}(x'-x)^2$ yields the per-point bound $\min(\epsilon_g, (\sqrt{|a|^2+2|b|t}-|a|)/|b|)$ with $a=f'(x_i)$ and $b=f''(x_i)$, with a separate linear case when $f''(x_i)=0$. For multivariate QoIs the same machinery runs on a variable-separated decomposition $F = C + \sum_i \alpha_i f(x_i)$, where a deterministic bound (Theorem 5.3) and a sub-Gaussian concentration bound (Theorem 5.4) set the per-point QoI tolerance $t$ before Algorithm 2 converts it into data error bounds. The framework's second key component is the correction loop of Section 5.4, which detects decompressed points whose QoI error exceeds $\tau$ and losslessly stores the original values of those points, turning an approximate bound into a hard guarantee.
What would settle it
Run QPET with a loose global error bound (e.g., $\epsilon = 10^{-2}$) on a strongly autocorrelated smooth field, with QoI the block average of $x^3$, and measure the fraction of points the validator must store losslessly and the final bit rate against the best parameter-search baseline. If the outlier fraction rises substantially above 1% or the bit rate no longer beats the baseline, the sub-Gaussian assumption underlying Theorem 5.4 is the point of failure.
Extended reading notes
Core claim
QPET's central claim is that preserving a QoI under lossy compression can be reduced to a generic numerical problem. For a univariate QoI $f$, it uses a second-order Taylor expansion to solve, per data point $x_i$, the largest local error bound $\epsilon_i$ such that any decompressed value within $\epsilon_i$ keeps $|f(x')-f(x)|$ under the threshold $t$, with closed forms in Theorems 5.1 and 5.2. For multivariate QoIs it separates variables via a first-order differential or a linear decomposition and allocates per-point tolerances using a deterministic triangle-inequality bound (Theorem 5.3) and a sub-Gaussian concentration bound (Theorem 5.4). A global error-bound auto-tuner (Algorithm 4) then crops the point-wise bounds to reduce storage, and a QoI validator losslessly stores and corrects the few outliers so that the final output satisfies both $\|X-D\|_\infty \le \epsilon$ and $\|Q(X)-Q(D)\|_\infty \le \tau$. The strict guarantee comes from this correction step; the Taylor and probabilistic steps only make the correction overhead small.
Load-bearing premise
The load-bearing premise is the probabilistic model behind Theorem 5.4: that the per-point QoI errors are independent, symmetric, sub-Gaussian with a known variance proxy, so that only a small fraction of points need lossless correction; the paper states this is not always true.
Editorial extensions
If this is right
- A single QoI-preserving layer can be dropped onto interpolation-based compressors (SZ3, HPEZ) and a wavelet-based compressor (SPERR), so future compressors can gain QoI preservation without being redesigned.
- Users can specify a threshold on a derived quantity such as $\tanh x$, $x^3$, block averages, or vector magnitude, and QPET will find per-point error bounds automatically rather than requiring an analytic solution for each QoI.
- For multivariate QoIs with many variables, the concentration bound allows per-point tolerances to exceed the overall threshold, which is where the largest compression-ratio gains come from.
- The strict QoI guarantee is maintained even when the bounds are only estimates, because the validator losslessly corrects outliers; the practical cost is small as long as the fraction of outliers stays below about 1%.
- Parameter-search approaches that repeatedly compress and validate are replaced by a single forward pass of bound estimation, which explains the reported 2x-10x speedups.
Reading between the lines
- The framework's architectural contribution is the separation of QoI preservation into an estimation problem and a correction problem, so the exactness of the final guarantee does not depend on the Taylor or sub-Gaussian assumptions being perfectly true.
- On data with strongly correlated or heavy-tailed compression errors, such as smooth fields compressed with loose error bounds, the sub-Gaussian assumption should weaken, the outlier fraction should grow, and the compression-ratio advantage would shrink; a data-adaptive estimate of the variance-proxy parameter $c$ would be a direct extension.
- The same point-wise error-bound plus validation-and-correction pattern could be applied to non-differentiable QoIs by replacing Taylor bounds with finite-difference or automatic-differentiation surrogates, at the cost of more validation effort.
- Preserving several QoIs at once would amount to taking the point-wise minimum of the error bounds computed for each QoI, a combination the paper does not explore.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes QPET, a framework layer that can be inserted into existing error-bounded lossy compressors to preserve differentiable univariate and multivariate quantities of interest. QPET computes per-point error bounds using Taylor expansions (Theorems 5.1 and 5.2), uses a concentration inequality for variable-separated multivariate QoIs (Theorem 5.4), auto-tunes a global error bound (Algorithm 4), and losslessly corrects outlier points so that the final decompressed data satisfies both the data error bound and the QoI error threshold. The authors integrate QPET with SZ3, HPEZ, and SPERR and report 2x to 10x speedups over parameter-search and QoI-SZ3/QoI-HPEZ baselines, as well as compression-ratio improvements of up to 1000% over general-purpose compressors on selected datasets.
Significance. If the reported behavior holds, QPET is a genuinely useful contribution: it replaces hand-derived analytical error bounds for each QoI with a numerical routine, works with multiple compressor archetypes, and its correction step provides a strict end-to-end QoI guarantee that the Taylor and concentration estimates alone do not provide. The paper also includes an ablation study, a public artifact, and testing on six real-world datasets. The main uncertainty is not the correctness of the framework's final guarantee but the size of the correction overhead: the strong compression-ratio claims rely on the assertion that the losslessly stored outlier set is below 1%, and this assertion is never directly measured or reported.
major comments (5)
- [Section 5.4, Algorithm 1 lines 17-19] The strict QoI guarantee is carried entirely by the lossless outlier correction step, yet the paper never reports the size of the outlier set X_o or its bit-rate contribution. The statement that outliers are "only a tiny portion (<1%)" of the input is an unsupported assertion. Please add a table or figure reporting |X_o|/|X| and the bit-rate contributed by the losslessly compressed outliers for the configurations in Table 5 and Figures 5-7, especially at large error thresholds and for QoIs such as sin 10x and tanh x where the paper already notes limited compression gains.
- [Section 5.2.1, Theorem 5.4 and Eq. (7)] The proof of Theorem 5.4 sets the variance proxy to sigma_i = t/c with c around 2 to 3. For a bounded random variable with |D_i| <= t, Hoeffding's lemma gives a variance proxy of at most t, not t/c; choosing c > 1 is an additional distributional assumption that is not stated in the theorem's hypotheses. The uniform-distribution example with c = sqrt(3) conflates the standard deviation (t/sqrt(3)) with the sub-Gaussian variance proxy, which for a uniform variable is t. As written, the theorem's t values are optimistic unless a precise sub-Gaussian proxy assumption is stated and verified. Please either state the exact proxy assumption, justify the c values from measured error distributions, or present the concentration bound using the conservative proxy sigma = t.
- [Section 5.2.1, paragraph before Theorem 5.4] The independence and symmetry assumptions on the per-point QoI errors D_i are load-bearing for the concentration bound. For SZ3 and HPEZ, which are prediction-based compressors, compression errors are spatially correlated; the paper acknowledges this but provides no autocorrelation measurements or independence diagnostics for the six test datasets. Positive correlation increases the variability of the weighted sum sum_i alpha_i D_i relative to the independence-based bound, making the computed t too large and potentially increasing the outlier fraction. Please include autocorrelation of D_i or an alternative validation such as comparing predicted outlier rates with observed outlier rates across the experimental configurations.
- [Section 5.2.2, Eq. (8)] For non-variable-separated multivariate QoIs, the pointwise error-bound computation retains only the first-order Taylor term and discards higher-order and cross-partial terms. The resulting bounds are estimates whose bias is not quantified. Since the correction step in Section 5.4 is the only mechanism that provides a strict guarantee, the overhead of that step could grow for strongly nonlinear F. Please quantify this by reporting outlier fractions for the vector QoIs in Figure 7, or by augmenting Eq. (8) with a second-order remainder bound that controls the cross-derivative terms.
- [Section 6.1.3 and Figure 8(c)] The reported compression-ratio gains are sensitive to the free parameters c and beta, and the paper sets them differently per compressor (c = 2 for SZ3/HPEZ, c = 3 for SPERR) and linearly decreases c as the error threshold increases. Figure 8(c) shows that c = 3 gives substantially better compression ratio than c = 0 or c = 1 for the tested configuration, but no principled selection rule or cross-validation procedure is given. Please provide a default-selection criterion for c and beta, or a sensitivity analysis over datasets and QoIs, so that the reported advantages can be reproduced without per-dataset tuning.
minor comments (5)
- [Section 5.1, Eq. (3)] The Taylor expansion in Eq. (3) uses x0 in the remainder term while the surrounding text uses x_i; please standardize the notation to avoid confusion.
- [Section 5.2.1, Theorem 5.4] The theorem statement says "taking point-wise QoI error threshold t = max |f(x'_i) - f(x_i)| = ...", but t is a threshold to be set, not the maximum of the actual errors; the wording should be "setting t = ...".
- [Section 6.2.1] There is a typo in "QoI-preseving" in the first paragraph of Section 6.2.1.
- [References] References [15] and [16] are the same paper; please remove the duplicate and renumber.
- [Table 3 and Figures 5-7] The dataset name is written inconsistently as "Scale-LetKF" and "SCALE-LetKF" in different places; please unify.
Circularity Check
No significant circularity: QPET's strict QoI guarantee is enforced by the Section 5.4 lossless correction step, while the Taylor and concentration bounds are explicitly approximate and the self-citations are empirical, not load-bearing for correctness.
full rationale
Walking the derivation chain, the strict QoI guarantee is not claimed to follow from The estimation theorems. Theorem 5.1 is explicitly an 'estimation for the best-fit data error bound', and Theorem 5.4 is introduced under assumptions that the paper itself admits are 'not always true'. The actual guarantee is constructive and self-contained: Section 5.4 computes QoI values on the decompressed data, compares them with the original QoI values, and losslessly stores any out-of-constraint data values. This is a definitional property of the algorithm, not a prediction fitted to its own output. The concentration bound in Theorem 5.4 is a valid Hoeffding consequence under its stated sub-Gaussian condition; the choice sigma0 = t/c is a modeling assumption, and the paper self-cites [35,36,58] for empirical support, but the correctness of QPET does not rest on that assumption because violations are corrected. Algorithm 4 tunes the global error bound by performing actual compression tests on the data, so the reported compression ratios are measured optima rather than forward predictions of QoI error. The main evidence gaps, namely the unmeasured outlier fraction in Section 5.4 ('only a tiny portion (<1%)') and the possible violation of independence for interpolation-based compressors, are limitations of the efficiency claim rather than circular reductions. No load-bearing step reduces to its own inputs by construction, so the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- c (variance-proxy divisor in Theorem 5.4) =
2 for SZ3 and HPEZ, 3 for SPERR, linearly decreasing to 1.0 as tau grows from 1e-3 to 1e-2
- beta (confidence level in Theorem 5.4) =
0.999 for HPEZ and SPERR, 0.99999 for SZ3
- quantile set in Algorithm 4 =
{0.2, 0.1, 0.05, 0.02, 0.01, 0.005, 0.0025}
- dropping threshold coefficient c0 in Algorithm 4 =
not specified, <=1
assumptions (5)
- domain assumption QoI functions are second-order differentiable on the data range
- ad hoc to paper Per-point QoI errors D_i are independent, symmetric, and sub-Gaussian with variance proxy sigma <= t/c
- ad hoc to paper Compression error distributions are near-uniform or near-Gaussian and autocorrelation drops rapidly at fine error bounds
- domain assumption The fraction of outlier points requiring lossless correction is below 1%
- standard math Third derivative f''' is bounded on [x_i - eps_i, x_i + eps_i]
Cite this review
Pith. "Pith review of QPET: A Versatile and Portable Quantity-of-Interest-Preservation Framework for Error-Bounded Lossy Compression." pith.science (2026). https://pith.science/paper/PCGQASMW
@misc{pith2026241202799,
author = {Pith},
title = {Pith review of: QPET: A Versatile and Portable Quantity-of-Interest-Preservation Framework for Error-Bounded Lossy Compression},
year = {2026},
howpublished = {\url{https://pith.science/paper/PCGQASMW}},
note = {Machine review of arXiv:2412.02799}
}
read the original abstract
Error-bounded lossy compression has been widely adopted in many scientific domains because it can address the challenges in storing, transferring, and analyzing unprecedented amounts of scientific data. Although error-bounded lossy compression offers general data distortion control by enforcing strict error bounds on raw data, it may fail to meet the quality requirements on the results of downstream analysis, a.k.a. Quantities of Interest (QoIs), derived from raw data. This may lead to uncertainties and even misinterpretations in scientific discoveries, significantly limiting the use of lossy compression in practice. In this paper, we propose QPET, a novel, versatile, and portable framework for QoI-preserving error-bounded lossy compression, which overcomes the challenges of modeling diverse QoIs by leveraging numerical strategies. QPET features (1) high portability to multiple existing lossy compressors, (2) versatile preservation to most differentiable univariate and multivariate QoIs, and (3) significant compression improvements in QoI-preservation tasks. Experiments with six real-world datasets demonstrate that integrating QPET into state-of-the-art error-bounded lossy compressors can gain 2x to 10x compression speedups of existing QoI-preserving error-bounded lossy compression solutions, up to 1000% compression ratio improvements to general-purpose compressors, and up to 133% compression ratio improvements to existing QoI-integrated scientific compressors.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
[n.d.]. HDF5. http://www.hdfgroup.org/HDF5. Last Accessed: 2025-05-25
work page 2025
-
[2]
[n.d.]. nvCOMP. https://github.com/NVIDIA/nvcomp. Last Accessed: 2025-05- 25
work page 2025
-
[3]
EXAALT: Malecular Dynamics at the Exascale
2020. EXAALT: Malecular Dynamics at the Exascale. https://www. exascaleproject.org/wp-content/uploads/2019/10/EXAALT.pdf. Online, Last Accessed: 2025-05-25
work page 2020
-
[4]
2021. Team at Princeton Plasma Physics Laboratory employs DOE super- computers to understand heat-load width requirements of future ITER de- vice. https://www.olcf.ornl.gov/2021/02/18/scientists-use-supercomputers-to- study-reliable-fusion-reactor-design-operation. Online, Last Accessed: 2025-05- 25
work page 2021
-
[5]
Nitin Agrawal and Ashish Vulimiri. 2017. Low-latency analytics on colossal data streams with summarystore. In Proceedings of the 26th Symposium on Operating Systems Principles. 647–664
work page 2017
-
[6]
Mark Ainsworth, Ozan Tugluk, Ben Whitney, and Scott Klasky. 2019. Multilevel techniques for compression and reduction of scientific data-quantitative control of accuracy in derived quantities. SIAM Journal on Scientific Computing 41, 4 (2019), A2146–A2171
work page 2019
-
[7]
Jyrki Alakuijala, Andrea Farruggia, Paolo Ferragina, Eugene Kliuchnikov, Robert Obryk, Zoltan Szabadka, and Lode Vandevenne. 2018. Brotli: A general-purpose data compressor. ACM Transactions on Information Systems (TOIS) 37, 1 (2018), 1–30
2018
-
[8]
Andrei Arion, Angela Bonifati, Ioana Manolescu, and Andrea Pugliese. 2007. XQueC: A query-conscious compressed XML database. ACM Transactions on Internet Technology (TOIT) 7, 2 (2007), 10–es
work page 2007
Show all 59 references
-
[9]
Benjamin Bross, Ye-Kui Wang, Yan Ye, Shan Liu, Jianle Chen, Gary J Sullivan, and Jens-Rainer Ohm. 2021. Overview of the versatile video coding (VVC) stan- dard and its applications. IEEE Transactions on Circuits and Systems for Video Technology 31, 10 (2021), 3736–3764
2021
-
[10]
Zhiyuan Chen, Johannes Gehrke, and Flip Korn. 2001. Query optimization in com- pressed database systems. In Proceedings of the 2001 ACM SIGMOD international conference on Management of data . 271–282
2001
-
[11]
Yann Collet. 2015. Zstandard – Real-time data compression algorithm. http://facebook.github.io/zstd/ (2015)
2015
-
[12]
L Peter Deutsch. 1996. GZIP file format specification version 4.3
1996
-
[13]
Elmagarmid, Emmanuel Cecchet, Walid G
Hazem Elmeleegy, Ahmed K. Elmagarmid, Emmanuel Cecchet, Walid G. Aref, and Willy Zwaenepoel. 2009. Online Piece-Wise Linear Approximation of Nu- merical Streams with Precision Guarantees. Proc. VLDB Endow. 2, 1 (Aug. 2009), 145–156
2009
-
[14]
Wozniak, Wei Xu, and Shinjae Yoo
Ian Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Youl Choi, Emil Constantinescu, Philip E Davis, Sheng Di, Wendy Di, Hanqi Guo, Scott Klasky, Kerstin Kleese Van Dam, Tahsin Kurc, Qing Liu, Abid Malik, Kshi- tij Mehta, Klaus Mueller, Todd Munson, Geo...
2017
-
[16]
Qian Gong, Xin Liang, Ben Whitney, Jong Youl Choi, Jieyang Chen, Lipeng Wan, Stéphane Ethier, Seung-Hoe Ku, R Michael Churchill, C-S Chang, et al
-
[17]
Wassily Hoeffding. 1963. Probability Inequalities for Sums of Bounded Random Variables. J. Amer. Statist. Assoc. 58, 301 (1963), 13–30. http://www.jstor.org/ stable/2282952
1963
-
[18]
In Smoky Mountains Computational Sciences and Engineering Conference
Maintaining trust in reduction: Preserving the accuracy of quantities of interest for lossy compression. In Smoky Mountains Computational Sciences and Engineering Conference. Springer, 22–39
-
[19]
Søren Kejser Jensen, Torben Bach Pedersen, and Christian Thomsen. 2018. Mod- elardb: Modular model-based time series management with spark and cassandra. Proceedings of the VLDB Endowment 11, 11 (2018), 1688–1701
2018
-
[20]
Lawrence Ibarria, Peter Lindstrom, Jarek Rossignac, and Andrzej Szymczak. 2003. Out-of-core compression and decompression of large n-dimensional scalar fields. In Computer Graphics Forum, Vol. 22. Wiley Online Library, 343–348
2003
-
[21]
Søren Kejser Jensen, Torben Bach Pedersen, and Christian Thomsen. 2019. Scal- able Model-Based Management of Correlated Dimensional Time Series in Mode- larDB+. arXiv e-prints (2019), arXiv–1903
2019
-
[22]
Pu Jiao, Sheng Di, Hanqi Guo, Kai Zhao, Jiannan Tian, Dingwen Tao, Xin Liang, and Franck Cappello. 2022. Toward Quantity-of-Interest Preserving Lossy Com- pression for Scientific Data. Proceedings of the VLDB Endowment 16, 4 (2022), 697–710
2022
-
[23]
Fabian Knorr, Peter Thoman, and Thomas Fahringer. 2021. ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUs. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1–14
2021
-
[24]
Fabian Knorr, Peter Thoman, and Thomas Fahringer. 2021. ndzip: A high- throughput parallel lossless compressor for scientific data. In 2021 Data Com- pression Conference (DCC). IEEE, 103–112
2021
-
[25]
Shaomeng Li, Peter Lindstrom, and John Clyne. 2023. Lossy scientific data compression with SPERR. In 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 1007–1017
2023
-
[26]
Lazaridis and S
I. Lazaridis and S. Mehrotra. 2003. Capturing sensor-generated time series with quality guarantees. In Proceedings 19th International Conference on Data Engineer- ing (Cat. No.03CH37405). 429–440. https://doi.org/10.1109/ICDE.2003.1260811
2003 arXiv
-
[27]
Xin Liang, Sheng Di, Franck Cappello, Mukund Raj, Chunhui Liu, Kenji Ono, Zizhong Chen, Tom Peterka, and Hanqi Guo. 2022. Toward feature-preserving vector field compression. IEEE Transactions on Visualization and Computer Graphics 29, 12 (2022), 5434–5450
2022
-
[28]
Panagiotis Liakos, Katia Papakonstantinopoulou, and Yannis Kotidis. 2022. Chimp: efficient lossless floating point compression for time series databases. Proceedings of the VLDB Endowment 15, 11 (2022), 3058–3070
2022
-
[29]
Xin Liang, Sheng Di, Dingwen Tao, Sihuan Li, Shaomeng Li, Hanqi Guo, Zizhong Chen, and Franck Cappello. 2018. Error-Controlled Lossy Compression Opti- mized for High Compression Ratios of Scientific Datasets. In 2018 IEEE Interna- tional Conference on Big Data . IEEE
2018
-
[30]
Xin Liang, Sheng Di, Dingwen Tao, Zizhong Chen, and Franck Cappello. 2018. An efficient transformation scheme for lossy data compression with point-wise relative error bound. In 2018 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, 179–189
2018
-
[31]
Gok, Jiannan Tian, Junjing Deng, Jon C
Xin Liang, Kai Zhao, Sheng Di, Sihuan Li, Robert Underwood, Ali M. Gok, Jiannan Tian, Junjing Deng, Jon C. Calhoun, Dingwen Tao, Zizhong Chen, and Franck Cappello. 2023. SZ3: A Modular Framework for Composing Prediction-Based Error-Bounded Lossy Compressors. IEEE Transactions ...
2023
-
[32]
Xin Liang, Hanqi Guo, Sheng Di, Franck Cappello, Mukund Raj, Chunhui Liu, Kenji Ono, Zizhong Chen, and Tom Peterka. 2020. Toward Feature-Preserving 2D and 3D Vector Field Compression.. In PacificVis. 81–90
2020
-
[33]
Peter G Lindstrom et al. 2017. Fpzip. Technical Report. Lawrence Livermore National Lab.(LLNL), Livermore, CA (United States)
2017
-
[34]
Peter Lindstrom. 2014. Fixed-rate compressed floating-point arrays. IEEE trans- actions on visualization and computer graphics 20, 12 (2014), 2674–2683
2014
-
[35]
Jinyang Liu, Sheng Di, Kai Zhao, Xin Liang, Zizhong Chen, and Franck Cappello
-
[36]
Chunwei Liu, Hao Jiang, John Paparrizos, and Aaron J Elmore. 2021. Decom- posed bounded floats for fast compression and queries. Proceedings of the VLDB Endowment 14, 11 (2021), 2586–2598
2021
-
[37]
Kiyoshi Masui, Mandana Amiri, Liam Connor, Meiling Deng, Mateus Fandino, Carolin Höfer, Mark Halpern, David Hanna, Adam D Hincks, Gary Hinshaw, et al. 2015. A compression scheme for radio data in high performance computing. Astronomy and Computing 12 (2015), 181–190
2015
-
[38]
Tuomas Pelkonen et al. 2015. Gorilla: A Fast, Scalable, in-Memory Time Series Database. Proc. VLDB Endow. 8, 12 (Aug. 2015), 1816–1827
2015
-
[39]
Jinyang Liu, Sheng Di, Kai Zhao, Xin Liang, Sian Jin, Zizhe Jian, Jiajun Huang, Shixun Wu, Zizhong Chen, and Franck Cappello. 2024. High-performance effective scientific error-bounded lossy compression with auto-tuned multi- component interpolation. Proceedings of the ACM on M...
2024
-
[40]
Zhaoyuan Su, Sheng Di, Ali Murat Gok, Yue Cheng, and Franck Cappello. 2022. Understanding impact of lossy compression on derivative-related metrics in sci- entific datasets. In 2022 IEEE/ACM 8th International Workshop on Data Analysis and Reduction for Big Scientific Data (DRB...
2022
-
[41]
Vivienne Sze, Madhukar Budagavi, and Gary J Sullivan. 2014. High efficiency video coding (HEVC). In Integrated circuit and systems, algorithms and architec- tures. Vol. 39. Springer, 40
2014
-
[42]
X Carol Song, Preston Smith, Rajesh Kalyanam, Xiao Zhu, Eric Adams, Kevin Colby, Patrick Finnegan, Erik Gough, Elizabett Hillery, Rick Irvine, et al. 2022. Anvil-system architecture and experiences from deployment and early user operations. In Practice and experience in advanc...
2022
-
[43]
Dingwen Tao, Sheng Di, Hanqi Guo, Zizhong Chen, and Franck Cappello. 2019. Z-checker: A framework for assessing lossy compression of scientific data. The International Journal of High Performance Computing Applications 33, 2 (2019), 285–303. https://doi.org/10.1177/1094342017737147
2019 doi
-
[44]
David S Taubman and Michael W Marcellin. 2002. JPEG2000: Standard for interactive imaging. Proc. IEEE 90, 8 (2002), 1336–1357
2002
-
[45]
Dingwen Tao, Sheng Di, Zizhong Chen, and Franck Cappello. 2017. Significantly improving lossy compression for scientific data sets based on multidimensional prediction and error-controlled quantization. In 2017 IEEE International Parallel and Distributed Processing Symposium ....
2017
-
[46]
Jiannan Tian, Sheng Di, Kai Zhao, Cody Rivera, Megan Hickman Fulp, Robert Underwood, Sian Jin, Xin Liang, Jon Calhoun, Dingwen Tao, and Franck Cap- pello. 2020. cuSZ: An Efficient GPU-Based Error-Bounded Lossy Compression Framework for Scientific Data. InProceedings of the ACM...
2020
-
[47]
Robert Underwood, Jon C Calhoun, Sheng Di, Amy Apon, and Franck Cap- pello. 2022. OptZConfig: Efficient Parallel Optimization of Lossy Compression Configuration. IEEE Transactions on Parallel and Distributed Systems (2022)
2022
-
[48]
Jiannan Tian, Sheng Di, Xiaodong Yu, Cody Rivera, Kai Zhao, Sian Jin, Yunhe Feng, Xin Liang, Dingwen Tao, and Franck Cappello. 2021. cuSZ (x): Optimizing Error-Bounded Lossy Compression for Scientific Data on GPUs. CoRR (2021)
2021
-
[49]
Roman Vershynin. 2018. High-dimensional probability: An introduction with applications in data science . Vol. 47. Cambridge university press
2018
-
[50]
Martin J Wainwright. 2019. High-dimensional statistics: A non-asymptotic view- point. Vol. 48. Cambridge university press
2019
-
[51]
Robert Underwood, Sheng Di, Jon C Calhoun, and Franck Cappello. 2020. Fraz: A generic high-fidelity fixed-ratio lossy compression framework for scientific floating-point data. In 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 567–577
2020
-
[52]
Thomas Wiegand, Gary J Sullivan, Gisle Bjontegaard, and Ajay Luthra. 2003. Overview of the H. 264/AVC video coding standard.IEEE Transactions on circuits and systems for video technology 13, 7 (2003), 560–576
2003
-
[53]
Xuan Wu, Qian Gong, Jieyang Chen, Qing Liu, Norbert Podhorszki, Xin Liang, and Scott Klasky. 2024. Error-controlled Progressive Retrieval of Scientific Data under Derivable Quantities of Interest. In 2024 SC24: International Conference for High Performance Computing, Networkin...
2024
-
[54]
Gregory K Wallace. 1991. The JPEG still picture compression standard. Commun. ACM 34, 4 (1991), 30–44
1991
-
[55]
Lin Yan, Xin Liang, Hanqi Guo, and Bei Wang. 2023. TopoSZ: Preserving topol- ogy in error-bounded lossy compression. IEEE Transactions on Visualization and Computer Graphics (2023)
2023
-
[56]
Feng Zhang, Zaifeng Pan, Yanliang Zhou, Jidong Zhai, Xipeng Shen, Onur Mutlu, and Xiaoyong Du. 2021. G-TADOC: Enabling efficient GPU-based text analyt- ics without decompression. In 2021 IEEE 37th International Conference on Data Engineering (ICDE). IEEE, 1679–1690
2021
-
[57]
Mingze Xia, Sheng Di, Franck Cappello, Pu Jiao, Kai Zhao, Jinyang Liu, Xuan Wu, Xin Liang, and Hanqi Guo. 2024. Preserving Topological Feature with Sign- of-Determinant Predicates in Lossy Compression: A Case Study of Vector Field Critical Points. In 2024 IEEE 40th Internation...
2024
-
[58]
Tonellot, Zizhong Chen, and Franck Cappello
Kai Zhao, Sheng Di, Maxim Dmitriev, Thierry-Laurent D. Tonellot, Zizhong Chen, and Franck Cappello. 2021. Optimizing Error-Bounded Lossy Com- pression for Scientific Data by Dynamic Spline Interpolation. In 2021 IEEE 37th International Conference on Data Engineering (ICDE) . 1...
2021
-
[60]
Feng Zhang, Jidong Zhai, Xipeng Shen, Onur Mutlu, and Wenguang Chen. 2018. Efficient document analytics on compressed data: Method, challenges, algorithms, insights. Proceedings of the VLDB Endowment 11, 11 (2018), 1522–1535
2018
-
[2022]
In 2022 SC22: International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
Dynamic quality metric oriented error bounded lossy compression for scientific datasets. In 2022 SC22: International Conference for High Performance Computing, Networking, Storage and Analysis (SC) . IEEE Computer Society, 892– 906
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.