{"id":"492f6f75-6338-406f-b806-3621e6105e8d","arxiv_id":"2608.08386","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FZ-VIS provides an interactive environment for exploring QoI-aware lossy compression configurations with linked metric, spatial, and feature-preservation views.","lead":"FZ-VIS is a web-based visual analytics framework that lets scientists interactively design, compare, and refine lossy compression settings while tracking application-specific quantities of interest. It demonstrates how linked views of compression metrics, spatial error, and topological or spectral features can guide compressor selection.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'helps users' claim rests solely on narrative case studies; without a controlled task-based comparison or interaction-latency measurements, the central contribution is not empirically established.","rationale":"The paper is a plausible and well-integrated systems contribution: the architecture is coherent, the source code is released under MIT, the supplemental materials include a demonstration video, and the case studies illustrate concrete workflows that combine configuration, evaluation, and QoI correction. The technical content, including the MSz and FFCz integration and the linked-view design, is internally consistent and does not appear to contain a correctness error. However, the central claim is not about algorithmic correctness; it is about whether the framework helps users make better compression decisions. The only evidence for that claim is narrative. The reader's weakest assumption correctly identifies this gap: the case studies show that a user could perform the workflow, not that FZ-VIS outperforms existing scripted workflows in time, accuracy, or usability. I agree with the reader's conditional verdict. I would not reject the paper, because the framework has independent value as a demonstrated integration and the code is available for further evaluation, but acceptance should be conditioned on either a controlled study or a careful weakening of the 'helps users' framing to 'supports the illustrated workflows.' The verification step I propose is feasible precisely because the source code is released, making a head-to-head task comparison against LibPressio-based scripting possible.","tokens_in":18838,"tokens_out":2718,"duration_ms":32483,"concrete_test":"Run a pre-registered within-subjects task study with 12–16 participants spanning the three user groups (novice, developer, domain scientist). Each participant performs the three case-study tasks once with FZ-VIS and once with a scripted baseline built on LibPressio plus Matplotlib and Z-checker, in counterbalanced order. Record completion time, number of configurations explored, decision accuracy against expert-defined ground truth, and System Usability Scale score. Pre-specify thresholds, e.g., FZ-VIS must not be significantly slower and must show higher decision accuracy or a significant usability improvement. If no such difference is observed, the 'enables users' claim in Section 8 is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FZ-VIS's central claim, stated in Section 8, is that it 'enables users to explore compressor variants, examine trade-offs, and evaluate feature preservation for downstream analysis.' The only support offered is the narrative case studies in Section 6. These demonstrate that the intended workflows can be constructed, but they do not measure whether the framework improves decision-making over the scripted LibPressio/Z-checker/Foresight + Matplotlib workflows that the paper itself positions as the status quo. There is no task-completion time, no decision accuracy metric, no expert-defined ground truth, and no baseline comparison in the main text. Section 7 lists limitations about composition flexibility, QoI coverage, and scalability, but conspicuously omits the lack of empirical user evaluation, even though the introduction explicitly claims FZ-VIS 'demonstrably helps users efficiently navigate complex design spaces.' The linked-view benefit asserted in Section 5.5 is an unmeasured assumption, not a demonstrated result. The Table 1 workflow comparison is a self-assessment on qualitative rows and does not substitute for measured user performance. For a visual analytics systems paper, this may be acceptable as a design contribution, but it does not support the stronger 'helps users' claim without an empirical task study or at least reproducible quantitative workflow metrics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents FZ-VIS, a web-based visual analytics framework for quantities-of-interest (QoI)-aware lossy compression. The system integrates interactive compressor configuration, batch generation of configuration variants, linked quantitative and spatial comparison views, inspection of intermediate compressor outputs, and QoI-specific analysis modules for Morse-Smale segmentation and power-spectrum fidelity, including correction workflows based on the authors' previously published MSz and FFCz algorithms. The paper claims that FZ-VIS unifies the compressor design space and the scientific evaluation space, enabling novice users, compressor developers, and domain scientists to explore compression trade-offs and make informed decisions. The utility is demonstrated through three narrative case studies on CESM, Hurricane Isabel, Atmospheric Rivers, and NYX data.","tokens_in":19236,"tokens_out":2516,"duration_ms":26663,"significance":"If the central claim is validated, FZ-VIS would address a genuine gap: existing tools such as LibPressio, Z-checker, Foresight, and ParaView do not provide an integrated environment for QoI-aware compression exploration. The framework is thoughtfully designed around a three-stage workflow, and the authors explicitly connect requirements to interface features. The paper also provides reproducibility artifacts, including a workflow graph translated to a JSON specification, open-source code, and supplementary materials. The underlying QoI-correction algorithms (MSz, FFCz) are previously published and are not fitted to the results of this paper. However, the paper's strongest claim—that FZ-VIS 'enables users to explore compressor variants, examine trade-offs, and evaluate feature preservation' (Section 8)—is supported only by illustrative case studies, not by controlled empirical evaluation. The case studies show that the intended workflows can be constructed, but they do not measure whether the framework improves decision-making relative to existing scripted workflows. This limits the current evidentiary basis for the human-in-the-loop benefit that motivates the work.","major_comments":[{"comment":"The central claim that FZ-VIS 'enables users to explore compressor variants, examine trade-offs, and evaluate feature preservation for downstream analysis' is supported only by narrative case studies. The three case studies in Section 6 demonstrate that the system can be used to perform specific workflow sequences, but they do not include task-completion time, decision accuracy, expert-defined ground truth, or any comparison with the scripted LibPressio/Z-checker/Foresight workflows that the paper positions as the status quo. Since the conclusion (Section 8) repeats this claim as a demonstrated result, the manuscript either needs a controlled task-based user study or a substantially tempered statement of what has been established.","section":"Section 6 and Section 8"},{"comment":"The workflow-level comparison in Table 1 is a self-assessment on qualitative rows such as 'Linked quantitative and spatial comparison' and 'Single-environment workflow (no external scripting).' The table does not report measured performance of any tool by any user. As presented, it cannot support the claim that FZ-VIS reduces coordination overhead or improves workflow efficiency over existing tools; it only states the authors' design intent. A quantitative comparison of task completion or at least a reproducible scripted workflow comparison would be needed to make this row meaningful.","section":"Appendix A, Table 1"},{"comment":"The Limitations and Discussion section lists composition flexibility, QoI coverage, and scalability as limitations, but it does not acknowledge the absence of any empirical user evaluation. Given that the introduction explicitly claims FZ-VIS 'demonstrably helps users efficiently navigate complex design spaces,' the lack of measured evidence is a load-bearing omission. The limitations section should either report the absence of a controlled evaluation or point to a supplementary evaluation that provides such evidence.","section":"Section 7"},{"comment":"The favorable QoI-correction results in Sections 6.3 and 6.4 are obtained using MSz and FFCz, which are the authors' own previously published algorithms. This does not make the results incorrect, but it means the case studies partly showcase the authors' own techniques rather than independently established methods. The paper should explicitly state this relationship and discuss how the results would differ if independent correction algorithms were used, or at least clearly frame the case studies as validations of the integrated workflow rather than of the correction algorithms themselves.","section":"Sections 6.3 and 6.4"}],"minor_comments":[{"comment":"The abstract states that case studies 'demonstrate the utility' of FZ-VIS, while Section 6 says the case studies are 'intended to illustrate integrated workflow support.' These statements should be aligned; 'illustrate' is more accurate given the absence of controlled evaluation.","section":"Abstract and Section 6"},{"comment":"The JSON specification example in Figure 5 contains a mixed quotation mark in the metrics array (\"psnr\", “mse\"). This is likely a typesetting artifact, but it should be corrected for consistency.","section":"Figure 5"},{"comment":"The Overall Compression Ratio (OCR) is defined in the text, but the definition appears only after Figure 13 is referenced. Placing the definition before the figure reference would improve readability.","section":"Section 6.3"},{"comment":"The text states that ZFP keeps the spectral relative error 'below 0.0020%' within the ROI. Given that this is a relative metric, it would be helpful to clarify the baseline against which this percentage is computed.","section":"Section 6.4"},{"comment":"The paper cites several accepted-but-not-yet-published works (e.g., [26], [39]). For a journal submission, providing DOIs or stable URLs for accepted works would be helpful.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid systems contribution with a clear architecture and reproducible artifacts, but the gap between the claims and the evidence is substantial. The case studies are well-illustrated, yet they do not meet the bar for 'demonstrably helps users' in a visual analytics venue. I would encourage the editor to request a controlled user study or a clear repositioning of the contribution as a design paper with more modest claims. The circularity concern regarding MSz and FFCz is real but manageable through explicit framing; it is not a reason for rejection by itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The quick read: FZ-VIS is a real, working system—not a mockup or a wireframe. Source code is out (MIT), OSF supplements include a video and formative study, and the architecture genuinely integrates LibPressio-based compressor composition, batch configuration generation, linked quantitative/spatial views, intermediate-stage inspection, and QoI modules for topology and spectral analysis. That integration is new for this problem: Z-checker, Foresight, and LibPressio give you pieces, but not the full iterate-compare-inspect-correct loop. The four case studies (CESM, Hurricane Isabel, atmospheric rivers, NYX) show concrete workflows and report real numbers like compression ratios and spectral errors. A systems paper with this much working code deserves to be taken seriously.\n\nThe soft spot is the gap between what is shown and what is claimed. Abstract and conclusion say FZ-VIS 'helps users efficiently navigate complex design spaces' or 'demonstrably helps users,' but the evidence is narrative case studies. There is no task-completion time, no decision-accuracy metric, no baseline comparison against scripted LibPressio/Z-checker workflows, no user study in the main text. Section 7 lists limitations but silently skips the evaluation gap. That's a real omission, and the stress-test note lands.\n\nI don't think this sinks the paper. For a VIS systems paper, a working implementation plus illustrative case studies is often sufficient to claim a design contribution—but the authors overreach when they claim to demonstrate user benefit. The fix is straightforward: either run a modest controlled comparison (even a small expert study with three to five users) or rewrite the claims to say 'supports the intended workflow' rather than 'provably helps users.' The formative user study mentioned in the supplements could also be reported in the main text if it contains any measurable findings.\n\nOn the QoI correction modules: MSz and FFCz are the authors' own prior algorithms, so the favorable QoI results partly showcase their own work. That's not circular in the fitting sense, and the citations are appropriate, but it's worth being explicit that the correction modules come from the same group.\n\nBottom line: this deserves peer review, not a desk reject. The design contribution is solid and reproducible; the evaluation claim needs to be brought back in line with the evidence. I'd send it out with a request for either a user study or a softened claim, and I'd probably cite the system if I work on compressor evaluation.","headline":"FZ-VIS is a genuine, well-integrated system for QoI-aware compression design, but its 'helps users' claim outruns the narrative-case-study evidence.","tokens_in":19612,"tokens_out":4246,"would_cite":true,"duration_ms":42285,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FZ-VIS claims that quantities-of-interest-aware lossy compression design can be unified into a single interactive visual analytics workflow, and supports the claim with case studies across three user groups.","keywords":["lossy compression","quantities of interest","visual analytics","Morse-Smale segmentation","power spectrum","human-in-the-loop","error-bounded compression","interactive exploration"],"falsifier":"Conduct the same compressor-selection and QoI-validation tasks in a controlled study with two groups, one using FZ-VIS and one using scripted configuration plus standard metric and QoI scripts, and compare task completion time, number of configurations attempted, and the QoI preservation of the final choice. If the FZ-VIS group is not faster or does not end with comparable-or-better configurations, the framework's central value claim fails.","tokens_in":18693,"feed_emoji":"📊","tokens_out":6765,"duration_ms":63649,"temperature":0.7,"pith_summary":"The paper tries to establish that choosing and tuning a lossy compressor for scientific data should be treated as a human-in-the-loop visual analytics problem, and that a single interactive environment can replace the usual sequence of scripts, benchmark tables, and separate analysis tools. It presents FZ-VIS, a web-based framework that couples the compressor design space (module choices, error bounds, execution variants) with a scientific evaluation space (fidelity metrics, spatial comparison, and application-specific quantities of interest). The motivating gap is that error-bounded compressors may preserve pointwise error while destroying downstream features such as topological segmentations or power spectra, and existing evaluation tools do not make those failures visible during compressor selection. If the central claim holds, scientists can judge a compressor by whether the features their analysis depends on survive, and can correct failures in the same workflow.","feed_headline":"Compression tuning now checks whether scientific features survive","feed_subtitle":"A web-based framework links rate, fidelity, and quantities-of-interest views so users pick compressors by downstream impact.","key_machinery":"The central mechanism is the linked-view workflow that ties the compressor design space to the scientific evaluation space. Concretely: users compose a reference compressor pipeline and branch it into configuration families via parallel coordinates and a force-directed graph; a batch runner executes variants and stores stage-wise intermediates; quantitative views (rate-distortion diagrams, rankings, spectral-error curves) are linked to spatial views (error maps, region-of-interest comparisons); and pluggable QoI modules compute topology and spectrum. The topological module uses Morse-Smale segmentations - partitions of a scalar field into regions sharing the same pair of extrema - to count false structures, and the spectral module compares power spectra $P(k)=\\sum_{u^2+v^2+w^2=k^2}|X'_{u,v,w}|^2$. Correction is provided by MSz, an edit-based method that adjusts decompressed values within the error bound to restore Morse-Smale segmentations, and FFCz, which projects reconstruction errors onto the intersection of spatial and frequency-domain constraints.","core_discovery":"FZ-VIS's central claim is that unifying compressor configuration, batch evaluation, and quantities-of-interest analysis in one workflow lets users make better-informed compression decisions. The framework structures work in three stages - setup, composition and execution, and analysis and refinement - with generated JSON specifications preserving provenance. Its linked views connect quantitative summaries (rate-distortion curves, ranked candidates) to spatial comparisons (original/decompressed/error maps and region-of-interest detail), and its QoI modules compare Morse-Smale segmentations and power spectra between original and reconstructed data, with optional correction stages that repair topology or spectrum within the same environment. The case studies show a novice selecting a compressor under a compression-ratio and DSSIM target, a developer diagnosing why a regression predictor underperforms the Lorenzo predictor by inspecting residual and quantization-index distributions, a domain scientist finding that tightening error bounds alone does not preserve atmospheric-river Morse-Smale segmentations and then applying a correction, and a cosmology scientist trading spectral fidelity against compression ratio and recovering a high-ratio candidate through a frequency-domain correction.","pith_inferences":["A reasonable next step not explored in the paper is to use the accumulated linked metric and QoI results to train a lightweight surrogate that predicts QoI preservation from compressor parameters, letting users skip the interactive loop for routine datasets.","The correction-module pattern suggests a broader design principle: any lossy compressor paired with a post hoc feature-preserving corrector can be exposed as a single tunable unit, so the framework's value may grow as more correctors are built for other quantities.","The case studies leave open whether the benefit is largest for novices, who gain guided exploration, or for domain scientists, who gain QoI awareness; a controlled experiment could test which group's decisions improve most."],"forward_implications":["Users can determine, in one session, whether any candidate compressor preserves the topological structures or power spectra their downstream analysis needs, rather than learning this after committing to a configuration.","Compressor developers can trace a performance difference to a specific pipeline stage by comparing residual histograms, quantization-index spreads, and stage-wise runtimes across variants.","QoI failures that pointwise error bounds cannot cure are not dead ends: correction modules can repair Morse-Smale segmentations or spectrum within the same workflow, at a visible cost in compression ratio.","Novice users can navigate a large parameter space by branching from a reference pipeline and filtering candidates on rate, fidelity, and QoI criteria, turning blind parameter tuning into structured exploration.","Because QoI modules are pluggable stages, the same workflow pattern extends to other application-specific quantities beyond topology and spectrum."],"supporting_citations":[{"why":"Supplies the unified compressor configuration interface that FZ-VIS builds on for interactive pipeline composition.","marker":"[47]"},{"why":"Provides the standard lossy-compression evaluation framework whose metric-focused scope FZ-VIS extends with QoI analysis.","marker":"[46]"},{"why":"Represents the extensible benchmarking pipeline that, alongside other tools, motivates the single-workflow design.","marker":"[16]"},{"why":"Supplies the MSz correction algorithm the topology module uses to restore Morse-Smale segmentations within the error bound.","marker":"[25]"},{"why":"Supplies the FFCz correction algorithm the spectral module uses to preserve the power spectrum under frequency error bounds.","marker":"[39]"},{"why":"Defines the modular prediction-based compressor whose predictor, quantizer, encoder, and lossless options are explored in the case studies.","marker":"[30]"},{"why":"Defines the transform-based compressor used as the comparison baseline in the selection and QoI case studies.","marker":"[31]"},{"why":"Establishes Morse-Smale segmentations of Integrated Vapor Transport as the quantity of interest for atmospheric-river analysis.","marker":"[22]"}],"fun_headline_variants":["Feature-aware compression tuning with visual analytics","See if lossy compression preserves your key features","Interactive framework balances compression ratio and fidelity","Compression that cares about your quantities of interest","Tune lossy compression by what matters: your features"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that interactive, linked-view exploration genuinely improves users' compression decisions compared with existing scripted or benchmark-based workflows; the paper supports this with narrative case studies rather than a controlled measurement of decision time or quality.","fun_headline_variants_meta":{"raw":{"variants":["Feature-aware compression tuning with visual analytics","See if lossy compression preserves your key features","Interactive framework balances compression ratio and fidelity","Compression that cares about your quantities of interest","Tune lossy compression by what matters: your features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1351,"prompt_tokens":955,"completion_tokens":396,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":327}},"tokens_in":571,"tokens_out":396,"duration_ms":4359,"temperature":1.0,"reasoning_tokens":327,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:36:03.343429+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct the same compressor-selection and QoI-validation tasks in a controlled study with two groups, one using FZ-VIS and one using scripted configuration plus standard metric and QoI scripts, and compare task completion time, number of configurations attempted, and the QoI preservation of the final choice. If the FZ-VIS group is not faster or does not end with comparable-or-better configurations, the framework's central value claim fails.","supporting_citations":[{"cited_title":"Underwood, V","cited_arxiv_id":null,"evidence_quote":"Supplies the unified compressor configuration interface that FZ-VIS builds on for interactive pipeline composition."},{"cited_title":"Soler, M","cited_arxiv_id":null,"evidence_quote":"Provides the standard lossy-compression evaluation framework whose metric-focused scope FZ-VIS extends with QoI analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the FFCz correction algorithm the spectral module uses to preserve the power spectrum under frequency error bounds."},{"cited_title":"Liang, K","cited_arxiv_id":null,"evidence_quote":"Defines the modular prediction-based compressor whose predictor, quantizer, encoder, and lossless options are explored in the case studies."}],"review_version":1}