{"id":"32ea4208-88cb-48cc-a900-a46fcc221caf","arxiv_id":"2604.02220","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.5,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Composable visual decoding operators measured on PDF/CDF tasks predict bias and variance of scatterplot mean estimation under a pre-registered project-twice mean strategy with no parameters fit to the target data.","lead":"The paper shows that visualization reading can be broken into reusable perceptual operators whose measured bias and variance compose to predict performance on new chart-task pairs without refitting. This offers a path to empirical visualization research whose findings transfer beyond the exact conditions tested.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified; the transfer claim is supported by zero-parameter predictive match under pre-registered composition.","rationale":"The paper's strongest claim is an existence proof of compositional transfer: operators isolated on one chart family generate calibrated, zero-parameter predictions of both bias and variance on a structurally different chart \times task pair. The reader's weakest assumption correctly identifies a modeling incompleteness (no slope-dependent or aspect-ratio terms), but that incompleteness is already acknowledged in the text and is not required for the existence proof to hold. The successful strategy match under the simpler operators already demonstrates that findings can compose and predictions can extend beyond the measured conditions. No internal inconsistency or hidden free parameter appears in the transfer pipeline. Therefore the CONDITIONAL verdict with high confidence remains appropriate; no adjustment is warranted.","tokens_in":22642,"tokens_out":522,"duration_ms":5836,"concrete_test":"Re-fit the hierarchical ProjectToAxisY model of Section 4.2.1 / 5.4.1 after adding a first-order local-slope term (or an aspect-ratio covariate) to the multiplicative variance; recompose the project-twice mean strategy on the same Moritz stimuli with zero free parameters. If the predictive density overlays of Figure 14 remain calibrated (observed responses still fall inside the 68 % intervals across the four variability conditions), the transfer claim is robust to the missing correction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (multiplicative visual-angle variance scaling and Normal/Weibull families estimated on 600\times450 PDF/CDF charts remaining valid without slope or aspect-ratio corrections on Moritz scatterplots) is real but does not undermine the central claim. Section 4.2.1 explicitly notes that models condition variance on projection distance but not local slope, calling slope-dependent first-order error propagation a natural future refinement. Yet the existence proof only requires that the learned operators, when composed, produce calibrated posterior predictive distributions of bias and variance on the held-out task with no parameters fit to the response data. One of the six strategies (project-twice mean) does exactly that (Figure 14); the five alternatives fail distinguishably. Because the successful strategy already matches observed spread and location under the uncorrected operators, residual misspecification of slope or aspect ratio is not load-bearing for the feasibility claim. Sample size, non-pre-registered exclusions in Study 1, and post-hoc factorial completion are secondary and already flagged by the reader.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes visual decoding operators—perceptual primitives with estimable bias and variance in visual-angle space—as a composable unit of analysis for quantitative visualization interpretation. Using PDF and CDF charts, it isolates five operators (horizontal/vertical projection, highest-point, max-slope, bisect-area) via hierarchical Bayesian models and reports a sensor-fusion account of PDF median estimation. It then reuses the learned ProjectToAxis_Y operator, without refitting, under six composition strategies to predict Moritz et al.’s scatterplot mean-estimation responses (different chart type, size, aspect ratio, and analytic goal). One strategy (project-twice mean) matches observed bias and variance; five alternatives fail in distinguishable ways. The authors present this as an existence proof that empirical findings in visualization can compose and transfer beyond the conditions in which they were measured.","tokens_in":22936,"tokens_out":1231,"duration_ms":20608,"significance":"If the result holds, the paper supplies a concrete alternative to channel rankings and task taxonomies: units that carry quantitative error profiles and compose into falsifiable predictions for untested chart × task pairs. Strengths that support this contribution include hierarchical Bayesian estimation with posterior predictive and PIT-ECDF checks, virtual-chinrest calibration into visual angle, open materials and analysis scripts on OSF, and a transfer evaluation that fits no parameters to the held-out response data. The sensor-fusion finding for PDF median is an additional explanatory contribution. Even as a scoped existence proof on position encodings, the work reframes how empirical visualization research can accumulate reusable primitives rather than isolated ordinal comparisons.","major_comments":[{"comment":"§5.3.1 and footnote 4: the pre-registered strategies were project-once×mean, project-once×median, and project-twice×sensor-fusion. The strategy that captures both bias and variance (project-twice×mean; Fig. 14 top) was added when the full 2×3 factorial was completed post hoc. The abstract and §5.4.3 frame the result as arising under a pre-registered analysis plan. Please restate the confirmatory vs. exploratory status of each cell explicitly in the abstract, §5.4, and discussion, and qualify claims of pre-registered success so they attach only to the three pre-specified strategies (none of which fully matches both location and spread).","section":"§5.3.1, footnote 4, abstract, §5.4.3"},{"comment":"§4.1.3 and §4.1.4: Study 1 recruited 24 participants and excluded 8 under criteria that the paper states were not pre-registered (eye-to-screen distance <20 cm or Pearson r <0.5), leaving n=16 for hierarchical models with participant-level bias and scale parameters. Study 2 applied the same criteria after pre-registration and excluded none. The transfer claim is the paper’s load-bearing result; please report sensitivity of the Study 1 operator posteriors (and of the composed predictions that depend on them) to the exclusion rule, or justify why the non-preregistered exclusions do not materially affect the operator parameters used in §5.","section":"§4.1.3–4.1.4, §5.4.1"}],"minor_comments":[{"comment":"§4.2.1 notes that projection variance is conditioned on distance but not local slope, with slope-dependent first-order propagation left as future work. A short forward reference in §5.4.3 would help readers see that residual aspect/slope misspecification is acknowledged and not claimed to be resolved by the existence proof.","section":"§4.2.1, §5.4.3"},{"comment":"Fig. 14 and §E.2.2–E.2.3: project-once strategies produce substantially narrower predictive intervals. A one-sentence quantitative summary (e.g., interval width relative to observed SD) in the main text would make the distinguishable failures easier to assess without the appendix.","section":"§5.4.2, Fig. 14"},{"comment":"Table 1 and operator naming: the manuscript mixes prose names (ProjectHorizontally, HighestPoint, BisectArea) with ad-hoc symbols. A single consistent operator notation table early in §3 would reduce cognitive load when reading the composition equations in §5.3.1.","section":"§3, Table 1, §5.3.1"},{"comment":"Several figure captions and in-text references use placeholder or garbled tokens (e.g., visual-angle equations and operator names rendered as boxes). Ensure the camera-ready version restores readable math and labels throughout Figs. 3–5 and 7–14.","section":"Figs. 3–5, 7–14"},{"comment":"§6.1’s four validity conditions for an operator are useful; consider stating them earlier (end of §3) so readers can evaluate the five operators against those criteria as they are introduced.","section":"§3, §6.1"}],"recommendation":"minor_revision","confidential_remarks":"The science is solid for an existence-proof paper; the main risk is overclaiming pre-registration for the winning cell. If the authors cleanly separate confirmatory and exploratory strategy cells, this is appropriate for TVCG. Sample sizes are modest but typical for carefully instrumented perception work with hierarchical models; I would not reject on n alone given the zero-parameter transfer design."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real news is the transfer result. They isolate five operators on PDF/CDF tasks, fit hierarchical Bayesian error models in visual-angle space, then compose the projection operator under six candidate strategies to predict Moritz et al. scatterplot mean estimation. No parameters are fit to the Moritz responses. Project-twice mean captures both location and spread; the other five fail in distinguishable ways. That is a genuine existence proof for compositional, predictive empirical work in visualization, not another ranking table.\n\nWhat is new is the unit itself: chart-agnostic decoding operators that carry estimable bias and variance and compose. Prior channel rankings, Amar-style taxonomies, MA-P, ACT-R graph models, and Talbot’s slope-ratio work do not give you reusable quantitative primitives that transfer without refitting. The inverse-MSE sensor-fusion account of PDF median (mode + area bisection) is a nice secondary result and fits better than a mixture. Methods are careful: virtual chinrest, hierarchical models, posterior predictive and PIT-ECDF checks, public OSF materials, and a pre-registered transfer plan.\n\nSoft spots are real but secondary. Sample sizes are modest (n=16 after exclusion in study 1, n=20 in study 2). The first-study exclusion rule was not pre-registered. They completed the 2\times3 strategy factorial after looking at the three pre-registered cells; they are transparent about it and the pipeline is deterministic, but it is still post-hoc. Projection variance is scaled only by distance, not local slope or aspect ratio; they flag this themselves as future work. None of these overturn the zero-parameter match on the held-out task.\n\nMath and data look solid; citations are appropriate and not self-serving. This is for people who care about theory accumulation and predictive design tools rather than one-off effectiveness rankings. It deserves a serious referee. I would bring it to reading group and expect to cite the transfer result.","headline":"Clean existence proof that chart-agnostic operators with quantitative error can be isolated, composed, and transfer zero-parameter to a different chart\times task; one of six strategies matches bias and variance.","tokens_in":23526,"tokens_out":494,"would_cite":true,"duration_ms":5491,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Visual decoding operators with measurable bias and variance can be composed to predict perception on new charts and tasks without refitting.","keywords":["Visualization Theory","Visual Decoding Operator","Composable Models","Sensor Fusion","Perceptual Effectiveness","Graphical Perception","Hierarchical Bayesian Modeling"],"falsifier":"Hold operator parameters fixed from the isolation tasks and apply them to mean estimation on scatterplots (or other held-out chart × task pairs) of new sizes, aspect ratios, or mark types; if no composition strategy produces predictive distributions whose bias and variance match observed responses, or if the previously successful project-twice mean strategy systematically mispredicts under those changes, the transfer claim fails.","tokens_in":23532,"feed_emoji":"📊","tokens_out":992,"duration_ms":22120,"temperature":0.7,"pith_summary":"This paper claims that rankings of visual channels and taxonomies of tasks cannot predict how accurately people will read a new chart for a new goal, because those units lack quantitative structure that composes. It proposes instead that quantitative chart reading is sequences of reusable perceptual operations—projecting a point to a curve or axis, finding a peak, judging the steepest slope, bisecting an area—each with bias and variance measured in visual-angle units. Using PDF and CDF charts, the authors isolate five such operators and characterize them with hierarchical Bayesian models; when a task admits more than one operator, inverse-MSE sensor fusion fits better than a simple mixture. They then recompose projection operators under six candidate strategies to predict mean estimation on scatterplots that differ in type, size, aspect ratio, and analytic goal, with no parameters fit to the new responses. One strategy (project-twice mean) captures both bias and variance of observed answers; the other five fail in distinguishable ways. The result is an existence proof that empirical visualization findings can transfer beyond the conditions in which they were measured.","feed_headline":"Operators predict chart reading on new tasks without refitting","feed_subtitle":"Five perceptual primitives compose into mean estimates that match bias and variance on unseen scatterplots.","key_machinery":"Visual decoding operators: chart-agnostic perceptual primitives (project-to-curve, project-to-axis, highest-point, max-slope, bisect-area), each modeled as a distribution with estimable bias and variance in visual-angle space so that multi-step tasks accumulate error traceably when operators are sequenced.","core_discovery":"Individually measured visual decoding operators can be composed under candidate strategies to generate calibrated posterior predictive distributions of both bias and variance for a structurally different chart × task pair—scatterplot mean estimation—with no parameters fit to the response data. Of six strategies, project-twice mean captures observed bias and variance; the alternatives fail in distinguishable ways. That transfer constitutes an existence proof that findings in graphical perception can compose rather than requiring a new experiment for every pairing.","pith_inferences":["Axis-reading and color-extraction operators would compose directly with the projection operators already measured and are the natural next primitives for the library.","The same composition rules could serve as cost functions inside automated visualization recommenders that optimize predicted decoding error rather than channel orderings alone.","If variance scales with projection distance only approximately when local slope is ignored, adding first-order slope-dependent terms should tighten out-of-sample predictions without changing the operator concept.","Classic channel-ranking studies could be re-cast as operator-isolation experiments so historical results become reusable parameters instead of ordinal tables."],"forward_implications":["Design tools can enumerate candidate visualizations, decompose each into an operator sequence, and report predicted error distributions rather than only ordinal channel rankings.","Once operator-level perceptual error is known as a baseline, remaining error can be attributed more cleanly to non-perceptual sources such as task mistranslation or elicitation artifacts.","A fuller operator library would let empirical findings transfer across chart types and display contexts without a new experiment for every visualization × task pairing.","When multiple operators apply to one task, inverse-MSE weighting can model how viewers integrate strategies rather than selecting a single one.","A failed composed prediction points to which specific operator needs revision instead of leaving an undifferentiated residual."],"fun_headline_variants":["Composed decoding ops predict scatterplot means with no refit","Five perceptual ops transfer to match bias and variance on new charts","Visual operators compose to forecast novel chart-task performance","Reusable decoding primitives predict unseen mean estimation without fit","Chart-agnostic ops generate calibrated posteriors for new scatterplots"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the projection error models—especially multiplicative scaling of variance with visual-angle distance—learned on one set of PDF and CDF charts remain valid without slope or aspect-ratio corrections when the same operators are applied to scatterplots of different size and shape.","fun_headline_variants_meta":{"raw":{"variants":["Composed decoding ops predict scatterplot means with no refit","Five perceptual ops transfer to match bias and variance on new charts","Visual operators compose to forecast novel chart-task performance","Reusable decoding primitives predict unseen mean estimation without fit","Chart-agnostic ops generate calibrated posteriors for new scatterplots"]},"model":"grok-4.5","effort":"low","cost_usd":0.0055,"raw_usage":{"total_tokens":1530,"prompt_tokens":824,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":55000000,"prompt_tokens_details":{"text_tokens":824,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":640,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":824,"tokens_out":66,"duration_ms":5268,"temperature":1.0,"reasoning_tokens":640,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T13:55:07.042117+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold operator parameters fixed from the isolation tasks and apply them to mean estimation on scatterplots (or other held-out chart × task pairs) of new sizes, aspect ratios, or mark types; if no composition strategy produces predictive distributions whose bias and variance match observed responses, or if the previously successful project-twice mean strategy systematically mispredicts under those changes, the transfer claim fails.","supporting_citations":[],"review_version":1}