{"id":"dd86f39a-8536-4e0d-b6ce-6cb4f8a0ea5d","arxiv_id":"2608.00433","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In two preregistered experiments with 1,380 participants, line-based multiple forecast visualizations beat confidence intervals at communicating the shape of forecast distributions, while linewidth and opacity weighting did not clearly hurt distribution perception.","lead":"This paper ran two large online experiments on how people read climate forecast charts. It found that showing individual forecast lines instead of shaded confidence bands helps readers see the true shape of forecast uncertainty, and that weighted lines can encode extra information without clearly damaging that perception.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Q1/Q2 response options are generated as MFV-style line sets, so the main MFV-over-CI advantage may be a response-format matching artifact, a confound the authors acknowledge in §7.2.","rationale":"The reader's weakest assumption is the same one I find most load-bearing, and the manuscript's own §7.2 limitation supports it. However, this does not warrant rejection: the experiments are preregistered, the OSF materials are public, the Bayesian analyses are reported with credible intervals, and many contrasts (e.g., Uniform, Q1 H1A) show large effects in the predicted direction. The problem is specifically that the dependent measure is not symmetric between conditions. A targeted replication with matched response formats can settle it. Until then, the central claim should be treated as conditional. I also note that the Experiment 2 claim that weighting 'preserves' distribution perception rests on null contrasts (H4: 1/60 credible positive; H5 not supported); this is a secondary weakness that an equivalence test or Bayes factor for absence of effect would address, but it does not supersede the response-format confound for the paper's main comparison. Since the reader's verdict is already CONDITIONAL, no adjustment is needed.","tokens_in":23434,"tokens_out":4826,"duration_ms":42222,"concrete_test":"Run a preregistered follow-up of Experiment 1 with response-option format as a between-subjects factor: half of participants receive the current MFV-style options (15 lines; strip plots for Q2) and half receive CI-style options (e.g., median line with 95% and 50% shaded bands for Q1; interval/quartile summaries for Q2), for the same four data distributions and four visualization conditions. Test the visualization × response-format interaction in the Bayesian multinomial model. The central claim requires the MFV-over-CI accuracy advantage to remain credibly positive when response options are CI-style; if the advantage shrinks or reverses in that arm, the format-matching account is confirmed and the headline conclusion must be scaled back to 'MFV aids identification when the response format matches MFV.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that Q1/Q2 measure distribution perception in a format-neutral way. In §4.5, all five response options for Q1 are generated as MFV-style displays (15 forecast lines), and Q2 options are strip plots of 15 values; the task instructions tell participants these are 'simplified representation of the forecast values.' Participants in the CI conditions therefore compare a shaded interval/band to a set of individual lines, while participants in the MFV conditions compare lines to lines. This visual format matching can inflate MFV accuracy independently of any genuine perceptual advantage. The authors explicitly flag this in §7.2: 'we designed the response option for Q1 to have a similar shape to MFV, which might lead participants to use format-matching and advantage the MFV condition.' Since the central claim (MFV more perceptually accurate than equivalent CIs) rests on the H1A/H1C contrasts measured only through this MFV-shaped response format, the confound is load-bearing; if format matching is at work, the main result could be a measurement artifact rather than evidence about the visualizations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether multiple forecast visualizations (MFV) communicate underlying forecast distributions more accurately than confidence interval (CI) plots, in the context of climate forecast projections. Two preregistered online experiments are reported. Experiment 1 (480 participants) compares CI 95, CI 95+50, MFV 22, and MFV 9 across four data distributions (Uniform, Normal, Bimodal, GCM) on three tasks: overall distribution identification (Q1), single-year distribution identification at 2099 (Q2), and upper-bound judgment (Q3). Experiment 2 (900 participants) tests weighted MFV 9 variants using linewidth and alpha opacity, relative to a no-weighting baseline. The paper reports that MFV 22, and to a lesser extent MFV 9, lead to more accurate distribution identification and fewer 'Normal' misclassifications than CI plots, while weighted MFV 9 displays generally preserve distribution perception. The authors conclude that MFV affords a more perceptually accurate foundation for communicating both the distribution and the differences of multiple forecasts.","tokens_in":23671,"tokens_out":4727,"duration_ms":42044,"significance":"If the findings hold, the work is practically significant: it offers an evidence-based alternative to CI-based displays for forecast communication, which is relevant to climate reporting, weather prediction, and decision support. The study is well-designed in many respects: both experiments are preregistered, materials and analysis scripts are publicly available, the Bayesian multilevel modeling is appropriate for the multinomial outcome data, and the paper reports contrasts with credible intervals rather than relying solely on p-values. The paper also takes seriously the trade-off between clutter and distributional information by testing a downsampled MFV 9 variant. The main empirical claims are plausible and the authors explicitly acknowledge several limitations, including the response-format matching concern. However, as detailed below, the central claim rests on a measurement confound that needs to be addressed before the conclusion can be accepted at face value.","major_comments":[{"comment":"The central claim that MFV is more perceptually accurate than CI is supported primarily by Q1 and Q2, but the response options in both tasks are generated as MFV-style displays: Q1 options are 15 forecast lines and Q2 options are strip plots of 15 values. Participants in CI conditions therefore compare a shaded interval or band to collections of individual lines, while MFV participants compare lines to lines. This format-matching confound affects all H1A–H2D contrasts; if participants are using low-level visual similarity between the stimulus and response options, the observed MFV advantage could be a measurement artifact rather than evidence about the relative perceptual accuracy of the visualization formats. The authors acknowledge this in §7.2 ('we designed the response option for Q1 to have a similar shape to MFV, which might lead participants to use format-matching and advantage the MFV condition') and list it as a limitation, but the limitation is load-bearing for the paper's main conclusion. To support the claim, the paper needs a format-neutral outcome measure (e.g., a task in which response options are matched to the stimulus format for each condition, or a task that does not rely on visual analogies between the stimulus and the answer set), or a supplemental analysis demonstrating that the MFV advantage persists when format-matching is controlled.","section":"§4.5 and §7.2"},{"comment":"The no-weighting baseline in Experiment 2 is not a concurrent control: the paper states that 'the baseline group of Experiment 2 (no weighting) was taken from Experiment 1.' This means that comparisons between weighting conditions and baseline in H4–H6 are not based on random assignment within the same experiment; they could be confounded by experiment session, recruitment differences, attrition, or subtle differences in administration. This is particularly relevant because the key positive claim of Experiment 2 is that weighting 'preserves' distribution perception (i.e., no harmful effect versus baseline). A null difference against a non-concurrent baseline is weak evidence of preservation. The authors should either analyze Experiment 2 only with participants drawn from the same session, or provide a sensitivity analysis and explicit justification for why the cross-experiment baseline is valid.","section":"§4.6 and §6.1"}],"minor_comments":[{"comment":"The label 'SSP-8.5' appears to be a typo; it should read 'SSP5-8.5'.","section":"Figure 4"},{"comment":"The text contains garbled superscript text such as '15 90' that appears to be a formatting artifact; please clean these instances.","section":"§4.6"},{"comment":"H2C and H2D predict similarity between CI and MFV for Normal distributions, but the analysis uses a difference-based categorical regression with no equivalence testing or ROPE criterion. The 'Not supported' interpretation in Table 2 is therefore not a proper evaluation of a similarity hypothesis; consider using an explicit equivalence region or acknowledge that the data only show a directional difference.","section":"Table 2 and §4.8"},{"comment":"The power-transformation exponent of 1.6 used for consensus weights is stated but not justified; a brief rationale or citation for this value would help reproducibility.","section":"§4.3"},{"comment":"The phrase 'no consistent proof for H4' is informal; consider 'no consistent evidence for H4'.","section":"§6.1"}],"recommendation":"major_revision","confidential_remarks":"The reader's stress-test concern about response-format matching is real and is acknowledged by the authors; I agree that it is load-bearing for the main claim. The paper has many strengths—preregistration, open materials, Bayesian modeling, careful reporting of contrast uncertainty—but the main conclusion would be significantly stronger if the authors could report a format-neutral analysis or a control condition that addresses the matching confound. The cross-experiment baseline in Experiment 2 is a second issue that should be addressed. I do not think the paper should be rejected, but a revision with additional analysis (or a reframing of the central claim to be more modest) is needed. No issues with citation practice or novelty disclosure were apparent."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a well-executed, honestly reported pair of experiments, but the central comparison—MFV better than CI for distribution perception—is built on a response format that looks like MFV for everyone. The authors know this and say so in §7.2, which is to their credit, but it means the abstract's confidence overshoots the evidence.\n\nWhat's genuinely new and useful: they take the MFV program beyond trust and judgment into distribution perception, using realistic CMIP6-derived stimuli with controlled shapes (Uniform, Normal, Bimodal, GCM). Both experiments are preregistered, the Bayesian multilevel analysis is appropriate, and the OSF materials include data, analysis scripts, and stimuli. That kind of transparency is real and makes reanalysis possible. The weighting manipulations (linewidth, alpha) are also a sensible first test of whether you can layer extra attributes onto a downsampled display without wrecking perception.\n\nThe soft spots, in rough order of severity. First, the response-format confound. Q1 and Q2 options are all rendered as 15-line sets or strip plots—MFV-style representations—so CI participants have to match a shaded band to a set of lines while MFV participants match lines to lines. The paper's own limitation section concedes this could advantage MFV. For a claim this size, that's a load-bearing issue: I can't tell how much of the MFV advantage is shape perception and how much is format matching. Second, the \"weighting preserves perception\" conclusion is a null result presented as an affordance. They did not do equivalence tests, so \"no credible difference\" isn't \"no meaningful difference.\" Third, H2C-D didn't go as predicted: for Normal data, CI produced fewer normal responses than MFV, which is awkward for the normality-implies-CI story but not fatal.\n\nWho's this for? Researchers working on uncertainty visualization and forecast communication will find the methods and questions useful, and the OSF artifacts are a nice asset. The design recommendation—use MFV or weighted MFV instead of CIs for shape—needs a follow-up with matched response formats before I'd put it in front of practitioners.\n\nMy recommendation: engage with it. It deserves a serious referee and probably a major revision focused on re-running at least the key H1A/H1C contrasts with response options that don't carry the format signal. If the authors can do that and add equivalence tests for the weighting claims, this becomes a solid contribution.","headline":"Well-run, transparent preregistered study, but the headline MFV-over-CI result rests on an acknowledged response-format confound, so the design claim is weaker than the abstract suggests.","tokens_in":24184,"tokens_out":2944,"would_cite":false,"duration_ms":26563,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Showing each forecast line separately lets readers identify the true shape of a forecast distribution more accurately than confidence-interval bands do, and weighted line displays preserve that advantage.","keywords":["uncertainty visualization","multiple forecast visualization","graphical perception","line charts","confidence intervals","forecast distribution","climate forecasts","visual weighting"],"falsifier":"A replication that swaps the answer options so that CI viewers choose among shaded interval bands while MFV viewers choose among line sets, or that uses format-neutral responses such as direct estimates of variance and modality, would settle whether the MFV accuracy advantage survives without format matching; if it disappears or reverses, the paper's central claim is undermined.","tokens_in":23257,"feed_emoji":"📈","tokens_out":10035,"duration_ms":83845,"temperature":0.7,"pith_summary":"When several models produce different forecasts, communicators often compress them into a confidence-interval band. This paper argues and tests the opposite: drawing the individual forecast lines, a multiple forecast visualization (MFV), lets readers perceive the actual distribution of forecasts, while a band nudges them toward assuming a normal, bell-shaped distribution even when none exists. In two preregistered experiments using climate temperature projections, readers who saw 22 forecast lines identified the true generating distribution more accurately than readers who saw CI 95 or CI 95+50 displays, and were less likely to answer 'Normal' for non-normal data. Downsampling to 9 lines retained most of the benefit, and encoding forecast weights through line thickness or opacity did not hurt distribution perception. If the claim holds, forecast graphics used in climate, weather, and health communication could replace or supplement interval bands with weighted line sets to show both the shape of uncertainty and the differences among models.","feed_headline":"Forecast lines beat confidence bands at conveying true shape","feed_subtitle":"Shaded bands make readers assume a bell curve; individual lines, weighted or not, keep the real shape.","key_machinery":"The central object is the multiple forecast visualization (MFV): a line chart that draws each forecast trajectory as its own mark rather than collapsing the set into a summarized band. The argument runs on two perceptual mechanisms. First, preserving individual marks keeps distributional features such as clusters, gaps, multimodality, and skew visible, whereas a confidence-interval band encodes only a center and spread and visually suggests a symmetric bell shape. Second, readers use an extent heuristic, treating the furthest visible edge of the marks as the plausible upper bound, so CI bands that hide outer forecasts lead to narrower perceived ranges than full MFV. The tested variants include a 22-line full display, a 9-line downsample chosen by percentile-based selection, and weighted 9-line displays in which consensus weights computed from a kernel density estimate are mapped to linewidth or opacity, both magnitude channels.","core_discovery":"The paper's central claim is that, compared with equivalent confidence-interval plots, multiple forecast visualizations afford a more perceptually accurate foundation when the goal is to communicate both the distribution and the differences of multiple forecasts. In Experiment 1, posterior contrasts showed that MFV 22 produced higher predicted probabilities of correctly identifying the generating distribution (Uniform, Normal, Bimodal, or the original global climate model distribution) than CI 95+50 and CI 95, with the largest effect for the Uniform distribution in Q1 ($\\Delta=0.64$, 95% credible interval $[0.43, 0.81]$). CI displays elicited more 'Normal' responses when the true distribution was not normal, and MFV viewers were more likely to select the higher upper-bound point, consistent with an extent heuristic in which readers treat the furthest visible mark as the plausible upper limit. In Experiment 2, weighting a 9-line MFV by consensus-derived opacity or linewidth did not credibly improve distribution identification (hypothesis H4 not supported) but also did not harm it, and the faintest $\\alpha$ levels reduced selection of the highest upper-limit response. The authors conclude that summary-based CI displays imply normality and obscure clusters, gaps, and multimodality, whereas downsampled and weighted MFV can communicate additional forecast attributes without sacrificing readers' perception of the underlying forecast distribution.","pith_inferences":["Extension: if the normality bias of CI bands is as consistent as Experiment 1 suggests, public forecast graphics that show only a central projection plus a range may systematically overstate model consensus; replacing bands with weighted line sets could reduce the tendency to treat a single pathway as inevitable.","Extension: a direct test of the paper's acknowledged format-matching limitation would give CI viewers CI-shaped response options and MFV viewers line-set options; if CI accuracy rises to MFV levels, part of the reported advantage is response-option design rather than perceptual fidelity.","Extension: because the weights were derived from the same distribution being visualized (consensus), the null weighting results leave open whether weights encoding independent attributes, such as historical accuracy, would help or hurt identification; that is the natural next experiment.","Extension: opacity appears to be the more reliable channel for controlling perceived forecast extent, so designers who want to downweight outlier forecasts should prefer alpha over linewidth until the effect is tested at their own chart sizes and line densities."],"forward_implications":["Forecast graphics that currently use CI bands to summarize model agreement can adopt full MFV, or a 9-line downsample when clutter is a concern, to keep distribution shape visible.","CI-based displays should be understood as communicating central tendency plus spread, not an accurate rendering of the forecast distribution: they will overstate normality and hide multimodality or clusters.","Downsampling from 22 to 9 forecasts is a viable decluttering strategy, since distribution perception is largely preserved and visual channels are freed for other attributes.","Visual weighting by linewidth or opacity can encode model-level attributes such as reliability or consensus without undermining readers' perception of the overall forecast distribution.","De-emphasizing extreme forecast lines with low opacity shifts readers' upper-limit judgments, so alpha weighting can steer perceived forecast range, an effect linewidth produced less reliably."],"supporting_citations":[{"why":"It defines multiple forecast visualizations and recommends roughly nine forecasts, supplying the MFV concept and the MFV 9 design.","marker":"[41]"},{"why":"It frames the aggregation spectrum for ensemble visualization that motivates comparing less-aggregated line displays with more-aggregated interval summaries.","marker":"[49]"},{"why":"It provides prior evidence that summary visualizations can imply normal distributions, motivating hypotheses H2A-H2D.","marker":"[39]"},{"why":"It documents the extent heuristic in MFV reading that underlies Q3, H3, and the Experiment 2 upper-limit results.","marker":"[44]"},{"why":"It shows that viewers are influenced by interval boundaries in hurricane forecast displays, supporting the prediction that CIs bias distribution perception.","marker":"[53]"},{"why":"It demonstrates that summary-based representations suppress information about the frequency and arrangement of individual forecasts, which the paper extends.","marker":"[55]"},{"why":"It provides prior work on non-expert perception of uncertainty in line charts that the paper complements with its normality-bias finding.","marker":"[59]"},{"why":"It motivates weighting climate models by performance and independence, the real-world practice behind Experiment 2's weighting manipulation.","marker":"[30]"},{"why":"It provides perceptual thresholds for line attributes used to select the alpha and linewidth levels.","marker":"[57]"}],"fun_headline_variants":["Forecast lines, not bands, show true distribution shape","Weighted forecast lines preserve shape perception","Lines beat shaded bands for reading forecast distributions","Multiple forecast lines cut clutter, keep real shape","Forecast distribution: lines outperform confidence intervals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the multiple-choice answer options, which were drawn as sets of forecast lines or strip plots in the same visual language as the MFV displays, measure perception of the underlying forecast distribution fairly across all visualization conditions.","fun_headline_variants_meta":{"raw":{"variants":["Forecast lines, not bands, show true distribution shape","Weighted forecast lines preserve shape perception","Lines beat shaded bands for reading forecast distributions","Multiple forecast lines cut clutter, keep real shape","Forecast distribution: lines outperform confidence intervals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1663,"prompt_tokens":1100,"completion_tokens":563,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":716,"completion_tokens_details":{"reasoning_tokens":494}},"tokens_in":716,"tokens_out":563,"duration_ms":5087,"temperature":1.0,"reasoning_tokens":494,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:19:48.852804+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that swaps the answer options so that CI viewers choose among shaded interval bands while MFV viewers choose among line sets, or that uses format-neutral responses such as direct estimates of variance and modality, would settle whether the MFV accuracy advantage survives without format matching; if it disappears or reverses, the paper's central claim is undermined.","supporting_citations":[{"cited_title":"Knutti, J","cited_arxiv_id":null,"evidence_quote":"It motivates weighting climate models by performance and independence, the real-world practice behind Experiment 2's weighting manipulation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides perceptual thresholds for line attributes used to select the alpha and linewidth levels."}],"review_version":2}