{"id":"434fbf0e-86f1-4494-a934-ef4ed6a92734","arxiv_id":"2507.17024","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"No single fast method reproduces free-response visualization takeaways, but combining ranking and rating methods approximates some affordances, and GPT-4o only aligns with humans on salience ratings.","lead":"This paper compares four ways of asking people what they take away from a chart, and tests whether a large language model can imitate those answers. Ranking and rating methods only partly match free-response findings, so researchers should combine methods and treat LLMs cautiously.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's central claim—that combinations of ranking and rating methods can serve as an effective proxy—is never directly tested; Studies 2–4 are evaluated only individually, and the proposed combination in Sec. 10 lacks a defined ensemble or benchmark comparison.","rationale":"The reader's conditional verdict already flags that the combined-proxy claim is overstated, but their stated weakest assumption is the completeness and validity of the five-factor taxonomy. I agree that the taxonomy is a real limitation, since all structured methods are closed under those five factors and the free-response coding also uses them. However, the most load-bearing gap for the paper's strongest claim is that no combination is ever formed or tested: the abstract promises that combinations can serve as an effective proxy, yet the empirical sections report only per-method results, several of which contradict each other (e.g., Shape is rare in free response but highest-ranked in Study 3; Study 4 finds no chart-specific interaction). The proposed test would settle this directly by defining an ensemble and scoring it against the Study 1 benchmark. Because the paper is otherwise carefully designed and the individual results are informative, I would not reject it; the conditional verdict stands, but the condition should explicitly require a direct evaluation of the proposed combination.","tokens_in":21207,"tokens_out":5334,"duration_ms":62164,"concrete_test":"Use the existing response-level data from Studies 1–4 to construct a concrete ensemble. For example, predict each of the 45 charts' Study 1 factor-elevation labels from Study 3 first-rank proportions and Study 4 mean salience ratings per factor and chart type, using a pre-specified combination rule, and compare the ensemble's precision and recall against all significant standardized residuals from Study 1 (Fig. 4). The central claim holds only if the ensemble recovers the full affordance pattern—at least heatmap→Clusters and line→Small Trends without assigning a spurious affordance to dot plots—and if it outperforms each individual method on a held-out set of datasets. If no reasonable ensemble achieves this, the abstract should be revised to say that individual methods selectively match free response, not that combinations are an effective proxy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim—that combinations of ranking and rating methods can serve as an effective proxy at a broad scale—is not supported by any direct test in the manuscript. Studies 2–4 are analyzed individually, and each has a distinct failure mode: Study 2 (Rank Charts) rankings are dominated by a line-chart preference correlated with familiarity (Sec. 6.2); Study 3 (Rank Conclusions) finds Shape ranked highest across all chart types even though Shape was the least common free-response factor (Sec. 7.2); and Study 4 (Rate Salience) reports no significant chart type × factor interaction (Sec. 8.2). The Discussion (Sec. 10) suggests that salience ratings and conclusion rankings 'could triangulate' the free-response affordances, but no ensemble, weighting, or decision rule is specified and no comparison of such a combination against the Study 1 benchmark is reported. Thus the central claim rests on an inference from partial overlaps rather than on evidence that a combination actually approximates the free-response affordance pattern. A related but secondary limitation is that all structured methods present conclusions drawn only from the five-factor taxonomy in Sec. 4.3, so any combination would inherit the taxonomy's boundaries. These are not internal inconsistencies; they are gaps between the claim and the evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares four methods for eliciting visualization affordances (free response, chart ranking, conclusion ranking, salience rating) across line charts, dot plots, and heatmaps, and additionally evaluates GPT-4o as a human proxy. A preliminary study derives a five-factor taxonomy of takeaways (Points, Small Trends, Shape, Large Trends, Clusters). Study 1 establishes free-response benchmark affordances: heatmaps afford Clusters, line charts afford Small Trends, dot plots afford Shape. Studies 2–4 evaluate the structured methods; each partially aligns with the benchmark but shows distinct biases (familiarity-driven line-chart preference, inflated Shape rankings, and absence of chart-specific interaction, respectively). The GPT-4o case study finds salience ratings most closely match human patterns. The abstract claims that combinations of ranking and rating methods can serve as an effective proxy, but no combination is actually tested.","tokens_in":21430,"tokens_out":7344,"duration_ms":69606,"significance":"If its central claim were supported, this paper would provide a valuable, cost-effective alternative to labor-intensive free-response coding in visualization affordance research. The empirical work has notable strengths: a priori power analyses, a large participant sample (N≈1,350 across studies), double-coding with κ=0.73, chi-square residual analyses, and a systematic human–LLM comparison. The proposed chart-specific affordances (heatmaps–Clusters, line charts–Small Trends) are concrete and falsifiable. However, the main contribution is undermined by the untested combination claim and the unvalidated, underreported taxonomy; the paper is a solid descriptive comparison but not yet a demonstration of the proxy method.","major_comments":[{"comment":"The abstract's central claim—that combinations of ranking and rating methods can serve as an effective proxy at a broad scale—is never directly tested. Studies 2–4 are analyzed individually, and each has a distinct failure mode: Study 2 rankings are dominated by a line-chart preference correlated with familiarity (Sec. 6.2); Study 3 ranks Shape highest across all chart types even though Shape was the least common free-response factor (Sec. 7.2); Study 4 reports no significant chart type × factor interaction (Sec. 8.2). Section 10 only suggests that salience ratings and conclusion rankings \"could triangulate\" patterns, without specifying an ensemble, weighting, or decision rule, and without comparing such a combination against the Study 1 benchmark. The inference from partial overlaps to an effective combined proxy is therefore unsupported. Please either soften the claim to a hypothesis or add an explicit post-hoc combination analysis—for example, a decision rule that aggregates first-ranked conclusions from Study 3 and top-rated factors from Study 4 per chart type and evaluates agreement with Study 1 factor frequencies.","section":"Abstract; Sec. 10"},{"comment":"The five-factor taxonomy from Sec. 4.3 is load-bearing: it defines the coding scheme for Study 1 and the conclusion options in Studies 2–4 and the GPT-4o prompts. Yet the factor analysis is underreported: no factor loadings, eigenvalues, BIC values, or model-comparison statistics are provided, and the choice of the five-factor model is justified only as a \"balance of attributes.\" The taxonomy is not validated against any external coding scheme or independent dataset. The authors acknowledge in Sec. 11 that the factors \"may not fully reflect the range of interpretations,\" but this limitation is more central than presented: if the taxonomy misses or merges affordance categories, the benchmark and all method comparisons inherit that bias, and the \"proxy\" claim is only about the five-factor space. Please report the factor-analysis details or validate the taxonomy independently, and reframe the contribution accordingly.","section":"Sec. 4.3; Sec. 11"}],"minor_comments":[{"comment":"The sentence \"About 97% of human conclusions were accurate, as opposed to 63% of GPT-4o conclusions contained\" is grammatically incomplete; please rewrite (e.g., \"contained inaccuracies\").","section":"Sec. 9.2.1"},{"comment":"The chi-square value reported for the GPT-4o free-response analysis (χ² = 46.3, df = 8) is identical to the value reported for human responses in Sec. 5.2; this appears to be an error and should be verified and corrected.","section":"Sec. 9.2.1"},{"comment":"The parenthetical in \"Based on the five factors from the preliminary study (Sec. 4.3, we created...\" is missing a closing parenthesis; Sec. 2.1 also contains the typo \"scattplots.\"","section":"Sec. 6.1; Sec. 2.1"},{"comment":"The embedded figure text contains garbled spacing (e.g., \"Shap e\", \"P oints\", \"S m all T ren d s\"); please ensure the final PDF text is correctly vectorized.","section":"Fig. 1"},{"comment":"The phrase \"GPT-4o generated overall poorly aligned relative takeaway rankings as compared to humans\" is awkwardly worded; consider \"GPT-4o's relative takeaway rankings aligned poorly with humans' rankings.\"","section":"Sec. 9.2.3"},{"comment":"The paper would benefit from clarifying at the outset that the comparison operates within a five-factor affordance space; the current abstract's \"broad scale\" overstates the scope given the Sec. 11 limitation to three time-series chart types.","section":"Abstract; Sec. 11"}],"recommendation":"major_revision","confidential_remarks":"The paper has a clear empirical core and the comparison of four methods is useful, but the abstract's central claim overreaches the evidence. The duplicate chi-square value in Sec. 9.2.1 suggests a proofreading error that should be fixed. The taxonomy validation concern is deeper and should be addressed either by reporting more factor-analysis detail or by reframing the contribution. The paper fits the journal's scope as a methods-focused empirical study, but it requires major revision before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe bottom line: this is a careful, honest methods paper with one bad headline. If you care about visualization affordances or about using LLMs as stand-ins for human readers, you'll want to know the details.\n\nThe good is substantial. Four crowdsourced studies with reasonable samples and power analyses; a free-response baseline double-coded at kappa=0.73; chi-square residual analyses that pin down chart-specific affordances; and a GPT-4o case study run with the same prompts as the human studies. The main empirical pattern—heatmaps afford Clusters, line charts afford Small Trends, dot plots afford Shape/Points—shows up consistently wherever the methods manage to work. The paper is also admirably transparent about each method's failure: Study 2's line-chart familiarity bias, Study 3's 'Shape dominates everything' artifact, Study 4's null interaction. That honesty is real.\n\nThe soft spot is not honesty but the abstract: 'combinations of ranking and rating methods can serve as an effective proxy at a broad scale' is not tested. Nobody combines anything. Studies 2–4 are analyzed individually, and each has a distinct weakness. The Discussion floats the idea that salience ratings and conclusion rankings could triangulate, but there's no ensemble, no weighting, no comparison against the Study 1 benchmark. The claim is an inference from partial overlaps, not a result. Fixing this is not trivial—you'd need to define the combination and show it beats the individual methods against the free-response benchmark. The stress-test note is right on this point.\n\nA quieter but important issue: the five-factor taxonomy comes from the authors' own earlier free-response data, and then all structured methods force conclusions into that taxonomy. Any combination inherits those boundaries. The paper acknowledges this in Limitations, but it's more central than they let on. Also, the coding scheme and factor analysis live in supplements that weren't accessible; for a methods paper, that's a reproducibility gap.\n\nWho should read this: anyone designing crowdsourced affordance studies or evaluating LLMs as research participants. The GPT-4o salience-rating result is the most interesting part—it's the only method where the LLM partially aligns with humans, and that's worth discussing. The paper deserves a serious referee; the empirical work is strong enough that the overclaim can be fixed by actually testing the combination or softening the abstract. I'd send it to review, with a clear request to address the combination claim head-on.","headline":"Well-run empirical comparison with an abstract that overclaims: the 'combined proxy' is never actually combined or tested.","tokens_in":21990,"tokens_out":2222,"would_cite":true,"duration_ms":25860,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Structured chart tasks can approximate, but not replace, free-response studies of what viewers take away from charts.","keywords":["visualization affordances","elicitation methods","free response","chart ranking","conclusion ranking","salience rating","large language models","chart takeaways"],"falsifier":"A reader could test this by conducting an independent free-response study in which participants are not shown the five factors, then having independent coders label the responses without using the taxonomy. If new takeaway categories emerge (for example, uncertainty judgments or distribution comparisons), or if the coders disagree on where responses like 'revenue peaked in Year 5' fit, the taxonomy's completeness would be shown to be too narrow to support the method comparison.","tokens_in":1555,"feed_emoji":"📊","tokens_out":1501,"duration_ms":38131,"temperature":0.7,"pith_summary":"This paper asks whether cheaper, more structured tasks can replace labor-intensive free-response studies for identifying visualization affordances: the link between a chart design and the takeaways readers form. It compares four elicitation methods across dot plots, line charts, and heatmaps, and finds that no single structured method fully reproduces the affordances seen in free-response data. However, combinations of ranking and rating methods can serve as an effective proxy at a broad scale, especially for heatmaps (Clusters) and line charts (Small Trends). A case study with GPT-4o shows the model aligns best with humans on a salience-rating task but is otherwise a poor proxy, emphasizing that method choice changes the affordances you will observe.","feed_headline":"Cheaper chart tasks only partly capture what viewers take away","feed_subtitle":"No structured method beats free-response for affordances, but ranking plus rating works broadly for heatmaps and line charts.","key_machinery":"The load-bearing object is the five-factor taxonomy of chart takeaways: Points, Small Trends, Shape, Large Trends, and Clusters. This taxonomy was derived from exploratory factor analysis of participants' free-response conclusions in a preliminary study, and it drives the entire comparison: stimuli are built from it, coding in Study 1 uses it, and the takeaways ranked or rated in Studies 2 through 4 are constructed to represent it. If the taxonomy is incomplete or does not generalize, every subsequent method is evaluated only against its five categories, and any affordance outside those categories would be invisible to all four methods.","core_discovery":"The central claim is that while no method fully replicates the affordances observed in free-response conclusions, combinations of ranking and rating methods can serve as an effective proxy at a broad scale. Across four crowdsourced studies, free-response data showed that heatmaps afford Clusters, line charts afford Small Trends, and dot plots afford Shape or Points depending on the method. Chart-ranking tasks were biased by familiarity and overall preference for line charts, conclusion-ranking tasks partially aligned with free-response for heatmaps and line charts but inflated the salience of Shape, and salience ratings captured overall patterns like Small Trends but missed chart-specific affordances. The paper also claims that GPT-4o, prompted with the same instructions as human participants, performs best as a human proxy for the salience-rating methodology but suffers from severe constraints, including inaccuracies, low diversity, and strong default preferences, in the other three methods.","pith_inferences":["The paper's method comparison is a template that could be applied to other chart families (bar charts, scatterplots, area charts) to test whether the ranking-plus-rating proxy generalizes beyond line charts, dot plots, and heatmaps.","The factor analysis that produced the five-factor taxonomy was conducted on time-series data with a limited set of trends; an independent replication on a wider range of data shapes could either validate the taxonomy or reveal that some affordances (e.g., distributional comparisons, uncertainty judgments) are missing.","A natural extension would be to give GPT-4o the same combination task (e.g., rank-then-rate) and measure whether the combination reduces the model's default preferences, since the paper only evaluates each method in isolation.","The observed discrepancy between free-response (Shape rarely mentioned) and conclusion-ranking (Shape ranked highest) suggests that comparative tasks may systematically alter perceived affordances; this could be tested directly by asking participants to generate their own takeaways after completing a ranking task and comparing the two outputs."],"forward_implications":["If the ranking and rating combination works as claimed, researchers studying these three chart types can substitute structured tasks for free-response coding to detect broad affordance patterns at lower cost.","The familiarity bias documented in chart ranking (line charts ranked highest and correlated with self-reported familiarity) implies that ranking-based affordance signals should be interpreted cautiously when chart familiarity varies.","The salience-rating method's failure to find chart-specific interactions suggests that what is visually salient is not equivalent to what readers spontaneously take away, a distinction that matters for selecting evaluation metrics.","The GPT-4o case study implies that out-of-the-box LLMs are not yet reliable proxies for human free-response or ranking data, but their salience ratings may be usable as a coarse screening tool for candidate takeaways.","Because the five-factor taxonomy is used as the common coding scheme, any new chart type or dataset that produces a different kind of takeaway would require extending the taxonomy before the method comparison would apply."],"supporting_citations":[{"why":"Provides the chart-selection approach for capturing visualization affordances, which motivates the ranking methodology in Study 2.","marker":"[25]"},{"why":"Documents lexical and semantic ambiguities in free-response chart takeaways, justifying the search for alternative elicitation methods.","marker":"[76]"},{"why":"Defines visualization affordances and motivates the choice of dot plots, line charts, and heatmaps as canonical chart types.","marker":"[10]"},{"why":"Supplies the comparison framework and prompting parameters used in the GPT-4o case study, and prior evidence that LLM chart takeaways diverge from human ones.","marker":"[74]"},{"why":"Provides the MASSVIS targets393 dataset from which the 15 data patterns for the expanded stimuli set were derived.","marker":"[11]"},{"why":"Specifies the R package used for the exploratory factor analysis that produced the five-factor takeaway taxonomy.","marker":"[55]"},{"why":"Identifies GPT-4o as the out-of-the-box model evaluated in the case study and establishes its baseline capabilities.","marker":"[1]"}],"fun_headline_variants":["No shortcut fully captures chart takeaways; combine ranking and rating","GPT-4o works as a proxy only for salience ratings, not others","Combined ranking and rating best approximate free-response findings","Ranking plus rating as a broad proxy for chart affordances","Free response still king; ranking and rating are viable proxies"],"cache_read_input_tokens":24064,"weakest_assumption_plain":"The entire comparison depends on the five-factor taxonomy (Points, Small Trends, Shape, Large Trends, Clusters) being a complete and valid representation of the takeaways readers actually form from these charts; if the factor analysis missed or merged categories, every method inherits that bias.","fun_headline_variants_meta":{"raw":{"variants":["No shortcut fully captures chart takeaways; combine ranking and rating","GPT-4o works as a proxy only for salience ratings, not others","Combined ranking and rating best approximate free-response findings","Ranking plus rating as a broad proxy for chart affordances","Free response still king; ranking and rating are viable proxies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00065,"raw_usage":{"total_tokens":3009,"prompt_tokens":1000,"completion_tokens":2009,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":1923}},"tokens_in":616,"tokens_out":2009,"duration_ms":15850,"temperature":1.0,"reasoning_tokens":1923,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:58:02.625242+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could test this by conducting an independent free-response study in which participants are not shown the five factors, then having independent coders label the responses without using the taxonomy. If new takeaway categories emerge (for example, uncertainty judgments or distribution comparisons), or if the coders disagree on where responses like 'revenue peaked in Year 5' fit, the taxonomy's completeness would be shown to be too narrow to support the method comparison.","supporting_citations":[{"cited_title":"Xiong, V","cited_arxiv_id":null,"evidence_quote":"Documents lexical and semantic ambiguities in free-response chart takeaways, justifying the search for alternative elicitation methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the comparison framework and prompting parameters used in the GPT-4o case study, and prior evidence that LLM chart takeaways diverge from human ones."},{"cited_title":"Revelle and M","cited_arxiv_id":null,"evidence_quote":"Specifies the R package used for the exploratory factor analysis that produced the five-factor takeaway taxonomy."}],"review_version":1}