{"id":"2ab90b33-8715-42fd-b4d9-9cc8e4c6b7b1","arxiv_id":"2505.06702","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM agents can partially reproduce human ratings in visualization studies, and their agreement is highest for basic conclusions that expert evaluators predict with high confidence.","lead":"This paper tests whether GPT-4V-based agents can predict how humans will rate visualizations by re-running six published user studies with agents in place of participants. The agents matched humans on simple, high-confidence conclusions but missed most nuanced findings, so the authors propose using agents only to prototype experiments, not to replace user studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sec. 4.6's confidence–alignment correlation relies on a non-significant, non-independent 3/5 vs 1/12 split and should be framed as exploratory, not demonstrated.","rationale":"The paper is a careful empirical study with several real strengths: it reuses OSF materials, discloses protocol modifications (between-to-within subjects, local image batches), tests multiple models including GPT-4o, and openly discusses data pollution. I am not objecting to the negative results or the RAG/caution findings. The load-bearing issue is specifically the quantitative confidence-alignment claim, which is the basis for the abstract's 'provided that... high-confidence hypotheses' qualification. Because that correlation is derived from a small, non-independent, post-hoc-coded table, the headline result is fragile even if the individual replications are accurate. The reader's weakest assumption focused on the reliability of author-assigned alignment labels and the genuineness of expert 'pre-experiment' confidence; I agree with those concerns and add that even taking the labels and confidence at face value, the 3/5 vs 1/12 split does not clear conventional significance, so the correlation claim fails on its own reported numbers. The correct fix is not to reject the paper but to re-report RQ2 as an exploratory pattern pending an independent coding pass and a significance test; this is exactly the reader's CONDITIONAL verdict, so no change in verdict is needed.","tokens_in":22976,"tokens_out":6918,"duration_ms":68541,"concrete_test":"Run an exact test on Table 1 at the level of unique conclusions: have two coders independently label each row's H-A and C-A match (Y/N/P) from the paper's text and appendices, blind to the Confidence column, compute inter-rater agreement, deduplicate the duplicated C1.1/C3.1/C4.2 rows, then apply Fisher's exact test to the resulting high-vs-nonhigh match counts. If p < 0.05 with high agreement and with confidence values obtained from experts who were blinded to published outcomes, the correlation is supported; otherwise the central claim must be downgraded to exploratory. As a secondary check, recompute the split using only conclusions from studies published after GPT-4V's training cutoff or using held-out OSF stimuli.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Sec. 4.6/RQ2—that agent-human alignment correlates with expert pre-experiment confidence—is supported only by a post-hoc comparison of 17 author-coded rows in Table 1 (3/5 matches in high-confidence rows vs 1/12 elsewhere). This split is not statistically significant: a Fisher exact test on that 2x2 table gives a one-sided p of about 0.053, and a two-sided test is larger. The rows are not independent: C1.1, C3.1, and C4.2 appear as duplicated rows for different rating types, and all conclusions come from only five papers, so the effective sample is smaller than 17. The Y/N/P alignment labels come from a single coding pass with no inter-rater reliability, and the confidence values were assigned after the agent experiments by five experts who were not demonstrably blinded to the published outcomes; 'pre-experiment' is asserted, not verified. The correlation is also confounded with conclusion specificity: the high-confidence matches are broad, basic findings (e.g., noise affects fit, horizontal timelines preferred) that are more likely to be over-represented in GPT-4V's training data and in expert priors, while the low-confidence rows are detailed pattern claims. The paper's own data-pollution discussion (Sec. 7) acknowledges this confound for the LLM but does not apply it to the expert-confidence explanation. Therefore the abstract's 'demonstrates' overstates what the data show; at most the study suggests an exploratory correlation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether GPT-4V-based agents can simulate human ratings in visualization experiments across three studies. Study I replicates a published CHI time-series user study and finds that agent ratings can mimic human-like reasoning but fail to capture user diversity. Study II replicates six human-subject experiments from five OSF-sourced papers, introduces external expert confidence coding, and reports that agent-human alignment positively correlates with experts' pre-experiment confidence. Study III tests input preprocessing and knowledge injection, finding that these techniques can improve or bias agent ratings. The paper concludes that agents can potentially simulate human ratings when guided by high-confidence expert hypotheses, while explicitly cautioning that agents cannot replace user studies.","tokens_in":23143,"tokens_out":2910,"duration_ms":30792,"significance":"If the central claim holds, the paper makes a useful contribution to visualization evaluation methodology by identifying when LLM agents can serve as low-cost proxies in formative rating studies. The work uses external published human studies as benchmarks, avoiding circular validation, and the authors provide open code, detailed replication protocols, and candid discussions of limitations such as data pollution and modified experimental procedures. The proposed fast-prototyping scenario in Sec. 6 is a concrete, falsifiable use case. However, the significance is bounded by the small dataset (five papers, six experiments) and the fragility of the headline confidence-alignment correlation, which is supported by a small non-independent sample and coding procedures that need better verification.","major_comments":[{"comment":"The central claim that agent-human alignment positively correlates with expert confidence is supported by a single 2x2 split: 3/5 match in high-confidence rows versus 1/12 elsewhere. This split is not statistically significant (one-sided Fisher exact test yields p about 0.053, and a two-sided test is larger), and the rows are not independent because C1.1, C3.1, and C4.2 are duplicated for different rating types, while all rows come from only five papers. The abstract and conclusion state that the study 'demonstrates' this correlation; given the evidence, this is an overstatement. Please report an analysis that accounts for clustering by paper, or explicitly reframe the finding as exploratory and hypothesis-generating.","section":"Sec. 4.6, Table 1"},{"comment":"The expert confidence judgments were elicited after the agent experiments, and the paper does not describe any blinding of experts to the published outcomes, so the label 'pre-experiment' is asserted rather than verified. Furthermore, the protocol of lowering confidence to 'low' for experts who proposed opposing hypotheses is a post hoc recoding that can inflate the apparent agreement between high confidence and agent-human alignment. In addition, the Y/N/P alignment labels in Table 1 come from a single coding pass without inter-rater reliability. Please provide the full elicitation protocol, blinding details, inter-rater reliability statistics, or at minimum a sensitivity analysis that treats the recoding alternative.","section":"Sec. 4.6, Confidence Coding"},{"comment":"The data-pollution discussion in Sec. 7 acknowledges that GPT-4V's training data may inflate alignment for broad, well-known findings, but the same confound is not applied to the expert-confidence explanation. The three high-confidence matches (C1.1, C2.1, C4.1) are broad basic findings ('noise affects fit', 'horizontal timelines are preferred') that are likely overrepresented both in LLM training data and in expert prior knowledge, while the low-confidence rows tend to be detailed pattern claims. The paper does not distinguish the expert-confidence account from a training-data-familiarity account; this needs to be addressed in the revised discussion and ideally in the analysis.","section":"Sec. 7, Data pollution; Sec. 4.6"}],"minor_comments":[{"comment":"The prompt text contains a typo: 'answe' should be 'answer'.","section":"Appendix 1.1, Prompt 3"},{"comment":"The code block for the Imputation for Uncertainty execution is duplicated verbatim; one copy should be removed.","section":"Appendix 1.2"},{"comment":"The model name is written inconsistently as 'GPT-4V' and 'GPT-4v'; please standardize to 'GPT-4V' throughout.","section":"Appendix 2.5 and Sec. 5.2"},{"comment":"Several rows (C2.4, C4.3, C5.1, C5.2, C6.2) have blank entries under H-C, H-A, or C-A; please clarify whether these denote 'no hypothesis' or 'not applicable' in the table caption.","section":"Table 1"},{"comment":"The abstract says the second study 'repeated six human-subject studies' while Sec. 4.1 states that 'we identified five suitable papers'; please clarify that six experiments were drawn from five papers.","section":"Abstract and Sec. 4.1"},{"comment":"The sentence 'Our response to RQ2 suggests...' should read 'Our answer to RQ2 suggests...' for consistency with the RQ phrasing used elsewhere.","section":"Sec. 4.6, Implications"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for the journal and the authors have made a good-faith effort to use external benchmarks and open code. The main concern is that the headline claim of a positive correlation between expert confidence and agent-human alignment is built on a small, non-independent, and not statistically significant sample, and the confidence coding protocol has post hoc elements that could bias the result. The paper would be better framed as an exploratory study with clear caveats, and the abstract should not use 'demonstrates' for this correlation. I would support publication after the authors reframe the claim, report appropriate statistical analyses, and strengthen the methodological transparency of the expert-confidence elicitation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth your time, but read the abstract with a grain of salt. The headline claim—that agent-human alignment on visualization ratings correlates with expert pre-experiment confidence—is a plausible idea supported by a weak, post-hoc pattern, not by the 'demonstrates' in the abstract.\n\nWhat's actually new: it re-runs six published subjective-rating studies using GPT-4V agents and compares conclusion-level alignment, rather than measuring chart QA or literacy. That targets a genuinely underexplored feedback type. The expert-confidence coding and the RAG injection test are also new, and the paper is admirably transparent about its protocol modifications and limitations. The appendix gives prompts and code, and the data-pollution discussion in Sec. 7 is honest. These are real strengths.\n\nThe soft spots are load-bearing but not fatal. The central correlation in Sec. 4.6 rests on a 3/5 vs 1/12 split in Table 1. Fisher's exact one-sided p is about 0.053, so it's marginal even before counting the non-independence: rows for C1.1, C3.1, and C4.2 appear twice, and all conclusions come from five papers. The Y/N/P alignment labels are a single coding pass with no inter-rater reliability. The five experts were consulted after the agent runs, and nothing demonstrates they were blind to the published outcomes; 'pre-experiment confidence' is asserted, not verified. And the high-confidence matches are broad, basic findings—noise affects fit estimation, horizontal timelines preferred—which are exactly the conclusions most likely to be over-represented in LLM training data and in expert priors. The low-confidence rows are more detailed pattern claims. The paper acknowledges the data-pollution confound for the LLM but doesn't apply it to the expert-confidence explanation.\n\nSo the right reading is: an exploratory correlation that deserves follow-up, not a validated regularity. That said, the negative results—texture, magnitude judgment, open-source models—are useful, and the potential scenario for prototyping experimental parameters (Sec. 6) is a nice practical contribution.\n\nThis deserves serious peer review. I would send it out, but ask the authors to add a chance-level baseline for the match scores, report inter-rater reliability, clearly flag Sec. 4.6 as exploratory, and either blind the experts or caveat the confidence coding. With that, it becomes a solid, citable paper for visualization evaluation researchers and anyone building LLM-as-participant pipelines.","headline":"A useful empirical study of LLM agents as rating proxies, but the central expert-confidence correlation is a marginal post-hoc pattern—closer to a hypothesis than a demonstration.","tokens_in":23778,"tokens_out":3484,"would_cite":true,"duration_ms":31179,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language model agents can rate visualizations like human users, but this paper shows their agreement with human conclusions is largely limited to experiments where expert hypotheses are confident.","keywords":["large language models","visualization evaluation","human-agent alignment","user study simulation","subjective ratings","expert confidence","retrieval-augmented generation"],"falsifier":"Take the six replicated experiments and have two independent coders re-label, from the raw transcripts, whether agent feedback matches each hypothesis and each conclusion, then recompute the match-rate difference between high- and low-confidence rows; if inter-coder agreement is low or the gap disappears, the paper's central correlation does not hold. A complementary test would pre-register expert hypotheses and confidence before a new set of visualization user studies and check whether agent-human agreement rises monotonically with pre-registered confidence.","tokens_in":22682,"feed_emoji":"📊","tokens_out":9251,"duration_ms":88630,"temperature":0.7,"pith_summary":"This paper asks whether a large language model agent can take the place of human participants in rating visualizations, and it answers with a conditional yes. Across three studies, the authors replicated six published human-subject experiments by feeding the original stimuli and instructions to a multimodal language model that rated the charts on confidence, ease of use, readability, or aesthetics. The central finding is that the agent's ratings reproduce the human conclusion when external visualization experts were able to form consistent, high-confidence hypotheses about the experiment; for moderate-confidence or hypothesis-free questions, alignment drops sharply or disappears. The paper also finds that common enhancement techniques—changing image inputs, wording comparisons explicitly, or injecting web-retrieved knowledge—can shift agent ratings, but can also inject new biases. The bottom line, stated by the authors, is that agent simulation can complement and inform user studies but cannot replace them.","feed_headline":"LLM agents match human ratings only when experts are confident","feed_subtitle":"Six replicated user studies show agent ratings match human data chiefly where experts already predict the outcome","key_machinery":"The argument is carried by a comparison table in which every conclusion from the six replications carries three labels: whether the original study's hypothesis (H) matched its conclusion (C), whether the agent feedback (A) matched the hypothesis, and whether it matched the conclusion. Onto these rows the paper grafts a confidence score per conclusion, obtained by asking five external visualization experts to form their own predictions for each experiment and rate their certainty on a three-point scale, with opposed expert hypotheses downgraded to low; the five scores are then averaged and rounded. The agent runs themselves use a fixed replication protocol—the original instructions and stimuli fed to a multimodal language model, with between-subject designs converted to within-subject batches because the model cannot hold a stable rating standard across separate sessions.","core_discovery":"The paper claims that alignment between language-model agents and human raters in visualization experiments is predictable from expert confidence. In six replications, the match rate between agent and human conclusions was 3 out of 5 for high-confidence expert hypotheses, but 1 out of 12 elsewhere; the one high-confidence case that failed was a magnitude-judgement task where the agent consistently inverted the human interpretation by reading axis labels as text. The authors interpret this pattern as evidence that the agent behaves like an aggregator of historically trained knowledge: it can reproduce basic, well-established perceptual findings, but it does not possess the visual and cognitive machinery to track the subtle or contested effects that human experiments are usually run to discover.","pith_inferences":["Because the replication protocol converted between-subject designs into within-subject batches, the reported alignment may be an upper bound for how an agent would do when each condition is judged alone, as in most field use.","The high-confidence alignments are consistent with the agent reproducing patterns it has seen during training; a decisive test would run the same protocol on experiments published after the model's knowledge cutoff.","The five-expert confidence coding suggests a practical gate: pre-register expert confidence for a new visualization question and only trust agent simulation where confidence is high.","The failure to model demographic profiles and aesthetic/texture judgments implies agent ratings describe an average reader, not a population; studies in which individual differences drive ratings are the least suitable for simulation."],"forward_implications":["When experts have low confidence in a hypothesis, agent ratings are unlikely to reproduce the human conclusion, so low-confidence conditions still require human participants.","For basic, high-confidence findings—noise changes perceived fit, horizontal timelines are most readable—an agent can serve as a quick pre-check before a small pilot study.","Prompt and input choices are consequential: explicitly comparing variables or removing images pushes the agent toward textual stereotypes, so agent-based evaluations must report and validate their exact prompt configuration.","Web-retrieved knowledge injection can repair a specific reasoning error, such as aggregating icicle children, but the retrieved knowledge is uncertain in relevance and cannot yet be treated as a general fix.","If an agent already matches a validated user study, changing only experimental parameters, such as a more extreme decentering condition, can provide a preliminary preview of the new outcome and speed up iterative design."],"supporting_citations":[{"why":"Supplies the time-series experiment replicated in Study I and the conclusions C1.1/C1.2 that first show agent-human alignment.","marker":"[1]"},{"why":"Supplies the fit-estimation studies whose high-confidence conclusion C2.1 anchors the confidence-alignment pattern.","marker":"[39]"},{"why":"Supplies the uncertainty-imputation study whose confidence ratings C3.1/C3.2 test alignment under moderate and low expert confidence.","marker":"[40]"},{"why":"Supplies the timeline readability experiment whose conclusion C4.1 is the other fully aligned high-confidence case.","marker":"[10]"},{"why":"Supplies the texture aesthetics and readability studies whose low-confidence conclusions the agent fails to reproduce.","marker":"[19]"},{"why":"Supplies the magnitude-judgement study, the high-confidence case where the agent opposes the human conclusion.","marker":"[4]"},{"why":"Specifies the multimodal language model used as the rating agent throughout the replications.","marker":"[37]"},{"why":"Provides the open-sourced materials, instructions, and stimuli that make the six replications possible.","marker":"[38]"},{"why":"Motivates the web-retrieval knowledge-injection procedure tested in Study III.","marker":"[13]"}],"fun_headline_variants":["LLM raters align with humans only when experts are confident","Agent ratings match humans only for high-confidence hypotheses","Visualization agent ratings hinge on expert confidence","LLM agents fail on subtle visualization judgements","Agent-human rating alignment predicted by expert confidence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The correlation between expert confidence and agent alignment is only as strong as the hand-assigned labels in the comparison table and the experts' self-reported confidence; if those labels or confidence ratings are unreliable, the correlation may be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["LLM raters align with humans only when experts are confident","Agent ratings match humans only for high-confidence hypotheses","Visualization agent ratings hinge on expert confidence","LLM agents fail on subtle visualization judgements","Agent-human rating alignment predicted by expert confidence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1455,"prompt_tokens":934,"completion_tokens":521,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":450}},"tokens_in":550,"tokens_out":521,"duration_ms":5041,"temperature":1.0,"reasoning_tokens":450,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:34:45.815527+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the six replicated experiments and have two independent coders re-label, from the raw transcripts, whether agent feedback matches each hypothesis and each conclusion, then recompute the match-rate difference between high- and low-confidence rows; if inter-coder agreement is low or the gap disappears, the paper's central correlation does not hold. A complementary test would pre-register expert hypotheses and confidence before a new set of visualization user studies and check whether agent-human agreement rises monotonically with pre-registered confidence.","supporting_citations":[{"cited_title":"GPT-4 Technical Report, 2023","cited_arxiv_id":null,"evidence_quote":"Specifies the multimodal language model used as the rating agent throughout the replications."},{"cited_title":"Open Science Framework - OSF, 2024","cited_arxiv_id":null,"evidence_quote":"Provides the open-sourced materials, instructions, and stimuli that make the six replications possible."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the web-retrieval knowledge-injection procedure tested in Study III."}],"review_version":1}