{"id":"3ba79fea-a16c-437d-9d83-59c58d51fca9","arxiv_id":"2508.03713","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Attention patterns in chart viewing correlate with visualization literacy, enabling literacy prediction from a single attention map and literacy-aware saliency modeling.","lead":"According to the abstract, the paper links visualization literacy to eye movements and builds models that predict attention from literacy and literacy from attention. It is relevant because it could lead to visualization tools that adapt to a reader's expertise, but only the abstract could be reviewed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Supplied full text is an unrelated tax-reform paper, so the claimed 235-participant study and 86% accuracy cannot be verified from the review package; the central claim's evidence is absent.","rationale":"The reader's verdict of UNVERDICTED is appropriate because the supplied full text does not match the abstract. My stress-test identifies the same underlying problem but routes it through a more fundamental concern: the central claim's evidence is entirely absent from the review package. The reader's stated weakest assumption was about generalizability of the attention patterns to real visualization contexts; that is a legitimate secondary concern, but it presupposes that the study exists and was analyzed as claimed. The primary load-bearing issue is verification: no portion of the supplied manuscript describes the reported experiments, model training, or evaluation. Without those details, the 86% accuracy figure cannot be assessed for methodology (e.g., whether it is a held-out prediction or an in-sample fit), and the claimed superiority of Lit2Sal over state-of-the-art models cannot be checked. Therefore, I agree with keeping the verdict UNVERDICTED, and I would not change the reader's assessment. The concrete test I propose is to retrieve the actual paper and, if it matches the abstract, perform an independent re-analysis of the 86% accuracy with proper cross-validation; if it does not match, the central claim remains unsupported.","tokens_in":12005,"tokens_out":2058,"duration_ms":22884,"concrete_test":"Fetch the actual arXiv source for 2508.03713 and verify that its abstract and full text correspond. If the full text describes the 235-participant eye-tracking study with mini-VLAT, CALVI, SGL, and the Sal2Lit/Lit2Sal models, then run a second check: recompute the 86% accuracy using leave-one-participant-out cross-validation on the released data, and compare against the majority-class baseline. If the full text is the tax-reform paper, the abstract and full text are inconsistent, and the central claim is unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract asserts a strong empirical result: a 235-participant user study across mini-VLAT, CALVI, and SGL, two predictive models (Lit2Sal and Sal2Lit), and an 86% accuracy claim for literacy prediction from a single attention map. The full text supplied for arXiv:2508.03713 is an entirely different manuscript on tax reform as a constrained optimization problem, with no mention of visualization literacy, eye tracking, saliency models, or the named tests. Consequently, the review package contains no methods, no participant demographics, no eye-tracking protocol, no model architecture, no training/evaluation split, no baseline comparisons, and no raw results that could substantiate the 86% figure. This is not a dispute about generalizability—it is a failure of basic verification: the artifact under review does not contain the study it claims to report. The most load-bearing assumption of the central claim is that the study actually took place and was analyzed as described; with the supplied text, that assumption cannot be checked at all. The reader's focus on transfer to real visualization contexts is reasonable, but it presumes the study exists. The more immediate concern is that the supplied evidence is missing entirely, so the central claim is unsupported in this review package.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript under review, arXiv:2508.03713, presents an abstract claiming a 235-participant user study on visualization literacy and eye tracking, and proposes two computational models (Lit2Sal and Sal2Lit) with a reported 86% accuracy for predicting visual literacy from a single attention map. However, the full text supplied for this arXiv ID is an entirely different manuscript on tax reform as a constrained optimization problem, with no connection to visualization, eye tracking, literacy, or the named tests. The review package therefore contains none of the methods, experimental protocol, results, or analyses that would substantiate the abstract's claims.","tokens_in":12237,"tokens_out":3471,"duration_ms":35534,"significance":"If the abstract's claims were supported by an actual study, the work would be significant for personalized visualization design: predicting literacy from a single attention map could yield a fast assessment tool, and literacy-conditioned saliency could adapt visualizations to individual users. The proposed direction of accounting for individual differences in saliency models is well motivated. However, because the submission contains no supporting study, no model details, and no evaluation results, the significance cannot be assessed from this document, and the central claims remain entirely unverified. No credit can be given for reproducible code or machine-checked proofs, as none appear in the submitted text for the claimed topic.","major_comments":[{"comment":"The full text supplied for arXiv:2508.03713 is the manuscript \"Tax reform as a constrained optimization problem: a piecewise-linear framework and software implementation\" by Verhagen, Schellekens, and Garstka. It contains no mention of visualization literacy, eye tracking, saliency models, mini-VLAT, CALVI, SGL, the 235-participant study, or the Lit2Sal/Sal2Lit models. The central claims of the abstract are therefore entirely unsupported by the submitted manuscript. There is no methods section, no experimental protocol, no data analysis, no model architecture, and no evaluation results from which the reported 86% accuracy could be checked. This is a load-bearing failure: the review package does not contain the study it purports to report.","section":"Full Text (entire)"},{"comment":"Even taken on its own, the abstract's headline quantitative claim is insufficiently specified: \"86% accuracy\" is reported without any definition of the classification task (e.g., binary vs. multi-class literacy levels), the chance baseline, cross-validation scheme, or confidence interval. Similarly, the claim that Lit2Sal \"outperforms state-of-the-art saliency models\" is made without naming baselines or reporting effect sizes. These omissions would need to be addressed in any complete submission; they are noted here because the full text fails to provide them.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase \"two-way prediction\" in the title is evocative but not defined; context suggests it refers to predicting attention from literacy (Lit2Sal) and literacy from attention (Sal2Lit), which could be stated explicitly.","section":"Abstract"},{"comment":"The claim that Sal2Lit \"only takes less than a minute\" should be accompanied by the measurement procedure (e.g., one trial duration, number of stimuli) to be interpretable.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"This appears to be a case where the wrong full text was uploaded for the arXiv ID: the abstract describes a visualization literacy study, while the body is an unrelated operations-research paper on tax reform. I cannot evaluate the claimed study because its content is entirely absent. If the correct manuscript is provided, a fresh review would be needed. The editor may wish to contact the authors about the mismatch, but the current submission cannot receive a positive recommendation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take on arXiv:2508.03713. The review packet is broken: the attached full text is a completely different manuscript about tax reform optimization, not the visualization-literacy and eye-tracking study described in the abstract. That means I can only assess the abstract, and every substantive claim—the 235-participant study, the expert-focus/novice-exploration pattern, Lit2Sal outperforming SOTA, Sal2Lit's 86% accuracy—is unsupported by the materials actually supplied. This is not a generalizability dispute; it is a basic verification failure. The reader's worry about transfer to real visualization contexts is legitimate but premature.\n\nThat said, what is on the page is worth taking seriously. The abstract proposes something genuinely new: two models in the reverse direction, Lit2Sal (literacy-conditioned saliency) and Sal2Lit (attention-to-literacy prediction), and claims a practical payoff—a sub-minute literacy screen from a single attention map. If the real paper delivers those models with honest evaluation, it is a solid contribution to personalized visualization and could be useful to the HCI/vis community. The choice of three standard literacy tests (mini-VLAT, CALVI, SGL) is sensible, and the 86% figure, if accompanied by proper cross-validation and confidence intervals, would be meaningful.\n\nThe soft spots are about evidence, not speculation. No methods, no participant demographics, no eye-tracking protocol, no model architecture, no train/test split, no baseline table, no error bars. The abstract alone cannot distinguish a well-fit prediction on a held-out set from optimistic in-sample accuracy. And because the supplied full text is not this paper, I cannot check whether the authors cite prior literacy-aware saliency work properly or whether their 'overlooked' claim is fair. Those questions need the actual manuscript.\n\nRecommendation: if this is a metadata mix-up and the correct version of the paper is available, it absolutely deserves a serious peer review—send it to referees. But on this package, the right editorial move is to desk-reject or, better, ask for the correct submission. I wouldn't cite it or bring it to reading group until the real text is in hand.","headline":"The supplied full text is an unrelated tax-reform paper, so the actual visualization-literacy study is unverifiable from this package; the abstract is promising but cannot be reviewed.","tokens_in":12762,"tokens_out":2363,"would_cite":false,"duration_ms":24415,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Eye-movement patterns predict visualization literacy with 86% accuracy.","keywords":["visualization literacy","visual attention","saliency model","eye tracking","individual differences","Lit2Sal","Sal2Lit","user study"],"falsifier":"If high-literacy readers on a held-out set of visualizations, outside the three tests, show diffuse rather than focused attention, or if Sal2Lit's accuracy drops to near chance when tested on a new chart type, the claimed two-way link would fail to generalize.","tokens_in":11817,"feed_emoji":"👀","tokens_out":5429,"duration_ms":48187,"temperature":0.7,"pith_summary":"The paper tries to establish a two-way link between visual attention during data exploration and visualization literacy. Analyzing eye-movement data from 235 participants across three literacy tests (mini-VLAT, CALVI, SGL), it reports that high scorers show strong attentional focus while low scorers explore more diffusely. On that basis it builds two models: Lit2Sal predicts an observer's attention map from their literacy level, and Sal2Lit predicts literacy from a single attention map, reaching 86% accuracy in under a minute. The payoff, if the link holds, is fast literacy assessment and visualization designs that adapt to the reader.","feed_headline":"Eye-movement patterns predict visualization literacy with 86% accuracy","feed_subtitle":"Two new models link gaze patterns to chart-reading skill, making literacy screening a one-minute task.","key_machinery":"The load-bearing objects are the two named models. Lit2Sal is a visual saliency model that takes a literacy score as input and outputs a predicted attention map, operationalizing the claim that literacy shapes gaze. Sal2Lit is its inverse: it consumes one attention map and returns a predicted literacy level, which is how the paper obtains its 86% accuracy figure. The empirical foundation is the measured gaze contrast between high and low scorers on the three visualization tests.","core_discovery":"The central discovery is that attention patterns during chart reading encode visualization literacy reliably enough for two-way prediction. On three standard tests, experts' gaze is concentrated while novices scan more broadly, and this contrast serves as the signal: Lit2Sal generates the attention map a reader of a given literacy level would produce, and Sal2Lit inverts that mapping to estimate literacy from gaze. The paper reports that Sal2Lit achieves 86% accuracy from a single attention map and that the literacy-aware saliency model Lit2Sal outperforms existing saliency models that ignore literacy.","pith_inferences":["A testable extension the paper does not pursue is whether Sal2Lit transfers from the three study tests to real-world dashboards; if it does not, the 86% figure may describe task-specific gaze rather than a stable literacy trait.","The two-way link raises the possibility of using gaze as a real-time comprehension sensor, flagging when a viewer's attention wanders from the data.","If attention patterns are partly trainable, then gaze training could become an intervention to build visualization literacy, not just a way to measure it; this is a speculation beyond the paper's claims."],"forward_implications":["Literacy screening could shrink from a full test battery to a single short gaze recording, since one attention map predicts literacy with 86% accuracy in under a minute.","Visualization tools could estimate a reader's literacy on the fly and adjust chart complexity or annotation accordingly.","The expert-focus, novice-explore pattern suggests a concrete instructional target: training novices to allocate attention more selectively may raise literacy.","Saliency models that ignore literacy likely miss a systematic source of individual variation in where people look."],"supporting_citations":[],"fun_headline_variants":["Gaze alone reveals chart-reading skill with 86% accuracy","Literacy from looks: 86% accuracy from one attention map","Two-way model links gaze and visualization literacy","From gaze to literacy: 86% accuracy in under a minute","Attention patterns decode visualization literacy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the gaze patterns recorded in the lab with these three tests and 235 participants remain stable enough to predict literacy and attention outside the study.","fun_headline_variants_meta":{"raw":{"variants":["Gaze alone reveals chart-reading skill with 86% accuracy","Literacy from looks: 86% accuracy from one attention map","Two-way model links gaze and visualization literacy","From gaze to literacy: 86% accuracy in under a minute","Attention patterns decode visualization literacy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000648,"raw_usage":{"total_tokens":2945,"prompt_tokens":882,"completion_tokens":2063,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1986}},"tokens_in":498,"tokens_out":2063,"duration_ms":17300,"temperature":1.0,"reasoning_tokens":1986,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:57:54.031864+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If high-literacy readers on a held-out set of visualizations, outside the three tests, show diffuse rather than focused attention, or if Sal2Lit's accuracy drops to near chance when tested on a new chart type, the claimed two-way link would fail to generalize.","supporting_citations":[],"review_version":1}