{"id":"27da349a-633f-458c-ac29-e72da2f8b29a","arxiv_id":"2501.14327","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A five-intent taxonomy (searching, observing, traversing, comparing, exploring) with distinct gaze signatures holds for both low vision and sighted image viewers.","lead":"This study tracked the eye movements of 20 low vision and 20 sighted participants during image viewing, then used playback interviews to define five visual intents: searching, observing, traversing, comparing, exploring. The findings are an early step toward assistive tools that detect what a low vision user is trying to see and adapt support in real time.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quantitative intent differences may be confounded by segment duration: ETRTA segments are defined by verbal descriptions rather than fixed windows, and Sections 4.3-4.5 do not control for duration, so entropy and object-count effects could be artifacts of how segments were cut.","rationale":"The paper's central claim is that five bottom-up visual intents exist, have distinct gaze signatures, and are moderated by visual ability. The strongest support is the convergence of qualitative coding (Cohen's Kappa = 0.73), quantitative effects, and gaze data quality validation. However, the quantitative support inherits a segmentation loop: the same researcher who hears the retrospective description also places the segment boundaries and later labels segments from those transcripts. The absence of a duration covariate is the narrowest point where the argument could break. If intent categories differ systematically in how long a described episode lasts, the entropy and object-count measures used to distinguish them are confounded. The paper's own Section 5.1 admits there is no control over segment count or length, and Sections 4.3-4.5 never report segment lengths or include them in the model, so the concern is not speculative. I considered two other risks: information level (question type) confounding intent, and the acknowledged overlap between visual acuity and peripheral vision groups. The information-level confound is real but partially mitigated by the qualitative demonstration that each intent appears in multiple task types; the visual-ability overlap is explicitly listed as a limitation and mainly affects the moderation sub-claim. The segment-duration confound attacks the core distinct-gaze-patterns evidence in the central claim, so it is the most load-bearing. The required check is inexpensive and would settle the issue. Because the concern is testable and the rest of the empirical contribution, including the taxonomy, qualitative examples, and fixation/saccade findings, is substantial, the reader's CONDITIONAL verdict should stand with no change in verdict category.","tokens_in":28565,"tokens_out":7118,"duration_ms":68483,"concrete_test":"","verdict_should_be":"UNCHANGED","load_bearing_attack":"Reader's concern about label validity is the right axis, and there is a specific mechanism that makes it load-bearing. In the ETRTA procedure (Section 3.3), segment boundaries are set by the researcher while listening to retrospective descriptions and watching the gaze playback, then refined with the participant; Section 3.4.1 states that unresolved segments were discarded. The paper itself notes in Section 5.1 that there is \"no control over how many segments they produce or how long each segment is.\" None of the models in Sections 4.3-4.5 includes segment duration or fixation count as a covariate. This matters because two of the headline outcomes are mechanically duration-sensitive: stationary entropy and number of objects visited. Short observing episodes such as \"I was identifying the ingredient on the donut\" will trivially have low entropy and few objects visited, regardless of intent. If observing segments are shorter on average, the reported \"observing stands out\" pattern is partly a restatement of how segments were cut rather than evidence about stable mental categories. Fixation duration and saccade amplitude are less vulnerable to this confound, so the taxonomy may survive, but the quantitative corroboration in Sections 4.3-4.5 is not currently clean.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a retrospective think-aloud eye-tracking study with 20 low vision and 20 sighted participants who viewed static images while answering questions at three information levels. From participants' verbal reflections and gaze replays, the authors derive a taxonomy of five visual intents—searching, observing, traversing, comparing, and exploring—shared by both groups, and they report low-vision-specific behaviors such as color-perception recalibration and visual confusion. Using fixation, saccade, spatial-entropy, and object-level gaze measures, the manuscript compares gaze behavior across intents, between low vision and sighted groups, and across visual acuity and peripheral vision subgroups. The central claim is that these five intents are associated with distinct, measurable gaze patterns and that visual abilities modulate those patterns, laying a foundation for intent-aware low vision assistive technology.","tokens_in":28813,"tokens_out":3830,"duration_ms":37493,"significance":"If the central claim holds, this is the first bottom-up, cross-task visual intent taxonomy for low vision users, and the paper's mixed-methods design is a real strength: the qualitative taxonomy is grounded in participants' own descriptions rather than imposed solely from gaze statistics, the sample is diverse in visual conditions, and the authors take care with calibration, gaze-data quality, and object-level annotations. The paper also makes concrete, falsifiable predictions about which gaze measures differ across intents and between ability groups. However, the quantitative corroboration in Sections 4.3–4.5 currently rests on segment labels whose construction and selection are not fully reported, and on gaze measures that may be sensitive to segment duration; those issues are fixable and are the main barrier to accepting the quantitative claims as they stand.","major_comments":[{"comment":"The headline quantitative contrasts—especially stationary entropy and number of objects visited—are potentially confounded by segment duration. In the ETRTA procedure, segment boundaries are set by the researcher while listening to retrospective descriptions and refined with the participant, and Section 5.1 states that there is \"no control over how many segments they produce or how long each segment is.\" If observing segments are systematically shorter (for example, single-object identification episodes), they will trivially have lower stationary entropy and fewer visited objects regardless of intent. None of the LME/ART models in Sections 4.3–4.5 includes segment duration or fixation count as a covariate. Please report the distribution of segment duration and fixation count per intent, and re-run the entropy and object-count analyses with duration control or duration-matched segment comparison. If the observing effect is reduced or disappears, the qualitative taxonomy can still stand, but the quantitative corroboration would need to be reframed accordingly.","section":"Sections 3.3, 3.4.2, 4.3–4.5"},{"comment":"The validity of the intent labels is asserted through full agreement, but unresolved gaze segments were discarded to achieve complete agreement, and the paper does not report how many segments were discarded, whether discarding was systematic with respect to intent or participant group, or inter-coder agreement on the final segment labels. The reported Cohen's Kappa of 0.73 is for the initial codebook on three sample transcripts, not for the final labeling of all 510 segments. Because the quantitative analyses in Sections 4.3–4.5 inherit these labels, selection bias in segment labeling could produce the reported gaze differences. Please report the number and proportion of discarded segments, a reliability analysis of the final segment labels, and a sensitivity analysis that includes or models the unresolved segments.","section":"Section 3.4.1"},{"comment":"The statistical inference is based on many correlated gaze measures and numerous post-hoc contrasts, but the paper does not state the total number of tests performed or apply a family-wise or false-discovery-rate correction across measures. At least one headline result is only a trend (fixation rate, p=0.067), and the effect-size confidence intervals reported for the LME results (e.g., eta-squared CI [0.30, 1.0]) are implausible as printed and need correction. Please clarify how many hypotheses were tested, apply an appropriate correction at the measure level, and correct the effect-size reporting.","section":"Sections 4.3–4.5"},{"comment":"The visual ability analysis is weakened by overlap between the low-acuity and peripheral-vision-loss groups: Section 5.5 states that six participants had both conditions. Section 4.5 nevertheless reports main effects of visual acuity and peripheral vision without reporting the cell sizes of the 2x2 ability grouping or a sensitivity analysis that excludes the overlapping group. Given the modest sample size (20 low vision participants), these effects should be interpreted with more caution, and the paper should either report the 2x2 cell sizes or explicitly model the overlap to show that the saccade-amplitude and object-count effects are not driven by the same six participants.","section":"Sections 4.5 and 5.5"}],"minor_comments":[{"comment":"\"Stardgart's disease\" should be \"Stargardt's disease.\"","section":"Section 2.3"},{"comment":"The 20/100 visual acuity threshold is introduced without justification; since it is used to split the low vision group, please cite the prior source more clearly and, if possible, report sensitivity to the threshold.","section":"Section 3.4.4"},{"comment":"The figure caption and the text describe examples with participant IDs, but several panels (e.g., 4j) are not explicitly referenced in the prose; please ensure every panel is mentioned or state that some panels illustrate an intent collectively.","section":"Section 4.2, Figure 4"},{"comment":"No data or analysis-code availability statement is included. Given the central quantitative claims, making at least aggregate data and analysis scripts available would substantially strengthen reproducibility.","section":"Overall"},{"comment":"The phrase \"The mean angular errors from the 5-dot validation was...\" should be \"The mean angular error... was...\" for subject-verb agreement.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and timely empirical study for ASSETS, and I do not see grounds for rejection: the taxonomy is grounded in qualitative data and the quantitative concerns are addressable with additional analyses. The main risk is that the current quantitative corroboration is not yet clean enough to support the strong claim that the five intents have distinct, measurable gaze patterns. I would encourage the editor to request the duration-control and label-reliability analyses as part of the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This is the first cross-task visual intent taxonomy built bottom-up from low vision users' retrospective descriptions, and that is a genuine contribution. But the quantitative gaze contrasts in Sections 4.3–4.5 are partly confounded by segment length and the labeling procedure, so treat the qualitative taxonomy as the main result and the gaze statistics as suggestive rather than clean.\n\nWhat the paper does well: bottom-up method, 20 low vision and 20 sighted participants, a sensible range of image contexts and question types, standard qualitative coding with substantial intercoder agreement, and quantitative gaze measures triangulated with the qualitative labels. The five intents—searching, observing, traversing, comparing, exploring—are plausible and connect to prior task-specific categories. The low-vision-specific behaviors (color recalibration, boundary scanning with peripheral field loss) are vivid and useful for design. Prior work predefined intents for one task or studied sighted users only; this fills that gap.\n\nSoft spots, in order of severity. First, the segment-duration confound. Segmentation is based on retrospective descriptions (Section 3.3), and Section 5.1 explicitly says there is no control over how many segments participants produce or how long each segment is. No model in Sections 4.3–4.5 includes segment duration or fixation count as a covariate. Observing is defined by concentrated attention; shorter observing segments would trivially yield lower stationary entropy and fewer objects visited. Fixation duration and saccade amplitude are less vulnerable, so the taxonomy probably survives, but the entropy and object-count differences are not currently clean evidence. Second, label validity: labels come from participants' retrospective think-aloud and researcher segmentation, and unresolved segments were discarded (Section 3.4.1). That is a defensible method, but it is not ground truth. Third, multiple post-hoc comparisons are not globally controlled, and the visual acuity and peripheral-vision groups overlap (six participants have both), which the authors acknowledge in Section 5.5. Fourth, no data or materials are released, so independent reanalysis is hard.\n\nThe citation pattern is fine; the authors build on their own prior gaze work and cite relevant task-specific intent recognition literature. The paper is honest about limitations. This deserves a serious referee; the issues are addressable in revision. I'd want a covariate analysis or segment-length-matched comparison before fully trusting the quantitative claims. Who it's for: accessibility/HCI researchers working on gaze-based assistive tech and anyone designing intent-aware support for low vision users.","headline":"A solid bottom-up taxonomy paper; the qualitative core is new and useful, but the quantitative gaze contrasts are partly confounded by segment length and the labeling procedure.","tokens_in":29329,"tokens_out":2743,"would_cite":true,"duration_ms":25573,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that low vision and sighted viewers share five in-the-moment visual intents—searching, observing, traversing, comparing, exploring—and that each leaves a distinct gaze signature modulated by visual acuity and peripheral…","keywords":["eye tracking","low vision","visual intent","gaze behavior","retrospective think-aloud","visual acuity","peripheral vision loss","assistive technology"],"falsifier":"Fresh coders blind to the original labels could re-segment the same gaze recordings from the retrospective think-aloud sessions; if their five-intent labels do not agree with the paper's segments at better than chance, the gaze differences attributed to intents are an artifact of the labeling procedure rather than a stable property of visual intent.","tokens_in":28399,"feed_emoji":"👁","tokens_out":9692,"duration_ms":81068,"temperature":0.7,"pith_summary":"The paper tries to establish that during static image viewing, people's in-the-moment visual goals fall into five shared categories—searching, observing, traversing, comparing, and exploring—and that each category carries a measurable gaze signature. The authors derive this taxonomy from the bottom up using eye tracking with retrospective think-aloud on 20 low vision and 20 sighted participants, rather than starting from pre-defined intents. They show that gaze behavior differs across the five intents, that low vision participants scan more broadly than sighted participants during traversing and exploring, and that lower visual acuity and peripheral vision loss shorten saccades and increase the number of objects visited. If the classification is right, it gives intent-aware assistive technology a first cross-task vocabulary for recognizing what a low vision user is trying to see, and it says plainly that such recognition should combine gaze with visual ability and image context.","feed_headline":"Low vision and sighted viewers share five gaze intents","feed_subtitle":"Intent-aware assistive systems could read what a person is trying to do from eye movements, visual ability, and image context.","key_machinery":"The central object is the visual intent, defined as a distinct gaze pattern that reflects an in-the-moment, meta-level objective and is agnostic to the overall visual task. The procedure that carries the argument is the eye-tracking-based retrospective think-aloud protocol: participants viewed images while their gaze was tracked, then watched a playback of their own gaze trajectory and described what they were doing, and researchers segmented and labeled each trajectory into intent segments, discarding unresolved segments so the final set had complete labeling agreement. On those segments the authors computed fixation duration and rate, saccade amplitude, stationary entropy $H_s = -\\sum_i \\pi_i \\log_2 \\pi_i$ over an $8\\times5$ grid of areas of interest, number of distinct objects visited, object attention variability, and foreground attention ratio, and compared them across intents, between low vision and sighted groups, and across visual abilities. These comparisons are what bind each intent label to a measurable gaze signature and what connect visual acuity and peripheral vision to those signatures.","core_discovery":"On the paper's own terms, the central discovery is a five-category visual intent taxonomy shared by low vision and sighted viewers. Searching is a sequence of fixations aimed at locating a target; Observing concentrates fixations on one object or person to identify identity, details, or activity; Traversing moves fixations across adjacent objects, typically for counting or reading; Comparing shifts fixations back and forth between two or more objects to judge relationships; Exploring spreads fixations widely to gather overall context. The quantitative analyses tie these categories to distinct gaze signatures: observing shows longer fixations, shorter saccades, lower stationary entropy, and more uneven object attention; searching and exploring spend more time on background; comparing has longer saccades than traversing and more uneven attention. The paper further reports that low vision participants scanned more broadly during traversing and exploring, and that low visual acuity and peripheral vision loss were associated with shorter saccades and a larger number of objects visited, which the authors read as compensatory scanning. It also documents low-vision-specific gaze behaviors outside the five intents, such as briefly fixating a high-contrast object to restore confidence in color perception and confusion-driven irregular revisits after misidentifying an object.","pith_inferences":["A natural next experiment is to recruit balanced groups with acuity loss only and peripheral field loss only; because the paper's low vision groups overlap on both conditions, the separate effect of each visual ability remains provisional.","The 'palette cleanser' behavior—briefly fixating a high-contrast object to restore confidence in color perception—suggests a testable assistive design: detect gaze returning to a high-contrast region during a low-contrast judgment and offer color or contrast enhancement at that moment.","If the taxonomy is stable, a held-out classifier using gaze plus image context should label new low vision users' gaze segments with the five intent labels; the paper does not build such a model, so an accurate classifier would be the strongest confirmation of the taxonomy's usefulness.","The difficulty of separating searching, traversing, and exploring on low-level gaze measures alone suggests these categories may be graded or context-dependent rather than discrete, and a future recognition system might score them as soft labels instead of exclusive states."],"forward_implications":["An assistive system could use the taxonomy to choose support on the fly: magnify the fixated object during observing, highlight the relevant objects during comparing, and preview content beyond the visual field when a user with peripheral vision loss scans toward the boundary.","Visual intent recognition models for low vision users should be trained with visual ability and image context as inputs, because searching, traversing, and exploring are not reliably separated by gaze metrics alone.","The retrospective think-aloud protocol is feasible with low vision users but is too time-intensive and produces too variable segment lengths to build large training sets; standardized single-intent trials will be needed for model training.","The taxonomy is a starting point for static 2D viewing without assistive tools; dynamic content, real-world 3D scenes, screen magnifiers, and contrast enhancements are likely to introduce new intents and gaze behaviors such as smooth pursuit.","Gaze-based support must be personalized to the user's visual profile: peripheral vision loss and low visual acuity change how many objects a person visits and how far their saccades travel, so sighted gaze norms are not a safe default."],"supporting_citations":[{"why":"Supplies the accessible gaze calibration and visual field test procedure used with low vision participants, and the reading-gaze findings this study extends.","marker":"[95]"},{"why":"The only prior intent-aware low vision aid, and the source of the visual acuity threshold and design context that this paper generalizes beyond reading.","marker":"[94]"},{"why":"Represents prior predefined task-intent recognition and the standardized single-intent data-collection protocol the discussion proposes for future model training.","marker":"[98]"},{"why":"Introduces the eye-tracking retrospective think-aloud method that the study adapts for low vision participants.","marker":"[12]"},{"why":"Validates stimulated retrospective think-aloud against eye tracking, supporting the reliability of the labeling procedure.","marker":"[30]"},{"why":"The velocity-based eye-movement classification algorithm used to generate fixations for the quantitative gaze measures.","marker":"[16]"},{"why":"The image segmentation model whose masks define the salient objects used for object-level gaze measures like number of objects visited.","marker":"[72]"},{"why":"Provides the standard definitions of fixations, saccades, and stationary entropy used in the analysis.","marker":"[59]"},{"why":"The standard letter acuity charts used to measure participants' visual acuity at study intake.","marker":"[21]"},{"why":"The dispersion-based real-time fixation detection algorithm used in the gaze display and segmentation interface.","marker":"[43]"}],"fun_headline_variants":["Five visual intents tie low vision and sighted gaze","Eye tracking reveals 5 shared gaze intents for low vision","Low vision gaze: 5 intents from eye tracking","Gaze patterns decode visual intents in low vision users","Five intents describe low vision and sighted viewing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the intent labels placed on gaze segments from the retrospective think-aloud procedure are valid ground truth; if those labels do not correspond to stable mental categories, the gaze differences reported between intents could be an artifact of labeling and segmentation.","fun_headline_variants_meta":{"raw":{"variants":["Five visual intents tie low vision and sighted gaze","Eye tracking reveals 5 shared gaze intents for low vision","Low vision gaze: 5 intents from eye tracking","Gaze patterns decode visual intents in low vision users","Five intents describe low vision and sighted viewing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000332,"raw_usage":{"total_tokens":1877,"prompt_tokens":1006,"completion_tokens":871,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":790}},"tokens_in":622,"tokens_out":871,"duration_ms":7294,"temperature":1.0,"reasoning_tokens":790,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:14:16.274683+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fresh coders blind to the original labels could re-segment the same gaze recordings from the retrospective think-aloud sessions; if their five-intent labels do not agree with the paper's segments at better than chance, the gaze differences attributed to intents are an artifact of the labeling procedure rather than a stable property of visual intent.","supporting_citations":[],"review_version":1}