{"id":"8324f5c0-9b6b-4322-bc08-b7ddc020d5aa","arxiv_id":"2504.15007","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Radiologists' eye movements differ slightly between real and AI-generated chest X-rays, especially for longest and shortest fixations, but none of the differences are backed by inferential statistics.","lead":"This paper measured where 16 radiologists look when viewing real versus AI-generated chest X-rays, and reports small differences in gaze behavior. It is an early empirical step toward understanding how synthetic medical images might affect doctors' perception, but the study lacks statistical tests to support its claims.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Paper asserts 'significant differences' in gaze between real and fake X-rays, but reports only descriptive statistics; with 16 participants the observed differences are within sampling error, so the central claim is unsupported.","rationale":"The reader's weakest assumption focused on pooling across heterogeneous radiologists, which is a genuine concern (experience is known to affect visual search). However, the more immediate and load-bearing flaw is the absence of any inferential statistics connecting the observed descriptive differences to the claim of significance. The paper explicitly uses the word 'significant' in the abstract, Section 3.4, and Section 5, yet no test of statistical significance is performed. This is not a matter of disagreeing with field consensus; it is an internal validity problem: the central conclusion cannot be evaluated from the reported evidence. The proposed permutation test addresses both the lack of hypothesis testing and the need to respect subject-level clustering. If the test fails to reject, the paper would need substantial revision, including effect sizes and confidence intervals, before its central claim could be accepted. Thus the reader's REJECT verdict remains appropriate, though I would emphasize the missing inferential framework over the pooling issue as the primary reason.","tokens_in":9439,"tokens_out":3470,"duration_ms":32452,"concrete_test":"Perform a participant-level permutation test on the reported metrics. For each of the 16 radiologists, randomly permute the real/fake labels of their trials, then recompute the group-level differences for the key statistics: mean and median saccade amplitude (Table 1) and mean longest and shortest fixation durations (Table 2). Build the null distribution of these differences under exchangeability across subjects. If the observed differences (e.g., 0.22 degrees, 16.6 ms) fall within the central 95% of the null distribution, the claim of significant differences is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract and conclusion, is that radiologists exhibit significant differences in gaze behavior between real and AI-generated chest X-rays. Yet Sections 3.4 and 3.5 provide only descriptive statistics (means, medians, standard deviations) and saliency metrics (CC, KL, SIM) without any p-values, confidence intervals, effect sizes, or hypothesis tests. For example, Table 2 shows longest-fixation mean times of 575.9 ms (fake) versus 559.3 ms (real), a difference of 16.6 ms with standard deviations of roughly 250 ms; Table 1 shows saccade-amplitude means of 5.71 vs 5.93 degrees. With only 16 radiologists, such small differences are plausibly explained by between-subject variability. Moreover, the analysis ignores the hierarchical structure of eye-tracking data: thousands of fixations and saccades are nested within participants and stimuli, so the effective sample size for the real-versus-fake comparison is at most 16, and treating events as independent would overstate precision. Without inferential statistics, the claim of 'significant differences' is not supported by the reported analyses.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an eye-tracking study in which 16 radiologists viewed real and AI-generated (RoentGen) chest X-rays, with the goal of detecting shifts in gaze behavior between the two image types. The authors analyze saccade amplitude and direction distributions, fixation durations for first, last, longest, and shortest fixations, and spatial bias maps quantified by saliency metrics (correlation coefficient, KL divergence, similarity). The abstract and conclusion claim that radiologists exhibit significant differences in visual behavior between real and fake images, particularly in longest and shortest fixations and in saccadic amplitude distributions.","tokens_in":9603,"tokens_out":2762,"duration_ms":25461,"significance":"If the central claim were properly supported, this would be a useful contribution to the growing literature on AI-generated medical images and their effect on expert visual search, with implications for radiologist training, image generation quality, and regulatory guidance. The paper's strengths are its novel research question, the construction of a paired real/fake chest X-ray dataset using matched reports, and the descriptive characterization of multiple gaze features including joint saccade distributions and temporal fixation subtypes. However, the study currently provides only descriptive statistics and aggregate saliency metrics, with no inferential statistics, so the headline claim of 'significant differences' is not established. The hierarchical structure of the data is also ignored, making the reported aggregate comparisons difficult to interpret at the population level.","major_comments":[{"comment":"The central claim of 'significant differences' in gaze behavior is not supported by any inferential statistics. Tables 1 and 2 report only means, medians, standard deviations, and ranges; Table 3 reports saliency metric values. No p-values, confidence intervals, effect sizes, or hypothesis tests are provided. For instance, the longest-fixation mean difference in Table 2 is 575.9 ms (fake) versus 559.3 ms (real) with standard deviations of roughly 250 ms, and the saccade-amplitude difference in Table 1 is 5.71 vs 5.93 degrees with standard deviations near 4.6 degrees. With 16 participants, these differences are plausibly within sampling error, so the word 'significant' in the abstract, Section 3.4, and Section 5 is unjustified as written.","section":"Abstract, §3.4, §3.5, §5"},{"comment":"The analysis pools thousands of fixations and saccades across participants and stimuli without accounting for the hierarchical structure of eye-tracking data. Fixations are nested within participants and within images, so the effective sample size for the real-versus-fake comparison is at most 16 radiologists (or the number of distinct stimuli). Treating each event as an independent observation would inflate precision. The manuscript should either use a mixed-effects model with random intercepts for participants and images, or a participant-level summary analysis, before drawing conclusions about radiologists as a population.","section":"§3.3, §3.4"},{"comment":"The saliency metrics comparing real and fake bias maps are presented without any uncertainty quantification or significance testing. The values (e.g., CC = 0.1938 for Longest, CC = 0.5158 for First) are descriptive comparisons of group-aggregated maps, but no confidence intervals or permutation tests are provided. The conclusion in Section 5 that 'the alignment weakens significantly' is therefore not supported by the reported analyses. Additionally, the interpretation of what constitutes a meaningful difference in CC, KL, or SIM is not defined.","section":"§3.5, Table 3"},{"comment":"The study pools radiologists with heterogeneous experience levels (2 with 0–5 years, 6 with 6–10 years, 4 with 10–20 years, 3 with over 20 years) and different subspecialties. Because expertise is known to affect visual search behavior, the pooled aggregate statistics may mask subgroup differences or be driven by a few individuals. The paper should report whether the observed patterns are consistent across experience levels or justify pooling, for example by including experience as a covariate or performing a subgroup analysis.","section":"§3.3, §3.4"}],"minor_comments":[{"comment":"The text says '30 reports being randomly selected from this set' but does not specify how many unique images were shown to each participant; please clarify whether each radiologist viewed all 30 real and 30 fake images in a single trial and how image order was randomized.","section":"§3.2.2"},{"comment":"The caption lists '(a+b)' three times; the third pair should refer to subplots (e) and (f) for the joint distributions.","section":"Figure 2"},{"comment":"In the bullet on shortest fixations, the text states 'quick glances are more common in real images,' but the descriptive statistics show only that the mean and median are slightly higher; this interpretive claim goes beyond the presented data.","section":"§3.4.4"},{"comment":"Several references are incomplete or inconsistent in formatting (e.g., [8] lacks publication venue details, [21] lacks page numbers, [33] is listed as both a journal article and an arXiv preprint). Please harmonize the bibliography to the journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an extended version of a workshop paper (ETRA 2025 GenAI Workshop). For a journal submission, the absence of any inferential statistics is a serious gap, but it is fixable with appropriate mixed-effects modeling or permutation-based testing. The research question is timely and the dataset construction is a useful contribution. I would encourage the editor to invite a revision that addresses the statistical support for the central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The topic is genuinely new: nobody has looked at radiologists' eye movements when viewing real versus AI-generated chest X-rays, and the paper does a careful job of collecting and describing that data. The eye-tracking setup is sound, the stimuli generation via RoentGen from MIMIC-CXR reports is sensible, and the descriptive breakdown by first, last, longest, and shortest fixations is clearly presented. I also think the internal pattern they report—that first and last fixations are more similar between real and fake, while longest and shortest fixations diverge—is interesting and could be a real effect worth testing properly.\n\nThat said, the central claim is not supported. The abstract and conclusion say \"significant differences,\" but there are no p-values, confidence intervals, effect sizes, or any inferential tests anywhere. Table 2 shows longest-fixation means of 575.9 ms for fake versus 559.3 ms for real, with standard deviations around 250 ms; Table 1 shows saccade amplitude means of 5.71 versus 5.93 degrees. With 16 observers, these differences are within the range of between-subject variability. The saliency metrics (CC, KL, SIM) are computed on pooled group-level maps of 16 people, so they also carry no inferential weight. The paper's own text sometimes hedges with \"slightly\" or \"suggests,\" which is honest, but the abstract and conclusion overstate.\n\nThe weakest assumption is pooling 16 radiologists with very different experience levels and subspecialties. Experience is known to change search behavior, and with one trial per participant per image, a few outliers could drive the group-level patterns. This matters because the analysis ignores the hierarchical structure of the data entirely.\n\nThere is no circularity or fitting here—the measurements are direct and the behavior is observable. But the absence of inferential statistics is a load-bearing flaw, not a minor omission, because the entire point of the paper is to claim a real shift in gaze behavior.\n\nWho is this for? Researchers in medical image perception, generative model evaluation, and human-AI interaction. They will find the descriptive patterns useful as preliminary evidence, but the paper needs revision before its headline can be taken seriously. A serious referee should engage with it and request proper mixed-effects models or permutation tests, and ideally public release of the gaze data.\n\nMy bottom line: it deserves peer review, but only with the understanding that the central claim currently rests on descriptive patterns and needs much stronger statistical support.","headline":"A genuinely first look at how radiologists' gaze differs between real and AI-generated chest X-rays, but the headline claim of 'significant differences' rests on descriptive statistics only and, with n=16, the observed effects could easily be sampling noise.","tokens_in":10265,"tokens_out":1737,"would_cite":false,"duration_ms":17960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Radiologists' gaze gives away AI-generated chest X-rays","keywords":["eye tracking","medical imaging","gaze behavior","AI-generated images","chest X-ray","fixation bias maps","saccades","radiologist visual search"],"falsifier":"Recomputing every metric for each radiologist individually would settle it: if only a few participants show the longest and shortest fixation shift or the saccade-amplitude change, or if the differences vanish after matching images for low-level difficulty, the claim of a general gaze shift on AI-generated X-rays fails.","tokens_in":9237,"feed_emoji":"🩻","tokens_out":5271,"duration_ms":45994,"temperature":0.7,"pith_summary":"This paper asks whether expert radiologists look at real and AI-generated chest X-rays differently, and reports that they do: gaze statistics diverge most at the extremes of viewing time, in longest and shortest fixations, while first and last fixations behave nearly identically. The authors build a paired set of real and synthetic chest X-rays generated from the same report text, record eye movements of sixteen radiologists, and compare saccade amplitudes, directions, fixation durations, and spatial bias maps. If the finding holds, gaze behavior is a measurable signal of how synthetic images alter visual search and cognitive load, with consequences for AI model evaluation, radiology training, and safety guidelines. The study is a new application of established eye-tracking methodology rather than a new theoretical claim.","feed_headline":"Radiologists' gaze gives away AI-generated chest X-rays","feed_subtitle":"Longest and shortest fixations shift on synthetic images, a behavioral marker radiologists reveal without knowing it.","key_machinery":"The mechanism is scanpath analysis of eye-tracking data, organized into fixation bias maps and saccade distributions. Fixations are defined by velocity and acceleration thresholds from a head-mounted eye tracker, and the analysis isolates four temporal landmarks, first, last, longest, and shortest fixations, plus saccadic amplitude, direction, and their joint distribution. The synthetic stimuli are produced by a latent-diffusion model conditioned on the same radiology report text as the real images, creating matched real-fake pairs. Spatial comparison uses saliency metrics, correlation coefficient, KL divergence, and similarity, between real and fake bias maps. This machinery turns a hard-to-verbalize perceptual difference into quantitative distributions that can be compared.","core_discovery":"The paper's central claim is that radiologists' visual search patterns shift measurably when they view AI-generated chest X-rays, and the shift is concentrated in the temporal extremes of attention rather than in initial engagement. Mean and median saccadic amplitudes are smaller for fake images while the maximum amplitude is larger, which the authors read as more cautious, uncertain scanning; longest fixations are slightly longer on fake images and shortest fixations slightly longer on real images, and the spatial bias maps for these two conditions correlate only about 0.19 to 0.20, versus 0.48 to 0.52 for first and last fixations, with higher KL divergence and lower similarity. First and last fixations are similar across image types, suggesting that entry and exit attention are captured in the same way, while prolonged and brief attention differ. The authors conclude that fake images may demand more scrutiny or present distinct visual features that require different viewing strategies.","pith_inferences":["Because the sixteen radiologists span very different experience levels, one extension not pursued here is computing the same metrics per radiologist; the pooled bias maps might hide a subgroup whose gaze differs sharply from the rest.","Gaze signals could be repurposed as a passive detector of AI-generated images: if fixation extremes are reliably different, a radiologist's pattern of eye movements might flag suspicious images even when the radiologist cannot consciously identify them.","A direct testable extension would link the gaze differences to diagnostic accuracy on the same images, asking whether longer longest fixations on fake images coincide with correct rejection or with confusion."],"forward_implications":["Radiologists' initial and final attention can be expected to behave the same on real and AI-generated chest X-rays, so tools that rely on entry-point gaze may transfer between the two image types.","Longest and shortest fixation patterns form a behavioral marker for synthetic-image processing, which could be used as a gaze-based check of whether generated images elicit natural search behavior.","AI-generated images that provoke larger maximum saccades and smaller mean saccades may impose higher cognitive load, and generative models could be tuned to reduce this gap.","Gaze statistics can complement image-level realism scores when evaluating whether synthetic medical images are acceptable for clinical workflows.","The paired real-fake dataset, with matched report text, provides a reusable stimulus set for eye-tracking studies of AI-generated medical images."],"supporting_citations":[{"why":"Supplies the latent-diffusion generator that creates the synthetic chest X-rays from report text.","marker":"[42]"},{"why":"Supplies the real chest X-rays and matching radiology reports that form the real stimuli and condition the generated ones.","marker":"[43]"},{"why":"Provides the saliency metrics (correlation coefficient, KL divergence, similarity) that quantify real-fake bias map divergence.","marker":"[46]"},{"why":"Defines the eye-tracker calibration and the velocity and acceleration thresholds that segment fixations from saccades.","marker":"[41]"},{"why":"Shows that gaze feedback improves nodule detection, grounding the premise that radiologists' eye movements carry diagnostic information.","marker":"[30]"},{"why":"Shows that observing experts' search behavior improves performance, supporting the relevance of gaze patterns to radiology skill.","marker":"[31]"},{"why":"Provides the saccade model used to contextualize the joint amplitude-direction distributions.","marker":"[44]"}],"fun_headline_variants":["Doctors' gaze extremes reveal AI chest X-rays","Extreme gaze shifts mark AI-generated chest X-rays","Shortest and longest fixations shift on fake X-rays","AI X-rays alter only the extremes of doctor gaze"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pooled analysis assumes that sixteen radiologists with different experience levels and subspecialties share one consistent gaze response to synthetic images, and that the group-level maps do not hide a minority driving the differences.","fun_headline_variants_meta":{"raw":{"variants":["Doctors' gaze extremes reveal AI chest X-rays","Extreme gaze shifts mark AI-generated chest X-rays","Shortest and longest fixations shift on fake X-rays","AI X-rays alter only the extremes of doctor gaze"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001601,"raw_usage":{"total_tokens":6332,"prompt_tokens":854,"completion_tokens":5478,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":5414}},"tokens_in":470,"tokens_out":5478,"duration_ms":34761,"temperature":1.0,"reasoning_tokens":5414,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:34:53.961617+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recomputing every metric for each radiologist individually would settle it: if only a few participants show the longest and shortest fixation shift or the saccade-amplitude change, or if the differences vanish after matching images for low-level difficulty, the claim of a general gaze shift on AI-generated X-rays fails.","supporting_citations":[{"cited_title":"Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports,","cited_arxiv_id":null,"evidence_quote":"Supplies the real chest X-rays and matching radiology reports that form the real stimuli and condition the generated ones."},{"cited_title":"Saliency benchmarking made easy: Separating models, maps and metrics,","cited_arxiv_id":null,"evidence_quote":"Provides the saliency metrics (correlation coefficient, KL divergence, similarity) that quantify real-fake bias map divergence."},{"cited_title":"The eyelink toolbox: eye tracking with matlab and the psychophysics toolbox,","cited_arxiv_id":null,"evidence_quote":"Defines the eye-tracker calibration and the velocity and acceleration thresholds that segment fixations from saccades."},{"cited_title":"Computer-displayed eye position as a visual aid to pulmonary nodule interpretation,","cited_arxiv_id":null,"evidence_quote":"Shows that gaze feedback improves nodule detection, grounding the premise that radiologists' eye movements carry diagnostic information."},{"cited_title":"Viewing another person’s eye movements improves identification of pulmonary nodules in chest x-ray inspection.,","cited_arxiv_id":null,"evidence_quote":"Shows that observing experts' search behavior improves performance, supporting the relevance of gaze patterns to radiology skill."},{"cited_title":"Saccadic model of eye movements for free-viewing condition,","cited_arxiv_id":null,"evidence_quote":"Provides the saccade model used to contextualize the joint amplitude-direction distributions."}],"review_version":1}