{"id":"2457e581-20e9-4d83-8ef1-aaca675b2dd2","arxiv_id":"2508.01668","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"The submission combines an attention prediction abstract with a federated learning body, so the claimed result cannot be evaluated from the provided material.","lead":"The abstract describes an AI method to predict where and when pathologists look while grading prostate cancer slides. The uploaded full text is a different paper about federated learning, so this submission cannot be reviewed as the claimed study.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The submitted body is a different paper (FedVTC, arXiv:2508.01669), so the abstract's scanpath-prediction method, experiments, and results are entirely absent and unverifiable.","rationale":"I read the submission as attempting to present an attention-scanpath prediction method for pathology, and the abstract's claim is the strongest statement in that direction. For that claim to hold, the manuscript body would need to contain the described fixation extraction algorithm, the two-stage transformer architecture, the WSI dataset from 43 pathologists, and experiments comparing the model against chance and baselines. None of these are present: the body is the complete text of a federated learning paper (FedVTC). This is not a subtle modeling assumption or a baseline weakness; it is a complete absence of the claimed method and evidence. The reader's verdict of UNVERDICTED is therefore the correct disposition, and my stress-test does not move it. I do not fully agree with the reader's chosen weakest_assumption, which focuses on whether viewport center plus magnification faithfully represents attention. That is a legitimate secondary concern that would become relevant if the method existed, but the decisive issue is that the submission contains no method or experiments at all. My concrete test would settle the concern by verifying the absence of all attention-related content in the body and confirming the body's identity with a different paper. I raise no concerns about author integrity; the issue is confined to the submission's content not supporting its stated contribution.","tokens_in":17532,"tokens_out":2131,"duration_ms":24504,"concrete_test":"Perform a term search over the submitted PDF body (excluding the abstract): 'fixation', 'scanpath', 'pathologist', 'whole slide', 'prostate', 'attention heatmap', 'magnification', 'two-stage', 'autoregressive'. If none of these terms appear in the body, the manuscript contains no evidence for the central claim. As a second step, compare the body byte-for-byte or section-by-section with arXiv:2508.01669; if the bodies match and no supplementary material accompanies the submission, then the abstract's attention-prediction method and results are not present in the submission and the claim is unverifiable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims a two-stage transformer for attention scanpath prediction on prostate whole slide images using 43 pathologists and 123 WSIs, with fixation extraction, intermediate attention heatmaps, and autoregressive next-fixation generation, and states that the model outperforms chance and baselines. The full text of this submission is the complete text of arXiv:2508.01669, 'Bridging Generalization Gap of Heterogeneous Federated Clients Using Generative Models' (FedVTC). That paper contains no mention of whole slide images, pathologists, fixations, scanpaths, attention heatmaps, magnification, or the proposed two-stage architecture. Consequently, the central claim in the abstract has no supporting method section, dataset description, experimental protocol, or result table in this manuscript. The claimed contribution cannot be checked for correctness, reproducibility, or internal consistency because none of its components are present. This is the load-bearing concern: the submission as a whole does not provide evidence for the contribution it announces. Even if the viewport-based attention surrogate were defensible, there is no body text to defend it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript announces a study that measures and predicts where and when pathologists focus their visual attention while grading whole slide images (WSIs) of prostate cancer. The abstract describes a dataset of x, y, and magnification viewport trajectories from 43 pathologists over 123 WSIs, a fixation extraction algorithm, a two-stage transformer model for static attention heatmap prediction followed by autoregressive scanpath generation, and experimental results claimed to outperform chance and baseline models. However, the supplied full text is the complete text of a different paper, arXiv:2508.01669 (FedVTC), which is about model-heterogeneous federated learning and contains no mention of WSIs, pathologists, fixations, scanpaths, attention heatmaps, or magnification. The announced method, experiments, and results are therefore entirely absent from the manuscript.","tokens_in":17679,"tokens_out":4948,"duration_ms":55584,"significance":"If the claimed contribution existed as described, it could be valuable for pathology training and for understanding expert visual search in gigapixel images, and the reported dataset of 43 pathologists and 123 WSIs would be a useful resource. However, the current submission provides no method description, experimental protocol, baseline definitions, evaluation metrics, or results for the announced task. The full text appears to be a complete, self-contained paper on a different topic; that work may have its own merits, but it does not support or even address the abstract's claims. No code, data, or machine-checked proofs are provided for the attention-prediction contribution. I concur with the stress-test concern: the central claim is unverifiable because the manuscript does not contain its supporting evidence.","major_comments":[{"comment":"The central claim that the proposed scanpath prediction model outperforms chance and baseline models is unsupported because the supplied full text is the complete text of a different paper, arXiv:2508.01669 (FedVTC), on model-heterogeneous federated learning; it contains no mention of whole slide images, pathologists, fixations, scanpaths, attention heatmaps, magnification, or the proposed two-stage architecture. Consequently, no method, dataset description, experimental protocol, baseline comparison, or result table for the announced attention-prediction task exists in this manuscript.","section":"Abstract / Full Text"},{"comment":"The abstract states that the fixation extraction algorithm 'preserv[es] semantic information' and that the scanpath targets are constructed from viewport centers, but the manuscript never defines the fixation extraction thresholds, windowing parameters, or any validation that extracted fixations correspond to actual attentional fixations; this is load-bearing because the model is trained and evaluated against these self-defined targets.","section":"Abstract (fixation extraction)"},{"comment":"The only experiments reported in the full text are FedVTC's generalization accuracy, communication cost, and memory consumption on MNIST, CIFAR10, CIFAR100, and Tiny-ImageNet (Tables 1-4); these are unrelated to where-and-when attention prediction on prostate whole slide images, so the abstract's empirical claim cannot be checked.","section":"Full Text, Section 4"},{"comment":"The two-stage transformer architecture is only named in the abstract; the body's methodology section describes variational transposed convolution for federated learning, not attention heatmap prediction or autoregressive next-fixation generation, so the proposed model's inputs, outputs, loss functions, and hyperparameters are entirely absent.","section":"Full Text, Section 3"}],"minor_comments":[{"comment":"The appendix states that large language models were used only to correct spelling and grammatical mistakes, but this statement appears to refer to the FedVTC text rather than the announced attention-prediction manuscript; a resubmission would need a statement covering the actual manuscript.","section":"Appendix C"},{"comment":"The abstract refers to 'different magnifications' but does not define the magnification levels or their relationship to WSI pyramid levels; this should be specified in any revised version.","section":"Abstract (magnification levels)"},{"comment":"The manuscript lacks a data availability statement for the viewport trajectories and any code, and it does not cite prior work on eye tracking in pathology or scanpath prediction; such details would be needed for reproducibility and context.","section":"General"}],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract and the full text is total: the body is the complete text of arXiv:2508.01669 (FedVTC). This is not a missing appendix or a formatting issue; the announced method, experiments, and results are entirely absent from the submission. I would treat this as an editorial mismatch/integrity issue and desk-reject, as the normal review process cannot assess the announced contribution because none of its components are present."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this submission is not a paper about predicting pathologists' attention. The abstract describes a two-stage transformer for scanpath prediction on prostate whole slide images, using viewport trajectories from 43 pathologists across 123 WSIs. The full text is the complete, unmodified text of arXiv:2508.01669, a federated learning paper (FedVTC) about heterogeneous clients and variational transposed convolution. The two documents share not a single method, dataset, or result. So the central claim—that the scanpath model outperforms chance and baselines—has no supporting evidence in this submission.\n\nWhat's new: I can only speak to the abstract. The idea of predicting static attention heatmaps per magnification and then autoregressively generating the scanpath conditioned on those heatmaps is a reasonable two-stage design, and the dataset size (43 pathologists, 123 WSIs) is meaningful if the data collection is real. A fixation extraction algorithm that reduces viewport trajectories to fixations while preserving semantic content is a sensible preprocessing step, though the abstract gives no details.\n\nThe soft spots are not subtle. First and foremost, the body is a different paper. There is no method section, no experiments, no baselines, no dataset description for the attention task. A referee cannot check correctness, reproducibility, or even whether the model exists. This alone makes the submission un-reviewable in its current form. Second, the abstract defines attention as viewport center plus magnification. That is a plausible proxy but not the same as eye fixation; pathologists can move their eyes within a stationary viewport, and the abstract does not validate the surrogate against eye tracking. The reader's circularity concern about deriving fixations from the same trajectories used for training is also worth noting, but it is secondary to the mismatch.\n\nWho is this for? Nobody, as submitted. The abstract might interest someone working in computational pathology or medical image training, but the body gives them nothing to evaluate. My recommendation: desk reject, and if the authors have the actual scanpath paper, they should resubmit that. Do not send this to referees; it would waste their time.","headline":"The abstract announces a pathology attention scanpath model, but the submitted body is an unrelated federated learning paper, so the claimed contribution is entirely absent.","tokens_in":18300,"tokens_out":2010,"would_cite":false,"duration_ms":20813,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-stage transformer predicts where and when pathologists look while grading prostate cancer whole-slide images, the paper claims.","keywords":["visual attention prediction","scanpath prediction","whole slide images","prostate cancer grading","digital pathology","fixation extraction","transformer","pathologist training"],"falsifier":"Record eye-tracking data from pathologists while they grade the same prostate cancer WSIs and compare gaze fixations with the fixations extracted from viewport-center trajectories. If many gaze fixations occur at positions that differ from the viewport center while the viewport is stationary, or if predicted scanpaths match viewport-derived fixations but not gaze fixations, then the model is predicting a self-defined proxy rather than expert attention.","tokens_in":17277,"feed_emoji":"🔬","tokens_out":4099,"duration_ms":51156,"temperature":0.7,"pith_summary":"The paper tries to establish that a pathologist's visual attention during prostate cancer grading can be predicted as a dynamic scanpath, not just as a static map of where experts look. It defines attention as the trajectory of viewport centers and magnification levels, compresses those trajectories into fixations with a new extraction algorithm, and trains a two-stage transformer model on data from 43 pathologists viewing 123 whole-slide images. The first stage predicts attention heatmaps at different magnifications, and the second stage autoregressively predicts the next fixation points starting from the WSI center. If the claim holds, expert-like attention becomes a learnable output that could support tools for training pathology residents.","feed_headline":"Scanpaths of pathologists predicted during prostate cancer grading","feed_subtitle":"Viewport-motion and magnification traces from 43 pathologists train a model that beats chance and baselines.","key_machinery":"The central object is the attention scanpath, represented as a temporal sequence of viewport positions (x, y, m), where m is the magnification level at which the pathologist is viewing. A fixation extraction algorithm compresses raw viewport trajectories into fixations while preserving semantic information, and this preprocessed data is what the model learns from. The predictive machinery is a two-stage transformer: the first sub-network predicts attention heatmaps as static attention across different magnifications, and the second sub-network uses these heatmaps as multi-magnification feature representations to autoregressively predict the next fixation points, starting from the WSI center.","core_discovery":"The central claim is that expert attention in whole-slide image reading is predictable at the scanpath level. The paper defines an attention trajectory as the sequence of viewport center coordinates (x, y) together with magnification m, collects this from 43 pathologists across 123 prostate cancer WSIs, and introduces a fixation extraction algorithm that simplifies each trajectory while preserving semantic information. A two-stage transformer then models the scanpath: the first stage produces attention heatmaps across magnifications, and the second stage consumes multi-magnification features from the first stage to autoregressively predict successive fixation points, beginning at the WSI center. The abstract reports that the resulting scanpath predictions outperform chance and baseline models, with the intended payoff being a training tool that helps pathology trainees allocate attention the way experts do.","pith_inferences":["If viewport trajectories are genuinely predictable, they likely encode diagnostic strategy, so predicted scanpaths could be used as additional features for automated WSI grading models.","The same approach could be tested on other cancer types or on screening tasks where search strategy, not just final diagnosis, is the object of interest.","A caution specific to the received manuscript: the full text supplied is about a different topic, heterogeneous federated learning, so the abstract's experimental claims could not be checked against the body in this version; verification requires the actual attention-prediction paper."],"forward_implications":["Digital pathology training platforms could show trainees where an expert would look next as they scan a slide.","The first-stage attention heatmaps could serve as visual summaries of diagnostically relevant regions, independent of the scanpath.","The two-stage autoregressive design could be applied to other image-heavy domains where experts navigate large images, such as radiology.","Reliable scanpath prediction could generate synthetic expert-viewing demonstrations for education when real expert time is scarce."],"supporting_citations":[],"fun_headline_variants":["AI forecasts pathologists' scanpaths in prostate cancer slides","Model predicts where and when pathologists focus on WSIs","Transformer predicts pathologist attention scanpaths in cancer slides","Predicting pathologists' view path during prostate slide review","Scanpath prediction: where will a pathologist look next?"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The approach assumes that where a pathologist moves the viewport center, plus the magnification setting, is a faithful record of where their attention is, even though eyes can move within a stationary viewport and not every viewport center is a true fixation.","fun_headline_variants_meta":{"raw":{"variants":["AI forecasts pathologists' scanpaths in prostate cancer slides","Model predicts where and when pathologists focus on WSIs","Transformer predicts pathologist attention scanpaths in cancer slides","Predicting pathologists' view path during prostate slide review","Scanpath prediction: where will a pathologist look next?"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001188,"raw_usage":{"total_tokens":4918,"prompt_tokens":976,"completion_tokens":3942,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":3862}},"tokens_in":592,"tokens_out":3942,"duration_ms":29014,"temperature":1.0,"reasoning_tokens":3862,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:28:32.372433+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record eye-tracking data from pathologists while they grade the same prostate cancer WSIs and compare gaze fixations with the fixations extracted from viewport-center trajectories. If many gaze fixations occur at positions that differ from the viewport center while the viewport is stationary, or if predicted scanpaths match viewport-derived fixations but not gaze fixations, then the model is predicting a self-defined proxy rather than expert attention.","supporting_citations":[],"review_version":1}