{"id":"250803be-642f-406c-9ed9-9bfd77df8055","arxiv_id":"2607.28466","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Progressive finding-to-frame recovery from routine colonoscopy reports yields a vision–language model that outperforms biomedical encoders on retrieval, multi-centre classification, and structured report generation.","lead":"EndoCLIP turns 280,000 routine colonoscopy reports into 125,756 lesion-level image–text pairs and beats biomedical vision–language models on retrieval, six clinical tasks, and structured reporting. It shows that recovering which frame matches which finding can replace per-task image annotation with language-specified clinical targets.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The multi-lesion matching threshold and same-archive retrieval benchmark leave open whether gains reflect true finding-to-frame semantics or self-reinforced archive style.","rationale":"The paper’s strongest claim is empirical and well-scoped: progressive recovery from routine reports produces lesion-level pairs that beat strong VL baselines on retrieval, six clinical tasks, and frozen-decoder structured reporting, with a reader study on pathology. Stage ablations, the naive pairing control, multi-centre EndoVL tasks, and uncertainty on EndoReport100 are real supports. The single soft spot that still carries the claim is exactly the reader’s weakest assumption—correctness and lack of bias in model-selected pairs, especially multi-lesion matches gated by a fixed similarity threshold, evaluated on a same-institution retrieval set while source signal remains in embeddings. That does not overturn the results as reported, but it caps confidence and keeps the appropriate verdict CONDITIONAL pending pair-quality audit, threshold sensitivity, and stricter leakage/de-duplication checks. No stronger internal inconsistency is evident; I do not move to REJECT or to unconditional ACCEPT.","tokens_in":17808,"tokens_out":654,"duration_ms":14584,"concrete_test":"Have clinicians audit a stratified sample of ~500 multi-lesion pairs retained at threshold 0.28 (and at 0.25/0.30/0.35) for true finding–frame match; retrain the final stage on clinician-verified pairs only (or report precision vs threshold). Separately re-evaluate global/intra-case retrieval after removing any EndoReport100 cases with near-duplicate frames or report phrasing relative to the pretraining archive. If audited pair precision is low (<~70%) or verified-only / de-duplicated R@1 collapses toward the naive control, the load-bearing assumption fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that progressive recovery (Methods §4.4.2) produces sufficiently correct lesion-level pairs that EndoCLIP’s gains over PMC-CLIP/BiomedCLIP on EndoReport100 retrieval (global mean R@1 14.3% vs 2.8%), zero-shot/linear-probe classification, and structured reporting reflect clinical semantics rather than residual same-archive cues or circular matching. Multi-lesion pairs (~56.2k of 125.8k) are kept only when EndoCLIP-III similarity exceeds a hand-chosen threshold 0.28 on a held-out multi-lesion validation set; EndoReport100 is independent of pretraining but from the same institutional archive; Discussion notes source identity remains locally predictable in embeddings (Supp. Fig. S5, Table S18). The naive case-level control (global R@1 6.5%) shows progressive recovery helps, but does not separate true correspondence quality from style/source memorization that would also boost same-archive retrieval and transfer. If a non-trivial fraction of multi-lesion pairs are wrong or style-correlated, the headline “recovering finding-to-frame correspondence yields scalable supervision” is only partly supported.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces EndoCLIP, a CLIP-style dual-encoder foundation model for colonoscopy trained on 125,756 lesion-level image–text pairs recovered from 280,476 routine reports. Because reports summarise whole procedures rather than caption frames, the authors propose a three-stage progressive correspondence-recovery pipeline (case-level evidence localisation → single-lesion anchor selection → multi-lesion disambiguation with a similarity threshold) and train with symmetric InfoNCE. They release EndoReport100 for lesion-level retrieval and evaluate on global/intra-case retrieval, six multi-centre classification tasks (zero-shot, linear probe, end-to-end), embedding kNN structure, structured report generation with a frozen Qwen3-14B decoder, stage ablations, a naive case-level pairing control, and a blinded 12-endoscopist pathology reader study. EndoCLIP outperforms general-purpose and biomedical VL encoders on these benchmarks, with a frozen linear probe approaching expert-tier accuracy on benign-versus-malignant classification.","tokens_in":18102,"tokens_out":1605,"duration_ms":47503,"significance":"If the recovered pairs largely reflect true finding-to-frame clinical semantics, the work is a substantial contribution: it shows that routine endoscopy documentation can supply scalable vision–language supervision without per-task dense annotation, and that clinical targets can be specified in language. Strengths that should be credited include the multi-protocol evaluation design (bootstrap CIs and paired tests on EndoReport100; multi-seed label-efficiency curves; frozen-decoder structured reporting), the stage-wise ablation separating coarse polyp detection from descriptive retrieval, the naive case-level control, public release of EndoReport100 and code, and the blinded reader comparison. These make the empirical package stronger than typical medical VL pretraining reports and give the field a concrete benchmark for multi-lesion report-grounded retrieval.","major_comments":[{"comment":"Methods §4.4.2 (Stage III): multi-lesion pairs (~56.2k of 125.8k) are retained only when EndoCLIP-III similarity exceeds a fixed threshold of 0.28 chosen on a held-out multi-lesion validation set (Supp. Fig. S6, Table S22). The central claim—that recovering finding-to-frame correspondence yields scalable clinical supervision—depends on these matches being sufficiently correct and unbiased. The manuscript does not report a human audit of match precision/error modes (wrong lesion, normal mucosa, style-correlated near-misses) at the operating threshold, nor sensitivity of EndoReport100 and EndoVL gains to that threshold. A modest clinician-rated sample of accepted/rejected multi-lesion pairs, plus a brief threshold sweep on downstream metrics, is needed to ground the axiom that model-selected pairs encode clinical semantics rather than self-reinforced similarity.","section":"Methods §4.4.2; Supp. Fig. S6 / Table S22"},{"comment":"Results §2.2 and Discussion: EndoReport100 is independent of the pretraining split but drawn from the same institutional archive, and the Discussion notes that source identity remains locally predictable in the embeddings (Supp. Fig. S5, Table S18). Global mean Recall@1 of 14.3% vs 2.8% for PMC-CLIP is a strong headline, but same-archive style or acquisition cues could inflate retrieval and some transfer without proving lesion-level semantic correspondence. The naive case-level control (global R@1 6.5%) shows progressive recovery helps, yet does not separate true correspondence quality from archive memorization. Please either (i) quantify how much retrieval/classification signal survives controls that disrupt clinical text while preserving style (e.g., shuffled descriptors within site, or site-only prompts), or (ii) more tightly bound claims that rest primarily on EndoReport100 versus th","section":"Results §2.2; Discussion; Supp. Fig. S5 / Table S18"},{"comment":"Results §2.3 / reader study: the frozen EndoCLIP probe is reported to approach expert-tier accuracy on benign-versus-malignant classification (0.852±0.010 vs experts 0.846±0.014), with exploratory McNemar tests. Clarify whether the Zhongshan pathology images and patients are disjoint from the 280k-report pretraining archive (same centre; Methods §4.1–4.3). If overlap or near-overlap is possible, the reader comparison and in-house pathology AUC cannot be read as fully external. State the separation rule explicitly and, if full disjointness cannot be guaranteed, mark the reader study as same-centre and lean on EndoVL for external claims.","section":"Results §2.3; Methods §4.1–4.3; Supp. Table S15"}],"minor_comments":[{"comment":"Fig. 1 caption begins with a spaced typo: “W eak report–image alignment”. Fix.","section":"Fig. 1"},{"comment":"Abstract and main text say 125,756 pairs / 280,476 records; Fig. 1b funnel text uses 125.8k / 280.5k and 104.5k cases—align rounding and the curated-case count with Supplementary Table S1 everywhere.","section":"Abstract; Fig. 1b"},{"comment":"Methods §4.4.3 Eq. (1): state whether τ is the learnable CLIP logit scale (as in §4.4.4) or a fixed temperature; the two subsections should use identical notation.","section":"Methods §4.4.3–4.4.4"},{"comment":"Table 1 and Fig. 4: define the exact clinical prompt templates in the main text or point more visibly to Supplementary Table S2 when discussing zero-shot failures of PMC-CLIP (e.g., size sensitivity 0.002).","section":"Table 1; Fig. 4; Supp. Table S2"},{"comment":"Fig. 6g: the claim that 91% of Paris misclassifications are adjacent on the protrusion continuum is important; add the numeric off-diagonal counts in the supplement if not already fully tabulated.","section":"Fig. 6g"},{"comment":"Data/code availability statements are clear and appreciated; ensure the GitHub URL and Figshare DOIs remain stable through review and that EndoVL split manifests document any unavoidable patient-ID gaps noted in the Discussion.","section":"§5–6 Data/Code availability"}],"recommendation":"major_revision","confidential_remarks":"Solid empirical medical-AI paper with unusually thorough evaluation and released assets. The progressive-recovery idea is the novelty; the main risk is over-claiming that gains equal proven correspondence quality when multi-lesion pairs are model-selected and the primary retrieval benchmark is same-archive. I view this as fixable with a pair audit and tighter claim scoping rather than a conceptual flaw—hence major_revision, not reject. Fit for a methods/AI-in-medicine audience is good if external multi-centre results remain the load-bearing evidence after revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is not another CLIP on medical images. It is a concrete three-stage way to turn procedure-level colonoscopy reports into lesion-level pairs, plus evidence that those pairs actually buy you descriptive retrieval and label-efficient clinical transfer.\n\nWhat is new is the correspondence pipeline: coarse “polyp present / no finding” localisation, single-lesion sentence anchors, then multi-lesion matching with a similarity gate. That is the right problem for endoscopy, where reports summarise whole cases. They ship scale (~126k recovered pairs from ~280k records), stage checkpoints, a naive case-level control (global R@1 6.5% vs 14.3%), EndoReport100 with clinician links, multi-centre EndoVL tasks, frozen-decoder structured reporting, and a 12-reader pathology study. Baselines are the right ones (OpenAI CLIP, BiomedCLIP, PMC-CLIP). Stage ablations are honest: I/II detect polyps but sit at chance on descriptive retrieval; semantics show up only after sentence–frame training. Code and the retrieval benchmark are public. Math is standard InfoNCE; citations look appropriate, not padded.\n\nSoft spots, in proportion. Multi-lesion pairs (~45% of the set) are model-selected above a hand-tuned 0.28 threshold, so some circularity in “recovered correspondence” is real. EndoReport100 is held out from training but same institutional archive, and they admit source identity is still locally predictable in embeddings. That weakens the purest reading of the retrieval headline more than the EndoVL and pathology transfer story. Full pretraining data stay restricted, which is normal clinically but limits external audit. None of this looks load-bearing enough to dismiss the central claim; the progressive control and external classification curves still point the same way.\n\nThis is for people building endoscopy foundation models, report-grounded supervision, or label-efficient GI AI—not for pure theory. I would bring it to reading group, cite it if I work in this lane, and send it to referees. Worth engaging; ask for threshold sensitivity and clearer multi-site separation in revision, not a rewrite.","headline":"Solid applied VL paper: progressive report-to-frame recovery is the real contribution, and the multi-centre results mostly carry it despite same-archive retrieval caveats.","tokens_in":18797,"tokens_out":545,"would_cite":true,"duration_ms":18522,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Recovering which colonoscopy frames match which report sentences turns routine procedure notes into scalable training signal for a vision–language model that beats general biomedical encoders on retrieval, classification, and structured rep","keywords":["colonoscopy","medical foundation model","vision-language contrastive learning","finding-to-frame correspondence","zero-shot classification","structured report generation","EndoCLIP","lesion-level retrieval"],"falsifier":"A prospective, patient-grouped multi-institution test where EndoCLIP is scored on retrieval and zero-shot morphology or pathology labels drawn from centres, report styles, and scopes never seen in pretraining or EndoReport100; failure to beat the same biomedical baselines under that separation would refute the central claim.","tokens_in":18640,"feed_emoji":"🔬","tokens_out":1067,"duration_ms":25089,"temperature":0.7,"pith_summary":"Colonoscopy reports already describe lesion look, size, and location, but they summarise whole procedures, not individual frames, so the clinical text is only weakly tied to the pictures. This paper argues that you can recover those missing finding-to-frame links in stages—first finding polyp-bearing frames, then anchoring single-lesion descriptions, then disambiguating multi-lesion cases—and train a dual encoder (EndoCLIP) on the resulting 125,756 pairs drawn from about 280,000 routine records. The claim is that this recovered supervision is enough for the model to outperform general-purpose and biomedical vision–language encoders on lesion-level image–text retrieval, six multi-centre clinical classification tasks in zero-shot and linear-probe regimes, and structured report generation from frozen visual features. On benign-versus-malignant reading, a linear probe on the frozen encoder approaches expert endoscopists in a blinded comparison with twelve readers. If that holds, clinical targets can be stated in language instead of hand-labelling a new dataset for every task.","feed_headline":"Routine colonoscopy notes train a model that nears experts","feed_subtitle":"Recovering which frames match which findings turns 280k reports into lesion-level supervision.","key_machinery":"Three-stage progressive correspondence recovery: case-level evidence localisation with coarse “polyp present / no finding” prompts, single-lesion anchor selection that pairs each finding sentence with its best-matching frame, and multi-lesion disambiguation that keeps sentence–frame matches only above a similarity threshold—then standard CLIP-style contrastive training on the recovered pairs.","core_discovery":"Progressively recovering lesion-level image–text pairs from routine colonoscopy reports yields scalable cross-modal supervision: EndoCLIP trained on 125,756 recovered pairs outperforms general-purpose and biomedical vision–language encoders on lesion-level retrieval (global mean Recall@1 14.3% versus 2.8% for the strongest comparator), six multi-centre clinical classification tasks in zero-shot and linear-probe settings, and structured report generation, with a frozen linear probe approaching expert readers on benign-versus-malignant classification.","pith_inferences":["If source identity remains readable in the embeddings, apparent multi-centre gains may partly track acquisition or reporting style; hard site-holdout and style-matched negatives would separate true semantics from site cues.","The largest label-efficient gains on rare descriptors (villous, lobulated) suggest the method is most valuable precisely where new hand-labelled cohorts are hardest to build.","A natural next measurement is calibration and uncertainty of prompt scores under real time pressure, not only rank and AUC on static benchmarks.","Thresholded multi-lesion matching could systematically drop hard or atypical lesions; auditing discarded pairs against expert links would show whether the training set is skewed toward easy cases."],"forward_implications":["Clinical concepts such as Paris type, surface pattern, size threshold, or malignancy can be specified as text prompts instead of building a separate labelled classifier for each target.","Routine procedure documentation becomes a large, reusable training resource once finding-to-frame links are recovered, reducing dependence on dense frame-level annotation.","Frozen endoscopy-specific visual features can drive structured JSON drafts (diameter, Paris class, descriptor sets) for endoscopist review more accurately than general biomedical encoders under the same decoder.","Language-based lesion retrieval can support case review and teaching by locating frames from textual descriptions in multi-lesion procedures.","The same correspondence-recovery idea is proposed as transferable to other many-frame, many-finding procedure records such as upper GI endoscopy."],"fun_headline_variants":["EndoCLIP recovers 125k lesion pairs from 280k routine colonoscopy reports","Finding-to-frame recovery turns routine notes into colonoscopy VLM supervision","Lesion-level pairs from reports let EndoCLIP near experts on malignancy calls","Routine colonoscopy documentation trains a model that approaches expert readers","EndoCLIP outperforms biomedical VLMs after recovering report-to-frame links"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The automatically chosen frame–sentence matches, especially multi-lesion pairs kept only when similarity clears a hand-set threshold, are accurate and unbiased enough that later gains reflect real clinical meaning rather than archive style or self-reinforced matching errors.","fun_headline_variants_meta":{"raw":{"variants":["EndoCLIP recovers 125k lesion pairs from 280k routine colonoscopy reports","Finding-to-frame recovery turns routine notes into colonoscopy VLM supervision","Lesion-level pairs from reports let EndoCLIP near experts on malignancy calls","Routine colonoscopy documentation trains a model that approaches expert readers","EndoCLIP outperforms biomedical VLMs after recovering report-to-frame links"]},"model":"grok-4.5","effort":"low","cost_usd":0.003931,"raw_usage":{"total_tokens":1229,"prompt_tokens":752,"num_sources_used":0,"completion_tokens":99,"cost_in_usd_ticks":39308000,"prompt_tokens_details":{"text_tokens":752,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":378,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":752,"tokens_out":99,"duration_ms":9235,"temperature":1.0,"reasoning_tokens":378,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T06:29:03.987436+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A prospective, patient-grouped multi-institution test where EndoCLIP is scored on retrieval and zero-shot morphology or pathology labels drawn from centres, report styles, and scopes never seen in pretraining or EndoReport100; failure to beat the same biomedical baselines under that separation would refute the central claim.","supporting_citations":[],"review_version":1}