{"id":"6d194609-671d-4c24-b9ec-9d5e0bde87c0","arxiv_id":"2607.09102","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"CAPRA calibrates image-derived semantic proxy axes on a small labeled split into a reusable subgroup interface for failure auditing and domain-dependent robust transfer without deployment metadata.","lead":"CAPRA builds a calibrated, image-derived subgroup interface so medical imaging models can be audited and made more robust when demographic and acquisition metadata are missing at deployment. It matters because strong average accuracy often hides clinically critical failures once those labels disappear.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The load-bearing soft spot is whether small-cohort calibrated proxy posteriors stay faithful enough under shift to underwrite both audit and robust transfer.","rationale":"The reader already isolates the right load-bearing assumption: small metadata-labeled calibration plus cross-fitting must turn imperfect image-derived proxies into posteriors that stay useful under deployment shift for both audit and transfer. The manuscript’s own evidence boundaries (Discussion/Limitations; single mBRSET external; domain-dependent RQ1 gains; in-domain RQ5 budget) make that the softest link in the strongest claim, not a peripheral limitation. I do not find a deeper internal inconsistency in the method construction itself—proxy axes, soft posteriors, and axis-selective risk are coherent for the missing-metadata setting—so the concern does not push past CONDITIONAL. It confirms the reader’s verdict rather than changing it: denser multi-site external audit (and clearer when the interface helps vs hurts BA) is still what would convert CONDITIONAL into a stronger accept. Agreement is full on the weakest assumption; no separate soundness hole (e.g., leakage or metric redefinition) is more load-bearing than this faithfulness-under-shift premise.","tokens_in":16092,"tokens_out":736,"duration_ms":8344,"concrete_test":"Hold out a second external fundus site (or multi-hospital CheXpert-style protocol split) never used in D_cal/selection; recompute Table 2-style BA/WGA/Macro-F1 and Table 3 Pur./NMI/Agr./Gap for CAPRA vs CAPRA-anchor and ExMap after calibrating only on source D_cal. If external WGA/Macro-F1 gains reverse or Agr. with the dominant failure axis falls below image-only/ExMap, the shift-faithfulness premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim needs the calibrated interface I={t,q,r,α} (Eqs. 6–9, Alg. 1) to remain a faithful enough surrogate for hidden subgroup risk when true metadata A are unavailable at deployment. That faithfulness rests on a small D_cal (m≪ n), patient-level cross-fitting, temperature+confusion shrinkage (Eq. 7), and trust×disparity axis weights (Eq. 9). The paper itself shows the assumption is fragile: RQ2/Table 2 is only BRSET\to mBRSET (one handheld fundus shift); on matched BRSET the simpler anchor beats full CAPRA on BA/WGA, while full CAPRA only wins under that single external shift. RQ1/Table 1 further shows domain-dependent and sometimes negative BA deltas (e.g., DFR/CheXpert −8.5 BA, GSR/BRSET −4.3 BA), so the same interface can hurt average performance. RQ5/Fig. 4 only sweeps calibration budget on BRSET/HAM10000 in-domain, not under multi-site acquisition/protocol shift. If proxy teachers encode source-specific shortcuts that D_cal cannot fully correct, both the “remains informative under dataset shift” and “reusable for robust transfer” parts of the strongest claim weaken, leaving mainly in-domain semantic-alignment evidence (RQ4).","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes CAPRA, a calibrated proxy-axis framework for hidden-subgroup analysis when demographic, acquisition, and quality metadata are unavailable at deployment. It learns image-derived semantic proxy axes, calibrates their posteriors on a small metadata-labeled cohort with patient-level cross-fitting (temperature scaling plus confusion-matrix shrinkage), and packages tokens, posteriors, reliability scores, and trust×disparity axis weights into a reusable interface I={t,q,r,α}. That interface is used both for standalone failure-aware adaptation and as input to downstream robust learners (GroupDRO, JTT, DFR, DPE, GSR). Experiments on BRSET, external handheld mBRSET, HAM10000, and CheXpert address six RQs covering downstream transfer, external shift, failure-axis portability, semantic alignment versus image-only/ExMap partitions, calibration budget, and interface ablations. The authors report improved support-filtered worst-group accuracy in most settings, stronger alignment with explicit failure axes than latent baselines, and domain-dependent reuse gains, while acknowledging that proxy axes are not ground-truth metadata and that external validation remains limited.","tokens_in":16537,"tokens_out":1843,"duration_ms":35419,"significance":"If the calibrated interface remains a faithful enough surrogate for hidden subgroup risk under realistic missing-metadata deployment, CAPRA would fill a genuine gap between group-robust optimization (which usually assumes known groups) and latent slice discovery (which often yields clinically opaque clusters). The problem setting is clinically relevant: metadata routinely disappear in imaging workflows, and aggregate metrics can mask failure modes. Strengths that should be credited include a clear axis-selective risk formulation that avoids unstable full Cartesian products; explicit leakage control (external audit never used for calibration/selection); patient-level splits; multi-modality evaluation on public datasets; systematic RQs with means±std; and an honest Limitations section on domain-dependent gains and non-universal failure taxonomies. The reusable-interface framing (audit plus optional robust transfer) is more useful than yet another standalone robust learner. The main significance risk is that the dual claim—deployment-time audit under shift and reusable robust transfer—currently rests on thin external evidence and mixed average-case deltas.","major_comments":[{"comment":"RQ2 / Table 2: The abstract and contributions claim that CAPRA “remains informative under dataset shift,” but the only external deployment-shift result is BRSET→mBRSET (one handheld fundus acquisition change). On matched BRSET the simpler CAPRA anchor is stronger on BA and WGA; full CAPRA wins only under that single external set. One modality-matched shift is insufficient to underwrite a general shift claim for a multi-domain methods paper. Either add at least one further external/protocol shift (e.g., multi-site CXR or dermoscopy device shift) with the same no-leakage protocol, or substantially soften the abstract/intro/conclusion language to “informative under the tested handheld fundus shift.”","section":"§6.2 RQ2, Table 2"},{"comment":"RQ1 / Table 1: CAPRA improves WGA in 14/15 comparisons, but BA deltas are mixed and sometimes large and negative (DFR/CheXpert −8.5 BA, −3.2 WGA; GSR/BRSET −4.3 BA; JTT/CheXpert −2.4 BA). The text correctly notes domain dependence, yet the reusable-transfer claim still needs a clearer failure analysis: under what measurable conditions does attaching I hurt average performance or even WGA? Without that, readers cannot decide when to deploy the interface versus leave a strong baseline (e.g., DFR on CheXpert) alone. A short diagnostic—e.g., relating negative transfer to axis-trust u_k, calibration ECE, or overlap between proxy buckets and true error mass—would make the central “reusable interface” claim operational rather than post-hoc.","section":"§6.1 RQ1, Table 1"},{"comment":"§4.1 and Algorithm 1: Proxy teachers T_k are load-bearing for every subsequent object (tokens, q, r, α), but the manuscript only says CAPRA “trains or obtains” teachers that produce logits over V_k. Architecture, training data, whether teachers see the same D_tr images, label source for each axis (age/sex vs quality/focus/artifacts), and teacher accuracy/ECE on D_cal are not reported. Without this, the faithfulness assumption behind Eq. (7) cannot be audited, and the pipeline is not reproducible. Please specify teacher construction per dataset/axis and report teacher calibration quality before and after the temperature+shrinkage step.","section":"§4.1 Calibrated Proxy-Axis Interface, Alg. 1"},{"comment":"RQ5 / Figure 4 and Eq. (9): The calibration-budget sweep (gains leveling after ~2–5% on BRSET/HAM10000) and the QUALITY-trust / AGE-weight pattern are only shown in-domain. The skeptic concern is exactly whether small-cohort calibrated posteriors remain faithful once acquisition/protocol shift changes which slices are hard. Because axis weights α_k ∝ u_k Δ_k(h_0) are fit on D_cal from the source, a source-specific shortcut that D_cal cannot correct would propagate into both audit and reweighting. A minimal stress test—re-estimate or freeze α_k under the mBRSET (or another) shift and report posterior reliability / worst-axis BA as a function of calibration fraction under shift—would directly address the weakest assumption of the paper.","section":"§6.5 RQ5, Fig. 4, Eq. (9)"}],"minor_comments":[{"comment":"Abstract line “reveals disparity patterns missed by metadata-only slicing” is stronger than the main evidence, which primarily shows alignment with explicit failure axes and different dominant axes across domains (Fig. 2, Table 3). Clarify what metadata-only slicing would have missed in each dataset.","section":"Abstract"},{"comment":"Eq. (4) and §5.2: π_min and the default support rule for WGA versus the fixed @20 threshold in Fig. 2 should be stated numerically per dataset so support-filtered metrics are reproducible.","section":"§3 Eq. (4), §5.2"},{"comment":"Figure 3 is described as “aligned 2D subgroup partitions” but the embedding method (e.g., UMAP/t-SNE on which features) is not specified in the main text; a one-sentence caption or methods note would help.","section":"Fig. 3, §6.4"},{"comment":"Table 1 footnote marks GroupDRO* as a proxy-group variant; consider also stating in the table header or caption that no method receives oracle subgroup labels at training, to prevent misreading the large +CAPRA jumps as oracle GroupDRO gains.","section":"Table 1"},{"comment":"Several related-work citations appear only loosely connected to medical imaging subgroup audit (e.g., control/PHD-filter and federated quantization references). Trim or relocate non-essential citations to keep the related-work narrative focused.","section":"§2 Related Work, References"},{"comment":"Notation: Z_k is introduced as proxy semantic axes sharing vocabulary with A_k, but hard predictions â_k and soft q^(k) are both used; a short glossary of I={t,q,r,α} early in §4 would reduce reader load.","section":"§3–§4"}],"recommendation":"major_revision","confidential_remarks":"The contribution is real and the writing is unusually careful about domain dependence and leakage, which is above average for this area. The main risk for the journal is over-claiming on external shift from a single handheld fundus transfer while advertising a multi-domain deployment interface. If the authors add one more external shift or clearly scope the shift claim, and fully specify proxy teachers, this could become a solid methods paper. I do not see fabrication or circular evaluation; the soft spot is evidence breadth, not internal inconsistency. Fit is appropriate for a medical imaging / health-informatics venue that values deployment-facing evaluation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is that CAPRA is not another worst-group optimizer. It is a calibrated proxy-axis interface (tokens + soft posteriors + reliability + trust×disparity weights) for auditing and optionally reweighting when deployment metadata are gone. That packaging is the actual contribution.\n\nWhat is new is the interface object itself, not any single component. Proxy teachers, temperature/confusion calibration, and excess-risk reweighting are known; tying them to patient-level cross-fitting, support-filtered worst-axis risk, and a reusable I={t,q,r,α} for both audit and transfer is the move. The paper does this cleanly: problem setting is honest about not recovering full Cartesian groups; Algorithm 1 is readable; leakage control is stated; limitations admit domain-dependent gains and non-universal failure axes.\n\nEmpirically it earns credit. Three modalities (BRSET, HAM10000, CheXpert), public data, patient-level splits, means±std. RQ4 alignment is the strongest result—CAPRA purity/NMI and agreement beat image-only and ExMap while keeping a usable gap. RQ1 shows WGA up in 14/15 learner-dataset pairs, with honest negative BA cases (DFR/CheXpert, GSR/BRSET). RQ3’s failure maps make the portability claim precise: the interface transfers, the dominant axis need not. RQ5’s 2–5% calibration budget is practical.\n\nSoft spots, in proportion: external shift is one handheld fundus set (mBRSET). On matched BRSET the simpler anchor sometimes wins; full CAPRA wins under that single shift. Downstream BA can drop. No code/artifacts. Free parameters (τ, λ, π_min, clip ρ) are real but not hidden. The stress-test concern is fair but not load-bearing for the interface claim—the paper already frames gains as domain-dependent and does not claim universal robustness. Circularity is low: metadata are held out for calibration/audit only.\n\nWho it is for: medical imaging fairness, post-market monitoring, and people who need interpretable slices when labels vanish. Not a theory paper; not a field reorganizer. Math is standard calibration/reweighting, citations cover GroupDRO/JTT/DFR/Domino/ExMap/DPE and medical VLM work without obvious gaps.\n\nI would send it to peer review. Engage if you care about missing-metadata audit interfaces; treat external multi-site faithfulness as the open question, not a reason to dismiss the work.","headline":"Solid methods paper that packages proxy axes, calibration, and reweighting into a reusable missing-metadata audit interface; multi-domain evidence is real, external shift evidence is thin.","tokens_in":17127,"tokens_out":640,"would_cite":true,"duration_ms":8336,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"CAPRA recovers auditable medical-imaging subgroup failures from calibrated image-derived proxy axes when deployment metadata are missing.","keywords":["Medical Imaging","Subgroup Discovery","Robust Learning","Missing Metadata","Proxy Axes","Calibration","Hidden Stratification","Multimodal Learning"],"falsifier":"On a fresh multi-site external cohort never used for calibration or selection, CAPRA partitions show no better agreement with known failure axes and no higher support-filtered worst-group accuracy than image-only clustering or ExMap, even when the source calibration cohort still looks well calibrated.","tokens_in":17015,"feed_emoji":"🦷","tokens_out":555,"duration_ms":19736,"temperature":0.7,"pith_summary":"Medical imaging models are often deployed without the demographic, acquisition, and quality labels needed to check whether they fail on clinically important subgroups. CAPRA answers that gap by predicting semantic axes from the images themselves, calibrating those predictions on a small metadata-labeled cohort with patient-level cross-fitting, and packaging the calibrated posteriors into a reusable subgroup interface. The interface supports both deployment-time failure analysis and optional robust reweighting without requiring subgroup labels at deployment. Across fundus photography, dermoscopy, and chest radiography, it surfaces disparity patterns that metadata-only slicing misses, stays informative under external handheld-camera shift, and yields partitions that track explicit failure axes more closely than image-only or latent-slice baselines. The same interface can be reused by downstream robust learners, with gains that depend on the domain and how much subgroup structure the baseline already captured.","feed_headline":"Images alone recover hidden medical AI failure modes","feed_subtitle":"Calibrated proxy axes expose subgroup risk and feed robust learners without labels at deployment.","key_machinery":"CAPRA (Calibrated Proxy-Axis Risk Auditing): the calibrated subgroup interface of proxy-axis posteriors (temperature scaling plus confusion-matrix shrinkage under patient-level cross-fitting), reliability scores, and trust-weighted axis selection that concentrates robustness mass on axes that are both well calibrated and failure-relevant.","core_discovery":"Hidden subgroup analysis under missing metadata can be recast as calibrated proxy-axis risk estimation: image-derived semantic axes, after patient-level cross-fitted calibration on a small labeled split, form an interpretable interface that both exposes clinically meaningful failure modes and supplies group structure to robust learners without subgroup labels at deployment.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["CAPRA recovers hidden medical AI failures from images alone","Calibrated proxy axes expose subgroup risk without metadata","Image-derived axes unmask medical AI disparity patterns","CAPRA builds calibrated subgroup interface from images","Proxy axes feed robust learners without deployment labels"],"cache_read_input_tokens":128,"weakest_assumption_plain":"A small metadata-labeled calibration set, handled with patient-level cross-fitting, is enough to turn imperfect image-based proxy predictions into posteriors that stay faithful under real deployment shift.","fun_headline_variants_meta":{"raw":{"variants":["CAPRA recovers hidden medical AI failures from images alone","Calibrated proxy axes expose subgroup risk without metadata","Image-derived axes unmask medical AI disparity patterns","CAPRA builds calibrated subgroup interface from images","Proxy axes feed robust learners without deployment labels"]},"model":"grok-4.5","effort":"low","cost_usd":0.006698,"raw_usage":{"total_tokens":1673,"prompt_tokens":739,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":66980000,"prompt_tokens_details":{"text_tokens":739,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":880,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":739,"tokens_out":54,"duration_ms":9958,"temperature":1.0,"reasoning_tokens":880,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T05:23:04.357424+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a fresh multi-site external cohort never used for calibration or selection, CAPRA partitions show no better agreement with known failure axes and no higher support-filtered worst-group accuracy than image-only clustering or ExMap, even when the source calibration cohort still looks well calibrated.","supporting_citations":[],"review_version":1}