{"id":"d3c92ea0-2c52-46b3-8f2a-08ef1168b1e1","arxiv_id":"2505.24108","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Federated self-supervised pretraining with masked autoencoders produces GI endoscopy representations that outperform single-site training and approach centralized training on classification, detection, and segmentation.","lead":"This paper trains a vision model for gastrointestinal endoscopy images in a federated way, so hospitals share model updates but not patient images. It reports that the federated model approaches centrally trained performance on classification, detection, and segmentation tasks, which matters for privacy-preserving medical AI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FedFound's advantage over the lower bound is confounded with data volume and domain mix; the paper's own upper bound shows centralized full-data training beats FedFound, so federation itself is not shown to cause the gain.","rationale":"The strongest claim is that federated pretraining yields better representations than single-site pretraining and comes close to centralized pretraining. The reader's weakest assumption—that the lower bound is confounded with data volume—is exactly the load-bearing issue. Section IV.A assigns the server 10,000 images and each of the five clients 18,938 images in the homogeneous split, so FedFoundAvg trains on at least 104,690 images versus 10,000 for the lower bound, while the stated pool is 125,846 images. The heterogeneous split similarly gives the lower bound only the 10,000 HyperKvasir server images, while FedFoundAvg sees all data including Kvasir-Capsule. Thus the reported 8.9-point classification gap could reflect data scale and domain diversity rather than the aggregation mechanism. The paper's upper bound is informative: centralized training on all data achieves 81.3%, above FedFoundAvg's 78.0%, demonstrating that access to the full data suffices to explain the improvement. To attribute the gain to federated collaboration, the authors must add a control that matches total data and compute while varying only the training protocol—for example, a centralized run with the same number of optimizer steps as the federated run. The proposed test settles this. The verdict remains CONDITIONAL: the framework is relevant and the results are plausible, but the central causal claim is not yet supported. This critique concerns experimental design, not author integrity; no error bars, the data-accounting discrepancy, and epoch selection on downstream performance compound the issue but are secondary to the confound.","tokens_in":13633,"tokens_out":8124,"duration_ms":94091,"concrete_test":"Train a centralized MAE on the full pretraining pool (all 125,846 images) using the same optimizer, batch size, and total number of gradient steps as the FedFoundAvg run (or the same wall-clock time on identical hardware), then fine-tune it under the identical downstream protocol across at least 3 seeds. If this data-matched centralized model reaches at least FedFoundAvg's 78.0% accuracy (and comparable detection/segmentation metrics), the 8.9-point gain over the lower bound is attributable to data volume and domain mix rather than federation. Additionally, resolve the data-split totals in Section IV.A, where the homogeneous split sums to 104,690 images instead of the stated 125,846.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the comparison in Section V.A: FedFoundAvg reaches 78.0% classification accuracy versus 69.1% for the lower bound and 81.3% for the upper bound. However, the lower bound (Section IV.A) is pretrained only on the server's 10,000 images, while FedFoundAvg is pretrained on the distributed pool of at least 104,690 images—and the paper states the full pool is 125,846 images, revealing an accounting discrepancy in the homogeneous split. The lower bound also sees only HyperKvasir data, whereas FedFoundAvg sees both HyperKvasir and Kvasir-Capsule, so the comparison conflates data volume and domain diversity with the effect of federated aggregation. The paper's own upper bound is the decisive control: centralized training on the full data achieves 81.3%, exceeding FedFoundAvg's 78.0%. This means access to the full data is sufficient to explain the improvement over the lower bound; federated aggregation itself contributes no measurable benefit. Without a control that matches total data and computational budget while varying only the federation mechanism, the claim that federated collaboration improves over single-site pretraining is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a federated learning framework for pretraining masked-autoencoder foundation models on unlabeled GI endoscopy images distributed across five clients and a server. It compares four FL aggregation algorithms (FedAvg, FedAvgM, FedAdam, FedAdagrad) under homogeneous and heterogeneous data splits, then fine-tunes the pretrained encoders for classification on GastroVision and detection/segmentation on EDD2020. The central claimed result is that FedFoundAvg improves classification accuracy from 69.1% (lower bound) to 78.0%, approaching the 81.3% centralized upper bound, with analogous gains in detection and segmentation.","tokens_in":13879,"tokens_out":5762,"duration_ms":60597,"significance":"The paper addresses a relevant and timely problem: privacy-preserving pretraining of medical foundation models. Its strengths are the systematic comparison of four FL algorithms, two distribution settings, three downstream tasks, and a frozen-feature evaluation, all grounded in existing public datasets. If the federated-collaboration benefit were cleanly isolated, this would be a useful benchmark. At present, however, the comparison to the lower bound conflates federation with data volume and domain mix, and the absence of error bars and the downstream-based epoch selection weaken the quantitative claims. The contribution is therefore promising but not yet established.","major_comments":[{"comment":"The homogeneous split description states that 18,938 images are assigned to each of five clients and 10,000 images to the server, but these numbers sum to 104,690, not the 125,846 images implied by 87,970 + 37,876. This leaves 21,156 images unaccounted for and makes the pretraining data pool and the lower-bound data volume ambiguous. This matters because the main comparison in Section V.A is between FedFound models and a lower bound trained on only the server's 10,000 images; if the accounting is off, the data-volume confound is even harder to evaluate.","section":"Section IV-A"},{"comment":"The central claim that federated collaboration improves over the lower bound is confounded with data volume and domain mix. The lower bound is pretrained on the server's 10,000 HyperKvasir images, whereas FedFoundAvg is pretrained on the distributed pool of about 125,846 images from both HyperKvasir and Kvasir-Capsule. The paper's own upper bound (centralized full-data training) reaches 81.3% accuracy, exceeding FedFoundAvg's 78.0%, so full data access is sufficient to explain the improvement. To support the claim that the federated aggregation mechanism itself is beneficial, the authors need a control that varies only the federation mechanism while matching total data and computational budget, such as a single-site model trained on the full pool of 125,846 images.","section":"Section V.A versus Section IV.A"},{"comment":"The paper states that pretraining was run for 1,000 epochs, 'a value selected based on performance on downstream tasks.' This is a form of evaluation leakage: choosing pretraining epochs using the downstream test data can bias the comparison, especially if the epoch count was tuned to maximize FedFound performance. The authors should either select the epoch count using a validation split that is disjoint from the test set, or report downstream performance across a range of pretraining epochs for all methods so that the selection rule is transparent and fair to the baselines.","section":"Section IV-B"},{"comment":"All reported metrics are single runs with no error bars, confidence intervals, or significance tests. Several important comparisons are close: for homogeneous segmentation, FedFoundAvgM achieves DSC 0.6959 versus FedFoundAdam 0.6617 and upper bound 0.6488; for homogeneous detection, FedFoundAvg achieves mAP 18.649 versus upper bound 20.860. Without repeated seeds and variance estimates, the claims that FedFound models 'consistently outperform' the lower bound and that one FL algorithm is better than another are not statistically supported. Please provide mean and standard deviation over at least three seeds for the main results, and use appropriate tests or confidence intervals when ranking algorithms.","section":"Tables B.1 and B.2, Section V"},{"comment":"Equation (4) defines a sample-size-weighted federated objective with a 1/n sum of n_k terms, but Algorithm A.1 line 11 aggregates client updates as a uniform average 1/|S| times the sum of delta_{t,i}. If the experiments use Algorithm A.1, the implemented objective differs from Eq. (4); if they use weighted averaging, Algorithm A.1 is incorrectly specified. The distinction is non-trivial here because the server holds 10,000 images while clients hold approximately 18,938 each. Please state exactly which aggregation rule was used in the experiments and align the algorithm description and the mathematical objective.","section":"Algorithm A.1 versus Eq. (4)"}],"minor_comments":[{"comment":"The left-hand side should be theta* = arg min_theta ..., not L(theta*) = arg min_theta ...; as written, the equation defines the loss value circularly.","section":"Eq. (1)"},{"comment":"There is a typo: 'HypeKvasir' should be 'HyperKvasir.'","section":"Section IV-A"},{"comment":"The word 'acehives' should be 'achieves.'","section":"Section V-B"},{"comment":"The self-supervised target is written with a tilde in Eq. (1) and as 'fX' in Eq. (4); please use consistent notation.","section":"Eqs. (1) and (4)"},{"comment":"The captions state that underlined values represent the second-highest results, but no underlining is visible in the manuscript text; please ensure the formatting is present in the final version.","section":"Tables B.1 and B.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript explicitly states that the pretraining epoch count was selected based on downstream task performance; this is an unusual admission and should be treated as a substantive methodological issue in revision, not merely a wording problem. The data-split arithmetic discrepancy in Section IV-A also deserves close editorial attention during the revision process."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid empirical demonstration, not a conceptual breakthrough. The authors federate MAE pretraining over two GI endoscopy datasets, compare FedAvg/FedAvgM/FedAdam/FedAdagrad in homogeneous and heterogeneous splits, and evaluate on classification, detection, segmentation, plus frozen-feature probes. FedAvg gets closest to the centralized upper bound. That is a useful data point for federated medical imaging.\n\nWhat it does well: the setup is non-trivial—four FL algorithms, two distribution settings, three downstream tasks, and an explicit upper/lower bound framing. The frozen-feature evaluation is a nice extra. Results are internally consistent, and the limitations section is honest.\n\nThe soft spots are real but not fatal. The headline claim—FedFound beats the lower bound—conflates federation with data volume and domain diversity. The lower bound is pretrained on the server's 10k HyperKvasir images; FedFound sees ~105k images from two datasets. The upper bound (centralized on the full pool) beats FedFound, so the gain is mostly from having more data, not from the aggregation mechanism itself. This is not a fatal flaw—access to distributed data is exactly the point of FL—but the paper should be worded more carefully and ideally add a control matched on data volume (e.g., a central model on the same 105k subset) to isolate the FL contribution. The stress-test's stronger claim that federation itself contributes nothing is overstated; the paper only claims to approach, not beat, centralization.\n\nSecond, experimental rigor: single runs, no error bars or significance tests. The pretraining epoch count was 'selected based on downstream performance'—true test-set leakage if the same test set was used for selection. They need to clarify. Third, an arithmetic discrepancy in the homogeneous split: the pooled total (125,846) doesn't match the sum of client and server assignments (104,690). That needs fixing. Finally, no code or split definitions, which hurts reproducibility.\n\nCitations are appropriate, though they don't benchmark against Federated EndoViT or EndoDINO directly.\n\nBottom line: worth sending to serious reviewers. It's a reasonable empirical contribution that would be much stronger with variance estimates, a data-volume-matched baseline, and a cleanup of the accounting. I'd conditional-accept after major revision, not desk reject.","headline":"A useful but not groundbreaking empirical study of federated MAE pretraining for GI endoscopy; the main claim is confounded by data volume, and the paper needs variance estimates and cleaner baselines before it can be fully trusted.","tokens_in":14427,"tokens_out":5262,"would_cite":false,"duration_ms":55826,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Federated pretraining of a masked autoencoder on unlabeled endoscopy images distributed across hospital clients yields representations that transfer to classification, detection, and segmentation nearly as well as centralized training.","keywords":["federated learning","foundation model","masked autoencoder","gastrointestinal endoscopy","medical image analysis","self-supervised learning","Vision Transformer","transfer learning"],"falsifier":"Train the same MAE model on the full 125,000 images for the same number of optimizer steps and epochs as the federated run (the current upper bound uses fewer compute hours); if this centralized model reaches or exceeds FedFoundAvg's downstream accuracy, the gap the paper attributes to federation is really a data-volume or compute effect.","tokens_in":13456,"feed_emoji":"🏥","tokens_out":6007,"duration_ms":52294,"temperature":0.7,"pith_summary":"The paper proposes a federated framework for pretraining a foundation model for gastrointestinal endoscopy images, where unlabeled images stay at local hospital nodes and only model updates are shared. Using a masked autoencoder objective, the authors pretrain a Vision Transformer backbone across five clients and a server, then fine-tune on classification (GastroVision) and on detection and segmentation (EDD2020). Their FedFoundAvg model improves classification accuracy by 8.9 points over a single-site lower bound (78.0% vs 69.1%), approaching the centralized upper bound (81.3%), with similar gains in detection and segmentation. The claim is that federated, label-free pretraining can build a general-purpose medical vision model without aggregating sensitive data.","feed_headline":"Federated training builds endoscopy AI nearly as good as pooled data","feed_subtitle":"Hospitals can jointly pretrain an endoscopy foundation model without sharing data, outperforming single-site training.","key_machinery":"The central object is the federated masked autoencoder (MAE): a Vision Transformer encoder trained to reconstruct masked image patches with an L1 loss, where pretraining is distributed across N client nodes plus a server node and aggregated by a federated algorithm (FedAvg, FedAvgM, FedAdam, or FedAdagrad). This combines the task-agnostic representation learning of MAE with the privacy-preserving update exchange of federated learning; the encoder's frozen or fine-tuned features are then evaluated through task-specific heads (an MLP for classification, ViTDet with Mask R-CNN losses for detection and segmentation).","core_discovery":"The central claim is that federated self-supervised pretraining with masked autoencoding is a viable route to general-purpose GI endoscopy representations. Across both homogeneous and heterogeneous client splits, the federated models (especially FedFoundAvg, which uses FedAvg aggregation) consistently outperform a lower bound pretrained on a single site's data and approach an upper bound trained on all pooled data. Concretely, FedFoundAvg reaches 78.0% classification accuracy (lower bound 69.1%, upper bound 81.3%), detection mAP of 18.649 (lower 14.008, upper 20.860), and segmentation Dice of 0.6301 (lower 0.5202, upper 0.6488) in the homogeneous setting; in the heterogeneous setting it reaches 77.8% accuracy, 20.353 mAP, and 0.6909 Dice, the latter exceeding the upper bound. The paper interprets these results as showing that federated foundation models can approach the performance of centralized training while preserving data privacy.","pith_inferences":["Because the lower bound is pretrained on one-tenth of the data that FedFound sees, the reported gains may overstate the value of federation per se; a matched-data-volume control is needed to isolate the collaboration effect.","The result that FedFoundAvg beats the upper bound on heterogeneous segmentation Dice hints that federated averaging may act as an implicit regularizer or benefit from diverse data ordering, a hypothesis testable by comparing ensembles of local models against their average.","A practical extension is to apply the same protocol to non-medical image benchmarks with deliberately matched compute and data volume to quantify how much federation costs or saves relative to centralized training.","Another testable extension is to inject differential privacy or secure aggregation into the update exchange and measure the resulting accuracy drop, giving deployers a privacy-utility tradeoff curve."],"forward_implications":["If the claim holds, hospitals can collaboratively build a shared endoscopy foundation model without exporting patient data, sidestepping a major barrier to medical foundation models.","FedAvg, the least sophisticated aggregation tested, delivers the best overall transfer, suggesting that simple averaging is sufficient for federated MAE pretraining on medical images.","Frozen features from FedFoundAvg support lightweight classifiers (MLP, SVM, XGBoost) with accuracy near the upper bound, so institutions with limited labels or compute can still benefit.","The heterogeneous-split results, where clients hold non-overlapping imaging domains, indicate the approach degrades gracefully under realistic domain shift, supporting multi-institutional deployment.","The early-epoch convergence behavior of FedFoundAvg suggests that federated pretraining can reach usable representations faster than single-site training."],"supporting_citations":[{"why":"Supplies the masked autoencoder pretraining objective that the federated foundation model is built on.","marker":"[12]"},{"why":"Provides the FedAvg aggregation algorithm, which yields the best-performing FedFoundAvg model.","marker":"[16]"},{"why":"Provides the ViTDet architecture used for fine-tuning on detection and segmentation downstream tasks.","marker":"[33]"},{"why":"Provides the Mask R-CNN multi-task loss used for detection and segmentation fine-tuning.","marker":"[34]"},{"why":"HyperKvasir, one of the two unlabeled datasets that supply pretraining images.","marker":"[35]"},{"why":"Kvasir-Capsule, the other unlabeled dataset supplying pretraining images.","marker":"[36]"},{"why":"GastroVision, the dataset used for the classification downstream task.","marker":"[1]"},{"why":"EDD2020, the dataset used for the detection and segmentation downstream tasks.","marker":"[37]"}],"fun_headline_variants":["Federated endoscopy AI rivals pooled-data training","Privacy-preserving endoscopy AI nears centralized performance","Endoscopy foundation model trains without sharing patient data","Federated pretraining beats single-site for endoscopy AI","Hospitals jointly build endoscopy AI without data leaving"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The lower-bound baseline is pretrained on only the server's 10,000 images, whereas FedFound uses roughly 125,000 images, so the improvement the paper credits to federation could instead come from simply seeing more data.","fun_headline_variants_meta":{"raw":{"variants":["Federated endoscopy AI rivals pooled-data training","Privacy-preserving endoscopy AI nears centralized performance","Endoscopy foundation model trains without sharing patient data","Federated pretraining beats single-site for endoscopy AI","Hospitals jointly build endoscopy AI without data leaving"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1452,"prompt_tokens":975,"completion_tokens":477,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":591,"tokens_out":477,"duration_ms":4793,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:34:32.294579+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same MAE model on the full 125,000 images for the same number of optimizer steps and epochs as the federated run (the current upper bound uses fewer compute hours); if this centralized model reaches or exceeds FedFoundAvg's downstream accuracy, the gap the paper attributes to federation is really a data-volume or compute effect.","supporting_citations":[{"cited_title":"Hyperkvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy,","cited_arxiv_id":null,"evidence_quote":"HyperKvasir, one of the two unlabeled datasets that supply pretraining images."},{"cited_title":"Kvasir-capsule, a video capsule endoscopy dataset,","cited_arxiv_id":null,"evidence_quote":"Kvasir-Capsule, the other unlabeled dataset supplying pretraining images."},{"cited_title":"Gastrovision: A multi-class endoscopy image dataset for computer aided gastrointestinal disease detection,","cited_arxiv_id":null,"evidence_quote":"GastroVision, the dataset used for the classification downstream task."},{"cited_title":"Endoscopy disease detection and segmentation (edd2020),","cited_arxiv_id":null,"evidence_quote":"EDD2020, the dataset used for the detection and segmentation downstream tasks."}],"review_version":1}