{"id":"1f225180-c7fe-4623-ad99-edeab3ce5a50","arxiv_id":"2505.03380","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"RCMed, trained on 20M automatically described medical image-mask pairs, segments 177 biomedical tasks across 9 modalities from text prompts and matches interactive baselines on external cancer data.","lead":"RCMed is a medical AI system that turns text prompts into pixel-level segmentations across 177 medical imaging tasks, trained on 20 million image-mask-description triplets. Its color-based description strategy and vision-language loop aim to give clinicians a prompt-only tool that works on many modalities without expert hand-labeling.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reader's overlap concern is real but only partially settles the issue; the larger unaddressed gap is the absence of released inference code and weights, which makes the 38.93-point held-out margin impossible to audit.","rationale":"The reader's weakest assumption (overlap of AbdomenAtlas, DDTI, uwaterloo with SA-Med2D-20M training data) is a legitimate risk to the external-validation claims. I partially agree: it is a plausible source of inflated external numbers, and Sec. 2.3's assertion that 'all datasets were held out' is under-evidenced because SA-Med2D is itself an aggregation of public datasets and the paper does not list its constituent sources. That concern does not, however, directly threaten the primary internal held-out headline, since the paper explicitly discloses a leak only for MedSAM, not for RCMed's own training or BiomedParse's. The more decisive issue is that no code or weights are released, so neither internal nor external numbers can be audited. Data availability is also confusing: it promises in-house datasets from GDPH (ESCC, PTC, CRC, GC, LC, BC, Lymphoma, NSCLC-HQ) while the evaluation describes a different in-house set from Sun Yat-sen Memorial Hospital, with no mapping between them. The architecture description is coherent and the CRD method is plausibly effective, but the headline claims rest on unreleased artifacts. A conditional verdict is appropriate: the claims are credible in principle but not yet checkable.","tokens_in":22107,"tokens_out":1434,"duration_ms":14369,"concrete_test":"Download the released code and model weights (or, while unavailable, run the interactive demo and a small public subset) and reproduce per-task DSC for a sample of the 177 held-out tasks using category-name prompts with the released RCMed pipeline and the public BiomedParse checkpoint. If the reported 38.93-point average margin cannot be reproduced on a multi-modality, multi-task subset, the central claim is unverifiable as presented.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that RCMed outperforms BiomedParse by 38.93 DSC on 835,081 held-out samples. The reader identifies potential training overlap with external sets; this is worth testing, but the internal held-out claims are the quantitative core. Here, the load-bearing gap is verifiability: the paper reports code and data as 'will release upon publication,' with no released model weights, no inference pipeline, and no per-task prompt specification. Without an executable implementation, the 38.93-point advantage cannot be checked against a public benchmark. The comparison also admits a known confound: Sec. 2.1 states that 'due to the inconsistent training data, some of our held-out data is also involved in training MedSAM'—only MedSAM, not BiomedParse—but BiomedParse is trained on BiomedParseData, which is distinct from SA-Med2D-20M; so dataset overlap cannot itself explain the BiomedParse gap. The 70.86-point lead over MedSAM no-prompt is less informative because a prompt-free SAM is a suboptimal off-label baseline. The decisive scientific question is whether a user with only the paper's category names and the released inference code would reproduce RCMed's large margin over BiomedParse; that cannot be inspected from this preprint.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes RCMed, a vision-language assistant for medical image analysis that performs text-prompted segmentation, classification, and localization. The authors construct a large training corpus (RCMedData) of 20 million image-mask-description triplets by applying a Color Region Description (CRD) strategy that uses the VLM InternVL-1.5 to describe the shapes and relative positions of colored anatomical masks. RCMed is built on a LLaVA-style architecture with a SAM encoder and Vicuna-7B LLM, trained end-to-end with a V-L and L-V projection loop. The authors report a held-out evaluation on 835,081 samples across 177 tasks in 9 modalities, claiming an average DSC improvement of 38.93 points over BiomedParse and superiority over MedSAM in prompt-free and loosely-prompted modes. They also report external validation on 33 datasets, including an in-house multi-cancer set from Chinese and Egyptian hospitals, with 20 cancer types, several of which are claimed to be unseen during training, plus a radiologist user study.","tokens_in":22485,"tokens_out":8787,"duration_ms":81227,"significance":"If the reported results are reproducible, RCMed would be a substantial advance: it presents a scalable way to generate language-driven medical segmentation datasets and demonstrates that a single text-promptable model can handle a wide range of clinical tasks across modalities. The CRD strategy is a practical contribution, and the in-house multinational cancer evaluation is a genuine attempt to assess generalization in a clinical setting. The held-out evaluation is large (835k samples) and includes per-task breakdowns. However, the lack of released code or weights, the absence of an overlap analysis between RCMedData and the public external datasets, and the absence of strong per-task supervised baselines currently prevent verification of the state-of-the-art claims.","major_comments":[{"comment":"The assertion in §4.3 that 'all datasets were held out and did not appear during model training' is unsubstantiated for the public external datasets because RCMedData is built directly from SA-Med2D-20M, which is itself an aggregation of public medical segmentation datasets. The manuscript does not list the component datasets of SA-Med2D-20M or perform an overlap analysis against AbdomenAtlas, DDTI, and uwaterloo. If any of these public external sets are contained in the training corpus, the external validation results in Fig. 2a do not demonstrate generalization. The authors must provide an explicit overlap analysis or a complete list of constituent datasets, and remove from the external evaluation any dataset that overlaps with training.","section":"§4.3, §2.3, Fig. 2"},{"comment":"The central quantitative claims—the 38.93-point average DSC improvement over BiomedParse and the 70.86-point improvement over MedSAM with no prompt—cannot be independently verified because the code, model weights, and a complete inference specification (exact text prompts for all 177 tasks, preprocessing steps, and the CRD description generation protocol) are not provided. The paper states that code and weights will be released 'upon publication,' but no demo or repository link is active for review. Without an executable implementation, the reported numbers are not auditable. The authors should release the code and weights, or provide a sufficiently detailed experimental protocol to allow replication, before final acceptance.","section":"Data/Code Availability"},{"comment":"The comparison with BiomedParse may not be fair because the prompt format for BiomedParse is not specified. The paper states that category names are used as text prompts for both RCMed and BiomedParse, but BiomedParse is a foundation model trained with specific prompt templates (e.g., 'Segment <class> in the image'). In Table 1, BiomedParse achieves exactly 0.00 DSC on numerous tasks (e.g., adrenal gland left, brainstem, gluteus maximus left, all rib left/right entries). A zero DSC across many tasks suggests that BiomedParse may be failing to parse the prompt or return an empty mask, rather than producing an incorrect but nonempty segmentation. The authors should report the exact prompt format for each baseline and discuss whether zero-DSC cases arise from prompt mismatch; if so, those tasks should be excluded or the baseline should be re-run with its recommended prompt.","section":"§2.2, Table 1"},{"comment":"The claim of 'state-of-the-art' performance is not supported by comparison with strong per-task supervised baselines such as nnU-Net (ref. [18]), which is the standard benchmark for medical image segmentation. The paper compares only with foundation models (BiomedParse, MedSAM) and a few classification/localization models. If the claim is intended to mean 'state-of-the-art among text-promptable foundation models,' that should be stated explicitly. Otherwise, the authors should add per-task comparisons with nnU-Net or equivalent supervised methods on a representative subset of the 177 tasks to justify the unqualified 'state-of-the-art' wording.","section":"Abstract, §2.2"},{"comment":"The reported evaluation split is internally inconsistent. Section 2.1 states 'We held out 20% of the RCMedData data to comprehensively evaluate the model's performance,' while Section 4.3 states the data was 'randomly split into 80%, 10%, and 10% as training, tuning, and validation.' Neither percentage matches the reported test size of 835,081 samples, which is about 4.2% of 20 million. The authors must clarify the exact split procedure, the number of samples in each split, and why the validation set used for comparisons differs from the stated split. This inconsistency undermines the precision of the central held-out evaluation.","section":"§2.1 vs §4.3"}],"minor_comments":[{"comment":"The sentence 'The model undergoes end-to-end training for 5 iterations' is almost certainly a typo and should read '5 epochs'; otherwise the model would not converge on 20 million samples.","section":"§4.4"},{"comment":"The row 'clavicula right' appears twice with identical values, and the table header says 'Dice Coefficient Similarity' instead of 'Dice Similarity Coefficient.' Please clean the table and correct the metric name.","section":"Table 1"},{"comment":"The caveat 'some of our held-out data is also involved in training MedSAM' is important and should appear earlier, ideally in the Results overview or a dedicated note, so that readers understand the MedSAM comparison on the held-out set is partially confounded.","section":"§2.1"},{"comment":"The Data Availability section mentions 'in-house datasets from Guangdong Provincial People’s Hospital (GDPH),' but Section 4.1 states the data were collected from Sun Yat-sen Memorial Hospital, Sun Yat-sen University. Please clarify the source institution and the relationship between the two names.","section":"Data Availability"},{"comment":"The abstract claims 'a 23.5% relative improvement in cell segmentation from microscopy images over prior art,' but the prior art is never named and no dataset or metric is specified. Please provide a precise comparison with the baseline and cite the source.","section":"Abstract"},{"comment":"References [7] and [28] appear to refer to the same paper (BiomedParse by Zhao et al.). Please merge or differentiate them.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is promising: the dataset construction via CRD is scalable, and the in-house multi-cancer external evaluation is a genuine strength. However, I cannot recommend acceptance without (1) an overlap analysis between RCMedData and the public external sets, (2) released code and weights for auditability, (3) better-justified baselines for the state-of-the-art claim, and (4) resolution of the internal inconsistencies. The issues are fixable but require substantial additional work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a serious empirical effort, not a toy. The CRD strategy—coloring masks and asking an off-the-shelf VLM to describe shapes and positions—is a genuinely useful trick for turning any segmentation dataset into language-driven training data. That alone is worth a look. The scale is real: 20M triplets, 177 tasks, 835k held-out samples, plus in-house multi-cancer external validation. If the numbers are right, RCMed is a meaningful step toward a single text-promptable medical segmentation model.\n\nThe architecture is a reasonable GLaMM/LLaVA-style extension with a SAM decoder and a bidirectional V-L/L-V projection. Nothing revolutionary, but sensible and clearly described.\n\nSoft spots, ordered by severity.\n\nFirst, verifiability. The paper says code will be released upon publication, no weights, no demo that works, no per-task prompt specification. The headline number—38.93 DSC over BiomedParse—cannot be checked. That is the load-bearing issue. The stress-test note is right: dataset overlap alone can't explain the BiomedParse gap, because BiomedParse's training data is distinct from SA-Med2D. So the reader's overlap concern, while worth testing, is not the strongest objection. The strongest objection is that there is no executable artifact to audit.\n\nSecond, baseline fairness. The MedSAM comparisons use bounding boxes derived from ground truth; that's a strong oracle baseline, and the no-prompt mode is off-label. The paper admits some held-out data leaked into MedSAM's training, which further muddies those specific numbers. The comparison against BiomedParse is cleaner, but we don't know the prompt template used for BiomedParse or whether the evaluation harness was shared.\n\nThird, small but real: no error bars except t-test p-values; no per-task variance; the external public datasets (AbdomenAtlas, DDTI, uwaterloo) could overlap with SA-Med2D components—the paper asserts disjointness but gives no overlap analysis. This is a moderate concern, not fatal.\n\nThe circularity burden is low: the numbers are empirical, not derived from the method itself. The bigger risk is self-referential evaluation through training-set leakage, which needs a careful audit.\n\nWho should read it: anyone working on medical vision-language models or segmentation foundation models. The CRD idea is citable even if the full claims don't hold up.\n\nMy recommendation: send it to peer review, but with a strong request for code/weights and a documented eval harness. Without those, the SOTA claims should be treated as unverified. I'd want the reviewers to push on the leakage analysis and the prompt specifications. The paper deserves a serious referee; it's just not there yet in terms of reproducibility.","headline":"Big empirical bet on text-promptable medical segmentation; the CRD dataset trick is the real contribution, but the central 38.93-point claim is unauditable until code and weights ship.","tokens_in":22945,"tokens_out":1938,"would_cite":true,"duration_ms":19106,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One model, prompted by plain text, outperforms prior medical segmentation assistants in 165 of 177 tasks across nine imaging modalities.","keywords":["medical image segmentation","vision-language model","text-promptable segmentation","medical foundation model","self-reinforcing correlation","color region description","cancer segmentation","multimodal medical AI"],"falsifier":"Inspect the component datasets inside SA-Med2D-20M and test whether AbdomenAtlas, DDTI, or the uwaterloo dermoscopy images (or their patient-level volumes) appear in the training corpus; if any are found, the external-validation Dice gains would no longer demonstrate generalization to unseen data.","tokens_in":21874,"feed_emoji":"🩻","tokens_out":6852,"duration_ms":63702,"temperature":0.7,"pith_summary":"The paper claims that a single text-promptable model can handle medical segmentation, localization, and classification across nine imaging modalities, because it is trained on 20 million image-mask-description triplets that teach it the shape and position of each structure rather than just its name. If the claim holds, clinicians and non-specialists could obtain precise lesion boundaries with natural-language prompts, without drawing bounding boxes or possessing radiology expertise. The reported evidence is a held-out evaluation of 835,081 samples across 177 tasks, where the model outperforms BiomedParse by an average of 38.93 Dice points and beats MedSAM's prompt-free mode by 70.86 points, plus external and in-house cancer datasets that were allegedly unseen during training.","feed_headline":"One AI model wins 165 of 177 medical segmentation tasks","feed_subtitle":"Plain-text prompts segment organs and lesions across nine modalities, beating box-guided rivals in speed.","key_machinery":"The load-bearing mechanism is a self-reinforcing vision-language correlation loop, aided by a Color Region Description (CRD) annotation strategy. CRD turns each segmentation mask into a set of colored regions and asks an off-the-shelf vision-language model to describe their shape and relative position, so the training text encodes spatial morphology instead of a bare class name. In the network, a vision-to-language projection carries image features into the Vicuna language model, while a language-to-vision projection carries conditioned embeddings back into the mask decoder; the <SEG> token triggers mask generation. This loop is what converts a text prompt like 'liver tumor' into a pixel-precise mask, and it is also what the paper credits for generalization to unseen diseases.","core_discovery":"On the paper's own terms, the central discovery is that the weakness of previous medical AI assistants is a weak vision-language correlation, and that a closed-loop architecture trained on richly described masks fixes it. RCMed runs visual features through a language model and then feeds language-conditioned features back into a SAM-style mask decoder, so text semantics can steer pixel-level attention while image details sharpen the text. The Color Region Description strategy generates the needed supervision by converting masks into colored patches and asking a vision-language model to describe their shapes and relative positions, creating the 20-million-triplet RCMedData. This combination reportedly yields state-of-the-art Dice scores on 165 of 177 held-out tasks, a 23.5 percent relative gain on microscopy cell segmentation, and competitive results on external cancer segmentation, including classes the model never saw, which the authors attribute to learned knowledge of normal anatomy.","pith_inferences":["The large 38.93-point gap over BiomedParse may partly reflect the weakness of the BiomedParse baseline itself; against MedSAM with a tight ground-truth box, RCMed wins on prompt-free usability but often not on raw Dice, so the practical claim is best read as 'text prompting can approach box-guided accuracy with far less input effort.'","Because CRD descriptions are generated by an off-the-shelf vision-language model from synthetic colored masks, the RCMedData supervision inherits whatever shape-description errors that generator makes; measuring segmentation performance against description quality would reveal the ceiling of the whole pipeline.","A straightforward testable extension is to train the same architecture on the same 20 million triplets but with class-name-only prompts, which would isolate how much of the gain is CRD text versus the closed-loop network design."],"forward_implications":["Text-only prompting could replace box-and-click interaction for routine organ and lesion segmentation, lowering the expertise barrier for using medical AI.","If the external numbers hold, the model's learned normality model lets it flag anomalies it was never named, such as acoustic neuroma or ovarian cancer.","The CRD annotation pipeline converts any existing image-mask dataset into language-driven training data, so scaling to new modalities or tasks requires no manual caption writing.","One-shot training-free adaptation offers a route to new classes without retraining, which the paper identifies as still limited and subject to catastrophic forgetting."],"supporting_citations":[{"why":"BiomedParse: the main text-driven baseline and dataset it must beat; supplies BiomedParseData with 3.4 million samples across 82 tasks.","marker":"[28]"},{"why":"SA-Med2D-20M: the underlying image-mask corpus that RCMedData converts into 20 million triplets.","marker":"[32]"},{"why":"MedSAM: the interactive segmentation baseline whose no-prompt, point, and box modes define the comparison.","marker":"[36]"},{"why":"SAM: supplies the image encoder, prompt encoder, and mask decoder that RCMed initializes and fine-tunes.","marker":"[31]"},{"why":"LLaVA: the architectural template for connecting the vision encoder to the Vicuna language model.","marker":"[29]"},{"why":"Vicuna-7B: the language model backbone that integrates image features and generates the <SEG> instruction.","marker":"[30]"},{"why":"GLaMM: source of the two-layer MLP V-L and L-V projection design.","marker":"[40]"},{"why":"Medical Segmentation Decathlon: source of the liver-tumor training and held-out data from France used in cross-race analysis.","marker":"[37]"}],"fun_headline_variants":["Closed-loop vision-language model wins 165 of 177 medical tasks","Text-guided pixel attention boosts medical segmentation accuracy","Color-described anatomy triplets train AI to precise delineation","Self-reinforcing AI links vision and language for diagnosis","RCMed matches rivals on 165 tasks, excels on unseen cancers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the external public datasets used to prove generalization truly were absent from the 20-million-image training corpus; the paper asserts this but does not report any overlap analysis.","fun_headline_variants_meta":{"raw":{"variants":["Closed-loop vision-language model wins 165 of 177 medical tasks","Text-guided pixel attention boosts medical segmentation accuracy","Color-described anatomy triplets train AI to precise delineation","Self-reinforcing AI links vision and language for diagnosis","RCMed matches rivals on 165 tasks, excels on unseen cancers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000336,"raw_usage":{"total_tokens":1862,"prompt_tokens":950,"completion_tokens":912,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":830}},"tokens_in":566,"tokens_out":912,"duration_ms":8770,"temperature":1.0,"reasoning_tokens":830,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:52:38.117087+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the component datasets inside SA-Med2D-20M and test whether AbdomenAtlas, DDTI, or the uwaterloo dermoscopy images (or their patient-level volumes) appear in the training corpus; if any are found, the external-validation Dice gains would no longer demonstrate generalization to unseen data.","supporting_citations":[{"cited_title":"Biomedparse: a biomedical foundation model for image parsing of everything everywhere all at once,","cited_arxiv_id":null,"evidence_quote":"BiomedParse: the main text-driven baseline and dataset it must beat; supplies BiomedParseData with 3.4 million samples across 82 tasks."},{"cited_title":"Segment anything in medical images,","cited_arxiv_id":null,"evidence_quote":"MedSAM: the interactive segmentation baseline whose no-prompt, point, and box modes define the comparison."},{"cited_title":"Segment anything,","cited_arxiv_id":null,"evidence_quote":"SAM: supplies the image encoder, prompt encoder, and mask decoder that RCMed initializes and fine-tunes."},{"cited_title":"Glamm: Pixel grounding large multimodal model,","cited_arxiv_id":null,"evidence_quote":"GLaMM: source of the two-layer MLP V-L and L-V projection design."},{"cited_title":"The medical segmentation decathlon,","cited_arxiv_id":null,"evidence_quote":"Medical Segmentation Decathlon: source of the liver-tumor training and held-out data from France used in cross-race analysis."}],"review_version":1}