{"id":"68d053bc-c41f-4bc3-a51e-9e976e3ec365","arxiv_id":"2507.01055","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A broad survey that organizes prompt mechanisms for medical image generation, segmentation, and classification into a two-dimensional taxonomy of core technologies and clinical applications.","lead":"This paper surveys the many ways AI models in medical imaging can be guided by prompts: text descriptions, visual clicks or boxes, and learnable vectors, across image generation, segmentation, and classification. It gives a research field a structured map of methods, applications, and open problems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central claim to systematically survey prompt-based medical imaging is not auditable: no search protocol, inclusion criteria, or data-retrieval details are reported, so the corpus behind the taxonomy and trend analysis cannot be reconstructed.","rationale":"I read the paper as a survey whose value is organizational: it assembles prompt-based medical imaging works, groups them by task and prompt type, and claims to identify trends and challenges. For that central claim to hold, the assembled set must be representative of the field. The manuscript gives no way to check this. The reader's conditional verdict is therefore appropriate. I do not object to the existence of the taxonomy or the individual summaries; many are consistent with the cited works, and the paper provides useful entry points. The problem is epistemic: a self-described systematic review without a reproducible search is not yet a systematic review. My proposed check would settle the question by testing recall and trend stability. I am not alleging selective citation; the absence of methodology is sufficient to make the claim unverifiable. No change to the reader's conditional verdict is needed.","tokens_in":31420,"tokens_out":3214,"duration_ms":36337,"concrete_test":"Reconstruct the review by running a defined search across PubMed, Scopus, arXiv, and IEEE Xplore for 2022-2025 using stated prompt-related query terms, then compare the retrieved set against the references in Tables 1-5 and recompute the annual counts in Figure 2 from the retrieved set. If recall of representative key works is below roughly 80%, or if the recomputed trend changes the claimed steep increase, the representativeness claim fails and the manuscript needs a documented search protocol or a downgraded 'systematic' label.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that the review 'systematically surveyed the landscape' and that the trend in Figure 2 supports the field's rapid growth. This depends entirely on the completeness and unbiasedness of the literature corpus. The Introduction says only that the authors conducted 'an extensive literature search'; no databases, query terms, date range, screening rules, or inclusion/exclusion criteria are given, and no PRISMA-style flow or search log is provided. Consequently, the reference set in Tables 1-5 cannot be independently reproduced, and the quantitative curve in Figure 2 is not a verifiable measurement. Since the proposed taxonomy and the future-directions analysis are inferences drawn from this unstated corpus, the central organizational claim is currently unfalsifiable. The paper is still a useful narrative compilation with a large reference list, but the term 'systematic' carries evidential weight that is not supported by the manuscript's methods.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of prompt-based mechanisms in medical imaging. It proposes that prompts—text, visual, learnable, structural, and multimodal—can guide deep learning models toward improved generation, segmentation, and classification, particularly in data-scarce and distribution-shift settings. The paper is organized around the three clinical tasks, with additional sections on theoretical foundations, medical image foundation models, and future directions, and it includes five reference tables covering representative works in generation, segmentation, and classification. The central claims are that prompt mechanisms are a transformative advance in medical imaging and that the paper systematically surveys this landscape through a novel classification framework and trend analysis.","tokens_in":31573,"tokens_out":5555,"duration_ms":59599,"significance":"If taken as a narrative review, the manuscript is useful: it collects a large and recent body of work, organizes it in a mostly sensible task-based manner, and highlights relevant challenges such as prompt brittleness, lack of standardization, and clinical deployment. The discussion of foundation models in pathology, radiology, and ophthalmology is a helpful complement to the task-based tables. The self-citations do not appear to drive the conclusions, and the descriptions of specific methods are broadly consistent with the cited literature. However, the paper's value as a systematic survey is currently limited because the corpus behind the taxonomy, tables, and trend analysis is not auditable, and the promised two-dimensional classification framework is not actually operationalized in the body. The manuscript is best viewed as a broad narrative compilation with useful reference tables rather than as a reproducible systematic review.","major_comments":[{"comment":"The abstract and the contributions list describe the work as a systematic review, but the only methodological statement is 'Our approach involves an extensive literature search' in the Introduction. No databases, query terms, date range, screening rules, inclusion or exclusion criteria, or quality-assessment procedure are reported. As a result, the reference sets in Tables 1-5 cannot be reconstructed and the quantitative trend in Figure 2 is not a verifiable measurement. Please add a methods section or appendix with a search protocol and corpus-selection details, or revise the 'systematic' claim to describe a narrative survey.","section":"Introduction"},{"comment":"The Introduction promises a classification system with two dimensions: 'the core technologies underpinning prompt mechanisms' and 'clinical application paradigms.' However, the body is organized exclusively by clinical task (generation, segmentation, classification) and by prompt modality; no section defines or applies the core-technology dimension (design, generation, integration, transfer/multimodal synergy), and the tables list only Reference, Year, Task, Prompt Type, and Link. Please add an explicit taxonomy with definitions and assign each included work to its categories, or revise the claim to describe a task-and-modality-based organization rather than a two-dimensional classification framework.","section":"Theoretical Foundations and Taxonomies of Prompt Mechanisms"},{"comment":"Figure 2 reports the 'Number of papers' per half-year from 2022-Q1/2 to 2025-Q1/2. Since the manuscript is dated June 2025, the final interval is necessarily partial, and the monotone growth shown cannot be checked without knowing the corpus, retrieval date, and inclusion rules. The figure should be accompanied by the underlying counts, the search and screening protocol that produced them, and a clear label that 2025-Q1/2 is a partial interval; otherwise the trend claim is not independently verifiable.","section":"Figure 2"}],"minor_comments":[{"comment":"The phrase 'intense learning models such as convolutional neural networks' should be 'deep learning models such as convolutional neural networks.'","section":"Introduction"},{"comment":"In the paragraph comparing visual and text prompts, the sentence beginning 'On the other hand, text-prompt-based models can be trained with multiple types and modalities of data simultaneously...' appears twice, with slightly different wording; please remove the duplicate.","section":"Medical Image Segmentation (application section)"},{"comment":"The sentence 'Representative works in this area include references187, 188, 188, 189, 189, 190' contains repeated citation numbers; please clean up the citation list.","section":"Medical Image Foundation Models"},{"comment":"References 21 and 60 (and references 13 and 67) appear to be the same works cited under different numbers; please merge or cross-reference them to avoid duplication.","section":"References"},{"comment":"The row for Huang et al. 141 has an empty Prompt Type entry; please fill in the value or mark it as not applicable so that the table is self-consistent.","section":"Table 4"},{"comment":"Please state the retrieval date for the literature corpus and label the 2025-Q1/2 bar as partial data, since the manuscript is dated June 2025.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would be strengthened by adding a short methods appendix on search and screening; if the target journal does not require systematic-review methodology, the authors should remove the terms 'systematic' and 'extensive literature search' and reposition the paper as a narrative review. The duplicated text and reference inconsistencies are easy to fix but should be corrected before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on arXiv:2507.01055. It's a survey worth reading as a map of prompt-based medical imaging, not as a systematic measurement of the field. The two-axis framework (prompt technology vs clinical task) is a genuinely helpful expository device, and the three task sections (generation, segmentation, classification) cover a broad and current set of papers, with tables that include links. Newcomers will find it a fast way into the literature.\n\nThe paper does a decent job distinguishing prompt types — text, visual, learnable, structural — and the segmentation section's discussion of where prompts enter the pipeline (input, fusion, decoding) is more thoughtful than most surveys. I saw no serious mischaracterizations of the cited methods, which is the main thing you want from a survey.\n\nNow the soft spots. The 'systematic review' label is not supported. There is no search protocol, database list, inclusion/exclusion criteria, or PRISMA-style flow. That means the reference set and especially the growth curve in Figure 2 are not independently checkable. That matters because the paper leans on Figure 2 to justify urgency. It is what it is: a curated narrative review, and that's fine, but it should say so. The editorial quality also has rough edges: a duplicated sentence in the segmentation section ('On the one hand... On the other hand...' repeated with near-identical wording), and several references are listed twice under different numbers (e.g., 13 and 67 both are FLAIR; 21 and 60 are the same arXiv preprint; the list 187–190 has repeats). These are minor but they erode confidence if they appear in the final version.\n\nThe taxonomy itself is the main contribution, and it's a modest one — similar surveys with different axes exist in the broader prompt-learning literature, and the paper doesn't position itself against them. That's a gap in the related work, not a fatal flaw.\n\nMy bottom line: the paper is a solid, useful compilation. It deserves a serious peer review, but it needs a methodology section (or a retraction of the 'systematic' claim), cleanup of the duplicated content, and a paragraph positioning it against other prompt surveys. For a reader new to this area, it's a good starting point; for an expert, it's mostly a reference index. I'd be comfortable citing it as a recent survey once the editorial issues are fixed.","headline":"A useful narrative survey with a workable taxonomy, but the 'systematic' label is unsupported and the editorial slips need cleanup before it can be trusted as a reference index.","tokens_in":32093,"tokens_out":2552,"would_cite":true,"duration_ms":29958,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prompt-based mechanisms are becoming the default way to adapt medical imaging AI, steering generation, segmentation, and classification with text, clicks, boxes, and learned embeddings.","keywords":["prompt engineering","medical image segmentation","medical image classification","medical image generation","vision-language models","foundation models","visual prompts","learnable prompts"],"falsifier":"A systematic replication using a defined search protocol (databases, dates, inclusion and exclusion criteria) that identifies a substantial body of prompt-based medical imaging work the survey omits, or a period of publications whose distribution over prompt types and tasks contradicts the reported trend lines, would show the review's synthesis is not representative. The same replication could check the claim of exponential growth by counting prompt-driven medical imaging papers per quarter in a fixed set of venues and comparing to Figure 2.","tokens_in":31252,"feed_emoji":"🩻","tokens_out":5215,"duration_ms":56415,"temperature":0.7,"pith_summary":"This survey argues that prompt-based mechanisms — textual instructions, visual clicks or boxes, and learnable embedding vectors — have become a decisive way to adapt deep learning models to medical imaging, especially large vision-language and segmentation foundation models. The authors claim that prompting improves image generation, segmentation, and classification by injecting domain knowledge without full retraining, and that this makes models more accurate, more robust to data scarcity and distribution shifts, and more interpretable. They organize the field with a two-dimensional classification framework: the core technology behind a prompt (how it is designed, generated, and integrated) and the clinical task it serves. The review's central message is that prompt engineering is the path from single-task medical AI to flexible, multimodal generalist systems. The paper is a survey, so this claim rests on the completeness and representativeness of its literature synthesis.","feed_headline":"Survey: prompts make medical imaging AI adaptable without retraining","feed_subtitle":"Text, clicks, and learned prompts let one model handle many imaging tasks, but prompt brittleness still blocks clinical use.","key_machinery":"The organizing apparatus is a two-axis classification framework. One axis categorizes prompt mechanisms by their core technology: textual prompts encoded by biomedical language models, visual prompts such as points, boxes, and masks encoded by prompt encoders, and learnable or adaptive prompts such as prompt tokens, self-prompt modules, and domain-specific vectors. The other axis categorizes by clinical application: image generation, segmentation, and classification. Within this framework, the recurring technical machinery is a frozen pre-trained backbone — a vision-language model such as CLIP or a segmentation model such as SAM — with prompts injected at the input stage, through cross-attention, or into decoders, so that the prompt carries task-specific knowledge the backbone did not train on. The framework is what lets the authors compare methods that otherwise look very different and extract cross-cutting trends such as the shift from static text templates to learnable prompt vectors and the emergence of text-to-visual prompt conversion.","core_discovery":"On the paper's own terms, the discovery is that a small external conditioning signal — a phrase, a point, a box, a learnable vector — can steer a pre-trained model toward a specific medical task while leaving most of the network frozen. Across the surveyed studies, prompts appear as a common mechanism in three task families: text-conditioned generation of chest X-rays, CT, MRI, pathology, and ultrasound; visual- and text-guided segmentation of organs, lesions, and nuclei, often built around SAM-style prompt encoders; and zero- or few-shot classification mediated by vision-language encoders such as CLIP and its medical variants. The paper claims this convergence is not incidental: prompting constrains the model's hypothesis space, modulates attention, and projects features toward task-relevant regions, which is why it can deliver task adaptation with parameter efficiency and improved robustness. The review further claims that the main obstacles are now prompt brittleness across scanners and institutions, non-standardized prompt design, and the need for scalable clinical deployment, and that future work should move toward multimodal, personalized, and automatically optimized prompts.","pith_inferences":["The survey's taxonomy implies a testable ranking: prompt mechanisms can be compared by how much task knowledge they transfer per trainable parameter, a metric the reviewed papers rarely report uniformly.","The centrality of SAM and CLIP architectures suggests that medical-prompt research will co-evolve with generalist vision models; if those backbones change, prompt designs may need to be rebuilt.","An independent replication of the literature search — with explicit inclusion criteria — could verify whether the reported trends, such as the dominance of text prompts, reflect the actual publication distribution or a sampling bias.","Prompt brittleness across scanners is presented as a challenge, but the reviewed low-frequency and domain-prompt methods suggest a concrete research program: characterizing which prompt types are invariant to which distribution shifts."],"forward_implications":["If the claim is right, prompt-based adaptation becomes a default interface for medical foundation models, allowing one model to serve multiple tasks and modalities by swapping prompts instead of retraining.","Prompt-driven generation can supply synthetic medical images for data augmentation and education, directly targeting the data-scarcity bottleneck described in the survey.","Text-guided segmentation and classification should keep improving with richer medical text encoders and hierarchical prompts, enabling zero-shot and few-shot workflows where pixel-level labels are scarce.","The identified obstacles — prompt brittleness under distribution shift and lack of standardized prompt evaluation — become the field's core research agenda, with prompt optimization and benchmarks as expected next steps.","Multimodal and personalized prompting should move medical AI toward systems that integrate images, reports, genomics, and patient history, supporting diagnostics and treatment planning."],"supporting_citations":[{"why":"Supplies the contrastive vision-language backbone whose text-prompt paradigm most surveyed classification and generation methods build on.","marker":"[170]"},{"why":"Defines the promptable segmentation model with point, box, and mask prompts that many medical segmentation studies adapt.","marker":"[49]"},{"why":"Establishes the medical adaptation of SAM with box prompts, a central exemplar of the visual-prompt segmentation category.","marker":"[42]"},{"why":"Demonstrates text-prompt-driven universal organ and tumor segmentation, anchoring the text-guided segmentation line.","marker":"[7]"},{"why":"Provides the medical vision-language contrastive framework used for text-guided classification and generation.","marker":"[22]"},{"why":"Supplies a large-scale biomedical vision-language pretraining model whose text encoders encode medical prompts.","marker":"[21]"},{"why":"Shows parameter-efficient medical image segmentation via learnable prompt tokens, anchoring the learnable-prompt category.","marker":"[4]"},{"why":"Introduces collaborative domain prompts for medical image classification, anchoring the domain-generalization application.","marker":"[5]"}],"fun_headline_variants":["Prompts let medical imaging AI adapt without retraining","Small prompts steer medical AI across imaging tasks","Survey: prompt-based adapters for medical imaging","Prompt engineering: key to flexible medical imaging AI","One model, many scans: prompts make it possible"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the literature search is comprehensive and unbiased enough that the taxonomy, trends, and challenge list faithfully represent the whole field of prompt-based medical imaging; the paper does not disclose a search protocol, inclusion criteria, or quality assessment that would allow this to be checked.","fun_headline_variants_meta":{"raw":{"variants":["Prompts let medical imaging AI adapt without retraining","Small prompts steer medical AI across imaging tasks","Survey: prompt-based adapters for medical imaging","Prompt engineering: key to flexible medical imaging AI","One model, many scans: prompts make it possible"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1354,"prompt_tokens":953,"completion_tokens":401,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":569,"tokens_out":401,"duration_ms":5090,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:57:59.024358+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic replication using a defined search protocol (databases, dates, inclusion and exclusion criteria) that identifies a substantial body of prompt-based medical imaging work the survey omits, or a period of publications whose distribution over prompt types and tasks contradicts the reported trend lines, would show the review's synthesis is not representative. The same replication could check the claim of exponential growth by counting prompt-driven medical imaging papers per quarter in a fixed set of venues and comparing to Figure 2.","supporting_citations":[],"review_version":1}