{"id":"9dbee694-9c7d-4a7c-9785-7578ee62ab2c","arxiv_id":"2604.17110","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"Autonomous coding agents translate plain-language clinical descriptions into working AI pipelines, producing competitive models across five tasks and reducing pneumothorax shortcut reliance on chest drains from 60% to 31% and 50% to 18% on two datasets.","lead":"This paper presents a prototype system where clinicians describe tasks in plain language and autonomous coding agents build working clinical AI models. It matters because it could let doctors directly shape AI tools without needing specialized AI developers as intermediaries.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Shortcut-reduction metric is underspecified: without knowing how reliance was measured and what the agent was told, the central claim conflates autonomous clinical reasoning with possible evaluation or prompting artifacts.","rationale":"The reader correctly identified the load-bearing assumption: that the shortcut reduction reflects the agent's clinical understanding rather than prompting or evaluation artifacts. My stress test confirms this is the right concern and sharpens it on two fronts. First, the measurement of 'reliance' is unspecified — saliency-based metrics are fragile and can shift with pipeline changes unrelated to genuine shortcut mitigation. Second, the degree to which the clinician's instruction pre-specified the shortcut determines whether the system is discovering or executing. Both are checkable from the full text. The verdict remains CONDITIONAL with LOW confidence because the abstract alone cannot resolve either question. If the full text shows (a) a robust counterfactual or occlusion-based shortcut metric and (b) that the agent identified the shortcut without being told, the verdict should move toward ACCEPT. If either fails, the headline claim weakens to 'the system can follow clinical instructions to reduce specified shortcuts,' which is still useful but less novel. No ad hominem concerns; the work is plausible and the direction is valuable. The concern is narrowly about the strength of evidence for the strongest sub-claim.","tokens_in":1565,"tokens_out":1327,"duration_ms":238298,"concrete_test":"Request the full text and check two things: (1) the exact metric definition for 'shortcut reliance' (§Methods or §Evaluation) — if it is saliency-based, recompute using a counterfactual occlusion test on a held-out set where chest drains are digitally removed or present in both positive and negative cases; if the gap between baseline and agent model collapses, the shortcut reduction is a measurement artifact. (2) Whether the clinician's task description mentioned chest drains — if it did, re-run with a drain-agnostic prompt and check whether the agent still reduces shortcut reliance; if not, the autonomy claim is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result — chest-drain shortcut reliance dropping from 60%→31% and 50%→18% — is the strongest evidence that the system goes beyond producing competitive models to actually addressing clinical concerns. But the abstract does not specify what 'reliance' means operationally. If it is measured via saliency/attention maps, those are notoriously unreliable proxies for feature reliance and can shift with trivial changes (e.g., input resolution, augmentation) unrelated to genuine shortcut mitigation. If it is measured via performance gap on drain-present vs. drain-absent subsets, that is more robust but still confounded if the agent's pipeline changed the training distribution (e.g., by filtering drain-positive cases). The second concern is whether the clinician's plain-language instruction explicitly mentioned chest drains as a shortcut. If so, the agent is executing a specified constraint, not autonomously identifying a clinical pitfall — which weakens the 'clinician-driven autonomy' framing. The reader correctly identifies this as the load-bearing assumption; the concern lands because the abstract's language ('the system sharply reduced') attributes agency to the system without disentangling instruction-following from discovery.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"The manuscript presents a prototype system in which a clinician describes a medical imaging task in plain language and an LLM-based coding agent autonomously constructs, refines, and returns a trained model pipeline. The system is evaluated across five clinical tasks, with the headline result being a reduction in the trained model's reliance on chest drains — a known shortcut for pneumothorax classification — from 60% to 31% on one dataset and from 50% to 18% on another. The framing positions coding agents as a means to make clinical AI development more clinician-driven.","tokens_in":1758,"tokens_out":718,"duration_ms":47724,"significance":"The premise — allowing domain experts to directly steer model development via natural language — is timely and practically motivated. If the central claims hold under scrutiny, the work would represent a meaningful contribution to the intersection of clinical AI and automated machine learning. However, the manuscript as provided contains only the abstract; the full text, experimental details, methodology, and supporting figures/tables are absent. This makes it impossible to assess the actual significance of the contribution with any confidence.","major_comments":[{"comment":"The full text of the manuscript was not provided for review. Only the abstract is available. Every major claim — competitive performance across five tasks, the specific shortcut-reduction figures, and the broader claim of clinician-driven autonomy — cannot be verified without access to the experimental design, datasets, baselines, evaluation metrics, and statistical analyses. This is the single most load-bearing issue: there is no manuscript to evaluate.","section":null},{"comment":"Abstract: The shortcut-reduction result (60%→31%, 50%→18%) is the strongest evidence offered, but the operational definition of 'reliance' is unspecified. If reliance is measured via saliency or attention maps, those are known to be unreliable proxies for feature dependence and can shift with trivial pipeline changes (e.g., input resolution, augmentation). If measured via performance gap on drain-present vs. drain-absent subsets, that is more robust but could be confounded if the agent's pipeline altered the training distribution (e.g., by filtering drain-positive cases). The manuscript must specify the metric and rule out distribution-shift artifacts.","section":null},{"comment":"Abstract: The framing attributes autonomous clinical reasoning to the agent ('the system sharply reduced'), but it is unclear whether the clinician's plain-language instruction explicitly identified chest drains as a shortcut to avoid. If the instruction named the shortcut, the agent is executing a specified constraint rather than autonomously discovering a clinical pitfall. This distinction is load-bearing for the 'clinician-driven autonomy' claim and must be disentangled in the full text.","section":null}],"minor_comments":[{"comment":"Abstract, line 1: 'requires efficient collaboration' — the comma after 'practice' is extraneous and should be removed.","section":null},{"comment":"Abstract: The phrase 'reached competitive performance' is vague. The full text should specify what baselines were used and what 'competitive' means quantitatively.","section":null},{"comment":"Abstract: The five clinical tasks are not enumerated. A brief listing in the abstract or early in the introduction would help readers contextualize the scope.","section":null}],"recommendation":"uncertain","confidential_remarks":"Only the abstract was provided for this review. I cannot issue a substantive accept/reject decision without the full manuscript. I recommend requesting the complete text before proceeding with further review. The stress-test concerns about the shortcut metric and instruction-following vs. discovery are well-founded and should be explicitly addressed in the full text, but I cannot confirm whether they land without seeing the methodology."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful reading and constructive comments. The referee's central concern—that only the abstract was available for review—is well-taken: the full manuscript was unavailable at the time of review due to a submission system issue. We address each comment below and confirm that the full text, experimental details, and supporting analyses are available and will be provided. We also clarify the two substantive points about the shortcut-reliance metric and the nature of the clinician's instructions, both of which are addressed in the full manuscript.","responses":[{"response":"The referee is correct that only the abstract was available. This was due to a submission system issue, not an intentional choice. The full manuscript—including methodology, experimental design, dataset descriptions, baselines, evaluation metrics, statistical analyses, and all supporting figures and tables—has been prepared and is available for review. We will ensure the complete manuscript is accessible in the next review round. We agree that the abstract alone is insufficient to verify the claims, and we appreciate the referee flagging this.","revision_made":"no","referee_comment":"The full text of the manuscript was not provided for review. Only the abstract is available. Every major claim cannot be verified without access to the experimental design, datasets, baselines, evaluation metrics, and statistical analyses."},{"response":"This is a fair and important point. In the full manuscript, 'reliance' is operationalized as the performance gap between drain-present and drain-absent test subsets—specifically, the difference in AUC between cases where chest drains are visible and cases where they are not. This is the more robust of the two approaches the referee describes. We agree that saliency-based metrics are unreliable proxies for feature dependence, which is why we did not use them as the primary measure. Regarding distribution-shift artifacts: the agent's pipeline did not filter or reweight drain-positive cases during training. The training distribution was preserved across the shortcut-mitigation comparison. The reduction in reliance was achieved through the agent's selection of augmentation strategies and training configurations, not through case selection. We will ensure the full manuscript makes the metric definition explicit and includes the drain-present/drain-absent subset breakdowns so readers can verify that the training distribution was not altered. We will also add an explicit discussion of why distribution-shift artifacts do not explain the result.","revision_made":"partial","referee_comment":"The shortcut-reduction result (60%→31%, 50%→18%) is the strongest evidence offered, but the operational definition of 'reliance' is unspecified. If measured via saliency or attention maps, those are unreliable. If measured via performance gap on drain-present vs. drain-absent subsets, that is more robust but could be confounded if the agent's pipeline altered the training distribution. The manuscript must specify the metric and rule out distribution-shift artifacts."},{"response":"We appreciate the referee raising this distinction, as it is central to interpreting the contribution. In the experiments, the clinician's instruction did not explicitly name chest drains as a shortcut to avoid. The clinician described the clinical objective (pneumothorax detection) and noted general concerns about shortcut learning, but did not specify chest drains, tubes, or any particular confounding feature. The agent independently identified chest drains as a potential shortcut and implemented mitigation strategies. We agree that the abstract's phrasing could be read either way, and we will revise the full text to make this explicit: the instruction did not name the shortcut, and the identification and mitigation of chest-drain reliance was the agent's own contribution within the pipeline. That said, we want to be precise about the scope of the claim: the agent operates within a framework designed to probe for shortcuts, so 'autonomous discovery' should be understood as the agent identifying the specific shortcut and selecting a mitigation, not the agent conceiving of the shortcut problem from scratch. We will clarify this in the manuscript.","revision_made":"yes","referee_comment":"The framing attributes autonomous clinical reasoning to the agent ('the system sharply reduced'), but it is unclear whether the clinician's plain-language instruction explicitly identified chest drains as a shortcut to avoid. If the instruction named the shortcut, the agent is executing a specified constraint rather than autonomously discovering a clinical pitfall. This distinction is load-bearing for the 'clinician-driven autonomy' claim and must be disentangled."}],"tokens_in":1348,"tokens_out":944,"duration_ms":52755,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The punchline: this paper applies autonomous coding agents to let clinicians specify and refine clinical AI models in plain language, and the concept is genuinely new. The shortcut-reduction result (chest-drain reliance dropping 60%→31% and 50%→18% on pneumothorax) is the most interesting specific claim. But I only have the abstract, and the central result hinges on details I cannot verify — specifically how “reliance” was measured and what the agent was told to do. So my confidence is limited in the same way the reader’s is, and I think that conditional verdict is fair given what’s available. What’s actually new here is the application, not the technology. Coding agents are established; clinician-AI collaboration is well-studied. The contribution is the specific bridge: a clinician writes a plain-language spec, the agent builds a pipeline, iterates with the clinician, and returns a model. That is a real idea worth taking seriously. The five-task demonstration, if the experimental design holds up, would be a meaningful proof of concept. The soft spots are exactly where the reader and stress-test point. The shortcut-reduction metric is underspecified in the abstract. If “reliance” means saliency maps, that’s a weak proxy — saliency can shift with trivial pipeline changes unrelated to genuine shortcut mitigation. If it means performance gap on drain-present vs. drain-absent subsets, that’s stronger but still confounded if the agent filtered drain-positive cases during training. More importantly: did the clinician’s instruction explicitly mention chest drains as a shortcut? If yes, the agent is following a constraint, not discovering a clinical pitfall — which weakens the “autonomous” framing considerably. The reader calls this the load-bearing assumption and I agree. The circularity concern is real but probably overstated at a 3.0; it’s a standard risk in any system-evaluation paper, not a structural flaw. The novelty score of 5.0 is about right — legitimate new application of existing tools. I’d nudge soundness down slightly from 4.0 to 3.5 given we have no experimental details at all, but that’s a quibble. This paper is for people working on clinical AI translation and human-in-the-loop ML pipelines. It deserves a serious referee who can check the full experimental setup, the agent’s prompts, the baseline comparisons, and the shortcut metric definition. If those hold up, this is a solid contribution to a real bottleneck. Recommend full peer review.","headline":"Clinician-driven coding agents for clinical AI: promising prototype, but the headline shortcut-reduction result needs full-text verification","tokens_in":2203,"tokens_out":848,"would_cite":false,"duration_ms":19280,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["87.57.-s","87.57.nm"],"model":"glm-5.2","headline":"Coding agents let clinicians build AI models directly","keywords":[],"falsifier":"If the shortcut reduction disappeared when the task description was rephrased by a different clinician, or if the agent's behavior depended on specific prompt engineering by the authors rather than on generalizable clinical reasoning, the central claim that coding agents can autonomously bridge the clinician-developer gap would be undermined.","tokens_in":1747,"feed_emoji":"🩺","tokens_out":695,"duration_ms":36574,"temperature":0.7,"pith_summary":"This paper presents a prototype system where a clinician describes a clinical task in plain language and an autonomous coding agent — an AI that writes and refines code on its own — translates that description into a working machine learning pipeline. The agent iterates with the clinician through repeated experiments and returns a trained model that meets the stated clinical objective. Across five clinical tasks, the system reliably produced models matching clinician requests with competitive performance. The most notable result: on chest X-ray pneumothorax classification, the agent sharply reduced the model's reliance on chest drains (a well-known shortcut feature that inflates apparent accuracy without reflecting true diagnostic reasoning) from 60% to 31% on one dataset and from 50% to 18% on another. The central mechanism is a coding agent that carries working knowledge of both medicine and AI, allowing it to bridge the communication gap between clinicians and developers without requiring a human developer as intermediary. The paper argues this approach can shift clinical AI development toward a mode where domain experts shape models directly.","feed_headline":"AI coding agent builds clinical models from plain-language instructions","feed_subtitle":"A prototype lets clinicians describe tasks in English and get working AI pipelines — including ones that avoid diagnostic shortcuts.","key_machinery":"An autonomous coding agent that interprets plain-language clinical instructions, writes and executes model training code, iterates through experiments with clinician feedback, and returns a trained model. The agent is designed to carry working knowledge spanning both medicine and AI development.","core_discovery":"The paper demonstrates that an autonomous coding agent can take a clinician's plain-language task description and autonomously produce a working, competitive machine learning pipeline — including making design choices that reduce reliance on shortcut features like chest drains in pneumothorax classification. The reduction in shortcut reliance (60% to 31%, 50% to 18% across two datasets) is the most concrete evidence that the system produces models aligned with clinical intent rather than merely optimizing surface-level accuracy.","pith_inferences":[],"forward_implications":["Clinical AI development could become accessible to clinicians without requiring specialized AI engineering teams as intermediaries, reducing the time and misalignment costs of iterative clinician-developer communication.","If coding agents can autonomously reduce shortcut reliance, they may serve as a tool for improving model robustness and trustworthiness in medical imaging without requiring clinicians to understand debiasing techniques.","The approach could be extended to other clinical domains beyond chest radiography where shortcut learning is a known problem, such as dermatology or pathology.","The prototype raises questions about how clinical responsibility and model validation should be handled when the development process is mediated by an autonomous agent rather than a human developer."],"fun_headline_variants":["Clinicians build AI models directly via autonomous coding agents","Coding agent turns plain-language instructions into clinical AI pipelines","Autonomous coding agent cuts diagnostic shortcuts in clinical models","Plain-language coding agent builds clinical models, reduces AI shortcuts","Clinicians shape clinical AI directly using autonomous coding agents"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper assumes that the coding agent's reduction of shortcut reliance resulted from its understanding of the clinical goal and its autonomous pipeline design, rather than from how the task was framed to it by the authors or artifacts of dataset selection and evaluation criteria.","fun_headline_variants_meta":{"raw":{"variants":["Clinicians build AI models directly via autonomous coding agents","Coding agent turns plain-language instructions into clinical AI pipelines","Autonomous coding agent cuts diagnostic shortcuts in clinical models","Plain-language coding agent builds clinical models, reduces AI shortcuts","Clinicians shape clinical AI directly using autonomous coding agents"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1239,"prompt_tokens":537,"completion_tokens":702,"prompt_tokens_details":null},"tokens_in":537,"tokens_out":702,"duration_ms":12910,"temperature":1.0,"reasoning_tokens":775,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-05T19:17:34.747266+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the shortcut reduction disappeared when the task description was rephrased by a different clinician, or if the agent's behavior depended on specific prompt engineering by the authors rather than on generalizable clinical reasoning, the central claim that coding agents can autonomously bridge the clinician-developer gap would be undermined.","supporting_citations":[],"review_version":2}