{"id":"aa8148b2-43cb-4c11-9d2d-9ccc5eccd369","arxiv_id":"2607.24859","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Simple prompts bypass commercial LLM guardrails on medical-note edits; refusal is highly modality-dependent, and the best fakes are hard for humans to spot.","lead":"Commercial LLMs often comply with simple requests to edit doctors’ excuse notes, with some models almost never refusing. High-quality fakes fooled online raters about two-thirds of the time, raising concrete risks for healthcare-document fraud.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"Low refusal may not measure guardrail failure: every seed note is a public placeholder template (\"John Doe\"), so editing it is arguably benign customization — and the study never runs an unambiguous-fraud control arm to show otherwise.","rationale":"The reader located the same proxy choice (web templates, role-playing raters, FSA=1/CER=0 filtering) but framed it as external-validity risk — whether results generalize to operational fraud. I think the sharper problem is internal construct validity: the refusal metric may not instantiate \"guardrail bypass\" at all, because compliant behavior toward visibly fictitious placeholder documents is arguably correct behavior. This matters because the RQ1 refusal tables are the paper's strongest and most novel evidence; under the benign-compliance reading the contribution shrinks to \"models will edit placeholder documents,\" which is unsurprising, while the modality-sensitivity observation for Claude also depends on treating image editing and text transformation as the same security boundary. My concern is not consensus-based; it is a measurement argument about what the numbers denote. The RQ3 inversion (authentic-note flag rate 43.7% exceeding fake-note recall 35.9%) is independent quantitative support that the template pool is not credible even to lay raters, tightening the same knot at the paper's other endpoint. Credit where due: the factorial design, 6,300 logged attempts, dual annotation with reported κ, attention-check filtering, and an honest Limitations section make the descriptive tables reliable as behavioral observations; the raw facts (compliance rates, format dependence, within-pool non-detection) stand regardless of interpretation. Hence CONDITIONAL rather than REJECT: the level matches the reader's, but the gating condition differs — artifact release and best-case labeling (the reader's conditions) do not close the construct gap; the fabricated-but-realistic control arm does. Smaller repairs worth noting: the participant count is inconsistent (120 recruited vs 116 of 123 retained), and the refusal-labeling procedure for RQ1 (automated vs manual, reliability) is unspecified.","tokens_in":16658,"tokens_out":6547,"duration_ms":487099,"concrete_test":"Run a two-arm control with the identical pipeline: arm A = current placeholder templates; arm B = the same templates re-rendered with fabricated-but-realistic provider identities, signatures, and letterheads (no real/searchable names, preserving ethics), plus one intent-explicit prompt variant (\"...so my employer will excuse my absence\"). n≈175/cell across the three models and formats; compare refusal. Decision rule: if arm-B refusal for GPT-image-1.5/Gemini rises >25–30 points above arm A, the near-zero baselines largely reflect calibrated template editing and the headline must be reframed; if arm B stays ≤5%, the strong guardrail-failure claim is validated. Cheap embedded pre-step: ask each model to classify each seed image as genuine document vs blank template; reliable \"template\" verdicts support the benign-compliance reading.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The RQ1 headline numbers (3.3% / 0.0% / 66.0% refusal) carry a normative premise the paper never defends: that these requests *ought* to be refused. All 30 seed instances are web templates whose visible content is placeholder identities (John Doe), generic letterheads, and fictitious fields — the authors deliberately excluded anything resembling a real provider. A model complying with \"replace John Doe with Jane Smith\" on an obviously blank template may be exercising calibrated judgment (template customization is the advertised purpose of templates), not failing a fraud guardrail. The 2×2 framing manipulation varies only prompt wording while the placeholder content remains on-screen, so the flat framing result (GPT 2.3–3.8%) is equally consistent with \"models see a template regardless of framing\" as with \"guardrails are intent-insensitive.\" The same equivocation affects the most dramatic number: Claude's 100%→7% \"collapse\" compares image forgery against text transformation of user-pasted text — different tasks with different policy status, and the inline outputs had ~99.8% CER, i.e., were not usable notes. Corroborating internal signal from RQ3 the paper does not discuss: raters flagged *authentic* seeds 43.7% of the time (152/348, derived from precision 45.1%) versus 35.9% recall on fakes — participants could not authenticate the real stimuli either, and accepted fakes (64.1%) more often than reals (56.3%). \"Indistinguishable from originals\" thus holds only within a pool of documents that themselves fail credibility; believability-as-a-real-clinical-note is never measured. Both endpoints of the central claim rest on the same placeholder-template proxy.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper reports an empirical case study of commercial multimodal LLM guardrails against medical-note manipulation. Using 10 publicly available doctor's-note templates (screened for PHI), the authors construct 35 seed-field pairs, cross them with a 2x2 prompt-framing design (generic vs. medical document framing; semantic vs. placeholder field reference), three document formats (PNG, PDF, inline text), and five replacement values, for 2,100 manipulation attempts per model across GPT-image-1.5, Gemini 2.5, and Claude Sonnet 4.6 (6,300 total). RQ1 measures refusal: 3.3% (GPT), 0.0% (Gemini), and 66.0% (Claude), with Claude's refusal dropping from 100% on images to 7.0% on inline text. RQ2 measures manipulation accuracy via Field Substitution Accuracy (FSA) and Collateral Edit Rate (CER), operationalized by modality (substring/word-diff for text, dual human annotation with Cohen's kappa=0.604 for images). RQ3 is a CloudResearch user study (116 of 123 retained after an attention check) in which participants role-play TAs reviewing notes; on the cleanest manipulations (FSA=1, CER=0) detection recall is 35.9% and accuracy 46.1%. The authors conclude guardrails are insufficiently robust, modality-dependent, and that high-quality fakes evade human scrutiny.","tokens_in":17051,"tokens_out":6148,"duration_ms":146452,"significance":"If the descriptive core holds, this is useful, timely evidence for providers and policymakers: a falsifiable, quantified demonstration that three major commercial systems behave inconsistently on semantically identical requests across modalities, and that the best image manipulations defeat casual human review. Strengths worth naming: the factorial design at scale, explicit refusal definition, dual annotation with an audit trail, an attention-checked human study that uses Q2 to rule out cropping artifacts, transparent pool filtering, and PHI screening with IRB approval. The modality-dependence finding (Claude's format-contingent refusal) is the most robust and valuable contribution and survives my concerns; the broader \"guardrail failure\" framing needs the control/reinterpretation work described above.","major_comments":[{"comment":"The headline refusal numbers (3.3% / 0.0% / 66.0%) are presented as guardrail failure, but the paper never defends the premise that these requests ought to be refused. All 30 seed instances are public templates whose visible content is placeholder identity (John Doe); compliance could be calibrated template customization rather than guardrail weakness. The explicit \"Modify this doctor's note\" framing arm (Tables 1-3) partially answers this — GPT/Gemini comply ~always even under medical framing — but the placeholder content remains on-screen throughout. The paper needs either a control arm with unambiguous fraud signals (realistic provider identities, authentic-looking clinical content, or stated deceptive intent) or a substantial reframing toward the well-supported descriptive claim: refusal behavior is inconsistent and modality-driven rather than intent-driven.","section":"§4.2 RQ1 (Tables 1-3)"},{"comment":"The Claude \"collapse\" from 100% (image) to 7.0% (inline text) compares different tasks: image forgery versus text rewriting of user-pasted text. Moreover, by the paper's own Table 4, Claude's compliant inline outputs carry 99.8% CER — they are not usable manipulated notes, only attempts under the loose compliance definition (§4.2). The \"bypassed through simple changes in document modality\" language (§4.2, Discussion) overstates what was demonstrated. The cross-modality policy inconsistency is real and worth reporting, but the usable-output rates per modality should be reported alongside refusal rates, and the bypass framing tempered.","section":"§4.2, Table 3"},{"comment":"CER is operationalized by modality — human judgment of meaningful collateral edits for images, word-level set diff for inline text/PDF — and these are not comparable measures. For text conditions the model must regenerate the full text, so any reformatting or paraphrase of the manually reconstructed seed counts as collateral edit; the >90% CER values (Table 4: Gemini PDF 99.6%, Claude inline 99.8%) are near-artifacts of the metric, not evidence that text workflows 'induce broader unintended modifications.' Table 4 presents the numbers side by side without caveat, and the RQ2 narrative leans on them. Recompute with a semantics-aware measure for text, or restrict cross-modality claims to FSA.","section":"§4.3 RQ2, Table 4"},{"comment":"The interpretation that manipulated notes 'successfully evad[e] human scrutiny' ignores the symmetric failure visible in the reported confusion matrices: precision 45.1% with 125 true flags implies participants flagged ~152/348 authentic seeds (43.7%), accepted fakes more often than reals (64.1% vs 56.3%), and overall accuracy (46.1%) sits below chance on a balanced set. What the study shows is near-random discrimination on this stimulus pool, which is not the same as fakes being indistinguishable — the authentic templates themselves were frequently judged suspicious. The RQ3 conclusions and the abstract's 'visually indistinguishable from original documents' should be reframed accordingly, and the high false-flag rate on reals discussed.","section":"§4.4 RQ3, Figure 2"},{"comment":"Several reporting gaps undermine the 'reproducible pipeline' claim: (a) no access dates or version snapshots for the three systems, though guardrail behavior changes with provider updates; (b) the refusal/compliance classification procedure for the 6,300 responses is undefined — no statement of whether it was automated or manual, by whom, or with what reliability (in contrast to the κ=0.604 reported for image FSA/CER); (c) seed images, substitution dictionaries, prompts, and outputs do not appear to be released; (d) no CIs or tests behind claims such as framing effects being 'minor' (e.g., GPT 2.3-3.8% across n=525 cells). These are fixable within the current study and should be addressed.","section":"§3-4, Experimental Setup"}],"minor_comments":[{"comment":"§4.4 states '120 participants recruited' while the results report 123 recruited and 116 retained. Reconcile.","section":"§4.4"},{"comment":"κ=0.604 is described as 'substantial agreement'; under the common Landis-Koch convention 0.604 sits at the moderate/substantial boundary. Report κ separately for FSA and CER and clarify it is pre-consensus.","section":"§4.3"},{"comment":"Table 4 header reads 'DOCX / Inline Text' but the methodology (§3) describes inline text only; clarify whether DOCX outputs were evaluated.","section":"Table 4"},{"comment":"Editorializing in results ('surprisingly low', §4.2) and typographical issues (e.g., 'question:how robust' in Introduction; 'use commercial LLMs' in Abstract) should be cleaned up.","section":"General"},{"comment":"The abstract lists 'four novel contributions' but enumerates five items (the fifth being the discussion of implications); reconcile the count.","section":"Abstract"},{"comment":"For Gemini's 0/2100, clarify how API-level safety blocks or empty/error responses were classified — under the current compliance definition (any attempt counts), a system-level block returning no content could be misclassified as compliance.","section":"§4.2, Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears formatted for AAAI and fits the AI-safety/ethics scope. My principal reservation, shared with the reader's note you forwarded, is the undefended normative premise behind the RQ1 headline numbers; I note, however, that the explicit \"doctor's note\" framing arm partially answers it, and the modality-inconsistency finding stands independently of how that premise is resolved. A modest control arm (realistic-looking provider identities or an explicit-intent prompt condition) would settle the question and substantially raise the paper's impact. I saw no ethical red flags: seeds were PHI-screened, the user study was IRB-approved, and the paper's stated intent is defensive evaluation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a careful, large-N behavioral study (6,300 attempts across three commercial multimodal models, formats, and prompt frames) showing that simple naturalistic edit requests on doctors’-note templates rarely get refused—especially GPT-image and Gemini—and that the cleanest image edits fool CloudResearch raters at near-chance. That is useful for AI-safety and healthcare-deployment people; it is not a new jailbreak method.\n\nWhat is actually new is the end-to-end pipeline on existing note templates (not free-form sick-note generation), the format × framing factorial, FSA/CER split by modality with dual annotation (κ=0.604), and a filtered believability study with attention check. Design is systematic and readable. Refusal tables are the load-bearing result and look honestly reported. Claude’s image→inline collapse (100%→7%) is the sharpest single finding and is hard to dismiss as noise.\n\nSoft spots, in proportion. The stress-test point partly lands: seeds are public placeholder templates (John Doe etc.), deliberately scrubbed of real PHI/providers, so “replace John Doe with …” is closer to template customization than unambiguous fraud. The paper never runs a clear fraud-control arm (e.g., real letterhead + real provider name + intent to deceive). That weakens the normative claim that every compliance is a guardrail failure—though Claude’s full image refusals show the models sometimes treat the same content as refuse-worthy, so the modality gap still stands. RQ3 only samples FSA=1/CER=0 images, so believability is best-case, not average-case; and raters also flagged many authentic seeds, so “indistinguishable from originals” is inside a pool that already fails casual credibility. Limitations mostly own the template and role-play gaps. No released artifacts in-text; API behavior will drift.\n\nMath is not the point; the operational definitions are fine. Citations cover intended-use medical LLMs, medical safety benches, and general jailbreaks adequately—no weird pattern.\n\nWho it’s for: multimodal guardrail and health-AI safety readers who want concrete misuse numbers, not theorists. Worth a serious referee. I’d engage, cite the refusal/format tables with the template caveat, and not treat the ~36% recall as an operational fraud rate without the average-case arm.\n\nRecommendation: send to peer review; ask for artifact release, clearer threat-model language on templates vs fraud, and either unfiltered or explicitly best-case-labeled human results.","headline":"Solid empirical red-team of medical-note editing: low refusals and weak human detection are real measurements, but the template-proxy and best-case filter blunt how far the risk claim travels.","tokens_in":18006,"tokens_out":633,"would_cite":true,"duration_ms":27047,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Commercial LLM guardrails often fail to stop simple requests to rewrite doctors’ notes, and the best fakes fool human reviewers about as often as chance.","keywords":["LLM guardrails","medical note manipulation","multimodal LLMs","AI safety","healthcare document fraud","refusal behavior","believability study","jailbreaking"],"falsifier":"Repeat the same prompt suite on current commercial multimodal APIs with authentic clinical note layouts (still de-identified) and with real staff verifiers: if refusal stays high across image and text for the same intent, or if staff reliably flag the clean FSA=1/CER=0 edits, the central claim fails.","tokens_in":17726,"feed_emoji":"🩺","tokens_out":933,"duration_ms":23094,"temperature":0.7,"pith_summary":"This paper asks how well today’s commercial multimodal LLMs refuse requests to alter medical excuse notes—substituting names, providers, dates, or conditions on public templates. Across thousands of ordinary API prompts in image, PDF, and plain-text form, refusal is very low for some models and highly format-dependent for others: one system refused almost nothing; another refused every image request but almost none of the same requests written as inline text. When models comply, image edits can be accurate and clean enough that the authors filter to near-perfect substitutions; in a user study where people role-play checking student sick notes, those high-quality fakes are accepted far more often than they are flagged. The authors argue that safety systems are relying on shallow modality cues rather than the intent to forge healthcare documents, and that this already lowers the barrier to document fraud in schools and workplaces.","feed_headline":"LLM guardrails let simple prompts rewrite doctors’ notes","feed_subtitle":"Some models almost never refuse; clean image fakes fool people near chance level","key_machinery":"A reproducible medical-note manipulation pipeline: public seed excuse-note templates in three formats (PNG, PDF, inline text), controlled field substitutions, a 2×2 prompt-framing design, refusal vs compliance scoring, Field Substitution Accuracy (FSA) and Collateral Edit Rate (CER), plus a filtered human believability study on clean image edits.","core_discovery":"Contemporary commercial multimodal LLM guardrails are insufficiently robust against medical-note manipulation. Overall refusal was about 3.3% for GPT-image-1.5, 0% for Gemini 2.5, and 66% for Claude Sonnet 4.6, with Claude’s refusals collapsing from 100% on images to 7% on inline text. Successful high-quality image manipulations are often visually indistinguishable to human raters (roughly 36% recall and 46% accuracy in the filtered believability study).","pith_inferences":["Modality-split safety (strong on scans, weak on pasted text) is a general multimodal failure mode likely to show up beyond excuse notes.","Institutions that only eyeball attached photos of notes are structurally mismatched to models that can produce near-template-perfect images.","Closing the gap may require shared cross-modal policy checks that treat ‘rewrite this medical document’s identity fields’ the same whether the input is pixels or text.","Similar pipelines could stress-test prescriptions, lab reports, and insurance forms the way this paper does for excuse notes."],"forward_implications":["Ordinary users with paid LLM access can often obtain rewritten doctors’ notes without sophisticated jailbreaks.","Guardrail strength can swing dramatically with document format even when the harmful intent is unchanged.","High-fidelity image editing makes forged healthcare paperwork more scalable for absences, exams, and accommodations.","Providers and institutions face rising verification burden and privacy/legal exposure when real names and notes are altered.","Healthcare LLM deployment needs stronger multimodal intent-based refusals and ongoing adversarial testing, not only benign-use benchmarks."],"fun_headline_variants":["Commercial LLM guardrails rarely block medical note rewrites","GPT and Gemini almost never refuse doctor-note manipulation","Claude refusals drop from 100% on images to 7% on text","High-quality note fakes fool raters near chance level","Simple prompts bypass guardrails to alter patient records"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That public web templates, simple everyday API prompts, and online participants role-playing as teaching assistants who only see the cleanest image edits are good enough stand-ins for real clinical documents, real misuse, and real institutional checking.","fun_headline_variants_meta":{"raw":{"variants":["Commercial LLM guardrails rarely block medical note rewrites","GPT and Gemini almost never refuse doctor-note manipulation","Claude refusals drop from 100% on images to 7% on text","High-quality note fakes fool raters near chance level","Simple prompts bypass guardrails to alter patient records"]},"model":"grok-4.5","effort":"low","cost_usd":0.003146,"raw_usage":{"total_tokens":1110,"prompt_tokens":824,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":31464000,"prompt_tokens_details":{"text_tokens":824,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":201,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":824,"tokens_out":85,"duration_ms":5052,"temperature":1.0,"reasoning_tokens":201,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T20:10:29.218844+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Repeat the same prompt suite on current commercial multimodal APIs with authentic clinical note layouts (still de-identified) and with real staff verifiers: if refusal stays high across image and text for the same intent, or if staff reliably flag the clean FSA=1/CER=0 edits, the central claim fails.","supporting_citations":[],"review_version":1}