{"id":"6d55a7fe-1758-4fb6-83e4-90d937190d09","arxiv_id":"2504.13475","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An evaluation framework and dataset show that GPT-4 and other LLMs frequently fail to adjust diagnoses when key patient information is perturbed, achieving only 5.28% accuracy on such changed cases.","lead":"This paper tests whether AI medical chatbots notice when key patient details change, such as age, gender, symptoms, or test results. It builds a new question set from medical exam problems and finds that even the best model often misses these changes, sometimes answering as if nothing had changed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 5.28% DAS result rests on physician-assigned labels for perturbed questions, but no annotation reliability is reported and many labels are non-diagnosis meta-options, so low scores could partly reflect labeler judgment rather than LLM insensitivity.","rationale":"The reader's weakest assumption identifies the same foundation: the derived dataset labels must be valid for the sensitivity measurements to mean what the paper claims. This is indeed the most load-bearing condition. The internal numeric contradiction in Section 6.4 (3.08%/3.65% quoted instead of 4.13%/5.25% for checkup) is real but secondary, because either set of values leaves the aggregate accuracy near 5% and does not alter the qualitative conclusion. The lack of a human baseline is not itself decisive, since the physician labels define the target, but the lack of annotation reliability evidence is. The paper deserves credit for releasing code and data, evaluating five models, and using deterministic decoding, but the central quantitative claim is only as strong as the DAS labels. The proposed re-annotation check is feasible and would settle whether the headline figure is robust. If the check passes, conditional acceptance stands; if it fails, the central claim would need to be re-scoped from 'LLMs are insensitive to key medical information' to 'LLMs perform poorly on these specific clinician-labeled perturbed questions.'","tokens_in":12086,"tokens_out":6356,"duration_ms":58821,"concrete_test":"Use the released DiagnosisQA derived datasets and have two independent clinicians, blinded to the original labels, re-annotate a stratified random sample of 200 items from each of the three largest DAS subsets (symptom change, checkup change, and gender change), assigning each item to Same Answer vs Different Answer and selecting the correct option. Compute Cohen's kappa and the label-switch rate. If kappa is below 0.7 or more than 10% of DAS labels move to SAS or to a different option, recompute GPT-4's DAS accuracy on the stable subset and check whether the headline 5.28% figure changes materially.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that LLMs are insensitive to key medical information is quantified by GPT-4's 5.28% accuracy on the Different Answer Subset (DAS). That number is load-bearing only if the DAS ground-truth labels are correct and objective. Section 6 explains that when a perturbation removes the original correct option, the authors add 'None of the above', and when the perturbation creates a clinical impossibility, they add 'The question contains inconsistency'. The gender-change subset alone contributes 378 of the 6,528 DAS items and 121 of GPT-4's 345 correct answers; these are not cases where the diagnosis changed but cases where the model must select a meta-option indicating inconsistency. For the much larger symptom and checkup DAS subsets (2,921+1,068 and 1,720+411 items respectively), whether a single perturbation actually changes the correct diagnosis is a clinical judgment. The paper reports only that 'four professional physicians and experts' assisted with annotation; it gives no inter-annotator agreement, no number of independent labels per item, no conflict-resolution procedure, and no validation of the regex/keyword extraction (Section 3.2) that determines whether the intended perturbation was actually applied. If a nontrivial fraction of DAS labels are annotator-dependent, the reported low accuracies could reflect label noise or benchmark ambiguity rather than LLM insensitivity to key information.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LLMSenEval, a framework for evaluating the sensitivity of large language models to key medical information in clinical diagnosis. The authors construct DiagnosisQA, a set of 4,603 case-based multiple-choice questions derived from MedQA, and generate eight derived datasets by applying two perturbation strategies (change and removal) to four types of key information (gender, age, symptoms, checkup results). They evaluate GPT-3.5, GPT-4, Gemini, Claude3, and LLaMA2-7b on these datasets, partitioning each derived dataset into Same Answer Subset (SAS) and Different Answer Subset (DAS) based on physician annotation. The central empirical claim is that current LLMs, including the best-performing GPT-4, show poor sensitivity to key medical information: GPT-4 answers only 345 of 6,528 DAS questions correctly (5.28%). The paper concludes that LLMs are not yet reliable for clinical diagnosis because they fail to update diagnoses when key information is perturbed.","tokens_in":12365,"tokens_out":4114,"duration_ms":37250,"significance":"The work addresses a genuine and underexplored gap: existing sensitivity evaluations of LLMs focus on irrelevant context and option-order artifacts, whereas clinical reliability depends on sensitivity to clinically meaningful changes. The proposed framework is systematic, the dataset and code are publicly released, multiple LLMs are compared, and the authors include a thoughtful limitations section. If the ground-truth annotations for the perturbed datasets are valid, the reported low DAS accuracies constitute a meaningful and cautionary empirical finding for the medical AI community. The framework itself is reusable and could be extended to other key information types and perturbation strategies. The paper is not circular: the evaluation criteria are defined independently of model behavior, and the numerical claims are falsifiable against the released dataset.","major_comments":[{"comment":"There is an internal numeric inconsistency in the GPT-4 checkup DAS results. The text states that GPT-4 correctly answers 71 and 22 questions from the 1,720 checkup-change and 411 checkup-removal questions, 'with an accuracy of 3.08% and 3.65%, respectively,' but Table 7 reports GPT-4 DAS accuracies of 4.13% for change and 5.25% for removal. The correct percentages from the stated counts are 71/1720 = 4.13% and 22/411 = 5.35%, so the table value 5.25% is also slightly off. This inconsistency affects the reported headline numbers and must be corrected; the text and the table should be reconciled and a single consistent set of values used throughout.","section":"Section 6.4, Table 7"},{"comment":"The DAS ground-truth labels are load-bearing because every DAS accuracy, including the headline 5.28% (345/6,528), depends on the correctness of the physician relabeling of perturbed questions. The paper reports only that 'four professional physicians and experts' reviewed and corrected the derived datasets; it gives no inter-annotator agreement, no number of independent annotations per item, no conflict-resolution procedure, and no adjudication details. Without this information, a nontrivial fraction of DAS labels could be annotator-dependent, and the low accuracies could partly reflect label noise or ambiguous questions rather than LLM insensitivity. Please report annotation reliability statistics (e.g., Cohen's kappa or Fleiss' kappa) and describe the labeling protocol.","section":"Section 4.1 and Section 6 (annotation reliability)"},{"comment":"The DAS mixes two different types of questions: those where the perturbation changes the correct diagnosis to another medical option, and those where the correct label is a meta-option ('None of the above' or 'The question contains inconsistency'). For example, the gender-change DAS has 378 items and GPT-4 answers 121 correctly, but many of these are cases where the model must recognize a clinically impossible scenario (e.g., a male patient with menstruation) rather than reason to an alternative diagnosis. The aggregated 5.28% conflates these abilities, and the low accuracy may partly reflect models' difficulty with meta-options rather than insensitivity to diagnostic information. Please report DAS accuracies separately for diagnosis-change items and meta-option items, or justify why pooling is appropriate.","section":"Section 6, DAS composition and meta-options"},{"comment":"The regex/keyword extraction described in Section 3.2 determines which questions are included in each derived dataset and which key-information value is perturbed, but no validation of this extraction is reported. The derived datasets exclude questions that could not be perturbed due to missing key information, yet the manuscript does not quantify how many questions were excluded per derived dataset or whether the exclusions introduce selection bias in the SAS/DAS partition. Please provide precision/recall or a manual review sample for the extraction, and report the exclusion counts and their overlap across the eight derived datasets.","section":"Section 3.2 and Table 2 (perturbation extraction and dataset filtering)"},{"comment":"Several DAS cells contain only a handful of items (gender removal n=4, age change n=4, age removal n=22), yet the text draws comparative conclusions from these, e.g., 'Claude3 achieves the highest sensitivity to age removal' based on one correct answer out of 22. These percentages have very wide confidence intervals and are not statistically reliable. The central claim about LLM insensitivity should be based primarily on the aggregate DAS or the larger subsets, and the small-sample cells should be explicitly labeled as unreliable or removed from comparative claims.","section":"Section 6.1 and 6.2 (small DAS subsets)"}],"minor_comments":[{"comment":"The paragraph on checkup SAS results says GPT-4 exhibits 'an increase of 0.37% in ∆ accuracy for checkup changes,' but Table 7 shows GPT-4's SAS ∆ accuracy as 0 for change and +0.37 for removal. The text and table contradict each other; please correct the description.","section":"Section 6.4, text vs Table 7"},{"comment":"The derived dataset abbreviations (DGC, DGR, DAC, DAR, DSC, DSR, DCC, DCR) are introduced in Section 3.2 but Table 1 lists them without a separate legend; adding a short expansion or one-sentence definition in the caption would improve readability.","section":"Table 1 and Section 3.2"},{"comment":"Equation (3) defines FIR as #followedInstructionR / #validR, but the text states LLaMA2-7b has FIR=0 while its RR is high; it would be clearer to state explicitly whether plain-text responses (which are valid but not instruction-following) are the cause of the zero FIR.","section":"Section 4.3"},{"comment":"The discussion of 'low sensitivity' vs 'high sensitivity' is somewhat confusing because SAS accuracy stability and DAS accuracy are opposite notions of sensitivity. A short definition of the intended semantics before the discussion would help readers interpret the figure and the qualitative claims.","section":"Section 7 and Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The central claim is defensible and the framework is a useful contribution, but the annotation validity and the internal numeric inconsistency are serious enough that the paper should not be accepted as-is. The lack of inter-annotator agreement and the unvalidated perturbation extraction are the main risks; they are fixable within the manuscript's scope by additional analysis or by softening the claims. The 'first work' novelty claim is somewhat strong but not a blocker. I also note that the paper's own Section 9 limitations are honest and should be kept, perhaps expanded with the annotation reliability caveat."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the framework is a good idea, the dataset is a real resource, and the main finding is worth taking seriously. But the accuracy numbers that get the headline treatment are only as good as the human labels behind them, and right now we have no reliability data for those labels.\n\nWhat's new: this is the first evaluation I know of that systematically perturbs four types of clinically decisive information (gender, age, symptoms, checkup results) with change and removal strategies, and the DiagnosisQA dataset plus its eight variants are public. The SAS/DAS split is a clean way to separate stability from sensitivity, and the result that GPT-4 gets only 5.28% on the Different Answer Subset, with other models near zero on many subsets, is a concrete safety flag for medical AI.\n\nThe soft spots are real but mostly fixable. First, Section 6.4 has a clear internal inconsistency: the body text reports GPT-4 checkup DAS accuracies as 3.08% and 3.65%, while Table 7 correctly shows 4.13% and 5.25% (those body-text numbers belong to the symptom section). A referee would catch that immediately. Second, the DAS ground truth is physician-assigned, but the paper reports only that four physicians were involved. No inter-annotator agreement, no number of independent labels per item, no conflict-resolution procedure. The stress-test concern holds: for symptom and checkup perturbations, whether the correct diagnosis actually changes is a clinical judgment call, so the DAS subsets could contain ambiguous or mislabeled items. The gender-change subset, where the correct answer is often the meta-option 'The question contains inconsistency,' is more objective, but that's also where GPT-4 earns most of its correct answers, so the overall 5.28% is not as clean a summary as it appears. Third, the regex/keyword extraction that decides which questions get perturbed is not validated, so some derived items may not be perturbed as intended. The age DAS subsets are tiny (n=4 and n=22), and the paper does acknowledge that.\n\nNone of this sinks the paper. The central claim that LLMs are insensitive to key medical information is plausible and independently suggested by the qualitative examples. The problems are in the rigor of measurement, not the direction of the result.\n\nWho this is for: people building or evaluating clinical LLMs, and benchmark designers. It deserves serious peer review, with requests for label reliability metrics, a numeric cleanup, and some error analysis on DAS items.\n\nI'd engage with it and cite it as related work, but I wouldn't take the exact DAS percentages at face value before the annotation process is documented.","headline":"A useful new sensitivity benchmark for clinical LLMs, but the headline DAS number rests on unvalidated human labels and Section 6.4 has a swapped-numbers error.","tokens_in":12884,"tokens_out":3769,"would_cite":true,"duration_ms":33322,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current LLMs fail to stay sensitive to diagnosis-changing medical details; GPT-4 answers only 5.28% of altered cases correctly.","keywords":["LLM sensitivity evaluation","clinical diagnosis","key medical information","perturbation strategies","DiagnosisQA","MedQA","multiple-choice question answering","GPT-4 reliability"],"falsifier":"Have an independent panel of physicians re-annotate a random sample of, say, 300 questions from the Different Answer Subset without seeing the paper's labels, and measure agreement on the correct option; if agreement falls below roughly 90%, the reported 5.28% GPT-4 accuracy is contaminated by label noise and is not a clean measure of LLM sensitivity.","tokens_in":11903,"feed_emoji":"🩺","tokens_out":7051,"duration_ms":58625,"temperature":0.7,"pith_summary":"The paper asks whether large language models behave like careful physicians when a small but diagnosis-relevant detail changes. To answer it, the authors build DiagnosisQA from medical exam questions and create eight derived datasets that change or remove four kinds of key medical information: age, gender, symptoms, and checkup results. The central finding is that all five tested models are unreliable on this test: GPT-4, the best, answers only 345 of 6,528 questions correctly (5.28%) when the perturbation changes the correct diagnosis, and several models also lose accuracy on changes that should leave the diagnosis unchanged. The paper argues that accuracy on standard medical benchmarks does not measure clinical reliability, and that LLM development should aim for explicit sensitivity to key information.","feed_headline":"GPT-4 misses 95% of diagnosis-changing edits in medical questions","feed_subtitle":"A new perturbation test on 4,603 clinical cases shows high exam scores do not make an LLM clinically reliable.","key_machinery":"The load-bearing mechanism is a perturbation matrix: four types of key medical information ($K = \\{gender, age, symptom, checkup\\}$) crossed with two perturbation strategies (change and removal), producing eight derived datasets. Each perturbed question is then classified as Same Answer Subset (SAS) or Different Answer Subset (DAS) through physician re-annotation, and two extra options—'None of the above' and 'The question contains inconsistency'—are added when the perturbation removes the correct option or creates a clinically impossible case. Sensitivity is read as the accuracy difference on SAS before and after perturbation (should be near zero) and accuracy on DAS (should be high). The framework also tracks response rate and instruction-following rate to separate format compliance from diagnostic sensitivity.","core_discovery":"The paper's central claim is that current LLMs do not have the sensitivity profile a clinician needs: they should notice when key information changes the diagnosis, and stay stable when it does not, but they do neither reliably. The evidence is a set of eight derived datasets in which each question's gender, age, symptom, or checkup result is either changed or removed, with correct answers re-annotated by physicians. Questions are split into a Same Answer Subset, where the diagnosis is unchanged, and a Different Answer Subset, where the correct answer changes. GPT-4 answers only 345 of 6,528 DAS questions correctly (5.28%), far above the other models but still far too low for clinical use, while models such as Gemini are easily thrown off by symptom changes that should not alter the diagnosis. From this the paper concludes that high performance on medical licensing benchmarks is not evidence of clinical reliability.","pith_inferences":["Beyond the reported five models, the same datasets could be used to test whether medical fine-tuning or larger scale narrows the DAS gap, since the paper does not show that any training recipe repairs it.","The 5.28% figure likely mixes two distinct failures—not noticing that a key fact changed, and noticing but failing to update the diagnosis—and a follow-up that asks models to first state which fact changed would separate them.","The perturbation design could be recycled as a training signal: exposing models to DAS examples with the corrected labels might improve sensitivity without requiring new data collection."],"forward_implications":["Medical LLM evaluation should include sensitivity stress tests, not only accuracy on unperturbed exam questions.","A high score on MedQA-style benchmarks does not imply a model will notice when a patient's gender, symptom, or test result changes the diagnosis.","Models need two opposite skills at once: stability on diagnosis-irrelevant changes and responsiveness to diagnosis-relevant ones; current models fail at one or both.","Instruction-following and sensitivity are separate failure modes: LLaMA2-7b has a 0% followed-instruction rate, while GPT-4 follows instructions almost perfectly yet still misses 94.72% of DAS questions."],"supporting_citations":[{"why":"Supplies the MedQA exam questions from which DiagnosisQA is filtered and perturbed.","marker":"Jin et al., 2021"},{"why":"Provides the negation and keyword extraction method used to locate key information values for perturbation.","marker":"Chapman et al., 2001"},{"why":"Documents LLM sensitivity to option order, the prior sensitivity focus the paper contrasts with key-information sensitivity.","marker":"Pezeshkpour and Hruschka, 2023"},{"why":"Shows LLMs exhibit option bias in multiple-choice questions, motivating a clinical sensitivity test.","marker":"Zheng et al., 2023"},{"why":"Shows LLMs can be distracted by irrelevant context, the counterpart to insensitivity to relevant changes.","marker":"Shi et al., 2023"},{"why":"Reports near-human accuracy on medical knowledge benchmarks, the background against which the paper's low sensitivity results are surprising.","marker":"Singhal et al., 2023"}],"fun_headline_variants":["GPT-4 flunks 95% of diagnosis-change cases","New test: GPT-4 scores 5% on diagnosis-shift questions","LLMs miss 95% of diagnosis-altering medical edits","GPT-4 accurate on just 5% of diagnosis-changing queries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measurement stands or falls on the physician re-annotations of the eight derived datasets being correct and complete, and on the perturbation-and-extraction rules (including the added inconsistency option) producing exactly one valid correct answer for every question.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 flunks 95% of diagnosis-change cases","New test: GPT-4 scores 5% on diagnosis-shift questions","LLMs miss 95% of diagnosis-altering medical edits","GPT-4 accurate on just 5% of diagnosis-changing queries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001018,"raw_usage":{"total_tokens":4283,"prompt_tokens":916,"completion_tokens":3367,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":3292}},"tokens_in":532,"tokens_out":3367,"duration_ms":24167,"temperature":1.0,"reasoning_tokens":3292,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:06:45.879035+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent panel of physicians re-annotate a random sample of, say, 300 questions from the Different Answer Subset without seeing the paper's labels, and measure agreement on the correct option; if agreement falls below roughly 90%, the reported 5.28% GPT-4 accuracy is contaminated by label noise and is not a clean measure of LLM sensitivity.","supporting_citations":[],"review_version":1}