{"id":"3bf2fd87-6796-41e3-a082-44ed0b1a7451","arxiv_id":"2506.22808","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"MedEthicsQA is a 10,974-question medical ethics benchmark with a 4P-26C-256G taxonomy, and evaluation shows MedLLMs decline on ethics tasks relative to their foundation models.","lead":"This paper introduces MedEthicsQA, a benchmark of 5,623 multiple-choice and 5,351 open-ended questions for evaluating medical ethics in large language models. It finds that medical fine-tuned models generally score lower on ethics questions than their base models, suggesting ethics alignment is neglected in current MedLLM training.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline ↓4.4% rests on GPT-4o-generated reference answers scored by GPT-4o-mini, with no human rescoring of the MedLLM-vs-foundation deltas, no error bars, and mixed per-model signs; style bias could explain the decline.","rationale":"The dataset is a valuable resource with a transparent taxonomy, multi-stage filtering, and a validated 1,158-question challenge subset; the concern is not about provenance or integrity but about whether the reported performance decline is a true finding or a measurement artifact. The reader's identification of GPT-4o-generated gold answers is correct and is the weakest point; the stress test sharpens it by adding the judge-bias and no-significance-testing angle. The proposed experiment is feasible: human experts can be blinded and the sample is small. If the re-scoring reproduces the negative delta, the paper's conditional acceptance is warranted; if not, the claim should be downgraded to a benchmark-description-only finding.","tokens_in":23020,"tokens_out":3443,"duration_ms":39091,"concrete_test":"Run a blinded human rescoring experiment: for the 9 MedLLM–foundation pairs in Table 2, sample 50 open-ended questions per pair (450 responses total), and have three medical-ethics experts independently score each response against (a) the original GPT-4o key-point references and (b) key points the experts themselves write from the underlying PubMed passage. Compare the MedLLM-minus-foundation delta from GPT-4o-mini scoring against the human-scored deltas. Also compute the two-sided Wilcoxon signed-rank test over all 5,351 open-ended questions for the average MedLLM-vs-base RS difference. If the human-scored delta is not negative, or if the signed-rank p-value exceeds 0.05, the headline claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that medical fine-tuning degrades (↓4.4%) medical ethics performance—depends entirely on the open-ended Relative Score (RS), which is computed by a GPT-4o-mini checklist judge against key points that GPT-4o extracted from PubMed text (Sec. 2.2.2, Sec. 3.1). The human validations cover data quality (1,200 open-ended items, Tab. 16) and judge reasonableness (200 items, App. E), but no human experts re-score the model responses used in the MedLLM-vs-foundation comparison. The scoring rubric awards points only when the judged response 'conveys similar meaning' and is 'concrete, detailed' (Fig. 20), so response length and phrasing—not ethical correctness—can drive RS. Foundation chat models often produce longer, enumerative answers that mechanically cover more key points, while fine-tuned MedLLMs may give concise, equally ethical answers that score lower; GPT-4o-mini is known to favor such stylistic features. Table 2 itself shows per-model deltas with mixed signs (e.g., Aloe-8b-beta ↑2.1, Meditron3-8b ↓3.7), and the paper reports no standard errors, confidence intervals, or paired significance tests, so the aggregate ↓4.4% (and the parallel ↓4.4 RS decline) could plausibly be a metric artifact rather than a true ethics-alignment gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MedEthicsQA, a benchmark of 5,623 multiple-choice and 5,351 open-ended questions for evaluating medical ethics in LLMs. The authors construct a 4P-26C-256G taxonomy from global medical ethics documents, collect MCQs from medical QA sets and question banks, and synthesize open-ended questions from PubMed passages using GPT-4o with filtering and human validation. They evaluate a range of foundation models, MedLLMs, and proprietary models, and report that medical fine-tuning is associated with an average 4.4% decline in the overall Ethics Score relative to the corresponding foundation models, which they interpret as a lack of medical-ethics alignment in MedLLM training. The paper also reports category-level analyses and a validated 'challenge' subset whose findings are consistent with the main results.","tokens_in":127,"tokens_out":5263,"duration_ms":75529,"significance":"If the central finding holds, the benchmark would be a valuable resource for the medical safety community, particularly because it combines MCQ and open-ended formats and a taxonomy grounded in official international ethics documents. The authors ship a relatively large, publicly released dataset with multi-stage filtering, expert validation on 1,200 items, LLM-consensus checks, and a human-verified judge analysis on 200 items; these are concrete strengths. However, the headline claim that MedLLMs decline by 4.4% is supported only by an internal GPT-4o-to-GPT-4o-mini evaluation pipeline, with no significance testing or human rescoring of the actual model responses used in the comparison. The resource itself may still be useful independent of that particular interpretation, but the main empirical statement needs stronger evidence before it can be accepted.","major_comments":[{"comment":"The headline claim that medical fine-tuning induces an overall performance degradation of 4.4% is an average over nine paired models with mixed signs (e.g., Aloe-8b-beta improves by 2.1 ES points while Meditron3-8b drops by 3.7 points). The paper reports no standard errors, confidence intervals, paired significance tests, or effect sizes. Because the per-model differences are small relative to plausible measurement noise, the aggregate decline may not be statistically meaningful. Please provide a paired bootstrap or permutation test for both Acc and RS, and report per-model confidence intervals.","section":"Section 3.3, Table 2"},{"comment":"The open-ended reference answers are generated by GPT-4o extracting key points from PubMed passages, and the scores are assigned by GPT-4o-mini using a checklist with criteria 'similar meaning' and 'concrete, detailed' (Fig. 20). Human validation checks question/answer quality on a difficulty-biased 22.4% subset (1,200 items) and judge reasonableness on 200 items, but no human expert rescoring is performed on the model outputs used to compute the MedLLM-vs-foundation deltas. As a result, the reported 4.4% decline in relative score could reflect stylistic preferences of the judge (e.g., favoring longer, enumerative responses) rather than ethical quality. I request a stratified human rescoring of the actual model responses for a random sample of paired models, plus a length-controlled sensitivity analysis.","section":"Sections 2.2.2, 3.1, Appendix F"},{"comment":"The MCQ pool is filtered by removing questions 'unanimously answered correctly by all the small-scale models used in Section 3.2.' Since some of those small-scale checkpoints (e.g., Llama2-7b, Llama3-8b) also serve as foundation models in the evaluation, this filtering step removes exactly the items on which base and fine-tuned models might be expected to agree, potentially biasing the observed MCQ accuracy decline. Please list the screening models explicitly and rerun the main comparison on a random unfiltered sample or on the full MCQ pool before filtering.","section":"Section 2.2.1"}],"minor_comments":[{"comment":"The Introduction gives an expert-validation error rate of '2.47%', while the Abstract and Table 16 report '2.72%'. Please clarify the discrepancy and state which fraction corresponds to which annotation aspect.","section":"Introduction vs. Abstract/Table 16"},{"comment":"The text contains 'Ses Figure 10' and should read 'See Figure 10'.","section":"Section 3.1"},{"comment":"The last row of the base-model mapping table does not clearly indicate that Huatuo-o1-70b is based on Llama3.1-70b; please reformat the mapping table to make each MedLLM-to-foundation pairing explicit.","section":"Appendix D, Table 9"},{"comment":"The benchmark name alternates between 'MedEthicsQA' and 'MedEthicQA' (e.g., Figure 3, Table 2). Standardize to one spelling throughout.","section":"Various"},{"comment":"Table 12 contains the typo 'Medirton3-8b' instead of 'Meditron3-8b'.","section":"Appendix D, Table 12"},{"comment":"The caption says 'validated ratings' but the figure itself is not legible in the submitted manuscript; provide a higher-resolution version.","section":"Appendix E, Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The benchmark resource is potentially useful, but the main empirical conclusion needs stronger statistical and human-validation support. I would recommend inviting a revision that adds significance testing, human rescoring of the actual paired model outputs, and a clearer account of the filtering decisions, rather than accepting or rejecting as is. The taxonomy construction and dataset release are clear strengths."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on MedEthicsQA. The resource is real: 5.6k MCQs and 5.4k open-ended ethics questions with a 4P-26C-256G taxonomy, and the paper's quantification that under 2% of widely used medical benchmarks are ethics-related is a useful gap analysis. The curation pipeline is standard but careful: keyword search, three-model consensus filtering, dedup, and removal of easy questions. The open-ended synthesis from PubMed with key-point extraction is a reasonable way to get grounded questions at scale. Credit where due: the dataset is large, openly licensed, and includes a challenge subset; the human validation that exists (1,200 open-ended items, plus judge-reasonableness on 200 pairs) is thoughtfully designed, even if limited.\n\nThe soft spots are the ones the stress-test highlights. The headline 4.4% decline rests on open-ended RS scores assigned by GPT-4o-mini against key points extracted by GPT-4o. The rubric rewards 'concrete, detailed' phrasing, so a concise but ethically correct answer can score lower than a verbose one. That is a real confound, and the paper does not report confidence intervals or significance tests on any of the model deltas. The per-model results are mixed (Aloe-8b-beta up, Meditron3-8b down), so the aggregate could be driven by a few models. That said, the stress-test may overstate the vulnerability: the MCQ subset, which is not subject to the style confound, also shows an overall 4.0-point decline for MedLLMs, and the challenge-subset validation replicates the direction with a 0.98 correlation. So the central finding is plausible, but not proven at the reported precision.\n\nOne more soft spot: human validation covers 22.4% of open-ended questions and samples the hardest ones, which is fine for estimating a lower bound on low quality, but it does not verify that the reference key points are ethically correct ground truth. The authors acknowledge this implicitly by calling it synthesis rather than gold-standard construction.\n\nWho is this for? Anyone building or evaluating medical LLMs with an eye on safety. The benchmark is a useful measuring stick; the empirical claim about fine-tuning tax is secondary and should be read cautiously. I would send it to peer review: the resource deserves scrutiny and the authors have shipped data. But the revision needs error bars, separate reporting for MCQ and open-ended subsets with significance or effect sizes, and preferably release of the full human annotations. As is, treat the 4.4% as a directional signal, not a measured fact.","headline":"MedEthicsQA is a genuinely useful new benchmark with a real gap analysis, but the headline 4.4% decline of MedLLMs is a directional signal that needs error bars and fuller human validation before being treated as fact.","tokens_in":23896,"tokens_out":2684,"would_cite":true,"duration_ms":28956,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Medical fine-tuning cuts LLM ethics scores by 4.4%","keywords":["medical ethics","benchmark","large language models","MedLLM","ethics alignment","fine-tuning tax","LLM-as-Judge","question answering"],"falsifier":"Re-score the 1,158 human-validated challenge questions using reference answers written independently by clinicians rather than extracted by GPT-4o; if MedLLMs no longer underperform their foundation models under those expert-authored references, the fine-tuning-tax claim fails. A simpler check would compute the Ethics Score gap on only the human-validated subset and see whether the 4.4% decline still persists.","tokens_in":1577,"feed_emoji":"⚖️","tokens_out":1647,"duration_ms":86404,"temperature":0.7,"pith_summary":"The paper introduces MedEthicsQA, a benchmark of 5,623 multiple-choice and 5,351 open-ended questions for testing whether large language models follow medical ethics. Its central finding is that medical-domain fine-tuning does not help and on average hurts: medical large language models (MedLLMs) score about 4.4% lower than their foundation models on the overall Ethics Score, even though the same MedLLMs do better than their foundations on ordinary medical-knowledge benchmarks. The authors attribute this decline to a neglect of medical ethics alignment in current training, possibly a fine-tuning tax in which heavy medical-knowledge training crowds out general ethical reasoning. A careful reader should care because it suggests clinical AI safety cannot be inferred from medical QA performance, and ethical behavior needs to be measured and trained as its own capability.","feed_headline":"Medical fine-tuning cuts LLM ethics scores by 4.4%","feed_subtitle":"An 11k-question benchmark shows MedLLMs beat base models on medical knowledge but fall behind on ethics.","key_machinery":"The benchmark itself is the central object. MedEthicsQA combines 5,623 multiple-choice questions filtered from existing medical QA datasets and question banks with 5,351 open-ended questions synthesized by GPT-4o from 14k pages of PubMed medical-ethics literature, each open-ended question carrying a reference answer broken into key points. All questions are organized under a hierarchical taxonomy, 4P-26C-256G: four pillar principles (beneficence, non-maleficence, autonomy, and justice), 26 categories, and 256 detailed guidelines drawn from worldwide medical codes of conduct. Evaluation uses a checklist-based LLM-as-Judge protocol in which GPT-4o-mini awards partial credit for each reference key point the model response covers, and the overall Ethics Score averages MCQ accuracy with that relative score. The argument's load-bearing comparison is the paired difference between each MedLLM and the foundation model it was fine-tuned from.","core_discovery":"The paper claims that current medical-domain fine-tuning does not improve, and on average degrades, performance on medical ethics questions. On MedEthicsQA, the eight MedLLMs evaluated show an average decline of 4.4% in the overall Ethics Score relative to their foundation models, with MCQ accuracy down about 4.0 points and open-ended relative scores down about 4.4 points. The authors argue this decline reflects a neglect of medical ethics alignment in training and a possible fine-tuning tax, where overtraining on medical knowledge causes models to forget general ethics knowledge useful in ethical dilemmas. They report that the decline persists under few-shot prompting and on a human-validated challenge subset, which strengthens the claim that the result is not merely an artifact of the synthetic open-ended questions.","pith_inferences":["An extension the authors do not test is whether adding ethics-specific instruction tuning to a MedLLM closes the 4.4% gap while preserving medical MCQ accuracy; the benchmark would support such a before-and-after study.","Because the open-ended reference answers are LLM-extracted key points from PubMed text, the absolute relative scores partly measure agreement with GPT-4o's judgment; re-scoring a sample with clinician-authored key points would test whether the MedLLM-versus-foundation gap survives a change in reference authorship.","The same fine-tuning-tax pattern may appear in other professional domains with normative content, such as legal or financial ethics; a parallel benchmark in those fields would show whether the effect is specific to medicine or general to fine-tuning on factual knowledge."],"forward_implications":["Medical-domain fine-tuning alone does not confer ethical safety, so ethics must be evaluated separately from medical knowledge in any clinical deployment pipeline.","Ethics alignment should be an explicit objective in MedLLM training, since the evidence indicates it is currently being crowded out by medical-knowledge training.","Evaluation of medical models should include both multiple-choice and open-ended formats, because improvements on one format observed in several models did not generalize to the other.","The benchmark's long-tail category distribution suggests that patient-centered ethical considerations dominate current medical-ethics data, leaving physician-centered categories such as reporting misconduct and managing conflicts of interest relatively undersupplied.","The 4P-26C-256G taxonomy provides a reusable structure for linking model behavior to globally recognized medical ethical standards rather than a single country's code."],"supporting_citations":[{"why":"Supplies the four pillar principles of medical ethics that structure the benchmark's taxonomy and category labels.","marker":"Beauchamp, 1979"},{"why":"Defines the medical-safety notion the paper expands and provides the MedSafetyBench comparison used in the benchmark statistics.","marker":"Han et al., 2024a"},{"why":"Provides the checklist-based LLM-as-Judge protocol the paper adapts for scoring open-ended responses against reference key points.","marker":"Que et al., 2024"},{"why":"Contributes the consensus-based LLM classification and filtering practice used to identify ethics-related MCQs and validate synthetic questions.","marker":"Li et al., 2024b"},{"why":"One of the evaluated MedLLMs and the citation for the observation that even incorporating medical guidelines into training still leads to degraded ethics performance.","marker":"Chen et al., 2023"},{"why":"Supplies the fine-tuning-tax concept the paper uses to explain why medical fine-tuning hurts ethics scores.","marker":"Yuan et al., 2024"},{"why":"Supports the reported scaling-law pattern in which overall Ethics Score improves with model size.","marker":"Kaplan et al., 2020"},{"why":"Identifies GPT-4o as the model used for open-ended question synthesis and GPT-4o-mini as the judge, both load-bearing for dataset construction and evaluation.","marker":"Achiam et al., 2023"}],"fun_headline_variants":["Medical fine-tuning degrades LLM ethics by 4.4%","LLM ethics slip after medical specialization","MedLLMs lose ground on ethics benchmarks","Fine-tuning doctors not ethicists: LLM ethics drop","Medical fine-tuning costs AI ethics performance"],"cache_read_input_tokens":25984,"weakest_assumption_plain":"The whole open-ended evaluation, including the reported 4.4% decline, rests on the assumption that GPT-4o-generated reference answers extracted from PubMed passages are a valid ground truth for medical ethics, with human experts checking only 1,200 of the most challenging questions (22.4%) to confirm this.","fun_headline_variants_meta":{"raw":{"variants":["Medical fine-tuning degrades LLM ethics by 4.4%","LLM ethics slip after medical specialization","MedLLMs lose ground on ethics benchmarks","Fine-tuning doctors not ethicists: LLM ethics drop","Medical fine-tuning costs AI ethics performance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000493,"raw_usage":{"total_tokens":2386,"prompt_tokens":877,"completion_tokens":1509,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":1436}},"tokens_in":493,"tokens_out":1509,"duration_ms":12811,"temperature":1.0,"reasoning_tokens":1436,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:56:24.736109+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the 1,158 human-validated challenge questions using reference answers written independently by clinicians rather than extracted by GPT-4o; if MedLLMs no longer underperform their foundation models under those expert-authored references, the fine-tuning-tax claim fails. A simpler check would compute the Ethics Score gap on only the human-validated subset and see whether the 4.4% decline still persists.","supporting_citations":[],"review_version":1}