REVIEW 3 major objections 4 minor 18 references
In open-ended medical conversations with deleted clinical detail, measured safety and model rankings shift with the LLM judge, with a same-provider bump surviving adjustment and all tested judges more lenient than clinicians.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 14:09 UTC pith:SI5TMZ4T
load-bearing objection A transparent, valuable empirical study of evaluator effects in open-ended medical AI evaluation, but its headline same-provider ordering flip currently rests on raw rates rather than severity-adjusted predictions. the 3 major comments →
Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The authors claim that in open-ended medical conversation under missing information, the evaluator is part of what is being measured. Two findings carry the argument. First, judge choice changes apparent safety: the four judges reach only moderate agreement, and a positive same-provider association remains after controlling for general severity, large enough that one model appears to over-commit least under its own provider's judge and near-worst once that judge is excluded. Second, LLM judges are systematically more permissive than clinicians, crediting appropriate uncertainty on roughly two-thirds to five-sixths of items where a stricter clinician credits about half; the authors therefore
What carries the argument
The central mechanism is a missing-information probe: delete the latter half of the final user turn in each clinical conversation, then grade the model's response against a single binary criterion — did it acknowledge missing information, express appropriate uncertainty, or ask a clarifying question, rather than give a confident definitive answer? Because this criterion is not deterministically checkable, scoring is done by LLM judges, and the study treats the judge as part of the measurement. The load-bearing analysis is a vote-level logistic regression with subject-model and judge fixed effects plus a same-provider term, which separates genuine self-preference from the confound of general
Load-bearing premise
The measurement stands on treating 'appropriate response to missing information' as a single binary — acknowledge, express uncertainty, or ask a clarifying question — scored by LLM judges; the paper itself flags that this does not grade the safety of interim advice, so a response that safely conditionally advises can be scored as over-commitment.
What would settle it
Take the stored model responses, strip all stylistic and textual markers that reveal which provider produced them, and re-run the four-judge panel; if the same-provider coefficient drops to zero and the leave-own-provider-out re-ranking disappears, the evaluator-dependence claim would collapse. Alternatively, run the probe on a prespecified 200-item stratified sample with clinician labels on every item: if the LLM judges' permissiveness gap against the stricter clinician vanishes, the optimism claim would collapse.
If this is right
- Any LLM-judged open-ended safety ranking in medicine should report judge reliability and same-provider sensitivity, because the apparent best model can change when a model's own-provider judge is removed.
- Absolute LLM-judged 'appropriate uncertainty' rates in free-text medical answers should be read as optimistic until calibrated against clinicians; the gap here was roughly 14 to 32 percentage points versus the stricter clinician.
- A cross-provider judge panel buys reliability but not validity: aggregating by tie-positive majority matched clinician judgment no better than the worst single judge, while unanimity or excluding the subject's own provider looked more conservative in this sample.
- High closed-ended accuracy can coexist with over-commitment when information is missing; accuracy and safe abstention are separate axes that need separate measurement.
Where Pith is reading between the lines
- The same-provider bump likely generalizes beyond this probe: any rubric-based evaluation that lets a model or its provider's sibling grade its own outputs may over-credit them, so re-scoring existing open-ended medical benchmark results with cross-provider panels is a cheap, testable extension.
- The two independent clinicians disagreed with each other almost as much as the LLM judges disagreed with the stricter clinician (about 52% vs 70% appropriate rates), suggesting that 'safe response to missing information' is partly a normative policy choice; a graded rubric that makes the clarify-before-advising versus conditional-advice policy explicit could shrink both judge-human and clinician-c
- If evaluator choice can flip rankings at this sample size, a practical regulatory implication is that safety claims for conversational medical AI should be accompanied by judge-panel data and a leave-own-judge-out analysis; otherwise a provider could unknowingly rank itself first by judging with its own model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends missing-information stress-testing to open-ended HealthBench conversations: the latter half of the final user turn is deleted, and four LLMs (Claude Opus 4.8, GPT-5.5, Grok 4.3, Gemini 3.5 Flash) are scored on whether the response appropriately acknowledges missing information, expresses uncertainty, or asks a clarifying question. The primary contributions are measurement-oriented. A four-provider judge panel exhibits only moderate agreement (Fleiss' kappa = 0.65), and a vote-level logistic regression with subject and judge fixed effects finds a residual same-provider association (shared +0.35 log-odds, permutation p = 0.04); raw leave-one-provider-out rates change GPT-5.5's apparent standing from second-best to near-worst. A blinded 50-item human subsample shows LLM judges are more lenient than the stricter of the two independent clinicians. A MedQA anchor shows high accuracy and no option-order effect for three of four models, locating the gap in calibration rather than knowledge. The paper is notable for a candid prespecified/post hoc ledger and full release of prompts, completions, judge votes, human labels, and audit files.
Significance. If the ordering claim is placed on a severity-adjusted basis, the paper would make a solid contribution to medical-AI evaluation. Its strengths include a crossed judge grid rather than a self-preference-only design; fixed-effects separation of judge severity from same-provider association; clustered bootstrap CIs; an exact permutation test; and full reproducibility artifacts. The same-provider association is measured, not assumed, so the design is not circular. The main residual risks are that the headline rank flip is not shown after severity adjustment, that the human reference is internally split, and that the binary endpoint may not capture clinical safety. These are fixable with additional analysis or rephrasing.
major comments (3)
- [§3.3, Tables 4–6] The paper's headline ordering shift — GPT-5.5 from 0.14 inappropriate under its own judge to 0.30 under leave-own-provider-out — is computed from raw judge rates. Table 4 shows GPT-5.5 is also the most lenient judge overall (0.84 vs 0.72–0.81), so the raw contrast conflates same-provider preference with general severity. The regression adjustment is never applied to the ordering: the severity-adjusted difference-in-differences (+0.11 for GPT-5.5) and the shared same-provider coefficient (+0.35 log-odds) are reported only as coefficients. Moreover, the exact permutation test has only 24 bijective assignments, so p=0.04 is the coarsest possible resolution, and the GPT-5.5-specific coefficient is not individually significant (CI [−0.10, +1.35]). Please either recompute the leave-one-provider-out ordering after putting all judges on a common severity scale (e.g., residualizing judge fixed ef
- [§3.4, Tables 7–8] The human reference is not stable enough to support the unqualified claim that 'LLM judges are more permissive than clinicians.' The two independent clinicians disagree substantially (O↔G kappa = 0.47; appropriate rates 0.52 vs 0.70). All four judges are significantly more lenient than O, but only GPT-5.5 is significantly more lenient than G; Grok is directionally stricter than G. The paper's own §4.2 concedes the O/G gap may be legitimate policy variation rather than noise. The abstract and Discussion privilege the 'stricter' clinician, but that choice is post hoc. Please frame the permissiveness conclusion as reference-dependent, report which clinician is used, and avoid implying a single clinician standard.
- [§2.3, §5 #5, Appendix D] The open-ended outcome is a single disjunctive binary that can score clinically prudent responses as 'unsafe over-commitment.' Appendix D.1 (postpartum-depression plan) and D.3 (harm-reduction response that explicitly flags the truncated question) are examples where the judge reasons that the model 'proceeds as if it has enough information,' yet the response includes appropriate caveats. Since the paper's title and abstract frame the metric as 'safety' and 'over-commitment,' this construct-validity gap is load-bearing. The limitation is acknowledged, but it is not addressed analytically (e.g., by re-scoring with a graded rubric or by separating recognition of missing information from the safety of interim advice). Please either add a sensitivity analysis with a graded rubric or consistently describe the endpoint as 'appropriate-response recognition' rather than 'safety.'
minor comments (4)
- [§3.3] Report a confidence interval for Fleiss' kappa; currently only a point estimate over 198 items is given.
- [§3.1 / Table 2] The stratum 'admin/rewriting only' is described with n=8 at item level while §3.1 reports 7/33 at prompt level; clarify the relationship between item-level and prompt-level strata.
- [Abstract / §3.3] When reporting p=0.04, specify that the exact permutation test has only 24 bijective assignments; otherwise readers may interpret it as a conventional continuous p-value.
- [Throughout] Minor typographical issues: non-standard ligatures in 'insufficent' and 'coefficient' appear in several places; use standard spelling.
Circularity Check
No significant circularity: the evaluator-dependence findings are measured, disclosed, and not derived from their own assumptions.
full rationale
The paper's central claims are empirical measurements, not derivations that assume their targets. The same-provider association is estimated from a vote-level logistic regression with subject-model fixed effects, judge fixed effects, and a same-provider term; this is a standard decomposition of observed ratings, and the permutation test enumerates all 24 bijective own-judge assignments, so the p-value is not fixed by construction. The leave-one-provider-out ordering shift is explicitly labeled a sensitivity analysis, not a common-scale correction, in Sections 2.6, 3.3, and 5, and the severity-adjusted coefficient is reported separately, so no fitted quantity is renamed as a prediction. The human-permissiveness finding compares each LLM judge against independent clinician labels on a blinded subsample; the stricter clinician O provides a reference independent of the author, and the author-influenced consensus is explicitly secondary. The coarse open-ended criterion and the author's dual role as rater/auditor are disclosed limitations (Limitations #2, #3, #5), not circular reductions. No load-bearing self-citation occurs; the paper cites external prior work for the existence of self-preference bias and does not invoke a uniqueness theorem to force its design. The fact that GPT-5.5 judged its own outputs in the primary run is a potential bias that the paper treats as the object of study, not as a premise that guarantees the conclusion.
Axiom & Free-Parameter Ledger
free parameters (3)
- equivalence region margin =
±5 percentage points
- first 50 HealthBench items =
n=50 consensus subset, first items
- first 100 MedQA items =
n=100 test split, first items
axioms (4)
- domain assumption An LLM judge's rating under the provided 'appropriate response' criterion is a meaningful proxy for clinical safety behavior.
- domain assumption Two independent clinicians' labels on a 50-item subsample provide a valid human reference for the safety criterion.
- domain assumption The author's perturbation audit (33/50 unique prompts) correctly identifies clinical-underdetermined items.
- domain assumption Deleting the latter half of the final user turn is a meaningful test of missing-information behavior.
read the original abstract
Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement. We stress-test four models - three flagships (Claude Opus 4.8, GPT-5.5, Grok 4.3) and one mid-tier model (Gemini 3.5 Flash) - by deleting the latter half of the final user turn in HealthBench conversations, grading responses with a four-provider LLM-judge panel and a blinded clinician-anchored reference. Two evaluator-facing results are robust. First, judge choice materially changes apparent safety: inter-judge agreement is only moderate (Fleiss' kappa = 0.65), and after adjusting for each judge's general leniency (vote-level logistic regression), a positive same-provider association remains (exact permutation p = 0.04; GPT-5.5 ~ +0.10 on the probability scale) - large enough to change which model appears to over-commit least once its own-provider judge is excluded. Second, LLM judges are more permissive than clinicians on a blinded 50-item subsample: all four are significantly more lenient than the stricter independent clinician (crediting appropriate uncertainty on 66-84% of items vs 52%), and three of four than the author-influenced consensus (Grok directional only; judge-vs-consensus kappa = 0.20-0.43). On the author-audited clinical-underdetermined subset the permissiveness gap widened and the point-estimate model ordering held. A closed-ended MedQA anchor confirms accuracy is high and option-order effects are within a +/-5-point equivalence region for three of four models, so the safety gap is about calibration, not knowledge. We release the harness, prompts, per-item outputs, judge panel, perturbation audit, and human-annotation protocol.
Figures
Reference graph
Works this paper leans on
-
[1]
Evaluating the robustness and readiness of large frontier models in health AI applications
Gu, Y., Fu, J., Liu, X., Valanarasu, J.M.J., Codella, N.C.F., Tan, R., Liu, Q., Jin, Y., Zhang, S., et al. Evaluating the robustness and readiness of large frontier models in health AI applications. Nature Medicine, 2026. doi:10.1038/s41591-026-04501-8. (Preprint: The Illusion of Readiness in Health AI , arXiv:2509.18234, 2025.)
arXiv 2026
-
[2]
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., Szolovits, P. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams (MedQA). Applied Sciences 11(14):6421, 2021. arXiv:2009.13081
Pith/arXiv arXiv 2021
-
[3]
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., et al. Large language models encode clinical knowledge. Nature 620:172–180, 2023. doi:10.1038/s41586-023-06291-2
-
[4]
Large Language Models Are Not Robust Multiple Choice Selectors
Zheng, C., Zhou, H., Meng, F., Zhou, J., Huang, M. Large Language Models Are Not Robust Multiple Choice Selectors. ICLR 2024. arXiv:2309.03882
Pith/arXiv arXiv 2024
-
[5]
Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
Pezeshkpour, P., Hruschka, E. Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions. arXiv:2308.11483, 2023
Pith/arXiv arXiv 2023
-
[6]
Language Models (Mostly) Know What They Know
Kadavath, S., Conerly, T., Askell, A., et al. Language Models (Mostly) Know What They Know. arXiv:2207.05221, 2022
Pith/arXiv arXiv 2022
-
[7]
Uncertainty-Based Ab- stention in LLMs Improves Safety and Reduces Hallucinations
Tomani, C., Chaudhuri, K., Evtimov, I., Cremers, D., Ibrahim, M. Uncertainty-Based Ab- stention in LLMs Improves Safety and Reduces Hallucinations. arXiv:2404.10960, 2024
Pith/arXiv arXiv 2024
-
[8]
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L., Chiang, W.-L., Sheng, Y., et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 (Datasets & Benchmarks Track). arXiv:2306.05685
Pith/arXiv arXiv 2023
-
[9]
Self-Preference Bias in LLM-as-a-Judge
Wataoka, K., Takahashi, T., Ri, R. Self-Preference Bias in LLM-as-a-Judge. arXiv:2410.21819, 2024
Pith/arXiv arXiv 2024
-
[10]
LLM Evaluators Recognize and Favor Their Own Generations
Panickssery, A., Bowman, S.R., Feng, S. LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. arXiv:2404.13076
Pith/arXiv arXiv 2024
-
[11]
Arora, R.K., Wei, J., Soskin Hicks, R., Bowman, P., Quiñonero-Candela, J., et al. (Ope- nAI). HealthBench: Evaluating Large Language Models Towards Improved Human Health. arXiv:2505.08775, 2025
Pith/arXiv arXiv 2025
-
[12]
Probable Inference, the Law of Succession, and Statistical Inference
Wilson, E.B. Probable Inference, the Law of Succession, and Statistical Inference. Journal of the American Statistical Association 22(158):209–212, 1927
1927
-
[13]
Measuring nominal scale agreement among many raters
Fleiss, J.L. Measuring nominal scale agreement among many raters. Psychological Bulletin 76(5):378–382, 1971
1971
-
[14]
A Coefficient of Agreement for Nominal Scales
Cohen, J. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1):37–46, 1960
1960
-
[15]
Addressing Benchmarking Gaps in 24 Large Language Models for Health and Medicine with Dynamic Red-Teaming
Pan, J., Jian, B., Hager, P., Zhang, Y., Liu, C., et al. Addressing Benchmarking Gaps in 24 Large Language Models for Health and Medicine with Dynamic Red-Teaming. Nature Health,
-
[16]
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
Pombal, J., Rei, R., Martins, A.F.T. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models. arXiv:2604.06996, 2026
Pith/arXiv arXiv 2026
-
[17]
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
Philipp, W., et al. Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking. arXiv:2607.01103, 2026. 25
Pith/arXiv arXiv 2026
-
[2026]
doi:10.1038/s44360-026-00152-8. (Preprint arXiv:2508.00923.)
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.