REVIEW 3 major objections 4 minor 1 cited by
A frontier LLM can grade open-ended medical answers with almost physician-level agreement, but it never shows the clinical caution physicians exercise, and it favors its own model lineage.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-08-03 01:59 UTC pith:7XEZO4ZA
load-bearing objection Useful new German open-response benchmark and a solid abstention/bias analysis, but the headline 'alignment with the physician ceiling' is not robust because the LLM is scored against a 9-rater consensus while the physician ceiling uses 8-rater leave-one-out consensuses. the 3 major comments →
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the MedQADE benchmark, the paper's central discovery is that an automated judge can land within the noise of human-level agreement while lacking the two properties that make human grading trustworthy. The top LLM evaluator achieved κ=0.694 (95% CI ≈ [0.62, 0.75]) against the full physician consensus, overlapping the leave-one-out physician ceiling of κ=0.709 (95% CI ≈ [0.67, 0.75]). Yet on the full 3,800-item set, physicians scaled abstention with item difficulty, while frontier models assigned Correct or Incorrect in every instance. Additionally, all models showed significant self-enhancement and intra-family bias, preferring outputs from their own architecture by up to 11–16 percentage
What carries the argument
The load-bearing machinery is a corrected human-reference comparison plus a bias decomposition. MedQADE provides 3,800 open-response clinical items with physician consensus labels; on a 200-item core, the paper computes a leave-one-out physician ceiling by scoring each held-out physician against the majority consensus of the other eight, which puts the human baseline on the same estimand as an LLM scored against the full consensus. To isolate bias, it defines paired-difference metrics Δself and Δfamily that subtract the average score from independent peer annotators of other lineages from the target model's assigned scores, before and after swapping in architectural siblings. These tools tur
Load-bearing premise
The headline alignment claim depends on the 200-item consensus core being representative of the full benchmark and on comparing a single LLM against the full nine-physician consensus being equivalent to comparing a physician against an eight-physician leave-one-out consensus; wide confidence intervals mean the 'consistent with the ceiling' could also be consistent with meaningfully sub-human performance.
What would settle it
Enlarge the consensus core (e.g., all nine physicians rate 500–1,000 items instead of 200) and recompute the leave-one-out ceiling and the top judge's κ; if the 95% confidence intervals no longer overlap, the alignment claim collapses. Alternatively, present deliberately ambiguous clinical items to the same judges: if the LLM abstains at physician-like rates on those items, the 'absent caution' claim would be refuted.
If this is right
- Automated scoring of open-response German medical answers is feasible at scale, since a single frontier judge can approach the measured human ceiling.
- Benchmarks that report only agreement statistics will miss safety-relevant failures; abstention behavior and lineage-bias metrics should become standard reporting.
- Models from one lineage should not serve as judges for outputs from the same lineage, because measured scores can reflect stylistic affinity rather than medical accuracy.
- Small open-weights models fall well short of a licensing-style 60% accuracy threshold on open German items, so they are not yet suitable for unsupervised clinical use.
- Ensembling multiple LLM judges did not beat the single best judge, suggesting that adding cheaper models can dilute rather than sharpen expert consensus.
Where Pith is reading between the lines
- If the abstention deficit generalizes beyond German neurology, automated judges may silently overrule genuine uncertainty in other clinical languages and specialties, making overconfidence a systems-level risk rather than a mere prompt artifact.
- A testable extension is whether explicit uncertainty prompting or refusal training can restore physician-like abstention without sacrificing kappa; the paper motivates but does not test this.
- The lineage-bias result implies that current model rankings produced by same-family judges may be systematically inflated; cross-family judge panels would provide a stronger audit.
- Because the physician ceiling itself varies per rater (κ from 0.60 to 0.79), part of the remaining gap for any judge may be irreducible noise in the gold standard, not a correctable model deficiency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MedQADE, a German open-response clinical benchmark of 3,800 items drawn from Ankizin, annotated by nine practicing neurologists (plus a pediatrician tiebreaker) and nine LLM evaluators. The central claim is that the best LLM judge, Gemini 3 Flash, reaches alignment consistent with the physician ceiling (κ=0.694 vs κ=0.709), despite near-absent abstention behavior and systematic lineage-based scoring bias. The authors argue that statistical alignment does not ensure clinical caution and that evaluator independence requires explicit verification. The benchmark and annotations are publicly released.
Significance. If the conclusions hold, this is a valuable contribution: it provides the first standardized German open-response clinical evaluation set, with expert annotations and a detailed analysis of LLM-as-a-judge behavior. The abstention analysis and the quantification of self/family bias are informative and likely robust; the finding that LLM evaluators rarely abstain and favor architectural siblings is a useful cautionary result for medical benchmarking. However, the headline alignment claim — that an LLM judge matches the physician ceiling — rests on a methodological asymmetry and on a small 200-item core with wide confidence intervals. The paper's strengths include a publicly released dataset, explicit annotation guidelines, and a transparent bootstrap-based uncertainty assessment, but the central claim requires reanalysis before the conclusions can be accepted as stated.
major comments (3)
- [LLM Annotator Alignment (p. 10, 14) and Discussion (p. 17)] The headline comparison is not on the same estimand. The physician ceiling is computed by scoring each physician against the leave-one-out consensus of the other N−1=8 physicians, while the LLM is scored against the full N=9-physician consensus. A 9-rater majority is a less noisy gold standard than an 8-rater majority, so a single rater's kappa against it is systematically inflated relative to the leave-one-out ceiling. With bootstrapped CIs of width ~0.13, this asymmetry could easily account for the observed Δκ=−0.016. The Discussion's statement that 'the leave-one-out comparison places the LLM and the human reference on the same estimand' is therefore incorrect. Please re-run the comparison with the LLM scored against an 8-physician leave-one-out consensus (or the ceiling computed against the full 9-physician consensus) and report both results.
- [Abstract and LLM Annotator Alignment (p. 13-14)] The claim that Gemini 3 Flash reached 'alignment consistent with the physician ceiling' is too strong given the data. The LLM's bootstrapped 95% CI is [0.619, 0.754] against a ceiling of 0.709 with CI [0.667, 0.746]. Overlap of confidence intervals is not evidence of equivalence; the LLM's CI extends to 0.619, which is substantially below the ceiling. The paper acknowledges the wide CIs, but the abstract and conclusion still frame the result as alignment. The authors should report a confidence interval for Δκ, perform an equivalence test if appropriate, or rephrase the finding as 'cannot be distinguished from the physician ceiling' rather than 'consistent with' it.
- [Annotation Matrix and Rater Distribution (p. 7-8) and LLM Student Performance (p. 9)] In the 3,600 split-annotated items, if one of the two raters abstains, the remaining non-abstaining rater's label becomes the gold label. This effectively creates single-rater ground truth on a subset of items, adding idiosyncratic noise to the student-accuracy estimates and to the difficulty-stratified abstention analysis. The manuscript does not report how often this situation arises or how it is handled in downstream analyses. Please report the frequency of single-rater gold labels (e.g., by item difficulty) and either require two non-abstaining raters or model the missingness. This is a limitation for benchmark validity, even though the alignment analysis on the 200-item core is not directly affected.
minor comments (4)
- [Dataset Formulation and Processing (p. 5)] Minor typo: 'T arget Answer' should be 'Target Answer' in the example.
- [Introduction (p. 3)] Missing space in 'AsLLMs enter clinical documentation' — should read 'As LLMs enter'.
- [Human Inter-Rater Reliability (p. 11)] The phrase 'inter-rater reliability (inter-rater reliability (IRR))' repeats the term; simplify to 'inter-rater reliability (IRR)'.
- [Quantitative Analysis (p. 10)] The definition of the leave-one-out ceiling is clear, but the bootstrap procedure should specify whether resampling is over items or raters; currently it says 'bootstrap resampling (1,000 iterations)' without unit of resampling. Please clarify.
Circularity Check
No significant circularity: the headline alignment, abstention, and lineage-bias results are direct empirical measurements against an external human gold standard.
full rationale
The derivation chain is anchored at an external human gold standard: 3,800 answers were generated by five independent student models, then scored by nine practising physicians (with a tenth tiebreaker) and by nine LLM annotators. The headline alignment estimate (Gemini 3 Flash kappa = 0.694 vs leave-one-out physician ceiling kappa = 0.709) is a direct comparison of an independent LLM's labels to human consensus labels; no parameter was fitted to the human labels to produce this number, and no prediction is generated from the quantity it is compared with. The abstention and lineage-bias analyses are likewise descriptive measurements defined from the annotation matrix (e.g., Delta_self compares a model's own scores against a peer-consensus baseline) and are not used as evidence for their own definitions. The only references to prior work are ordinary support for corpus provenance and known self-enhancement phenomena; Ankizin citations are external resources, and the self-enhancement finding is re-measured here rather than assumed. A reviewer concern that the LLM is scored against a 9-rater consensus while the physician ceiling uses an 8-rater leave-one-out consensus is a possible estimand-comparability limitation, but it is not circular: both quantities are measured from the same independent human ratings, and the comparison is a methodological choice rather than a definitional identity. No self-definitional step, fitted-input-called-prediction, or load-bearing self-citation chain was found.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Ankizin cloze items are a valid proxy for German clinical knowledge
- domain assumption Physician panel (nine neurologists) majority provides a reliable gold standard for correctness
- ad hoc to paper Lexical-match status separates item difficulty
- ad hoc to paper Peer annotators from different architectural lineages provide an unbiased baseline for bias estimation
- domain assumption The 200-item core subset is representative of the 3,800-item benchmark for alignment estimates
Cite this review
Pith. "Pith review of Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking." pith.science (2026). https://pith.science/paper/7XEZO4ZA
@misc{pith2026260701103,
author = {Pith},
title = {Pith review of: Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking},
year = {2026},
howpublished = {\url{https://pith.science/paper/7XEZO4ZA}},
note = {Machine review of arXiv:2607.01103}
}
read the original abstract
Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration and caution, however, remains untested. We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large Language Model (LLM) evaluators. The top-performing evaluator model, Gemini 3 Flash, reached alignment consistent with the physician ceiling (\k{appa} = 0.694 vs. \k{appa} = 0.709), though wide confidence intervals limit interpretation. Despite this statistical alignment, automated evaluators exhibited near-absent clinical metacognition: physicians scaled abstention with item difficulty, while frontier models assigned definitive scores in every case. We additionally quantified systematic lineage-dependent biases, where models preferentially scored architectural siblings, an effect independent of language. These results show that statistical alignment does not ensure clinical caution, and that evaluator independence requires explicit verification.
Figures
Forward citations
Cited by 1 Pith paper
-
Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
In open-ended medical conversations with missing information, LLM judges are more lenient than clinicians and a model's same-provider judge can skew apparent safety rankings.
Reference graph
Works this paper leans on
-
[1]
Fudickar S, Bantel C, Spieker J, Töpfer H, Stegeman P, Schiphorst Preuper HR, et al. Natural Language Processing of Referral Letters for Machine Learning–Based Triaging of Patients With Low Back Pain to the Most Appropriate Intervention: Retrospective Study. Journal of Medical Internet Research. 2024 Jan;26:e46857. doi:10.2196/46857
-
[2]
Klug K, Beckh K, Antweiler D, Chakraborty N, Baldini G, Laue K, et al. From admission to discharge: a systematic review of clinical natural language processing along the patient journey. BMC Medical Informatics and Decision Making. 2024 Aug;24(1). doi:10.1186/s12911-024-02641-w
-
[3]
Maarseveen TD, Glas HK, Veris-van Dieren J, van den Akker E, Knevel R. Improving musculoskeletal care with AI enhanced triage through data driven screening of referral letters. npj Digital Medicine. 2025 Feb;8(1). doi:10.1038/s41746-025-01495-4
-
[4]
Current applications and challenges in large language models for patient care: a systematic review
Busch F, Hoffmann L, Rueger C, van Dijk EH, Kader R, Ortiz-Prado E, et al. Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine. 2025 Jan;5(1). doi:10.1038/s43856-024-00717-2
-
[5]
Development and evaluation of a clinical note summarization system using large language models
Oliveira JD, Santos HDP, Ulbrich AHDPS, Couto JC, Arocha M, Santos J, et al. Development and evaluation of a clinical note summarization system using large language models. Communications Medicine. 2025 Aug;5(1). doi:10.1038/s43856-025-01091-3
-
[6]
Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication
Mandal S, Wiesenfeld BM, Szerencsy AC, Small WR, Major V, Richardson S, et al. Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication. npj Digital Medicine. 2025 Oct;8(1). doi:10.1038/s41746-025-01972-w
-
[7]
PubMedQA: A Dataset for Biomedical Research Question Answering
Jin Q, Dhingra B, Liu Z, Cohen W, Lu X. PubMedQA: A Dataset for Biomedical Research Question Answering. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language August 3, 2026 19/30 Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics; 2019. p....
-
[8]
MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
Pal A, Umapathi LK, Sankarasubbu M. MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering. arXiv preprint arXiv:220314371. 2022. doi:10.48550/ARXIV.2203.14371
-
[9]
Jin D, Pan E, Oufattole N, Weng WH, Fang H, Szolovits P. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences. 2021 Jul;11(14):6421. doi:10.3390/app11146421
-
[10]
Moëll B, Farestam F, Beskow J. Swedish Medical LLM Benchmark: development and evaluation of a framework for assessing large language models in the Swedish medical domain. Frontiers in Artificial Intelligence. 2025 Jul;8. doi:10.3389/frai.2025.1557920
arXiv 2025
-
[11]
Evaluation of the performance of GPT-3.5 and GPT-4 on the Polish Medical Final Examination
Rosoł M, Gąsior JS, Łaba J, Korzeniewski K, Młyńczak M. Evaluation of the performance of GPT-3.5 and GPT-4 on the Polish Medical Final Examination. Scientific Reports. 2023 Nov;13(1). doi:10.1038/s41598-023-46995-z
-
[12]
Luo D, Liu M, Yu R, Liu Y, Jiang W, Fan Q, et al. Evaluating the performance of GPT-3.5, GPT-4, and GPT-4o in the Chinese National Medical Licensing Examination. Scientific Reports. 2025 Apr;15(1). doi:10.1038/s41598-025-98949-2
-
[13]
GerMedIQ: A Resource for Simulated and Synthesized Anamnesis Interview Responses in German
Hofenbitzer J, Schöning S, Sebastian B, Lammert J, Modersohn L, Boeker M, et al. GerMedIQ: A Resource for Simulated and Synthesized Anamnesis Interview Responses in German. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). Association for Computational Linguistics; 2025. p. 1...
-
[14]
Doll N, Buschhoff JS, Satheesh S, Abdelwahab H, Allende-Cid H, Klug K. Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?. arXiv; 2026. Available from:https://arxiv.org/abs/2604.19394. doi:10.48550/ARXIV.2604.19394
-
[15]
Measuring Massive Multitask Language Understanding
Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, et al. Measuring Massive Multitask Language Understanding. In: Proceedings of the International Conference on Learning Representations (ICLR); 2021. . August 3, 2026 20/30
2021
-
[16]
The effect of testing versus restudy on retention: A meta-analytic review of the testing effect
Rowland CA. The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin. 2014;140(6):1432-63. doi:10.1037/a0037559
-
[17]
Procop GW, McCarthy T, Schlinsog A, Ghofrani M. A comparison of free-response and multiple-choice questions on the American Board of Pathology primary certification examinations. Academic Pathology. 2026 Apr;13(2):100248. doi:10.1016/j.acpath.2026.100248
arXiv 2026
-
[18]
Pharmacy Student Performance on Constructed-Response Versus Selected-Response Calculations Questions
Sheaffer EA, Addo RT. Pharmacy Student Performance on Constructed-Response Versus Selected-Response Calculations Questions. American Journal of Pharmaceutical Education. 2013 Feb;77(1):6. doi:10.5688/ajpe7716
-
[19]
The pitfalls of multiple-choice questions in generative AI and medical education
Singh S, Alyakin A, Alber DA, Stryker J, Tong APS, Sangwon K, et al. The pitfalls of multiple-choice questions in generative AI and medical education. Scientific Reports. 2025 Nov;15(1). doi:10.1038/s41598-025-26036-7
-
[20]
Cocchieri A, Ragazzi L, Tagliavini G, Moro G. ReMedQA: Are We Done With Medical Multiple-Choice Benchmarks? In: Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics; 2026. p. 2706-38. doi:10.18653/v1/2026.eacl-long.124
-
[21]
Evaluating clinical AI summaries with large language models as judges
Croxford E, Gao Y, First E, Pellegrino N, Schnier M, Caskey J, et al. Evaluating clinical AI summaries with large language models as judges. npj Digital Medicine. 2025 Nov;8(1). doi:10.1038/s41746-025-02005-2
-
[22]
Chen Y, Wen B, Zulkernine F. A Multiagent Summarization and Auto-Evaluation Framework for Medical Text: Development and Evaluation Study. JMIR AI. 2025 Dec;4:e75932-2. doi:10.2196/75932
-
[23]
LLM Evaluators Recognize and Favor Their Own Generations
Bowman S, Feng S, Panickssery A. LLM Evaluators Recognize and Favor Their Own Generations. In: Advances in Neural Information Processing Systems 37. NeurIPS 2024. Neural Information Processing Systems Foundation, Inc. (NeurIPS); 2024. p. 68772-802. doi:10.52202/079017-2197
-
[24]
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
Ackerman C, Panickssery N. Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct. arXiv; 2024. doi:10.48550/ARXIV.2410.02064
-
[25]
Gemini 3 Flash – Model Card; 2025
Google DeepMind. Gemini 3 Flash – Model Card; 2025. https://deepmind.google/models/model-cards/gemini-3-flash/. August 3, 2026 21/30
2025
-
[26]
Ankizin: Digitale Karteikarten für das Medizinstudium; 2024
Ankizin Project Team, Bundesvertretung der Medizinstudierenden in Deutschland (bvmd) e V . Ankizin: Digitale Karteikarten für das Medizinstudium; 2024. Accessed: 2026-05-18. https://www.ankizin.de
2024
-
[27]
Ankizin Leitfaden 2024 / 2025 / 2026; 2024
Ankizin Project Team, Bundesvertretung der Medizinstudierenden in Deutschland (bvmd) e V . Ankizin Leitfaden 2024 / 2025 / 2026; 2024. Accessed: 2026-07-29. https://docs.google.com/presentation/d/1llOCWYl9SHc_QfiplK0a-Vu-atvpwjWOkqqWLNR_ZEU
2024
-
[28]
Gemini Team, Google DeepMind. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. arXiv; 2025. doi:10.48550/ARXIV.2507.06261
-
[29]
OpenAI. GPT-5 System Card. arXiv; 2025. doi:10.48550/ARXIV.2601.03267
-
[30]
Gemma Team, Google DeepMind. Gemma 3 Technical Report. arXiv; 2025. doi:10.48550/ARXIV.2503.19786
-
[31]
Qwen Team, Alibaba Group. Qwen3 Technical Report. arXiv; 2025. doi:10.48550/ARXIV.2505.09388
-
[32]
Better Zero-Shot Reasoning with Role-Play Prompting
Kong A, Zhao S, Chen H, Li Q, Qin Y, Sun R, et al. Better Zero-Shot Reasoning with Role-Play Prompting. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics; 2024. p. 4099-113. doi:10.18653/v1/202...
-
[33]
Large Language Models Understand and Can be Enhanced by Emotional Stimuli
Li C, Wang J, Zhang Y, Zhu K, Hou W, Lian J, et al.. Large Language Models Understand and Can be Enhanced by Emotional Stimuli. arXiv; 2023. doi:10.48550/ARXIV.2307.11760
-
[34]
The Temperature Feature of ChatGPT: Modifying Creativity for Clinical Research
Davis J, Van Bulck L, Durieux BN, Lindvall C. The Temperature Feature of ChatGPT: Modifying Creativity for Clinical Research. JMIR Human Factors. 2024 Mar;11:e53559. doi:10.2196/53559
-
[35]
GPT-5.4 Thinking System Card; 2026
OpenAI. GPT-5.4 Thinking System Card; 2026. https://openai.com/index/gpt-5-4-thinking-system-card/
2026
-
[36]
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Kim S, Shin J, Cho Y, Jang J, Longpre S, Lee H, et al.. Prometheus: Inducing Fine-grained Evaluation Capability in Language Models. arXiv; 2023. doi:10.48550/ARXIV.2310.08491
-
[37]
Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models
Bosma M, Chi E, Ichter B, Le QV, Schuurmans D, Wang X, et al. Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models. In: Advances in Neural Information Processing Systems August 3, 2026 22/30
2026
-
[38]
Neural Information Processing Systems Foundation, Inc
NeurIPS 2022. Neural Information Processing Systems Foundation, Inc. (NeurIPS); 2022. p. 24824-37. doi:10.52202/068431-1800
-
[39]
Byrt T, Bishop J, Carlin JB. Bias, prevalence and kappa. Journal of Clinical Epidemiology. 1993 May;46(5):423-9. doi:10.1016/0895-4356(93)90018-v
-
[40]
The Measurement of Observer Agreement for Categorical Data
Landis JR, Koch GG. The Measurement of Observer Agreement for Categorical Data. Biometrics. 1977;33(1):159-74. doi:10.2307/2529310
doi:10.2307/2529310 1977
-
[41]
Computing Krippendorff’s Alpha-Reliability
Krippendorff K. Computing Krippendorff’s Alpha-Reliability. University of Pennsylvania; 2011
2011
-
[42]
Approbationsordnung für Ärzte (ÄApprO); 2002
Bundesministerium der Justiz. Approbationsordnung für Ärzte (ÄApprO); 2002. Zuletzt geändert durch Art. 1 V v. 12.1.2023 (BGBl. 2023 I Nr. 18). https://www.gesetze-im-internet.de/_appro_2002/. Supporting information S1 Appendix. System prompt for LLM students.The following system prompt (in German) was used for all LLM student models to generate cloze que...
2002
-
[43]
Liefere a u s s c h l i e s s l i c h das fehlende Wort oder den fe hl en de n kurzen F a c h b e g r i f f als deine Antwort
-
[44]
Verwende au ss ch li e ß lich die korrekte deutsche m e d i z i n i s c h e F a c h t e r m i n o l o g i e
-
[45]
Gib genau eine pr ä zise und s p e z i f i s c h e Antwort
-
[46]
Beispiel : Eingabe : Das Hormon , das den B l u t z u c k e r s p i e g e l senkt , ist ___
Keine zus ä tzlichen Erkl ä rungen , S ä tze , Anf ü hrungszeichen , Nummern , S a t z z e i c h e n oder F o r m a t i e r u n g e n . Beispiel : Eingabe : Das Hormon , das den B l u t z u c k e r s p i e g e l senkt , ist ___ . E rw ar te te Ausgabe : Insulin August 3, 2026 23/30 English translation: You are an e x p e r i e n c e d medical student in a...
2026
-
[47]
Deliver only the missing word or short te ch ni ca l term as your answer
-
[48]
Use only correct German medical t e r m i n o l o g y
-
[49]
Give exactly one precise and specific answer
-
[50]
Example : Input : The hormone that lowers blood sugar is ___
No a d d i t i o n a l explanations , sentences , qu ot at io n marks , numbers , punctuation , or f o r m a t t i n g . Example : Input : The hormone that lowers blood sugar is ___ . Expected output : Insulin S2 Appendix. Human annotation guidelines and LLM annotator prompt.The following system prompt (in German) was used for all LLM rater models to eval...
2026
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.