Pith. sign in

REVIEW 5 major objections 5 minor 76 references

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Correct medical answers can hide flawed reasoning; pruning to physician-relevant sentences exposes the gap.

desk verdict A genuinely new and useful relevance-label dataset for medical QA, but the headline pruning-accuracy claim needs a random-removal control before it can be trusted. read the letter →

arxiv 2505.24040 v1 pith:NMZI7E4Y submitted 2025-05-29 cs.CL cs.AI

classification cs.CLcs.AI
keywords MedPAIRmedicalquestionansweringrelevancealignmentsentence-levelannotationspuriousratephysiciantraineeslargelanguagemodelscontextattribution
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that correct answers in medical question answering do not reveal whether a model reasoned the way a clinician would: large language models systematically disagree with physician trainees about which sentences in a patient case are relevant. It introduces MedPAIR, a dataset of 1,300 clinical QA pairs with sentence-level relevance labels from 36 physician trainees, and measures how often model-assigned relevance matches those labels. The central result is that when the context is pruned to only the sentences physician trainees judged highly relevant, accuracy improves for the trainees and for every tested language model. The paper reads this as evidence that models often lean on spurious or distracter information, and it proposes a Spurious Rate metric to quantify how often a correct answer flips to wrong after such pruning.

What carries the argument

The load-bearing object is the majority-vote sentence relevance label: each sentence in a patient vignette is classified by physician trainees as high, low, or irrelevant, and these labels define the set $S^+$ that physicians deem sufficient to answer the question. The Spurious Rate is then the fraction of questions a model answers correctly on the full context but incorrectly on $S^+$ alone, with the numerator counting correct-to-wrong flips and the denominator counting all correct answers on the full context. The paper also maps numerical context-attribution scores to the same three categories by keeping the top-$k$ sentences where $k$ is the number of physician-relevant sentences, which is what makes model-human concordance comparable.

What would settle it

Run the same pruning experiment but remove an equal number of sentences selected at random instead of by physician labels; if random pruning produces the same or larger accuracy gains on the same 1,300 questions, then the reported improvements do not depend on the content of the relevance labels.

Watch

Extended reading notes

Core claim

MedPAIR records, for each of 1,300 multiple-choice medical questions drawn from four existing benchmarks, sentence-level trinary labels (high relevance, low relevance, irrelevant) assigned by physician trainees and aggregated by majority vote. Comparing those labels with the relevance estimates produced by four large language models—three via a context-attribution score and one via self-reported sentence judgments—the paper finds agreement on only about 45 to 66 percent of sentences marked highly relevant by clinicians. When the patient case is cut down to the physician-relevant sentences, accuracy rises for both the human labelers and every tested model, with gains as large as roughly 20 percentage points on the hardest benchmark; the Spurious Rate, which counts the fraction of originally correct answers that become wrong after pruning, ranges from about 2 to 18 percent across models. The paper concludes that exam-style accuracy overstates model competence and that expert-labeled relevance exposes a previously unmeasured misalignment.

Load-bearing premise

The paper assumes that the sentences physician trainees mark as highly relevant are by themselves sufficient to answer the question correctly, so any model that is correct on the full case but wrong after pruning must have been relying on something doctors would call irrelevant.

Editorial extensions

If this is right

  • Correct answers on medical QA benchmarks can overstate reasoning quality, because a model may reach the right answer by leaning on sentences physicians mark as irrelevant.
  • Filtering out physician-irrelevant sentences is itself a low-cost intervention that raises accuracy for both humans and models, so context curation should be evaluated alongside model improvements.
  • The Spurious Rate gives a numeric handle on whether a model's correct answers are robust to distractor removal, making it a candidate diagnostic for clinically deployable systems.
  • The largest gains appear on the benchmark with the longest cases and the largest option set, suggesting that relevance misalignment matters most in information-dense clinical scenarios.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The pruning experiment lacks a random-removal control, so part of the reported gain could come from shorter context alone; comparing physician label pruning with equal-length random pruning on the same questions would settle this.
  • The labeled set only includes questions physician trainees answered correctly, so the measured disagreement is conditional on trainee success; alignment on cases that trainees themselves miss is left open.
  • A testable extension is to train a model to predict the physician relevance labels rather than prune at inference time; if the labels carry causal signal, such training should improve accuracy on full unpruned cases.
  • The near-flat position profile of the attribution scores relative to expert labels suggests the attribution estimator itself may rank sentences weakly, so comparing two attribution methods on the same cases could separate model misalignment from estimator noise.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper introduces MedPAIR, a dataset of 1,300 medical QA pairs with sentence-level relevance labels collected from 36 physician trainees, alongside relevance scores from three open-source LLMs via ContextCite and from GPT-4o via self-reporting. The authors measure concordance between physician labels and LLM relevance labels (Table 2), report that concordance is below two-thirds for every model, and then study what happens when low/irrelevant sentences identified by physician trainees are removed before answering (Section 4.3, Figure 3, Table 7). They also define a Spurious Rate (SR) metric that counts correct-to-incorrect flips after pruning (Section 3.2). The paper's main claims are that LLMs are poorly aligned with physician relevance estimates and that filtering out physician-labeled irrelevant sentences improves accuracy for both physician trainees and LLMs.

Significance. If the concordance measurement is taken at face value, MedPAIR is a valuable community resource: it provides a new kind of annotation—sentence-level physician relevance judgments for multiple medical QA benchmarks—and it makes both the annotations and the evaluation code publicly available. The concordance table is a simple, honest descriptive result, and the paper does not fit parameters to produce its main quantities. The causal pruning claim, however, is not currently supported by the experiments, because the improvement is measured only against the full context and not against a random-removal control. The dataset contribution is significant enough to warrant publication after the analysis is corrected or reframed.

major comments (5)
  1. [§4.3 / Figure 3 / Table 7] The central causal claim that pruning low/irrelevant sentences improves accuracy is not supported without a control. The experiment compares the full context S with the physician-selected subset S+ only; there is no condition that removes the same number of sentences at random or matched on length/perplexity. This matters because Table 1 shows the removed sentences are systematically shorter and have higher perplexity than high-relevance sentences, and Table 4 shows that three different pruning criteria (physician labels, ContextCite, GPT-4o self-report) all produce gains on most datasets, with particularly large gains on MedXpertQA. The gains may therefore be a context-length or text-difficulty effect rather than a consequence of physician relevance labels. Please add a random-removal baseline (with multiple seeds) and a length/perplexity-matched removal baseline, and report per-dataset deltas with error bars.
  2. [§3.2 / §4.3 / Table 7] The Spurious Rate definition and the interpretation of correct-to-incorrect flips assume that S+ alone is sufficient for a human to answer correctly. That assumption is not met: Table 7 reports physician trainee accuracy of only 67.2% on the S+ version of the 248-QA subset, so a model can flip from correct to incorrect after pruning simply because S+ is insufficient. Please report human accuracy on S+ for the full 1,300-QA set, or on the same QA subset used for LLM evaluation, and separate cases where S+ is sufficient from cases where it is not; otherwise SR conflates spurious reliance with information insufficiency.
  3. [§3.1.1 / Appendix B] There is a direct inconsistency in the description of the annotation labels. Section 3.1.1 states that each QA is annotated by at least three physician trainees, while Appendix B states that after excluding annotations from labelers who answered incorrectly, each item received between one and three valid annotations. The final 'majority-vote' labels are therefore sometimes single-annotator or two-annotator decisions, not true majority votes. Because all downstream comparisons use these labels as ground truth, please report the distribution of valid labels per QA, the inter-annotator agreement (e.g., Fleiss kappa on the trinary labels), and the number of QAs at each valid-label count.
  4. [§3.3 / Table 2] The ContextCite concordance measurement in Table 2 depends on a matching procedure that is not fully specified. For each QA, k is set to the number of physician 'high' labels, and the k highest ContextCite sentences are labeled 'high'; the remaining sentences are then assigned to 'low' or 'irrelevant' based on score order, but the cutoff between low and irrelevant is not defined. This procedure forces the model to have exactly k high labels and can inflate or deflate agreement relative to the model's natural threshold. Please specify the low/irrelevant boundary, report sensitivity to the choice of k and to that boundary, and consider a rank-based metric that does not depend on the human count.
  5. [§4.1 / Figure 1] The dataset is constructed only from QAs on which physician trainees answered correctly (2,918 of 6,224 labels; 1,300 final QAs). As a result, the relevance labels and all accuracy comparisons are conditioned on physician correctness. The paper's claims about misalignment and pruning should be explicitly scoped to QAs where physician trainees are correct, and the discussion should address whether the conclusions plausibly extend to cases in which physicians themselves err.
minor comments (5)
  1. [Table 4] Table 4 reports only percentage gains without standard deviations, confidence intervals, or significance markers; please add uncertainty estimates.
  2. [Figure 3 / §4.3] The text says 'in round 2, physician trainee only annotated 248 QAs,' but Table 7 appears to report LLM results on the full 1,300 QAs; please clarify exactly which subset each model group was evaluated on in Round 2.
  3. [§3.1.2] Please report the decoding parameters (temperature, top-p, number of samples) used for GPT-4o self-reported labels and for the ContextCite generation runs, because Section 2.2 argues that LLM labels are sensitive to stochastic decoding.
  4. [§3.1 / §4.1] Section 3.1 says 2,000 QA pairs were sampled, but the final dataset is 1,300; please state the sampling fractions and explain the 104-QA exclusion earlier in the main text rather than only in Section 4.1.
  5. [Table 1 / Appendix E.2] Minor typos: 'MedPair' in the Table 1 caption should be 'MedPAIR', and 'an 58.6%' in Appendix E.2 should be '58.6%'.

Circularity Check

1 steps flagged · score 2.0 of 10

No significant circularity: the main accuracy and concordance results are independent human-judgment measurements; only the Spurious Rate metric's causal label is definitional.

  1. self definitional [Section 3.2 (Problem Formulation), Section 4.3, Table 3 caption]
    "f(S)=Y, f(S +)̸=Y : By removing irrelevant sentences, we flip a correct prediction to an incorrect one. This indicates that the model may have been relying on spurious information in S − (i.e. information for which a human deems irrelevant) to make its predictions. ... A higher SR indicates greater reliance on spurious or irrelevant information."

    SR is defined as the proportion of cases where f(S)=Y and f(S+)≠Y, and the paper then labels this event 'reliance on spurious or irrelevant information' (Table 3: 'LLM depend on sentences annotated as low-relevance or irrelevant to arrive at the correct answer'). The label is not an independent finding: the same event occurs whenever S+ alone is insufficient for the model, a possibility the paper assumes away ('under the assumption that the set f(S+) is sufficient for a human to answer the question correctly') and that its own Table 7 contradicts for humans (67.2% accuracy on S+ in the 248-QA subset). Thus the 'spurious reliance' conclusion is by construction the metric's numerator, not an empirical discovery.

full rationale

The paper's principal quantities—physician sentence-level relevance labels, LLM accuracy before/after pruning, and label concordance—are not fitted to the outcomes they predict. The labels were collected from 36 physician trainees before the accuracy comparisons, and the accuracy results are direct empirical measurements on the original and filtered contexts. The top-k normalization for ContextCite (Section 3.3) uses the human count k only as a budget and does not force which sentences are selected, so the concordance in Table 2 remains informative. The pruning experiments compare full context to the high-relevance subset; although there is no random-removal control, that is an experimental confound, not circularity. The only definitional issue is the Spurious Rate metric: it defines 'spurious reliance' as the flip event f(S)=Y, f(S+)≠Y and then reports that event as evidence of dependence on irrelevant sentences, without establishing that S+ alone is sufficient. This is a secondary metric, not the paper's central claim. No load-bearing self-citation or ansatz-smuggling occurs; self-citations in the related-work section are not used to justify the main conclusions. Overall circularity score: 2.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new physical or conceptual entities. Its free parameters are the thresholding choices in the label mapping and binarization. The key assumptions are clinical: that physician trainee labels are a valid gold standard and that S+ retains sufficient information for correct answers. The ContextCite-as-relevance assumption is acknowledged by the authors as imperfect.

free parameters (2)
  • per-QA top-k threshold for ContextCite high-relevance mapping = number of physician-labeled relevant sentences for that QA
    Section 3.3: k is set equal to the count of physician-majority 'high relevance' sentences; the top-k ContextCite scores are then labeled high. This forces the cardinality of model 'high' labels to match humans and affects the concordance measurement.
  • relevance label binarization thresholds = 0.66 and 0.33 on average numeric score
    Appendix B: labels are converted to 1.0/0.5/0.0, averaged, and then thresholded at 0.66 and 0.33 to produce high/low/irrelevant. These thresholds are arbitrary and affect all downstream label-based results.
assumptions (3)
  • domain assumption S+ (physician-labeled relevant sentences) is sufficient for a human to answer the question correctly.
    Section 3.2 states: 'under the assumption that the set f(S+) is sufficient for a human to answer the question correctly.' This is the premise that makes the pruning experiment interpretable as a test of reliance on distracting information.
  • domain assumption Physician trainee majority-vote labels are a valid ground truth for sentence relevance.
    The entire evaluation treats these labels as gold standard (Section 3.1.1). Labels are collected with knowledge of the correct answer and only from trainees who answered correctly, so they embed the correct answer and may not generalize to cases where physicians err.
  • domain assumption ContextCite scores reflect the model's input relevance for QA.
    Section 3.1.2 uses ContextCite to represent open-source LLM relevance. The Limitations section (Section 6) itself notes that 'ContextCite scores do not always accurately capture the relevance of each sentence,' weakening this assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering." pith.science (2026). https://pith.science/paper/NMZI7E4Y

@misc{pith2026250524040,
  author       = {Pith},
  title        = {Pith review of: MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NMZI7E4Y}},
  note         = {Machine review of arXiv:2505.24040}
}
read the original abstract

Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logic, and models may reach accurate conclusions through flawed processes. In this study, we introduce the MedPAIR (Medical Dataset Comparing Physicians and AI Relevance Estimation and Question Answering) dataset to evaluate how physician trainees and LLMs prioritize relevant information when answering QA questions. We obtain annotations on 1,300 QA pairs from 36 physician trainees, labeling each sentence within the question components for relevance. We compare these relevance estimates to those for LLMs, and further evaluate the impact of these "relevant" subsets on downstream task performance for both physician trainees and LLMs. We find that LLMs are frequently not aligned with the content relevance estimates of physician trainees. After filtering out physician trainee-labeled irrelevant sentences, accuracy improves for both the trainees and the LLMs. All LLM and physician trainee-labeled data are available at: http://medpair.csail.mit.edu/.

Figures

Figures reproduced from arXiv: 2505.24040 by the authors.

Figure 1
Figure 1. Study Design. We consolidated four QA data sources into two main components: the patient profile and the query. In the first step, 36 physician trainees and 4 LLMs independently selected the most appropriate answer. In the second step, physician trainees annotated the relevance of each sentence within the patient profile, excluding annotations linked to incorrect answers. Majority voting was used to produce binary r… view at source ↗
Figure 2
Figure 2. Aligning Physician Trainee Annotations with LLM ContextCite Raw Scores Using an Identical Input Context Budget. To compare the LLM-generated ContextCite scores (numerical) with the relevance labels assigned by the physician trainee (three categories) for each sentence, we established a matching metrics between ternary labels and ContextCite scores to map the numerical scores to the categorical labels. For each QA pa… view at source ↗
Figure 3
Figure 3. Effect of Filtering Context on Final Performance. GPT-4o outperforms all tested open-source language models. After removing irrelevant and low-relevance sentences, LLaMA 70B and Qwen 14B demonstrated the most substantial accuracy improvements. In contrast, Qwen 72B occasionally experiences performance drops following the removal process [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Sentence Position Analysis. Plot (a) Distribution of physician trainees’ majority-vote relevance labels by sentence position. Plot (b) Distribution of GPT-4o self-reported relevance labels by sentence position. Plot (c) ContextCite scores across the context for three o…
Figure 5
Figure 5. Figure 5: Centaur Labs Labeling Interface. The physician trainee labelers first answer the classification question, then provide high relevance, low relevant, and not relevant labels to each sentence. MMLU Precision Medicine (193 QAs) JAMA Clinical Challenge (582 QAs) Med Bullet…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

76 extracted references · 58 canonical work pages

  1. [1]

    Can We Use Large Language Models to Fill Relevance Judgment Holes?, May 2024

    Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi. Can We Use Large Language Models to Fill Relevance Judgment Holes?, May 2024. arXiv:2405.05600 [cs]

  2. [2]

    Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answer- ing.Transactions of the Association for Computational Linguistics, 12:681–699, 2024

    Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answer- ing.Transactions of the Association for Computational Linguistics, 12:681–699, 2024. Place: Cambridge, MA Publisher: MIT Press

  3. [3]

    Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing, April 2025

    Jihyun Janice Ahn and Wenpeng Yin. Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing, April 2025. arXiv:2504.01282 [cs] version: 1

  4. [4]

    LLM Stability: A detailed analysis with some surprises.CoRR, January 2024

    Berk Atil, Alexa Chittams, Liseng Fu, Ferhan Ture, Lixinyu Xu, and Breck Baldwin. LLM Stability: A detailed analysis with some surprises.CoRR, January 2024

  5. [5]

    Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, pages 1–16, New York, NY , USA, May 2021. Associatio...

  6. [6]

    LLMs with Chain-of-Thought Are Non-Causal Reasoners.CoRR, January 2024

    Guangsheng Bao, Hongbo Zhang, Linyi Yang, Cunxiang Wang, and Yue Zhang. LLMs with Chain-of-Thought Are Non-Causal Reasoners.CoRR, January 2024

  7. [7]

    Bernd Bohnet, Vinh Q. Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, Tom Kwiatkowski, Ji Ma, Jianmo Ni, Lierni Sestorain Saralegui, Tal Schuster, William W. Cohen, Michael Collins, Dipanjan Das, Donald Metzler, Slav Petrov, and Kellie Webster. Attribute...

  8. [8]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making.Proc. ACM Hum.-Comput. Interact., 5(CSCW1):188:1–188:21, April 2021

Show all 76 references
  1. [9]

    Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians.JAMA Internal Medicine, 184(5):581–583, May 2024

    Stephanie Cabral, Daniel Restrepo, Zahir Kanjee, Philip Wilson, Byron Crowe, Raja-Elie Abdulnour, and Adam Rodman. Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians.JAMA Internal Medicine, 184(5):581–583, May 2024

  2. [10]

    Barnhill, Mar Llamas-Velasco, Gabriela Poch, Sören Korsing, Wiebke Sondermann, Frank Friedrich Gellrich, Markus V

    Tirtha Chanda, Katja Hauser, Sarah Hobelsberger, Tabea-Clara Bucher, Carina Nogueira Garcia, Christoph Wies, Harald Kittler, Philipp Tschandl, Cristian Navarrete-Dechent, Sebastian Podlip- nik, Emmanouil Chousakos, Iva Crnaric, Jovana Majstorovic, Linda Alhajwan, Tanya Foreman...

  3. [11]

    Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

    Hanjie Chen, Zhouxiang Fang, Yash Singla, and Mark Dredze. Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,Proceedings of the 2025 Conference of the Nations of the Americas Chapte...

  4. [12]

    Reasoning Models Don’t Always Say What They Think

    Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato Carson Denison, John Schul- man, Arushi Somani, Peter Hase, Misha Wagner Fabien Roger Vlad Mikulik, Sam Bowman, Jan Leike Jared Kaplan, and others. Reasoning Models Don’t Always Say What They Think

  5. [13]

    Cheng-Han Chiang and Hung-yi Lee. Can Large Language Models Be an Alternative to Human Evaluations? In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ...

  6. [14]

    What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions, May 2024

    Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing. What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influen...

  7. [15]

    Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models, May 2024

    Jaekeol Choi. Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models, May 2024. arXiv:2405.06931 [cs]

  8. [16]

    SelfCite: Self-Supervised Align- ment for Context Attribution in Large Language Models, February 2025

    Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Zejiang Shen, Zhaofeng Wu, Hu Xu, Xi Victoria Lin, James Glass, Shang-Wen Li, and Wen-tau Yih. SelfCite: Self-Supervised Align- ment for Context Attribution in Large Language Models, February 2025. arXiv:2502.09604 [cs]

  9. [17]

    Learning to Attribute with Attention, April 2025

    Benjamin Cohen-Wang, Yung-Sung Chuang, and Aleksander Madry. Learning to Attribute with Attention, April 2025. arXiv:2504.13752 [cs]

  10. [18]

    ContextCite: Attributing Model Generation to Context

    Benjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, and Aleksander Madry. ContextCite: Attributing Model Generation to Context. November 2024

  11. [19]

    Current and future state of evaluation of large language models for medical summarization tasks.npj Health Systems, 2(1):1–13, February 2025

    Emma Croxford, Yanjun Gao, Nicholas Pellegrino, Karen Wong, Graham Wills, Elliot First, Frank Liao, Cherodeep Goswami, Brian Patterson, and Majid Afshar. Current and future state of evaluation of large language models for medical summarization tasks.npj Health Systems, 2(1):1–...

  12. [20]

    Skinner, Ariel Dora Stern, and David Wennberg

    David Cutler, Jonathan S. Skinner, Ariel Dora Stern, and David Wennberg. Physician Beliefs and Patient Preferences: A New Look at Regional Variation in Health Care Spending.American Economic Journal: Economic Policy, 11(1):192–221, February 2019

  13. [21]

    Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. ERASER: A Benchmark to Evaluate Rationalized NLP Models. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors,Proceedings of the 58th Annua...

  14. [22]

    RAGAs: Automated Evaluation of Retrieval Augmented Generation

    Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Nikolaos Aletras and Orphee De Clercq, editors,Proceedings of the 18th Conference of the European Chapter of the Association for Computational ...

  15. [23]

    Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, June 2024

    Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, June 2024. Publisher: Nature Publishing Group

  16. [24]

    CiteBench: A Benchmark for Scientific Citation Text Generation

    Martin Funkquist, Ilia Kuznetsov, Yufang Hou, and Iryna Gurevych. CiteBench: A Benchmark for Scientific Citation Text Generation. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, page...

  17. [25]

    Enabling Large Language Models to Generate Text with Citations

    Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. Enabling Large Language Models to Generate Text with Citations. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6465–6488, S...

  18. [26]

    Koch, Matthias F

    Susanne Gaube, Harini Suresh, Martina Raue, Eva Lermer, Timo K. Koch, Matthias F. C. Hudecek, Alun D. Ackery, Samir C. Grover, Joseph F. Coughlin, Dieter Frey, Felipe C. Kita- mura, Marzyeh Ghassemi, and Errol Colak. Non-task expert physicians benefit from correct explainable ...

  19. [27]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, A...

  20. [28]

    Large Language Models lack essential metacognition for reliable medical reasoning.Nature Communications, 16(1):642, January 2025

    Maxime Griot, Coralie Hemptinne, Jean Vanderdonckt, and Demet Yuksel. Large Language Models lack essential metacognition for reliable medical reasoning.Nature Communications, 16(1):642, January 2025. Publisher: Nature Publishing Group

  21. [29]

    A Survey on LLM-as-a-Judge, March 2025

    Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Yuanzhuo Wang, Wen Gao, Lionel Ni, and Jian Guo. A Survey on LLM-as-a-Judge, March 2025. arXiv:2411.15594 [cs]

  22. [30]

    McKone, Daniel K

    Yuexing Hao, Jason Holmes, Jared Hobson, Alexandra Bennett, Elizabeth L. McKone, Daniel K. Ebner, David M. Routman, Satomi Shiraishi, Samir H. Patel, Nathan Y . Yu, Chris L. Hallemeier, Brooke E. Ball, Mark Waddle, and Wei Liu. Retrospective Comparative Analysis of Prostate Ca...

  23. [31]

    Measuring Massive Multitask Language Understanding, January 2021

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring Massive Multitask Language Understanding, January 2021. arXiv:2009.03300 [cs]

  24. [32]

    Spurious

    Ai Ishii, Naoya Inoue, Hisami Suzuki, and Satoshi Sekine. Analysis of LLM‘s “Spurious” Correct Answers Using Evidence Information of Multi-hop QA Datasets. In Russa Biswas, Lucie-Aimée Kaffee, Oshin Agarwal, Pasquale Minervini, Sameer Singh, and Gerard de Melo, editors,Proceed...

  25. [33]

    RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning

    Congyun Jin, Ming Zhang, Weixiao Ma, Yujiao Li, Yingbo Wang, Yabo Jia, Yuliang Du, Tao Sun, Haowen Wang, Cong Fan, Jinjie Gu, Chenfei Chi, Xiangguo Lv, Fangzhou Li, Wei Xue, and Yiran Huang. RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinic...

  26. [34]

    What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, January 2021

    Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, January 2021. Number: 14 Publisher: Multidisciplina...

  27. [35]

    PubMedQA: A Dataset for Biomedical Research Question Answering

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. PubMedQA: A Dataset for Biomedical Research Question Answering. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors,Proceedings of the 2019 Conference on Empirical Methods in Natural Language...

  28. [36]

    Effective Context Selection in LLM- Based Leaderboard Generation: An Empirical Study

    Salomon Kabongo, Jennifer D’Souza, and Sören Auer. Effective Context Selection in LLM- Based Leaderboard Generation: An Empirical Study. In Amon Rapp, Luigi Di Caro, Farid Meziane, and Vijayan Sugumaran, editors,Natural Language Processing and Information Systems, pages 150–16...

  29. [37]

    GPT versus Resident Physicians — A Benchmark Based on Official Board Scores.NEJM AI, 1(5):AIdbp2300192, April 2024

    Uriel Katz, Eran Cohen, Eliya Shachar, Jonathan Somer, Adam Fink, Eli Morse, Beki Shreiber, and Ido Wolf. GPT versus Resident Physicians — A Benchmark Based on Official Board Scores.NEJM AI, 1(5):AIdbp2300192, April 2024. Publisher: Massachusetts Medical Society

  30. [38]

    Baleen: robust multi-hop reasoning at scale via condensed retrieval

    Omar Khattab, Christopher Potts, and Matei Zaharia. Baleen: robust multi-hop reasoning at scale via condensed retrieval. InProceedings of the 35th International Conference on Neural Information Processing Systems, NIPS ’21, pages 27670–27682, Red Hook, NY , USA, December 2021....

  31. [39]

    Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S

    Shuyue S. Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S. Ilgen, Emma Pierson, Pang W. Koh, and Yulia Tsvetkov. MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning.Advances in Neural Information Processing Systems, 37:28858–28888, Dece...

  32. [40]

    AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution

    Fengyuan Liu, Nikhil Kandpal, and Colin Raffel. AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution. October 2024

  33. [41]

    Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

    Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. October 2022

  34. [42]

    Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R

    Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, Le Hou, Yong Cheng, Yun Liu, S. Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R. Webster, Ewa ...

  35. [43]

    Context Example Selection for LLM Generated Relevance Assessments

    Jack McKechnie, Graham McDonald, and Craig Macdonald. Context Example Selection for LLM Generated Relevance Assessments. InAdvances in Information Retrieval: 47th European Conference on Information Retrieval, ECIR 2025, Lucca, Italy, April 6–10, 2025, Proceedings, Part I, page...

  36. [44]

    Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, A

    OpenAI, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, A. J. Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, Alex ...

  37. [45]

    MedMCQA: A Large- scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

    Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. MedMCQA: A Large- scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering. InPro- ceedings of the Conference on Health, Inference, and Learning, pages 248–260. PMLR, April

  38. [46]

    Bowman, and Shi Feng

    Arjun Panickssery, Samuel R. Bowman, and Shi Feng. LLM Evaluators Recognize and Favor Their Own Generations. November 2024

  39. [47]

    Wan Beom Park, Seok Hoon Kang, Yoon-Seong Lee, and Sun Jung Myung. Does Objective Structured Clinical Examinations Score Reflect the Clinical Reasoning Ability of Medical Students?The American Journal of the Medical Sciences, 350(1):64–67, July 2015

  40. [48]

    Qwen2.5 Technical Report, January 2025

    Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le ...

  41. [49]

    Benchmarking Prompt Sensitivity in Large Language Models

    Amirhossein Razavi, Mina Soltangheis, Negar Arabzadeh, Sara Salamat, Morteza Zihayat, and Ebrahim Bagheri. Benchmarking Prompt Sensitivity in Large Language Models. InAdvances in Information Retrieval: 47th European Conference on Information Retrieval, ECIR 2025, Lucca, Italy,...

  42. [50]

    Towards Human-Centered Explainable AI: A Survey of User Studies for Model Explanations.IEEE Trans

    Yao Rong, Tobias Leemann, Thai-Trang Nguyen, Lisa Fiedler, Peizhu Qian, Vaibhav Unhelkar, Tina Seidel, Gjergji Kasneci, and Enkelejda Kasneci. Towards Human-Centered Explainable AI: A Survey of User Studies for Model Explanations.IEEE Trans. Pattern Anal. Mach. Intell., 46(4):...

  43. [51]

    Thomas Savage, Ashwin Nayak, Robert Gallo, Ekanath Rangan, and Jonathan H. Chen. Di- agnostic reasoning prompts reveal the potential for large language model interpretability in medicine.npj Digital Medicine, 7(1):1–7, January 2024. Publisher: Nature Publishing Group

  44. [52]

    Jung, Maria Zerlik, Waldemar Hahn, Martin Sedlmayr, and Brita Sedlmayr

    Katharina Schuler, Ian-C. Jung, Maria Zerlik, Waldemar Hahn, Martin Sedlmayr, and Brita Sedlmayr. Context factors in clinical decision-making: a scoping review.BMC Medical Informatics and Decision Making, 25(1):133, March 2025

  45. [53]

    Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation, June 2017

    Shikhar Sharma, Layla El Asri, Hannes Schulz, and Jeremie Zumer. Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation, June 2017. arXiv:1706.09799 [cs]

  46. [54]

    Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, April 2025

    Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush V osoughi. Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, April 2025. arXiv:2406.07791 [cs]. 16

  47. [55]

    Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R. Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H. Chen, Nigam H. Shah, Sami Lachgar, Philip Andrew M...

  48. [56]

    Don’t Use LLMs to Make Relevance Judgments.Information Retrieval Research, 1(1):29–46, March 2025

    Ian Soboroff. Don’t Use LLMs to Make Relevance Judgments.Information Retrieval Research, 1(1):29–46, March 2025. Number: 1

  49. [57]

    RadQA: A Question Answer- ing Dataset to Improve Comprehension of Radiology Reports

    Sarvesh Soni, Meghana Gudala, Atieh Pajouhi, and Kirk Roberts. RadQA: A Question Answer- ing Dataset to Improve Comprehension of Radiology Reports. In Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hi- to...

  50. [58]

    Eric Strong, Alicia DiGiammarino, Yingjie Weng, Andre Kumar, Poonam Hosamani, Jason Hom, and Jonathan H. Chen. Chatbot vs Medical Student Performance on Free-Response Clinical Reasoning Examinations.JAMA internal medicine, 183(9):1028–1030, September 2023

  51. [59]

    Lawler, Jimmy Ba, Rahul G

    Augustin Toma, Patrick R. Lawler, Jimmy Ba, Rahul G. Krishnan, Barry B. Rubin, and Bo Wang. Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue- Based Knowledge Encoding, August 2023. arXiv:2305.12031 [cs]

  52. [60]

    Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs

    Li Wang, Xi Chen, XiangWen Deng, Hao Wen, MingKe You, WeiZhi Liu, Qi Li, and Jian Li. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. npj Digital Medicine, 7(1):1–9, February 2024. Publisher: Nature Publishing Group

  53. [61]

    MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs, April 2025

    Juncheng Wu, Wenlong Deng, Xingxuan Li, Sheng Liu, Taomian Mi, Yifan Peng, Ziyang Xu, Yi Liu, Hyunjin Cho, Chang-In Choi, Yihan Cao, Hui Ren, Xiang Li, Xiaoxiao Li, and Yuyin Zhou. MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs, April 2025. a...

  54. [62]

    An automated framework for assessing how well LLMs cite relevant medical references.Nature Communications, 16(1):3615, April 2025

    Kevin Wu, Eric Wu, Kevin Wei, Angela Zhang, Allison Casasola, Teresa Nguyen, Sith Ri- antawan, Patricia Shi, Daniel Ho, and James Zou. An automated framework for assessing how well LLMs cite relevant medical references.Nature Communications, 16(1):3615, April 2025. Publisher: ...

  55. [63]

    CARES: A Comprehensive Benchmark of Trust- worthiness in Medical Vision Language Models.Advances in Neural Information Processing Systems, 37:140334–140365, December 2024

    Peng Xia, Ze Chen, Juanxi Tian, Yangrui Gong, Ruibo Hou, Yue Xu, Zhenbang Wu, Zhiyuan Fan, Yiyang Zhou, Kangyu Zhu, Wenhao Zheng, Zhaoyang Wang, Xiao Wang, Xuchao Zhang, Chetan Bansal, Marc Niethammer, Junzhou Huang, Hongtu Zhu, Yun Li, Jimeng Sun, Zongyuan Ge, Gang Li, James ...

  56. [64]

    Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems

    Qian Yang, Yuexing Hao, Kexin Quan, Stephen Yang, Yiran Zhao, V olodymyr Kuleshov, and Fei Wang. Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ...

  57. [65]

    A survey of datasets in medicine for large language models.Intelligence & Robotics, 4(4):457–478, December 2024

    Deshiwei Zhang, Xiaojuan Xue, Peng Gao, Zhijuan Jin, Menghan Hu, Yue Wu, and Xiayang Ying. A survey of datasets in medicine for large language models.Intelligence & Robotics, 4(4):457–478, December 2024. Publisher: OAE Publishing Inc

  58. [66]

    Meyer, and Steffen Eger

    Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. Mover- Score: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. 17 In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors,Proceedings of the 2019 Con...

  59. [67]

    Gonzalez, and Ion Stoica

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.Advances in Neural Information Processing Sys...

  60. [68]

    Melton, James Zou, and Rui Zhang

    Shuang Zhou, Mingquan Lin, Sirui Ding, Jiashuo Wang, Canyu Chen, Genevieve B. Melton, James Zou, and Rui Zhang. Explainable differential diagnosis with dual-inference large language models.npj Health Systems, 2(1):1–9, April 2025. Publisher: Nature Publishing Group

  61. [69]

    high relevance,

    Yuxin Zuo, Shang Qu, Yifei Li, Zhangren Chen, Xuekai Zhu, Ermo Hua, Kaiyan Zhang, Ning Ding, and Bowen Zhou. MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding, February 2025. arXiv:2501.18362 [cs]. 18 A Dataset Explanation Massive Multitask Language Und...

  62. [73]

    A 29-year-old female presents with low back pain of five days’ duration

  63. [74]

    Her new job involves walking several miles daily across a large facility

  64. [75]

    The pain is localized without radiation; no traumatic history

  65. [76]

    Her new job involves walking several miles daily across a large facility

    Medications: only oral contraceptives. Question:What is the most likely diagnosis? Options: A. bilateral sacral extension B. bilateral sacral flexion C. sacral base posterior D. right-on-right sacral torsion E. sacral base anterior 22 F. right-on-left sacral torsion G. unilate...

  66. [77]

    Towards Digital Sustainability in Health Care: Developing Digital Health Products through Data-Driven User Insights

    Institutional Review Board (IRB) Approvals or Equivalent for Research with Human Subjects 27 Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approv...

  67. [2023]

    arXiv:2212.08037 [cs]

  68. [2024]

    Association for Computational Linguistics

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.