Pith. sign in

Paper Citation Record · LEDGER

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

As of 7 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 3 inbound Pith citation observations for arXiv:2505.24040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24040 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:53.123407Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T17:02:30.696295Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:19:34.994254Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy46
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 896e37bd-4420-440b-b58a-955634184988 · outbound

This paper cites Can We Use Large Language Models to Fill Relevance Judgment Holes?.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Can We Use Large Language Models to Fill Relevance Judgment Holes?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:47.373137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:47.373137Z digest=sha256:d530d099841afc5a3d1b561f08e2ab1a88b5e073446d201f54685082e8e3413c

Observation 4ba8765b-c6d5-4427-a016-69206bdb2628 · outbound

This paper cites Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answer- ing.Transactions of the Association for Computational Linguistics, 12:681–699, 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answer- ing.Transactions of the Association for Computational Linguistics, 12:681–699, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.540335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.454079Z digest=sha256:71e5005a0d386878e9c1c6d3acb4f18288c7b1b449adea0e9910797cb5a1f8c4

Observation edb08fed-c0a0-4444-a431-21e84c342fa6 · outbound

This paper cites Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:47.520925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:47.520925Z digest=sha256:7417d036ea495d9f41a83fc7d6f55b5954e059edfa82d9654b1381516f485345

Observation dcc6e987-7e02-414b-b536-0475620f76c7 · outbound

This paper cites LLM Stability: A detailed analysis with some surprises.CoRR, January 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering LLM Stability: A detailed analysis with some surprises.CoRR, January 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.374152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.573430Z digest=sha256:d7036067df94aecbabc70023863ccdb7b58069b734135d0b6d26d8f95b89e1a4

Observation a7088955-a11b-40fc-b0f9-5cbba743a2c1 · outbound

This paper cites Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.204581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.618744Z digest=sha256:487f2d76637c3ac641d5c6edcd430a4e01fca67be5b56415c3e7c969b1114d17

Observation 5c5091e5-d6b4-4eb1-9bbc-097a5aa693bf · outbound

This paper cites LLMs with Chain-of-Thought Are Non-Causal Reasoners.CoRR, January 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering LLMs with Chain-of-Thought Are Non-Causal Reasoners.CoRR, January 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:01.026316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.677371Z digest=sha256:f65b09fdfdb1d4749534dc39a5d26a1c7ce8ab8d661ec301d15cedb27ac74cd3

Observation c962acac-362b-44e9-a52c-ab7da4651209 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:00.853146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.733566Z digest=sha256:cc10055988d77a7c55d9975c1f20d84d1b8d39596ee0a1d56ad7821d71bbbdad

Observation 2dab7938-f13c-4679-b569-71c0f1da83aa · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:00.766014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.825672Z digest=sha256:f796a300eae582caaa0d9063c12a125f329923c98199317ebdf650b0ae2ac8e6

Observation 605ec00f-5428-4b24-8de1-0fb383a0b60b · outbound

This paper cites Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians.JAMA Internal Medicine, 184(5):581–583, May 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians.JAMA Internal Medicine, 184(5):581–583, May 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.641860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.859493Z digest=sha256:9fd5aa2c9d4896fa938a306820e241cb33771cbb2a29074987ff4b386b1645d3

Observation 0022eabe-eae4-4e43-a1d3-c13a3cc5b603 · outbound

This paper cites Barnhill, Mar Llamas-Velasco, Gabriela Poch, Sören Korsing, Wiebke Sondermann, Frank Friedrich Gellrich, Markus V.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Barnhill, Mar Llamas-Velasco, Gabriela Poch, Sören Korsing, Wiebke Sondermann, Frank Friedrich Gellrich, Markus V

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.520737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.899250Z digest=sha256:a42aedc30c0bef2fac463677fcf1feee34c704cfb0035c07912a159045086f23

Observation 1aa9cc7f-e276-443d-bc5e-9ac402ac6630 · outbound

This paper cites Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.370899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:47.933942Z digest=sha256:074825e4f86ff0ae6c2b69bc69d945c6201b574d7f31e29bfa926868e2c5e5cc

Observation 60895be4-10d9-4b8e-81bf-5fa59db03965 · outbound

This paper cites Reasoning Models Don’t Always Say What They Think.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Reasoning Models Don’t Always Say What They Think

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.262806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.003702Z digest=sha256:16539b064528de50d3818f1a39c86545d7ba4bd95b231d80d495a2a8b2f5094f

Observation a333b942-e8db-48f9-a0b5-f255f65bab64 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:43:00.169216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.037851Z digest=sha256:d5ca25d9f817fceee5244fd13881abed0e94894208aa12601a437b68951f25d0

Observation 3c51a8e8-fc08-426b-8379-074619e98521 · outbound

This paper cites What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.072916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.072916Z digest=sha256:bdf59b3f19dc7a4e7fad196fa0f4331fd32fcfa9dc0acb9f494f0613a4522483

Observation 3bebac39-b22f-4dc8-b55c-b2c82906396f · outbound

This paper cites Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:42:53.527784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.135694Z digest=sha256:0f9c72086ac44dca70ddf4699cefd08ec8d5a93be08fce937fd3e284b5bd056a

Observation e558dbca-b791-4367-8d0d-00a90e38525e · outbound

This paper cites SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.196183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.196183Z digest=sha256:1b48ce9066c1036823f1be9e9387bc04d20600bde417f9eb60f6a84d667af6a2

Observation 92323745-e460-41c4-859c-9603845f493c · outbound

This paper cites Learning to Attribute with Attention.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Learning to Attribute with Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.271464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.271464Z digest=sha256:2aec7b40d31dba1daa311bc8ba060f2a16d9526c0cce28dcbc5a4e06e963bff4

Observation dbbf625c-d841-4454-9cb7-1168e8422158 · outbound

This paper cites ContextCite: Attributing Model Generation to Context.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering ContextCite: Attributing Model Generation to Context

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:00.051715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.295228Z digest=sha256:0f80b53f561031ecfcf1cff3022fa121dfb791603c916cc55e8c386dbe6695f8

Observation 25c4cf0f-5e78-460e-a01f-5b4faf721827 · outbound

This paper cites Current and future state of evaluation of large language models for medical summarization tasks.npj Health Systems, 2(1):1–13, February 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Current and future state of evaluation of large language models for medical summarization tasks.npj Health Systems, 2(1):1–13, February 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.957124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.349220Z digest=sha256:657e68110d1bfd086dc56be1a61b7e6236bda52a17a328561006e214fa42b55d

Observation 5cd13f48-69b2-4551-8d11-b8913902161d · outbound

This paper cites Skinner, Ariel Dora Stern, and David Wennberg.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Skinner, Ariel Dora Stern, and David Wennberg

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.790928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.427677Z digest=sha256:085623eb9c2977a597c3edfafc168eaa3ee486e2392751b78e54c454772f69be

Observation 5f85e1c6-8d58-4b93-af8a-3d4ccc12c41c · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:59.646744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.515752Z digest=sha256:aa07e284ee157c6f843adbcfa596450c001bb75b65575215b669f5903c9a7534

Observation 8646b9b0-66ce-46f4-a2ce-7eab95d78398 · outbound

This paper cites RAGAs: Automated Evaluation of Retrieval Augmented Generation.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering RAGAs: Automated Evaluation of Retrieval Augmented Generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.489396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.596357Z digest=sha256:2b171bf698bb2757f7994edcd99e27328a0dccd7bc78d28aadde055cbf5c1f5d

Observation ebedf70f-cde7-4367-bdd8-4953545d8098 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, June 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, June 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.358407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.762620Z digest=sha256:f752322b5836b3e3b7e8e13845d59a9d7841bf7e26dfc538ba4e1699e43d358f

Observation 956e18ff-2a35-4c26-9997-0ba048f0648e · outbound

This paper cites CiteBench: A Benchmark for Scientific Citation Text Generation.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering CiteBench: A Benchmark for Scientific Citation Text Generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.208232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.836182Z digest=sha256:0da9553e950bbb9c16b6e01d4e1a6b7bea16a58b7e751a047d274ede4b3cf380

Observation d2fafe29-dff1-4c5e-8cf7-779d624989c9 · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Enabling Large Language Models to Generate Text with Citations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.103626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:48.958146Z digest=sha256:84c94b5347b0b801a25c951aa2a02cd3f77a222eee016ea23baaf2d15162096c

Observation 7ac6e3ac-3e43-45eb-9f47-5d1ed96a87bc · outbound

This paper cites Koch, Matthias F.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Koch, Matthias F

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:59.021226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.059462Z digest=sha256:9b4110aaa40e4421e7238d379c8fd39bc6716b2cf5f3f9958af931769ad660fe

Observation 933f896e-63d6-4172-a219-43e18b2ff4bb · outbound

This paper cites The Llama 3 Herd of Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:49.147633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:49.147633Z digest=sha256:8423786a84afc9a0aa12dc175b391201d79913e9f7c47d664c7b07e947cce5e1

Observation 3e4e0f94-6992-457e-849c-f11294198349 · outbound

This paper cites Large Language Models lack essential metacognition for reliable medical reasoning.Nature Communications, 16(1):642, January 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Large Language Models lack essential metacognition for reliable medical reasoning.Nature Communications, 16(1):642, January 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.926035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.223937Z digest=sha256:db12af576357e0d8d4e5c90359f1097d67b9048307ebb9f19f9e726478e296d8

Observation 858e417c-5da7-4210-8a11-adbd5dc72bad · outbound

This paper cites A Survey on LLM-as-a-Judge.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering A Survey on LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:49.331240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:49.331240Z digest=sha256:46c75309c1367b883a4169dbabd24dcb03074cac19677f021a36009c6cd8cc6a

Observation 414ad4e3-3197-4a1c-9a74-43f0ede88883 · outbound

This paper cites McKone, Daniel K.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering McKone, Daniel K

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.849701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.421050Z digest=sha256:b025d8d5f295b46eaefc4f5354a72be64c3fffee066994fb485378bc43ce0a7d

Observation 31db672f-c43e-4c3f-88fc-749c9f84ae2d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Measuring Massive Multitask Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:49.523183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:49.523183Z digest=sha256:8c3f8b4d1dd2a3b9d991f859a4a74f603a766747169be2ac775518c97815e3bb

Observation 49b3ef4f-48b4-4ec8-89a4-4895b76b0d5b · outbound

This paper cites Spurious.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Spurious

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.731232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.623543Z digest=sha256:a3352aa70806811ea84551fb23827c23a7601a1cb45bfa8aefe33979397eae35

Observation 554300d1-91e4-4681-a9c6-7bc73458aaec · outbound

This paper cites RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.692217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.696190Z digest=sha256:39586575287eb4cc440fdc19c22d3c5c0089723b55fdbe55757f4c7fb59738e0

Observation 00d8c488-aa8c-4d0a-9c97-98bf7547a73a · outbound

This paper cites What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, January 2021.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, January 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.634175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.775976Z digest=sha256:9d9cfc88230caa9582ad816cf815eaef598278c68ea96c75a4d572ec33b4d959

Observation e5109b91-6a33-4bdf-baf7-1efaf4f3b2da · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.548413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.864087Z digest=sha256:a6cfa11e3cd6ff9dcd44d09675d52388d189c109cd3e5bd38828104be4be5318

Observation 3e55d466-c684-44d0-8684-7aa9cee9bae2 · outbound

This paper cites Effective Context Selection in LLM- Based Leaderboard Generation: An Empirical Study.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Effective Context Selection in LLM- Based Leaderboard Generation: An Empirical Study

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.475273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:49.945421Z digest=sha256:e10a39b25f84b30aaa2f7f911b2fc50edbed438c1d6889fe7febc8a84cbef3ed

Observation e2f9176e-86cf-4e67-948e-bbbf3a2d91a8 · outbound

This paper cites GPT versus Resident Physicians — A Benchmark Based on Official Board Scores.NEJM AI, 1(5):AIdbp2300192, April 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering GPT versus Resident Physicians — A Benchmark Based on Official Board Scores.NEJM AI, 1(5):AIdbp2300192, April 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.398558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.051191Z digest=sha256:5d6877892c991485bdf09a192b7101d5ac6fe13a908fa46c5d86b0a37e362396

Observation 1f555e5e-8ec6-48f4-a962-e52dba343d5a · outbound

This paper cites Baleen: robust multi-hop reasoning at scale via condensed retrieval.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Baleen: robust multi-hop reasoning at scale via condensed retrieval

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.326631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.133218Z digest=sha256:fb7b1c3fa0cb7e0080d4b706256ec5b37362e673b983006c02c61e98dbad2c3c

Observation f6434b9f-01aa-484f-a74c-930e59a325b9 · outbound

This paper cites Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.278636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.234128Z digest=sha256:9f96d3d19eba98d6d99e3486069d56c641136c3e30fdf6bc8d70a41396c8c9f0

Observation 1f67cda6-ea80-4875-82ee-58408f191d64 · outbound

This paper cites AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.213188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.304300Z digest=sha256:5d6f8aff773bdf26e97a7c51fa27ac1c0f28d51b89f1287892d6ace319b7437d

Observation 56fd67af-3220-4adc-af55-53276c8e0353 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.169801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.380832Z digest=sha256:9b6f31a597da80200a6373610cdba8fc3b8d7d98dc12ef6389ccaa155a4eac24

Observation 64a4ff52-f181-4bd4-9cf0-d4020f078ffa · outbound

This paper cites Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:58.127920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.460951Z digest=sha256:2bcf38db8a54f4613a9f9ed6f051fc93ff5e6c06189c3d1bfb8cb79c151770d4

Observation 850b2af6-8012-430e-9571-07616fc036ad · outbound

This paper cites Context Example Selection for LLM Generated Relevance Assessments.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Context Example Selection for LLM Generated Relevance Assessments

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.985087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.551402Z digest=sha256:ca1154e9c710528d8594ca91c2f2491b4f572839c788f8add5fce45770b05f06

Observation 2a0765b8-280c-4ef4-b802-90f2dee564db · outbound

This paper cites GPT-4o System Card.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering GPT-4o System Card

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:50.653743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:50.653743Z digest=sha256:26cf8a96f14cbf8c002b162546e86d8794b572addb130a3d730b3f137ef774d4

Observation 6203c3c4-7570-4e33-a864-ccb56c174205 · outbound

This paper cites MedMCQA: A Large- scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedMCQA: A Large- scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.903870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.733849Z digest=sha256:c80a5bfdc0f83ee99f9636d8e09fec23487b9346a57cd7821fd8e300ebe76de6

Observation ecd5867d-7b9f-47f0-9424-6e4c838f586b · outbound

This paper cites Bowman, and Shi Feng.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Bowman, and Shi Feng

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.665395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.789115Z digest=sha256:49a09097e6ecaf9af8ea04549e373154c1773cdb2dd70817fc1f6c1b047b2518

Observation 41b30a8b-73d9-463b-89dd-a2fee267dff3 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:57.510099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.838304Z digest=sha256:aaf2d2d2e57e42173f41c5109d357fd28c28865d62167501de56ea6d73eb2a76

Observation 572f7938-3816-440f-8d60-01db2dac20db · outbound

This paper cites Qwen2.5 Technical Report.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:50.889423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:50.889423Z digest=sha256:39d6dee4b65ecd9a73b06a1067896ab3a8f1503ff688855261f446c4620e1de1

Observation 2a2b63cc-2784-4256-b394-ab564a89868e · outbound

This paper cites Benchmarking Prompt Sensitivity in Large Language Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Benchmarking Prompt Sensitivity in Large Language Models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.268717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:50.989585Z digest=sha256:e4586dd5b5ed5bb53bc2c53ecb093df7df9761e9162b3d65bddf86be8839c441

Observation 9ad4cffd-c04b-41da-8b46-f9574415308d · outbound

This paper cites Towards Human-Centered Explainable AI: A Survey of User Studies for Model Explanations.IEEE Trans.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Towards Human-Centered Explainable AI: A Survey of User Studies for Model Explanations.IEEE Trans

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:57.096053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.056410Z digest=sha256:ddb9cf87fda21e95de8de3c4d954231adeca8ae9edafd0bd352a2809457e2aa1

Observation 03c0d856-2e1d-4243-ae7b-523cd8d01e8a · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:57.017659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.118620Z digest=sha256:c18468e0a72cf536c7ee59eea4144d9a1dacfb38067d04882c45063faf85f69d

Observation b6a919f8-20ca-478d-b579-b69def8345e1 · outbound

This paper cites Jung, Maria Zerlik, Waldemar Hahn, Martin Sedlmayr, and Brita Sedlmayr.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Jung, Maria Zerlik, Waldemar Hahn, Martin Sedlmayr, and Brita Sedlmayr

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.844972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.187616Z digest=sha256:1bbe8390fd971452d721f89a355be072890581dda5799b8897cee05e56495443

Observation 8e97d18c-f69c-4a7f-a6a0-6b9cd9d21e2a · outbound

This paper cites Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.272998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.272998Z digest=sha256:7e97b2ac5da078991bfd019cc35fd2382c26c3674f27751effca75507442cdad

Observation 618559c9-d48e-4adc-84e1-0c84343e0e0d · outbound

This paper cites Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, April 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, April 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.346660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.346660Z digest=sha256:fe6d4a184005f95b0b468f233c3d761cefc78ce30fd9d1b25e5df39e3f4f7ba6

Observation 5f00423a-8a68-4d61-805b-db539a1b02dc · outbound

This paper cites Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.543815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.410573Z digest=sha256:32584fc405e06b279d741909b7e33fe9b1700bc20d6b8f50231102123ed93721

Observation 00ce585b-331c-4895-a2de-ba017a47fc50 · outbound

This paper cites Don’t Use LLMs to Make Relevance Judgments.Information Retrieval Research, 1(1):29–46, March 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Don’t Use LLMs to Make Relevance Judgments.Information Retrieval Research, 1(1):29–46, March 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.273646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.483199Z digest=sha256:e53b2fd3aa56f23eb61f8cf2d24c9f09e3057a01935fe16fbf98ee94fd4727f3

Observation 332a4770-99a4-4ef3-a481-e6fccb35aeae · outbound

This paper cites RadQA: A Question Answer- ing Dataset to Improve Comprehension of Radiology Reports.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering RadQA: A Question Answer- ing Dataset to Improve Comprehension of Radiology Reports

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:56.005774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.499100Z digest=sha256:e52b917778a210b0feaf3e6fd12c57cfd4e14c8ddcfac84cc2d0ef0a43d9e3c5

Observation 131b2e92-d6ca-4370-ae6e-8b32163ac763 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:55.813354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.519320Z digest=sha256:4987314533f132f03bbb19fd358cc8016e1c6a1ba65c21ba94bea7b5f925c319

Observation 5d75952b-8710-4d0a-be25-59fb2c34cfd8 · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.574225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.574225Z digest=sha256:df755c3f662308f9eb2e4429248c27becd1fe759522c13354382c6e571dcc272

Observation 4718ff6a-cf92-411f-8310-476c7ff3d371 · outbound

This paper cites Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.610609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.692771Z digest=sha256:255359e241292598f2b1f49705e539eee0bfb1b9c941665a3ac0d4a06bc6df28

Observation 9773ce0e-f307-480e-abf7-c1e95ab90dc7 · outbound

This paper cites MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:51.752953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:51.752953Z digest=sha256:0be63b3dde6e9cd2dacae3bcc844c2ae72ffaa0789d0de5625ea276d2cfd8d2d

Observation 012df2de-7e4e-421c-8f92-ba936bbb4b9f · outbound

This paper cites An automated framework for assessing how well LLMs cite relevant medical references.Nature Communications, 16(1):3615, April 2025.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering An automated framework for assessing how well LLMs cite relevant medical references.Nature Communications, 16(1):3615, April 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.489658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.855274Z digest=sha256:282c16e4d1052581bc1a5e4fa6f15e5b973c65825f4f16250515a498d3dae4a6

Observation 1758cd8c-4c30-4e3c-9b32-823590261821 · outbound

This paper cites CARES: A Comprehensive Benchmark of Trust- worthiness in Medical Vision Language Models.Advances in Neural Information Processing Systems, 37:140334–140365, December 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering CARES: A Comprehensive Benchmark of Trust- worthiness in Medical Vision Language Models.Advances in Neural Information Processing Systems, 37:140334–140365, December 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.353717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:51.968438Z digest=sha256:d9dfd6a4a2bce76f56465d30df388c83d565c05ed5b47fe0a90fe354d9c821c1

Observation fa70b999-29d9-4a59-a506-b046a9b0b448 · outbound

This paper cites Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.221327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.018572Z digest=sha256:58d7496c73d03ed6e7748e18ed871514a9d8c4a839f39c86089137434b150605

Observation edca9806-32b8-4907-b8a0-62b4c8b149a1 · outbound

This paper cites A survey of datasets in medicine for large language models.Intelligence & Robotics, 4(4):457–478, December 2024.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering A survey of datasets in medicine for large language models.Intelligence & Robotics, 4(4):457–478, December 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:55.092113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.106367Z digest=sha256:714013bd3ec9549c552a4866bf684196afb829d9d3b97dda536b39e58290e128

Observation ed446d7e-86b8-46f0-bad2-1b772a81272b · outbound

This paper cites Meyer, and Steffen Eger.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Meyer, and Steffen Eger

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:54.974189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.206653Z digest=sha256:0f56dde3822ce88901c1ca836547a7e64bbd976116c75cacb6fd76d5eb57184e

Observation 30d7ea68-3a86-48f9-8f68-d268b451553a · outbound

This paper cites Gonzalez, and Ion Stoica.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Gonzalez, and Ion Stoica

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:52.346225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:52.346225Z digest=sha256:a48889e7c4e64cf7cdcdd570cf8805bf6f6072a79421c21c0f5830cbe0efd6fe

Observation 1ee80f7e-a2d4-4c6b-a799-220692457244 · outbound

This paper cites Melton, James Zou, and Rui Zhang.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Melton, James Zou, and Rui Zhang

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:54.796467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.455186Z digest=sha256:08ad17ffaf08a8ab99a30616f81d2052031e3ff654c78a43db003d4ab0677475

Observation f548da07-6d9c-438d-bc51-f2528ac72cc9 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:52.565185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:52.565185Z digest=sha256:e422b0947d66acc2fdb6e75e781fd2de9122128b70d87af1b1d5212f9fed3efd

Observation 3bf35d46-e107-48e7-bb1b-8f8c404da2b8 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:54.605443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.655992Z digest=sha256:37d2e896255c96f535f4e92478fd9edc4bbb0a5f5d09ef2aeafaefcac7445376

Observation e67608e9-17db-4bdc-8e0d-51711bed76e4 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:54.420889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.759621Z digest=sha256:8fc67723fc145c1a9def25a2572ef1273ee7a040f4474f1d21860621fca4aaaf

Observation 753df1bd-8613-42a7-8c2a-9c5e9cbedbed · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:54.180946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.868840Z digest=sha256:9c635b16e9708de94e62441422f93525d6cf7af3e0f60cdce61e1127769d90cb

Observation 76b99462-dda6-4b93-ac42-1f74f8952046 · outbound

This paper cites Her new job involves walking several miles daily across a large facility.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Her new job involves walking several miles daily across a large facility

Reference 76

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:42:54.061125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:52.971780Z digest=sha256:9037c231ad98e06965ebd86603b042aba1a7fc7ead5b3a97d001ead6f11fc06f

Observation 12000119-02f6-4bc8-b0e2-c5e33039fb90 · outbound

This paper cites Towards Digital Sustainability in Health Care: Developing Digital Health Products through Data-Driven User Insights.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Towards Digital Sustainability in Health Care: Developing Digital Health Products through Data-Driven User Insights

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:53.800495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:42:53.123407Z digest=sha256:e833a888cc7ce632d5f051d84143a7d0d08a43d6aefe39f544584b7c3d5be809

Observation d72d57a3-c8b8-4355-b7c5-2d08ac677448 · outbound

This paper cites Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:47.761883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:47.761883Z digest=sha256:2edfa155f5be1167948aa29e9d023fe72b5f6164a8f59f0a1a7569aec2b299df

Observation e2972f1e-6fff-451e-b275-05a5a6bfbc06 · outbound

This paper cites an unresolved cited work.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:48.690834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:48.690834Z digest=sha256:25fa3191f71a1b367c645048c3eab7ed34900340ea828fd89de420bd15aaaec1

Pith citing papers

Observation c2b5d90d-a592-4cfb-a612-656ac73f0cb7 · inbound

Treatment, evidence, imitation, and chat cites this paper.

Treatment, evidence, imitation, and chat MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.892747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T08:21:05.698812Z digest=sha256:c499eb6f1becc25b20913a3b515af145ea94ac68d95fec5c23d4077760faf643

Observation 754b740e-9581-46b2-9fe4-fbb513d5b8b2 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.604352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:e91736a178f4a68f87151e5e4a5d0472b9db21db52ac9ef6336b8cf11a57bac1

Observation 3911e0e9-f2f5-4471-a71b-d8a61893a70b · inbound

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining cites this paper.

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:34.996501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:02:30.696295Z digest=sha256:0094b0625a1d4dd76259f78804607301481e6482f26bf67ff01e011ed86ee9c5