Pith. sign in

Paper Citation Record · LEDGER

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

As of 7 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2606.12291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.12291 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T09:30:47.726480Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact22
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c261a26-5299-4181-9e32-924ec78a814a · outbound

This paper cites The evaluation illusion of large language models in medicine.npj Digital Medicine, 8(1):600, 2025.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context The evaluation illusion of large language models in medicine.npj Digital Medicine, 8(1):600, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:f0975dd64fe34e4dc8361fceb8474bcf79616770f773406eb30760c27a808289

Observation b80a1a4e-76e0-4561-956d-3f49a38a89a9 · outbound

This paper cites Introducing Claude Sonnet 4.6, February 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Introducing Claude Sonnet 4.6, February 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:f3abf8102f65c5c97a0ee195116d799312519bbd57aa42d86b53cf9328789aad

Observation f5a1742e-396c-4ff5-8470-868ed1648851 · outbound

This paper cites Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum.JAMA internal medicine, 183(6):589–596, 2023.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum.JAMA internal medicine, 183(6):589–596, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:b2d3050920cdcad0133940a25124a9c9d353b0f62ee41e9ae7ba4fb7e424e0d1

Observation 67e647d4-03be-4e73-b46b-8dba7ca2f750 · outbound

This paper cites Fries, Michael Wornow, Akshay Swami- nathan, Lisa Soleymani Lehmann, Hyo Jung Hong, Mehr Kashyap, Akash R.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Fries, Michael Wornow, Akshay Swami- nathan, Lisa Soleymani Lehmann, Hyo Jung Hong, Mehr Kashyap, Akash R

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.713725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:6de294d23e1a4c5058c05f0056481b3b77f1f71d582da4145ddde5f500dab2ab

Observation be36c7f0-cd1a-4761-99c8-606b73022289 · outbound

This paper cites and Saba, Luca and Hadamitzky, Martin and Kather, Jakob Nikolas and Truhn, Daniel and Cuocolo, Renato and Adams, Lisa C.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context and Saba, Luca and Hadamitzky, Martin and Kather, Jakob Nikolas and Truhn, Daniel and Cuocolo, Renato and Adams, Lisa C

Reference 5

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.715756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:3c85c4daa6459e57517f6dcbc62a3d9878387cc6b2e9c42d1b533a243a3fead1

Observation 7e605dd6-e4a8-4331-bce0-cb9334308c6b · outbound

This paper cites MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.726832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:b778e3042edf369509bf7eddab62d7d98a887811332e3bcf77f8eaa7e6361307

Observation d577bdeb-c2e2-452a-bf3e-cec76a00e335 · outbound

This paper cites Singer, Xuguang Ai, Po-Ting Lai, Zhizheng Wang, et al.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Singer, Xuguang Ai, Po-Ting Lai, Zhizheng Wang, et al

Reference 7

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.729745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:08886a81df25b584c9a6a95823f2855079a231e72a1ddf5d2754358e9ae809b3

Observation 1e241908-2353-4149-87e4-a074900d5322 · outbound

This paper cites M ed R isk E val: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context M ed R isk E val: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings

Reference 8

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.703992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:e0e8859601a621c727b667ac18319fc2b82a830a755bf6a64f02b4734c2bb3a7

Observation 719ec87d-f4b3-4372-be84-e1568a1de664 · outbound

This paper cites Nour, Seth Spielman, Samuel F.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Nour, Seth Spielman, Samuel F

Reference 9

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.717626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:232138c807957aa2e056abcd3de2c7fe597a80ddab03d87fb27dc2128d799dfc

Observation 5351bece-dab5-4883-a7cc-6ef534825137 · outbound

This paper cites Openseeker: Democratizing frontier search agents by fully open-sourcing training data.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Openseeker: Democratizing frontier search agents by fully open-sourcing training data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.720548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:4141d7b503a643326fbe15171f9a236a047c5385feccffb0d8a50fb2ae9bbfc8

Observation 540d2000-a428-44b9-ab3d-16b6a5dfff03 · outbound

This paper cites Gemini 3.1 Flash-Lite model card, March 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Gemini 3.1 Flash-Lite model card, March 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:2bc9347b82cf515aee038e5857fec0a0a03d0cec08507db6e6cd7e8986462d99

Observation c5699b5a-6319-465d-be68-7fab4520da9c · outbound

This paper cites Gemini 3.1 Pro model card, February 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Gemini 3.1 Pro model card, February 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:c66532a5c88ebccde55ffa0f7c4fb5f10c9d6fab0ab03a68fa572052359c8159

Observation 27ac4c67-ad69-4624-9585-ed5d35d10b72 · outbound

This paper cites Gemma 4 model card, April 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Gemma 4 model card, April 2026

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:09f047cb5da4d8eb41f8c33dfc8168c308f76685c13e7d7de574841618e23ac3

Observation 7a4ed9e4-0b01-4fca-b6dc-362ca410d226 · outbound

This paper cites InProceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec @ CCS 2023).

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context InProceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec @ CCS 2023)

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:40:47.665521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:ae44ffcbaf2d53cd45ff26af78e35b59069dc28620a5cdfa18a74a8a05ef081c

Observation 32c75ced-a565-4b9f-9ea8-459d041a706b · outbound

This paper cites Bressem, Jakob Niko- las Kather, and Daniel Truhn.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Bressem, Jakob Niko- las Kather, and Daniel Truhn

Reference 15

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.671155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:6c649da9accfd7a6953511e8c00a13bb3257ef18cdf982d9e203020bd0b39ed2

Observation 1c4768a3-0fc8-408c-9191-6c9e5ae83fe9 · outbound

This paper cites MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.682854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:06b6b9e76a2e0f3a4fea0cabc57b0b29bb727532bf62ead18a32d5811057941a

Observation 4a51a8e8-4057-411b-a30a-9ff8d7640f71 · outbound

This paper cites What disease does this patient have? A large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context What disease does this patient have? A large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421

Reference 17

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.679630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:2ff2e88d2accf42a8f465293a7f4e3ea9b24b8e746f907ae9aa6edb1c89ad025

Observation 58b18d2e-4f74-41c8-aaf5-84660a3342ab · outbound

This paper cites H., et al.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context H., et al

Reference 18

Resolution
metadata mismatch
doi, observed 2026-06-27T09:40:47.657248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:1a098b5c88e556d8f2f8d5d76327a95f2257727e1a6bca58325a699696fb78b8

Observation 89a3a225-a350-4a6e-a4bc-33a2f0d0ab4e · outbound

This paper cites Evaluating clinical competencies of large language models with a general practice benchmark.Nature Communications, 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Evaluating clinical competencies of large language models with a general practice benchmark.Nature Communications, 2026

Reference 19

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.668222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:12575abf1f2bdc4690dc53b53adc0af0e701326192a921dc59905c70bcd70620

Observation 1a051535-a6c0-4d00-84d3-dc346cc526ad · outbound

This paper cites an unresolved cited work.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.659138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:f1b5b3767e6182e05f19ed4ba78fba694bd6ec4769c72383ec1f54efd459c6d1

Observation 9e798e51-8f1c-4fb7-9a37-9cf94bc46b4f · outbound

This paper cites Chen, Yining Hua, Peilin Zhou, Junling Liu, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, and David A.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Chen, Yining Hua, Peilin Zhou, Junling Liu, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, and David A

Reference 21

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.661074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:c2d8568999de2156f43fb2b5738424b3729974ce8fa827dcb66797879651103e

Observation 3b68e594-aaad-45d0-851a-2c5d26fc447e · outbound

This paper cites Benchmarking large language models on CMExam - A comprehensive chinese medical exam dataset.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Benchmarking large language models on CMExam - A comprehensive chinese medical exam dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:569175c4363f0310ac24c0478317edd8463cd096bde3c6c88f0d24c43ef900d3

Observation 38bcec30-d012-4b2e-a991-839b04c20f97 · outbound

This paper cites Carrero, Xiaofeng Jiang, Dyke Ferber, Georg Wölflein, Li Zhang, Sanddhya Jayabalan, Tim Lenz, Zhouguang Hui, et al.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Carrero, Xiaofeng Jiang, Dyke Ferber, Georg Wölflein, Li Zhang, Sanddhya Jayabalan, Tim Lenz, Zhouguang Hui, et al

Reference 23

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.653523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:a29da6e25c074e34d7724b7afcb14b58496f010a9108ae3391e6849a70a54421

Observation b64375e8-78c7-4159-b8ff-7db14240cf13 · outbound

This paper cites Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:40:47.697055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:599883dca03584e883ab8998772a8b6b9b5c24a9a7cec50ba193e0e49180b893

Observation 66d22251-a305-46e8-86ef-e56236d08886 · outbound

This paper cites Persuasion with Large Language Models: A Survey of Empirical Evidence, Study Methodologies, and Ethical Implications.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Persuasion with Large Language Models: A Survey of Empirical Evidence, Study Methodologies, and Ethical Implications

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:40:47.674649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:fc6531a284e2332ed84d351a843e14f3d7e314df258064c1b885d6b1ded6e840

Observation 1d8fef7f-8bf1-4cd9-8e4e-b071d2fb5167 · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Capabilities of GPT-4 on Medical Challenge Problems

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:40:47.694621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:682bd64fee4c20b42bce7a2db9b83d3f8e94807364ec904518192d36e9fec756

Observation d88b5189-5ad4-4e53-9185-0cd2cf898024 · outbound

This paper cites Wieler, Alexander W.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Wieler, Alexander W

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.688112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:4eba01f58867a7612d673d1c461f44173d8b0cb0b2d6a1a6cd6f3152f3b049f5

Observation 5148f2fd-5fa3-4acf-8a4f-09af51f4093f · outbound

This paper cites Introducing HealthBench, May 2025.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Introducing HealthBench, May 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:025d99d08757c235642e10f33352e36f93889181ae94cbbf5905d0b73e680040

Observation a7a54846-05e7-43af-ab70-764dbaacd8e6 · outbound

This paper cites Introducing ChatGPT health, January 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Introducing ChatGPT health, January 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:e412047727144c4d0b96211b6f1ff138485de207785600528082e21413c89115

Observation 941e6f5f-006e-42ee-abf5-c930b369fb1a · outbound

This paper cites Introducing GPT-5.4, March 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Introducing GPT-5.4, March 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:ef7c6f9dd9df53b41ef5868ab89355d4e09fd68a5b2a173cee4232736f5e013d

Observation 49d70b12-958a-434a-ab60-021c82fb2059 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:8db9e207ee6815d19e04387417e45d309ea095b4de1e1c4f8093f30709f4c110

Observation c8402bc5-06b9-45e1-b2f0-a77ca98e549d · outbound

This paper cites Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

Reference 32

Resolution
metadata mismatch
doi, observed 2026-06-27T09:40:47.692104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:f672ca773a98a0536313f38dc3657ae5d36dc3378ea58fc280d075d2ddcdf7c3

Observation 0ab6ef48-f502-416c-9af1-027bd2548c5f · outbound

This paper cites A benchmark of expert-level academic questions to assess ai capabilities.Nature, 649(8099):1139–1146, 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context A benchmark of expert-level academic questions to assess ai capabilities.Nature, 649(8099):1139–1146, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:90b7579fb58fc7a522d5d2dde96d67ed062ef264d525a048cd2f5961a5b7e42e

Observation 544ca31a-b033-459b-9e6a-345ab9c37178 · outbound

This paper cites Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:67d7cd03cb8954f7e76502439c20e16f94df5005ce62f9c56f8b948f439c2d43

Observation d825c699-3e00-43b2-b048-d8e047d718f7 · outbound

This paper cites Rao, Kaiz P.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Rao, Kaiz P

Reference 35

Resolution
malformed identifier
doi_truncated, observed 2026-06-27T09:40:47.689981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:9f684a415bea01fa969358a5d396be8b3de29f397380e283d468428ca529a307

Observation b4f2261e-e04c-4420-869d-b162aede1ea0 · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:40:47.651478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:3558dd0168c6e3bc672e837f5a8a56bdb5fff3cb9808fc548f2d505291a8e4bb

Observation 75acb012-7971-4e53-ac1e-1ca640badf4b · outbound

This paper cites MedGemma Technical Report.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedGemma Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:28:04.906945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:50d751f48957d4174362fa096c5732adafc2e5fe7573a99b6cd565b2e3c502b0

Observation 88941836-0bfd-4bf4-97ac-daacae0e5e87 · outbound

This paper cites MedGemma Technical Report.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedGemma Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:40:47.701506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:fae3bdcbc0299c79828f9d03ec781b44129b678577fe9e1481a181c25d864a01

Observation 6bad8ac1-4d39-4408-94d6-2d1db9c1f87d · outbound

This paper cites Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:34f0fe06ccafcaf57bb666e672afdcaaaf05691f75ab80e0714790a75fbf6e5c

Observation e86c6f93-bb8b-420f-adee-b5903f9810c6 · outbound

This paper cites Large Language Models Encode Clinical Knowledge.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Large Language Models Encode Clinical Knowledge

Reference 40

Resolution
metadata mismatch
doi, observed 2026-06-27T09:40:47.685587Z

Source-reported events for the cited work

correction dated 2023-07-27. Source: crossref record 10.1038/s41586-023-06455-0->10.1038/s41586-023-06291-2:correction, observed 2026-07-11T03:08:19.417011+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:38403244725897aba0b643cab056672b39a7a45c4ef52dda6ce5e6ccc91e5bde

Observation 1f9089c2-00cb-4e96-81e9-ee9bd8b32776 · outbound

This paper cites Toward expert-level medical question answering with large language models.Nature medicine, 31(3):943–950, 2025.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Toward expert-level medical question answering with large language models.Nature medicine, 31(3):943–950, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:a2594c3ecc70f5c7ab9dd5fe2ff262d4919e49d80b78bc1cc36b0015f9ac7e4c

Observation 06564744-66fb-43d8-9e4a-068065fa7b05 · outbound

This paper cites J., Ting, D.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context J., Ting, D

Reference 42

Resolution
metadata mismatch
doi, observed 2026-06-27T09:40:47.677277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:c817b2757a385ba106e161e09b310dce857d2eefd340f5c0fb2fb9d5862803f0

Observation 2b451404-2d67-4e8f-8c24-2c9344267af2 · outbound

This paper cites Scientists invented a fake disease.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Scientists invented a fake disease

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:924a808853813aa358fb1306f1ce154df73e1df65aad8d701b7b9fa30bceb2f8

Observation 56da7f00-9808-4575-ae31-8c6dedc8d968 · outbound

This paper cites Sara Mahdavi, Christoph er Semturs, Juraj Gottweis, Joelle Barral, Katherine Chou, Greg S.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Sara Mahdavi, Christoph er Semturs, Juraj Gottweis, Joelle Barral, Katherine Chou, Greg S

Reference 44

Resolution
verified exact
doi, observed 2026-06-27T09:40:47.665957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:f51daa734e3e5052d4dec7c928217254eefc0a94e174e58947acf8eee6fc2b6b

Observation 246cfae3-a8c3-4950-b0f6-6006f339bc29 · outbound

This paper cites A novel evaluation benchmark for medical llms illuminating safety and effectiveness in clinical domains.npj Digital Medicine, 2025.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context A novel evaluation benchmark for medical llms illuminating safety and effectiveness in clinical domains.npj Digital Medicine, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:fcdba36f38fbf53c977f128a3a5be5c51b7562c4ac6ecce8ff45586e15b2a1c3

Observation 6c145b93-6c3a-40ee-81fb-08fbd133268c · outbound

This paper cites Understanding the infodemic and misinformation in the fight against COVID-19, 2026.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Understanding the infodemic and misinformation in the fight against COVID-19, 2026

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:52f9489d483408584ba895c07511fb98a4ee993c6d90922dd368fa855cf6304b

Observation b3f29295-7fea-479c-ab2e-e5a41d1454a2 · outbound

This paper cites Medjourney: Benchmark and evaluation of large language models over patient clinical journey.Advances in Neural Information Processing Systems, 37:87621–87646, 2024.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Medjourney: Benchmark and evaluation of large language models over patient clinical journey.Advances in Neural Information Processing Systems, 37:87621–87646, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:f21fc68d54968ab86a7a578c9e7373e9ab11337cca607196957956edc1d7cb01

Observation 02ea82fb-5be2-413e-813c-a31f49c1b078 · outbound

This paper cites npj Health Systems 2025 2:1 2:2-.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context npj Health Systems 2025 2:1 2:2-

Reference 48

Resolution
metadata mismatch
doi, observed 2026-06-27T09:40:47.655715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:b80cf410da913cf630df5afa210f0d2b9aefd4ea913d9408a2778373bb24a6f1

Observation 375f9888-66cb-4d10-8354-53ba2a5d62de · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context React: Synergizing reasoning and acting in language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:dd612d963285be3247287a0bc614bddef615bb96797bc657fcf4076d402b4b14

Observation 1476ca49-19fc-46c6-930f-65b89229fa44 · outbound

This paper cites PoisonedRAG: Knowledge corruption attacks to Retrieval-Augmented generation of large language models.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context PoisonedRAG: Knowledge corruption attacks to Retrieval-Augmented generation of large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T09:30:47.726480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:c42d57f3613ee163df4ab882d6b8c419cdf3fb117af246198e7827d7925167ba

Observation 85ae5e85-f44d-490e-8969-4645295422f7 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 51

Resolution
malformed identifier
local_arxiv, observed 2026-07-03T11:28:04.910150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:88a5e5f70a83ac75d2e4073e4460a40ef81a66a10e27905aa03283f30b0c2710

Pith citing papers

No inbound Pith citation observations are available.