Pith. sign in

Paper Citation Record · LEDGER

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.28677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28677 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:47:08.786613Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de96c79e-4a40-40a5-8179-5e64eac28911 · outbound

This paper cites Performance of a large language model on the reasoning tasks of a physician.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Performance of a large language model on the reasoning tasks of a physician

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:05.752339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:05.752339Z digest=sha256:8fd5722a6f190a5a9aff4cac473132506a75ffa3ec5b2870fb14d19de6efa245

Observation b74099ce-fab3-4452-b685-1b63fff2d913 · outbound

This paper cites Reliability of LLMs as medical assistants for the general public: a randomized preregistered study.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:05.840374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:05.840374Z digest=sha256:b7e66f38a9e70d9c4abe5290cfcc9f48f803b4fc0996fcc0fe0ce5c158fe9965

Observation 7b945728-8bc2-4755-9509-9281d59801e5 · outbound

This paper cites Measuring what Matters: Construct Validity in Large Language Model Benchmarks2025 November 3, 2025.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Measuring what Matters: Construct Validity in Large Language Model Benchmarks2025 November 3, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:05.945108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:05.945108Z digest=sha256:89c5b6c42af78df4d5b0caad97bf1a25a85a5ce3431c63c65014b0f4d5fff014

Observation fce241fd-572f-4cc1-9c1d-660504b4d9fd · outbound

This paper cites On the robustness of medical term representations in locally deployable language models.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support On the robustness of medical term representations in locally deployable language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.111582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.111582Z digest=sha256:786d2396d1d017431c8a6fcc25fa58e49e8b2b243eb3aae79ba6309af39295a4

Observation f2784f09-c73c-48ff-b9bf-53c9ba25b0d8 · outbound

This paper cites Health AI needs meaningful human involvement: lessons from war.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Health AI needs meaningful human involvement: lessons from war

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.220549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.220549Z digest=sha256:22eff07d3e9d22cda14f5a0bad0c55108a9379859ad4192aeff5acc88b7a8c1e

Observation ecee4c67-07d8-45fd-bd8a-8b4fe65eb86f · outbound

This paper cites From Concept to Clinic: Real World Evidence for Autonomous AI Deployment in Primary Care Telemedicine.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support From Concept to Clinic: Real World Evidence for Autonomous AI Deployment in Primary Care Telemedicine

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.389839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.389839Z digest=sha256:c74881db090274c1d143088d2473ce9024b2332d3983d5b5513a5191ff74d21a

Observation d5498374-9fa5-4b33-a0d7-26f93bedc42d · outbound

This paper cites Towards autonomous medical artificial intelligence agents.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Towards autonomous medical artificial intelligence agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.486690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.486690Z digest=sha256:6ab3d653213659941b8ed387b9ee83814a0bc517962486271098509e1aa691c1

Observation bfb0457c-5416-4141-822b-4a2cc3be5c05 · outbound

This paper cites A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.542200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.542200Z digest=sha256:4a6edc4e59641a5428d7384ce65fac28a366785a0802cfa0afa57726689f6f15

Observation a2dccb91-62a7-4e8d-932a-6e10c5fb2df4 · outbound

This paper cites Safety of a large language model-based clinical decision support system in African primary healthcare.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Safety of a large language model-based clinical decision support system in African primary healthcare

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.666268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.666268Z digest=sha256:829ecda1bff9eab8dfd73cca72d7d22886ed78bcf9ccaca591480249979548fd

Observation 20518bf9-356d-4738-a5ad-8f65ee5cae15 · outbound

This paper cites AI-based Clinical Decision Support for Primary Care: A Real-World Study.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support AI-based Clinical Decision Support for Primary Care: A Real-World Study

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.754803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.754803Z digest=sha256:afa81779ac04f5ca1f34861eb2b3ae5e4f7b4150c319613935f4a33645dbd094

Observation 035c0676-8880-4aaf-bec3-b24b63b55bd4 · outbound

This paper cites Medical errors in large language models revealed using 1,000 synthetic clinical transcripts.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Medical errors in large language models revealed using 1,000 synthetic clinical transcripts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.851847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.851847Z digest=sha256:9ae07045275a74817727d30b577b88c3f2c7380694417ef384d2c31dba3b66d7

Observation c00148ae-0141-4595-8dba-67898da83f87 · outbound

This paper cites Large Language Models lack essential metacognition for reliable medical reasoning.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Large Language Models lack essential metacognition for reliable medical reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.907210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.907210Z digest=sha256:e9cc29ee1917083b9c512d3f7427a76a3cfcd596bc5d5e1a3cef8f5bd003c73c

Observation d85b0cde-3c7c-4c81-b0ec-817a1e640eee · outbound

This paper cites Towards conversational diagnostic artificial intelligence.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Towards conversational diagnostic artificial intelligence

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.069220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.069220Z digest=sha256:106e50424712a81d4a8f56a5801850ec8d2a0da1eb19a12d5e01fb5a035d4726

Observation 8dbb94b9-03f0-4e4f-a837-c683fd606076 · outbound

This paper cites Towards Conversational AI for Disease Management.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Towards Conversational AI for Disease Management

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.234978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.234978Z digest=sha256:ed5a69ba373f8bde65e5e004abfd65da2ddadeb3eca6fd24d0134774318349fd

Observation 0b016d29-21a5-4973-a476-dc71e2ed4388 · outbound

This paper cites Testing and Evalua- tion of Health Care Applications of Large Language Models: A Systematic Review.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Testing and Evalua- tion of Health Care Applications of Large Language Models: A Systematic Review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.401396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.401396Z digest=sha256:b77c1f30ed7e50bbb0a77849dd5f74217f6963f3da2394eaff02beb787bd146a

Observation b74fbffc-c8cd-4bce-a7db-44304f367166 · outbound

This paper cites LLM-assisted systematic review of large language models in clinical medicine.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support LLM-assisted systematic review of large language models in clinical medicine

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.513681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.513681Z digest=sha256:91a94ca6a212e78309caf4da50e022802c0c3f9a008a7a4d9366cf5b9e8faeb9

Observation 0a5bcce6-3244-4658-b3c3-83e1f25ad143 · outbound

This paper cites Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.674935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.674935Z digest=sha256:cb9f9a8adbc7344d6cf1b5218469938eb123233afafb8b817096101d71eb43fe

Observation 471de8cb-5145-4435-8013-9643d29439e8 · outbound

This paper cites MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.840702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.840702Z digest=sha256:1b55524e46fe4cf4bdb3f5e3830aa84243c69360fcaee64af5d57abee722491a

Observation b5515b52-0bcc-4319-bd09-816bbdb412ef · outbound

This paper cites Training language models to be warm can reduce accuracy and increase sycophancy.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Training language models to be warm can reduce accuracy and increase sycophancy

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.900627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.900627Z digest=sha256:82737f24321c8c287707351b15ba7bb135735c9a5819516c6c56266ba2b05315

Observation a4b24a9d-51dc-4923-96b4-ea1e1d9911cb · outbound

This paper cites Competing Biases underlie Overconfidence and Underconfidence in LLMs.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Competing Biases underlie Overconfidence and Underconfidence in LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.008719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.008719Z digest=sha256:1a8b751095bb0cc26b7f72ff28bfc6d2518673e3c517dc224e78bc3bc115031a

Observation 4a28b52e-1033-4391-bfe9-df546d31a1dc · outbound

This paper cites Assessment of Large Language Models in Clinical Reasoning: A Novel Benchmarking Study.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Assessment of Large Language Models in Clinical Reasoning: A Novel Benchmarking Study

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.039934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.039934Z digest=sha256:74ad24fb39afc2581549864851baadc88ba0fdf633b36d8fa138c357824ea09a

Observation 511ac63f-afcc-442a-a47c-4b86c6f4099d · outbound

This paper cites First, do NOHARM: towards clinically safe large language models.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support First, do NOHARM: towards clinically safe large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.117061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.117061Z digest=sha256:2a48fb77e96c9524ee83682f883192e681dce32e1ef9ae1bbf5ccb8634a6e5b4

Observation 14f78564-fe69-4c59-bb6f-5d2f97640ba6 · outbound

This paper cites MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.177007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.177007Z digest=sha256:8f6ff90a68c1325c603d2070780d729ff6ac953c2e527fa3e079af0516a6e117

Observation 827f14c3-925e-4fe8-af3f-f3090208de68 · outbound

This paper cites SymptomAI: Toward a Con- versational AI Agent for Everyday Symptom Assessment.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support SymptomAI: Toward a Con- versational AI Agent for Everyday Symptom Assessment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.285001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.285001Z digest=sha256:696310228addfa7419d392fa52f3fdb6c12caed921bea67f8a1bcbadf72ada10

Observation 91c4f6c5-71d2-4782-9134-7b3028b411a7 · outbound

This paper cites A POMDP Formulation of Preference Elicitation Problems.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support A POMDP Formulation of Preference Elicitation Problems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.396535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.396535Z digest=sha256:9620ccc69f82551f755bfbbfbd3af29cca3a535efc8b58877191491eb991f299

Observation c7d11616-466c-4f77-8618-b75311bf09eb · outbound

This paper cites A clinical environment simulator for dynamic AI evaluation.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support A clinical environment simulator for dynamic AI evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.504867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.504867Z digest=sha256:a1a57980a895b5faba52bae8b93f63ad2e750885a0e29fb8a3ad170f5f250499

Observation 993d9874-d12c-4727-b5dd-74d99e6f52a2 · outbound

This paper cites LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.564508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.564508Z digest=sha256:e88d2ddf5d63a42109f86f53c42d6aaea031f406772cb18aa01924295ac0d909

Observation d50a2074-7081-42a9-9b88-1b298f12bbc8 · outbound

This paper cites AI, Health, and Health Care Today and Tomorrow: The JAMA Summit Report on Artificial Intelligence.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support AI, Health, and Health Care Today and Tomorrow: The JAMA Summit Report on Artificial Intelligence

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.624687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.624687Z digest=sha256:fa1b715d7e6876fb0906241f7bf222ad889baa288e2f9202f12f427415fa7e42

Observation abf90368-e3e8-47f5-b310-f25a8c54446e · outbound

This paper cites Nature Medicine.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Nature Medicine

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.786613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.786613Z digest=sha256:78fce0628d0275053e04a8a938eb525975d11b1b8dd5480536f596ee8e51633a

Pith citing papers

No inbound Pith citation observations are available.