Pith. sign in

Paper Citation Record · LEDGER

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models

As of 12 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 0 inbound Pith citation observations for arXiv:2412.15748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15748 v2

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:11:20.582856Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 111 outbound references displayed

  • verified exact5
  • verified fuzzy21
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba171195-4212-4408-8d2a-dfd967a22076 · outbound

This paper cites Attention is all you need.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Attention is all you need

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.081483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.081483Z digest=sha256:0c99b95c109d294219d2934dc064abbdaf71515d00aa7ea25e39a13c487096d0

Observation 05ad7186-4003-4460-b4e9-ce0114ef6c26 · outbound

This paper cites DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.086727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.086727Z digest=sha256:fa1e04ae812b33578dfd16f62f8aca55f0419f66535949dbc956d717740eebd5

Observation 955fee99-12ec-4d45-b643-4386a5b3f9f8 · outbound

This paper cites Sequential Diagnosis with Language Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Sequential Diagnosis with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.092099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.092099Z digest=sha256:4325880a9ed914791549bd2ea404aaeeb46cce9762c7dce21da3b38fd32849b4

Observation be829d62-6e26-4adf-a602-27cd40dcf905 · outbound

This paper cites Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine.NPJ Digital Medicine, 7(1):20, 2024.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine.NPJ Digital Medicine, 7(1):20, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.097479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.097479Z digest=sha256:eee756fbca8c23cc8846892523577ec1dfdd850e410aff5b591863170ec88d34

Observation feea806c-f572-4e1f-a657-6bc674dc31a7 · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.102484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.102484Z digest=sha256:31cbaf768e38eb3760c5fb279311eac02d4e51b5e760c32bb32163be630797e0

Observation c2985272-7e53-4da9-96a1-aecac7f62f7d · outbound

This paper cites Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.107764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.107764Z digest=sha256:bdadf63d27c3139dad33cc6b82a66916fc023a6f10ba0e187546315d1d5c49fd

Observation 0db61431-7f41-418b-903d-b31ad7565bb9 · outbound

This paper cites Enhancing medical summarization with parameter efficient fine tuning on local cpus.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Enhancing medical summarization with parameter efficient fine tuning on local cpus

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.112970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.112970Z digest=sha256:fcd3990ea950e22a7bdf704777aa69b270b74f6131ed1dec0f8a50744afd9c97

Observation dafeb8fc-71f1-40c9-8986-109187c47e06 · outbound

This paper cites EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health Records.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health Records

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.117728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.117728Z digest=sha256:3d2681e9931f8a36eee8732b29a47569c0f5f5ab6bbaa05d39c1e89eebec9e27

Observation 1f77cb2f-5991-4b6b-9855-bf2ec20e281b · outbound

This paper cites genomicbert: A light-weight foundation model for genome analysis using unigram tokenization and specialized dna vocabulary.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models genomicbert: A light-weight foundation model for genome analysis using unigram tokenization and specialized dna vocabulary

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.122551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.122551Z digest=sha256:7dbe594ea6ab816869315fd272e8ddfa81bc96b1bd25d5fb8337d2990d37d609

Observation d315e8f4-8923-470c-8d59-ae09b978657d · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams, 2020.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models What disease does this patient have? a large-scale open domain question answering dataset from medical exams, 2020

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.127754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.127754Z digest=sha256:f26e17e25b5c067a2897b4d2a2baf28adecba3fac021b573a5bc137dc434fa27

Observation 72b6f34d-0a3e-4460-bb5f-05331277c91c · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.133527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.133527Z digest=sha256:3e28b4f76c0881ae7b92e1fedc18786ea466c86f9d15bf76c65e2e7f60bc4669

Observation d051557d-bbdf-4e66-ad37-b485ae81a077 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.138167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.138167Z digest=sha256:7b8c4966c7bfe3cb55b00fb801650332c6418482b79500ffe58abdda1ac2dbe7

Observation 294c32de-0fb3-4f2c-97bf-68283254967c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.142851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.142851Z digest=sha256:3452ece818c9ec3b39800d67643c3ace9b2fdc759e576ca268401dcfa28478f3

Observation 39159a6a-e6dd-4b95-a462-892d9fc5b5ee · outbound

This paper cites Ai legal innovations: The benefits and drawbacks of chat-gpt and generative ai in the legal industry.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Ai legal innovations: The benefits and drawbacks of chat-gpt and generative ai in the legal industry

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.147786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.147786Z digest=sha256:b52bea412a394e8285e9fb3fcd0e02c090c03e7bb7317b94d5e4132429c36c11

Observation d0bef824-d397-49a6-a102-2737b1b7c3ba · outbound

This paper cites Ai detection’s high false positive rates and the psychological and material impacts on students.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Ai detection’s high false positive rates and the psychological and material impacts on students

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.152245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.152245Z digest=sha256:44f689d239c3934e1d62c442f5231f919d8ac13aef1ef411317a02fe7a93c800

Observation e2b870e4-d47d-4d10-b25b-0470fc9ea0f5 · outbound

This paper cites Better alone than in bad company: Addressing the risks of companion chatbots through data protection by design.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Better alone than in bad company: Addressing the risks of companion chatbots through data protection by design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.156893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.156893Z digest=sha256:90f558184761a36d6ee06cec4e18c61873053280f64820d9b08df1e7f94c2115

Observation 8619aecb-5e1c-485b-b8cd-47dcbcb6626b · outbound

This paper cites The google engineer who thinks the company’s ai has come to life.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models The google engineer who thinks the company’s ai has come to life

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.161472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.161472Z digest=sha256:af8a4f2bace67f1d1fd963b3fe9044569c98454538a83221b51e63f05bf53abd

Observation e4716fe6-318e-4300-90e8-842ae5d51971 · outbound

This paper cites Artificial intelligence algorithm for predicting mortality of patients with acute heart failure.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Artificial intelligence algorithm for predicting mortality of patients with acute heart failure

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.166063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.166063Z digest=sha256:bc6029dedfd21741a28807b97f5f5ba3a9b153d53c07c53f0c610d1185df0449

Observation d6474ca3-7925-4f31-984d-494ff3187bee · outbound

This paper cites Axiomatic attribution for deep networks.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Axiomatic attribution for deep networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.170743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.170743Z digest=sha256:c6e80b6e60046cce231a8496c1b52d861ea102a717c94ee58ef7da6bea36751d

Observation 86fa1bf6-f40d-4822-8a7f-d489a53c4da5 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.175501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.175501Z digest=sha256:56ace866c9f254d2c976f2b72895c0d85799550893370685d6253dcc13577b65

Observation 4d92a129-f3f6-4298-b479-288a9ef7103a · outbound

This paper cites A novel visual interpretability for deep neural networks by optimizing activation maps with perturbation.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A novel visual interpretability for deep neural networks by optimizing activation maps with perturbation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.180021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.180021Z digest=sha256:7f9a5e4c6412d2917b260034647a512c7dbd6618b5e5ee63eb58f263a9b2d86d

Observation 9fdc667d-ed37-4a43-9aed-8e793c0c89b1 · outbound

This paper cites Ecco: An open source library for the explainability of transformer language models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Ecco: An open source library for the explainability of transformer language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.184519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.184519Z digest=sha256:95a255b8de5f3e0dfdfad560ccec83a719bb6e76a98ceda0b370328580a28037

Observation d0e1174a-4c39-48ab-886c-8106ca1877ff · outbound

This paper cites The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.189513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.189513Z digest=sha256:08636517c8c06d3d3663fed99596e92e6d473cf33f55376fb34ed2d344b582aa

Observation 264366b1-5250-48ea-a5ac-fa903e798f0c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.198696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.198696Z digest=sha256:e9f37c4045fa05f8cba1dd25ca247ea82cc383b5e956a0b8b14c54c0cb32449f

Observation 18e35f36-3ec8-4145-88ef-fc3bffafcf00 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik Narasimhan.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Griffiths, Yuan Cao, and Karthik Narasimhan

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.203491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.203491Z digest=sha256:999c73ffbf09f794b544f53ca1d8df07dca35ef427712b25ba56f07d57951973

Observation d268c457-5503-4dfc-b413-321c3328a715 · outbound

This paper cites A Survey of LLM-based Agents in Medicine: How far are we from Baymax?.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A Survey of LLM-based Agents in Medicine: How far are we from Baymax?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.207982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.207982Z digest=sha256:23e572e9681644c34ab967419d811bfc64643f776a90c5385b96039c2cb66390

Observation e8993ab3-6649-49f4-a9c7-f15cc4d32421 · outbound

This paper cites Clinicalagent: Clinical trial multi-agent system with large language model-based reasoning.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Clinicalagent: Clinical trial multi-agent system with large language model-based reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.212897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.212897Z digest=sha256:21e883d64b9b910b946221b842671b245383ca0d9f6ce31e986800813edc8e09

Observation 5fa5062a-0710-48da-9715-b8782ca16109 · outbound

This paper cites MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:11:21.411396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.217550Z digest=sha256:0d2b49303364e86150f0c494429c0a1a661b8c9338372750a60c11a6f931d0b7

Observation f46557a9-8c58-4c6f-b9a4-27d3c2a69d07 · outbound

This paper cites ArgMed-Agents: Explainable Clinical Decision Reasoning with LLM Disscusion via Argumentation Schemes.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models ArgMed-Agents: Explainable Clinical Decision Reasoning with LLM Disscusion via Argumentation Schemes

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:11:21.388444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.222600Z digest=sha256:b7fb2bf3bc2f1febcbafd7dbe8c7cc42dbacd3cdabf9b1315a4ee52114edd99b

Observation 4c4e26f4-cdea-4170-8e2e-12ab49a37f6c · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.227414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.227414Z digest=sha256:544a491516371b598f7a6018ae45a533c1c8f59935b819dc1615f75fb96ab709

Observation 6c07370f-fa77-4b7f-bc8e-5eeaf56bce06 · outbound

This paper cites Marco-o1: Towards open reasoning models for open-ended solutions, 2024.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Marco-o1: Towards open reasoning models for open-ended solutions, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.232703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.232703Z digest=sha256:a1c8d9118e5b52e1b093f1b9a58eb399c66922795ed575685cba77580e11838c

Observation 853219d8-dc98-4d92-b692-9e1abaae7aee · outbound

This paper cites Large language models encode clinical knowledge.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Large language models encode clinical knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.237196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.237196Z digest=sha256:7e5954e8ac991c781c94eccd47b7774d54c554e6ecf02ce2091790a208f92011

Observation b234839b-bef4-4d69-af20-7f323f437755 · outbound

This paper cites Deep reinforcement learning from human preferences.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Deep reinforcement learning from human preferences

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.241573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.241573Z digest=sha256:121ea908f5fac911db037453ffc0e079ea801cbb6c1853eb17ff290d11b9aad3

Observation 0eda1fd2-3438-4fea-bd96-6a95416c272e · outbound

This paper cites GPT-4 Technical Report.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models GPT-4 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.246676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.246676Z digest=sha256:e3937f452b07efa1df0753909536b9fdfd984b909e8c0b0468c62050fb94c753

Observation 8f5d2cae-254d-4472-be67-812189a55f62 · outbound

This paper cites Training language models to follow instructions with human feedback.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.252216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.252216Z digest=sha256:b0e5809ae6f65bc0e3b1493f49efa02d0c8af9e1fa9984f291327fb5f1336b34

Observation 21efd48e-75d9-482b-9b9c-bd342ebab675 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.256822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.256822Z digest=sha256:a15e478a2ffa3516955d0f0c3e5c07e38646700ec79334fd1da153fd4b3a44fc

Observation f59ebd8d-11fe-47c3-9558-111c7f83c6e5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.261716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.261716Z digest=sha256:3bbabd3172eefad2b346e44266c13a7ba41d1ec03b55645ca6be4dbce755e8e1

Observation 0e63cb37-d081-4013-8fdf-1ff1c909d913 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.266213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.266213Z digest=sha256:fdd394dbe79edb67dc842d528988fa58524ee6ac1e77f2b248cadcff45abf67d

Observation 8b20ec94-b24b-4b4f-88ee-0cd8221b8566 · outbound

This paper cites Directed acyclic graphs.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Directed acyclic graphs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.271409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.271409Z digest=sha256:63b79343a37887e0f45910fd62877a9ed03c3af64c8f5806c90354db703c10ed

Observation f12f40c9-dc6e-49e0-9c89-e11adddc54be · outbound

This paper cites Directed acyclic graphs: a tool for causal studies in paediatrics.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Directed acyclic graphs: a tool for causal studies in paediatrics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.276579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.276579Z digest=sha256:583841c438dd187fb55c87fa4d1a1d0779e8472856b8ff78605e305ffb3f00a6

Observation 7195783f-1484-4b07-ad31-4d11bab7c387 · outbound

This paper cites Causal Reasoning and Large Language Models: Opening a New Frontier for Causality.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Causal Reasoning and Large Language Models: Opening a New Frontier for Causality

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.281295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.281295Z digest=sha256:665b637ff7083630c72c3f37f984e783047ac90e744f0487a4c931adbfbb4870

Observation 358df73d-d1cd-4921-88d4-51724c0164d9 · outbound

This paper cites Applying large language models for causal structure learning in non small cell lung cancer.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Applying large language models for causal structure learning in non small cell lung cancer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.285872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.285872Z digest=sha256:f51055b2d5b0c798357c5de48f053a12bd5f94a5fe7c635924758da8e9e89f36

Observation ac635160-9a31-4ffd-a915-baa25b15dd8a · outbound

This paper cites Inferbert: a transformer-based causal inference framework for enhancing pharmacovigilance.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Inferbert: a transformer-based causal inference framework for enhancing pharmacovigilance

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.290548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.290548Z digest=sha256:f93b94f4124a0030a5540c4fdc93c4fe0aafb3e43f2e317c10e753548bc97424

Observation 9f1cb4e5-ae72-45bd-83cd-d842f23995ec · outbound

This paper cites Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.296204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.296204Z digest=sha256:e73ba86b32fc1ca389f2a647d86e6dcde267a7e5dc4a35f4b2148247c140cd0d

Observation 4bad19f2-ab68-4d83-8a72-c7d1db89f05c · outbound

This paper cites Cambridge Handbook Of Thinking And Reasoning Ebook.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Cambridge Handbook Of Thinking And Reasoning Ebook

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.301120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.301120Z digest=sha256:5c32910da780b785dddeb57cfd423a397c78229b4b9ee78019960510b36eebed

Observation ba2a0025-d682-4dd0-b265-5c0aec825012 · outbound

This paper cites Causal models: How people think about the world and its alternatives.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Causal models: How people think about the world and its alternatives

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.306846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.306846Z digest=sha256:57611b1d41585655d9c7e1c9659a8eab151ad90c7a999e644711c47030891df6

Observation fe87cc89-f4f7-4219-a92f-026da95d3214 · outbound

This paper cites Philosophy of mathematics, 2007.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Philosophy of mathematics, 2007

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.311484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.311484Z digest=sha256:d73f0dddb4622645dae185bfa876ce5b0a26068bf204d6d89125d9da03430796

Observation 269c7d34-ecc0-4c60-97e6-2fdcbd73fd60 · outbound

This paper cites A Survey of Reasoning with Foundation Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A Survey of Reasoning with Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.316040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.316040Z digest=sha256:cb1f9e541d3d9ac03cc0d98a46c5e385a38358e29bd174eb9c1b07b467708d6b

Observation 3c4280c5-1ebd-42a1-ab3e-73a1f9e7bfe1 · outbound

This paper cites IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:11:21.269260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.321819Z digest=sha256:359194a99dcd47d9bedcaaa62bb5fb13185f0a44c162ec7d35a1269a73fc18c7

Observation bd9fed90-6dd6-4098-a272-56b2df7c81b1 · outbound

This paper cites Causal reasoning and large language models: Opening a new frontier for causality, 2024.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Causal reasoning and large language models: Opening a new frontier for causality, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.173855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.327367Z digest=sha256:5e241b6e954da458cf48387ca9d35b162b5250cd2ac6030471067136b12d158f

Observation 63a6850f-40d1-4a2f-81b7-ac062c734ff6 · outbound

This paper cites Mycin: a knowledge-based consultation program for infectious disease diagnosis.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Mycin: a knowledge-based consultation program for infectious disease diagnosis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.157664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.332504Z digest=sha256:3ca0345e9ce3546630ffc331bbe7cfeb177837bf003a849d7da724e3eafcba0d

Observation 8288469d-f0aa-404e-b1b1-51637e9d051c · outbound

This paper cites Internist-i, an experimental computer-based diagnostic consultant for general internal medicine.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Internist-i, an experimental computer-based diagnostic consultant for general internal medicine

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.140943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.337623Z digest=sha256:04da1f441db8f98d029f06aeb04365cefee73a1640622d845477f60471491284

Observation a8396c6b-a1a5-4da7-8443-b423cc0ca500 · outbound

This paper cites Neurosymbolic artificial intelligence (why, what, and how).

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Neurosymbolic artificial intelligence (why, what, and how)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.122587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.343109Z digest=sha256:610b085b0da4f19df8f2debaf0ea16d9d2c7072bf9c20c0e1ea6f99d47f841dd

Observation c2d8ee80-d596-4fde-93be-caf39ede1205 · outbound

This paper cites A brief overview of chatgpt: The history, status quo and potential future development.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A brief overview of chatgpt: The history, status quo and potential future development

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.106784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.348574Z digest=sha256:34fcff00bace00d43a6025ecaddb9e3a50fa4746e340d52a82c7a64f8a9c22a1

Observation 5371fe2b-589e-4493-aaab-48202cd78a46 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.353694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.353694Z digest=sha256:7db1a1ba0ee585244986d561598e26a5100988aa70e9c66224caca950fd46687

Observation 866c0a75-0d1c-4933-aefa-ae6ed6668b3d · outbound

This paper cites Opportunities and risks of chatgpt in medicine, science, and academic publishing: a modern promethean dilemma.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Opportunities and risks of chatgpt in medicine, science, and academic publishing: a modern promethean dilemma

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.089868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.359674Z digest=sha256:532f410206906ea969b5c6d6d249b1d11124a7c0d4025b4e235cae7cef8333c9

Observation 45afeace-17ba-4402-92eb-b339f5bf1d8f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.364690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.364690Z digest=sha256:a58939da4c509f0d62cfb23f1d1bb0af4f0cc05f4a4a423e1e64fcbadb05690c

Observation 62bd610d-9343-4fdc-ad4d-8618e6a4bbfd · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.369960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.369960Z digest=sha256:c682f03c68bfa4b5181d633589dba8165dec6da3afb45fbf773fbc6166c4f780

Observation 91d6b0ee-0ddf-4c2d-89db-6d62bab6e81d · outbound

This paper cites A generalist medical language model for disease diagnosis assistance.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A generalist medical language model for disease diagnosis assistance

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.074140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.375049Z digest=sha256:1509e42cecc1e4bc2a3de935f49123b377452746badf11bc355e643b545ef0aa

Observation 32416109-0852-4155-9d6c-2262277eb2a7 · outbound

This paper cites Crossing the trust gap in medical ai: Building an abductive bridge for xai.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Crossing the trust gap in medical ai: Building an abductive bridge for xai

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.057712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.380283Z digest=sha256:93e2951fb053d3529b03acaee300dc0a0a40f7187cd2994cada65a7bf1458b03

Observation 1cd1173a-2321-4c59-ad08-3e27102bba90 · outbound

This paper cites Mimic-iii, a freely accessible critical care database.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Mimic-iii, a freely accessible critical care database

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.385150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.385150Z digest=sha256:6cc7548ce94ee55f66ca9f9794c280b345af84b2cfac825935d6465abb276d23

Observation bee3fef1-17f6-4da2-bf45-91b068429816 · outbound

This paper cites Large language models are clinical reasoners: Reasoning- aware diagnosis framework with prompt-generated rationales.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Large language models are clinical reasoners: Reasoning- aware diagnosis framework with prompt-generated rationales

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.389936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.389936Z digest=sha256:b0bdb2a28b078f2a272b89260cb34352836db26ae90e8941371a2d135ee3ccf9

Observation 305f9309-7d5e-43e3-bfdc-ddca5a06a718 · outbound

This paper cites MedDM:LLM-executable clinical guidance tree for clinical decision-making.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models MedDM:LLM-executable clinical guidance tree for clinical decision-making

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.394652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.394652Z digest=sha256:01b35722dadba8ead603ae524a3fb3feec63dd249d348a98297de120b50ac562

Observation 450d37eb-56d2-48be-accc-684509a50405 · outbound

This paper cites Patterson, Matthew M.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Patterson, Matthew M

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.019512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.399686Z digest=sha256:c2be00867c95aeef3c964e14f149a7f3a808b4871ec84e8b231eb192623dbb8f

Observation 58aff079-5a05-4ad2-859c-7460cde792f6 · outbound

This paper cites Interpretable Medical Diagnostics with Structured Data Extraction by Large Language Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Interpretable Medical Diagnostics with Structured Data Extraction by Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.405757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.405757Z digest=sha256:ba4674e0adf5468852dfa0a9f5b591386069a1318a3ba88f27762d4aa687c0de

Observation 3717f0fb-741b-4165-aa35-85ffdfce09da · outbound

This paper cites Towards conversational diagnostic ai.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Towards conversational diagnostic ai

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:22.002253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.410923Z digest=sha256:7f9b68e8822f2dfe37f160568b653e2c88e0bed94e399773e9028397d4c465c7

Observation 66a6273e-a963-4e3b-8cb8-79e741001662 · outbound

This paper cites Towards trustworthy automatic diagnosis systems by emulating doctors’ reasoning with deep reinforcement learning.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Towards trustworthy automatic diagnosis systems by emulating doctors’ reasoning with deep reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.986328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.415730Z digest=sha256:2f182f6ee55e35528175d6628fedfbbfd1fc7f3dfab3e6ae8dd5d2a00fde2e5b

Observation ec2f4fdd-f1b8-4614-a4a5-3c4b097d832b · outbound

This paper cites MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.420408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.420408Z digest=sha256:a95036872c5abdc33dc6c5d64b4461b528267507a2cfa86d5ba6daabad10949a

Observation d5a6d98e-52c2-4d79-ae52-7ea7cf4dae0b · outbound

This paper cites Causality extraction from medical text using large language models (llms), 2024.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Causality extraction from medical text using large language models (llms), 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.970129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.425371Z digest=sha256:7ee0f366adbc8a0593fc7958e6e83ffa7d665ed0d9dbd9edbc2c1357d6210ff2

Observation 8a9df475-bf98-44d4-bc78-8fcb5a4b443c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.430700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.430700Z digest=sha256:2cc1b4584315d863a203727ee7cf9981c243187c9676beb185df89830959308e

Observation aede93e2-5a21-42be-b9ee-da469bf1a0aa · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.436314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.436314Z digest=sha256:9985759e06160f661b145ee0ca7bebcaa5a93d9090d46a7355085576b8707dc8

Observation b984a95b-70ea-4cdc-975b-1c7b8b3ef0ab · outbound

This paper cites an unresolved cited work.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:11:21.953789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.441362Z digest=sha256:9df0ae3cc2befc78b0d5f2aba7aa4a900208a2d8d2b04e149069756a7f603f15

Observation d92d1362-c50c-4c70-9bc6-72222c35a3dd · outbound

This paper cites Sok: Memorization in general-purpose large language models, 2023.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Sok: Memorization in general-purpose large language models, 2023

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.937658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.447024Z digest=sha256:d230f8602704c375ab7694b13a2ea2596cc887fc5c27d247a392e49aff282dd8

Observation dfd99550-ba9a-4b69-8720-9d165517aaa5 · outbound

This paper cites Ddxplus: A new dataset for automatic medical diagnosis.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Ddxplus: A new dataset for automatic medical diagnosis

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.451796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.451796Z digest=sha256:3adc9e5fa9a8a6cd9168d2c799f902b42cefb9dbf4826fad95c667766d732b75

Observation d21d1301-403a-4feb-91e8-c7fdb7632add · outbound

This paper cites Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.456418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.456418Z digest=sha256:621cf1eb7e4306c78fc8316ef6bc4ba863bac34d7318d8d7bd68f8476051bc74

Observation 578093fa-f539-4011-af5b-b3a83c2b0424 · outbound

This paper cites Toward clinical generative ai: Conceptual framework.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Toward clinical generative ai: Conceptual framework

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.910742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.461133Z digest=sha256:cf21c61bfd962f4061657ca6927d991db77a9c3b92357ea6f8716af032b146a4

Observation 77095d08-1604-44e8-9604-5b64697ecf21 · outbound

This paper cites Towards metacognitive clinical reasoning: Benchmark- ing md-pie against state-of-the-art llms in medical decision-making.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Towards metacognitive clinical reasoning: Benchmark- ing md-pie against state-of-the-art llms in medical decision-making

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.891989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.465910Z digest=sha256:fedbe937d8d40cc2c6791358fa2acc69ce4ada884d551c48b9b244d97e5d7b2c

Observation 1500b680-71c4-469a-b548-4630025d1eea · outbound

This paper cites Clinical reasoning of a generative artificial intelligence model compared with physicians.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Clinical reasoning of a generative artificial intelligence model compared with physicians

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.876061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.470708Z digest=sha256:162f1a78e66474e18b99413d6782e39b495c00ac2b412ba50373ed9ecd67770d

Observation 7813af97-b55d-4b8a-8249-b9ebfafba695 · outbound

This paper cites Medical reasoning in llms: an in-depth analysis of deepseek r1.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Medical reasoning in llms: an in-depth analysis of deepseek r1

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.860336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.475968Z digest=sha256:6c5b74d409b08b9d8bbc7796ea28f9f77238332cd8896b4cfd0ba94de612e098

Observation 89338951-d8e3-47ff-a305-7f5e7bfd8d5d · outbound

This paper cites Automating Expert-Level Medical Reasoning Evaluation of Large Language Models.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.481528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.481528Z digest=sha256:9c0aeb703dffd4fb99065b995fab6959c30e26dc114ea6911b6afb77d0321b0d

Observation 9364561e-c929-4754-824e-5c25d4ae3a86 · outbound

This paper cites MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.487211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.487211Z digest=sha256:5867701285e029e72c9e6e4c13134f86579aade7654753d5b86f531f15c8ba3d

Observation 8170ad43-6136-49e0-89f3-97328b26bc02 · outbound

This paper cites Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.492320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.492320Z digest=sha256:f0aca928942c507245d2b3d78534d4cf97d51d9fe24992e447bb9ebccee8d29c

Observation de82f416-5ea8-4d6f-ab76-fca4bfbeb023 · outbound

This paper cites Applying large language models for causal structure learning in non small cell lung cancer.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Applying large language models for causal structure learning in non small cell lung cancer

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.844018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.497273Z digest=sha256:3e83d79d8381cb38cff1ba053bc853d21e00382401dca5d9ea23ca5491c0ecb7

Observation 8e190327-3c99-452d-a923-659f57060191 · outbound

This paper cites A Unified Approach to Interpreting Model Predictions.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A Unified Approach to Interpreting Model Predictions

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.501964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.501964Z digest=sha256:b4f33f414995d40209696376ba22deaa2d51da79f30e5783a62c89072f32a89d

Observation 709b182f-208e-4e1f-a350-6afad2e2b36c · outbound

This paper cites Doctor XAvIer: Explainable Diagnosis on Physician-Patient Dialogues and XAI Evaluation.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Doctor XAvIer: Explainable Diagnosis on Physician-Patient Dialogues and XAI Evaluation

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:11:21.051738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.507117Z digest=sha256:16c5305dc382d450055a79f09ea419bc40f94785a365be9c4d49e0866fd2f4d2

Observation d5483a22-a496-4ade-b9a4-3e580d0d5181 · outbound

This paper cites An interactive programming learning environment supporting paper computing and immediate evaluation for making thinking visible and traceable.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models An interactive programming learning environment supporting paper computing and immediate evaluation for making thinking visible and traceable

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.827475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.512096Z digest=sha256:30ed817caf255e0e452014f1c04b0e9fc27a17dd91efc05c5ee0ddefddddb274

Observation a0485a8c-fc0e-4f02-9185-c4d1f7fd4110 · outbound

This paper cites Interactive Natural Language Processing.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Interactive Natural Language Processing

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.516947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.516947Z digest=sha256:ccebd88d83cbed4585cf541820370f1e8cea2b551f0a19096743514fbe1c18a6

Observation 2d34a395-1831-49d0-9bf4-228778ffc92c · outbound

This paper cites A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:11:21.010417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.521909Z digest=sha256:b44292778e9a99d714e68ae3554296b99325e64a13003c4f120afd32f59ba30f

Observation 15deab5b-8dee-4d19-aa50-002ed1ad2bb0 · outbound

This paper cites Chain-of-thought reasoning without prompting, 2024.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Chain-of-thought reasoning without prompting, 2024

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.527916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.527916Z digest=sha256:19ca25bc6f354180f0d61300269c2d893098f913eeb6b348c26ab1068322c677

Observation 77de7b19-f217-4a34-bb6f-0ef31375ba0c · outbound

This paper cites Qwen2 Technical Report.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Qwen2 Technical Report

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.532807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.532807Z digest=sha256:59e103d6d1e01112752d4862ffe3841e3929ceeeda2b8e66bef86772581aa64b

Observation 4313abfd-2b56-4499-bb92-fc4f9406a39e · outbound

This paper cites Qwen2.5-VL Technical Report.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Qwen2.5-VL Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.537912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.537912Z digest=sha256:7e76401932f8c332198c1cdf4450cf7210dfc3b553bef751badce3c3f0aef548

Observation 117dbaed-0527-4c2a-b82a-5aef11be4611 · outbound

This paper cites Xgboost: A scalable tree boosting system.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Xgboost: A scalable tree boosting system

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.542946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.542946Z digest=sha256:ddb8f24068c2461f1320bbea45d22549ff186671392e60a2310fe25dbca990ab

Observation 398de49e-4654-4f28-9279-f8725e162618 · outbound

This paper cites Random forests.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Random forests

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.547688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.547688Z digest=sha256:6303ae5783bc6b41a7a1d507161cce2df41cfd1792b0e9759e2314fb3a3535e3

Observation d4b981c1-78d3-465e-87c0-cbc13256c74d · outbound

This paper cites Machine learning and complex biological data.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Machine learning and complex biological data

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.779372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.552498Z digest=sha256:2fc85c0cc1ebcd55cd72e8bde9972bb75a50bdef703a2db85e07d6f597087bf7

Observation 0048b474-4fc6-430d-b863-2dd3e4313101 · outbound

This paper cites From formal boosted tree explanations to interpretable rule sets.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models From formal boosted tree explanations to interpretable rule sets

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:11:21.763289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:11:20.557098Z digest=sha256:6648922d38dc7f12b7e42540dfe85b03fb4207187d727f24bfd14b14e8f4b724

Observation 685b4c00-a69f-4a65-b121-d8a65c569eda · outbound

This paper cites Dual-process theories of higher cognition: Advancing the debate.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Dual-process theories of higher cognition: Advancing the debate

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.561743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.561743Z digest=sha256:aed907a32797ebf7b1dce2eb34191d198350427452f3c761d10cfc7b6f3dc1d5

Observation d5355a64-bd51-4f73-87bb-9afb64fd12b3 · outbound

This paper cites On the Measure of Intelligence.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models On the Measure of Intelligence

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.567574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.567574Z digest=sha256:d2cfefdca7313f05258783c335eb96aaaf02f6146b68844809ff0e7e53cb2961

Observation 22609a65-6ae5-4f63-a4d2-a6e38f09bcb6 · outbound

This paper cites ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.572820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.572820Z digest=sha256:acf81e04abb5cb561dd7abba209f7e57ad49972f5f3ec9db89d4842173c80d41

Observation fbfb7c4d-3fa1-4eb5-b46d-71aaad2a423b · outbound

This paper cites Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.577706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.577706Z digest=sha256:dc734f8236b5a4037a9f02d3eb3ade839ff728c0e8b9730d6e984ed455409adf

Observation e0dbcf38-f032-4510-9c8e-1fa152e14886 · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge, 2025.

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models From generation to judgment: Opportunities and challenges of llm-as-a-judge, 2025

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T11:11:20.582856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:11:20.582856Z digest=sha256:040ddc2cd8ff7aef29f1f8fc327e73c4c2f654d0729849e71512e67ac9b3a2a5

Pith citing papers

No inbound Pith citation observations are available.