Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:58:40.219801Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 24 inbound Pith citation observations for arXiv:2501.14654.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:58:40.219801Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:21:26.013911Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
32 of 32 outbound references displayed
External citation measurements
4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 2aff5357-71d8-4deb-af1d-2c6ee951f63d · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b3061e-c962-4a9a-b7af-289d5f1aadfb · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Llm-based agentic systems in medicine and healthcare
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4b0caaeb-f480-4e6f-a926-1e38927c1b57 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents The rise of agentic ai teammates in medicine
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d867c574-4769-4410-b609-9e304243c858 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Implications of large language models for quality and efficiency of neurologic care: emerging issues in neurology
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2412c680-1b80-451a-8340-9d853e750636 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Large language models and artificial intelligence: a primer for plastic surgeons on the demonstrated and potential applications, promises, and limitations of chatgpt
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7da4501e-309c-4867-99c0-dbb72b28da8e · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Artificial intelligence in us health care delivery
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 30bf36f7-9447-4da1-8ebc-552822bda452 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents How artificial intelligence could transform emergency care
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b5793756-aad1-4139-9b6a-0c7b8c5abdd6 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents The elastic ehr: A five-tiered framework for applying ai to electronic health record maintenance, configuration, and use
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 27188816-d92d-4dca-a7f7-3882b78e4ac4 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Efficient healthcare with large language models: optimizing clinical workflow and enhancing patient care
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4de876c5-7960-4cc1-b41f-e780a78a6b2c · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AgentBench: Evaluating LLMs as Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8f3f8d-4614-49d1-888f-da52a08077b2 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d04ead3-5ae7-4f01-b39a-e77c4de052d7 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Gorilla: Large Language Model Connected with Massive APIs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5548ec96-b172-4a32-ad8b-6e13a923eb4c · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6408848f-c261-472e-81ab-cfee7b486a9a · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Cyber insecurity in healthcare: The cost and impact on patient safety and care, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e046209c-3ff1-4744-bda1-cb6f176dd28a · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Trust and medical ai: the challenges we face and the expertise needed to overcome them
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 50c45ef1-4ea8-41da-9a78-6008d65bdcc8 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Application of artificial intelligence in the health care safety context: opportunities and challenges
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8ce1b9ac-8a42-4266-b148-78c38247f182 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Ethical and regulatory challenges of ai technologies in healthcare: A narrative review
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 19755e88-4c29-4cc3-bb95-3acecabcb2bc · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d113ebc2-ac5a-4905-9c6c-64dd3ef19fa0 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Superhuman performance of a large language model on the reasoning tasks of a physician
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e7b8e0-29a6-43b5-8a31-093b507c7c3c · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Craft-md: A conversational evaluation framework for comprehensive assessment of clinical llms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 957fa9e0-e783-49fe-9512-d4544dc4106a · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9928ed1-ce87-412d-983d-aa17129fada7 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Large language models lack essential metacognition for reliable medical reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation facd6862-3440-427d-ae38-cab54b5df6bf · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde73f76-5b0f-482d-aa45-8a145ce9e2ba · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AutoBencher: Towards Declarative Benchmark Construction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87878499-c552-4e4a-a6a5-ca11aa6f94e6 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Allocation of physician time in ambulatory practice: a time and motion study in 4 specialties
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 533ecd96-b9b5-4cb6-9b1d-69c0dd3429bc · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Balancing act: the complex role of artificial intelligence in addressing burnout and healthcare workforce dynamics
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dab33dfd-8b8c-410c-9b7d-949c529df2c8 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents A new paradigm for accelerating clinical data science at Stanford Medicine
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566c462b-768c-4f74-92e5-1c45327b56d5 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Many-Shot In-Context Learning in Multimodal Foundation Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ff2b54-b40c-4a91-9014-f2db02eabcb5 · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511fd793-ec5c-4910-95c2-eae3aa897d6e · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99932c2b-7887-4500-b3d9-e6d9ef04860a · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db32833f-20ac-4c57-8c49-a4106a79c8ba · outbound
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Here is a list of functions in JSON format that you can invoke
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0b8e6533-2940-4938-bc2b-30cd95ddd63f · inbound
Large Language Model Agent: A Survey on Methodology, Applications and Challenges MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3659cc0b-64ed-42c2-8922-6e1b5b36e509 · inbound
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd81056e-e4f3-4d72-9174-e2f8aadf642b · inbound
BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4787371-308f-4d4d-86c7-dc604f5bea2f · inbound
Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees' Dialogue to Facilitate Nurse Communication Training MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 766ac164-8ae7-487c-9284-875414760586 · inbound
A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9af4b11-6ba5-424b-98b2-7eec7f0c8735 · inbound
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5c8bd682-a9b2-4fd6-84c2-078c266d92fd · inbound
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f9ee0f25-79ba-4d32-86d2-bf01797963ec · inbound
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a4f7beb2-ec53-4c51-983b-a2703a7ce921 · inbound
ClinQueryAgent: A Conversational Agent for Population Health Management MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a633ad67-e9b3-4c9b-b9af-8addec95764e · inbound
Design and Report Benchmarks for Knowledge Work MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5bb0b44e-83f6-4333-80b3-5cdc13f30f2c · inbound
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 49c534cf-00ca-4e6f-a63f-6bdc5a13088f · inbound
UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 72c61c02-8158-4888-868b-fdcef8f96452 · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1c4768a3-0fc8-408c-9191-6c9e5ae83fe9 · inbound
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5ba495e1-35b3-4e47-95ff-6e589c2b6bb4 · inbound
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0a341d0f-793c-4991-b85d-cf1c49841d6e · inbound
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 79b19494-493c-41c3-87a6-b2bf24459395 · inbound
AgentFairBench: Do LLM Agents Discriminate When They Act? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a206d07b-58fc-4504-8d1b-d6873ae52798 · inbound
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5035f6d7-57dc-4869-9ca7-57a148805abe · inbound
MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6040f6c6-0641-41fd-a0a6-4cd8d4a1bd1e · inbound
An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b3f8ce0b-193e-4b27-88d2-3383809cc627 · inbound
Cura 1T: Specialized Model for Agentic Healthcare MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d61818-c0ee-4ad4-94cd-98cf5826d025 · inbound
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · inbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ec4d251-36cf-45e9-8d5b-7c63b3544076 · inbound
ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.