Pith. sign in

Paper Citation Record · LEDGER

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

As of 22 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 24 inbound Pith citation observations for arXiv:2501.14654.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14654 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:58:40.219801Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:21:26.013911Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2aff5357-71d8-4deb-af1d-2c6ee951f63d · outbound

This paper cites Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.093491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.093491Z digest=sha256:de4df2a4e4bc366bb5a5aabf1934fb73df80825485036c27370efe94dbf87c81

Observation a8b3061e-c962-4a9a-b7af-289d5f1aadfb · outbound

This paper cites Llm-based agentic systems in medicine and healthcare.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Llm-based agentic systems in medicine and healthcare

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.590327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.097648Z digest=sha256:df2352ac835648994904078e8c39604b34aefc29f38c180c787fa4a4f1baad6c

Observation 4b0caaeb-f480-4e6f-a926-1e38927c1b57 · outbound

This paper cites The rise of agentic ai teammates in medicine.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents The rise of agentic ai teammates in medicine

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.580487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.101705Z digest=sha256:a426a9a8c2f4b83a89fd2dc913fcc7e752d152276d1f090eabb04d338d473dc6

Observation d867c574-4769-4410-b609-9e304243c858 · outbound

This paper cites Implications of large language models for quality and efficiency of neurologic care: emerging issues in neurology.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Implications of large language models for quality and efficiency of neurologic care: emerging issues in neurology

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.569955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.106195Z digest=sha256:667ef7ed375794fc4f522fc5a2c432c0be81d07f170d7971e2534886d77ff662

Observation 2412c680-1b80-451a-8340-9d853e750636 · outbound

This paper cites Large language models and artificial intelligence: a primer for plastic surgeons on the demonstrated and potential applications, promises, and limitations of chatgpt.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Large language models and artificial intelligence: a primer for plastic surgeons on the demonstrated and potential applications, promises, and limitations of chatgpt

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.557870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.110602Z digest=sha256:6413d7825f5f57172c7d4eebdf2b429a08727bcb355779b0324ff045056c0d97

Observation 7da4501e-309c-4867-99c0-dbb72b28da8e · outbound

This paper cites Artificial intelligence in us health care delivery.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Artificial intelligence in us health care delivery

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.545972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.115224Z digest=sha256:177b63da527b14cea001358d7db98bc1c8a5c84d7bcc5d688de41d6194794e08

Observation 30bf36f7-9447-4da1-8ebc-552822bda452 · outbound

This paper cites How artificial intelligence could transform emergency care.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents How artificial intelligence could transform emergency care

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.534665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.119979Z digest=sha256:d0e73b0e45c94e80c5d2c6eec2f056307898b44862f53d1014f9c538dec00b2b

Observation b5793756-aad1-4139-9b6a-0c7b8c5abdd6 · outbound

This paper cites The elastic ehr: A five-tiered framework for applying ai to electronic health record maintenance, configuration, and use.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents The elastic ehr: A five-tiered framework for applying ai to electronic health record maintenance, configuration, and use

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.522979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.124841Z digest=sha256:11b46c3e834039a1bbc031179fefdca43c145d8c8dff71c6631cabc100790b7f

Observation 27188816-d92d-4dca-a7f7-3882b78e4ac4 · outbound

This paper cites Efficient healthcare with large language models: optimizing clinical workflow and enhancing patient care.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Efficient healthcare with large language models: optimizing clinical workflow and enhancing patient care

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.510868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.128585Z digest=sha256:3bd07f1e1457352a81c6462628b1e8d240f1494939b28b6e428f3cc472dd5b01

Observation 4de876c5-7960-4cc1-b41f-e780a78a6b2c · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AgentBench: Evaluating LLMs as Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.132488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.132488Z digest=sha256:78368b1b3e44fbde8dcd3ed66beed1af5b3d92022093ea14a602a5c9ab2c0b4d

Observation 6b8f3f8d-4614-49d1-888f-da52a08077b2 · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.137175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.137175Z digest=sha256:5fa6622b8d69f3f66044a4368e0ca4dc400c5d1554e83e01757a8e375516d5e0

Observation 8d04ead3-5ae7-4f01-b39a-e77c4de052d7 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Gorilla: Large Language Model Connected with Massive APIs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.142284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.142284Z digest=sha256:ae65d759a4f7b29fd65f5c64edd0c8f2ad64f74c7a799091def4fec12ef9445a

Observation 5548ec96-b172-4a32-ad8b-6e13a923eb4c · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.146824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.146824Z digest=sha256:ed81461fd39ac88c250b92936b4fffc0a24332fbd363159818eb0361ee3c35dd

Observation 6408848f-c261-472e-81ab-cfee7b486a9a · outbound

This paper cites Cyber insecurity in healthcare: The cost and impact on patient safety and care, 2024.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Cyber insecurity in healthcare: The cost and impact on patient safety and care, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.499173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.151023Z digest=sha256:6cd8f3c4b0e02e0c67bcb91c10e741f406d6d79879f0ca99bb1214ee36569fd0

Observation e046209c-3ff1-4744-bda1-cb6f176dd28a · outbound

This paper cites Trust and medical ai: the challenges we face and the expertise needed to overcome them.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Trust and medical ai: the challenges we face and the expertise needed to overcome them

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.487477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.154344Z digest=sha256:2f275ba441596dbc9304eaf81eea6633636aae7821d7aa915cfd869aee9f62d8

Observation 50c45ef1-4ea8-41da-9a78-6008d65bdcc8 · outbound

This paper cites Application of artificial intelligence in the health care safety context: opportunities and challenges.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Application of artificial intelligence in the health care safety context: opportunities and challenges

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.476129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.158168Z digest=sha256:1daada0796295ebcf5e0a257d8052cc98710849c650edd0767cea0a5ad6d689a

Observation 8ce1b9ac-8a42-4266-b148-78c38247f182 · outbound

This paper cites Ethical and regulatory challenges of ai technologies in healthcare: A narrative review.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Ethical and regulatory challenges of ai technologies in healthcare: A narrative review

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.463970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.162109Z digest=sha256:a004dcae5a6b5b14aee6254348991aa41b590eff1c8be4c2dfe74aad7f21e490

Observation 19755e88-4c29-4cc3-bb95-3acecabcb2bc · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.165855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.165855Z digest=sha256:2928ab3d442283b922268f1e4b94c7855fae66a61e8f826ea3bd3f0af1fbf4f3

Observation d113ebc2-ac5a-4905-9c6c-64dd3ef19fa0 · outbound

This paper cites Superhuman performance of a large language model on the reasoning tasks of a physician.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Superhuman performance of a large language model on the reasoning tasks of a physician

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.169554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.169554Z digest=sha256:fc9ee9e97a8e3cde76652bda54caf0794733f99b99b44771b90614a8b63da602

Observation 67e7b8e0-29a6-43b5-8a31-093b507c7c3c · outbound

This paper cites Craft-md: A conversational evaluation framework for comprehensive assessment of clinical llms.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Craft-md: A conversational evaluation framework for comprehensive assessment of clinical llms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.445398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.173816Z digest=sha256:9b1936f187e497a5d23bc26bf3969c307a3f44387a53d1ef81a651abc807de0f

Observation 957fa9e0-e783-49fe-9512-d4544dc4106a · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.177525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.177525Z digest=sha256:31995d9ac16ede101e29831e54822d9d196d21f108328712df0201ee1516a0c4

Observation a9928ed1-ce87-412d-983d-aa17129fada7 · outbound

This paper cites Large language models lack essential metacognition for reliable medical reasoning.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Large language models lack essential metacognition for reliable medical reasoning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.432977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.181451Z digest=sha256:9f4881a0597280d792f9ad9602f21ae27bb9c09a37f47f48bdc840a003e9b614

Observation facd6862-3440-427d-ae38-cab54b5df6bf · outbound

This paper cites MMedAgent: Learning to Use Medical Tools with Multi-modal Agent.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents MMedAgent: Learning to Use Medical Tools with Multi-modal Agent

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.184816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.184816Z digest=sha256:c69a45d881a87bd3f1529e2b2efe44cb06277dd1e0419eb427fc93e89adf20e2

Observation dde73f76-5b0f-482d-aa45-8a145ce9e2ba · outbound

This paper cites AutoBencher: Towards Declarative Benchmark Construction.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents AutoBencher: Towards Declarative Benchmark Construction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.189118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.189118Z digest=sha256:805ffbf93d19b95150d376fbf7c465c6beec7b6d2437b0f760c032c76b4679d5

Observation 87878499-c552-4e4a-a6a5-ca11aa6f94e6 · outbound

This paper cites Allocation of physician time in ambulatory practice: a time and motion study in 4 specialties.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Allocation of physician time in ambulatory practice: a time and motion study in 4 specialties

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.418803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.192510Z digest=sha256:02a4f4ed309ef8bebe4d8d20ce87d1c1880597d5f57ca91a2da15f3e61f25fac

Observation 533ecd96-b9b5-4cb6-9b1d-69c0dd3429bc · outbound

This paper cites Balancing act: the complex role of artificial intelligence in addressing burnout and healthcare workforce dynamics.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Balancing act: the complex role of artificial intelligence in addressing burnout and healthcare workforce dynamics

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.406361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.196324Z digest=sha256:ac94676ac74f3ac7bda9a924ff51a7a0f71860c8bd23a2b7f09bf42ec8a91e0a

Observation dab33dfd-8b8c-410c-9b7d-949c529df2c8 · outbound

This paper cites A new paradigm for accelerating clinical data science at Stanford Medicine.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents A new paradigm for accelerating clinical data science at Stanford Medicine

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.199697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.199697Z digest=sha256:e54b82ef289a8d6a2b10197cd2169eb788cd428df3af4ddffbede1df755280e6

Observation 566c462b-768c-4f74-92e5-1c45327b56d5 · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.203221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.203221Z digest=sha256:11bf7d7b0fa042245c0d88307f4ae1636d399c9a0aee3a4cdbaa8ce973ea7f47

Observation 11ff2b54-b40c-4a91-9014-f2db02eabcb5 · outbound

This paper cites Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.207451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.207451Z digest=sha256:af3f89d614a54b7f2d32d49f9dbde5104759dfa4e55fdf27d6f724b7bbdb5f2c

Observation 511fd793-ec5c-4910-95c2-eae3aa897d6e · outbound

This paper cites an unresolved cited work.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:58:40.393427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.211550Z digest=sha256:ee03017e1985e3005a33bd4d75ba2441c6b7c6545a805b80c039ae42017cbd39

Observation 99932c2b-7887-4500-b3d9-e6d9ef04860a · outbound

This paper cites an unresolved cited work.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:58:40.380553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.215600Z digest=sha256:c1782d77e79d6717567fcf594a30a190e97df20fc93e91e553f122623b4c7fbd

Observation db32833f-20ac-4c57-8c49-a4106a79c8ba · outbound

This paper cites Here is a list of functions in JSON format that you can invoke.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Here is a list of functions in JSON format that you can invoke

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:58:40.368601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:58:40.219801Z digest=sha256:1e51ca5bb3136775318cbc74a05b1b805bad39418ec056da6b8cfa4aa7baa35b

Pith citing papers

Observation 0b8e6533-2940-4938-bc2b-30cd95ddd63f · inbound

Large Language Model Agent: A Survey on Methodology, Applications and Challenges cites this paper.

Large Language Model Agent: A Survey on Methodology, Applications and Challenges MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:52:10.740837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T21:51:34.309870Z digest=sha256:6bc9b6cbf4810653a30175855b97ccd2caa88e0e0d218d335635878a55cb7974

Observation 3659cc0b-64ed-42c2-8922-6e1b5b36e509 · inbound

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems cites this paper.

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:03.354525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:03.354525Z digest=sha256:1539540e5b7b0203683283099878d8e7b655a110db0c53f4c1555af59644de20

Observation dd81056e-e4f3-4d72-9174-e2f8aadf642b · inbound

BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum cites this paper.

BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:01.830371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:01.830371Z digest=sha256:dbf734875978569f3ffdc3b44354185bad94f58cb3414681bac78e9f5813b623

Observation f4787371-308f-4d4d-86c7-dc604f5bea2f · inbound

Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees' Dialogue to Facilitate Nurse Communication Training cites this paper.

Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees' Dialogue to Facilitate Nurse Communication Training MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.679079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.679079Z digest=sha256:ab4380cd7501cefa557a5ba381982c0c98c790731ba088735567b79c0f817b39

Observation 766ac164-8ae7-487c-9284-875414760586 · inbound

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models cites this paper.

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:29.834880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:29.834880Z digest=sha256:cfa57b9604325b5f9fa82eb6baf8a6a46fc4e8c776780decba0ad706c3385d03

Observation e9af4b11-6ba5-424b-98b2-7eec7f0c8735 · inbound

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors cites this paper.

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:18.019393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T22:07:58.614654Z digest=sha256:0844bf5cb466c9017800c28cb294f99e2cdf36c0bee1faf1fd32c4bdf5761e34

Observation 5c8bd682-a9b2-4fd6-84c2-078c266d92fd · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:08.854680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T10:22:34.260242Z digest=sha256:9648184922129d8f40205ad7ae4a3c757ddd12a105468273adee31a5f8c7a333

Observation f9ee0f25-79ba-4d32-86d2-bf01797963ec · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.124201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:33:52.182432Z digest=sha256:5994076651b07a4498472d3a1f7cba13048b52692ddc826c34ca0b919d81bd7d

Observation a4f7beb2-ec53-4c51-983b-a2703a7ce921 · inbound

ClinQueryAgent: A Conversational Agent for Population Health Management cites this paper.

ClinQueryAgent: A Conversational Agent for Population Health Management MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:33:55.639019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T01:31:07.031424Z digest=sha256:f7ca6dbb86b63d26009784fced6425a30c11621e6bbdce45bd24da43ce875543

Observation a633ad67-e9b3-4c9b-b9af-8addec95764e · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:23.554227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:9b4db53148f2dfd5da843c3e6e6324a7153635ffe6562c70a4f84c9f2a1d5f64

Observation 5bb0b44e-83f6-4333-80b3-5cdc13f30f2c · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.084048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:0a4204160c31a3684d6b7b4abd3cd0d6950ca2229fc6f967c53cdfb044fda021

Observation 49c534cf-00ca-4e6f-a63f-6bdc5a13088f · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.364770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:8a9773f581dd6848660831062a5bd67af9bb5d10358fbea30975c927b6444bbd

Observation 72c61c02-8158-4888-868b-fdcef8f96452 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.306082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:8a2201e48fbb5416b1451c6251322e0a873bdbcec32ac1b367de888c5fbfc995

Observation 1c4768a3-0fc8-408c-9191-6c9e5ae83fe9 · inbound

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context cites this paper.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.682854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:068affd04c428b1235333a8de7095185a8cda18646ea8ee019269460325be9bc

Observation 5ba495e1-35b3-4e47-95ff-6e589c2b6bb4 · inbound

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales cites this paper.

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.327347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T09:34:09.347912Z digest=sha256:a63bb15395779f9d45e7a00b9419a3c7e4437b75703bf63e9bf236dbc1cb5b80

Observation 0a341d0f-793c-4991-b85d-cf1c49841d6e · inbound

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility cites this paper.

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.136412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T06:41:41.799596Z digest=sha256:b0ccd0efd284b44a7633a075896e5dd39a7a333822ae22fcc263d85cbdfd70c8

Observation 79b19494-493c-41c3-87a6-b2bf24459395 · inbound

AgentFairBench: Do LLM Agents Discriminate When They Act? cites this paper.

AgentFairBench: Do LLM Agents Discriminate When They Act? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:44.494148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T03:53:38.554457Z digest=sha256:6bdcc1fae0c4d9a01edaf8d15365c5e5746f878af5cf4f9078bd251ad358731f

Observation a206d07b-58fc-4504-8d1b-d6873ae52798 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.673808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:c9f288b866b424858e2912f41e8c797168fcacd3a21adfc8679d6e1c7ac0ecfd

Observation 5035f6d7-57dc-4869-9ca7-57a148805abe · inbound

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes cites this paper.

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.396230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T09:33:50.358159Z digest=sha256:6c6e288a1e731b2c669c34cc18af581a96c44bda2114aa26b9ff45ef40d2e105

Observation 6040f6c6-0641-41fd-a0a6-4cd8d4a1bd1e · inbound

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios cites this paper.

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.381699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T06:51:43.219551Z digest=sha256:ac3c0b4eb52cb3b0ef0be96d33b4059f138b657cd5607eaca47327d8625edc13

Observation b3f8ce0b-193e-4b27-88d2-3383809cc627 · inbound

Cura 1T: Specialized Model for Agentic Healthcare cites this paper.

Cura 1T: Specialized Model for Agentic Healthcare MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T02:16:39.860668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:16:39.860668Z digest=sha256:890f8daedf1721927bc59d67de6bc458e97f21544626a93acfd6818c092a677f

Observation 74d61818-c0ee-4ad4-94cd-98cf5826d025 · inbound

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents cites this paper.

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T13:50:41.542608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:50:41.542608Z digest=sha256:6f732483ea2ed2e1c4fa1a91b6caf84bcb67cb5ad6a61cf7ea829eddbb7a7b70

Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · inbound

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks cites this paper.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.491553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.491553Z digest=sha256:28b4f861d7416f2fa299025193db78940d799067d354beacaa597e9052b7a589

Observation 8ec4d251-36cf-45e9-8d5b-7c63b3544076 · inbound

ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making cites this paper.

ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:21:26.013911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:21:26.013911Z digest=sha256:7a17c5f64f5240b9234ec8d28e914be4e312ed68bd07553ff57d48984bbba1c2