Pith. sign in

Paper Citation Record · LEDGER

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2501.14654.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14654 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:42:29.834880Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0b8e6533-2940-4938-bc2b-30cd95ddd63f · inbound

Large Language Model Agent: A Survey on Methodology, Applications and Challenges cites this paper.

Large Language Model Agent: A Survey on Methodology, Applications and Challenges MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:52:10.740837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:51:34.309870Z digest=sha256:1a6787462a13cd21ca6c5c8656b7f827ddb35006962ad68bf57b59d7e920b0fd

Observation 766ac164-8ae7-487c-9284-875414760586 · inbound

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models cites this paper.

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:29.834880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:29.834880Z digest=sha256:3afa59fb62f6765b973cd7d5f6828b1a1d7cedea4482fb4573c9ff7802959777

Observation e9af4b11-6ba5-424b-98b2-7eec7f0c8735 · inbound

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors cites this paper.

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:18.019393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:07:58.614654Z digest=sha256:ce4ff0a3eccb4fb3d5d1522a56aaa769cf6bc7497c0899527a708e522bd22b86

Observation 5c8bd682-a9b2-4fd6-84c2-078c266d92fd · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:08.854680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T10:22:34.260242Z digest=sha256:7ab358e7524d6974139a84f9a16147538bf2019051f99f0924a24fcc387407c4

Observation f9ee0f25-79ba-4d32-86d2-bf01797963ec · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.124201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:33:52.182432Z digest=sha256:1107547f8490ea74f16165ab148b64b36ad58127576e12460bea072feca94673

Observation a4f7beb2-ec53-4c51-983b-a2703a7ce921 · inbound

ClinQueryAgent: A Conversational Agent for Population Health Management cites this paper.

ClinQueryAgent: A Conversational Agent for Population Health Management MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:33:55.639019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T01:31:07.031424Z digest=sha256:8497f3447a5467dba27ffb9e80ea4d1049c89ab2604aadfa22b5dce7ad996c3e

Observation a633ad67-e9b3-4c9b-b9af-8addec95764e · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:23.554227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:750af6096aa6d3c489f184cb547b3574a60602cc9cfb07c7b034d584cffb7d68

Observation 5bb0b44e-83f6-4333-80b3-5cdc13f30f2c · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.084048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:3592d033156d7a9c4ad75f2cc87fb1f803248477a7eb9b03876f1de17cbc222f

Observation 49c534cf-00ca-4e6f-a63f-6bdc5a13088f · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.364770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:8fce93b1755204db829e9cc13a89df31b120e4d18f5fcc9abc14fd25211d5c65

Observation 72c61c02-8158-4888-868b-fdcef8f96452 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.306082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:2e7221e38be576c70fce692bd80b65525b3bf73785ca96a9b450a3cb10cd5551

Observation 1c4768a3-0fc8-408c-9191-6c9e5ae83fe9 · inbound

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context cites this paper.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.682854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:d609236e438b4c0db7d04108326c45174987e70c05864adf644c4c5397b46cf1

Observation 5ba495e1-35b3-4e47-95ff-6e589c2b6bb4 · inbound

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales cites this paper.

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.327347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:34:09.347912Z digest=sha256:b16bb138edfe6b8c904f887b8a5bdd9c760d155153f388a83a2fbc4eec2e7e50

Observation 0a341d0f-793c-4991-b85d-cf1c49841d6e · inbound

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility cites this paper.

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.136412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:41:41.799596Z digest=sha256:277e59da356826ed9edd10a26c626c8df3733c6f3ae16a968d11c1866287583a

Observation 79b19494-493c-41c3-87a6-b2bf24459395 · inbound

AgentFairBench: Do LLM Agents Discriminate When They Act? cites this paper.

AgentFairBench: Do LLM Agents Discriminate When They Act? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:44.494148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:53:38.554457Z digest=sha256:98748b8a8671f124ab70158a2ab6cb2f001db91d1f8f894d5b2120e6c25251e3

Observation a206d07b-58fc-4504-8d1b-d6873ae52798 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.673808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:b45dcaaf3d812f25db41f2df2d18417259c127aaf8d9d8b04fb093df00f2cf6f

Observation 5035f6d7-57dc-4869-9ca7-57a148805abe · inbound

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes cites this paper.

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.396230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:33:50.358159Z digest=sha256:c8789688f681b42a221b6f81f43944baba4f2999aa8afda0dd90d7cd9046f309

Observation 6040f6c6-0641-41fd-a0a6-4cd8d4a1bd1e · inbound

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios cites this paper.

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.381699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:51:43.219551Z digest=sha256:9154946f2b55fda2b3787a886ac98b4026a0c884d8a8c933b033483ffd81b34d

Observation b3f8ce0b-193e-4b27-88d2-3383809cc627 · inbound

Cura 1T: Specialized Model for Agentic Healthcare cites this paper.

Cura 1T: Specialized Model for Agentic Healthcare MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T02:16:39.860668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:16:39.860668Z digest=sha256:bc5deaebc250270d1bfa75c883ec28bd37e177915e8208046bf507bc911b5cd5

Observation 74d61818-c0ee-4ad4-94cd-98cf5826d025 · inbound

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents cites this paper.

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T13:50:41.542608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:50:41.542608Z digest=sha256:5c630306c97432287047142e300467b2d751ace98950139f25ed34c7f143e46d

Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · inbound

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks cites this paper.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.491553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.491553Z digest=sha256:c5eaf8f41a9558e934b647b39fdf1af8d184418b2b6e2a5681809af37565e998