Pith. sign in

Paper Citation Record · LEDGER

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

As of 8 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2607.25485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25485 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:21:23.281512Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved89
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3766f1de-5153-41a0-9f76-2114a82dc3f9 · outbound

This paper cites 2001 , publisher =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2001 , publisher =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:17.996665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:17.996665Z digest=sha256:678490d368538d514f9b4dab8c7267d7d81268a2c7775b0a8337a631ffbbd107

Observation 5ed8be90-8cd6-4524-a1ca-f98a0bdf0cbe · outbound

This paper cites 2021 , url =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2021 , url =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.069957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.069957Z digest=sha256:8e206208dec32cfe63900b4bb7604daf66e58b425bee08660bd4399aba0ba164

Observation 1afb11b6-a59a-48e2-b58e-8303e11fa8d5 · outbound

This paper cites 2022 , note =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2022 , note =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.148970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.148970Z digest=sha256:ff7c155b8848d2271d34cbca0845ef3076ac4dde86d2c3eb6743d0e07be97302

Observation 4aabc5c4-edab-4321-a214-5471f0c5f490 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.249494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.249494Z digest=sha256:566ed826f965ff84ab88e67da3c8ee11e779e978cfee3fa986b3b818a59a7c3c

Observation c82ab66f-a41e-4d8b-9e56-33f1f51c5570 · outbound

This paper cites BMJ Quality & Safety , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents BMJ Quality & Safety , volume =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.327124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.327124Z digest=sha256:b255c15bb76048e957cbed9bdf307337afd868d146c6e0933899059742a6bacb

Observation a54b7eb0-1321-436b-87ec-04e7062fadd2 · outbound

This paper cites 2023 , howpublished =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2023 , howpublished =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.401279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.401279Z digest=sha256:8e308875182317c872b3862ae0b8292c07628268cac6789911f4b83fef769546

Observation d6f90ebe-28b7-42cf-80bf-7718e915c2c6 · outbound

This paper cites 2015 , publisher =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2015 , publisher =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.556940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.556940Z digest=sha256:c949a3bc00aec45cb8503d7084d3a54f50005017bee16e98f369651043f68455

Observation 2d6fbe6c-a523-440f-afa9-0bbbdded1c0b · outbound

This paper cites HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.679824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.679824Z digest=sha256:eb5cc05d84239ee27ce6e61f5e9fc823227167fd0a9d7def6c3f1bbaef5e83ce

Observation 5133b7e0-ecc1-411a-8ae2-c76bb59ada90 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.739263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.739263Z digest=sha256:fc76ede967da1255f5925ce8d2a3cf729e0e76e9f4176f487dd795dfe6358840

Observation 82dc9853-a72c-41d4-9b8e-5c731005aeaf · outbound

This paper cites 2024 , url=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2024 , url=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.871029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.871029Z digest=sha256:3ec895d6c6e89082ed487b8dce348b1d215e3142dea5fbcd8586d45c1049f4cc

Observation b8540d1f-2cf4-4128-a553-661017145689 · outbound

This paper cites , journal=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents , journal=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.895834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.895834Z digest=sha256:c56f18c00ff2b8c1230c836d9b8fcedab456dcbb20ce117c197e4e8f8f482b0d

Observation 5145c540-da71-4c6a-a9a6-581558ec44e3 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:18.977740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:18.977740Z digest=sha256:d7214c598f5a0e5bf6cafcb3f4e0b9bf23ab015854f7664970f25c39a2f2ee95

Observation e2b3a5c0-21ca-4b6b-b4da-62c8126118a7 · outbound

This paper cites npj Digital Medicine , volume=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents npj Digital Medicine , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.053604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.053604Z digest=sha256:4687e374eeff557631342cb27da350ad529856ebee1e3a6618650b580ad7260c

Observation 2ddafd12-fac0-4475-94fc-0131f2e7749b · outbound

This paper cites Nature Biomedical Engineering , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature Biomedical Engineering , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.137495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.137495Z digest=sha256:1d3ce72abe8939fb604d6866b139b486d15a43c2a255abe8c21cf5cc36162be3

Observation 572748a5-1d98-47a9-b633-b3312eebc609 · outbound

This paper cites NPJ Digital Medicine , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents NPJ Digital Medicine , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.237080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.237080Z digest=sha256:9fd794659c269bfc6a37388551ef6e96a9606b0cab8a3f48878b6a0457dda772

Observation 4340bb15-0dbd-49fc-b6d7-07cf854f9fd7 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.280290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.280290Z digest=sha256:8fd26d34b57c18405a78b7da31186163c113c77300a0998bc1e3dabec2db2365

Observation 84c1942a-39a1-4798-ba74-105204196ce2 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.358934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.358934Z digest=sha256:b453f0cabedd00bf16497537301c4f965123d28ccda2c7f8f041f105d8b4de31

Observation 12ff1d21-88bf-4091-90ef-80a748d623ae · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.448307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.448307Z digest=sha256:6eb95953402398d0e8a361f5e308083d30d39b97cce2d721e330f0828b9b18bc

Observation 2b6e8cb1-25a5-40c3-b37c-a349317de069 · outbound

This paper cites Patient Education and Counseling , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Patient Education and Counseling , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.510948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.510948Z digest=sha256:db287d83906078762fe91fd62e30319bda8954eee4614918b0faa31da0f31b1e

Observation 6a204367-a294-4bf0-bad3-e84c670ff169 · outbound

This paper cites BMJ , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents BMJ , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.606673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.606673Z digest=sha256:8fa6cda770428b336bd6f766511ddbd99ca34c83853f955b18132e787ba6ba3c

Observation e595d080-e486-42c3-a7f9-d943664b6795 · outbound

This paper cites Nature Medicine , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature Medicine , year=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.660341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.660341Z digest=sha256:c1004a18919ef974aad1c6c0dbff47f0612c5782031bb7498bfc7394d01432db

Observation 4e189cb0-cdc7-47e4-a618-48d48e14c39e · outbound

This paper cites Artificial Intelligence Review , volume=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Artificial Intelligence Review , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.766031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.766031Z digest=sha256:1e79fc9535fc9b1b5171226fc07b580783bda6e3eff2a0ef0de9a41391044dd9

Observation 5bfc97ef-4af7-49d9-bd9e-a2d335722588 · outbound

This paper cites JAMA , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents JAMA , year=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.788320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.788320Z digest=sha256:e001da351c897431b5b373ae05a3bf333d7794da3b1e95217a9e009be265e17e

Observation 7fdc98d7-198d-4182-a7e4-758b084326f3 · outbound

This paper cites Nature , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.865287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.865287Z digest=sha256:7b0ce06191d95b3e58c2773c6fbdefbb4b326d38fd6d581eb3aea7f720373015

Observation cedf752f-0159-4c4a-ad6e-dfc543addee8 · outbound

This paper cites NEJM AI , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents NEJM AI , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.921549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.921549Z digest=sha256:a5eb177f17cdbc1409c6a17b2e64d03ea0735e3b43aed330dfcb558e015e0a1b

Observation 7271ba60-670a-4192-a716-d03538f15416 · outbound

This paper cites Science , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Science , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:19.974313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:19.974313Z digest=sha256:ea473ee2173c4bd5c0901061f65df5639a34ba97fe84e652df42305bc39cb0a7

Observation 771c716b-d1dd-45b6-b0ac-f1a3ae886674 · outbound

This paper cites The Lancet Digital Health , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents The Lancet Digital Health , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.031170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.031170Z digest=sha256:09304e6b8c513c75d64f4269dfba3ed0ad5f2242f9ec7b17bb5d4a1b941c3f01

Observation aed0c3c7-2485-448e-900e-c406ac43fea6 · outbound

This paper cites NPJ Digital Medicine , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents NPJ Digital Medicine , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.121386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.121386Z digest=sha256:03ef25163779326c8e86f6f4c404aa36fa811f646e7fc6bbea46327f2d4a56c4

Observation c47e663e-93b1-4917-ac52-efc274eba4ab · outbound

This paper cites Journal of the American Medical Informatics Association , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Journal of the American Medical Informatics Association , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.170241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.170241Z digest=sha256:78728009ce35fecb40c2bad90e2c5e869619181e221137201e27508cbe526bc1

Observation dd80af87-8579-4405-b175-98d7cf812c7d · outbound

This paper cites Health Affairs , volume=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Health Affairs , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.229275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.229275Z digest=sha256:a37aa33358882c3f755cfe52d02aa8eca249eec5e05b0884255730c3a6c6b836

Observation 7a43b0ee-7c2f-4355-8db6-767ef1df19bd · outbound

This paper cites 2007 , doi=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2007 , doi=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.304733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.304733Z digest=sha256:2f5060d9e56bc5e3ebe28c63e5757179d24c1fab5c420eecc5ce0a47fd331181

Observation 39c08e43-e7f2-4c98-bf2b-6245d031a19a · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.359573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.359573Z digest=sha256:93259e28ab4273896755cb2afc87f54b19c0ef2de56d7ed360aafc17b0893879

Observation 340f490a-9557-44c3-9e62-c9be6692e907 · outbound

This paper cites Nature Medicine , volume=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature Medicine , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.483768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.483768Z digest=sha256:d83ac89f95dd902ca9c7022b0265a4e5f1a31d74c3a21adf0c0eb3052562a669

Observation 2774a0b5-9ace-474d-b3db-4890fad0a887 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 34

Resolution
parse uncertain
no resolver link, observed 2026-08-01T02:21:20.580478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.580478Z digest=sha256:6698b2c2a6f9d620638706a28c8b93a0066a9aab603d88c2d3532c95efc94a08

Observation 83821be7-b68c-4cb1-90d5-3469c448df4f · outbound

This paper cites 2025 , url=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2025 , url=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.689928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.689928Z digest=sha256:1f17fdac8f200f45bb99663e995e9080e9bf87c9e8d4c727a3641ebf05f8cc9a

Observation 40530c6b-34a6-449a-a484-d44c8c753e4d · outbound

This paper cites 2024 , url=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2024 , url=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.834219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.834219Z digest=sha256:48fd9b27c6a651998b30d53839db260ff2b774838613d05e41c4d58e7bd29bf9

Observation 19f2360d-aafb-4665-90e9-21aab131394d · outbound

This paper cites 2025 , eprint=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2025 , eprint=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.895142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.895142Z digest=sha256:2187e57775ea802768a7289047cc3ae4120f0b58685c2a192a81b40d42d731d6

Observation 06bd87c1-47d3-407d-888b-4f128e16389a · outbound

This paper cites Preference Leakage: A Contamination Problem in.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Preference Leakage: A Contamination Problem in

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.974125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.974125Z digest=sha256:1a205b5c736dd47bf55425f1ff2effb9502d5ee393c5af2183fda0930ced4a1e

Observation 0658af29-991a-41e1-a698-6f18bbadb162 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Self-Preference Bias in LLM-as-a-Judge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.070711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.070711Z digest=sha256:2ad3d3b25d96f337f075e831fe5d8638b3e7f145207b17e4b1b4d1b524a799be

Observation 92d99b06-3f87-4c21-a904-a0c9eaa4440c · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents LLM Evaluators Recognize and Favor Their Own Generations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.118082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.118082Z digest=sha256:27917d8265a0429b6d109cf5b737dc774eeb18642b085c4e667ab8638fbecf2b

Observation 8b3fba1c-eb8e-4ca3-b266-5d23955820f0 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents International Conference on Learning Representations (ICLR) , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.166050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.166050Z digest=sha256:4b99c2a0da95ae68e95d9df21e093831d47602d0be524aba92f47df01d85e0b6

Observation 913a0366-970a-461d-a17a-14865e2d0ea9 · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.301494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.301494Z digest=sha256:e4f2ea2c84497173f1e956bfe99833f2f80574b33f55e954fb4d0051d5d1155b

Observation 57ce7a47-62f4-4c2c-87f6-463013feeb28 · outbound

This paper cites Scientific Data , volume=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Scientific Data , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.360242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.360242Z digest=sha256:8ffd944770f5eb6bd3879c36c584788c95205d2e39c12930076e591a501cf190

Observation e726d091-3c72-4722-9af6-8bb128dc51e4 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.379807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.379807Z digest=sha256:c6346652a1e1044d968094b3ac85c4788f8fca8baf57a9d67ef03390e45dabaf

Observation 919727a7-8220-416f-bfff-9ce2953554a0 · outbound

This paper cites ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.414682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.414682Z digest=sha256:06c7e35401bf782790f8adb9746f43799af1bafed89035235ccd6529a62b0d05

Observation 1750fc42-26ec-4ad0-9c54-366008523316 · outbound

This paper cites Applied Sciences , volume=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Applied Sciences , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.452299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.452299Z digest=sha256:9c833f96df6b46cee20db60c96f86b2e430254e16030594a68fe2d3f2ba214cf

Observation 1c1951a3-9aee-448d-8aa8-70628b357b14 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents International Conference on Learning Representations (ICLR) , year=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.486668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.486668Z digest=sha256:17f3efcd929181d11efd1584829085563af3824949cda81320eedaf98fff8153

Observation 9e80aeb1-43f5-4fac-95fe-d408bc533dbe · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents International Conference on Learning Representations (ICLR) , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.516041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.516041Z digest=sha256:6e273e11543cd84bbd0401138eda473f8e0b7252d2b020d7c23976573b973671

Observation 4355c556-d7ab-47f6-b90c-abe9f23a7ecd · outbound

This paper cites PARADISE : a framework for evaluating spoken dialogue agents.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents PARADISE : a framework for evaluating spoken dialogue agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.563135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.563135Z digest=sha256:e350dc64292f381b40c73a0f39e3e2906f9b1afc5ff041617d55ff2447025122

Observation 4bedea21-e4f8-4cfd-b789-0f3487d0d447 · outbound

This paper cites Towards an automatic Turing test: Learning to evaluate dialogue responses.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Towards an automatic Turing test: Learning to evaluate dialogue responses

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.612110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.612110Z digest=sha256:f544670a88d6c687dfbd92bb5af4f071f6ccfc41fc57c1b41ef2640843b96d23

Observation 81cfd0c9-d0ce-4fb6-a6a4-312c336c1c78 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.671396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.671396Z digest=sha256:8ef24ea2b337b4e5993c2d2e02b727b8514404bfa8eeeaa92b6b1db396ca43e7

Observation 2fc661a9-25c8-48d8-a067-79c2490aee2c · outbound

This paper cites Proceedings of the 31st International Conference on Computational Linguistics (COLING) , pages =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Proceedings of the 31st International Conference on Computational Linguistics (COLING) , pages =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.720085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.720085Z digest=sha256:1fa93b5e999788af189c99411148b2a92e1eaffc80688208ab0a0688f51ee059

Observation 889aafdc-abcd-4541-b19f-22104f8d90e2 · outbound

This paper cites arXiv preprint arXiv:2601.03023 , year =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents arXiv preprint arXiv:2601.03023 , year =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.756090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.756090Z digest=sha256:22804ef082bab351e9ca3f8e086f5331a1735550b8904e8ed57b8a572e617e8c

Observation 358a967d-b42d-4908-b5ce-b193a960e88f · outbound

This paper cites and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.800471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.800471Z digest=sha256:8c608ea1f6e5452fbc464cd66f1ddd99c87b38f2a82bf03de8cc3acbe0303a2b

Observation 192c1a86-9ffc-4b4a-a1d5-ce80e0710139 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track , volume =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.840303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.840303Z digest=sha256:1b56eb227ed397deff86d0750777947be4710a9a1e2736706082eeb3d4995593

Observation 7f1af8b4-ae71-4e87-864a-c6bfca68d5d0 · outbound

This paper cites and He, Junjun and Qiao, Yu , booktitle =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents and He, Junjun and Qiao, Yu , booktitle =

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.893864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.893864Z digest=sha256:aaff815b2489f83e5e9a2cf3c5f0efa27ef9e766c389c2dc8990fba3f8b4e306

Observation a30e99a9-1327-476c-b4bc-e0e6128631df · outbound

This paper cites ClinicalBench: Can.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents ClinicalBench: Can

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.934658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.934658Z digest=sha256:acfa8c30ef3600416b1552c8eada761475affba8b6049522da2ed24b6adb2bf5

Observation d88614ee-8082-4a69-9e1d-aa4c9220d53a · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.982167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.982167Z digest=sha256:5c4ccac42f2bfbc61b416893a18995af6ca63e44221217cf5a2a8f6e5a943d8f

Observation d3b2b9ca-5269-43b2-8cfa-aaefcdd9ef89 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.026928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.026928Z digest=sha256:c346ea7027d5cfa63add07b473e9e9e1288d65dae272bb87cd7e6c3f025e5241

Observation 68867839-324e-42d0-8aba-4a4586818b24 · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.083147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.083147Z digest=sha256:3387da57543fcf34fb4aa1b6445daa3372fbe213c4e46aca430abb899009249d

Observation 8fa8e008-fdcf-4db9-8819-51b6ca835571 · outbound

This paper cites an unresolved cited work.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.136648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.136648Z digest=sha256:5e9dc0994f2dfeb04b3e206e08fdf491a23f05cf2816b03682cc8a325cc6e3ab

Observation 63aec43f-d4ba-4da9-bc64-31b744f43b7d · outbound

This paper cites and Mao, Huanzhi and Yan, Fanjia and Ji, Charlie Cheng-Jie and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents and Mao, Huanzhi and Yan, Fanjia and Ji, Charlie Cheng-Jie and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.234690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.234690Z digest=sha256:9be8be05a2b5120661e8079ffe8cdbe43ebc663ad86886c483a0ef7b67fb9db8

Observation de1fbce5-bb35-481b-bb7e-5a7a6ded9627 · outbound

This paper cites Journal of Medical Internet Research , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Journal of Medical Internet Research , volume =

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.258825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.258825Z digest=sha256:6e70d1e5707d31bf8f22881f8de1af24900d5f5075289955879378c41036781e

Observation 5c80304c-4a3e-4169-b328-da3805979977 · outbound

This paper cites Journal of Medical Internet Research , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Journal of Medical Internet Research , volume =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.290120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.290120Z digest=sha256:c0b63e6fe249cb4bae9b10323496ff7ee87eae4ce47c321bff8ceeb431a31e1a

Observation 5467353b-d1c7-487a-9086-04c91a2817b6 · outbound

This paper cites Journal of the American Medical Informatics Association , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Journal of the American Medical Informatics Association , volume =

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.330553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.330553Z digest=sha256:95209600888c5836c447382791679c4c4429aebb7fccb0c4a6f771f639367697

Observation 3956f1df-4c78-4a65-bce7-0251293db865 · outbound

This paper cites Journal of Medical Internet Research , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Journal of Medical Internet Research , volume =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.381143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.381143Z digest=sha256:41e1e22f90faf9b2728237bd4affc32a1ada961c9aa1c600dbb1d238a25f91f9

Observation 99e3c458-7e76-4afc-ba8c-49657f8a1759 · outbound

This paper cites npj Digital Medicine , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents npj Digital Medicine , volume =

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.413787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.413787Z digest=sha256:b8f0cecd5dec7f722c477b035eff2fbc6446310926ed25ba35072775492439a1

Observation 621d281b-c99d-4e1b-a222-189c274ff6f8 · outbound

This paper cites Journal of Medical Internet Research , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Journal of Medical Internet Research , volume =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.442579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.442579Z digest=sha256:519fb0955a50a35d4c7049eb429362fc07d8888f7310910e2ba9e5bfd1601596

Observation 2e8d33b3-5744-4a1c-b2d5-aa4b5cd71ac2 · outbound

This paper cites PLOS ONE , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents PLOS ONE , volume =

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.475559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.475559Z digest=sha256:cb9cfa0867e16854fe4efb3eb46f2946061e5dc3a10e03c3457fb89ad9abca26

Observation daf8417d-68ec-4074-83f5-fa72605b7898 · outbound

This paper cites BMJ Open , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents BMJ Open , volume =

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.498424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.498424Z digest=sha256:ba2b7410594e4aea1820477e62014852f0b225a8b9f2705bd3160cd645b1b764

Observation f3c6c9d6-cbec-4b73-bb34-fc4e24f7b3d7 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.521574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.521574Z digest=sha256:52a76e58e27d2370863aa1dad0b71190551f04ec38fa647364da758419017200

Observation 1d8f8230-c9c7-4dc7-9701-d771716a0fb8 · outbound

This paper cites Nature , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature , volume =

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.554244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.554244Z digest=sha256:3938df52849a20bfe54aa5dfc0e09bbadda3d8fcbd3da34a87e41099ce123647

Observation 3aaf1d19-7dfc-4022-a09f-92b15b08ef08 · outbound

This paper cites Nature Medicine , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature Medicine , volume =

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.597328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.597328Z digest=sha256:191c33be8cfd4f2cb0c3d73c032418be7990f535f29938d4fc19cdbd89b94941

Observation ed50d4a6-c454-4aed-8004-7289d0df48b2 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.633720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.633720Z digest=sha256:5a9740320ce160ddb916fcc4006761724911033ccdbcc2750dff5bfb227fa82e

Observation b99d2425-4fbf-4224-b2fc-4855c7c9b477 · outbound

This paper cites Annals of Internal Medicine , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Annals of Internal Medicine , volume =

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.662260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.662260Z digest=sha256:b45acf8117ea2b29138bfb512f900f2156a0ab18f09e002eb99584da4e1d071d

Observation 1d8974f7-43a0-4a1b-aa11-0673bffc94a4 · outbound

This paper cites Annals of Family Medicine , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Annals of Family Medicine , volume =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.672978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.672978Z digest=sha256:cd2a651e7a311d1517f72df98a9c92f26e646e2491ceefd71e74357214ee33ef

Observation 24a98490-6834-41b2-89e2-18859d70cc3c · outbound

This paper cites 2026 , note =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2026 , note =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.701316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.701316Z digest=sha256:ba9eed067c54ac9c8464b0931a285e2381afad037f9a8fb27dc59a1187447568

Observation ffa6a04d-d682-4a41-917b-a8e4c1feecd1 · outbound

This paper cites 2025 , eprint =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents 2025 , eprint =

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.761287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.761287Z digest=sha256:118ba218a60473e82aec9e98cfdd49d53f5f123ec6e8ed33cb1601df0f6b119f

Observation a22b90e4-3f35-4165-8489-570cfa65064d · outbound

This paper cites arXiv preprint arXiv:2511.18491 , year =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents arXiv preprint arXiv:2511.18491 , year =

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.796414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.796414Z digest=sha256:567b45e1ad202e57d5deb003291a7847a77229818c3b29aa0b988d9f076d802e

Observation 44199887-95a3-4403-b665-af61c1d84a11 · outbound

This paper cites Nature , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature , volume =

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.826642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.826642Z digest=sha256:e8102261af77e7a2a2da3e4067c81260d4df7e73564c983fd0f050c6ce100e11

Observation 240469f4-8160-449b-8357-651204248943 · outbound

This paper cites Nature , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature , volume =

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.871152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.871152Z digest=sha256:70d1a42ede77a2c484426da20c859e50537d0218c9c9b133ab17e693eeecab61

Observation e0bde585-f68a-4e08-9288-f3a9a7e5aa17 · outbound

This paper cites Nature Medicine , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature Medicine , volume =

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.903895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.903895Z digest=sha256:4dba0d61a15766996a4fa821d251853d1eab98499a6ee33fd6379be4f3bf6504

Observation cd35f801-4760-4278-a5ea-17ce71bfde78 · outbound

This paper cites Nature Medicine , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Nature Medicine , volume =

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.933035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.933035Z digest=sha256:3bb77b27789abe4b9321f56b53d450d4377d78d59b0f3e994d60f1fa686d1fac

Observation 5c6e3de3-a465-40e3-8d67-dc441c1b39f8 · outbound

This paper cites JAMA , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents JAMA , volume =

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.965492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.965492Z digest=sha256:608025232e49ac4b87729fcddd4f6cd174fc70928a3d514e944d4f5d3b37af84

Observation aea814bd-a2cb-4763-af3d-ac003bdb41fa · outbound

This paper cites npj Digital Medicine , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents npj Digital Medicine , volume =

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:23.002220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:23.002220Z digest=sha256:30db7be889fb0ce9e846ecb52200f718f96e718783bf2d2c800096f5f6d7888d

Observation 9d9d1c7a-3ec4-4c60-9e32-e9356c140530 · outbound

This paper cites JAMA Network Open , volume =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents JAMA Network Open , volume =

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:23.045209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:23.045209Z digest=sha256:10aa21f5e2465c7f9054635fd2b6d1be4d316fdee775d5b7181e25fc5eccf6bb

Observation c6c3f53f-d6c6-40ff-8e58-005d6c899190 · outbound

This paper cites and Haber, Nick , booktitle =.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents and Haber, Nick , booktitle =

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:23.128253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:23.128253Z digest=sha256:72dbc4f89100ec8d1ead9fa1155471e84e74ef56334721f26e603dee049a067c

Observation d3ab3876-fa6a-44f1-8a0b-f4ccc7ee767c · outbound

This paper cites FHIR-AgentBench: Benchmarking.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents FHIR-AgentBench: Benchmarking

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:23.176529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:23.176529Z digest=sha256:87562a759bbd96bb1067e41c176e672ee902f5810231d2050931bd7d6deff247

Observation 8c1426ff-2c85-4654-b9c8-6801953e1ae9 · outbound

This paper cites Holistic Evaluation of Large Language Models for Medical Tasks with.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Holistic Evaluation of Large Language Models for Medical Tasks with

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:23.227686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:23.227686Z digest=sha256:1619a261bfe75d221801389230d011af3c83286db82cbf5a88677e8463f27cad

Observation fc251de3-d568-45e1-9bb0-3ab4d26b4773 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents gpt-oss-120b & gpt-oss-20b Model Card

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:23.281512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:23.281512Z digest=sha256:34183da64ef986f3cc432818eb51ab5347ed6249bf6f6ca281259a8cbb5549a3

Pith citing papers

No inbound Pith citation observations are available.