Pith. sign in

Paper Citation Record · LEDGER

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2606.01961.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01961 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T14:38:40.017263Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T06:30:16.612345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact31
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 916ad50a-591d-46ad-a0b8-aa8f0c293545 · outbound

This paper cites AAPM low-dose CT grand challenge (LDCT-SimNICT).https://www.aapm.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models AAPM low-dose CT grand challenge (LDCT-SimNICT).https://www.aapm

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:f233d1788d27838dfdd50a77b8d73ba4d830081da4f572e72fbadbd35e38eefd

Observation 20847565-e594-49a4-aabe-a05fffa844f7 · outbound

This paper cites Claude Opus 4.6 system card.https://www-cdn.anthropic.com/ 14e4fb01875d2a69f646fa5e574dea2b1c0ff7b5.pdf, February 2026.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Claude Opus 4.6 system card.https://www-cdn.anthropic.com/ 14e4fb01875d2a69f646fa5e574dea2b1c0ff7b5.pdf, February 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:0b236d188db79feb9915ec6383bd0340fe1eb2cd57115f92a5ee697d20cbd88b

Observation 3b9bf7c7-03a7-4b4a-9239-31771f03a560 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Lawrence Zitnick, and Devi Parikh

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:287e6a076856c315d2dc5a53f4efd363b65655600595775f327f0df48c5b581c

Observation d30e4737-c523-4b59-9660-980866c81c84 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.112444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:ac8cc4fb9a4f825d2ca426345ede723099d86ddaa89f3c9e286d4d69329c4040

Observation eb6849ab-a01a-4a43-9944-c909bae1c906 · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models METEOR: An automatic metric for MT evaluation with improved correlation with human judgments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:3d48e3360eb3c987f1148c9a7e424369793946d354b1f3e65355b15eba11d95e

Observation 09695644-9950-4a37-a589-a75502a21ca0 · outbound

This paper cites MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.116582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:e05167fb598b77cda65f6f8c8672a0238b41ff08323b5734353ecb2d85216806

Observation a2414a59-e6ae-4ca8-8ef6-58b7e0e1479c · outbound

This paper cites autonomy.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models autonomy

Reference 7

Resolution
metadata mismatch
doi, observed 2026-06-28T14:42:17.751463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:7f2239527e881fe4985226e7a10fd4dd11611d7b1e84044c27e77470ad578809

Observation cae89cbf-3895-4f36-925e-5a9d1b2734de · outbound

This paper cites HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.087558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:f7b77602729be8b237aa638a26f0e7ac040231bf8be3ff11d1b33214a3fbbce9

Observation 5c715ff7-8de7-4752-a93c-3885b397c829 · outbound

This paper cites PANTHER challenge: Public training dataset, 2025.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PANTHER challenge: Public training dataset, 2025

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:b0ab2f878085422e7c6fdb1ed9f8f8bc723ffd0426fdc9c966a9019de9214ccd

Observation 33ebc104-0da3-4908-9ce3-9a374e924c4a · outbound

This paper cites PanTS: Pancreatic tumor segmentation.https://huggingface.co/ datasets/BodyMaps/PanTSMini, 2024.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PanTS: Pancreatic tumor segmentation.https://huggingface.co/ datasets/BodyMaps/PanTSMini, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:7e9e2e5c5df023c3c7d15dc20be396efeb8c75942bf160e304eab778c43525ca

Observation 263ca44b-e83b-4018-970d-2b36f5261bef · outbound

This paper cites End-to-end object detection with transformers.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models End-to-end object detection with transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:1d9674bf96acce0f33bb3a05f7d28b8f8c5430bf5fe65197370299cd44468302

Observation 6fe69d4c-de91-49ea-96f7-fd0fa47fc359 · outbound

This paper cites CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.101309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:276fc4101869a76266312d8481016f18e8271f475811f6491c0832c2df7c7985

Observation d6843f89-c972-4670-aba2-0cf7307ec310 · outbound

This paper cites MLE-bench: Evaluating machine learning agents on machine learning engineering,.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MLE-bench: Evaluating machine learning agents on machine learning engineering,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:0a2841415102253343e689e948e739d9013a77b539a9e30ef15193d980984da4

Observation ac48592d-2c15-40fc-9674-b48ffeda9bf4 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.108710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:a79d14571cb80a22987aa5270bf429633c15e98629c3e23bc07c94d146431170

Observation 27f64636-a80c-4a41-be60-a781e87ae554 · outbound

This paper cites IEEE Transactions on Medi- cal Imaging36(8), 1597–1606 (Aug 2017).https://doi.org/10.1109/TMI.2017.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models IEEE Transactions on Medi- cal Imaging36(8), 1597–1606 (Aug 2017).https://doi.org/10.1109/TMI.2017

Reference 15

Resolution
metadata mismatch
doi, observed 2026-06-28T14:42:17.787534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:868343f7e0bfb0362cf6d1801c18428c75a51d07e34d73532cddda644e69aa53

Observation e1cef2d1-9781-48ea-a470-589e4d49cdd0 · outbound

This paper cites Generating radiology reports via memory-driven transformer.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Generating radiology reports via memory-driven transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:bf4cb3ac871751a71f02f44685e909d1184d30ac424632c8b3cbf8f7f0e1a8e2

Observation 0992bc6f-2349-485d-a05b-6de26f522c41 · outbound

This paper cites Demner-Fushman, M.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Demner-Fushman, M

Reference 17

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.771872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:a8521bb1124fd7e9fe6f38dc00e6c8178db753597b77a25d22478f971eef024e

Observation eae58070-a2cd-4bb4-bdf0-6979eb0d2301 · outbound

This paper cites Measures of the Amount of Ecologic Association Between Species.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Measures of the Amount of Ecologic Association Between Species

Reference 18

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.781892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:b1007059c36d3aee725e7b1394a58514f283431c4df65903024ca75de35fc26d

Observation af1e7e51-984e-4589-bead-e10f5d7fba14 · outbound

This paper cites Available: https://doi.org/10.1007/s11263-009-0275-4.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Available: https://doi.org/10.1007/s11263-009-0275-4

Reference 19

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.785816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:aae467049b0c7dbae962d6198b9d4ad37b2721390d677366f6d6697c54568163

Observation 4817b442-f21e-4837-aef3-ff6c2fb4512a · outbound

This paper cites FeTA challenge: Fetal brain tissue annotation and segmentation.https: //fetachallenge.github.io/, 2021.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models FeTA challenge: Fetal brain tissue annotation and segmentation.https: //fetachallenge.github.io/, 2021

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:f12bfacc3a0703078ef14ffbce883e0dc32e9b85c9b279475ac11f502919aecf

Observation 14d13a2c-375e-404e-a19f-0c3ccd88cd3e · outbound

This paper cites Camyla: Scaling Autonomous Research in Medical Image Segmentation.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Camyla: Scaling Autonomous Research in Medical Image Segmentation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.140386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:f8cd53260657a30d3036dbc54bf2c5172d57143a17050f8ae03340a78edb778b

Observation 09688aa8-68e0-45d8-ae06-6fc596454ba8 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models GLM-5: from Vibe Coding to Agentic Engineering

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.148476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:9a0e3bcaede28b60deb244e59740166d76802ae59807ad47ebd789e3e0b1235e

Observation a7825ae4-d91b-4005-a7ce-8bd4147c4765 · outbound

This paper cites Gemini 3.1 Pro model card.https://deepmind.google/ models/model-cards/gemini-3-1-pro/, February 2026.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Gemini 3.1 Pro model card.https://deepmind.google/ models/model-cards/gemini-3-1-pro/, February 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:59bc252430a9c624c8788c5454e809ae58ba6629e037be87171b5761e9792692

Observation 5bc04c3f-71e9-42bc-af5a-c36a51a954b2 · outbound

This paper cites DENTEX: Dental enumeration and diagnosis on panoramic x-rays.https://huggingface.co/datasets/ibrahimhamamci/DENTEX, 2023.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models DENTEX: Dental enumeration and diagnosis on panoramic x-rays.https://huggingface.co/datasets/ibrahimhamamci/DENTEX, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:9fde75f91f85a72df2677c8cc98e42f03bc446e6896045613531cfc25013efda

Observation fa211428-2eeb-4057-8a99-65e59c854822 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.157019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:5a1be62b070178803a392713dc7fbd4cd2f7e356aeb208659682a6d077e21ee6

Observation 686951db-7913-4476-9e5b-5d9c5e4320ba · outbound

This paper cites KiTS19: Kidney tumor segmentation challenge.https://kits19.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models KiTS19: Kidney tumor segmentation challenge.https://kits19

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:bfb5d5774dd775f895a0c3caa1896e8a4d599fe8aca26fc429db951e10ab4645

Observation 8d0573de-4d3e-48f0-8133-5888ffc15a49 · outbound

This paper cites HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.128547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:5f31edcba64b74566e6f194ddd11ef041d8a694ff90b1ec45aafec79ea6fd2e2

Observation 8e28c5cc-fafb-4984-abe5-13cad87dd011 · outbound

This paper cites MLAgentBench: Evaluating lan- guage agents on machine learning experimentation.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MLAgentBench: Evaluating lan- guage agents on machine learning experimentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:d02026190f99e2a08866400eee377cfed279b4386a66e4f4c962d19949917243

Observation 0f7f84d9-bf11-4090-9d38-d38adb202923 · outbound

This paper cites nnU-Net: A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models nnU-Net: A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation

Reference 29

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.775042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:2f1b894b7f9008b263cc428b69f98816557707740e6d57f68dc35143bc6aa318

Observation 4a8cf1d6-b39c-4c2f-99a7-3f947a85aa50 · outbound

This paper cites Q., Nguyen Duong , D., Bui, T., Chambon, P., Lungren, M., Ng, A., Langlotz, C., and Rajpurkar, P.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Q., Nguyen Duong , D., Bui, T., Chambon, P., Lungren, M., Ng, A., Langlotz, C., and Rajpurkar, P

Reference 30

Resolution
metadata mismatch
doi, observed 2026-06-28T14:42:17.759009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:d0f908c7b37b51362e06294e4c063c163def563ac8e7ba62696bf2a33a2d53ea

Observation 4f0dfe32-d5f1-4b67-b079-8ce63aaa6fc0 · outbound

This paper cites Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark

Reference 31

Resolution
metadata mismatch
doi, observed 2026-06-28T14:42:17.759497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:d862cae400d1e2c86570e9d5b51f6b23229a122851374538a769ae4ebfc9d96b

Observation 5bb0b44e-83f6-4333-80b3-5cdc13f30f2c · outbound

This paper cites MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.084048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:13a53e99f26177912886dd88eedff57a6ab7cc25c52913309c7f7caa08b442f7

Observation dc4ee18a-8f2e-4710-858c-23fe92c505c5 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:7761ea3569230987c0a36ff10b882c7987d70c430e4df7e3d6bb0f0e668e30d0

Observation fd9518b7-5c28-4fe0-9942-fe06e9d3a0a6 · outbound

This paper cites What disease does this patient have? A large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models What disease does this patient have? A large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421

Reference 34

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.757611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:5eb2e8f49b4729131623c195473c464d181a7f87022df9afe57f9b0a6f4c6a7c

Observation debc29f6-f8ab-4e62-a39e-55f42b366d8f · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 35

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.794977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:da5d87a540b03fe85d77a11f72755f18a6c791acaf92cbe39b0d58ff5c6c3426

Observation d295afe1-aa5b-4b02-af5b-5d49c0ec636a · outbound

This paper cites Johnson et al.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Johnson et al

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:be1d2c90dd2b7ae829390c477879026ec0afa087aa7d859694e6ba951391df41

Observation 14516445-2559-4eee-b310-f9d479f3e7e2 · outbound

This paper cites BCCD: Blood cell count and detection dataset.https://huggingface.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models BCCD: Blood cell count and detection dataset.https://huggingface

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:ea621e07d95e81c8b4f67982ef37d7dc3e984589151bf66269490efe71fc4d7c

Observation 95b142f1-7c7c-47b7-81d4-bfeedf8b185c · outbound

This paper cites URLhttps://www.nature.com/articles/sdata2018251.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models URLhttps://www.nature.com/articles/sdata2018251

Reference 38

Resolution
metadata mismatch
doi, observed 2026-06-28T14:42:17.779613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:51b45c80a2a995ed160b3669c52aeb9fe92a31e9e527373c027d9beb6a839dfa

Observation 96c5dbe3-c811-4997-bc51-0a4153af0235 · outbound

This paper cites Fhir-agentbench: Benchmarking llm agents for realistic interoperable ehr question answering.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Fhir-agentbench: Benchmarking llm agents for realistic interoperable ehr question answering

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.165966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:c8457166409fda94e37b24849aed307266d133781043e0a256fe681697b79168

Observation 0678e7a6-5cc3-4662-9312-efb1cc2f8ee7 · outbound

This paper cites Agen- tehr: Advancing autonomous clinical decision-making via retrospective summarization.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Agen- tehr: Advancing autonomous clinical decision-making via retrospective summarization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.161668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:b24eefea55515eacb291225f0fed63d498abc4e968efec66531ec98ca7939dbf

Observation fa3c2fa9-f376-443e-a6dc-d47b38dd8d10 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models ROUGE: A package for automatic evaluation of summaries

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:e1dc7d03a4b2e783f4efba50c968a147d4c5bd1dc0286ff037f9156ec7fa44b0

Observation bcc2cf39-8f1b-4ae4-9276-7523e367290f · outbound

This paper cites Artificial Intelligence in Medicine143, 102611 (2023).

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Artificial Intelligence in Medicine143, 102611 (2023)

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:42:17.770721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:777e8bb06af26d236886aa46806520aaa32f89412d48e731ccd572258b103ee9

Observation 0655a30a-a525-403a-9c3b-e189e3648aa1 · outbound

This paper cites van der Laak, Bram van Ginneken, and Clara I.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models van der Laak, Bram van Ginneken, and Clara I

Reference 43

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.777323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:0827cfdbf1971e79690a115799fff19fac23382d217d63885172760f8834db6d

Observation 048f4a82-b5da-444b-bad5-e7c41842c503 · outbound

This paper cites SLAKE: A semantically-labeled knowledge-enhanced dataset for medical visual question answering.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models SLAKE: A semantically-labeled knowledge-enhanced dataset for medical visual question answering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:fe886fe3ca5e710d6ecb4a2f262cf2eb65e024cd745d059ad89a370f065fa54b

Observation e17750a0-ac70-4e49-ad05-c689ee71065a · outbound

This paper cites URL https://arxiv.org/abs/2102.09542.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models URL https://arxiv.org/abs/2102.09542

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T14:42:17.777443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:d62701a1180d39adca9de653ec1cf0368d6f2969bb4770995d5222614cf14426

Observation a5c2a0a7-81d2-4edc-81a4-6ab1bcb48d34 · outbound

This paper cites PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.132384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:9fd10725246e5304eb278e9a069422b7d2701e997bea79a89a85ac587f23ef5b

Observation af86cd11-889c-4d79-9bb2-e5dd51d07dd3 · outbound

This paper cites Agentbench: Evaluating LLMs as agents.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Agentbench: Evaluating LLMs as agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:c04f81f8c4c2f065ea828fa54c4da3e52f1790d6da5736f4dff10e0d3e852577

Observation fa30672c-3eca-4589-9219-8bd1f94a5804 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.090745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:39bfb2a06a7198d65bf0f04e9e4d7e11218ca11e84608f288d56055055469a16

Observation 6de4b676-0d25-434b-804d-9b7c2be9a0eb · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.094083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:b40c23a09d55d420fcfd8dcd4fddb3babd0ce4aff1a7ee15e52191096eba308b

Observation 2221835f-028b-49fb-aa83-7f9ddd932217 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.097857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:24eb7325ff92cc3bd16dcc9441adca0b1358bfbff953f517bea86f596908797e

Observation c420184c-a22a-4de6-8317-a90931d0e54e · outbound

This paper cites GAIA: a benchmark for general AI assistants.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models GAIA: a benchmark for general AI assistants

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:ef0ae30596cb50fe5b075761ea086c305fb99f7bad29f1695a646de07c72d8f0

Observation 5ad15700-7806-4109-8db1-04cbc3e90238 · outbound

This paper cites The MiniMax-M2 series: Mini activations unleashing max real-world intelligence,.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models The MiniMax-M2 series: Mini activations unleashing max real-world intelligence,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:906f936575c12ef26c9f0406a628e5bc450db91af81b756bf2ee0e4be431903e

Observation 6de01c2b-910e-475c-8523-985dab72257b · outbound

This paper cites The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.105338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:cfc119e13d3a0d04f85ce4f8016965c5426a26a73a357f100291e96941bf44b4

Observation 4e579121-6056-48b5-a863-c1081be727e7 · outbound

This paper cites GRAZPEDWRI-DX: A pediatric wrist radiograph dataset.https:// figshare.com/articles/dataset/GRAZPEDWRI-DX/14825193, 2022.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models GRAZPEDWRI-DX: A pediatric wrist radiograph dataset.https:// figshare.com/articles/dataset/GRAZPEDWRI-DX/14825193, 2022

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.124512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:ac0e5bf6bc95163c69cfc2fe70b5c6f01a08bc15dad1e22b860e09a01bf3fbcc

Observation d00ab522-7069-4605-b6c4-1e72fcec9d72 · outbound

This paper cites MLGym: A new framework and benchmark for advancing AI research agents.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MLGym: A new framework and benchmark for advancing AI research agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:9e54e383184cb0d33a1995955fff864a99c70e1b1c369fe560385ba5e681c546

Observation 0f446f5a-7692-4f06-a8c7-b7233d2fe062 · outbound

This paper cites Nguyen et al.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Nguyen et al

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:cf0ddbec4232b4027c657aaa8910f031c648ff97478c6a8eeab012540e020ce2

Observation a1190a0f-7bf4-4208-86aa-9906fa448477 · outbound

This paper cites GPT-5.4 Thinking system card.https://deploymentsafety.openai.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models GPT-5.4 Thinking system card.https://deploymentsafety.openai

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:32e8094f33ec07c385d70d8fe0154807e1c7561c75abc61113bb0016394760d5

Observation a8614ddf-d82d-4c75-b192-4cd842a43129 · outbound

This paper cites MedMCQA: A large- scale multi-subject multi-choice dataset for medical domain question answering.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedMCQA: A large- scale multi-subject multi-choice dataset for medical domain question answering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:1dc4b66c43533cd12f9b40737dbbdecb719f7a37d19fd5e58cefe1bba07f1ad2

Observation c9d6593a-a47c-42da-9395-21615de50368 · outbound

This paper cites doi:10.3115/1073083.1073135 , editor =.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models doi:10.3115/1073083.1073135 , editor =

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T14:42:17.787208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:7e6697690d4895d18acade865a9623387306e55f87e82d9dcf73728f026f2c7b

Observation 2493b696-f497-44cb-920a-088e51287699 · outbound

This paper cites Qwen3.5: Towards native multimodal agents.https://qwen.ai/blog? id=qwen3.5, February 2026.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Qwen3.5: Towards native multimodal agents.https://qwen.ai/blog? id=qwen3.5, February 2026

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:b35e5b1234caad6bf5babcd7a3a3118b594e92bd2dc1b9f1b1acef70f6e08142

Observation 10655f46-a0a7-4dd6-b681-2021c5ed712d · outbound

This paper cites AeroPath: Airway segmentation dataset.https://github.com/ raidionics/AeroPath, 2023.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models AeroPath: Airway segmentation dataset.https://github.com/ raidionics/AeroPath, 2023

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:4625da47d7adf6219f22e937549c34c5e8c72911d4492b0859ee378fc6c0e96a

Observation 44d77347-b30d-4b3c-99c9-a2c40e0e1ecc · outbound

This paper cites 2016, in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 779–788, doi: 10.1109/CVPR.2016.91.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models 2016, in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 779–788, doi: 10.1109/CVPR.2016.91

Reference 62

Resolution
metadata mismatch
doi, observed 2026-06-28T14:42:17.792926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:fe68941f6c0e557a927a415d003eef053b5376f62bd2f656c21f135a7c230c2e

Observation 2f6e6479-e96c-4d59-a6e9-9353d5a8ba7e · outbound

This paper cites Faster R-CNN: Towards real-time object detection with region proposal networks.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Faster R-CNN: Towards real-time object detection with region proposal networks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:e070831dd4dccb4e88f0403807aff8079d490fd882e8635c2c1ea88344f141de

Observation 502e2985-f5d7-4ed1-b2d0-e4300536f05d · outbound

This paper cites U-Net: Convolutional networks for biomedical image segmentation.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models U-Net: Convolutional networks for biomedical image segmentation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:d48a4972bd7206476af6ca87fddc2216f383af3551061d3888a886e1d6cc43c6

Observation 3d584db4-dd51-498c-8ad2-204d81e17450 · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.136145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:87ea42f9d0f0fac1f989d004c2bf2e82e3531fe8d6b59de0f9f98c91703670b3

Observation adffbb9c-e4c7-4770-8d2d-637f9b4d5274 · outbound

This paper cites Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:6bcd76f1e9253cbafba170f79067bc0bd58d700b016cb0a7c5cd9fb97686e52c

Observation ba478fe3-80d4-4c2f-98f5-56617b76c36b · outbound

This paper cites Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, and Arvind Narayanan.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, and Arvind Narayanan

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:000f5d9a53231adf1a658ede6f37484c716a9ff2d1d565a11a9658e3055fc079

Observation 63f1eba0-6a5d-4c9d-9af3-1a65f843f1bb · outbound

This paper cites Large Language Models Encode Clinical Knowledge.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Large Language Models Encode Clinical Knowledge

Reference 68

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.789287Z

Source-reported events for the cited work

correction dated 2023-07-27. Source: crossref record 10.1038/s41586-023-06455-0->10.1038/s41586-023-06291-2:correction, observed 2026-07-11T03:08:19.417011+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:6d04edda04402c186011ce2b7a33f25b5d6eaf894a02157245cc733add37a797

Observation 7c25706f-4bc9-4912-bb3c-7f3ddaf77de6 · outbound

This paper cites PaperBench: Evaluating AI’s ability to replicate AI research,.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PaperBench: Evaluating AI’s ability to replicate AI research,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:e12f588adae7eb21f492dea426e37bcda2550e489c58d56882ef4fe8c48ad924

Observation d0cdbdbb-f07d-48ae-b274-11c0b1372d11 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.152849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:b8682ed6f0ae6216555d07c0c0aa1eaa10a0be4b29db5e2bfee1c81abbd3ef2b

Observation 4b232d4e-ed5c-43c4-bdb2-0c6aabe24215 · outbound

This paper cites PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.179269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:3a1b7b24aac5eb980073ad191a3812ef43c69ec302b29d36a3403d9036e985aa

Observation f30b508f-5116-4990-a5a4-44b5e17c7b49 · outbound

This paper cites MedXpertQA-MM: Multimodal expert medical question answering.https: //huggingface.co/datasets/TsinghuaC3I/MedXpertQA, 2024.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedXpertQA-MM: Multimodal expert medical question answering.https: //huggingface.co/datasets/TsinghuaC3I/MedXpertQA, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:0fe0bf1d3c00a1510da5d5b991697a9cc2c59ab3d6dcd1fbbff8105706e2808c

Observation aef18d50-9cf6-4af0-bfb7-ffefed375d28 · outbound

This paper cites Bovik, Hamid R.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Bovik, Hamid R

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:42:17.774755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:db33ded3bd9adbc30deabefb2486446ca9a4ad84d5e2a8586f8261be5e5c3fe4

Observation 7e723958-2c50-42dc-97fa-2370b1ae0c16 · outbound

This paper cites TotalSegmentator: Robust segmentation of 104 anatomical structures in CT.https://github.com/wasserth/TotalSegmentator, 2023.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models TotalSegmentator: Robust segmentation of 104 anatomical structures in CT.https://github.com/wasserth/TotalSegmentator, 2023

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:8d43bdf46a9c6714460bfc0581232631d03d125fc7b410e16c13c2aae2f1b4dd

Observation d25a1708-869f-4877-94de-43d5f06fc7b7 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.120068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:e77e5d25bbb07f762f236b1a6b3181cc6dd5e485c5faf250e2f67644219385d0

Observation fe0aff11-89f6-4512-8fc3-45181eb90bef · outbound

This paper cites Medmnist v2 - a large-scale lightweight benchmark for 2d and 3d biomedical image classification.Scientific Data, 10(1), January 2023.doi:10.1038/s41597-022-01721-8.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Medmnist v2 - a large-scale lightweight benchmark for 2d and 3d biomedical image classification.Scientific Data, 10(1), January 2023.doi:10.1038/s41597-022-01721-8

Reference 76

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.791073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:e030da35915c3e154c21c97ef57ec0cb38284484ef3647df74b14d0c72109067

Observation e4d8356c-9d47-449f-9de3-f7a8412bd025 · outbound

This paper cites AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.144889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:dd035455ea36b1eb5e48aaa635333b3bfa2549885aacbb4804324f96e3a2b906

Observation eb3e034f-7e32-4625-a8f4-680917d30ebc · outbound

This paper cites Medframeqa: A multi-image medical vqa benchmark for clinical reasoning.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Medframeqa: A multi-image medical vqa benchmark for clinical reasoning

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.169959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:851776308959be225a1070253d996761cc37b8f139c0df63a26bb94b5816dd19

Observation 62ecbb48-0675-42cc-8146-fbbea6d31200 · outbound

This paper cites fastMRI: An Open Dataset and Benchmarks for Accelerated MRI.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models fastMRI: An Open Dataset and Benchmarks for Accelerated MRI

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.173834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:ada92d299a2e2b911e38632803ac0ee7955e4f56618b6acb1311c3a496b641ea

Observation c5f3758a-fe32-4f43-b745-a231f9d1e3d3 · outbound

This paper cites open-source.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models open-source

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-28T14:38:40.017263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:872f3a4bcf5c3bdd8c2df9a270190b289936b0fcc0a20ed7dcbfd9bc678bee1f

Pith citing papers

Observation c9a37d92-a99c-4543-987f-6a83fcddc3fe · inbound

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy cites this paper.

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

Reference 149

Resolution
unresolved
no resolver link, observed 2026-07-14T06:30:16.612345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:30:16.612345Z digest=sha256:c7ab4b3e6b36152cc0dd9de3c9c76cc0310ac8b165a7c215db1df7d45330f208