Pith. sign in

Paper Citation Record · LEDGER

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

As of 23 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 88 inbound Pith citation observations for arXiv:2501.18362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18362 v3

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T14:26:38.517787Z

measured 118 of 118 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 88 of 88 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:20:20.277083Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1681f3dd-2327-4d38-a78c-18accf2e4868 · outbound

This paper cites org/CorpusID:268232499.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding org/CorpusID:268232499

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.600620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:23f6c26791664baf3d6f32cc7dadf70578d1ba5dcd71308a0c14bf5497c3dfc8

Observation bb8ecff3-5e41-4fc1-9a04-78ece23abc7d · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:26:38.534771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:da897ab4b6fec27472361c1884424311f3e181816d4f7645ea530327b6c6ae23

Observation 144bd2dc-e05d-46d3-b2fb-19978cdddf0d · outbound

This paper cites The reduction was successful, as indicated by follow-up x-rays.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding The reduction was successful, as indicated by follow-up x-rays

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.537822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:3d8c7f52c47cb7a9fa855dc89e2af135c1ee2d3eefc4b6e951829613e43d046d

Observation a24f93cd-838b-4b83-a028-bda24f51f104 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.540385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:169454f55c90c3bea7f10c4267604a3d9382320444aacbef2df6a3e4e9734d71

Observation a394bcdb-a3e7-44e2-a79c-3d95aba52d98 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.542719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:4515ae7fef69cf0f021c603bb7311c37a9cfbb89e96119dc88feff156b048f0a

Observation 494336fe-6a00-456f-9a4c-b004e1310925 · outbound

This paper cites If there is a suspicion of nerve injury, such as the axillary nerve in this case, an EMG would be helpful in confirming nerve dysfunction or damage.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding If there is a suspicion of nerve injury, such as the axillary nerve in this case, an EMG would be helpful in confirming nerve dysfunction or damage

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.545499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:e7b6cd795b79dcd85ef61fa917d85fb078c20f4bafa8e4cc1c45da5413f5ec27

Observation 13fc9d3c-116e-406d-888b-99931e708187 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.547874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:3762688fe0646eaefdd16fc8eadd592ed385e445e0e96e012f9ae057b50d5d0f

Observation 7b2860e3-3c64-417b-8ed4-e780e609a92d · outbound

This paper cites There are no distinct P waves visible before each QRS complex; instead, there is a disorganized electrical activity, which is typical of fibrillatory waves.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding There are no distinct P waves visible before each QRS complex; instead, there is a disorganized electrical activity, which is typical of fibrillatory waves

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.550655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:09c7749883a931237ac92b7f561fecfe02f828c3286f62448963ad9b05560368

Observation 3facc0df-cced-48ee-a493-8a4a4dbfe47b · outbound

This paper cites saw-tooth.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding saw-tooth

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.552858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:b1acb786e8f210ffca6ffe40320a560f6319a2a93c6a7bec0fe331f6a87d073e

Observation c1cf6fc9-dec9-49c2-91df-4edcf2cf25c9 · outbound

This paper cites saw-tooth.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding saw-tooth

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.555139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:0a1e173eeda6d720d556e3ce50e835a9e6b34c39ea9dcea7357e8a1506970775

Observation e134e893-1be6-4dbe-b30a-5c5e0436373d · outbound

This paper cites Analysis: The above response fails to fully grasp the question’s implications in several ways.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Analysis: The above response fails to fully grasp the question’s implications in several ways

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.557551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:c26a90f8438e8fa53f0e325d3262270df857667839eed6f246c7af76b4854f28

Observation 91e08cf8-73d3-4bf2-8f08-c83aa91de78b · outbound

This paper cites Aging is associated with alterations in sleep architecture, which may lead to reductions in total sleep duration, slow-wave (deep) sleep, and sleep efficiency.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Aging is associated with alterations in sleep architecture, which may lead to reductions in total sleep duration, slow-wave (deep) sleep, and sleep efficiency

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.560183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:e40301a3d4a96a1dfa83de4cdda8c4e875aea767708b8efba9aac5679640cefd

Observation 44bc7d9a-0b50-45ea-843f-7d49e3d7fae5 · outbound

This paper cites His sleep appears peaceful, and he experiences no disruptive symptoms such as snoring or awakenings.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding His sleep appears peaceful, and he experiences no disruptive symptoms such as snoring or awakenings

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.562466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:84fa927d4a292c85faf19528e8ccf3a895cbd26ee7b7aad71aa5f578106ba840

Observation 03263f01-f0ef-406f-8075-78a8ca491727 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.564519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:a987294ae21522fdacb60722ae1569fe355946c7f411d5dfc513d3e5d78ff66a

Observation 048946b8-4d64-4023-8398-86d91d47169a · outbound

This paper cites This can result in less restorative sleep and a less refreshed feeling upon waking.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding This can result in less restorative sleep and a less refreshed feeling upon waking

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.566963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:1602c1ed6c3d2929508f304d18f1de3b15002eb4b1ad810f1a236c7b02711223

Observation af6f1089-91f7-4814-98ef-0fac92b684fa · outbound

This paper cites acute on chronic.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding acute on chronic

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.569119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:12d27b41fca9d11a65114b5a6e648628892a2a2b20a6282d7899dd3c9c9580c4

Observation 0b7fce81-589c-4ef7-98af-3934cf197471 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.571248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:0bc6ced1e7b0788b473cd8d8e0fc3122ab6e4e4145e27383485e66f72875fa7f

Observation 63adbc35-7d0f-4985-af8d-914ce78c9627 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.573509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:6c15a81234708b487781c6a11fac29d99003b936c77c791c120637eb75c10c88

Observation 1ad086e1-5e90-4a2e-a789-8a6433785b81 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.575528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:c41a456c5a3603bc2a7dc9e20d0ffd5a57a4750f6f4b268484581e2e616dcb06

Observation 7e3fced3-a82e-42d0-b134-291150473ca8 · outbound

This paper cites Rigorously ensure clarity and avoid ambiguity.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Rigorously ensure clarity and avoid ambiguity

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.577739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:6f0f2774e81651bf4203df3e2b89f412b053eefd40176172aeeba06946e1aa2e

Observation 7856cec6-a23f-4446-bc6d-471b0a40c916 · outbound

This paper cites Do not change, add, or delete any factual information.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Do not change, add, or delete any factual information

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.579861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:3a5d7b9bd3852d416d32e800077e668dd651045cafb3dff3d9f198920df507e3

Observation a5e981b8-a5e7-413e-aa37-c69650314810 · outbound

This paper cites Pay special attention to keep any tabular data in completely the same format as the original.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Pay special attention to keep any tabular data in completely the same format as the original

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.582069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:36bc1ba96f66045a02bcc476f7e7f7a4d3f4470930ef73c1849c83787e63c0ec

Observation 04249930-be57-43ac-90f2-a29cc6aeb007 · outbound

This paper cites Answer Choices: (A) [Option A] (B) [Option B].

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Answer Choices: (A) [Option A] (B) [Option B]

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.584097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:2a0b1058b86566000eab2a7ed028afc054296ccec1bc256edff429b47f365672

Observation c0c44156-0ac3-4f5f-b984-0c339de9d2d4 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.587307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:0dba618f58d2411dcfbd38371e912b2ef41c5947ed7ed6bb1a43740ed14b46ed

Observation dc22a6d1-3331-4cd4-a04c-14de87b7b8d7 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.589630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:8120c20ca799106a0fc9f1c2cc119d50a426603c0eb6799f0b30563a318f8052

Observation 8eaa5b77-a94d-4497-ad15-9f09edad1f54 · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.591786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:7b6f4140563f2d201d1aeb115d52909248bbeb0c41b7b589a28627ed5122ed70

Observation 33df4d60-d8a3-49d9-8873-34883d712933 · outbound

This paper cites They should be clear, concise, and professionally worded.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding They should be clear, concise, and professionally worded

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.594021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:ea27af6f8b032764e3406c0fc27630e2c26b733e53e08ecfd64e005806bf6ec0

Observation 4ce60eb3-9ac6-41ff-b88f-cfd712f2b10c · outbound

This paper cites an unresolved cited work.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-16T14:26:38.596171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:0d838fc045b4d9e4549e82a5ce4b4e12a3de07d1c494963c16cc68ec31ba3249

Observation e5cd1a33-ff7a-433f-8547-1c28770ffc76 · outbound

This paper cites Avoid options that are overtly illogical or unsupported.,→.

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Avoid options that are overtly illogical or unsupported.,→

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.598395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:1ec30e5520ddd9a88ac10e66fe422bff83b3ac278d8c66805c93a64d8233234e

Observation 135efca3-6d51-4b18-bfe8-1f01ca87ed36 · outbound

This paper cites Answer Choices: (A) [Option A] (B) [Option B].

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Answer Choices: (A) [Option A] (B) [Option B]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:26:38.602926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:38.517787Z digest=sha256:6bcb2fa9a25708f311a40791fff70341ed68305127c894894b84c6429ee06348

Pith citing papers

Observation dce3a636-1c70-416a-be7a-0d7582ac3b3c · inbound

ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification cites this paper.

ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:20:20.277083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:20:20.277083Z digest=sha256:f543c770d25760b410bab1daa2eb295ab2b0237d8368653d856bcf417f949d3c

Observation 00c3fdb2-6f00-44be-bef9-dc76724c0ff6 · inbound

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science cites this paper.

BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:33:24.554504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:33:24.554504Z digest=sha256:74c59d2bb05f73a5d3ef92de31d49d1726cbfafe9e74e602de7b6631248459ce

Observation a6deb072-9fe2-4cc4-ba96-89e80df68f58 · inbound

Disentangling Reasoning and Knowledge in Medical Large Language Models cites this paper.

Disentangling Reasoning and Knowledge in Medical Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 50

Resolution
malformed identifier
no resolver link, observed 2026-08-15T20:57:56.132621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:57:56.132621Z digest=sha256:3b7fd812bbd344a24597944751edfc38f98bd9d8c4296ed4f928ab12e93962e1

Observation ee58effb-e067-41e3-bf66-4c5f270dcafe · inbound

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models cites this paper.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.004807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.004807Z digest=sha256:976cb6d740eff17492e37b055559d1fd8324153022679957099bc090af92430a

Observation f4b35547-716f-4666-acfb-e821b41c857e · inbound

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL cites this paper.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.423359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.423359Z digest=sha256:e1739363ae39824553390c03f994d9f7f907ac2450d9b466c54618a8ff555b96

Observation b7b2aba6-1284-48d0-ae5a-629b81a01a6b · inbound

TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification cites this paper.

TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:49.370355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:39:49.370355Z digest=sha256:6f1bbdb4acd45246fbb4d9c03bfbd387d0b6c75dc0a4ee895bc7ab6c1ead9e1b

Observation 247bc90c-215b-453a-ad87-185f4f606d53 · inbound

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning cites this paper.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.605024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.605024Z digest=sha256:1731d68bef4c06e739ed0a24071746c42e1f1a83ef985589f4ca1d9a70da0476

Observation 7f4aa9d2-33be-4bf6-8def-c628972511ad · inbound

Towards Large Reasoning Models for Agriculture cites this paper.

Towards Large Reasoning Models for Agriculture MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:42.988989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:42.988989Z digest=sha256:a4930ebe19d680f90876af1d91f4b5b03649fc570d28f14906022609db759f16

Observation a4d01c4d-8b62-4bf9-a0f4-2baa36503564 · inbound

Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making cites this paper.

Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:43.187164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:43.187164Z digest=sha256:8a0311616db38f1597e6c686d9603e8d3b08efa6c756bd5a9954551bfee05357

Observation 361fc017-c4d4-4484-a06f-d00e895cb5f3 · inbound

Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble cites this paper.

Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:57:58.191080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:57:58.191080Z digest=sha256:35d2946bbeaa0437fb780ac0087688d58c2473a0998054e7c4886d2fb92eba45

Observation f548da07-6d9c-438d-bc51-f2528ac72cc9 · inbound

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering cites this paper.

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:52.565185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:52.565185Z digest=sha256:58b01e8e1f953b35b0aefff81aeafdbf677d95e28b3164075c290519fe45a66d

Observation b373415e-d784-4a85-8e22-d02a0c3ab2c4 · inbound

Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation cites this paper.

Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:36.112254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:36.112254Z digest=sha256:2bee93a9d510342ee10eb6621c6eb54255edef9150297ac7ae9d4d4de6e7c0b5

Observation 5851fe81-99b1-4ede-8416-5cae63b5bf82 · inbound

DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models cites this paper.

DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:12.658854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:12.658854Z digest=sha256:4081ccde740bfa53a7b70ea333f2db023dd9e178cfc3be1f8464f73ff4d7b4ba

Observation b88cc0e5-d93c-4fef-8a9f-90ed663e0d5e · inbound

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains cites this paper.

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:34.809777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:34.809777Z digest=sha256:316b29bdcabfbb1eba19d524577d2c941629e6e6b6bd42564017d4619425f4e7

Observation b21cae6e-9d43-4d4a-81f2-714733e39145 · inbound

Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning cites this paper.

Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:15.260193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:59:15.260193Z digest=sha256:b457955752d4ad57a3391f036580cbc62f886f124af2768898a41f1c448a10ce

Observation 79acf541-ad73-46d9-9e9d-714dce4bdc6e · inbound

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs cites this paper.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:56.122012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:56.122012Z digest=sha256:5891b109029c63711083e0282d0d7323ba2a4bff1f41cf963ba43e704a1c2d4f

Observation 05f1c7c1-6992-4164-925c-2b7b3d5f5976 · inbound

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports cites this paper.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.721022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.721022Z digest=sha256:13854d8b2e16851ccac8c15d2c9e8ef28ee6645e70fa0aaa652d933489903496

Observation d22a01de-73f5-42e5-a047-00fe398d3a8f · inbound

MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation cites this paper.

MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 19

Resolution
malformed identifier
local_arxiv, observed 2026-05-21T23:10:44.659252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T23:10:38.936021Z digest=sha256:147d2b9c4b97acb531d2e114d6f09208c4825dfc9cc472edce118b6d4c1fc12e

Observation 0c6493ec-dbf9-4033-8a9e-b6e045fa6b1d · inbound

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming cites this paper.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:40.971770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:40.971770Z digest=sha256:17ebcd1f7feed113e4e11ae4ad819818247198552775b67b27a74aa2a52fddfb

Observation 8297b2ac-d6e7-40a3-8f9e-69e7730b7534 · inbound

Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach cites this paper.

Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.311823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:37:35.311823Z digest=sha256:d70b497b862ac6dc16215dc0bd028a95b3f7310cd074a326d939e7c9d98079ab

Observation 095c61c5-541d-402c-bfe5-81fd8a7f004f · inbound

HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways cites this paper.

HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:15:09.811969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:15:09.811969Z digest=sha256:936111d3bec3a189aa656128ac4a4e5eb0c7314d5828362a353aa71c8abe5014

Observation f8be9093-1622-40c3-9f40-0cdeebdd0e72 · inbound

Capabilities of GPT-5 on Multimodal Medical Reasoning cites this paper.

Capabilities of GPT-5 on Multimodal Medical Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:48.110957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:48.110957Z digest=sha256:5db1bc7782037f058f18ad2ffa9d31202fa18bfea1b79063b84fdda4a5e6e60d

Observation 85fbb256-4269-4d16-b2cd-7529460147fe · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:37.012685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:37.012685Z digest=sha256:138ffc7891d07673e38fb6d8c3b04f9361260864c2e369a3f1a67fa4deac726b

Observation 3c9f4076-f74a-45f2-8fea-44f78b386640 · inbound

Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks cites this paper.

Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:16:24.067374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T13:16:19.744864Z digest=sha256:ef8228859299f146c15b23567d9369dfdf310542b804f3f447002bd2adb52e8b

Observation 00f9f1d0-c640-4be0-9a55-37921d3b8350 · inbound

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification cites this paper.

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:53.456752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:53.456752Z digest=sha256:aefba6e1e3ba2fb9fe71d056ccf45a25ecb1b2810b0e7437ff749825f8a1f756

Observation 4c3a3772-61bb-422c-9f92-44370eb6e772 · inbound

MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution cites this paper.

MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:21:29.499231Z digest=sha256:3394e7ec74e27af17fb5b2d3d30034625a3dc953c0ee2954782eee672f9a4cf5

Observation d682ce68-1629-4254-b1eb-e48346c3f9d1 · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:d156b717a3efc7a119597d50f4cb79185c0d951d94d43bdcf1effba6c2b6acc6

Observation 9610ecc2-1b58-4e3c-940b-02d187915dc1 · inbound

Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI cites this paper.

Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T17:25:49.913986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:25:49.913986Z digest=sha256:7a01f34a60696dba3b1b1882ea270ad5c1c25e37c9cef44b1e17414f5ecf0b0f

Observation 8d58c0f0-32c1-481f-911f-b7d8b8dab98a · inbound

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence cites this paper.

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 68

Resolution
malformed identifier
no resolver link, observed 2026-08-02T17:19:40.131720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:19:40.131720Z digest=sha256:1a6de5830b131d6491baaa293e44d2dbbcd049f6b95d84939a8a68068e2b055e

Observation 692ea18e-394a-496b-8778-b3574823c5d8 · inbound

MedGemma 1.5 Technical Report cites this paper.

MedGemma 1.5 Technical Report MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:36:41.167684Z digest=sha256:a7362c93b8597206cebc086c3c0b3bca3d515c58b61d67d73a61473a8444fa41

Observation 7ae76ceb-4b9c-4968-bbd5-44ef72b8a87f · inbound

EXAONE 4.5 Technical Report cites this paper.

EXAONE 4.5 Technical Report MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:47:34.692414Z digest=sha256:75c8a25fc2e1cad922f8c431edb68442acc602cbe8769ccdb992f5592098b4e8

Observation b660a08e-c5e1-4234-a24f-06cbd6a39e56 · inbound

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering cites this paper.

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:03:14.408704Z digest=sha256:5aa1a38a1c5d9a421b2b28fed90933c4c5f0faf9e57aa8ed8165031c4325a008

Observation 45d2525e-cacb-4566-8379-fd6343f029f8 · inbound

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering cites this paper.

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T16:33:56.516696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:33:56.516696Z digest=sha256:1d1f506acfd812e6f1a47a1006e9ce4b95b8eda6d0a16478858ab8d27a25374b

Observation 13da444f-16bc-453a-8cb0-333e43f456fc · inbound

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging cites this paper.

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T12:47:27.551670Z digest=sha256:0c7bc0ee5dacbfdd522bb8106bb9721005785415f76fed4059bda484c6616863

Observation 68a546cd-ca2a-48e1-9666-add64b0b2e6c · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:050f27df10dd3718793d678370e76bb06afd87b640ad638af388060f33f09f65

Observation dfd5bf8c-66ef-4eae-b0f0-f248f94f6c7b · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:6713d2a9f6a47abbc96934b975a1db5071134ad50bb077320cea7a1572e53770

Observation 3ea0346c-9da3-4694-b66a-712a2985cd70 · inbound

Green Shielding: A User-Centric Approach Towards Trustworthy AI cites this paper.

Green Shielding: A User-Centric Approach Towards Trustworthy AI MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:43:54.896449Z digest=sha256:5f0a7510418c2a2fd3d5e3c4d654a3495bdf409f8b04417a034b9edfed7c06a5

Observation 50884898-2721-41f2-82d6-5828cf462e2f · inbound

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution cites this paper.

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T14:02:54.395566Z digest=sha256:35e47f8cf2aaaf73743c840e6a6e8e19e53c56a326f8d6f4cfb42bfb630c412b

Observation 9298d93c-375f-449f-abef-a1fdf5de2c7d · inbound

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution cites this paper.

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T00:53:53.394460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T00:50:15.010488Z digest=sha256:e7f5fdde50eff2c48ecddd9470809f0cc783421c20d484b2d2f512691a2f6bce

Observation 2ba01998-58db-4f36-a07f-b138d6a680a6 · inbound

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution cites this paper.

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 161

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:15:43.776509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T09:10:28.244690Z digest=sha256:4dce3451b7b3bfb77037770985a1b52e51a2a7f16a71b08319809c3b14512149

Observation 48872420-c0a9-49c1-a1b6-42bb76b17d3b · inbound

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution cites this paper.

MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 161

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T01:39:23.926310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-04T01:35:11.966506Z digest=sha256:c30588c79ab16c3e3763f80c373e1a81785a8df3e2a1ea4511a630a8248283fb

Observation 54dcbd29-0610-4074-a560-598177c4e162 · inbound

Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction cites this paper.

Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T10:15:03.503011Z digest=sha256:d3f7b5bfbbfd49eefc9e85906e3b1af843e8f727552e404dfcc5e2e62b1318af

Observation c3f0f6c6-a27d-476c-a3b3-c2901221c1d6 · inbound

AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization cites this paper.

AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T00:55:52.965890Z digest=sha256:d33c7737cc73168d7b798dce94b8c695712df1ab415b6eee3ac95763c48ea7d5

Observation 32adff1c-b1b1-4a8b-80d5-bc214f8c2403 · inbound

AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization cites this paper.

AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T23:35:07.613373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T23:29:12.545347Z digest=sha256:cb73b45940e8564c76f8a227964f8c2c54c098991b171af8e5904ebc6c182ed5

Observation 4212d254-c39a-40c9-9d29-38092cc7b065 · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:21d90e2c9a7621f9babf5104e1507ec6cc040184bd2e5680012280ab8d7fff8d

Observation 672cb521-246b-48bd-9dca-de6c0991ecbe · inbound

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies cites this paper.

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:37:22.985929Z digest=sha256:3e7c517064f69d221fa45fea6290e2a0103cfc40f1d3acc0b3d023a0b6911ba6

Observation c045fad3-5bf3-4edd-80fc-4c8670828acc · inbound

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA cites this paper.

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:48:09.071812Z digest=sha256:9e256bc9f60cafb24c0fae8020c7146e928cc2fce3dca924d6c6144f2aa32bb8

Observation dc6b6f94-f700-4f62-a434-6da05e27de3c · inbound

Large Language Models Lack Temporal Awareness of Medical Knowledge cites this paper.

Large Language Models Lack Temporal Awareness of Medical Knowledge MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 43

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T20:18:08.160768Z digest=sha256:89ffcb83ad89a55b6378f08fe3b622718554103635c0b50c065b0205762ec5c8

Observation d65abc5c-b48d-43f6-a642-ed0e3d3b4e88 · inbound

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation cites this paper.

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 39

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T14:26:38.603695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T18:47:46.239810Z digest=sha256:686bf87b17ff1cdce6da45406f445a93cdbb212417f900a32ab77c4aaa8dc5bf

Observation 55812b3a-6954-483a-8deb-28937209982f · inbound

Fully Open Meditron: An Auditable Pipeline for Clinical LLMs cites this paper.

Fully Open Meditron: An Auditable Pipeline for Clinical LLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:53:38.771817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T18:53:16.389210Z digest=sha256:9202ca2fbde671680d502a5ccb60250ff330247971296f90e6fee77ee6b1ed50

Observation 59e7a0e2-c6b9-41cb-a6ec-1749cb4a4634 · inbound

Fully Open Meditron: An Auditable Pipeline for Clinical LLMs cites this paper.

Fully Open Meditron: An Auditable Pipeline for Clinical LLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:15:00.117175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T19:14:06.301012Z digest=sha256:2df523787787def8e3c34513bfd628622b54902dd943efc198c89ae9c9f0f816

Observation 95391074-b241-441b-83ca-749b1519187f · inbound

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? cites this paper.

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 67

Resolution
malformed identifier
local_arxiv, observed 2026-05-20T17:48:48.822438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T17:45:02.896703Z digest=sha256:d9f58aa2d8f3d586671007500b9f34689e59e4c7f65990676ba7a15b88cc5c21

Observation ca86765d-829e-4960-845c-477932d03578 · inbound

ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning cites this paper.

ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:23:03.760948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T05:19:35.222068Z digest=sha256:0291b4288c6ced86a282d6112b2c5ecf51fa023944b4e05b98e1d5f9fda2bb2a

Observation 33e9182b-8544-4f99-8121-9d884276b0b6 · inbound

NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding cites this paper.

NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:59:45.784825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T06:54:55.254082Z digest=sha256:415c2a5be89fd6f72766015375e59aa1cffd1a45154db71a369fbb85add08f58

Observation c85c7064-5477-457d-8fcb-e18e82f47da0 · inbound

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs cites this paper.

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:55:23.447614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T04:53:33.343509Z digest=sha256:9a5b67e2b3311efdef73c520778a689b35427cff0940111d7087597a952a31f8

Observation 7816dc5c-541f-4b18-b5eb-35b694a079b3 · inbound

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning cites this paper.

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:53:27.009199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T12:43:58.323948Z digest=sha256:127f7fc767dffe02dd703a2b0f4d61258f8ed47505a3cc60238754cf542bd2d1

Observation 719eb01f-bf4f-4d10-ad0e-04ef4c9b6da5 · inbound

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning cites this paper.

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:01:11.077901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:01:11.077901Z digest=sha256:6675d5e82868021861c1e3921a3c79c1f44fe0cf68c46cb0ab9a9d03d210ef7a

Observation a7275da5-de27-4e80-8de7-df8efc98f7d5 · inbound

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection cites this paper.

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:13:30.113213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T14:08:40.968105Z digest=sha256:ac069be022950bde1eeb90ac2af82c591e1b8c5651ff0aac55aa533987d4ef96

Observation d2c790ea-3813-48b5-9379-580a9ed8a723 · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 132

Resolution
malformed identifier
local_arxiv, observed 2026-06-29T06:43:10.362776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:225ea632ab55f72b11e3f2c2eb54a6f404c556d60bb75a2c953b11563db4753f

Observation 49accd18-ae0e-4825-83b6-1fce3e5503ad · inbound

Overview of the ClinicalSkillQA 2026 Shared Task on Continuous Perception and Procedural Reasoning in Clinical Skill Assessment cites this paper.

Overview of the ClinicalSkillQA 2026 Shared Task on Continuous Perception and Procedural Reasoning in Clinical Skill Assessment MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T00:56:25.600920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T12:56:39.071730Z digest=sha256:44efd5c222c597e8e4554d01470d8c668c42e9bf240924e06b984acf324e8df0

Observation 44c24813-c531-46d1-8cdf-d36d4352ab14 · inbound

Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases cites this paper.

Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.197661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:33:50.348031Z digest=sha256:a3bfcec24216c25a476c7e4e0787d3e36bad2fa917c08233ccc14befecb67d45

Observation d8aa4c76-c0f8-4783-b41f-891c3a40aebe · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 114

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T22:31:21.528834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:2d51783f4edb68937c19afa11acfb0ddd945ec77a91306ef7c7dc37c32945380

Observation 6c41d075-8dfc-44e4-9406-73c55ae184a3 · inbound

PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow cites this paper.

PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:15:01.454125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T18:39:14.083052Z digest=sha256:8d0c8085835d0a299160c9b937716a0e8b8e6dba2e38032cdc13ef3b78d80b93

Observation 856940c0-4d29-44cf-8f71-bc9abe4e8525 · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:07:17.905485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:75d3b63b4f1c7def34aa0d77540e797b91175de4287ac2377c522302b99c6fb0

Observation 9e6b1dec-f544-42f4-a64f-aaf7f14127de · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 121

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:37:25.362554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:8a3a908b5a4f4662dce9f34480b3bb1ef737131497ac51da3ec33f98354312c6

Observation 306a91b9-acf7-444a-bc11-08bb9a15bd19 · inbound

Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory cites this paper.

Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:07:30.881909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T16:45:30.431403Z digest=sha256:49bc93028fc7d67f70fd246b46dd18611d39760a90c2aafae5462b3044e4c6b7

Observation a88060c2-2971-4941-94ae-5f5925a5127a · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 189

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:47:59.439147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:e399d8b77e083b27d3eb6f9c28b9089b9e9ed336c11de79c258417c8ba81167c

Observation 85ae5e85-f44d-490e-8969-4645295422f7 · inbound

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context cites this paper.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 51

Resolution
malformed identifier
local_arxiv, observed 2026-07-03T11:28:04.910150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:44ed7b14ba1141089eb244b9e1f00fcfd2ba07f0bf4fdf9ee3d3ca96f6466012

Observation a7d1cc88-8241-470f-ae98-baecc0da6daf · inbound

Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA cites this paper.

Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.773387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T21:00:06.185627Z digest=sha256:411e6f7d196b8ffe1cbbd65ac3a5dc8fe043d80a9756ba2a38250df033455032

Observation 23739f5c-0713-4072-a46d-d56d9764052a · inbound

CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays cites this paper.

CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 40

Resolution
malformed identifier
local_arxiv, observed 2026-07-04T05:59:37.423945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T15:05:13.422803Z digest=sha256:01dbac26bae3057f0a29cc06a0be308c56206a3d10faac1215ab5337e1ab1d7f

Observation 6b6aaf1a-4d1f-4df2-a0bf-edf760cd9648 · inbound

Latent Confidence Alignment for LLM Self-Assessment cites this paper.

Latent Confidence Alignment for LLM Self-Assessment MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:41.809911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T11:21:33.744610Z digest=sha256:a2287b231cd17f5e297975c88509e1caae916436b81d2e413ea03029bada9da1

Observation 22abc74e-6ac9-46d8-bf11-d81713b57c31 · inbound

MMGist: A Comprehensive Multimodal Benchmark for 2027 cites this paper.

MMGist: A Comprehensive Multimodal Benchmark for 2027 MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:49:41.621498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T11:05:14.573386Z digest=sha256:d4a5706523168d971419c82cbeef42a74e2549284e34dccf4da4313a54631be9

Observation 886ab2bc-75ad-485a-80fe-6de71f878af1 · inbound

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models cites this paper.

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T20:40:07.573527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T19:58:23.594907Z digest=sha256:014f7ab4273d98ddf56bdd6929b73e413ad71065fa87934b837bdf204fc26b8a

Observation c645a31a-7ca4-43fc-b4c5-8ad9b304d649 · inbound

Reasoning Quality Emerges Early: Data Curation for Reasoning Models cites this paper.

Reasoning Quality Emerges Early: Data Curation for Reasoning Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-26T05:08:59.952542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:1d5ef843150e5d0e2dafc9cf40aa0ca2b65a2263a29d765476d1eb6cd7e7046f

Observation bc6b8fa0-2b7f-475c-ae5c-6a95e4b23536 · inbound

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning cites this paper.

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:15:45.207200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T05:36:43.609602Z digest=sha256:a49c5a59c86db3f1d45cb4132fcf88c639b7f710d38f6a36b4a53474130030f1

Observation 910b5549-5573-4a10-b9e5-1d98a5bc533a · inbound

FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning cites this paper.

FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:57.474691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-03T21:08:02.793814Z digest=sha256:29ca8b08b23da303c4a298644e65ae72ac5a63a3acbc6e4595773c0a49516429

Observation 299e3775-ebab-4eb0-b926-14ce590cb344 · inbound

Gemma 4 Technical Report cites this paper.

Gemma 4 Technical Report MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T07:10:06.281832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:10:06.281832Z digest=sha256:2c17b1f772bbac9d570a3bdd8e58ef3c1641d8b129b318734cda004cc21ab3aa

Observation f77e7e17-e43d-46e9-b61f-9c5ddbc01eb0 · inbound

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs cites this paper.

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T00:57:35.205227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T00:57:35.205227Z digest=sha256:6049a80f700ce011940b5af6c30bea80787d29b37062d56c7ac07b4c45f9176c

Observation 92c4a64b-125d-4981-87df-64f0aa6714ac · inbound

Evidence-Grounded AI for Musculoskeletal Care cites this paper.

Evidence-Grounded AI for Musculoskeletal Care MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T06:32:29.811544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:32:29.811544Z digest=sha256:a9f0236e936739c54aed945be0bc0f143cada2cd851e4ee4e2df2df0e09b5788

Observation 8e42376f-5c2a-41f0-9723-1e65cace1f15 · inbound

Cura 1T: Specialized Model for Agentic Healthcare cites this paper.

Cura 1T: Specialized Model for Agentic Healthcare MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T02:16:39.947485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:16:39.947485Z digest=sha256:5f15016a9c4481e03681c535c0d19f2d6ce176af043a013f18ad03c16aad9c23

Observation 1a18290f-fb69-4e97-90c3-c689dd7a6708 · inbound

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering cites this paper.

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:29.120119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:04:29.120119Z digest=sha256:eace4cb59f0eaf888550c20c08ab3e6b848f843b25824897775c09c88c195cde

Observation 91269e56-7bc3-4f7c-9000-d352c58b5e09 · inbound

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts cites this paper.

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T13:52:07.924487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:52:07.924487Z digest=sha256:17dfe4c6193c3a1cf855c7102e485409ba8b073a1dacb52343a72ccf07971da9

Observation 0cc6bf25-81ce-4bee-b96c-b85c31023eb4 · inbound

Do Pathology Vision-Language Models Truly See Pathology? cites this paper.

Do Pathology Vision-Language Models Truly See Pathology? MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:05.181089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:05.181089Z digest=sha256:495adc666a65c3990b0797d69a2bd8caa795f63e5a9ad4a9833c0082b56f3403

Observation e284e0ee-8709-43c8-a2da-e330f6077dea · inbound

MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA cites this paper.

MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:07.519917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:12:07.519917Z digest=sha256:3430e62bf98ec7e37afc162f649eb871a75ca0599704c39e023333681f926ada

Observation 81cfd0c9-d0ce-4fb6-a6a4-312c336c1c78 · inbound

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents cites this paper.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.671396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.671396Z digest=sha256:f33e6f3de5245755143da270e0105409810f017a7374e2ca7f6c544cd996930b

Observation b6a1f3c6-bcce-4ae2-a7fa-3721e66fa241 · inbound

Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA cites this paper.

Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T05:36:47.916051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:36:47.916051Z digest=sha256:90ce9241f14095611384af02c2e4dcf1fec71ced611f50aa276f22df4d8b7c81

Observation d2990618-5134-47dd-948f-cd0808a2a1d8 · inbound

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers cites this paper.

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 206

Resolution
unresolved
no resolver link, observed 2026-08-15T14:34:14.608714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:34:14.608714Z digest=sha256:6d4e85985c8a8b3af5bd190926962f393b030df734339230fc00c2138df8cee5

Observation e68527a9-4ca7-46a3-b8f0-f5778ba0f14c · inbound

MIRA: Medical Image Reflection for Agentic Diagnosis cites this paper.

MIRA: Medical Image Reflection for Agentic Diagnosis MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:14.279760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:14.279760Z digest=sha256:f581996b8ac9f5a769fbafd20cf4e8ab2294d53c67a973aa34dc2ce0eb0ce188