Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T14:26:38.517787Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 88 inbound Pith citation observations for arXiv:2501.18362.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T14:26:38.517787Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:20:20.277083Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
30 of 30 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 1681f3dd-2327-4d38-a78c-18accf2e4868 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding org/CorpusID:268232499
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bb8ecff3-5e41-4fc1-9a04-78ece23abc7d · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 144bd2dc-e05d-46d3-b2fb-19978cdddf0d · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding The reduction was successful, as indicated by follow-up x-rays
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a24f93cd-838b-4b83-a028-bda24f51f104 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a394bcdb-a3e7-44e2-a79c-3d95aba52d98 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 494336fe-6a00-456f-9a4c-b004e1310925 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding If there is a suspicion of nerve injury, such as the axillary nerve in this case, an EMG would be helpful in confirming nerve dysfunction or damage
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 13fc9d3c-116e-406d-888b-99931e708187 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7b2860e3-3c64-417b-8ed4-e780e609a92d · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding There are no distinct P waves visible before each QRS complex; instead, there is a disorganized electrical activity, which is typical of fibrillatory waves
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3facc0df-cced-48ee-a493-8a4a4dbfe47b · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding saw-tooth
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c1cf6fc9-dec9-49c2-91df-4edcf2cf25c9 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding saw-tooth
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e134e893-1be6-4dbe-b30a-5c5e0436373d · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Analysis: The above response fails to fully grasp the question’s implications in several ways
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 91e08cf8-73d3-4bf2-8f08-c83aa91de78b · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Aging is associated with alterations in sleep architecture, which may lead to reductions in total sleep duration, slow-wave (deep) sleep, and sleep efficiency
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 44bc7d9a-0b50-45ea-843f-7d49e3d7fae5 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding His sleep appears peaceful, and he experiences no disruptive symptoms such as snoring or awakenings
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 03263f01-f0ef-406f-8075-78a8ca491727 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 048946b8-4d64-4023-8398-86d91d47169a · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding This can result in less restorative sleep and a less refreshed feeling upon waking
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation af6f1089-91f7-4814-98ef-0fac92b684fa · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding acute on chronic
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b7fce81-589c-4ef7-98af-3934cf197471 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 63adbc35-7d0f-4985-af8d-914ce78c9627 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1ad086e1-5e90-4a2e-a789-8a6433785b81 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e3fced3-a82e-42d0-b134-291150473ca8 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Rigorously ensure clarity and avoid ambiguity
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7856cec6-a23f-4446-bc6d-471b0a40c916 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Do not change, add, or delete any factual information
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a5e981b8-a5e7-413e-aa37-c69650314810 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Pay special attention to keep any tabular data in completely the same format as the original
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 04249930-be57-43ac-90f2-a29cc6aeb007 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Answer Choices: (A) [Option A] (B) [Option B]
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0c44156-0ac3-4f5f-b984-0c339de9d2d4 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dc22a6d1-3331-4cd4-a04c-14de87b7b8d7 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8eaa5b77-a94d-4497-ad15-9f09edad1f54 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 33df4d60-d8a3-49d9-8873-34883d712933 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding They should be clear, concise, and professionally worded
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ce60eb3-9ac6-41ff-b88f-cfd712f2b10c · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e5cd1a33-ff7a-433f-8547-1c28770ffc76 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Avoid options that are overtly illogical or unsupported.,→
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 135efca3-6d51-4b18-bfe8-1f01ca87ed36 · outbound
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding Answer Choices: (A) [Option A] (B) [Option B]
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dce3a636-1c70-416a-be7a-0d7582ac3b3c · inbound
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c3fdb2-6f00-44be-bef9-dc76724c0ff6 · inbound
BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6deb072-9fe2-4cc4-ba96-89e80df68f58 · inbound
Disentangling Reasoning and Knowledge in Medical Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee58effb-e067-41e3-bf66-4c5f270dcafe · inbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4b35547-716f-4666-acfb-e821b41c857e · inbound
Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b2aba6-1284-48d0-ae5a-629b81a01a6b · inbound
TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247bc90c-215b-453a-ad87-185f4f606d53 · inbound
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4aa9d2-33be-4bf6-8def-c628972511ad · inbound
Towards Large Reasoning Models for Agriculture MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d01c4d-8b62-4bf9-a0f4-2baa36503564 · inbound
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 361fc017-c4d4-4484-a06f-d00e895cb5f3 · inbound
Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f548da07-6d9c-438d-bc51-f2528ac72cc9 · inbound
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b373415e-d784-4a85-8e22-d02a0c3ab2c4 · inbound
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5851fe81-99b1-4ede-8416-5cae63b5bf82 · inbound
DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88cc0e5-d93c-4fef-8a9f-90ed663e0d5e · inbound
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b21cae6e-9d43-4d4a-81f2-714733e39145 · inbound
Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79acf541-ad73-46d9-9e9d-714dce4bdc6e · inbound
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05f1c7c1-6992-4164-925c-2b7b3d5f5976 · inbound
MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d22a01de-73f5-42e5-a047-00fe398d3a8f · inbound
MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0c6493ec-dbf9-4033-8a9e-b6e045fa6b1d · inbound
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8297b2ac-d6e7-40a3-8f9e-69e7730b7534 · inbound
Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095c61c5-541d-402c-bfe5-81fd8a7f004f · inbound
HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8be9093-1622-40c3-9f40-0cdeebdd0e72 · inbound
Capabilities of GPT-5 on Multimodal Medical Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85fbb256-4269-4d16-b2cd-7529460147fe · inbound
AdsQA: Towards Advertisement Video Understanding MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c9f4076-f74a-45f2-8fea-44f78b386640 · inbound
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 00f9f1d0-c640-4be0-9a55-37921d3b8350 · inbound
AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c3a3772-61bb-422c-9f92-44370eb6e772 · inbound
MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d682ce68-1629-4254-b1eb-e48346c3f9d1 · inbound
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9610ecc2-1b58-4e3c-940b-02d187915dc1 · inbound
Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d58c0f0-32c1-481f-911f-b7d8b8dab98a · inbound
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692ea18e-394a-496b-8778-b3574823c5d8 · inbound
MedGemma 1.5 Technical Report MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ae76ceb-4b9c-4968-bbd5-44ef72b8a87f · inbound
EXAONE 4.5 Technical Report MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b660a08e-c5e1-4234-a24f-06cbd6a39e56 · inbound
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 45d2525e-cacb-4566-8379-fd6343f029f8 · inbound
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13da444f-16bc-453a-8cb0-333e43f456fc · inbound
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 68a546cd-ca2a-48e1-9666-add64b0b2e6c · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dfd5bf8c-66ef-4eae-b0f0-f248f94f6c7b · inbound
Evaluation-driven Scaling for Scientific Discovery MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3ea0346c-9da3-4694-b66a-712a2985cd70 · inbound
Green Shielding: A User-Centric Approach Towards Trustworthy AI MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 50884898-2721-41f2-82d6-5828cf462e2f · inbound
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9298d93c-375f-449f-abef-a1fdf5de2c7d · inbound
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2ba01998-58db-4f36-a07f-b138d6a680a6 · inbound
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 161
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 48872420-c0a9-49c1-a1b6-42bb76b17d3b · inbound
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 161
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 54dcbd29-0610-4074-a560-598177c4e162 · inbound
Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c3f0f6c6-a27d-476c-a3b3-c2901221c1d6 · inbound
AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 32adff1c-b1b1-4a8b-80d5-bc214f8c2403 · inbound
AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4212d254-c39a-40c9-9d29-38092cc7b065 · inbound
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 672cb521-246b-48bd-9dca-de6c0991ecbe · inbound
MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c045fad3-5bf3-4edd-80fc-4c8670828acc · inbound
Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dc6b6f94-f700-4f62-a434-6da05e27de3c · inbound
Large Language Models Lack Temporal Awareness of Medical Knowledge MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d65abc5c-b48d-43f6-a642-ed0e3d3b4e88 · inbound
RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 55812b3a-6954-483a-8deb-28937209982f · inbound
Fully Open Meditron: An Auditable Pipeline for Clinical LLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 59e7a0e2-c6b9-41cb-a6ec-1749cb4a4634 · inbound
Fully Open Meditron: An Auditable Pipeline for Clinical LLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 95391074-b241-441b-83ca-749b1519187f · inbound
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ca86765d-829e-4960-845c-477932d03578 · inbound
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 33e9182b-8544-4f99-8121-9d884276b0b6 · inbound
NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c85c7064-5477-457d-8fcb-e18e82f47da0 · inbound
DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7816dc5c-541f-4b18-b5eb-35b694a079b3 · inbound
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 719eb01f-bf4f-4d10-ad0e-04ef4c9b6da5 · inbound
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7275da5-de27-4e80-8de7-df8efc98f7d5 · inbound
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d2c790ea-3813-48b5-9379-580a9ed8a723 · inbound
EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 49accd18-ae0e-4825-83b6-1fce3e5503ad · inbound
Overview of the ClinicalSkillQA 2026 Shared Task on Continuous Perception and Procedural Reasoning in Clinical Skill Assessment MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 44c24813-c531-46d1-8cdf-d36d4352ab14 · inbound
Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d8aa4c76-c0f8-4783-b41f-891c3a40aebe · inbound
Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6c41d075-8dfc-44e4-9406-73c55ae184a3 · inbound
PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 856940c0-4d29-44cf-8f71-bc9abe4e8525 · inbound
Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9e6b1dec-f544-42f4-a64f-aaf7f14127de · inbound
Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 306a91b9-acf7-444a-bc11-08bb9a15bd19 · inbound
Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a88060c2-2971-4941-94ae-5f5925a5127a · inbound
UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 189
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 85ae5e85-f44d-490e-8969-4645295422f7 · inbound
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a7d1cc88-8241-470f-ae98-baecc0da6daf · inbound
Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 23739f5c-0713-4072-a46d-d56d9764052a · inbound
CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6b6aaf1a-4d1f-4df2-a0bf-edf760cd9648 · inbound
Latent Confidence Alignment for LLM Self-Assessment MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 22abc74e-6ac9-46d8-bf11-d81713b57c31 · inbound
MMGist: A Comprehensive Multimodal Benchmark for 2027 MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 886ab2bc-75ad-485a-80fe-6de71f878af1 · inbound
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c645a31a-7ca4-43fc-b4c5-8ad9b304d649 · inbound
Reasoning Quality Emerges Early: Data Curation for Reasoning Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc6b8fa0-2b7f-475c-ae5c-6a95e4b23536 · inbound
Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 910b5549-5573-4a10-b9e5-1d98a5bc533a · inbound
FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 299e3775-ebab-4eb0-b926-14ce590cb344 · inbound
Gemma 4 Technical Report MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77e7e17-e43d-46e9-b61f-9c5ddbc01eb0 · inbound
Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c4a64b-125d-4981-87df-64f0aa6714ac · inbound
Evidence-Grounded AI for Musculoskeletal Care MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e42376f-5c2a-41f0-9723-1e65cace1f15 · inbound
Cura 1T: Specialized Model for Agentic Healthcare MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a18290f-fb69-4e97-90c3-c689dd7a6708 · inbound
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91269e56-7bc3-4f7c-9000-d352c58b5e09 · inbound
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc6bf25-81ce-4bee-b96c-b85c31023eb4 · inbound
Do Pathology Vision-Language Models Truly See Pathology? MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e284e0ee-8709-43c8-a2da-e330f6077dea · inbound
MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81cfd0c9-d0ce-4fb6-a6a4-312c336c1c78 · inbound
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a1f3c6-bcce-4ae2-a7fa-3721e66fa241 · inbound
Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2990618-5134-47dd-948f-cd0808a2a1d8 · inbound
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 206
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68527a9-4ca7-46a3-b8f0-f5778ba0f14c · inbound
MIRA: Medical Image Reflection for Agentic Diagnosis MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.