Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.489072Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 10 inbound Pith citation observations for arXiv:2505.14107.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.489072Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:38:11.263118Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T19:27:18.623344Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 56360b9f-9e2f-49b4-8488-d891ecf33e37 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation add5754f-9f3e-4d26-9342-ff214d10c0a2 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a188e5-a87d-4848-bc63-f8bab946a0fa · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1026facc-1e92-45cb-be99-89ba8001c0c8 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11adcf86-c5af-4869-896e-1239510694c2 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a1a466e-dcb9-4c1c-8e0c-34cc4d7966a2 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23b4fb40-e8a3-4f52-9475-dc012462c6f2 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50cd3b33-c295-4122-838b-a37faeff1e13 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f8a5f54-5954-4049-8445-244175106aa9 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ec658c3-c211-4854-bccd-9d19e1538875 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13095c2b-6ad1-4105-8fc5-ab73c61da6ea · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13da93f-2bb2-43c4-8a63-ad77e29ae2cb · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a3d7ed5-6f5a-4d77-bb14-e3551883f2f3 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7749b063-e0fb-4649-b3c8-43bf035ed78e · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8acdc184-0a55-44fe-a814-696c7e73685e · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1a278f-7fe7-45b2-84a6-794c38506942 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252428e3-ca33-4985-8889-e7b155bfcaf2 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fced3ee-b9e0-408b-9371-699cdb69fb9e · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50458228-0e58-4fb2-bd2e-832ae2561881 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7a65e49-e29f-4dbf-8a23-0a860b2877ba · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e988c9fd-a9fc-4a70-80cb-4fd7071f6588 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0aff7f-6605-46a4-80eb-1c7516581bba · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adb21f5c-2b62-4a06-b4c6-294df9f95d0c · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef8e2c66-11bc-490a-96b2-52c7ef4f4a05 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86d4f465-4b4c-4223-91a8-df8860c40727 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d753f713-4f60-428d-adfe-305ab0f63ce4 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0f2a2b3-80f0-445d-ad20-e698540706d5 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey: A Strategic Progress Report -- Part 1
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c52ef2d4-c7d1-47c7-8cb6-135021e83135 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b82acba-893c-41df-b3ed-d825fbb530bc · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 143ce6d6-de68-44ed-84b8-5b90c3a4a49a · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d0adc0-8f96-4fce-9a3e-a7ade274a3dd · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Qwen2.5 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb2780e-d949-4312-b15e-44740ee94f1b · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4bf1c95-d465-4baf-bf40-c34333cbcdc8 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec7c2477-4f05-4195-bce0-ed7d2be78f22 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8024d9a-767f-437d-91b5-064645d21feb · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f349e7e-9776-4a43-a14d-1feee6535dca · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models CMB: A Comprehensive Medical Benchmark in Chinese
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca70e4c-e4fc-44f5-9602-498d4d853e7c · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70375a0-142c-47ef-b790-23de3e044cce · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a2fb92-3829-4cc4-b075-bef9db82819a · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0d1f68f-f404-453c-94e5-88b1a5f37ff6 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e0f9c04-a308-4aa7-ab73-06a7f94f0cf9 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ecad0b3-74e9-43dd-a2f4-f0a0ed2b5f08 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models o1-Coder: an o1 Replication for Coding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdbf2e02-346c-4d5c-80a6-1ea18350862f · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee58effb-e067-41e3-bf66-4c5f270dcafe · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9f9bad-5ba3-4f51-a0fb-1237c3b00e20 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb762cc9-3243-49ea-adb5-5a1fa467d0ce · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Use appropriate Markdown syntax for headings based on the original heading hierarchy (e.g., ‘#‘, ‘##‘, ‘###‘)
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbd173d9-fbc5-4850-9362-586f09d2e2d9 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Completely delete citations and footnotes, including their in-text markers
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f00fc92c-ce94-42eb-95ef-b59b375244f5 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bf3ae46-805e-4e72-a9d4-2a191cdb0763 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Only perform formatting and cleaning adjustments without modifying the original content
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b025b0b-76e8-4009-a527-6f4414c86a56 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe1a29ce-ff78-487b-b7e8-383fd5d81997 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - **Prohibited Content**: Any direct diagnostic statements involving disease names
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc00d117-d35b-460e-9deb-1529740beb7b · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ec7ea28-49a4-404e-84cf-37fc9aa3e8ac · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ca475ba-f535-4936-94d2-b5a53b8afc8b · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models DiagnosisArena-MCQ Evaluation Prompt You are an expert in the field of rare diseases
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a19227c1-4dd3-450b-8594-6b35a146e585 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ae94c43-e39a-435b-b59a-1fd19ee19465 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C Cases C.1 Benchmark Cases In this case, Case Information, Physical Examination, and Diagnostic Tests are components of the patient’s medical record
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83fdf6e5-502b-413a-8de7-2fbffa51c1bf · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Highly mobile, filamentous, causing embolic symptoms or arrhythmias
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 406b80e0-21cc-494b-a5d9-0dee71f9d2ab · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Could be in LVOT, leading to similar symptoms
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dda23e77-1902-4c30-988e-5ec19956f861 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mobile and can cause obstruction
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5b39362-9d2f-44e3-8ac8-e07a2071d6d2 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cac5c5c8-f692-4bac-abe8-81a1ebf3cb64 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Wait, but the movement pattern and CT findings might make fibroelastoma more likely than Lambl’s
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b4dd91b-2b5a-48e0-a433-31f02bff4043 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models The key is the mobile mass in LVOT causing possible embolic events (leading to AFib) or obstruction
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d33dae21-4284-4cfc-8e96-560b5b0f64f2 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ffe65f1-7028-494e-b52a-c5be7314a30a · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bde34791-774d-430d-b597-946b2915abad · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Measuring Massive Multitask Language Understanding
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef940ea-ed06-44c9-ae68-bc31d44abad0 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Advances in neural information processing systems, 35:24824–24837
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c32e658-c443-46e1-aeed-e7fb5d426ae3 · outbound
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Baichuan-M1: Pushing the Medical Capability of Large Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09047aae-f843-40e3-85b4-cf933f609390 · inbound
Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7ec127c-9cc0-44a7-94c9-86527838c0db · inbound
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56606448-0adb-4b38-bf1b-6c8a2b8894c4 · inbound
From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb509866-ada7-44b7-a024-1d689965352d · inbound
EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c17d835-2f8e-4172-ae51-856250113229 · inbound
EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7eb391fb-79aa-4e59-ba84-88b6f2ba5083 · inbound
EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96c78045-c903-4f3b-afa0-f6f2d46e03af · inbound
Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d98ffec1-12da-4a37-ab73-142cbea8de42 · inbound
RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fbbe7ad0-62de-42bd-a871-203c3e9afd2d · inbound
DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3530e665-cd62-480a-96b0-5471f5a47e1d · inbound
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.