Pith. sign in

Paper Citation Record · LEDGER

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 10 inbound Pith citation observations for arXiv:2505.14107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14107 v4

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.489072Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:38:11.263118Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:27:18.623344Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved51
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56360b9f-9e2f-49b4-8488-d891ecf33e37 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:37.297073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:20.680647Z digest=sha256:c1ca8b34c31f225a5d4f4c835cfb5908f2b8f4873858bdac3169035b52b5523a

Observation add5754f-9f3e-4d26-9342-ff214d10c0a2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.789835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.789835Z digest=sha256:1bd3f9491e951eedc83366122291d674ee4e98e686264fa7341756982167b0fe

Observation 29a188e5-a87d-4848-bc63-f8bab946a0fa · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.896208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.896208Z digest=sha256:89b87066c5cab3ddb1b97c0594b2f9d5a2a3c1a67b8901fdd72e55ee1c75c7ac

Observation 1026facc-1e92-45cb-be99-89ba8001c0c8 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.985769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.985769Z digest=sha256:fb065cc3287bdbc8178c26df3728a007b9c7fa9065c01778b4a432af7ea83b7a

Observation 11adcf86-c5af-4869-896e-1239510694c2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:37.158163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:21.068068Z digest=sha256:4f2a91ccb47f3a497fb0abd9b25a1f7fe822e35e97d64d93d818584683dbe522

Observation 0a1a466e-dcb9-4c1c-8e0c-34cc4d7966a2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.899551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:21.167084Z digest=sha256:19e24cd5821b6d1d8e11c04e84343cb0a72e302a77962692f66a9a07eccf14bf

Observation 23b4fb40-e8a3-4f52-9475-dc012462c6f2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.704308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:21.281936Z digest=sha256:259618f955307b2cb1b044ba39967ae7d51690bd2e3258838c3b737383ac7626

Observation 50cd3b33-c295-4122-838b-a37faeff1e13 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.412901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.412901Z digest=sha256:3dadb2526c66ccb42d810b5cdf273bbc146f8b3be84c62c3549375df82066f67

Observation 2f8a5f54-5954-4049-8445-244175106aa9 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.512527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:21.517209Z digest=sha256:cda27ffe6720a1c74d2e09262152dccc30d01d3e1900033b212cfb255b64a38c

Observation 6ec658c3-c211-4854-bccd-9d19e1538875 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.748095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.748095Z digest=sha256:657c3e9ce897c3df4e7b75440b48cd55a1bc679cd7558892169dfff90fa3d69c

Observation 13095c2b-6ad1-4105-8fc5-ab73c61da6ea · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.824586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.824586Z digest=sha256:b2fc78b4df2b5c985586a99d2301283b4e6e472ea8bea8f80c3551fe64f3ab6e

Observation d13da93f-2bb2-43c4-8a63-ad77e29ae2cb · outbound

This paper cites O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.902894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.902894Z digest=sha256:5c03fd2866e3c42ed3d3dae78e54cbaf7464967b5a3aa2dd0eef0732e4517862

Observation 5a3d7ed5-6f5a-4d77-bb14-e3551883f2f3 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.963612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.963612Z digest=sha256:6aa5713ac77c0656258a7fc668120a1ff1a76afe434d2a978eb49097cdcbf9f5

Observation 7749b063-e0fb-4649-b3c8-43bf035ed78e · outbound

This paper cites Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:42:27.076360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:22.042716Z digest=sha256:f4160e0e89ea084476e1b9eafce6a2c72c7f7e37098e0b54fb1bcbb43b67514e

Observation 8acdc184-0a55-44fe-a814-696c7e73685e · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.163572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.163572Z digest=sha256:149045ab8ef50d4bdebf673eba76e75d573827b7eacd6b4aabdbe749f42807f0

Observation 3f1a278f-7fe7-45b2-84a6-794c38506942 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.245082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.245082Z digest=sha256:9cd280554437d6f85f6f1f3c46c002b5d2e07e09999d26a3f5568c8a18b3fff8

Observation 252428e3-ca33-4985-8889-e7b155bfcaf2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.278415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:22.318276Z digest=sha256:6afdd23cc22f95fba31d999fe7cb0a569d7416084b1216bb8d004ce732515f7a

Observation 5fced3ee-b9e0-408b-9371-699cdb69fb9e · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.014916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:22.418200Z digest=sha256:0b80448d124327f6684858381ec465902c34e22aa959c39b7a831e41627d526d

Observation 50458228-0e58-4fb2-bd2e-832ae2561881 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.775621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:22.509608Z digest=sha256:828e070965c2bc642b2b8c7095c4db3d435ef561006eab9cc3e04ac1c187e71b

Observation a7a65e49-e29f-4dbf-8a23-0a860b2877ba · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.583653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.583653Z digest=sha256:89b715aee650af7da597d612934a03362779e1549d7b224026bc1d7a7d29cdfb

Observation e988c9fd-a9fc-4a70-80cb-4fd7071f6588 · outbound

This paper cites From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.658240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.658240Z digest=sha256:76523678786cbc23d4c9a315e84a33697f0ac8fc94d31f16f41c9e4c8ee3a1a4

Observation ed0aff7f-6605-46a4-80eb-1c7516581bba · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.540492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:22.738801Z digest=sha256:d288ac5913dd394169e7386a37286273ad2d20af44509407d2e6676473e26879

Observation adb21f5c-2b62-4a06-b4c6-294df9f95d0c · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.305787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:22.826660Z digest=sha256:f2152ca09cf780b091751095e81a3d65e8cc38302fb23139bfc19f53b34601cd

Observation ef8e2c66-11bc-490a-96b2-52c7ef4f4a05 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.060820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:22.923453Z digest=sha256:ab5d8382becea63a6ed3e53efc022dffd363481b687d4b7c3cd41fcd17c1b8b3

Observation 86d4f465-4b4c-4223-91a8-df8860c40727 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.892989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:23.014155Z digest=sha256:13e188988f9d4ee6a04b3210fcce89bc8fd0910232283159820dff7e9f6fae7a

Observation d753f713-4f60-428d-adfe-305ab0f63ce4 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.737803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:23.136683Z digest=sha256:732a3e621258db31bdc610a3ebe949259aa6f88d4b6c469ecedf93cadd35f273

Observation e0f2a2b3-80f0-445d-ad20-e698540706d5 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.212775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.212775Z digest=sha256:b4de5cdeb299ae11c9c826c309ac90a62284baf3803645163c9887eecab7ff4e

Observation c52ef2d4-c7d1-47c7-8cb6-135021e83135 · outbound

This paper cites Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.341138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.341138Z digest=sha256:9c2f46a8778a577d9903df5a97d16e7244108af4ff4a8998ce84cd33bd4105df

Observation 7b82acba-893c-41df-b3ed-d825fbb530bc · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.518832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:23.423821Z digest=sha256:8ffbe79d2873363b7ace5efe290e9db84323add9e36aa0bd6b16ae18154dded2

Observation 143ce6d6-de68-44ed-84b8-5b90c3a4a49a · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.541553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.541553Z digest=sha256:fcce14684bc4df193ce5ed40497ab4cb38551df56d826e28902e21b9256025bc

Observation 03d0adc0-8f96-4fce-9a3e-a7ade274a3dd · outbound

This paper cites Qwen2.5 Technical Report.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.635382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.635382Z digest=sha256:7fc8aa751b0b3f283ab81ba3cf0dba887ea12f04c2658a645bbeafeadd9abacd

Observation dfb2780e-d949-4312-b15e-44740ee94f1b · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 32

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:34.329538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:23.729601Z digest=sha256:89227a5f85fcd51b9023708d059545a9ce51c23c0afedbd54669421c9cd2bd01

Observation b4bf1c95-d465-4baf-bf40-c34333cbcdc8 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.165599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:23.825815Z digest=sha256:ae37e3860fe15d33b03b4224206c8d537aed9ae26e2dee7518173629bbbe3671

Observation ec7c2477-4f05-4195-bce0-ed7d2be78f22 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.999351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:23.921620Z digest=sha256:bec574da9f7b0517a634acc086b718c266406c3364d5168ac47e53b17fcb983e

Observation c8024d9a-767f-437d-91b5-064645d21feb · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.112367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.112367Z digest=sha256:f9c07531a65dd9632a8cf1bc0612757557b199799fdad6944a2bda49022dbfed

Observation 1f349e7e-9776-4a43-a14d-1feee6535dca · outbound

This paper cites CMB: A Comprehensive Medical Benchmark in Chinese.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models CMB: A Comprehensive Medical Benchmark in Chinese

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.190092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.190092Z digest=sha256:a8d9bc073cac0d0a4a3525467021c10e5345ee0b0b6d32bdc1f46c7971dcafff

Observation 6ca70e4c-e4fc-44f5-9602-498d4d853e7c · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.277569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.277569Z digest=sha256:d9de47802c27598e86a202e9bcc8bedcfefbbc0c38e8a7fd8a0f871b3e58f067

Observation b70375a0-142c-47ef-b790-23de3e044cce · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.437348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.437348Z digest=sha256:516f55885c43dc273f6b77e1968f49ee6896fda4550f758d36ffd69d970d0a82

Observation a1a2fb92-3829-4cc4-b075-bef9db82819a · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.625676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:24.591939Z digest=sha256:ab1dc38af5fcb8db16eb84c724868800a3ad5232e2c19bd83d47ba6731d939f7

Observation c0d1f68f-f404-453c-94e5-88b1a5f37ff6 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.452128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:24.700756Z digest=sha256:7a4ffafcc71ad2d6d7314eeb91d828c7171703475f4774da46567579845a849a

Observation 8e0f9c04-a308-4aa7-ab73-06a7f94f0cf9 · outbound

This paper cites FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.782003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.782003Z digest=sha256:c1f888fc86269e892a10d85bdfbb88d7f07662f927b83bf2210ee18095d1e2c1

Observation 9ecad0b3-74e9-43dd-a2f4-f0a0ed2b5f08 · outbound

This paper cites o1-Coder: an o1 Replication for Coding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models o1-Coder: an o1 Replication for Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.855422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.855422Z digest=sha256:68d0f4a95d3f0cd5e0c28338a5ffc0c6407a858c41d52f9fcf65a5baad2cd9a5

Observation bdbf2e02-346c-4d5c-80a6-1ea18350862f · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.281918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:24.949716Z digest=sha256:32005c49fa673b058f0f0b2cd7a8c98c25ce12c8500eb6b7b2b261c1f800b1ad

Observation ee58effb-e067-41e3-bf66-4c5f270dcafe · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.004807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.004807Z digest=sha256:6e4015a72f71458cd23a1b7705b08d3496ea5c0fca9854a39c0db405b123692b

Observation 4f9f9bad-5ba3-4f51-a0fb-1237c3b00e20 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.096586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.067031Z digest=sha256:58ce63ce211505c856e83cc178c41f2bb84774b5210e58224bd71c9259cb54f0

Observation cb762cc9-3243-49ea-adb5-5a1fa467d0ce · outbound

This paper cites - Use appropriate Markdown syntax for headings based on the original heading hierarchy (e.g., ‘#‘, ‘##‘, ‘###‘).

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Use appropriate Markdown syntax for headings based on the original heading hierarchy (e.g., ‘#‘, ‘##‘, ‘###‘)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.955211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.150270Z digest=sha256:5b0648a79c3080e4fa21bfca22930d3ae1e872db27ab838c15f389acc8fc5299

Observation cbd173d9-fbc5-4850-9362-586f09d2e2d9 · outbound

This paper cites - Completely delete citations and footnotes, including their in-text markers.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Completely delete citations and footnotes, including their in-text markers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.770494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.216414Z digest=sha256:de0c1c9091a09d3ac447ff3dc29736aaabf1118a3c6a7f1a21818c406fca9c7e

Observation f00fc92c-ce94-42eb-95ef-b59b375244f5 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:32.596816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.292911Z digest=sha256:2bf090f3aeb3f665df70e2f2646751b342764691167c95e751fe3d45a825ec9f

Observation 0bf3ae46-805e-4e72-a9d4-2a191cdb0763 · outbound

This paper cites - Only perform formatting and cleaning adjustments without modifying the original content.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Only perform formatting and cleaning adjustments without modifying the original content

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.345656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.349954Z digest=sha256:6bbc9c4c9e4f795687352daa6290150d574b2dd7e245ded346f5dc4dae4a964d

Observation 3b025b0b-76e8-4009-a527-6f4414c86a56 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.993742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.419255Z digest=sha256:ded75abc5999d8e3ab6d7a501af95790a03977c0bcc996a07c4bc3308ebed1e7

Observation fe1a29ce-ff78-487b-b7e8-383fd5d81997 · outbound

This paper cites - **Prohibited Content**: Any direct diagnostic statements involving disease names.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - **Prohibited Content**: Any direct diagnostic statements involving disease names

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:31.554814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.491524Z digest=sha256:64709d09fdd756d4a4de90a80a0660523602b7034c737a5d0cc61ee442f5d2b6

Observation cc00d117-d35b-460e-9deb-1529740beb7b · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.228417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.553291Z digest=sha256:fc759c26f3b3950770243f5a3b04f0c88d22e2cd925563b137e4456afb8ff4ec

Observation 7ec7ea28-49a4-404e-84cf-37fc9aa3e8ac · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.015175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.614692Z digest=sha256:d566b4676df024575121fef8cb49679a62f7423a6da1b88bd51a8722c7905640

Observation 8ca475ba-f535-4936-94d2-b5a53b8afc8b · outbound

This paper cites DiagnosisArena-MCQ Evaluation Prompt You are an expert in the field of rare diseases.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models DiagnosisArena-MCQ Evaluation Prompt You are an expert in the field of rare diseases

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:30.685315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.692156Z digest=sha256:78f481debfb052218198ffad4ff83f5135a9226466cb5a643b7355feb74332c2

Observation a19227c1-4dd3-450b-8594-6b35a146e585 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 58

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:30.306619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.769469Z digest=sha256:d756d8f50b577037ff8f2a4e75d8d5c6aaf8e02991500a97b8b15286d36b0cd0

Observation 0ae94c43-e39a-435b-b59a-1fd19ee19465 · outbound

This paper cites C Cases C.1 Benchmark Cases In this case, Case Information, Physical Examination, and Diagnostic Tests are components of the patient’s medical record.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C Cases C.1 Benchmark Cases In this case, Case Information, Physical Examination, and Diagnostic Tests are components of the patient’s medical record

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:30.090590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.833573Z digest=sha256:307b6cf9f3ae1f1892227945fd7914fea907e146215cd162510232cd0207ce82

Observation 83fdf6e5-502b-413a-8de7-2fbffa51c1bf · outbound

This paper cites Highly mobile, filamentous, causing embolic symptoms or arrhythmias.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Highly mobile, filamentous, causing embolic symptoms or arrhythmias

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.800933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.898275Z digest=sha256:5e19c2560ec141c95ea5d302e287a882c892cd980f19a23cc7c13d83e43f125b

Observation 406b80e0-21cc-494b-a5d9-0dee71f9d2ab · outbound

This paper cites Could be in LVOT, leading to similar symptoms.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Could be in LVOT, leading to similar symptoms

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.553124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:25.969036Z digest=sha256:3c201a432fd91c0a3b15ee935673255a652f8d05c06fa0b03fab2b4836555049

Observation dda23e77-1902-4c30-988e-5ec19956f861 · outbound

This paper cites Mobile and can cause obstruction.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mobile and can cause obstruction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.300596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:26.033021Z digest=sha256:9f761ba8ab992d9969691e2547c7a91c0c4d140fce978177a5735dc4bd6985cf

Observation a5b39362-9d2f-44e3-8ac8-e07a2071d6d2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:28.890974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:26.090955Z digest=sha256:3d71888d1c9e23c47069618bff03360fde859216bae7b9c69659981e4a511438

Observation cac5c5c8-f692-4bac-abe8-81a1ebf3cb64 · outbound

This paper cites Wait, but the movement pattern and CT findings might make fibroelastoma more likely than Lambl’s.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Wait, but the movement pattern and CT findings might make fibroelastoma more likely than Lambl’s

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:28.599013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:26.152160Z digest=sha256:fbc4b33719a79e25d93f8ebcbd0ae9418eebe8e1a809ed7b7c8033f475d51565

Observation 9b4dd91b-2b5a-48e0-a433-31f02bff4043 · outbound

This paper cites The key is the mobile mass in LVOT causing possible embolic events (leading to AFib) or obstruction.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models The key is the mobile mass in LVOT causing possible embolic events (leading to AFib) or obstruction

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:27.907921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:26.229267Z digest=sha256:7910631930b4ed7c4a6b4bd7e28f99381f731c398e5f4e319eb3cc6d0b09031d

Observation d33dae21-4284-4cfc-8e96-560b5b0f64f2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 67

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:28.214479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:26.328940Z digest=sha256:bc3d8e4e47453bc246ec69231b3c5e58a4a42dbe1a42fad3b88a57aa53eb246c

Observation 5ffe65f1-7028-494e-b52a-c5be7314a30a · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:27.623039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:26.489072Z digest=sha256:e6c258c736fde04b8fd9db19eb8c9f5c5498799270499b65cf915818a1c14b23

Observation bde34791-774d-430d-b597-946b2915abad · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.630549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.630549Z digest=sha256:ef05fd2b659f6304298825411e972413ecd2e7def2a543eddcca5e3d72f68ac5

Observation 5ef940ea-ed06-44c9-ae68-bc31d44abad0 · outbound

This paper cites Advances in neural information processing systems, 35:24824–24837.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Advances in neural information processing systems, 35:24824–24837

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.810735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:24.508298Z digest=sha256:ceb11fa750461bc704373cfb77949b6cb9cc9d27f3f4fda51f1b0b48b31a3e66

Observation 4c32e658-c443-46e1-aeed-e7fb5d426ae3 · outbound

This paper cites Baichuan-M1: Pushing the Medical Capability of Large Language Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Baichuan-M1: Pushing the Medical Capability of Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.011816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.011816Z digest=sha256:8b0a237aa20c757c9a4334aad42ed244d683375e3acc41fe62a4d4d26ff175b8

Pith citing papers

Observation 09047aae-f843-40e3-85b4-cf933f609390 · inbound

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks cites this paper.

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:31:25.606181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T13:26:58.309566Z digest=sha256:a67524c9096bc7c546e58a52de607dc0fe95141cce4a0be6a3dbe73abbd83dbd

Observation b7ec127c-9cc0-44a7-94c9-86527838c0db · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:56:50.369713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:77b5e393f62e50e259f58c5f55df2b646b51947086994e730522a816985510ae

Observation 56606448-0adb-4b38-bf1b-6c8a2b8894c4 · inbound

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning cites this paper.

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:50.561904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:19:15.053725Z digest=sha256:a38e2259cbd3cffb1585e1c58497f6808981bcecf35d48788def8c38121e9928

Observation eb509866-ada7-44b7-a024-1d689965352d · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:37.715348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:03:34.773345Z digest=sha256:6a2e173308268aff991799ae72f88a68b7bb423960048ff8827fa157089e46d3

Observation 4c17d835-2f8e-4172-ae51-856250113229 · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:27:59.758936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:25:31.662152Z digest=sha256:e617fc92dd009c5121c1e7dde42fe164ac55afe93c2fa664e579d726035ea26f

Observation 7eb391fb-79aa-4e59-ba84-88b6f2ba5083 · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:43:10.455707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:065c3989e015a894c88351fbd0ec00f93dbe2d32533dc237fcef86a558099ee1

Observation 96c78045-c903-4f3b-afa0-f6f2d46e03af · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.921236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:4f69fd85e88d4d6815ea6c3abc940dae5aa177f3d13f78c10ea46b30c237ab7e

Observation d98ffec1-12da-4a37-ab73-142cbea8de42 · inbound

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation cites this paper.

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.624793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T19:21:44.653877Z digest=sha256:4512e6df7a49f12119c95e30eb0eefd4a9ab58204effc61c62184a083ec09fbb

Observation fbbe7ad0-62de-42bd-a871-203c3e9afd2d · inbound

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs cites this paper.

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:38:11.263118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:38:11.263118Z digest=sha256:18305c5de64f867829a9e10ccc894fbb8cae3cdc77249dc0f14677221074ac56

Observation 3530e665-cd62-480a-96b0-5471f5a47e1d · inbound

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases cites this paper.

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:59.356644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:07:59.356644Z digest=sha256:b1d95f0803cb638dc45fcca33bd262ffc946a053bd661ede5e1bdc9080e2af7d