Pith. sign in

Paper Citation Record · LEDGER

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 10 inbound Pith citation observations for arXiv:2505.14107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14107 v4

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.489072Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:38:11.263118Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:27:18.623344Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved51
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56360b9f-9e2f-49b4-8488-d891ecf33e37 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:37.297073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:20.680647Z digest=sha256:6b59436262e8a5a7411483f833355ae4ba044ea8737999c5d8aea427a6224c6e

Observation add5754f-9f3e-4d26-9342-ff214d10c0a2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.789835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.789835Z digest=sha256:1bd3f9491e951eedc83366122291d674ee4e98e686264fa7341756982167b0fe

Observation 29a188e5-a87d-4848-bc63-f8bab946a0fa · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.896208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.896208Z digest=sha256:89b87066c5cab3ddb1b97c0594b2f9d5a2a3c1a67b8901fdd72e55ee1c75c7ac

Observation 1026facc-1e92-45cb-be99-89ba8001c0c8 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.985769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.985769Z digest=sha256:fb065cc3287bdbc8178c26df3728a007b9c7fa9065c01778b4a432af7ea83b7a

Observation 11adcf86-c5af-4869-896e-1239510694c2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:37.158163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:21.068068Z digest=sha256:49e5b86a99b628d5e1873c8640e8c3869617c85ee5100388880c3b3247edd5f5

Observation 0a1a466e-dcb9-4c1c-8e0c-34cc4d7966a2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.899551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:21.167084Z digest=sha256:a1b67b8f13f87427bbcfe8ae539d5ad54fae15a37c3df994641df69ecac5ea98

Observation 23b4fb40-e8a3-4f52-9475-dc012462c6f2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.704308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:21.281936Z digest=sha256:723f5500dbdbb351ebf35380a0e3c03c3cceda996f280019b744ca1d9dc3a45e

Observation 50cd3b33-c295-4122-838b-a37faeff1e13 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.412901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.412901Z digest=sha256:3dadb2526c66ccb42d810b5cdf273bbc146f8b3be84c62c3549375df82066f67

Observation 2f8a5f54-5954-4049-8445-244175106aa9 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.512527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:21.517209Z digest=sha256:026dbebc8cca7a214afd2dfeb5a7ac9d3c9d20629134f3d514f3d2b890032190

Observation 6ec658c3-c211-4854-bccd-9d19e1538875 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.748095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.748095Z digest=sha256:ccedfea75bff9261ee22b983ea289b5ed7c7d8e64252b541321037e9680f14f9

Observation 13095c2b-6ad1-4105-8fc5-ab73c61da6ea · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.824586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.824586Z digest=sha256:b2fc78b4df2b5c985586a99d2301283b4e6e472ea8bea8f80c3551fe64f3ab6e

Observation d13da93f-2bb2-43c4-8a63-ad77e29ae2cb · outbound

This paper cites O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.902894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.902894Z digest=sha256:5c03fd2866e3c42ed3d3dae78e54cbaf7464967b5a3aa2dd0eef0732e4517862

Observation 5a3d7ed5-6f5a-4d77-bb14-e3551883f2f3 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.963612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.963612Z digest=sha256:6aa5713ac77c0656258a7fc668120a1ff1a76afe434d2a978eb49097cdcbf9f5

Observation 7749b063-e0fb-4649-b3c8-43bf035ed78e · outbound

This paper cites Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:42:27.076360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:22.042716Z digest=sha256:d83bd3ca7a7045ea8e7a7a24cf194c4f555cb864ee6bffa7d0fd5f32b768f541

Observation 8acdc184-0a55-44fe-a814-696c7e73685e · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.163572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.163572Z digest=sha256:149045ab8ef50d4bdebf673eba76e75d573827b7eacd6b4aabdbe749f42807f0

Observation 3f1a278f-7fe7-45b2-84a6-794c38506942 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.245082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.245082Z digest=sha256:9cd280554437d6f85f6f1f3c46c002b5d2e07e09999d26a3f5568c8a18b3fff8

Observation 252428e3-ca33-4985-8889-e7b155bfcaf2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.278415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:22.318276Z digest=sha256:0e6889b9444128c441224c6f8ec68b7f44091781ced292ba144ee2a3bb3e8f33

Observation 5fced3ee-b9e0-408b-9371-699cdb69fb9e · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.014916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:22.418200Z digest=sha256:baf574307acfad1bccd8a42ef92d007eea9c45906dea19729e6b26a0a6a50edf

Observation 50458228-0e58-4fb2-bd2e-832ae2561881 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.775621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:22.509608Z digest=sha256:299570daf4b2cc4769f9f6b349cd60955b5cf0f4ad76ac26979c66bdcfcb2546

Observation a7a65e49-e29f-4dbf-8a23-0a860b2877ba · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.583653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.583653Z digest=sha256:89b715aee650af7da597d612934a03362779e1549d7b224026bc1d7a7d29cdfb

Observation e988c9fd-a9fc-4a70-80cb-4fd7071f6588 · outbound

This paper cites From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.658240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.658240Z digest=sha256:76523678786cbc23d4c9a315e84a33697f0ac8fc94d31f16f41c9e4c8ee3a1a4

Observation ed0aff7f-6605-46a4-80eb-1c7516581bba · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.540492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:22.738801Z digest=sha256:a5756722b287daf56ae0ca489ca29d157ff381e70796739519f0600e48860602

Observation adb21f5c-2b62-4a06-b4c6-294df9f95d0c · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.305787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:22.826660Z digest=sha256:cebbb1e3177c07e4e9aac691c903fa4b301fbde5169d7cc7cde1fc942786cb8a

Observation ef8e2c66-11bc-490a-96b2-52c7ef4f4a05 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.060820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:22.923453Z digest=sha256:8bcb9301db790c7d3772c29822b8f7808df0c8a8c2e4b4182cfefd372e4f3960

Observation 86d4f465-4b4c-4223-91a8-df8860c40727 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.892989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:23.014155Z digest=sha256:25dd287d4ec1ae719286bd3bf507adcce5e884ab4656f1bc06a933aaecce0fc1

Observation d753f713-4f60-428d-adfe-305ab0f63ce4 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.737803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:23.136683Z digest=sha256:aa0ca7cd38df8f48aafe5e8a4a83e99424832eb357d50cd84d98be6baadf15ec

Observation e0f2a2b3-80f0-445d-ad20-e698540706d5 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.212775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.212775Z digest=sha256:b4de5cdeb299ae11c9c826c309ac90a62284baf3803645163c9887eecab7ff4e

Observation c52ef2d4-c7d1-47c7-8cb6-135021e83135 · outbound

This paper cites Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.341138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.341138Z digest=sha256:9c2f46a8778a577d9903df5a97d16e7244108af4ff4a8998ce84cd33bd4105df

Observation 7b82acba-893c-41df-b3ed-d825fbb530bc · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.518832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:23.423821Z digest=sha256:0bb073498d92ffe077395f69dbcde4c6312048c8be2a4a0e649270dc16aa1eaa

Observation 143ce6d6-de68-44ed-84b8-5b90c3a4a49a · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.541553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.541553Z digest=sha256:fcce14684bc4df193ce5ed40497ab4cb38551df56d826e28902e21b9256025bc

Observation 03d0adc0-8f96-4fce-9a3e-a7ade274a3dd · outbound

This paper cites Qwen2.5 Technical Report.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.635382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.635382Z digest=sha256:7fc8aa751b0b3f283ab81ba3cf0dba887ea12f04c2658a645bbeafeadd9abacd

Observation dfb2780e-d949-4312-b15e-44740ee94f1b · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 32

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:34.329538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:23.729601Z digest=sha256:04a67c37ff987906915f7de2cb3fbd2044a34c576d1bb2fcd1eeceb3ca79a438

Observation b4bf1c95-d465-4baf-bf40-c34333cbcdc8 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.165599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:23.825815Z digest=sha256:05a447b7846de991fb48e759a3e7b7cbe4b9c7f4e6b0fcd8d10607cca5bced96

Observation ec7c2477-4f05-4195-bce0-ed7d2be78f22 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.999351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:23.921620Z digest=sha256:ad3a2d1cba9d3a8ced6be7e79ec54abeb48130e80cd77a7f4a25511749d9ecce

Observation c8024d9a-767f-437d-91b5-064645d21feb · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.112367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.112367Z digest=sha256:f9c07531a65dd9632a8cf1bc0612757557b199799fdad6944a2bda49022dbfed

Observation 1f349e7e-9776-4a43-a14d-1feee6535dca · outbound

This paper cites CMB: A Comprehensive Medical Benchmark in Chinese.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models CMB: A Comprehensive Medical Benchmark in Chinese

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.190092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.190092Z digest=sha256:a8d9bc073cac0d0a4a3525467021c10e5345ee0b0b6d32bdc1f46c7971dcafff

Observation 6ca70e4c-e4fc-44f5-9602-498d4d853e7c · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.277569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.277569Z digest=sha256:d9de47802c27598e86a202e9bcc8bedcfefbbc0c38e8a7fd8a0f871b3e58f067

Observation b70375a0-142c-47ef-b790-23de3e044cce · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.437348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.437348Z digest=sha256:516f55885c43dc273f6b77e1968f49ee6896fda4550f758d36ffd69d970d0a82

Observation a1a2fb92-3829-4cc4-b075-bef9db82819a · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.625676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:24.591939Z digest=sha256:ed8f2dd1420a72ea31f05f72dff325f8dbcef8979da0e41a01ee1f33956d7d13

Observation c0d1f68f-f404-453c-94e5-88b1a5f37ff6 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.452128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:24.700756Z digest=sha256:90b8de01b965cb1c460ff0f767ce2de9e747f46db8aaa2a6b255271fa17d11d2

Observation 8e0f9c04-a308-4aa7-ab73-06a7f94f0cf9 · outbound

This paper cites FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.782003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.782003Z digest=sha256:c1f888fc86269e892a10d85bdfbb88d7f07662f927b83bf2210ee18095d1e2c1

Observation 9ecad0b3-74e9-43dd-a2f4-f0a0ed2b5f08 · outbound

This paper cites o1-Coder: an o1 Replication for Coding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models o1-Coder: an o1 Replication for Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.855422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.855422Z digest=sha256:68d0f4a95d3f0cd5e0c28338a5ffc0c6407a858c41d52f9fcf65a5baad2cd9a5

Observation bdbf2e02-346c-4d5c-80a6-1ea18350862f · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.281918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:24.949716Z digest=sha256:2bb440ad600a1c847e8f0e8f8feb9dfde37a1bc960ecffaa0c01826b9d41535a

Observation ee58effb-e067-41e3-bf66-4c5f270dcafe · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.004807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.004807Z digest=sha256:087d479015e8833061cfa667bee71a15cd9b51a6283a535cc8fc49119a2d0085

Observation 4f9f9bad-5ba3-4f51-a0fb-1237c3b00e20 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.096586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.067031Z digest=sha256:885878aaa609634def5a832c4a0af6066b92940d1896ca5de30fc15e1a9b3194

Observation cb762cc9-3243-49ea-adb5-5a1fa467d0ce · outbound

This paper cites - Use appropriate Markdown syntax for headings based on the original heading hierarchy (e.g., ‘#‘, ‘##‘, ‘###‘).

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Use appropriate Markdown syntax for headings based on the original heading hierarchy (e.g., ‘#‘, ‘##‘, ‘###‘)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.955211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.150270Z digest=sha256:3af12dbbd1140926ce06e691562f3fbca70b5feacd096259e3099985de18301d

Observation cbd173d9-fbc5-4850-9362-586f09d2e2d9 · outbound

This paper cites - Completely delete citations and footnotes, including their in-text markers.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Completely delete citations and footnotes, including their in-text markers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.770494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.216414Z digest=sha256:6b922b41ac3afbc64dad84dec93f07684d23c71e0157095e464004aa38a82ee4

Observation f00fc92c-ce94-42eb-95ef-b59b375244f5 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:32.596816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.292911Z digest=sha256:2a88c16300897db46d92c8a0b9f29125b80fa0ab41875d387a2f7bae9c782dac

Observation 0bf3ae46-805e-4e72-a9d4-2a191cdb0763 · outbound

This paper cites - Only perform formatting and cleaning adjustments without modifying the original content.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Only perform formatting and cleaning adjustments without modifying the original content

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.345656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.349954Z digest=sha256:13f0baa9ae19cc0bb08479d1ced9da767925fe5072a777794467a2bbbab0e127

Observation 3b025b0b-76e8-4009-a527-6f4414c86a56 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.993742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.419255Z digest=sha256:0064e406c84740d1deff1a74777c205e0cbb8ce0c80f6b20996883bd4ba7f4a9

Observation fe1a29ce-ff78-487b-b7e8-383fd5d81997 · outbound

This paper cites - **Prohibited Content**: Any direct diagnostic statements involving disease names.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - **Prohibited Content**: Any direct diagnostic statements involving disease names

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:31.554814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.491524Z digest=sha256:57499461960d53198d69d45db6211da3371f5fed66fe0935dcd4a96a1b64b5c0

Observation cc00d117-d35b-460e-9deb-1529740beb7b · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.228417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.553291Z digest=sha256:f16a50b43189630537ab3b8aac6b72de6f358a32b74c1ce795a971a9e642cd56

Observation 7ec7ea28-49a4-404e-84cf-37fc9aa3e8ac · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.015175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.614692Z digest=sha256:3935812240d22b3489c4211eed7e90daa606b23bdd78bf932c9e4c367ae3d6ab

Observation 8ca475ba-f535-4936-94d2-b5a53b8afc8b · outbound

This paper cites DiagnosisArena-MCQ Evaluation Prompt You are an expert in the field of rare diseases.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models DiagnosisArena-MCQ Evaluation Prompt You are an expert in the field of rare diseases

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:30.685315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.692156Z digest=sha256:efc07a16224c2990499b78c06121bb9018162bfde2bcd5e38d2e3c2ce153f1bb

Observation a19227c1-4dd3-450b-8594-6b35a146e585 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 58

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:30.306619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.769469Z digest=sha256:1e14c05eb407ff443f2f030508edc01249d9cb92270720a1cc9cd20c54be28b2

Observation 0ae94c43-e39a-435b-b59a-1fd19ee19465 · outbound

This paper cites C Cases C.1 Benchmark Cases In this case, Case Information, Physical Examination, and Diagnostic Tests are components of the patient’s medical record.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C Cases C.1 Benchmark Cases In this case, Case Information, Physical Examination, and Diagnostic Tests are components of the patient’s medical record

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:30.090590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.833573Z digest=sha256:60e6c2b29699ecc6b1c23a687385d6b858dc192ee597b9380e130e47f03d8c12

Observation 83fdf6e5-502b-413a-8de7-2fbffa51c1bf · outbound

This paper cites Highly mobile, filamentous, causing embolic symptoms or arrhythmias.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Highly mobile, filamentous, causing embolic symptoms or arrhythmias

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.800933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.898275Z digest=sha256:fd3d3e95ffae3ed9bd81baad745bd08f3de5aaaf111cfa1258a144cf9cb77b88

Observation 406b80e0-21cc-494b-a5d9-0dee71f9d2ab · outbound

This paper cites Could be in LVOT, leading to similar symptoms.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Could be in LVOT, leading to similar symptoms

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.553124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:25.969036Z digest=sha256:fe646d8874c69e157c72b96e4e3c5e854274c021daf7989e6222e26d2e75c58d

Observation dda23e77-1902-4c30-988e-5ec19956f861 · outbound

This paper cites Mobile and can cause obstruction.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mobile and can cause obstruction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.300596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:26.033021Z digest=sha256:9bec7ee1f7fbfb52384d35d49887a01751a47780f2158e7118017333e1a31b0e

Observation a5b39362-9d2f-44e3-8ac8-e07a2071d6d2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:28.890974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:26.090955Z digest=sha256:d4a940d119022346cc9160480815364b830ca4d3d11600901b7b2fd804803ee9

Observation cac5c5c8-f692-4bac-abe8-81a1ebf3cb64 · outbound

This paper cites Wait, but the movement pattern and CT findings might make fibroelastoma more likely than Lambl’s.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Wait, but the movement pattern and CT findings might make fibroelastoma more likely than Lambl’s

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:28.599013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:26.152160Z digest=sha256:8720017fc0d6ee5f805c67894eec31f17663e098fb0acd1a89db2a65c264fb16

Observation 9b4dd91b-2b5a-48e0-a433-31f02bff4043 · outbound

This paper cites The key is the mobile mass in LVOT causing possible embolic events (leading to AFib) or obstruction.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models The key is the mobile mass in LVOT causing possible embolic events (leading to AFib) or obstruction

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:27.907921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:26.229267Z digest=sha256:fdbefb7ed497a919b98581713b75a2e585b61c82fc4e858b2254359b0e9e6b7f

Observation d33dae21-4284-4cfc-8e96-560b5b0f64f2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 67

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:28.214479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:26.328940Z digest=sha256:1698dba17156f294260af432c87547aeb16695ff671bdf22d2a32580a1570933

Observation 5ffe65f1-7028-494e-b52a-c5be7314a30a · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:27.623039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:26.489072Z digest=sha256:33b80761813ebee07f078d893088aeb2d177af24188218e147df63632709a65f

Observation bde34791-774d-430d-b597-946b2915abad · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.630549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.630549Z digest=sha256:c28be6e32364a703f397d4bc521ad6b73d7da02ec01f0080a21c8c5065aae4ca

Observation 5ef940ea-ed06-44c9-ae68-bc31d44abad0 · outbound

This paper cites Advances in neural information processing systems, 35:24824–24837.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Advances in neural information processing systems, 35:24824–24837

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.810735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:42:24.508298Z digest=sha256:7d9d31caf34a3d47f63065b53557a6ae841b6f6356282d77b64961c892222da1

Observation 4c32e658-c443-46e1-aeed-e7fb5d426ae3 · outbound

This paper cites Baichuan-M1: Pushing the Medical Capability of Large Language Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Baichuan-M1: Pushing the Medical Capability of Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.011816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.011816Z digest=sha256:8b0a237aa20c757c9a4334aad42ed244d683375e3acc41fe62a4d4d26ff175b8

Pith citing papers

Observation 09047aae-f843-40e3-85b4-cf933f609390 · inbound

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks cites this paper.

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:31:25.606181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T13:26:58.309566Z digest=sha256:e007c59647c4813e1bda9f7fba240751407ec57140c8f5896eadc78972225876

Observation b7ec127c-9cc0-44a7-94c9-86527838c0db · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:56:50.369713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:92ea0db9befd09aca3885feb9761777840aad4cd009a2d1d76b9fc7cbeebdf40

Observation 56606448-0adb-4b38-bf1b-6c8a2b8894c4 · inbound

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning cites this paper.

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:50.561904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:19:15.053725Z digest=sha256:fc4bb95bdb14a2fafb81147cf0a75672d55a072731ffe555b579fabd3e2f71e1

Observation eb509866-ada7-44b7-a024-1d689965352d · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:37.715348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:03:34.773345Z digest=sha256:eaf34972bed02b60a6b37f6759807756b371a2a38b04e93b3dac208bfa8dd76b

Observation 4c17d835-2f8e-4172-ae51-856250113229 · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:27:59.758936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:25:31.662152Z digest=sha256:726d372ae19379cdb7d0ae288c28fa4f6e1cbe0fdd8affa1fb5db019a8a72367

Observation 7eb391fb-79aa-4e59-ba84-88b6f2ba5083 · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:43:10.455707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:a29a944aa0ca690bc8b3e731a5a1b05602793fbd8e0d7cc7ef0a2537ee1e942c

Observation 96c78045-c903-4f3b-afa0-f6f2d46e03af · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.921236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:4eb2ea334cf6eb26ad029535d369acb437881c4986837dc2da32ac28e8fc1bf8

Observation d98ffec1-12da-4a37-ab73-142cbea8de42 · inbound

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation cites this paper.

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.624793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T19:21:44.653877Z digest=sha256:f07749ccd23907111e40258ad4f843d2b5e917431887363626f76629bd2fac0d

Observation fbbe7ad0-62de-42bd-a871-203c3e9afd2d · inbound

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs cites this paper.

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:38:11.263118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:38:11.263118Z digest=sha256:18305c5de64f867829a9e10ccc894fbb8cae3cdc77249dc0f14677221074ac56

Observation 3530e665-cd62-480a-96b0-5471f5a47e1d · inbound

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases cites this paper.

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:59.356644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:07:59.356644Z digest=sha256:b1d95f0803cb638dc45fcca33bd262ffc946a053bd661ede5e1bdc9080e2af7d