Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:21:49.124883Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 3 inbound Pith citation observations for arXiv:2509.02594.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:21:49.124883Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T13:13:26.156464Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
20 of 20 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation d991f462-9a44-4780-8137-5e8a7f5e21eb · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries A systematic review of large language model (LLM) evaluations in clinical medicine
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b79ed5a8-d09e-40b7-aef8-362eb4c21ff7 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Scientific Evidence for Clinical Text Summarization Using Large Language Models: Scoping Review
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290ee266-ba18-49c9-abab-4ea3e5857813 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries AI Consult: Real-World Evaluation of Large Language Model-Based Clinical Decision Support in Primary Care
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 33444acb-a9b2-424e-a8c7-e527cfae02e1 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Capabilities of GPT-4 on Medical Challenge Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b98faa-fdc4-4107-9ed2-3157428ed2d3 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Large Language Models Encode Clinical Knowledge
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 829959f1-622c-4ff3-9bdb-b8d9d3171c2e · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physi- cians
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8b0c8d4e-4550-48d2-a392-92ea17dd261d · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries What disease does this patient have? A large-scale open domain question answering dataset from medical exams
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e983042-da58-4197-b2d2-682e58a4a3ca · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f448757c-4fb3-4412-b70c-cc824f6ded0a · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3df4048-aebb-460e-bbe9-103af12f27e8 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Sequential Diagnosis with Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90590ceb-b958-4354-9ca7-36c4b8b5aa1e · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c23685-9364-4969-ba23-1a77bd1666b3 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Dr.INFO: Agentic Clinical Assistant
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9c2cb5cc-37b6-4b91-9024-227d766dda89 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Leveraging long context in retrieval augmented language models for medical question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 45e1f2df-3ae4-4b01-8c72-e8685e9bafae · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries VISTA: A Rubric-based Visual Task Assessment Instruction-Specific Task Assessments
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b1c8718a-f500-42f2-80f9-5492d1609b8e · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd8a3431-2352-4886-9242-4b74c792ef2d · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7133c9-94f5-43dc-822f-5eb75b4da0f0 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Autonomous medical evaluation for guideline adherence of large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57df649d-5060-47bc-9a32-08469a0b455a · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6b30a2-6cc9-4ca2-ab04-21ac5fc7e70c · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Introducing GPT-5
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1a42e871-2b06-46fe-be0b-3a73808e7376 · outbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Unresolved cited work
Reference 295
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e86579b6-042c-4532-831e-1997ffce9ee8 · inbound
DR. INFO at the Point of Care: A Prospective Pilot Study of Physician-Perceived Value of an Agentic AI Clinical Assistant OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2123923a-01e5-4e2e-b43c-1c4145388c74 · inbound
MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 89d6a204-5f3c-47cd-bef8-a5549cc475ba · inbound
Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.