Pith. sign in

Paper Citation Record · LEDGER

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 3 inbound Pith citation observations for arXiv:2502.06666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06666 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:46:29.739206Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:24:43.178158Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:59:38.023398Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3bbbf59a-aebb-4d4b-820c-079cd9576153 · outbound

This paper cites online" 'onlinestring :=.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.518909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.518909Z digest=sha256:e1ef3ed6158d82af5954de99938f90c1128ddf8d7a46a53655f71d97d14aca0a

Observation 850463a7-aaa6-48be-9aa6-1c67d836d434 · outbound

This paper cites write newline.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.524615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.524615Z digest=sha256:dab3577461a5709bf96989fa403b4f9feb13d00c364a2629cf33fc6a00393d3e

Observation 1283dea7-4db4-486d-b74a-b7deb8fde1f0 · outbound

This paper cites Creating Trustworthy LLMs: Dealing with Hallucinations in Healthcare AI.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Creating Trustworthy LLMs: Dealing with Hallucinations in Healthcare AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.530844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.530844Z digest=sha256:5fbfb6ee636dce02a481ad70dbe099c19cee57f1462036aa6969b5c03691e415

Observation 5ccc4589-99dc-4076-aae0-cc38f0092862 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.803045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.536495Z digest=sha256:95a847c89d7f958d20a615e82cb0173b06a6c275e85d486a18f02eeb385fddd1

Observation 6aabf32b-0f38-4f7f-9150-377d5b946014 · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.541762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.541762Z digest=sha256:166b8858a96a37e33803771a1a1dfdb2cc0a8cd1a9574d704bf4e8b7034f744c

Observation 1af37629-8959-4fb4-98c3-8911f50eb717 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.784194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.547732Z digest=sha256:05daf757e8227e9b4748e4fa0427b74703f71f29b594166eb31cc4daae01917c

Observation 3388325f-8e7c-445d-9dc7-db424159e0cb · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.768023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.553071Z digest=sha256:1804830829d39d1d4bbb4d132e365152cbb4950836f09e6cfeed363b3361f871

Observation fd74d032-ca36-478f-a1cf-dca14b31099b · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.750939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.557764Z digest=sha256:14b97f1b3b6e7638a9ee00bc3cdd2458afd4a72b31ac5decf9b4f2f080fb5134

Observation 28edeeb4-31ab-4c99-a644-4632cc037651 · outbound

This paper cites Benchmarking large language models for biomedical natural language processing applications and recommendations.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Benchmarking large language models for biomedical natural language processing applications and recommendations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.562681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.562681Z digest=sha256:0df9617f16e45ee9cf3e8cbf1a612928f0d8698602dd4256b8ac153c0c0e05ca

Observation 7a4a866f-3a30-487f-8de7-f4156a649dae · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.733748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.567264Z digest=sha256:58cb152f7e332a84f8242988683aee22b57f37caa585a7b568e2a08ffcb6d482

Observation 9635c7ba-dacb-4501-b06f-eee813c20d76 · outbound

This paper cites Med42-v2: A Suite of Clinical LLMs.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Med42-v2: A Suite of Clinical LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.571758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.571758Z digest=sha256:05898bc9dee451f33e47d99fc7c26f4fa5cf408d647ac388f89e3ea7098f6c04

Observation 5c711ba1-461b-444e-8ed5-dfe3a5eded22 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.576417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.576417Z digest=sha256:7c87bb650bff5e4faf3e14d6b543e302cefc05568c48d35b4fa1d560b53a03d2

Observation 01300569-716a-4cac-824b-8bd5f3a6a6e7 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.714409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.580866Z digest=sha256:3eeb2092376fe3fe9b7ba04de4b8758d1c428821ef6514b093c382cc6b0351d3

Observation 440f50b2-e9ed-4cfc-b551-49af38c70ebf · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.699053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.585068Z digest=sha256:21a92eb2af7043f7a861f7b319f160611dc4e130f9c1d194146553b4d575e57e

Observation 8e935abc-8196-4705-a420-85dbca0cfb18 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.589392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.589392Z digest=sha256:1aaeac8b00f66a3bf8b8786cdd9719e81f70396ecb7099e35c444931baeb2643

Observation 44131d9a-7334-4e33-8906-33b8c62b6ce9 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.682768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.594081Z digest=sha256:9ffcd980b9af82293962198f3b7de7995ac750f0b5d84ddea966b557790aca3d

Observation 20418767-79d8-435c-b263-446c64329612 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.667239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.598397Z digest=sha256:eaec782dba37949bf188ea0bc1dee561948396d982c5217eab5673144704def6

Observation 07faaac9-8b9b-425d-b429-8a96bcfb4bf2 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.652222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.602500Z digest=sha256:ddace67dc784c4b26f5113e91c689a047ef7cd6cc558b03fd4082a12672d880c

Observation 1cb332e2-6ba8-4ea8-a7fd-0ac708d83133 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.607256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.607256Z digest=sha256:1883950e5da24495b34fd4a1227710349445cd5df175b67521dc3901836a9d4e

Observation 9bd9f22d-b078-43da-9ff9-f952e8c3dc0a · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.627441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.612072Z digest=sha256:49103950eb857a57c5b8ee308b5cc5249ad9e41ef1539d233d345c11b78dd098

Observation 4423f7ba-ae5a-4110-b5eb-b549d5e66654 · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.617054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.617054Z digest=sha256:d6a218d853045c494c72d676610a62162c0214738f79ba19c8df3ff448a2994a

Observation 392f9416-a4a5-4d90-8580-321bc43a5b50 · outbound

This paper cites OLAPH: Improving Factuality in Biomedical Long-form Question Answering.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering OLAPH: Improving Factuality in Biomedical Long-form Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.622140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.622140Z digest=sha256:5b0b1b87b5664c4e14e0a6f069bae32471af5fa333482f3d05b16cc8603a0dd5

Observation 964883e9-d484-40f2-b3b1-59070b34ff6f · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.626947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.626947Z digest=sha256:b633a3ae522626457004948175858d71e5db24c99ca37b8a57e65ae5dfc0448e

Observation 0d0431a9-aa5f-40de-ab3f-5a303d3d5f03 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.602584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.631736Z digest=sha256:fb768ac9af22eda2b0b444ff5738645d580d51a2327eb1123247ce5a9d76f8ae

Observation 3d3cc3b0-46af-4e9a-9014-4905e879e5b6 · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.636932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.636932Z digest=sha256:b684bb2345ac007e47a421e9fbff31f2d2c16d04c6758cf8e4bddc3e2638c6dd

Observation 5904f9a0-4020-4755-8ee1-a246be78e0f0 · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.642901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.642901Z digest=sha256:e8321af504d8a7e6f4bb596d98c23671143721a4f9e56ac64bc2de9c2f707711

Observation 4d478e3f-6162-4489-ac08-11f8185db780 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.649246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.649246Z digest=sha256:c6e1077d520012eec029b288a16e24f8e268f14c4184dadbb0c702a9798dcec2

Observation a1f7d08c-4030-4d6a-8592-d22be6238419 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.654391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.654391Z digest=sha256:69b0af616cd030150a6c6434b75ca62fb659feaaf8b8b73f2e3a7af1c481e3ce

Observation 06638153-b732-4a2b-b7cd-4d6773385d08 · outbound

This paper cites Can multiple-choice questions really be useful in detecting the abilities of LLMs?.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Can multiple-choice questions really be useful in detecting the abilities of LLMs?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.659804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.659804Z digest=sha256:605a4f6ac4f820c029ba57a5546383e71461f2c8559032941b2ca9392da11ff4

Observation 72b9fa63-f5b0-4111-b76f-2de3e5766a7c · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.665017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.665017Z digest=sha256:8a182ad2799ffa5cd83c80d7f86edf4b813c1df2b4445b15fa19033d8cce05d3

Observation 33493c1d-cef7-4509-a595-612aed3e0b27 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.567180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.669943Z digest=sha256:2d92e7f237568ee086112dc7ea61944eb1f8dcfe9d7eb0c60e60aec9f9e4035f

Observation 6d208892-cbb9-4c0b-9ffc-f8b009381b9f · outbound

This paper cites Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.674840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.674840Z digest=sha256:1e0690244a392f34181cf3cd8af3b65185ea53554b3dd642b73ceee7d638774a

Observation bb731652-7e56-4dad-9ed5-5907e3f4c78f · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 33

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T14:46:30.129401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.680052Z digest=sha256:8023d22214fa4fa1f45b15377a7e6b2b13d34d4aa1a2304dfa13b6205019964e

Observation 5acaf2af-1191-46c6-b54f-942db7a0b691 · outbound

This paper cites MedConceptsQA: Open Source Medical Concepts QA Benchmark.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering MedConceptsQA: Open Source Medical Concepts QA Benchmark

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:46:29.912044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.684247Z digest=sha256:35e075350cf97d8623f55f7f03547c776b746f1fd0e9856a3e51dd213850f525

Observation 00f716b3-e7bf-4484-b15f-2057fbb985da · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.688673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.688673Z digest=sha256:0069ca7b39aaaa67262a0f3495d26f94bcbffd24cb868e385ed8f681e50dbed3

Observation 46d4251d-b837-4d1d-b415-b098e7988428 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.693191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.693191Z digest=sha256:63ff36560d5b325854d8c27bda9b1694a2d4cc7c81b00720872de4bbb3e55b03

Observation dba1cb38-769e-414f-9922-53724a2b5405 · outbound

This paper cites Med-HALT: Medical Domain Hallucination Test for Large Language Models.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Med-HALT: Medical Domain Hallucination Test for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.697697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.697697Z digest=sha256:d6a8a99fe7105a8391311e5a88371fe2c93920271497a87db1c468ca276f9902

Observation e1bdcd6b-34f9-44d8-ab92-4e53efafbd26 · outbound

This paper cites Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.702216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.702216Z digest=sha256:28d463d6a32864e4a63d6cce0077803c16b99c34b19febef5d9fd6a5e51235be

Observation b8b97ea0-ae19-4636-9444-e8de7230ceea · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.550064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.706924Z digest=sha256:49906a9f63d5838a0a1780d7013c8416168dc3ba60541b67046c39da5648ec1c

Observation 2dc0e316-f4cc-42ef-acf0-53d56e2330f3 · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Chain-of-Thought Reasoning Without Prompting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.711348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.711348Z digest=sha256:8720244c378f9c3fdecb9ef3313a3fcd37d5ae54167497cab57f4f565f3534ec

Observation 47b87c55-fc5a-49fc-b430-c79d3cf665fb · outbound

This paper cites Qwen2 Technical Report.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Qwen2 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.716103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.716103Z digest=sha256:747c104c8c6833e83c61844320109e04f4609ceef3447d30f73a60b4edb5e152

Observation b92ceb5f-1482-4682-ad61-0c2c8d846ab5 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.533821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.720920Z digest=sha256:8ed6e640678458fdfac7cbe236b22abe90293e05a062eec4b4f12743339094ca

Observation 6d6a70cd-7650-4683-b5ac-fc082ea71c30 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Yi: Open Foundation Models by 01.AI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.725457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.725457Z digest=sha256:8d0f0b12d558d1965a8b27cdd75af910c66362edfb08bc9e6cf5fb42968da890

Observation 82c598b2-ef86-45c9-97a5-b2c12086f074 · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:46:30.517346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:46:29.730073Z digest=sha256:7821a56b495ad001575628f095ef60ce63c9ac813349a20ef31acb337b12c0bc

Observation 158089a1-6082-4287-ab4b-1ad3f226d83f · outbound

This paper cites an unresolved cited work.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.734491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.734491Z digest=sha256:eedaa1650328c0755ba8bf272b9fbc05e62e709d35cd874cc460164baa8da187

Observation cd1c6953-d553-4d06-b0cc-25c632d18d78 · outbound

This paper cites A Survey of Large Language Models in Medicine: Progress, Application, and Challenge.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering A Survey of Large Language Models in Medicine: Progress, Application, and Challenge

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.739206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.739206Z digest=sha256:26963e4b85ebe6ace4ff3e2400a9ffd3667ecc8ae2d915e4481725107cfdae53

Pith citing papers

Observation a7c90f96-cebf-48a7-badc-0e7119573cee · inbound

Essential-Web v1.0: 24T tokens of organized web data cites this paper.

Essential-Web v1.0: 24T tokens of organized web data Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:43.178158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:43.178158Z digest=sha256:2bfd9bde2338446e414d0addc2b236658849cc8633cc0da2464cf3854176eb50

Observation 1c069bbc-05f8-45c3-9434-ccb6027cc855 · inbound

SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model cites this paper.

SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:46.816118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:25:46.816118Z digest=sha256:293ad8c18460e520fc1ca8d691ad376950f4f2155c0e489f6def85464f068eb9

Observation aef5af40-d332-46b6-be89-6e4407f554e4 · inbound

Demystifying Numerical Instability in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL cites this paper.

Demystifying Numerical Instability in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:59:38.025037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T14:58:55.680525Z digest=sha256:09db338a7dbdb463183c706157076368eebc322e05ee5bef91a5fe2df176cca5