Pith. sign in

Paper Citation Record · LEDGER

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

As of 12 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 9 inbound Pith citation observations for arXiv:2508.04325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04325 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T00:40:34.440305Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:57:40.893213Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy43
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1882affb-a6d0-438d-9f31-9e717c4cdad5 · outbound

This paper cites Extending Internet Access Over LoRa for Internet of Things and Critical Applications.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Extending Internet Access Over LoRa for Internet of Things and Critical Applications

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T07:21:27.027337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:74dff223456221cee082bd6d6b904265c5989d8f9dce449dab0c72eccad31aec

Observation 6a12dde6-efe5-4dfd-a678-8026e8bc1c33 · outbound

This paper cites Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:e60cd5fdc959b4cdd86bf3af0a2ed3db0387919692d1ea79be71a4ce40ee7218

Observation c83eae94-12c4-46ae-9aec-cc52f427a514 · outbound

This paper cites A Survey on Data Contamination for Large Language Models.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models A Survey on Data Contamination for Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T07:21:27.058821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:e5305a6924c67b89b17cf709707fd849af84b9947889bb20988eb2a8a7d81ce6

Observation 967d7881-6b0a-438b-b28a-027b74ebfe1d · outbound

This paper cites ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:45:57.533893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:4398b9b3f4a0a8d1ff2b75186da96758181cc6f87095c3583c2e770547301e83

Observation c475641c-7c69-4ecf-aa9a-02a0b10a969f · outbound

This paper cites COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T07:21:27.071940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:b9fe5215f260f443d67029954c644e24a3d7c1632557d9369cba77b2488eb4cd

Observation 58f1b971-5115-4972-88cc-187979c2b949 · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:53:19.668674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:c682581dc32c30b0f721cedc118dab153691e05fa2e0856d824a4b2a04829f83

Observation cc344051-3cd2-4e62-9f3f-9bff7da01e77 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:37:42.029646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:3d8c45a7cd51b2ab5be9bc8f02bf199e4d3bfd3e47d1b97b34dde9d249666ce5

Observation 265a1e35-1630-4cbf-82af-bbea73f97d58 · outbound

This paper cites Ndepartment de- notes the number of medical departments, typi- cally referring to the medical specialties included in the model’s evaluation.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Ndepartment de- notes the number of medical departments, typi- cally referring to the medical specialties included in the model’s evaluation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.098121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:e3c3f0386cadb7b2a530e41efad270c20be21b36f35b42c19f1b61720b737fe9

Observation 254418af-93aa-4d7c-8828-d326e85dc3c7 · outbound

This paper cites an unresolved cited work.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-09T07:21:27.102979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:30fb9fb40cb238f431de3d3d5ee86ecc5cb5267c98cd5d07dc56eff012dce5d6

Observation 46a46edf-669b-4cf1-8613-c7a638dc6533 · outbound

This paper cites The remaining 72% (38 out of 53) did not conduct in- ternal consistency assessments.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models The remaining 72% (38 out of 53) did not conduct in- ternal consistency assessments

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.109994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:4419cd6ce330349459a98102112b409467e0cc9bfebc8c7aab40f78f15a6de33

Observation 2f15c379-6bad-4455-a54f-306ba260ab66 · outbound

This paper cites These findings highlight a lack of rigorous statistical stan- dards in current benchmark design.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models These findings highlight a lack of rigorous statistical stan- dards in current benchmark design

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.117463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:47855acbd3279c72351ed2e7e161e63bebc63b04b251476fddbbdd03a3c436bc

Observation cd55380f-431d-4d8f-8da6-218be68a6ebd · outbound

This paper cites • Justification: Clearly defined evaluation objectives can avoid ambiguity, facili- tating subsequent data collection, task design, and metric selection.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clearly defined evaluation objectives can avoid ambiguity, facili- tating subsequent data collection, task design, and metric selection

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-05-09T07:21:27.121027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:8760c715e8869b3d41358c84aad2c560f43224ab31ff2edf14ee8103580b6b03

Observation 6e578f29-ce02-4261-b9ba-924cee9406f3 · outbound

This paper cites • Justification: Linking the benchmark to real-world application scenarios en- sures that the results are meaningful.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Linking the benchmark to real-world application scenarios en- sures that the results are meaningful

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.124399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:711dabd3ea140149774b420128d67ffc99e353b6e385e257c36e1862465c1463

Observation d9c5e1c6-3b52-4773-8332-3a9af8507bd7 · outbound

This paper cites • Justification: Demonstrating the unique- ness demonstrates the necessity and jus- tification of the new benchmark.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Demonstrating the unique- ness demonstrates the necessity and jus- tification of the new benchmark

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.132540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:deb3db9f804bc65dbdd307c696f3bc46bb4ac65b2c8d9be36b2a0c4153a124ef

Observation 8612c382-dbd2-4134-ab22-61234d9a245c · outbound

This paper cites • Scoring: – 0: Does not define the target LLM capability.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Scoring: – 0: Does not define the target LLM capability

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.135286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:c9433948c911f48a792c0894d51eb5bac955b6b466b96cbe6d1fc1198a0db4bd

Observation 09c00e7b-7395-4424-b2e6-54d32ef3e37e · outbound

This paper cites • Justification: By clearly defining the medical scope, it helps users better un- derstand the breath and depth of the cov- erage of the benchmark.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: By clearly defining the medical scope, it helps users better un- derstand the breath and depth of the cov- erage of the benchmark

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.138820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:a113bdf0c9aa99db66ff00c7684b1d1b641a7a6447ae28f76534d8dca73885d3

Observation 115ae008-ebe8-4610-ac17-0acbdd4448d7 · outbound

This paper cites • Justification: An effective benchmark should serve the needs of users.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: An effective benchmark should serve the needs of users

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.143822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:91156bdb7537ec2f91473ce38bacf307bebf5560bb390151eaf53d8854a51595

Observation afa28656-4f21-458a-9a2f-5661ac20abdd · outbound

This paper cites • Justification: Due to the professionalism and rigor required in the medical field, the development of a benchmark must involve deep engagement from domain experts.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Due to the professionalism and rigor required in the medical field, the development of a benchmark must involve deep engagement from domain experts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.146898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:4c45c9fb879d019551d911e3e05267799d5891c0f2097582ce55d18f52e4f837

Observation af582f06-69b6-443b-bfa3-51f07b205f63 · outbound

This paper cites • Justification: Benchmark content should adhere to recognized, evidence- based medical knowledge sources.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Benchmark content should adhere to recognized, evidence- based medical knowledge sources

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.149259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:b1707f1aa4ff0b34769e287fb8924acc130f7b8e4013973bcd1f165950f4af5b

Observation 54b06b0b-bfd1-43e8-9dfc-3a22d65cfb31 · outbound

This paper cites • Justification: Adherence to medical standards ensures the clinical relevance and consistency, facilitating integration and comparison in reflect real-world medical practice.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Adherence to medical standards ensures the clinical relevance and consistency, facilitating integration and comparison in reflect real-world medical practice

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.154387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:bf3c50222dcb4305792dab650dd95003ad554a515ca015dd41b2f108d447629f

Observation 2a0d267c-a790-4a4e-8775-45d0295c4742 · outbound

This paper cites • Justification: Evaluation metrics di- rectly shape the interpretation of results.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Evaluation metrics di- rectly shape the interpretation of results

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.159888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:905835e6e701720a2e22373528e6879a64cbc8c9461505edaa57ac91e42992af

Observation 6d837b2a-c9ec-41d3-9b41-993b9ea59a8c · outbound

This paper cites • Justification: In high-stakes medical do- main, going beyond correctness is vi- tal.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: In high-stakes medical do- main, going beyond correctness is vi- tal

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.162167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:db9332b2ec5b38c1a61e78ce032cc8f59ee698e88a379979c1bbd625f2ea88b5

Observation 90ff61c2-0a41-4b02-ae99-c208046edd83 · outbound

This paper cites an unresolved cited work.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-09T07:21:27.164379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:eef093036a6448ae8e4688f674dd728a751ddf8303451173181325af96b33604

Observation a4f054ea-2641-43f2-ad86-1ad89a34ad89 · outbound

This paper cites • Justification: Clear and traceable data sources are critical to ensure trans- parency and ethical data usage, which is especially important when sensitive data is involved.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clear and traceable data sources are critical to ensure trans- parency and ethical data usage, which is especially important when sensitive data is involved

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.169014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:80271196b8eaf887a1816efb5656686f64f37ae8b0bae9a720983b15a936b781

Observation 2cd13a52-7c2e-44bf-8c0e-b6a43ef6218b · outbound

This paper cites • Justification: Collecting data from unre- liable sources may lead to invalid results for medical applications.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Collecting data from unre- liable sources may lead to invalid results for medical applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.171308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:6f447c52afe32cbdd76cff7ef0453f9f62567ad2eccec77819c4cbd13bd378df

Observation 26248f49-8d28-4ac9-83b7-388909d07d43 · outbound

This paper cites For synthetically generated data, the con- struction process and verification for its authenticity (e.g., expert review) should also be described.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models For synthetically generated data, the con- struction process and verification for its authenticity (e.g., expert review) should also be described

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.173632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:2708e7eb53db82b4be0a578ea5bfe1430046b4e6fa4e5fe165929c46b8ebce2f

Observation da50aa52-0560-44e8-8d6f-a1ac0c9946f0 · outbound

This paper cites • Justification: A benchmark that lacks representativeness may lead to bias in evaluation, reducing the clinical rele- vance, generalizability and fairness of the results.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A benchmark that lacks representativeness may lead to bias in evaluation, reducing the clinical rele- vance, generalizability and fairness of the results

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.178837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:91e19e31ec04fb561c630085167f1353f485a031a59a8f037c1cb9c033f6de2f

Observation 5994d96b-e7f1-4614-928e-423073eec76f · outbound

This paper cites • Justification: Ensuring the dataset cov- ers a variety helps comprehensively eval- uate the model’s generalization ability, reducing bias in the evaluation results.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Ensuring the dataset cov- ers a variety helps comprehensively eval- uate the model’s generalization ability, reducing bias in the evaluation results

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.182725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:82eec1bb36a5ca84ccdceaf38a5fc884358dd12e43fcfd9acaf07492d2df648b

Observation e42c3370-851a-4a00-b9b6-417e3ba13e45 · outbound

This paper cites • Justification: It ensures that the final dataset is well-structured, enhancing re- liability.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: It ensures that the final dataset is well-structured, enhancing re- liability

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.184919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:8b9a6fd696f28bc394b42d645f7ee098107b6e5a988993c543e411f7347b4973

Observation aa223e34-7247-4c85-95a6-cfdd5525fc8c · outbound

This paper cites Methods of de-identification should be described and compliance with relevant regulations (e.g., HIPAA) should be clearly stated.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Methods of de-identification should be described and compliance with relevant regulations (e.g., HIPAA) should be clearly stated

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.203571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:f6943cdea2152932807940d23ba2848c6f5b76ad2b9d6717f55d856229dce280

Observation 50fc5e83-57f2-4993-9827-6357b3053bb2 · outbound

This paper cites • Justification: A clear and consistent for- mat is essential for standardized evalua- tion.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A clear and consistent for- mat is essential for standardized evalua- tion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.217994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:5b2312057441e0414e242e75901966edcc2b179495eb811511de1bfe5d0ba9c2

Observation 4505880c-2f99-471d-812f-24e410f96e33 · outbound

This paper cites • Justification: A dataset construction pro- cess without review mechanism is prone to errors.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A dataset construction pro- cess without review mechanism is prone to errors

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.220262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:22110eb2b980a8f95e7c1baf609b14155282fcffd1b9a09a0407e2655f8eeeff

Observation bea2e980-b546-4663-b09a-922cebbd5a73 · outbound

This paper cites • Justification: Clear reference answers or scoring guidelines ensures transpar- ent and accurate evaluation.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clear reference answers or scoring guidelines ensures transpar- ent and accurate evaluation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.222414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:4d5e259e9d0e6f86413a8abf24c27993a6c10dbcf9eac07e06e106d94ae63c97

Observation a75d24a2-b9f0-46d3-a2a4-eea584efc0c5 · outbound

This paper cites • Justification: Data contamination may lead to inflated performance, which only reflect memorization instead of medical capability from the models.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Data contamination may lead to inflated performance, which only reflect memorization instead of medical capability from the models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.224666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:6a442c456821e884a93e2902239092c250793f3fa674d8de9d9562211e98466f

Observation b31b3c6d-4ae9-4238-bf66-8ce3282ff151 · outbound

This paper cites • Justification: It ensures that users can use the benchmark conveniently, promot- ing benchmark adoption and ensuring fair, transparent, and consistent evalua- tion.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: It ensures that users can use the benchmark conveniently, promot- ing benchmark adoption and ensuring fair, transparent, and consistent evalua- tion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.248332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:48081cc46f4778bcb1e4ffce99200468c4437e0d02047520ccfb87c68320065e

Observation 9e5e6522-2d9a-4b6f-928b-c0510cbf89a5 · outbound

This paper cites an unresolved cited work.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-09T07:21:27.253645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:24a93e067a3314b62aec77fcca3a09fb4642fa8318a89fe4fc81d4beba696810

Observation 90d60789-305f-4f8a-8114-8bcee9f2df84 · outbound

This paper cites • Justification: Providing different per- formance baselines allows comparison against the model’s performance, en- abling a deeper understanding and better interpretability.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Providing different per- formance baselines allows comparison against the model’s performance, en- abling a deeper understanding and better interpretability

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.258693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:af270c80d42d01b212382a8a1f897d64820bb0c5d5c753828fe652585b685262

Observation 0b27be8a-1277-44db-9437-4516e2592798 · outbound

This paper cites • Justification: In medical domain, un- derstanding the model’s decision-making process is just as important as the final answer.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: In medical domain, un- derstanding the model’s decision-making process is just as important as the final answer

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.260885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:5014921d2605c3adb13ca606334d78aa28f37e715a9092dc00d772733c39d180

Observation 8c4670c0-ee98-4236-a665-adff6ff18ef4 · outbound

This paper cites Robustness testing en- sures model’s output is consistent and reliable under different conditions.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Robustness testing en- sures model’s output is consistent and reliable under different conditions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.262979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:9a91f84d37dd6e4e387c1b4bfd0748972d5cbb1159bc538ed3c36c36aaf48917

Observation 4e559e94-892d-4108-9972-9c703ea4c069 · outbound

This paper cites an unresolved cited work.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-09T07:21:27.264831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:b3ed838c2023242765b4fe1f449e566aea5d3f6217683b3c1ce2f3f8af6e9991

Observation a87bf755-a425-4f44-bc89-28e74f1ad7f0 · outbound

This paper cites I don’t know.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models I don’t know

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.266632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:095b0ee8f1f57d0e1d8700b43bb4c8dbd23fbacda73f8a9eefcce50335d37eb8

Observation 75f581da-0c51-41ed-9e84-e3cb2fb9927b · outbound

This paper cites • Justification: It ensures that different types of models can be tested under the same interface.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: It ensures that different types of models can be tested under the same interface

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.268513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:14ad3ddf17baf08d0b05e461a8c215e0ec43f5610473287132c579fbaddd0d37

Observation 6938e0c3-ce32-4bac-9b7d-17a2803e2ef7 · outbound

This paper cites • Justification: Sufficient coverage of the core medical competencies the bench- mark aims to measure is the prerequisite for establishing content validity.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Sufficient coverage of the core medical competencies the bench- mark aims to measure is the prerequisite for establishing content validity

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.274950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:bbb2f26c1c386d10faf9ac4a27254abec6f088ec864afff92d69f8c7f9b204f0

Observation fe874795-7e6f-4106-919b-ae7a965300c3 · outbound

This paper cites • Justification: Ensuring that the evalu- ation task closely mirrors the targeted clinical practice in real-world scenarios enhance the relevance of the benchmark.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Ensuring that the evalu- ation task closely mirrors the targeted clinical practice in real-world scenarios enhance the relevance of the benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.284769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:61d94a46c20b01c565709ed1ec74f456a8d66ddd7870133c818119b4a8010747

Observation 237b8ce0-284e-4b9a-86ec-b3ae24123eb9 · outbound

This paper cites • Justification: An effective benchmark should be capable of differentiating and distinguishing models of varying capabil- ities.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: An effective benchmark should be capable of differentiating and distinguishing models of varying capabil- ities

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.287397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:c0fbe9bf0b1a5f7a4d76bc1408de2fcdeb28a55e505097649ff4b66b9af7b743

Observation 448593d2-abb6-4985-9034-af5d5e4f029b · outbound

This paper cites an unresolved cited work.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-09T07:21:27.289251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:86bc120154c8f4ec7fedc6f95f46397845e272d6dc81baffa35c4ae65bd38d32

Observation 07a93eb3-cf09-4d01-bfea-f096a83fd057 · outbound

This paper cites • Scoring: – 0: Does not mention or conduct any internal consistency measurement.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Scoring: – 0: Does not mention or conduct any internal consistency measurement

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.291391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:2441abdacf7e859b9d0a64062807b6988a64e3df8370809e600b381c64d126f4

Observation 57553101-f584-42f9-a6cb-227bb695d59f · outbound

This paper cites an unresolved cited work.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-09T07:21:27.293226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:d4b3a2441ff06de7740127a4c991ea514040724cc152ebe168011cf2142d8aeb

Observation 1f8582eb-a03b-49b5-9826-2d23b583ec2a · outbound

This paper cites • Justification: A complete and clear doc- umentation help users understand and use the benchmark properly, enhancing the usability, reproducibility and trans- parency.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A complete and clear doc- umentation help users understand and use the benchmark properly, enhancing the usability, reproducibility and trans- parency

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.295162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:c9a580fcf4416964f77d99bb7aa02f91e56fa93b1b2228f80cdbd0eaef087ebf

Observation e086260e-9abc-4c22-bbba-40f112dbed87 · outbound

This paper cites • Justification: Clear evaluation guide- lines help users better understand how model performance is quantified, ensur- ing a shared interpretation and under- standing.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clear evaluation guide- lines help users better understand how model performance is quantified, ensur- ing a shared interpretation and under- standing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.308599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:092a5e92287869c1799b794817497c4ed730c68703498a6a1c5b4b4544136666

Observation c6186465-86d9-4fe7-8d99-e82efc9804e6 · outbound

This paper cites • Justification: Disclosing limitations and risks demonstrates scientific rigor and responsibility.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Disclosing limitations and risks demonstrates scientific rigor and responsibility

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.313469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:cc392b134cc92229146c3d5bec4d36ff041b9b862686c04830af87c3b868b35c

Observation 7322dbd1-2315-4f61-981c-143649b8df95 · outbound

This paper cites • Justification: Going through peer review process means that the design, validity and results of a benchmark has been rig- orous evaluated.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Going through peer review process means that the design, validity and results of a benchmark has been rig- orous evaluated

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.315887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:61e637e2769a3fec8436619db38d2c4fed7af951b0addcd6d7e12a50b1bc6987

Observation 6d37a8ac-628e-4a4c-8b71-b39960abd0b8 · outbound

This paper cites on platforms like GitHub or Hugging Face) along with the applica- ble license.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models on platforms like GitHub or Hugging Face) along with the applica- ble license

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.317975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:e6a811f148eb7a477b7884e4ce12c4b1eb6dd780de32564406ea74dc4f3342e8

Observation b1a4caca-eddf-4e7d-9e5e-3b183afb2fb7 · outbound

This paper cites • Justification: Proper usage and citation guidelines help maintain academic in- tegrity.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Proper usage and citation guidelines help maintain academic in- tegrity

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.319871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:0c6542677ae566831e906eb6f04508fbea49a84a4ba6696baed52935ebbb7409

Observation 9e4bcdd5-65eb-4968-8231-a82ffecb67bf · outbound

This paper cites • Justification: Maintaining an effective feedback channel allows users to provide feedback when issues with the bench- mark are discovered.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Maintaining an effective feedback channel allows users to provide feedback when issues with the bench- mark are discovered

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.321986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:e3a41925b175c1264d63eb66c62d3d5ae8d46783f4e345064b87d1094ecea78e

Observation 42ffb1e8-6d51-4cbe-beb4-5696a4053aa4 · outbound

This paper cites • Justification: Maintaining an effective feedback channel allows users to provide feedback when issues with the bench- mark are discovered.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Maintaining an effective feedback channel allows users to provide feedback when issues with the bench- mark are discovered

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.324007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:0f2495cb509e6416c6936af647b185de44e8a2a55d103e4309938e56aeefcc5b

Observation c8c03501-839a-4f87-beea-c2fbe5820cc7 · outbound

This paper cites • Justification: Clarifying who holds long- term responsibility reassures the commu- nity that it will be actively supported and improved, ensuring usability and credi- bility.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clarifying who holds long- term responsibility reassures the commu- nity that it will be actively supported and improved, ensuring usability and credi- bility

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T07:21:27.328464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:6490216af6ad71ca0710ea91055f5eb92860ce103cca89d7185d66cd6a6c9dd1

Pith citing papers

Observation b1117514-0ac2-434e-b5a5-fdffb5bf688e · inbound

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction cites this paper.

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-08T17:23:40.483260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T17:22:17.266351Z digest=sha256:ec454a14b58cc91fba8e4656857f077d91fc3430ac35f4fccccf591550e06a51

Observation 2d9fc485-8d1a-47c7-a81e-953ad57fe0a4 · inbound

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction cites this paper.

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.474464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T23:55:22.768635Z digest=sha256:d14e5738c47ef35609027777dd26a0fe34482cebdf7705e1445390b402a948d2

Observation a063ba73-4ac4-4b98-a403-8fb3e350d976 · inbound

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare cites this paper.

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:56:36.315888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:26:28.419570Z digest=sha256:8da36ca93503c6c04ab3449a406d13700e80e1ddb6986831157a7779e321df4b

Observation b64375e8-78c7-4159-b8ff-7db14240cf13 · inbound

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context cites this paper.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:40:47.697055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:4f0951d7dbea7248d0a9ce20d002eee03603429b6d769fe8aa1b6b0e3508d63b

Observation f54b1aa8-b540-408b-abe8-e49fdb511d9f · inbound

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models cites this paper.

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:29:56.922691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T00:40:26.700516Z digest=sha256:7defd336f83c90b4239452f7d8903d01292660e5cb52fc986408210767498a48

Observation 2b6f817b-7124-4c0f-948e-1322b3e4f9e7 · inbound

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models cites this paper.

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:59:52.284424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T05:42:11.764973Z digest=sha256:b1755051e6a0ab345590afad1230fa2f5a6ade90d86387b4fb0d76d604598d3d

Observation 4772a3c0-f726-4610-868e-a35c4e2e6e0f · inbound

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models cites this paper.

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T12:36:00.189291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:36:00.189291Z digest=sha256:2de252e9dcad2aa83841bae6806277b6015e3b3a5917fd8c6e8d03c961027c40

Observation a0b7158f-dffb-4080-8b8c-4a42b6591ebf · inbound

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries cites this paper.

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:04:37.490635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-30T11:03:52.230759Z digest=sha256:084fd2fef48388e7d1130cf6967bd3935ebc2fc58253dfa926e761edcbbbf7ae

Observation cce5bdbf-ff82-4873-b6f3-3f0aba30d6eb · inbound

LLM-Assisted Review Prioritization for German Statutory Health Insurance Websites: A Multi-Stage Corpus Audit cites this paper.

LLM-Assisted Review Prioritization for German Statutory Health Insurance Websites: A Multi-Stage Corpus Audit Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:40.893213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:57:40.893213Z digest=sha256:610c39110713dc862867e2add4dfd2c334dfa8ca988708b5f6d05abfd4280f0a