Pith. sign in

Paper Citation Record · LEDGER

LLM Sensitivity Evaluation Framework for Clinical Diagnosis

As of 19 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2504.13475.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13475 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:10:35.941849Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b9d883c-279a-4465-8870-100ee51bf74b · outbound

This paper cites Large Language Models Are State-of-the-Art Evaluators of Translation Quality.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Large Language Models Are State-of-the-Art Evaluators of Translation Quality

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.896958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.896958Z digest=sha256:d0b4a49a5ecdb73ed9253b6765ea047a2246d26b97c9fbc0bd63ce4f38c0f47a

Observation 8835d984-0a6f-499d-93c1-1cb5cab1c9a3 · outbound

This paper cites Large Language Models Understand and Can be Enhanced by Emotional Stimuli.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Large Language Models Understand and Can be Enhanced by Emotional Stimuli

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.902565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.902565Z digest=sha256:0cb76c5bcd00ce78d6b5919f0282ab20b16efcb96967ef50092c663101249db0

Observation f8fc6580-3551-410a-859a-a66eb33dba92 · outbound

This paper cites GPT-4 Technical Report.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.907807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.907807Z digest=sha256:3cfba15423353fc37c478e195a550afdfe422e2fe7c9045ef6b992eb2af1f677

Observation bf92fdf1-6516-4668-aff5-daa45834ef69 · outbound

This paper cites Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.918172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.918172Z digest=sha256:76d1af8cbdf3e8ec885ec08b9f31147afee9f62afa26cfde1f8a701f0b6b1f90

Observation d301c822-3210-47fa-8aeb-80c9c8dd3cb0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.922951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.922951Z digest=sha256:5064d6ef2d98f4a0c5fd2e189d9bbd1e6887281cf8bb7a0f0b562f3f95054b76

Observation 47840c07-0c71-44c8-9abe-b9e03cdc7329 · outbound

This paper cites The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.927476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.927476Z digest=sha256:4a506922425e39be63a10be91153454bcce0df381cfe54845fc93be46094562d

Observation 0fc91d06-11cc-44f1-ba2b-2dc64ddfc8bd · outbound

This paper cites Enhancing Small Medical Learners with Privacy-preserving Contextual Prompting.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Enhancing Small Medical Learners with Privacy-preserving Contextual Prompting

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T12:10:36.020780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:10:35.932098Z digest=sha256:7a86f9ff499361084b67209998189055b682a863bdbf85811ede8d2e06142fd0

Observation 04aa36d6-dadf-418f-adf7-55bfaf4a8f37 · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Large Language Models Are Not Robust Multiple Choice Selectors

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.937101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.937101Z digest=sha256:2f4450fa699edef5dfac5bf8c715430437e559b038e0a32802194315fe61f02d

Observation ad76e021-1aea-42e8-8e20-4a992a0b2f88 · outbound

This paper cites MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.941849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.941849Z digest=sha256:00ba82c4ab80093116e6ae32cc3a7c744d326400a914e8bc34c1a24b1a998d0f

Observation 1706e4f7-6afc-40ad-b785-950f7bf90a1a · outbound

This paper cites Language Models are Few-Shot Learners.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.880103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.880103Z digest=sha256:b4bde8d2b501b5855787e5ea9d7de22d59f1b9bd24b2c7deec4249e406dc41d3

Observation ad5baca5-e8af-4451-9bbd-b265eb7ef7ad · outbound

This paper cites Training language models to follow instructions with human feedback.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.912710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.912710Z digest=sha256:17dc38a73fe71eb6b0f99ba298c86cbdedbf995de2c69acb49a066478f63fb7e

Observation 3337559f-3d8d-4f38-aa07-03494a23c98b · outbound

This paper cites MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.891660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.891660Z digest=sha256:07609cb4973262f9e4f4beaaf7affb7aefcb84b7f6209e4b3af74393812e4b91

Observation c398c02a-4530-4f0f-b024-57d4ca3cf920 · outbound

This paper cites Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4.

LLM Sensitivity Evaluation Framework for Clinical Diagnosis Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:35.886311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:10:35.886311Z digest=sha256:1673fca7c87f3cb1705f4d036bf31bb9befe235b739e9c1861eb2dba4a66b550

Pith citing papers

No inbound Pith citation observations are available.