Pith. sign in

Paper Citation Record · LEDGER

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

As of 20 August 2026, this Paper Citation Record lists 100 of 151 outbound references and 2 inbound Pith citation observations for arXiv:2506.17163.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17163 v1

Coverage vector

measured 100 of 151 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:15:22.050331Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T19:02:46.991897Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 151 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9e4ffd32-3ebf-4699-982f-5fcfb0c4ded6 · outbound

This paper cites Evaluating large language models on medical evidence summarization.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating large language models on medical evidence summarization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.695152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.695152Z digest=sha256:b133a8c04c19196d79a68528bfb39fd20d56df56494a8820c689ceada7a21f0c

Observation f4150240-e4d1-44a2-90c9-a862884d2a4d · outbound

This paper cites Adapted large language models can outperform medical experts in clinical text summarization.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Adapted large language models can outperform medical experts in clinical text summarization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.699276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.699276Z digest=sha256:a04ffd8013e0edd0de76eafd0db84356a1d445e845b2c111361c4c783f83e5ce

Observation bf3bde49-59c2-4ada-89d2-7e39b1f01e08 · outbound

This paper cites Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.703146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.703146Z digest=sha256:c9aa406b018d0eac03380591872a8f44fd894d292b09ff5d1aa9ca2f1af56b2e

Observation 2acbc252-7251-4b32-9fe8-e52d9b257a0b · outbound

This paper cites Llm-based agentic systems in medicine and healthcare.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Llm-based agentic systems in medicine and healthcare

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.706939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.706939Z digest=sha256:76ce8020d549151ed702ef15cdcdd57b8639d9f13430ca3cba5424006c6a9daa

Observation 07767479-b3f5-4f23-b0fc-55911580a52b · outbound

This paper cites MedDM:LLM-executable clinical guidance tree for clinical decision-making.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making MedDM:LLM-executable clinical guidance tree for clinical decision-making

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.710633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.710633Z digest=sha256:b64fcf9f61d79d40150603ae1424656a7f8b7b7fbad40539a1bae10b77daee8f

Observation f6e46a8b-8451-40f1-bb9c-4a5c6e56aae5 · outbound

This paper cites Large language models encode clinical knowledge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Large language models encode clinical knowledge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.714905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.714905Z digest=sha256:77816a97bf8a621f412c9ec2f7481223df1ace43392a21c87b5178de0c58ea02

Observation e9883c0e-2b68-4294-b0a1-ad3bc2e3d75a · outbound

This paper cites Toward expert-level medical question answering with large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Toward expert-level medical question answering with large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.718778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.718778Z digest=sha256:ebb017307b5e955ca40271db787025136aa89d1f4023ff8a357b7748004a9253

Observation 3032ad41-4e58-4e07-9730-0e3dda922e41 · outbound

This paper cites Luks and Zachary D.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Luks and Zachary D

Reference 8

Resolution
verified exact
raw_fallback, observed 2026-08-15T19:15:23.025016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:15:21.722269Z digest=sha256:c580a08ee38a4a25291a9b58444627a0289b20acf801c0d51fa1f6d1e90a0400

Observation 604dc004-ab51-4576-b071-c5aacd248574 · outbound

This paper cites Variability in language used on social media prior to hospital visits.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Variability in language used on social media prior to hospital visits

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.725660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.725660Z digest=sha256:a70683cf3da75b3e47a1adc60f41a660625bc1c303dec0a4f8fe4bf4ad8770c3

Observation 3086486e-829f-4e22-a92f-4f862604e9be · outbound

This paper cites How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.728874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.728874Z digest=sha256:3a27cdcfe1d0b3484abd9ec501fe6edb333b75f23d9bb4d3f63985a261c59d36

Observation f7f08073-c5a2-4ebf-810b-3166225e8530 · outbound

This paper cites Open medical llm leaderboard.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Open medical llm leaderboard

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.732634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.732634Z digest=sha256:b1b1abd925133046e50f0c4a963af5e5b5b7ab23fb2c914ba33e3dd949dd774d

Observation 163ba107-58bb-4b94-9fc5-0faa9c41c254 · outbound

This paper cites A rapid review of gender, sex, and sexual orientation documentation in electronic health records.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A rapid review of gender, sex, and sexual orientation documentation in electronic health records

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.736183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.736183Z digest=sha256:175644adcb7bb94e438cd843537cbb71c73c6336e061027f961b6e8005d55839

Observation c1dc0c61-a6ca-4f36-9ccd-88d90607b332 · outbound

This paper cites Health Care Experiences of Patients with Nonbinary Gender Identities.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Health Care Experiences of Patients with Nonbinary Gender Identities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.739883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.739883Z digest=sha256:3625d4966fce6d7801b657a720d960806d1474c1d47a03ed4ca950ac5d8ece57

Observation b054846d-c701-4048-9a16-1075165d6f0e · outbound

This paper cites Hoffmann, Roger B.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Hoffmann, Roger B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.747154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.747154Z digest=sha256:f9b27ab0448ed238c879f3ae8aa85a03f478b702a4132229ef5b40386407f22d

Observation 1cc38915-e466-45da-89d9-9a1fa99bc006 · outbound

This paper cites an unresolved cited work.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Unresolved cited work

Reference 15

Resolution
verified exact
doi, observed 2026-08-15T19:15:22.291183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:15:21.750666Z digest=sha256:37595ac3c1f4e7dfc8c6aea68e0b0643e600b362882f4786ab0a9d609897f0b3

Observation df9834ba-ed33-4765-be03-b5b8f5a6d95f · outbound

This paper cites Gender disparities in health care.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender disparities in health care

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.754161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.754161Z digest=sha256:fd3e193fc7dcc7690a84ef376196bcc1bbc802ed7f0b086527545e4e2560c28f

Observation b7a80d4a-d602-4a80-a0f8-ac13be4c3c0b · outbound

This paper cites Defining gender disparities in pain management.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Defining gender disparities in pain management

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.757503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.757503Z digest=sha256:94cd8e782fbbad91994f3a2bb8456e16b948461c47a0fe1310e9c0794babe74c

Observation 80b854cd-ef13-4eb3-91cf-389d54d53754 · outbound

This paper cites Gender differences in outcomes of a multimodal pain management program.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender differences in outcomes of a multimodal pain management program

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.760789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.760789Z digest=sha256:04fff25904974cb6da3fbc3187dda5b5390c9f099526dbfa8512551a43e7859c

Observation 9b8a5ea0-c979-47fc-b5b8-eeba7f4a7354 · outbound

This paper cites Health and healthcare disparities among us women and men at the intersection of sexual orientation and race/ethnicity: a nationally representative cross-sectional study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Health and healthcare disparities among us women and men at the intersection of sexual orientation and race/ethnicity: a nationally representative cross-sectional study

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.763778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.763778Z digest=sha256:055f2d9dca8c0579ef928e13b2a8c777166f88c09a5b766648bdab9891edbe71

Observation 3ca81320-71ba-4cbe-b4b8-4f2acbfcf452 · outbound

This paper cites Gender bias in transformers: A comprehensive review of detection and mitigation strategies.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias in transformers: A comprehensive review of detection and mitigation strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.767181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.767181Z digest=sha256:890d07627568592f2b9d166da66030c35fb5165ca72ba76468f3af83ff7d03c1

Observation 3818626d-6e34-46af-80ab-5e088208544f · outbound

This paper cites Gender bias in natural language processing and computer vision: A comparative survey.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias in natural language processing and computer vision: A comparative survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.770938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.770938Z digest=sha256:018f70a5aaef95075b9118db692f2071e75aec0cfd53806c3a2a88456dfe5ee0

Observation eae15c9e-3435-493d-b2cf-5b5970538e03 · outbound

This paper cites Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.774269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.774269Z digest=sha256:a67bb19167026bfb50b75c7720edf6d69d49d2a8271c572a44f6b5075f5242f3

Observation 03496300-554e-48c5-8674-3c11f4013c81 · outbound

This paper cites Addressing gender-related performance disparities in neural rankers.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Addressing gender-related performance disparities in neural rankers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.778046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.778046Z digest=sha256:2d026c9dc20d8552562b1809a37b3cfa3b4beb688c4dcc28ddb8e16d42f2815b

Observation 5ad8a1c1-fa49-4ea2-89ad-58928594874e · outbound

This paper cites Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:15:22.897241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:15:21.781255Z digest=sha256:09fee0025a27a34898d0ca543a81773376369667eb23ae2a30e0d0a85686e99b

Observation ae8aaac9-2fa7-4bc4-b193-cb82ec9fc32a · outbound

This paper cites Bias in bios: A case study of semantic representation bias in a high-stakes setting.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bias in bios: A case study of semantic representation bias in a high-stakes setting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.784946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.784946Z digest=sha256:66e01d1f279f5d55322077df4181ccbe939fc89b98c8b567e2496c52cbc05572

Observation b87c217d-1adc-4fbb-b025-c165642ae253 · outbound

This paper cites Assessing the potential of gpt-4 to perpetuate racial and gender biases in health care: a model evaluation study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Assessing the potential of gpt-4 to perpetuate racial and gender biases in health care: a model evaluation study

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:15:21.788803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.788803Z digest=sha256:0aa54512f063f6e1585300ba22b4e6dd894af7da8d27ae95e329db551e5f05e9

Observation 8f18a0f6-7c6b-42fc-b0db-22bd40ee9df7 · outbound

This paper cites Bias patterns in the application of LLMs for clinical decision support: A comprehensive study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bias patterns in the application of LLMs for clinical decision support: A comprehensive study

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.792493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.792493Z digest=sha256:2fff1137e5cd3567d704b07ba2af20cc7cac59dd4e91245574d0c3d15eaf6479

Observation a2e7c5b1-aa5d-46c0-b7c7-59e788268d01 · outbound

This paper cites An investigation into the impact of deep learning model choice on sex and race bias in cardiac mr segmentation.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making An investigation into the impact of deep learning model choice on sex and race bias in cardiac mr segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.796629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.796629Z digest=sha256:8e61d1c7acde54dc28773a91ca7bd6012f939f5f729e426cc161499e226caad8

Observation c5c3ee91-32a4-45ee-a99b-b1b5ed9739cf · outbound

This paper cites Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.799920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.799920Z digest=sha256:ab400825ee2e11aa66f139a31320db694c358637f3769b75a9e69beeab9e4c6f

Observation a6f76897-4016-44a2-a27b-cd88d1519c10 · outbound

This paper cites Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.803345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.803345Z digest=sha256:c7270723284182996eec6f9986af0f9b954569d34eb955564dfcaa9aae173bdd

Observation 06199c94-cb4c-4ba6-9fbb-ce320207cf17 · outbound

This paper cites Subbalakshmi.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Subbalakshmi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.806421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.806421Z digest=sha256:bab03cf344583a55d2694d76df462b57fb1b4acd7f65fdc492a01765347d6277

Observation 0028562f-ea4c-4883-9eba-f13b08177cfe · outbound

This paper cites Gender identification from e-mails.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender identification from e-mails

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.809840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.809840Z digest=sha256:1f6789dc1a15cc636e55e0d0ca292fde09204e3f2ed00a9b366482c418b84601

Observation e852443a-e92c-4b30-8e51-55e8b839eb9a · outbound

This paper cites Gender, pseudonyms, and cmc: Masking identities and baring souls.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender, pseudonyms, and cmc: Masking identities and baring souls

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.814000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.814000Z digest=sha256:498dc4596596246a61994b88d8de85fd8a440b6c544b28c5270d5cd6b94dec14

Observation 045e3bf8-d25b-43cf-b475-a2694ddc4bca · outbound

This paper cites Bensing, W.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bensing, W

Reference 34

Resolution
verified exact
doi, observed 2026-08-15T19:15:22.274909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:15:21.817266Z digest=sha256:2fcc9c368075b0f322894d9cf52a1607b7277b9ecead1cc421faffa0503ce9cd

Observation 0ba291eb-fc49-4ed3-8dc2-6a0a506eac05 · outbound

This paper cites Write it like you see it: De- tectable differences in clinical notes by race lead to differential model recommendations.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Write it like you see it: De- tectable differences in clinical notes by race lead to differential model recommendations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.821718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.821718Z digest=sha256:b39e30684d1a0217151556fdfb2c20d55c40f683451d1195d3d66fbaab99e9ab

Observation 6849da2c-e308-4cdc-8854-13aed97681d3 · outbound

This paper cites Peek, and Elizabeth L.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Peek, and Elizabeth L

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.824800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.824800Z digest=sha256:1cf5b5924fc501a7faa38cb5f2f5f422df722f15828904ff044b9848f880ff7f

Observation 03a31317-ae3e-4349-b2e1-c81e3dea94ec · outbound

This paper cites "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.828030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.828030Z digest=sha256:491a26dc2b67b51b075a3d680919657eb2cea4b12a0868ee7ba723d8da5ea0a0

Observation 83c0a3fe-383c-4e48-adee-edae9b217784 · outbound

This paper cites How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.832911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.832911Z digest=sha256:75c7471d62aea3bf78a3320e0fabf0029325196e6fa0d6e024a91810e5dbb449

Observation d213c9f1-f684-453a-b1cc-4664fba2219d · outbound

This paper cites Closing the gap between open source and commercial large language models for medical evidence summarization.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Closing the gap between open source and commercial large language models for medical evidence summarization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.836320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.836320Z digest=sha256:4c3b3d07ed8fb389acb50132bcace512c47d44e4d02882b9f2b2ca8cd599123f

Observation 0106d6ba-7ded-4ebd-ace1-abd7f29e770a · outbound

This paper cites Conversational ai in health: Design considerations from a wizard-of-oz dermatology case study with users, clinicians and a medical llm.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Conversational ai in health: Design considerations from a wizard-of-oz dermatology case study with users, clinicians and a medical llm

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.839288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.839288Z digest=sha256:aee6bdb56e2f5ae844b4a3094b68d5124b54bbea16e5fdce564ed209ef3d7226

Observation 94f25ef1-0ee1-4586-87c7-d207a25eee21 · outbound

This paper cites Guidelines for rigorous evaluation of clinical llms for conversational reasoning.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Guidelines for rigorous evaluation of clinical llms for conversational reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.842895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.842895Z digest=sha256:c1cd055b95b934efb3af93183d78e9a94337bd2500daf69549caf8a856587158

Observation 5548ce9c-7ed3-40ab-9143-3f46c509df84 · outbound

This paper cites Effectiveness of a chatbot for eating disorders prevention: a randomized clinical trial.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Effectiveness of a chatbot for eating disorders prevention: a randomized clinical trial

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.846373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.846373Z digest=sha256:b8372c50cb2849424bb884b1fafb1018393ab7cec37839a6910e9c7ed21614a1

Observation d3a0d68d-6fa2-49d8-8688-276dc405dcc2 · outbound

This paper cites Performance of chatgpt on free-response, clinical reasoning exams.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Performance of chatgpt on free-response, clinical reasoning exams

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.849693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.849693Z digest=sha256:e6d1c084c7a9576ca5a4c163c616071d09c455d7df4f0cb9928f6cd6d9665031

Observation 77facdb3-58ac-4d26-8c6b-ab36d281fba9 · outbound

This paper cites The next generation: chatbots in clinical psychology and psychotherapy to foster mental health–a scoping review.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The next generation: chatbots in clinical psychology and psychotherapy to foster mental health–a scoping review

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.852646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.852646Z digest=sha256:686946eb8a529a093974205967873ddb4b71c35888eb4ef30f8772c7d5b8beaa

Observation 95c9aecb-921f-43b6-b917-99702cc6590e · outbound

This paper cites Large language model influence on diagnostic reasoning: a randomized clinical trial.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Large language model influence on diagnostic reasoning: a randomized clinical trial

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.856000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.856000Z digest=sha256:ad546728ae9d23f2b651fea4fabf2323b1f793a74e6ab8081d894c91bc2c53bc

Observation aa7f93a6-87d2-4f5c-a4e9-0565fafd0eac · outbound

This paper cites Human-algorithmic interaction using a large language model-augmented artificial intelligence clinical decision support system.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Human-algorithmic interaction using a large language model-augmented artificial intelligence clinical decision support system

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.859010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.859010Z digest=sha256:d71f3fc56fefb3cb654e281aff399ad5a8b2083aa13425c4f457354490c41638

Observation 710a98b7-d38d-4063-b188-482e7675ed7b · outbound

This paper cites The impact of responding to patient messages with large language model assistance.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The impact of responding to patient messages with large language model assistance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.862248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.862248Z digest=sha256:8d170c04ad3bf61bf503b7ece568c3e4cb5342efc37624fe9dc51ba0c92c82c4

Observation 2fda06f4-ea88-4966-b4ff-ebb93d606c84 · outbound

This paper cites Com- paring physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Com- paring physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.865790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.865790Z digest=sha256:055106f0fe9de2c662bab4071e64a835e701c7a4587010cc86e6e5ae6d860c66

Observation 6394c0de-e839-4e4b-abff-3253bd9aacad · outbound

This paper cites An evaluation framework for clinical use of large language models in patient interaction tasks.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making An evaluation framework for clinical use of large language models in patient interaction tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.868987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.868987Z digest=sha256:1ab634120e4dd782b1b6b847f118af35fd7c193e75d5f01afef9726b991ddb35

Observation d492586e-0885-4e5f-8afb-d2f78a5bc812 · outbound

This paper cites The medium is the message: How non-clinical information shapes clinical decisions in llms.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The medium is the message: How non-clinical information shapes clinical decisions in llms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.872398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.872398Z digest=sha256:c354440733946f8fb123f375f91522b85c9939e95d86968075279a38274118e1

Observation 938a4e14-eb07-4e1d-b368-9f2a02f7e45a · outbound

This paper cites Chatgpt: the next-gen tool for triaging? The American journal of emergency medicine, 69:215–217, 2023.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Chatgpt: the next-gen tool for triaging? The American journal of emergency medicine, 69:215–217, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.875617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.875617Z digest=sha256:6e09e5a33358706aebfc65cd0b5e84f8d9f920ae5e3faf35103b1aa85b75323d

Observation 9328e306-038c-457b-a345-05db4a4a59a5 · outbound

This paper cites The diagnostic and triage accuracy of the gpt-3 artificial intelligence model.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The diagnostic and triage accuracy of the gpt-3 artificial intelligence model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.879040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.879040Z digest=sha256:c0bda4bcbaccd34a022e16694f247eeceb38da4533afee81f5e1e1e70823ce05

Observation 12b16d0d-408b-4619-a8ce-fcfff8ea9885 · outbound

This paper cites Triage performance across large language models, chatgpt, and untrained doctors in emergency medicine: comparative study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Triage performance across large language models, chatgpt, and untrained doctors in emergency medicine: comparative study

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.882949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.882949Z digest=sha256:1529a3589b3da543e1573a9704c256208e41e87facc920b53196604cf79a1793

Observation 68ccb041-be39-4937-9ebb-fa4b303f0fe6 · outbound

This paper cites Evaluating llm-based generative ai tools in emergency triage: A comparative study of chatgpt plus, copilot pro, and triage nurses.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating llm-based generative ai tools in emergency triage: A comparative study of chatgpt plus, copilot pro, and triage nurses

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.886354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.886354Z digest=sha256:e3dd1891f2265db5b830812c9e9d9be442a1d9de43a5aa32e5c99712ee85ab33

Observation e27deeb6-c602-4ad4-9298-7ce0c74f9acb · outbound

This paper cites Integration of customised llm for discharge summary generation in real-world clinical settings: a pilot study on russell gpt.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Integration of customised llm for discharge summary generation in real-world clinical settings: a pilot study on russell gpt

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.890590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.890590Z digest=sha256:011030346473f5961de5c2c663a29548c1aae420825a7175e8ff495e3955706e

Observation 126df0af-a510-4147-a76d-a8e726b53b44 · outbound

This paper cites A toolbox for surfacing health equity harms and biases in large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A toolbox for surfacing health equity harms and biases in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.894779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.894779Z digest=sha256:12e1b47c6a36d4537ff1541cf970060b40b3e75de95d52ab816a8f199b643acc

Observation 793693f4-dc90-4b61-af2a-c9f7fb331166 · outbound

This paper cites Can AI Relate: Testing Large Language Model Response for Mental Health Support.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Can AI Relate: Testing Large Language Model Response for Mental Health Support

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.898179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.898179Z digest=sha256:a1b41a0e0e1d96e5cdb44e34444192869f0a9df7d734696c280ffcfa9a141bab

Observation b9b7e40d-e11e-4c60-99bf-3baec94930de · outbound

This paper cites A systematic review of large language model (llm) evaluations in clinical medicine.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A systematic review of large language model (llm) evaluations in clinical medicine

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.901737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.901737Z digest=sha256:a830bb0085db26360f1ebd5908e935648e0b074bc17f98c2122eed7a2ef52638

Observation 94c15ff3-939c-417a-ac27-d71eee752d28 · outbound

This paper cites Evaluating the clinical benefits of llms.Nature Medicine, 30(9):2409–2410, 2024.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating the clinical benefits of llms.Nature Medicine, 30(9):2409–2410, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.904754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.904754Z digest=sha256:657f7d0f696dec5595f2a80e3ba39b69fc422706b8c3655e0a8604bbcdd221ab

Observation 99398c58-fd04-4ee2-ae43-53757bceb76d · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.908939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.908939Z digest=sha256:3ab6c9a766ea8a18216b60e852a33a577681a19e63d06965421da3400b012fa8

Observation c245a714-1587-4c8f-8e5c-6cf2b428a2d1 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.912397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.912397Z digest=sha256:a67caf8a08d1641c578508252d122064f030508e0517cebd19a6d48da1c8c99b

Observation 738c7030-4081-40ec-a74c-732def8691df · outbound

This paper cites DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models

Reference 62

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:15:21.916224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.916224Z digest=sha256:2f74ad4395a5d20816376b5f1f8bdc39aac201ec29c252313e903c57b5045425

Observation d9c327ef-6b38-42ba-8014-72870728a54a · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.919409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.919409Z digest=sha256:2be3b813b84f0074c05204cb73739f9610ab6fc1c462dae66997980a6517ae12

Observation 829fb1d9-8764-4bb3-946f-44676c6f096a · outbound

This paper cites Performance of large language models on medical oncology examination questions.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Performance of large language models on medical oncology examination questions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.922867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.922867Z digest=sha256:7a9491ffdb67239e9f95d8b7279dbffc62d185f2a6aa4b250f946aa76585212c

Observation 74f6f77b-b65e-4716-83ab-2d26855d1b0b · outbound

This paper cites Medical Large Language Model Benchmarks Should Prioritize Construct Validity.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Medical Large Language Model Benchmarks Should Prioritize Construct Validity

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.926256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.926256Z digest=sha256:3f567ba2f3e39567f0f19bc9cca8da277cd2f26209daae101ae68593b15c351e

Observation fd456a41-7389-4100-bd54-ddfb12d613e5 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.930184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.930184Z digest=sha256:bbc010e3086f944fa3b50dfa23a337f2da2d3d4d8982f775ee3628b695267a98

Observation f32569ca-49f8-4a11-bbe0-c12970b98b9b · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.934638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.934638Z digest=sha256:923b61f8ebe2a94dae0ba57ce7d3cfb97807eadd2f958ed674099156c61c1596

Observation b5873111-7447-490b-b0fd-4c507b035507 · outbound

This paper cites Automating evaluation of ai text generation in healthcare with a large language model (llm)-as- a-judge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Automating evaluation of ai text generation in healthcare with a large language model (llm)-as- a-judge

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.939396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.939396Z digest=sha256:8050ab8730749b69cdb2ec06baba9cfee2a15deabb84e897569bd27195689aa8

Observation aba09d39-5f6d-466b-944b-f6ba8c560a0f · outbound

This paper cites Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.942893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.942893Z digest=sha256:47422f01b3bebc47b854c9bb469d90d35235d90d3d6bef7e9a83136fad6a3bb4

Observation 9950664d-9951-4067-b49c-804a0c7f3f1f · outbound

This paper cites A Survey on LLM-as-a-Judge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A Survey on LLM-as-a-Judge

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.946197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.946197Z digest=sha256:b941a0fe394a0d4765e2e9846e65c9e4a6b3be52226fc57eb6060b82dbf21564

Observation 42676804-f7d6-42ac-bd6e-d2e393d58ebf · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Can Large Language Models Be an Alternative to Human Evaluations?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.950376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.950376Z digest=sha256:6e75bc79ef264ffbaed671c764b1b0a8274de57761a6fcc24969e2a3fd97976a

Observation 7d78a3f1-ba94-4a4e-965a-3d604a7b44f3 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.953810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.953810Z digest=sha256:e13bcf9259980d844ad601e5b419013ab64ee0489a66ca1a53e592f7c92449fb

Observation b06258f5-6ac2-47b3-8c31-6a6c4f36e4ad · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.957520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.957520Z digest=sha256:fe1d00fd297f41999c54edadb2e231f197efe84f030b32289f5e31c72395f9dd

Observation 23e4f448-4fe3-499c-a296-62d0bb0a3edd · outbound

This paper cites As- sessment of pathology domain-specific knowledge of chatgpt and comparison to human performance.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making As- sessment of pathology domain-specific knowledge of chatgpt and comparison to human performance

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.961112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.961112Z digest=sha256:5497ab05b2fae5c927f70fee74ed8dc226f4fed1b0f2c61a533e35d5c6dbfcd7

Observation 318db684-412b-4da4-8b6b-bd29aad69080 · outbound

This paper cites Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.964430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.964430Z digest=sha256:aba897df3d9fdfa68d7b705706718a405d721baa55a9b46d6527e536bc84e771

Observation e8d4c663-c964-4e59-b977-f379e7efeb7b · outbound

This paper cites Style Over Substance: Evaluation Biases for Large Language Models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Style Over Substance: Evaluation Biases for Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.967504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.967504Z digest=sha256:f26be0ed4d06baa6f436ca706ed61bda91024a72d3485a80f1137bfa00510a46

Observation ad2dbca1-09b4-414e-a234-3694e0db6415 · outbound

This paper cites Ehrnoteqa: An llm benchmark for real- world clinical practice using discharge summaries.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Ehrnoteqa: An llm benchmark for real- world clinical practice using discharge summaries

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.971049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.971049Z digest=sha256:7a90907288ff4b31c2eef84b69d5355f3c71e603cb4e62df647212fc6433a28c

Observation 4ad5a4b3-0cb3-4981-a583-42f4ac9cab04 · outbound

This paper cites Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Zhuo Ran Cai, Roxana Daneshjou, and Pranav Rajpurkar.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Zhuo Ran Cai, Roxana Daneshjou, and Pranav Rajpurkar

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.974313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.974313Z digest=sha256:b33fc07f9df76a7266f3eddc38cbb601162665d5bec5830dbed4dc186d22ce89

Observation a2fae7cf-5f85-49f8-b27f-5e8c2047f143 · outbound

This paper cites The Llama 3 Herd of Models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The Llama 3 Herd of Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.977984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.977984Z digest=sha256:768fe94b8ee72db5ede31f5779ba1eeda75d56cf1672fbdb8043722feec908ac

Observation 0d1bc52a-cb62-4ce3-b0e6-3c1024250754 · outbound

This paper cites Linguistic analy- sis of communication in therapist-assisted internet-delivered cognitive behavior therapy for generalized anxiety disorder.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Linguistic analy- sis of communication in therapist-assisted internet-delivered cognitive behavior therapy for generalized anxiety disorder

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.980919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.980919Z digest=sha256:2d3cae4ae0b258592c4f3f4375a543e8d7416a0203ede33e8f2c154a2a6bd422

Observation 34a263ac-a9f0-4e76-aab2-8184bb42c548 · outbound

This paper cites Toward linguistic recognition of generalized anxiety disorder.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Toward linguistic recognition of generalized anxiety disorder

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.984000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.984000Z digest=sha256:007c2075e601b1ecbc82d903f30f72b9297c1d2e11971614fec7808d4420edc6

Observation bab69140-78dd-404e-84cf-a0849dea1d7f · outbound

This paper cites Linguistic markers of anxiety and depression in somatic symptom and related disorders: Observational study of a digital intervention.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Linguistic markers of anxiety and depression in somatic symptom and related disorders: Observational study of a digital intervention

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.987812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.987812Z digest=sha256:cc03dc184cf908932f7ca0250322fe411b43c8824c33f6016bb6d5df5a7ec9f3

Observation 11dd3167-df4a-4bb8-b04d-32554d1dc644 · outbound

This paper cites Are patient linguistic tones associated with mental health and perceived clinician empathy? JBJS, 103(23):2181–2189, 2021.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Are patient linguistic tones associated with mental health and perceived clinician empathy? JBJS, 103(23):2181–2189, 2021

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.991152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.991152Z digest=sha256:8a7d52454f2db3e0708162999170176d78acd0d289c1d82c08831726f421af44

Observation 56185fba-c4b9-4e81-b80f-3eb7c82ac921 · outbound

This paper cites GPT-4 Technical Report.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making GPT-4 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.994385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.994385Z digest=sha256:dc2686030443de0a5f57f964b1133876aac40c9667afb3b7f78846f8fcd4a463

Observation 860ff5ca-935d-4850-898e-cd23879b55c0 · outbound

This paper cites Palmyra-med: Instruction-based fine-tuning of llms enhancing medical domain performance.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Palmyra-med: Instruction-based fine-tuning of llms enhancing medical domain performance

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.997880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.997880Z digest=sha256:b61554b65cad533b44151dae6880290fab3c9f938b4247d1db888affd3f18484

Observation 1d2fba3d-82ac-4d12-a340-e95bd0f2af14 · outbound

This paper cites The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.001808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.001808Z digest=sha256:3a1ac5c5f0ded95222bd5ba96a19829cfb360676d6616d892bdb8c3ec27f00b2

Observation 3e6b9251-d1e5-4672-8a36-51befbafa48c · outbound

This paper cites Multiple significance tests: the bonferroni method.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Multiple significance tests: the bonferroni method

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.005542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.005542Z digest=sha256:2927e233705bd41dd998ce805877eeacbc0353055e448272f60f969d669b1d05

Observation d9b4bed2-ba0b-42aa-a6a3-bf50ac4e2945 · outbound

This paper cites Note on the sampling error of the difference between correlated proportions or percentages.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Note on the sampling error of the difference between correlated proportions or percentages

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.008516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.008516Z digest=sha256:9e3f1c6a74b740fa072646c1f4af2045b0de9807da9f371684423faaef802e84

Observation b19016fc-cee7-4c16-b5a6-f9b4178cb3e8 · outbound

This paper cites Individual comparisons by ranking methods.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Individual comparisons by ranking methods

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.011752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.011752Z digest=sha256:02e905465ebea2a7c76ca1ef0d7cd28c93dc4a3be22f8ce68d173392c32dcedf

Observation 68fada91-d8d6-4e98-9c56-1fdf01e888d6 · outbound

This paper cites The measurement of observer agreement for categorical data.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The measurement of observer agreement for categorical data

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.015339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.015339Z digest=sha256:3579735e2dc03a36efef8be9755846a06d06eaaf2346446d07243f2e6e1d41a7

Observation 15692283-b172-4ac9-99ca-b5000a05cc08 · outbound

This paper cites On the use and interpretation of certain test criteria for purposes of statistical inference part i.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making On the use and interpretation of certain test criteria for purposes of statistical inference part i

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.018639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.018639Z digest=sha256:98cd80ee0c62f5093b8c6b4647eadd26c6a76d46fea38d89dc8bf0c7f3f55f01

Observation dcf64b46-7631-4460-b53a-9f9cc53a0df0 · outbound

This paper cites On a test of whether one of two random variables is stochastically larger than the other.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making On a test of whether one of two random variables is stochastically larger than the other

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.021993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.021993Z digest=sha256:c53ddb02f8c63e5709b6537e519980c2000b9114092f742b9d29032496050b9c

Observation a72ee610-6d01-4221-8918-e0fe277004d1 · outbound

This paper cites Gender bias and stereotypes in large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias and stereotypes in large language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.024908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.024908Z digest=sha256:da474f89f73446410e4dba4d7f1e790a8efb972d1c30e74ba5c20e48751b2433

Observation 0a2be227-7012-4856-9130-dbed21282f16 · outbound

This paper cites Llm evaluators recognize and favor their own generations.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Llm evaluators recognize and favor their own generations

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.027941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.027941Z digest=sha256:b30da36e5f75e6778c60224b289808d521a0097a24b1329b4d7f03d50b6a2280

Observation b8147dac-cbeb-41b5-8556-710015d964b2 · outbound

This paper cites Bender and Batya Friedman.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bender and Batya Friedman

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.031212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.031212Z digest=sha256:81e71c8fcc5c6daf946885122c283b086e5bb95069c87db5feed6ae548682a74

Observation 906af124-436a-4233-a10b-574b49c57199 · outbound

This paper cites Basic demographics, health practices, and health status of us medical students.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Basic demographics, health practices, and health status of us medical students

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.034556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.034556Z digest=sha256:a7f772bb0c2a2d03d4c5c8d3caa303e5e042e2e5fa19dc19f540b00bd5ee7659

Observation df7aa351-d8ab-4712-bdc0-a1a192dcd0cd · outbound

This paper cites International medical graduates in the us physician workforce and graduate medical education: current and historical trends.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making International medical graduates in the us physician workforce and graduate medical education: current and historical trends

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.037896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.037896Z digest=sha256:f49efa9004f2013300ee7db163904b9dc6a7eb97662e425fb5e4ac67f582a0a4

Observation ba0cffc4-169b-46d6-b428-a46af3745e46 · outbound

This paper cites Clinical reasoning education at us medical schools: results from a national survey of internal medicine clerkship directors.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Clinical reasoning education at us medical schools: results from a national survey of internal medicine clerkship directors

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.041997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.041997Z digest=sha256:02e2a506b2a3e9512bb35d9661b1adb6fb8b13a4421e2b18df1e8f95b420824b

Observation 27e35414-7bf1-44df-91fa-af90972b99f5 · outbound

This paper cites Teaching medical students the important connection between communication and clinical reasoning.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Teaching medical students the important connection between communication and clinical reasoning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.046929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.046929Z digest=sha256:60ffb8816eaaac035e0a8d70202baf27d6730466d925d203c32681931067fbbd

Observation 335521db-c88c-40fc-a86f-ff80cedd88af · outbound

This paper cites Factors associated with medical student clinical reasoning and evidence based medicine practice.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Factors associated with medical student clinical reasoning and evidence based medicine practice

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.050331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.050331Z digest=sha256:c09b51d2880d56742a37734cc9538a8b51c3c6ef31fdbb7e4bc83e3434c68f58

Pith citing papers

Observation fb758406-220f-4da9-b0e5-e28ca6c4b238 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.672097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:69783fb04053f8fc16780edc47a1a570e46d15f060866052356bf690d5a1521a

Observation 17c60d7f-00ef-42b6-be75-24438d6c3d87 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.549200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:5f29928664fdd7e930566caa351ffcb4c03a9b007ea0c238c4cc64548283ab36