LLMs routinely produce unsupported causal stories for personal sensing anomalies, and richer evidence or constrained prompts do not reliably eliminate this epistemic overreach.
Title resolution pending
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 7roles
background 1polarities
background 1representative citing papers
LLMs prompted as peer supporters for ADRD caregivers produce synthetic lived experience through narrative language that differs from human peers in first-person and past-tense usage, revealing a narrative authenticity gap.
Profile-conditioned LLMs achieve higher tacit alignment with humans on subjective spectra when traits match, as quantified by the new Tacit Understanding Index (TUX) from 241 humans and 200 agents.
Scoping review of 134 studies on LLM-as-a-Judge in healthcare finds concentration in clinical decision support and NLP, frequent use of OpenAI models with prompt engineering, and moderate-to-strong human alignment where validated.
Toxic prompt perturbations reduce LLM factual accuracy on three benchmarks and selectively amplify perturbation-sensitive nodes in attribution graphs.
LLUMI shows that open-source LLMs trained via SFT and DPO on Reddit community feedback can match proprietary GPT models on readability, empathy, connection, actionability, and safety for mental health support.
An algorithm audit finds that OpenAI moderation, Llama Guard, and Shield Gemma frequently flag content from real therapy sessions as undesirable.
citing papers explorer
-
Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
LLMs routinely produce unsupported causal stories for personal sensing anomalies, and richer evidence or constrained prompts do not reliably eliminate this epistemic overreach.
-
When AI Says "I have been in similar situations": Synthetic Lived Experience in Peer-Like Caregiver Support
LLMs prompted as peer supporters for ADRD caregivers produce synthetic lived experience through narrative language that differs from human peers in first-person and past-tense usage, revealing a narrative authenticity gap.
-
TUX: Measuring Human--AI Tacit Understanding
Profile-conditioned LLMs achieve higher tacit alignment with humans on subjective spectra when traits match, as quantified by the new Tacit Understanding Index (TUX) from 241 humans and 200 agents.
-
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
Scoping review of 134 studies on LLM-as-a-Judge in healthcare finds concentration in clinical decision support and NLP, frequent use of OpenAI models with prompt engineering, and moderate-to-strong human alignment where validated.
-
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
Toxic prompt perturbations reduce LLM factual accuracy on three benchmarks and selectively amplify perturbation-sensitive nodes in attribution graphs.
-
LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback
LLUMI shows that open-source LLMs trained via SFT and DPO on Reddit community feedback can match proprietary GPT models on readability, empathy, connection, actionability, and safety for mental health support.
-
AI Content Moderation in Therapy Conversations
An algorithm audit finds that OpenAI moderation, Llama Guard, and Shield Gemma frequently flag content from real therapy sessions as undesirable.