Across five subjective tasks and five open-source LLMs, demographic prompting improves human agreement only for 1–3 high-signal, directionally coherent attributes and degrades under the full attribute set.
Automated Ableism: An Exploration of Explicit Disability Biases in Sentiment and Toxicity Analysis Models
4 Pith papers cite this work, alongside 11 external citations. Polarity classification is still indexing.
fields
cs.CL 4representative citing papers
LLM counter-arguments in ChangeMyView convey more trust and social-intent signals than human opinion-changing comments, and crowdworkers prefer them as more persuasive.
Truthfulness grew from zero papers in 2021-2022 to the largest topic by 2025-2026, while explainability declined and then resurged in 2026 through mechanistic interpretability.
A survey that compiles and taxonomizes more than 32 existing hallucination mitigation techniques for LLMs while analyzing their challenges and limitations.
citing papers explorer
-
Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement
Across five subjective tasks and five open-source LLMs, demographic prompting improves human agreement only for 1–3 high-signal, directionally coherent attributes and degrades under the full attribute set.
-
"I understand your perspective": LLM Persuasion through the Lens of Communicative Action Theory
LLM counter-arguments in ChangeMyView convey more trust and social-intent signals than human opinion-changing comments, and crowdworkers prefer them as more persuasive.
-
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
Truthfulness grew from zero papers in 2021-2022 to the largest topic by 2025-2026, while explainability declined and then resurged in 2026 through mechanistic interpretability.
-
A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
A survey that compiles and taxonomizes more than 32 existing hallucination mitigation techniques for LLMs while analyzing their challenges and limitations.