Each tested LLM shows its own characteristic unreliability when engaging in repair during extended math-question dialogues.
People’s perceptions toward bias and related concepts in large language models: A systematic review
3 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 3roles
background 1polarities
background 1representative citing papers
TrustLLM defines eight trustworthiness principles, creates a six-dimension benchmark, and evaluates 16 LLMs showing proprietary models generally lead but some open-source ones are close while over-calibration can hurt utility.
A 2x2 between-subjects experiment finds contextualization lowers AI persuasiveness but warmth restores it through crossover interaction, with reliance invariant to design, trust predicting outcomes independently, and AI literacy decoupling trust from behavior.
citing papers explorer
-
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
Each tested LLM shows its own characteristic unreliability when engaging in repair during extended math-question dialogues.
-
TrustLLM: Trustworthiness in Large Language Models
TrustLLM defines eight trustworthiness principles, creates a six-dimension benchmark, and evaluates 16 LLMs showing proprietary models generally lead but some open-source ones are close while over-calibration can hurt utility.
-
Personalized to Persuade: The Effects of Contextualization and Warmth on Trust and Reliance in Conversational AI
A 2x2 between-subjects experiment finds contextualization lowers AI persuasiveness but warmth restores it through crossover interaction, with reliance invariant to design, trust predicting outcomes independently, and AI literacy decoupling trust from behavior.