Pith. sign in

REVIEW 3 cited by

How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07282 v2 pith:3AWA6UMQ submitted 2024-02-11 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmshelpfulnesshonestyhumanmodelsconversationallanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In day-to-day communication, people often approximate the truth - for example, rounding the time or omitting details - in order to be maximally helpful to the listener. How do large language models (LLMs) handle such nuanced trade-offs? To address this question, we use psychological models and experiments designed to characterize human behavior to analyze LLMs. We test a range of LLMs and explore how optimization for human preferences or inference-time reasoning affects these trade-offs. We find that reinforcement learning from human feedback improves both honesty and helpfulness, while chain-of-thought prompting skews LLMs towards helpfulness over honesty. Finally, GPT-4 Turbo demonstrates human-like response patterns including sensitivity to the conversational framing and listener's decision context. Our findings reveal the conversational values internalized by LLMs and suggest that even these abstract values can, to a degree, be steered by zero-shot prompting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training language models to be warm and empathetic makes them less reliable and more sycophantic

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Fine-tuning five LLMs for warmth raised errors by about 5 to 15 percentage points on safety-critical questions and increased sycophancy, especially when users sounded sad.

  2. Accelerating RLHF Training with Reward Variance Increase

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A new reward reshaping method provably increases reward variance for GRPO-based RLHF training, with an O(n log n) global optimization algorithm and preliminary speedups in experiments.

  3. Transition from Statistical to Hardware-Limited Scaling in Photonic Quantum State Reconstruction

    quant-ph 2026-03 unverdicted novelty 5.0 of 10

    Classical shadow tomography on integrated photonics shows a sharp transition from statistical O(M^{-1/2}) error scaling to a hardware-limited floor set by unitary spectral distortions.

Pith tools