Pith. sign in

REVIEW 13 cited by

Alignment for Honesty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07000 v2 pith:HISRIZNZ submitted 2023-12-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords honestyalignmentknowledgellmsmetricsmodelsresearchtraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent research has made significant strides in aligning large language models (LLMs) with helpfulness and harmlessness. In this paper, we argue for the importance of alignment for \emph{honesty}, ensuring that LLMs proactively refuse to answer questions when they lack knowledge, while still not being overly conservative. However, a pivotal aspect of alignment for honesty involves discerning an LLM's knowledge boundaries, which demands comprehensive solutions in terms of metric development, benchmark creation, and training methodologies. We address these challenges by first establishing a precise problem definition and defining ``honesty'' inspired by the Analects of Confucius. This serves as a cornerstone for developing metrics that effectively measure an LLM's honesty by quantifying its progress post-alignment. Furthermore, we introduce a flexible training framework which is further instantiated by several efficient fine-tuning techniques that emphasize honesty without sacrificing performance on other tasks. Our extensive experiments reveal that these aligned models show a marked increase in honesty, as indicated by our proposed metrics. We open-source all relevant resources to facilitate future research at \url{https://github.com/GAIR-NLP/alignment-for-honesty}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prompt Compression via Activation Aggregation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A learned weighted sum of intermediate-layer activations compresses an instruction prompt into a single patch vector that, injected at an early layer, recovers task accuracy within ~2% of the full prompt.

  2. The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Fine-tuning on model-generated text filtered by internal probes improves factual accuracy more than fine-tuning on gold documents across three long-form generation domains.

  3. Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Video-LLMs can be trained, via SFT or DPO on a new synthetic dataset UVQA, to refuse questions that cannot be answered from the video content, with modest cost to answerable QA performance.

  4. CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CausalAbstain filters multilingual self-feedback by comparing how much it changes the model's abstention decision, improving abstention accuracy over baselines on two benchmarks.

  5. Writing Like the Best: Exemplar-Based Expository Text Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Exemplar-based expository generation: a recurrent plan-then-adapt LLM framework converts a source topic's exemplar into target-topic text by question transfer, retrieval, and confidence-gated answering.

  6. Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SENATOR guides a language model through a knowledge graph, measures its uncertainty with structural entropy, and fine-tunes it on synthetic data chosen to fix its weak spots, gaining up to 12 percent average relative ...

  7. GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    GRAIT selects and reweights refusal-training examples using gradient influence, reporting lower hallucination rates and better helpfulness scores than prior refusal-aware tuning baselines.

  8. UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    UAlign improves LLM factuality alignment by adding predicted confidence and semantic entropy as input features to prompts and the reward model, helping the model answer known questions and refuse unknown ones.

  9. Improve Decoding Factuality by Token-wise Cross Layer Entropy of Large Language Models

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A new decoding-time method, END, uses per-token cross-layer entropy of prediction growth to boost factual tokens, improving truthfulness and informativeness on hallucination benchmarks.

  10. Aligning Large Language Models for Faithful Integrity Against Opposing Argument

    cs.CL 2025-01 conditional novelty 5.0 of 10

    An LLM is fine-tuned with DPO to make the strength of its stance in conversation match its self-estimated confidence, improving resistance to misleading arguments and receptiveness to corrections.

  11. How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Entity popularity and entity co-occurrence in Wikipedia correlate with LLM QA accuracy, confidence, and calibration, and combining them with confidence improves answer-correctness prediction by 5.24% on average.

  12. A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.

  13. Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.

Pith tools