Pith. sign in

REVIEW 6 cited by

Know Your Limits: A Survey of Abstention in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.18418 v3 pith:Z5W7HJVC submitted 2024-07-25 cs.CL

classification cs.CL
keywords abstentionframeworklanguagelargemodelsspecificsurveysystems
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Abstention, the refusal of large language models (LLMs) to provide an answer, is increasingly recognized for its potential to mitigate hallucinations and enhance safety in LLM systems. In this survey, we introduce a framework to examine abstention from three perspectives: the query, the model, and human values. We organize the literature on abstention methods, benchmarks, and evaluation metrics using this framework, and discuss merits and limitations of prior work. We further identify and motivate areas for future research, such as whether abstention can be achieved as a meta-capability that transcends specific tasks or domains, and opportunities to optimize abstention abilities in specific contexts. In doing so, we aim to broaden the scope and impact of abstention methodologies in AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PhantomFill: When the Form Demands an Answer, Language Models Invent One

    cs.LG 2026-06 conditional novelty 7.0 of 10

    Required JSON fields make LLMs invent answers to unanswerable questions 100% of the time in ten of thirteen models, even when an 'insufficient evidence' escape exists.

  2. Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A ground-truth-first synthetic memory benchmark shows that agent-memory architecture rankings invert with history length: short-horizon leaders lose at nine weeks.

  3. LLM Abstention Can Be a Prompt Artifact, in Addition to Genuine Uncertainty

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Adding an extra 'Unknown' option to True/False prompts causes LLMs to abstain on questions they can answer, and random words reproduce the effect, indicating abstention is partly a prompt artifact.

  4. Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Video-LLMs can be trained, via SFT or DPO on a new synthetic dataset UVQA, to refuse questions that cannot be answered from the video content, with modest cost to answerable QA performance.

  5. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.

  6. Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An automated LLM-based evaluator finds that eight LLMs rarely hallucinate factors in legal argument generation but often omit relevant factors and usually fail to abstain when no common ground exists.

Pith tools