Pith. sign in

REVIEW 2 cited by

Factuality of Large Language Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02420 v3 pith:WK4DVKWC submitted 2024-02-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords factualityllmsimprovinglanguagelargemodelsresearchsurvey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs), especially when instruction-tuned for chat, have become part of our daily lives, freeing people from the process of searching, extracting, and integrating information from multiple sources by offering a straightforward answer to a variety of questions in a single place. Unfortunately, in many cases, LLM responses are factually incorrect, which limits their applicability in real-world scenarios. As a result, research on evaluating and improving the factuality of LLMs has attracted a lot of attention recently. In this survey, we critically analyze existing work with the aim to identify the major challenges and their associated causes, pointing out to potential solutions for improving the factuality of LLMs, and analyzing the obstacles to automated factuality evaluation for open-ended text generation. We further offer an outlook on where future research should go.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conversational AI as a Catalyst for Informal Learning: An Empirical Large-Scale Study on LLM Use in Everyday Learning

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Most adults in a German Prolific sample report using large language models for informal learning, with four distinct learner profiles emerging from their usage patterns.

  2. AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    AssertBench measures how often LLMs keep the same true/false evaluation of a fact across contradictory user framings, and finds most tested models agree with the user's framing more when they do not know the fact.

Pith tools