Pith. sign in

REVIEW 60 cited by

FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14251 v2 pith:Y42PHZ7C submitted 2023-05-23 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords evaluationfactscoreatomicchatgpthumanmodelspublicautomated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Evaluating the factuality of long-form text generated by large language models (LMs) is non-trivial because (1) generations often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality inadequate, and (2) human evaluation is time-consuming and costly. In this paper, we introduce FACTSCORE, a new evaluation that breaks a generation into a series of atomic facts and computes the percentage of atomic facts supported by a reliable knowledge source. We conduct an extensive human evaluation to obtain FACTSCOREs of people biographies generated by several state-of-the-art commercial LMs -- InstructGPT, ChatGPT, and the retrieval-augmented PerplexityAI -- and report new analysis demonstrating the need for such a fine-grained score (e.g., ChatGPT only achieves 58%). Since human evaluation is costly, we also introduce an automated model that estimates FACTSCORE using retrieval and a strong language model, with less than a 2% error rate. Finally, we use this automated metric to evaluate 6,500 generations from a new set of 13 recent LMs that would have cost $26K if evaluated by humans, with various findings: GPT-4 and ChatGPT are more factual than public models, and Vicuna and Alpaca are some of the best public models. FACTSCORE is available for public use via `pip install factscore`.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 60 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text

    cs.CL 2026-08 conditional novelty 7.0 of 10

    A probe trained only on classical context-memory conflict data detects decomposition-induced contradictions (DI-CC) in an LLM's fact-splitting step, while self-consistency detection fails.

  2. SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

    cs.CL 2026-07 conditional novelty 7.0 of 10

    SERPO co-evolves per-question grading rubrics, Good-Normal-Bad response archives, and policy parameters so a language model can train itself at inference time without labels, gaining up to 20.6 points on open-ended me...

  3. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    cs.CL 2026-07 conditional novelty 7.0 of 10

    On a new 615-question business-case benchmark graded by AI against instructor rubrics, frontier LLMs score 87-88% partial credit but complete only about half the questions.

  4. Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Clinical RAG can attribute real evidence about drug Y to queried drug X at high rates under adversarial retrieval, a failure invisible to faithfulness and citation metrics but detectable by entity-attribution verification.

  5. RWGBench: Evaluating Scholarly Positioning in Related Work Generation

    cs.DL 2026-05 unverdicted novelty 7.0 of 10

    RWGBench measures related-work generation by citation choices, and shows citation-focused metrics expose failures that text-similarity and LLM-judge scores miss.

  6. More Is Not More: What Matters for Diversity in LLM Opinions?

    cs.CL 2026-05 conditional novelty 7.0 of 10

    Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.

  7. ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.

  8. Teaching Smaller Language Models To Generalise To Unseen Compositional Questions (Full Thesis)

    cs.CL 2024-11 conditional novelty 7.0 of 10

    Smaller language models can generalize to unseen compositional questions when trained and evaluated with retrieval-augmented contexts, and combining Wikipedia retrieval with LLM-generated rationales improves accuracy.

  9. Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CM-LRS is a workflow-output-layer reliability score for capital-markets LLM outputs; on five public-document workflows, frontier closed-source models score 4.09–4.31 and Llama 3.3 70B scores 3.15 under four LLM judges.

  10. Grounded verification of chemical and materials reasoning: detection is the bottleneck

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Database-grounded correction of LLM chemistry claims is detection-limited: repair of flagged errors succeeds 80-97%, while missed in-loop detection caps the accuracy gain.

  11. Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Cheaper LLM judges match frontier models on citation-quality F1 but differ substantially in false positive and false negative rates, meaning reward signal calibration matters more than model cost.

  12. Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Distilling an 8B reasoning teacher into a 0.6B student recovers most summary quality at ~50× speed, but teacher type—not scale alone—determines which capabilities transfer.

  13. Unsupervised Hallucination Detection by Inspecting Reasoning Processes

    cs.CL 2025-09 conditional novelty 6.0 of 10

    IRIS detects LLM hallucinations by training a lightweight probe on hidden states elicited during the model's own step-by-step verification, using the model's verbalized confidence as soft pseudolabels.

  14. MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A new benchmark with a three-way hallucination taxonomy, snapshot-based test cases, and an LLM judge shows LLM agents hallucinate at over 30% of risky decision points, with open and closed models closer than expected.

  15. Beyond Facts: Evaluating Intent Hallucination in Large Language Models

    cs.CL 2025-06 reject novelty 6.0 of 10

    The paper proposes a query-centric evaluation of LLM "intent hallucination" via constraint decomposition, but the headline metric comparison is undermined by a self-referential human evaluation design.

  16. Generating Grounded Responses to Counter Misinformation via Learning Efficient Fine-Grained Critiques

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MisMitiFact trains lightweight T5 critique models on fact-checking data to identify errors in numbers, entities, and topics, and uses their short critiques to refine LLM counter-responses at about 5x lower feedback cost.

  17. SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SUCEA improves adversarial fact-checking by decomposing claims into atomic sub-claims, editing each sub-claim toward retrieved evidence, and re-retrieving before predicting the final label.

  18. Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations

    cs.AI 2025-06 conditional novelty 6.0 of 10

    SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.

  19. Vid2Coach: Transforming How-To Videos into Task Assistants

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Vid2Coach converts how-to videos into a real-time, wearable task assistant that helps blind and low vision people cook with fewer errors.

  20. RAGPPI: RAG Benchmark for Protein-Protein Interactions in Drug Discovery

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RAGPPI introduces a QA benchmark for PPI biological impacts in drug target identification, with 500 expert-validated and 3,720 auto-labeled pairs (sum 4,220, though the abstract says 4,420).

  21. MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A clinician-validated taxonomy and 35-benchmark suite show that large language models vary widely across medical tasks, with reasoning models leading overall.

  22. Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PhantomCircuit traces knowledge overshadowing to attention circuits during training and prunes circuit edges to recover the overshadowed answer.

  23. A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Metalinguistic disagreements, where the dispute is over word meaning rather than facts, appear in LLM fact-checking against knowledge graphs, based on a 250-triple pilot study.

  24. Provence: efficient and robust context pruning for retrieval-augmented generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Provence prunes and reranks retrieved contexts in one pass, compressing 50-80% of the context while keeping question-answering accuracy close to the full-context baseline.

  25. FRAG: A Flexible Modular Framework for Retrieval-Augmented Generation based on Knowledge Graphs

    cs.CL 2025-01 conditional novelty 6.0 of 10

    FRAG routes knowledge-graph questions through a query-complexity classifier to BFS or shortest-path retrieval, improving KG-RAG accuracy without LLM fine-tuning or retrieval-time LLM calls.

  26. The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

    cs.CL 2025-01 conditional novelty 6.0 of 10

    FACTS Grounding is a benchmark and leaderboard that scores LLMs on producing long-form answers fully grounded in up to 32k-token documents, using a validated panel of judge models.

  27. MapExplorer: New Content Generation from Low-Dimensional Visualizations

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A new task and benchmark that generate contextually aligned text for arbitrary coordinates in 2D projection maps, evaluated with an LLM-based metric.

  28. Reducing Tool Hallucination via Reliability Alignment

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A reliability alignment framework, Relign, that expands the LLM tool-use action space with indecisive actions reduces tool hallucination rates and improves task success on the new RelyToolBench benchmark.

  29. A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A rubric-plus-question-answering LLM framework for literary translation evaluation beats traditional MT metrics but still trails human agreement, especially on Korean honorifics.

  30. FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A modular query-code-update pipeline reduces measurement hallucinations in chest X-ray reports by replacing model-generated numbers with measurements from specialized vision tools.

  31. Fact-Level Confidence Calibration and Self-Correction

    cs.CL 2024-11 conditional novelty 6.0 of 10

    The paper measures LLM confidence per atomic fact, weights correctness by relevance, and uses high-confidence facts from the same response to correct low-confidence facts.

  32. Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI

    cs.CL 2026-07 unverdicted novelty 5.0 of 10

    HALO is a layered oversight architecture that grounds, constrains, verifies, abstains, traces, and monitors LLM outputs to make hallucinations containable rather than eliminated.

  33. LLM-as-a-Verifier: A General-Purpose Verification Framework

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.

  34. Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Learning a margin-based confidence ranker for LLM judges improves agreement-target success in cascaded selective evaluation compared to heuristic confidence scores.

  35. vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

    cs.NI 2026-02 conditional novelty 5.0 of 10

    vLLM Semantic Router routes LLM requests by composing thirteen signal types into Boolean decision policies, with safety, caching, and model-selection plugin chains.

  36. Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

    cs.CL 2025-11 reject novelty 5.0 of 10

    Reasoning LLMs seem better at retrieving hierarchical facts not because they know more but because they navigate better; the key supporting RL experiment is missing from the paper.

  37. Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization

    cs.CL 2025-09 reject novelty 5.0 of 10

    PEFT adapters trained on high-resource summarization domains can improve Llama-3-8B's summaries on unseen domains, but the reported gains are weakened by test-set selection and missing significance tests.

  38. GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework

    cs.CL 2025-08 reject novelty 5.0 of 10

    GOSU globally merges semantic units from text chunks into a unit-centric knowledge graph and uses three-tier keyword retrieval to improve RAG generation quality, according to LLM-judge win rates.

  39. ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Adding information-gain-selected virtual views refined by video diffusion priors to 3D Gaussian Splatting improves arbitrary-view rendering quality.

  40. Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Tool-augmented LLM annotators improve agreement with ground-truth preferences on long-form factual and coding tasks, with mixed results on math, compared to standard LLM-as-a-judge baselines.

  41. "Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLMs ground answers in early context far more than later context, and chain-of-thought prompting or reasoning models reduce contextual grounding rather than improving it.

  42. RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new 6K-claim benchmark evaluates LLMs and multimodal LLMs on real-world fact-checking with an explicit 'unknown' option and shows web search and multimodal input improve performance.

  43. Improving Factuality for Dialogue Response Generation via Graph-Based Knowledge Augmentation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The paper proposes TG-DRG and GA-DRG, two graph-augmented frameworks that combine coreference resolution, knowledge selection, and graph encoding to improve factuality of dialogue responses, evaluated with a newly pro...

  44. MiniCPM4: Ultra-Efficient LLMs on End Devices

    cs.CL 2025-06 conditional novelty 5.0 of 10

    MiniCPM4-8B reportedly matches Qwen3-8B on standard benchmarks while using about 22% of the training tokens, and achieves large long-context speedups on edge devices.

  45. Bi-Fact: A Bidirectional Factorization-based Evaluation of Intent Extraction from UI Trajectories

    cs.AI 2025-02 conditional novelty 5.0 of 10

    Bi-Fact, a bidirectional fact-level LLM-based metric, reports higher agreement with human judgments than existing metrics when scoring intent extraction from GUI trajectories.

  46. LLMs to Support a Domain Specific Knowledge Assistant

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A synthetic QA dataset for IFRS sustainability reporting is created with LLMs and used to build and evaluate two QA pipelines, with the fully LLM-based pipeline scoring highest.

  47. RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance

    cs.LG 2025-01 reject novelty 5.0 of 10

    A framework using fine-tuned LLaVA and VILA models to score relevance and correctness in multimodal retrieval-augmented generation, with reported 88 percent test accuracy and 91 percent human agreement.

  48. Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM

    cs.CL 2024-12 reject novelty 5.0 of 10

    SumAutoEval is an LLM-based entity-level summarization evaluator with four dimensions; its claimed human-correlation advantage is not consistently supported by the experiments.

  49. Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A multi-agent, self-training LLM framework called MESA evaluates meeting summaries by detecting eight error types and reports higher correlation with human scores than existing automatic metrics.

  50. LLM Ensemble for RAG: Role of Context Length in Zero-Shot Question Answering for BioASQ Challenge

    cs.CL 2025-09 conditional novelty 4.0 of 10

    An ensemble of zero-shot LLMs with BM25 retrieval and semantic reranking ranked first in one BioASQ 13 yes/no batch, with longer contexts observed to hurt answer quality.

  51. Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection

    cs.CL 2025-09 reject novelty 4.0 of 10

    Singular values of label-grouped n-gram frequency tensors are used as MLP features for hallucination detection, with reported gains on HaluEval that rely on label-aware grouping.

  52. Holistic Artificial Intelligence in Medicine; improved performance and explainability

    cs.AI 2025-06 conditional novelty 4.0 of 10

    An extension of the HAIM multimodal framework that uses LLM-based retrieval and summarization to improve clinical prediction AUC from 79.9% to 90.3% and to generate document-grounded explanations.

  53. A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A research agenda calling for geo-temporal reasoning in deep research systems, with no experiments or system implementation.

  54. LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking

    cs.IR 2025-06 conditional novelty 4.0 of 10

    A KG-enhanced LlamaRec that feeds user-specific relation paths into a Llama-2 ranker reports modest MRR, NDCG, and Recall gains on two benchmarks.

  55. OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models

    cs.CL 2025-01 reject novelty 4.0 of 10

    OnionEval reports that small LLMs detect atomic fact hallucinations well but fail in layered contexts, a result undermined by prompt confounds and numeric inconsistencies.

  56. A Report on Financial Regulations Challenge at COLING 2025

    cs.CE 2024-12 conditional novelty 4.0 of 10

    A COLING 2025 shared task built nine financial regulation tasks, evaluated six submitted LLMs against three baselines, and found fine-tuning helps but closed models still lead.

  57. When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

    cs.CL 2026-01 conditional novelty 3.0 of 10

    Adding generic prompt rules to task-specific LLM prompts is not monotonic: in 15-20 case local suites, Llama 3 and Qwen 2.5 sometimes pass fewer extraction and RAG checks, so prompt changes should be tested per task.

  58. Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.

  59. A comprehensive taxonomy of hallucinations in Large Language Models

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A survey that organizes LLM hallucination types, causes, benchmarks, and mitigations, and restates the theorem that hallucination is inevitable for computable LLMs.

  60. eSapiens: A Platform for Secure and Auditable Retrieval-Augmented Generation

    cs.AI 2025-07 reject novelty 2.0 of 10

    The abstract promises benchmark wins for the eSapiens RAG platform, but the appendix tables do not contain those numbers and show higher hallucination rates than the baseline.

Pith tools