Pith. sign in

REVIEW 2 cited by

Systematic Assessment of Factual Knowledge in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.11638 v3 pith:B7B2CDGC submitted 2023-10-18 cs.CL

classification cs.CL
keywords knowledgellmsdomainsfactualevaluateframeworkgenericlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Previous studies have relied on existing question-answering benchmarks to evaluate the knowledge stored in large language models (LLMs). However, this approach has limitations regarding factual knowledge coverage, as it mostly focuses on generic domains which may overlap with the pretraining data. This paper proposes a framework to systematically assess the factual knowledge of LLMs by leveraging knowledge graphs (KGs). Our framework automatically generates a set of questions and expected answers from the facts stored in a given KG, and then evaluates the accuracy of LLMs in answering these questions. We systematically evaluate the state-of-the-art LLMs with KGs in generic and specific domains. The experiment shows that ChatGPT is consistently the top performer across all domains. We also find that LLMs performance depends on the instruction finetuning, domain and question complexity and is prone to adversarial context.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Adding information-gain-selected virtual views refined by video diffusion priors to 3D Gaussian Splatting improves arbitrary-view rendering quality.

  2. A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLM knowledge, measured by self-reported true/false checks on knowledge-graph triplets, shows homophily and degree correlations that a graph neural network exploits to select more effective fine-tuning data.

Pith tools