Pith. sign in

REVIEW 1 cited by

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.02058 v1 pith:KJFWHFVI submitted 2025-06-01 cs.CL cs.IRcs.LGstat.APstat.ME

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

classification cs.CL cs.IRcs.LGstat.APstat.ME
keywords knowledgeknowsumllmsobservedevaluationunseencapabilitiesdemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Accurate evaluation of large language models (LLMs) is crucial for understanding their capabilities and guiding their development. However, current evaluations often inconsistently reflect the actual capacities of these models. In this paper, we demonstrate that one of many contributing factors to this \textit{evaluation crisis} is the oversight of unseen knowledge -- information encoded by LLMs but not directly observed or not yet observed during evaluations. We introduce KnowSum, a statistical framework designed to provide a more comprehensive assessment by quantifying the unseen knowledge for a class of evaluation tasks. KnowSum estimates the unobserved portion by extrapolating from the appearance frequencies of observed knowledge instances. We demonstrate the effectiveness and utility of KnowSum across three critical applications: estimating total knowledge, evaluating information retrieval effectiveness, and measuring output diversity. Our experiments reveal that a substantial volume of knowledge is omitted when relying solely on observed LLM performance. Importantly, KnowSum yields significantly different comparative rankings for several common LLMs based on their internal knowledge.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UCS: Estimating Unseen Coverage for Improved In-Context Learning

    cs.LG 2026-04 unverdicted novelty 6.0

    UCS estimates the number of unrevealed latent clusters in candidate demonstration sets via Smoothed Good-Turing on embeddings to improve ICL performance by 2-6% when added to baselines.