Pith. sign in

REVIEW 6 cited by

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03492 v2 pith:WCX7CJS6 submitted 2024-10-04 cs.CL

classification cs.CL
keywords benchmarkuncertaintyevaluationllmsmodelsquantifyingreproduciblescore
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, partly due to the time and cost of repeated experiments. We use benchmarks designed for testing LLMs' capacity to reason about cardinal directions to explore the impact of experimental repeats on mean score and prediction interval. We suggest a simple method for cost-effectively quantifying the uncertainty of a benchmark score and make recommendations concerning reproducible LLM evaluation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Queries to Criteria: Understanding How Astronomers Evaluate LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A user study of an astronomy RAG bot identifies the question types and evaluation criteria astronomers actually use, and turns them into a 40-item benchmark.

  2. LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Zero-shot LLM Big Five labels show weak, genre-dependent signal that often collapses under lexical masking, and LEX-EC separates prevalence artifacts from recoverable psycholinguistic evidence.

  3. Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Large language models, especially GPT-4 with few-shot prompts, can classify topological spatial relations between WKT-encoded geometries with roughly 0.6 to 0.66 accuracy, though errors cluster near conceptually simil...

  4. ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

    cs.AI 2025-12 reject novelty 5.0 of 10

    LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.

  5. Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering

    cs.CR 2025-10 conditional novelty 5.0 of 10

    A Bayesian model that groups similar LLM test prompts into clusters gives better predictive scores than a no-clustering baseline but does not prove that it truly corrects prompt dependence.

  6. AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito

    cs.AI 2026-01 conditional novelty 4.0 of 10

    An AI agent combining GraphRAG, static Fortran analysis, and LLM code generation is reported to translate legacy Fortran finite-difference code into Devito, with Grade-A results claimed on roughly three-quarters of 13...

Pith tools