REVIEW 6 cited by
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, partly due to the time and cost of repeated experiments. We use benchmarks designed for testing LLMs' capacity to reason about cardinal directions to explore the impact of experimental repeats on mean score and prediction interval. We suggest a simple method for cost-effectively quantifying the uncertainty of a benchmark score and make recommendations concerning reproducible LLM evaluation.
Forward citations
Cited by 6 Pith papers
-
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs
A user study of an astronomy RAG bot identifies the question types and evaluation criteria astronomers actually use, and turns them into a 40-item benchmark.
-
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
Zero-shot LLM Big Five labels show weak, genre-dependent signal that often collapses under lexical masking, and LEX-EC separates prevalence artifacts from recoverable psycholinguistic evidence.
-
Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
Large language models, especially GPT-4 with few-shot prompts, can classify topological spatial relations between WKT-encoded geometries with roughly 0.6 to 0.66 accuracy, though errors cluster near conceptually simil...
-
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.
-
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering
A Bayesian model that groups similar LLM test prompts into clusters gives better predictive scores than a no-clustering baseline but does not prove that it truly corrects prompt dependence.
-
AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito
An AI agent combining GraphRAG, static Fortran analysis, and LLM code generation is reported to translate legacy Fortran finite-difference code into Devito, with Grade-A results claimed on roughly three-quarters of 13...
Discussion (0). Sign in to comment.