Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.
arXiv preprint arXiv:2306.10062 , year=
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
The benchmark score matrix of 84 models on 133 tasks is approximately rank-2; BenchPress recovers held-out scores to within 4.6 points and identifies 5-benchmark subsets that predict the full scorecard to within 3.93-4.55 points.
A ridge predictor using prompt-level agreement spread, label-assisted first-correct position, completion-length variance, and entropy reaches Spearman ρ=0.90 with observed best-of-N gains across three model families and six post-training methods.
LLM confidence judgments are dominated by a shared difficulty factor across models, with the confidence-performance link collapsing after removing agreed items, yielding no evidence for individuated metacognition.
Introduces a calibration framework for AI benchmarks using world-population probability levels on logarithmic scales derived from human test data and LLM extrapolation.
citing papers explorer
-
Will Scaling Improve Social Simulation with LLMs?
Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.
-
You Don't Need to Run Every Eval
The benchmark score matrix of 84 models on 133 tasks is approximately rank-2; BenchPress recovers held-out scores to within 4.6 points and identifies 5-benchmark subsets that predict the full scorecard to within 3.93-4.55 points.
-
Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics
A ridge predictor using prompt-level agreement spread, label-assisted first-correct position, completion-length variance, and entropy reaches Spearman ρ=0.90 with observed best-of-N gains across three model families and six post-training methods.
-
LLMs Show No Signs Of Individuated Metacognition
LLM confidence judgments are dominated by a shared difficulty factor across models, with the confidence-performance link collapsing after removing agreed items, yielding no evidence for individuated metacognition.
-
From Human-Level AI Tales to AI Leveling Human Scales
Introduces a calibration framework for AI benchmarks using world-population probability levels on logarithmic scales derived from human test data and LLM extrapolation.
- Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration