Pith. sign in

arXiv preprint arXiv:2306.10062 , year=

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

fields

cs.CL 3 cs.LG 3

years

2026 6

representative citing papers

Will Scaling Improve Social Simulation with LLMs?

cs.CL · 2026-07-02 · conditional · novelty 6.0

Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.

You Don't Need to Run Every Eval

cs.LG · 2026-06-22 · conditional · novelty 6.0

The benchmark score matrix of 84 models on 133 tasks is approximately rank-2; BenchPress recovers held-out scores to within 4.6 points and identifies 5-benchmark subsets that predict the full scorecard to within 3.93-4.55 points.

LLMs Show No Signs Of Individuated Metacognition

cs.LG · 2026-05-22 · unverdicted · novelty 5.0

LLM confidence judgments are dominated by a shared difficulty factor across models, with the confidence-performance link collapsing after removing agreed items, yielding no evidence for individuated metacognition.

From Human-Level AI Tales to AI Leveling Human Scales

cs.LG · 2026-02-21 · unverdicted · novelty 5.0

Introduces a calibration framework for AI benchmarks using world-population probability levels on logarithmic scales derived from human test data and LLM extrapolation.

citing papers explorer

Showing 6 of 6 citing papers.