Sequential testing with tailored stopping rules lets model evaluation halt early once statistical needs (CI width, significance, equivalence) are met, saving up to 80% compute on VLM leaderboards.
Efficient benchmarking (of language models)
3 Pith papers cite this work, alongside 7 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
background 1polarities
background 1representative citing papers
A dual local/global selection of ~100 non-agentic instances predicts four agentic benchmarks with LOOCV MAE under 4%, Spearman above 0.80, and ~85% pairwise ranking accuracy at under 1% of full agent cost.
PRISM supplies a geometric upper bound on LLM variant risk that splits drift into scale, shape, and head axes and doubles as a differentiable regularizer against forgetting.
citing papers explorer
-
Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data
Sequential testing with tailored stopping rules lets model evaluation halt early once statistical needs (CI width, significance, equivalence) are met, saving up to 80% compute on VLM leaderboards.
-
PACE: A Proxy for Agentic Capability Evaluation
A dual local/global selection of ~100 non-agentic instances predicts four agentic benchmarks with LOOCV MAE under 4%, Spearman above 0.80, and ~85% pairwise ranking accuracy at under 1% of full agent cost.
-
PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head
PRISM supplies a geometric upper bound on LLM variant risk that splits drift into scale, shape, and head axes and doubles as a differentiable regularizer against forgetting.