Pith. sign in

REVIEW 16 cited by

A Survey on Large Language Model Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.15361 v1 pith:MA5GLWWY submitted 2025-08-21 cs.CL

A Survey on Large Language Model Benchmarks

classification cs.CL
keywords benchmarksmodelcapabilitiesdevelopmentevaluationlanguagelargecore
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In recent years, with the rapid development of the depth and breadth of large language models' capabilities, various corresponding evaluation benchmarks have been emerging in increasing numbers. As a quantitative assessment tool for model performance, benchmarks are not only a core means to measure model capabilities but also a key element in guiding the direction of model development and promoting technological innovation. We systematically review the current status and development of large language model benchmarks for the first time, categorizing 283 representative benchmarks into three categories: general capabilities, domain-specific, and target-specific. General capability benchmarks cover aspects such as core linguistics, knowledge, and reasoning; domain-specific benchmarks focus on fields like natural sciences, humanities and social sciences, and engineering technology; target-specific benchmarks pay attention to risks, reliability, agents, etc. We point out that current benchmarks have problems such as inflated scores caused by data contamination, unfair evaluation due to cultural and linguistic biases, and lack of evaluation on process credibility and dynamic environments, and provide a referable design paradigm for future benchmark innovation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness

    cs.LG 2026-05 unverdicted novelty 7.0

    Benchmark-specific training maps to shift bribery and is NP-hard under Borda and mean win rate; mean win rate has the highest instance-level robustness (median 22 tasks on BBH) among tested aggregation rules.

  2. Efficient Ensemble Selection from Binary and Pairwise Feedback

    cs.GT 2026-05 unverdicted novelty 7.0

    The paper develops efficient algorithms for ensemble selection from binary and pairwise feedback, achieving (1-1/e) guarantees with query savings for coverage and PTAS-style results via submodular relaxation for theta...

  3. Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization

    cs.AI 2026-02 unverdicted novelty 7.0

    NLCO benchmark shows LLMs achieve reasonable feasibility on small natural-language CO tasks but degrade on larger instances, with set-based problems easier than graph-structured or bottleneck-objective ones.

  4. YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

    cs.CL 2026-07 unverdicted novelty 6.0

    YOMI-Bench is a new benchmark of four tasks for kanji reading and phonological understanding in LLMs, showing low performance even for Japanese-specific and commercial models.

  5. Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

    cs.AI 2026-06 unverdicted novelty 6.0

    EvalCards is a composable reporting schema and monitoring tool for AI evaluations, derived from 52 papers and 10 interviews, and applied to 5,816 models and 101,843 results to surface reporting gaps.

  6. LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling

    cs.LG 2026-05 conditional novelty 6.0

    LPDS quantifies difficulty of logic-preserving problem variations and searches for the hardest ones, producing up to 5x larger performance drops than random sampling and better robustness gains from fine-tuning on dif...

  7. Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

    cs.LG 2026-05 conditional novelty 6.0

    3-bit quantization induces new stereotypical biases in 6-21% of previously unbiased BBQ items across three LLMs, undetected by perplexity increases under 3%, with models declining in 'unknown' responses by 17.4%.

  8. Can LLM Teams Play What? Where? When?

    cs.CL 2026-05 unverdicted novelty 5.0

    Team interaction strategies improve LLM accuracy on recent ChGK questions by up to 20 points, reaching 44.23% and nearing some human team levels.

  9. Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints

    cs.CL 2026-04 unverdicted novelty 5.0

    Large reasoning models exhibit reasoning collapse, with accuracy dropping sharply beyond task-specific complexity thresholds in controlled versions of nine classical reasoning tasks using strict validity validators.

  10. MAVEN: Improving Generalization in Agentic Tool Calling

    cs.AI 2026-05 unverdicted novelty 4.0

    MAVEN is a modular verification scaffold that lifts an open 120b model's tool-calling accuracy from 48% to 71% on MAVEN-Bench without retraining.

  11. AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports

    cs.CL 2026-01 reject novelty 4.0

    An LLM-generated QA benchmark over World Bank African economic reports is evaluated on a small sample and found hard, but the abstract and body report conflicting dataset sizes, model names, and scores.

  12. Token-Operations-Oriented Inference Optimization Techniques for Large Models

    cs.SE 2026-06 unverdicted novelty 3.0

    The paper introduces a four-layer technical architecture for token-operations-oriented inference optimization in large models and reviews key technologies and industry status at each layer.

  13. Token-Operations-Oriented Inference Optimization Techniques for Large Models

    cs.SE 2026-06 conditional novelty 3.0

    A survey of large-model inference optimization, organized as a four-layer 'token-operations' taxonomy: multi-model fusion, model optimization, compute-model fusion, and compute-network-model fusion.

  14. Designing for Error Recovery in Human-Robot Interaction

    cs.RO 2026-04 unverdicted novelty 3.0

    Position paper calls for designing robotic AI to detect and recover from its own errors in continuous interactions, using nuclear glovebox operations as an illustrative case.

  15. When control meets large language models: From words to dynamics

    eess.SY 2026-02 unverdicted novelty 3.0

    The paper proposes a bidirectional continuum between LLMs and control systems, covering LLM-assisted controller design, control-based LLM steering, and state-space modeling of LLMs.

  16. The Necessity of a Unified Framework for LLM-Based Agent Evaluation

    cs.AI 2026-02 conditional novelty 3.0

    A position paper arguing that LLM-agent benchmarks are confounded by framework-specific choices and proposing a unified sandbox-and-methodology evaluation framework.