Pith. sign in

REVIEW 4 major objections 6 minor 39 references

LLM leaderboard scores leave a geometric blind spot so large that top rankings are often structurally indeterminate.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Public LLM leaderboards have effective dimension ~3–5, so the geometric blind spot between models with identical scores exceeds runner-up gaps by ~100× and makes top rankings structurally unreliable.

T0 review reviewed 2026-07-12 challenge →

load-bearing objection Solid multi-suite empirics and a usable coverage algorithm; the stereological bound is real geometry but the headline 100× indeterminacy sits on a width model the authors themselves flag as near its frontier limit. the 4 major comments →

arxiv 2606.05169 v1 pith:AP2TBU3K submitted 2026-04-15 cs.LG

The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models

classification cs.LG MSC 52A2062H2568T50
keywords LLM evaluationbenchmark coverageeffective dimensionalitystereologyHausdorff distancesubmodular selectionGardner problemranking reliability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper treats public LLM leaderboard scores as one-dimensional projections of unknown multi-dimensional capability profiles, the same way X-rays project a body. Using stereology and convex-geometry tools, it proves that when a suite of benchmarks only probes a few independent directions, many different capability profiles remain consistent with the same scores. On three independent competitive leaderboards the effective dimensionality is only about three to five, so the structural gap between models that look identical on the suite dwarfs both the runner-up score gap and ordinary statistical noise by one to two orders of magnitude. Half-split experiments confirm the practical consequence: randomly holding out half the benchmarks swaps the top-ranked model in most trials. The paper then supplies a greedy algorithm that selects a small stable core of benchmarks covering most of the observed variance, and as a pure-math side result resolves a 1995 open problem on how stably widths recover a smooth convex body. A sympathetic reader cares because the work turns a familiar unease about leaderboard rankings into a quantitative, falsifiable geometric claim with a concrete remedy.

Core claim

For any benchmark suite whose effective dimensionality is d_eff, the visible Hausdorff distance between two convex capability profiles that produce identical scores is at most ε plus a term that shrinks only as m to the power −1/(d_eff−1). Empirically d_eff lies in [2.86, 4.80] on three frontier leaderboards, so this structural blind spot exceeds the observed runner-up gap by roughly two orders of magnitude and dominates statistical noise by 52–127×; half-split trials swap the top model in 92 percent of draws.

What carries the argument

The indistinguishability bound (Theorem 2): after modelling each benchmark as a width (support-function) measurement of an origin-centred convex capability body, the Hausdorff distance that remains invisible to m directions in the effective subspace is controlled by the covering radius on the sphere of dimension d_eff−1, yielding the rate m^{−1/(d_eff−1)} with a matching Lipschitz lower bound.

Load-bearing premise

Each benchmark must be well approximated by a linear width measurement of a convex capability profile; when models become very similar or a task is highly non-linear, that linearisation residual can grow as large as ordinary score gaps and the geometric bound no longer cleanly separates structure from noise.

What would settle it

On a new frontier leaderboard, compute d_eff and the empirical covering radius; if the resulting visible indistinguishability radius is smaller than the top-pair score gap, or if random half-splits of the suite almost never reverse the top ranking, the central claim that structural blind spots dominate is falsified for that population.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper develops a stereological theory of LLM benchmark coverage: benchmarks are treated as width-like measurements of convex capability profiles, and for a suite with effective dimensionality d_eff the visible Hausdorff distance between profiles consistent with the same scores is bounded by ε + C R m^{-1/(d_eff-1)} (Theorem 2), with a matching Lipschitz lower bound and a smooth C^{1,α}/C² extension. Empirically, three leaderboard families have frontier d_eff ∈ [2.86, 4.80]; the structural radius is reported to exceed runner-up gaps by ~10² and statistical noise by 52–127×; half-split experiments swap top-1 in 92% of trials; a submodular greedy algorithm with the Nemhauser guarantee finds a stable 4-benchmark core and 7/12 for 90% coverage with strong temporal transfer; counterfactual removal/addition correlates with eigenstructure (ρ = −0.69, +0.38). As a second contribution, the authors claim a resolution of Gardner’s Problem 1.5 for C² support functions via optimal recovery on S^{D−1}, with rate Θ(R/(κ m^{2/(D−1)})).

Significance. If the stereological bound and its empirical calibration hold, the paper supplies a quantitative, geometry-based account of why public LLM rankings are structurally underdetermined even with infinite data per benchmark—distinct from item noise or decoding variance—and a practical coverage algorithm with classical approximation guarantees. Strengths that should be credited: multi-suite empirics (Open LLM v2, extended 12-bench, LiveBench, Arena categories, Epoch AI) with bootstrap CIs, MP/permutation eigenvalue checks, and half-split swap counts; full proof appendices for Theorems 1–4 and the Gardner-rate argument via n-widths/Jackson–Bernstein; explicit regime diagnostics for linearization (Prop. 1, App. I H.15); and a reproducibility path (pip install -e .; reproduce.py). The independent geometric-tomography contribution, if correctly scoped, is of interest beyond ML. The work is therefore potentially high-impact for evaluation methodology, provided the width-model application on the competitive frontier is tightened.

major comments (4)
  1. [Setup; Prop. 1; Table 2; App. I H.15] Setup / Prop. 1 and Table 2: The headline claim that the structural blind spot exceeds Δ₂ by two orders of magnitude (e.g. 123× on OLLM v2 frontier, 385× on Extended) is load-bearing and rests on the width representation after linearization. The manuscript itself reports that on the frontier (top 50%, n=148) median linear R² falls from 0.984 to 0.876 and the minimum to 0.710 (MATH Lvl 5), with TruthfulQA at 0.769 and highest η, and states that the model is at its regime boundary when the quadratic-vs-linear R² gap is comparable to standardized Δ₂ (gap ≤0.067 vs Δ₂≈0.072 on Extended). The practitioner formula δ_vis_H≈2R ω_emp drops ε; if residual η·diam² is O(Δ₂), the geometric covering term no longer cleanly dominates the observed gap as a pure stereological effect. Please either (i) recompute Table 2 and the 52–127× noise comparisons with ε including the measured linearization residual
  2. [§4 Theorem 2; Worked example; Corollary 1; App. H.21] Theorem 2 vs ranking claims (worked example, Corollary 1): δ_vis_H is a Hausdorff distance between convex bodies whose widths agree within ε on m directions in V_eff. The leap to “any ranking claim between models separated by less than 21.2 standardized units is structurally indeterminate” and to rank-equivalence classes (App. H.47) requires an explicit map from Hausdorff distance on profiles to aggregate score gaps under the actual aggregator. Half-split swap rates (92% top-1, mean 2.83/5 top-5) are strong and largely model-independent evidence of ranking instability; they should be cleanly separated from the stereological bound so that empirical instability does not silently substitute for the geometric theorem when the width residual is non-negligible.
  3. [Abstract; Theorem 3; App. E–G] Theorem 3 / App. G (Gardner’s Problem 1.5): The abstract and introduction claim a resolution of Gardner’s Problem 1.5 for C² support functions. Gardner’s original problem is stated for X-rays of planar convex bodies; the manuscript proves a minimax rate for width measurements and then argues X-ray parity for origin-symmetric bodies (Thm. 20 / App. G). That is a substantial and interesting contribution, but “resolution of Problem 1.5 as Gardner stated it” should be scoped precisely: state what is proved for widths, what is proved or conjectured for X-rays, and what remains open (exact constants, non-symmetric case, adaptive X-rays). Overclaiming the classical problem will invite specialist pushback and distract from the LLM results.
  4. [Corollary 1; Prop. 2; App. H G.1–G.4] Corollary 1 / App. G.2–G.3: The χ² top-1 formula depends on ambient D, which three estimators place in a wide range [6,184]; the paper correctly notes robustness of P(swap) because Δ₂ is small, but the calibration table shows the bound is loose at small r (14–17 pp) and prior-sensitive (isotropic 0.79 vs empirical 0.29 in simulation). Present the χ² bound strictly as a rejection threshold / optimistic isotropic case (Prop. 2), not as a calibrated probability, and avoid headline intervals [0.38,0.49] without stating the prior and D assumptions in the same sentence.
minor comments (6)
  1. [Table 1; App. I H.39] LiveBench frontier n=19 is very small; H.39 notes d_eff=4.74 lies above the bootstrap CI upper bound and Horn finds 0 signal eigenvalues. Soften LiveBench as a third independent confirmation or report only full-slice d_eff=2.63 with CI.
  2. [NeurIPS Paper Checklist] NeurIPS checklist answers are filled [Yes] with “Justification: Not applicable,” which is inconsistent with the guidelines text. Replace with brief real justifications or N/A where appropriate.
  3. [Figure 3] Figure 3 uses private-use Unicode glyphs that render as boxes in the manuscript text; replace with standard PDF text or vector labels.
  4. [§2 Notation; Theorem 2; Remark 2] Notation: m is both “visible benchmarks” and generic sample size in covering bounds; d_eff is non-integer while covering is on S^{⌈d_eff⌉−1}—Remark 2 helps, but a single consistent convention in Theorem 2 would reduce confusion.
  5. [§1 Related work; §7] Related work: BenchScope [Sha and Zhao, 2026] is concurrent; keep the diagnostic/credit split clear and ensure arXiv identifiers/dates are stable at camera-ready.
  6. [Table 2; App. I H.12] H.12 reports geometric vs statistical radius 8.95× while the abstract/main text say 52–127×; reconcile the two definitions (πR/k vs 2R ω_emp / bootstrap radius) in one place.

Circularity Check

1 steps flagged

No load-bearing circularity: Hausdorff rates, participation-ratio d_eff, and submodular coverage are classical derivations applied to leaderboard matrices; half-split swaps are independent empirical checks.

specific steps
  1. fitted input called prediction [Appendix H, G.3 (D estimation by three converging methods)]
    "Cross-validation against half-split swaps. Choosing the D that minimises the gap between the chi-squared prediction and the empirical half-split swap rate gives D = 6 (the minimum admissible)."

    One of the three ambient-dimension estimators is explicitly tuned so that the χ² top-1 formula matches the same half-split swap rates it is later compared to. That particular D is therefore not an independent prediction of the swap rate. The step is minor and non-load-bearing: the paper also reports power-law and parallel-analysis D, sweeps D∈[6,184], and the headline Hausdorff / d_eff / greedy claims do not depend on this calibrated D.

full rationale

Walk of the derivation chain shows the central results are not equivalent to their inputs by construction. d_eff is the standard participation ratio of Corr(S) (Definition 1 / Theorem 1), not a fit to ranking unreliability. Theorem 2’s visible Hausdorff bound follows from Lipschitz continuity of support functions plus Rogers covering on S^{d_eff−1}, with a matching lower bound via a Lipschitz construction; the practitioner form 2Rω_emp is an evaluation of that bound on measured loadings, not a re-labeling of Δ_2. Proposition 1’s width representation is an explicit modeling assumption with residual absorbed into ε and with reported R² diagnostics, not a self-definition of the bound. Greedy coverage uses classical submodularity of tr(P_T Σ) and the Nemhauser (1−1/e) guarantee (Theorem 4 / Das–Kempe). Gardner’s Problem 1.5 is resolved via Jackson–Bernstein / Kolmogorov n-widths on S^{D−1} (Theorem 3 / Appendix G), independent of the LLM empirics. The 92% half-split top-1 swap rate and the six-prior swap band [0.38, 0.49] are separate Monte Carlo / held-out experiments, not algebraic consequences of the fitted spectrum. The only mild proximity to circularity is that one of three D estimators (App. G.3) is chosen to minimize the gap between the χ² formula and half-split rates; that estimator is not load-bearing for the Hausdorff claim or the reported prior sweep, and the paper reports the full D∈[6,184] envelope. No self-citation uniqueness theorem or ansatz smuggling appears. Score 1 rather than 0 only for that optional D-calibration step.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 3 invented entities

The load-bearing move is modeling unknown model abilities as convex bodies whose widths are (linearized) benchmark scores, then reading leaderboard correlation eigenstructure as the effective measurement subspace. Free choices are mostly analysis knobs (frontier quantile, ambient D envelope, standardization). Background math (participation ratio, Marchenko–Pastur, Nemhauser, Rogers covering, Jackson–Bernstein, Kolmogorov n-widths) is standard. The main invented modeling layer is the capability-body / width-benchmark correspondence plus the visible-vs-orthogonal Hausdorff split used to turn PCA into a ranking-reliability bound.

free parameters (5)
  • frontier_threshold_q
    Main claims use top-50% by mean score; d_eff is monotone in q and tighter frontiers enlarge the blind spot (H.16). Not fitted to maximize the headline ratio, but it selects the population.
  • ambient_capability_dimension_D
    Enters χ² swap formulas via σ_hidden; three estimators give [6,184]. Headlines report robustness across a D grid rather than a single fitted D.
  • population_radius_R_and_standardization
    R is max standardized row norm (e.g. 7.30 on extended frontier). δ_vis_H/Δ₂ ratios change a lot across z-score/min-max/rank/raw (H.23/H.41), though qualitative dominance remains.
  • linearization_residual_epsilon_eta
    Absorbs width-model error; bounded empirically by quadratic-vs-linear R² gaps ≤0.067, but is a data-dependent tolerance in Theorem 2.
  • coverage_target_tau_and_core_frequency_cutoff
    90% coverage and >90% bootstrap appearance for the stable core of 4 are design choices for the treatment recommendations.
axioms (6)
  • domain assumption Benchmarks act as approximately linear, monotone, Lipschitz width measurements of capability profiles (Proposition 1).
    Justifies replacing scores by support/width geometry; verified by high linear R² on the studied suite but flagged as weaker for adversarial/safety tasks.
  • domain assumption Capability profiles may be treated as convex (or bounds apply to convex hulls); non-convexity only enlarges the blind spot.
    Required for Hausdorff–support-function control in Theorem 2; Theorems 1 and rank-reversal are stated more assumption-light.
  • standard math Effective dimensionality is the eigenvalue participation ratio of the score correlation matrix, optionally MP-truncated.
    Classical participation-ratio / RMT tool (Bell–Dean, Wegner, Marchenko–Pastur, BBP).
  • domain assumption Hidden capabilities follow a chi-squared / Gaussian projection model; isotropic Σ_hidden minimizes swap probability among fixed-trace covariances (Schur-convexity).
    Underpins Corollary 1 / Proposition 2 swap envelopes; sensitivity across six priors is reported.
  • standard math Coverage f(T)=tr(P_T Σ)/tr(Σ) is monotone submodular, so greedy gets (1−1/e) (Nemhauser; Das–Kempe).
    Standard experimental-design fact used for Theorem 4.
  • standard math C² support functions with curvature ≥κ have Jackson–Bernstein approximation rates on S^{D−1}; optimal recovery matches Kolmogorov n-widths.
    Background for Theorem 3 / Appendix G resolution of Gardner 1.5 for widths/C² bodies.
invented entities (3)
  • LLM capability profile as convex body in R^D with benchmark widths no independent evidence
    purpose: Turns leaderboard scores into geometric tomography so Hausdorff indistinguishability can bound ranking reliability.
    Modeling construct for this paper; independent handle is only indirect (linearization R², hull membership checks, counterfactual disruption correlations).
  • Visible Hausdorff blind spot δ_vis_H on the effective subspace V_eff independent evidence
    purpose: Separates measurable covering error from irreducible orthogonal 2R term reduced by greedy expansion of V_eff.
    Derived quantity from the model rather than a new physical entity; falsifiable via half-splits and covering-radius checks.
  • Stable core of benchmarks under bootstrap greedy (MUSR, GSM8K, IFEval, MMLU on extended suite) independent evidence
    purpose: Operational recommendation for permanent vs rotating evaluations.
    Data-dependent output of the algorithm on one suite family; transfer experiments provide partial external check.

reviewed 2026-07-12 · how reviews work

0 comments
Cite this review

Pith. "Pith review of The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models." pith.science (2026). https://pith.science/paper/AP2TBU3K

@misc{pith2026260605169,
  author       = {Pith},
  title        = {Pith review of: The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AP2TBU3K}},
  note         = {Machine review of arXiv:2606.05169}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We give a stereological theory of LLM benchmark coverage. For any suite with effective dimensionality d_eff, the visible Hausdorff distance between two convex capability profiles consistent with the same scores is bounded by epsilon + C R m^(-1/(d_eff-1)), with matching Lipschitz lower bound. Empirically, three independent leaderboards (Open LLM v2, an extended 12-benchmark suite, LiveBench) all have d_eff in [2.86, 4.80] on their competitive frontier; the structural blind spot exceeds the observed runner-up score gap by two orders of magnitude and dominates statistical noise by 52-127x. Under a chi-squared projection model, the isotropic prior is the optimistic case; across six hidden-capability priors and four ambient dimensions the simulated half-split swap rate of the top two models stays in [0.38, 0.49], and a 500-trial random visible/held-out split shows that 92% of trials swap the top-1 ranking with on average 2.83 of 5 top-5 models changing. A submodular greedy algorithm with the Nemhauser (1 - 1/e) guarantee finds a stable core of 4 benchmarks; 7 of 12 suffice for 90% coverage, and the trained subset transfers across temporal quarters with 93-97% retention. A counterfactual validation across 12 internal benchmarks and 27 Chatbot Arena categories confirms that the eigenstructure predicts which evaluations are irreplaceable (rho = -0.69, p = 0.013 for removal disruption) and which external evaluations bring new information (rho = +0.38). As a second, independent theoretical contribution, we resolve Gardner's Problem 1.5 (1995) for C^2 support functions, establishing the minimax rate Theta(R/(kappa m^(2/(D-1)))) in general dimension via optimal recovery theory on S^(D-1).

Figures

Figures reproduced from arXiv: 2606.05169 by Jason Z Wang.

Figure 1
Figure 1. Figure 1: Eigenvalue spectrum (full and frontier populations, both suites), with the Marchenko– [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: PCA biplot of the 12-benchmark extended frontier. PC1 is the residual g-factor; PC2 [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Left: cumulative coverage of the greedy subset on the extended frontier suite. Right: uncovered eigen-mass before versus after selecting 7 benchmarks; the blind spot shrinks by ≈ 90%. with the empirically most disruptive. Extending to 27 Chatbot Arena categories as external evalu￾ations, blind-spot alignment predicts ranking disruption from adding the external evaluation with ρ = +0.38 (p = 0.053; Appendix… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 1 linked inside Pith

  1. [1]

    Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices

    Jinho Baik, G \'e rard Ben Arous, and Sandrine P \'e ch \'e . Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices. Annals of Probability, 33 0 (5): 0 1643--1697, 2005

  2. [2]

    Atomic vibrations in vitreous silica

    R J Bell and P Dean. Atomic vibrations in vitreous silica. Discussions of the Faraday Society, 50: 0 55--61, 1970

  3. [3]

    On a short-coming of saaty's method of analytic hierarchies

    Valerie Belton and Tony Gear. On a short-coming of saaty's method of analytic hierarchies. Omega, 11 0 (3): 0 228--230, 1983

  4. [4]

    Revealing the structure of language model capabilities

    Ryan Burnell, Han Hao, Andrew R A Conway, and Jos \'e Hern \'a ndez-Orallo. Revealing the structure of language model capabilities. arXiv preprint arXiv:2306.10062, 2024

  5. [5]

    Approximation Theory and Harmonic Analysis on Spheres and Balls

    Feng Dai and Yuan Xu. Approximation Theory and Harmonic Analysis on Spheres and Balls. Springer, 2013

  6. [6]

    Submodular meets spectral: greedy algorithms for subset selection, sparse approximation and dictionary selection

    Abhimanyu Das and David Kempe. Submodular meets spectral: greedy algorithms for subset selection, sparse approximation and dictionary selection. Proceedings of ICML, 2011

  7. [7]

    Moduli of Smoothness

    Z Ditzian and V Totik. Moduli of Smoothness. Springer, 1990

  8. [8]

    Remarks on the analytic hierarchy process

    James S Dyer. Remarks on the analytic hierarchy process. Management Science, 36 0 (3): 0 249--258, 1990

  9. [9]

    A positive answer to the busemann-petty problem in three dimensions

    Richard J Gardner. A positive answer to the busemann-petty problem in three dimensions. Annals of Mathematics, 140 0 (2): 0 435--447, 1994

  10. [10]

    Geometric Tomography

    Richard J Gardner. Geometric Tomography. Cambridge University Press, 1995

  11. [11]

    Geometric Tomography (Second Edition)

    Richard J Gardner. Geometric Tomography (Second Edition). Cambridge University Press, 2006

  12. [12]

    Optimal rates of convergence for convex set estimation from support functions

    Adityanand Guntuboyina. Optimal rates of convergence for convex set estimation from support functions. Annals of Statistics, 40 0 (1): 0 385--411, 2012

  13. [13]

    The theory of ratio scale estimation: Saaty's analytic hierarchy process

    Patrick T Harker and Luis G Vargas. The theory of ratio scale estimation: Saaty's analytic hierarchy process. Management Science, 33 0 (11): 0 1383--1403, 1987

  14. [14]

    A rationale and test for the number of factors in factor analysis

    John L Horn. A rationale and test for the number of factors in factor analysis. Psychometrika, 30: 0 179--185, 1965

  15. [15]

    HypoSpace : underdetermination in enumerable hypothesis spaces, 2025

    HypoSpace Authors . HypoSpace : underdetermination in enumerable hypothesis spaces, 2025. Working paper

  16. [16]

    Dynabench : Rethinking benchmarking in NLP

    Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al. Dynabench : Rethinking benchmarking in NLP . In NAACL, 2021

  17. [17]

    metabench: A sparse benchmark to measure general ability in large language models

    Alex Kipnis, Konstantinos Voudouris, Luca M Schulze Buschoff, and Eric Schulz. metabench: A sparse benchmark to measure general ability in large language models. In ICLR, 2025

  18. [18]

    Intersection bodies, positive definite distributions, and the busemann--petty problem

    Alexander Koldobsky. Intersection bodies, positive definite distributions, and the busemann--petty problem. American Journal of Mathematics, 120 0 (4): 0 827--840, 1998

  19. [19]

    Near-optimal sensor placements in G aussian processes: theory, efficient algorithms and empirical studies

    Andreas Krause, Ajit Singh, and Carlos Guestrin. Near-optimal sensor placements in G aussian processes: theory, efficient algorithms and empirical studies. Journal of Machine Learning Research, 9: 0 235--284, 2008

  20. [20]

    Active evaluation acquisition for efficient llm benchmarking

    Yang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba, and Graham Horwood. Active evaluation acquisition for efficient llm benchmarking. In ICML, 2025

  21. [21]

    Intersection bodies and dual mixed volumes

    Erwin Lutwak. Intersection bodies and dual mixed volumes. Advances in Mathematics, 71 0 (2): 0 232--261, 1988

  22. [22]

    Diameters of compact sets in linear metric spaces

    G G Magaril-Il'yaev. Diameters of compact sets in linear metric spaces. Russian Mathematical Surveys, 34 0 (4): 0 1--40, 1979

  23. [23]

    Optimal recovery of operators and multidimensional carlson type inequalities

    G G Magaril-Il'yaev and K Yu Osipenko. Optimal recovery of operators and multidimensional carlson type inequalities. Journal of Complexity, 22 0 (5): 0 691--703, 2006

  24. [24]

    Distribution of eigenvalues for some sets of random matrices

    V A Mar c enko and L A Pastur. Distribution of eigenvalues for some sets of random matrices. Matematicheskii Sbornik, 114: 0 507--536, 1967

  25. [25]

    The Random Matrix Theory of the Classical Compact Groups

    Elizabeth S Meckes. The Random Matrix Theory of the Classical Compact Groups. Cambridge University Press, 2019

  26. [26]

    A survey of optimal recovery

    Charles A Micchelli and Theodore J Rivlin. A survey of optimal recovery. In Optimal Estimation in Approximation Theory, pages 1--54. Springer, 1977

  27. [27]

    An analysis of approximations for maximizing submodular set functions—i

    George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher. An analysis of approximations for maximizing submodular set functions—i. In Mathematical Programming, volume 14, pages 265--294, 1978

  28. [28]

    Efficient multi-prompt evaluation of llms

    Kanstantsin Pashkovich, Heather Pon-Barry, and Junyi Jessy Li. Efficient multi-prompt evaluation of llms. In NeurIPS, 2024

  29. [29]

    Projection bodies

    C M Petty. Projection bodies. Proceedings of the Colloquium on Convexity, pages 234--241, 1967

  30. [30]

    tinybenchmarks: evaluating llms with fewer examples

    Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. tinybenchmarks: evaluating llms with fewer examples. In ICML, 2024

  31. [31]

    Polynomial approximation on compact manifolds and homogeneous spaces

    David L Ragozin. Polynomial approximation on compact manifolds and homogeneous spaces. Transactions of the American Mathematical Society, 150: 0 41--53, 1970

  32. [32]

    Covering a sphere with spheres

    C A Rogers. Covering a sphere with spheres. Mathematika, 10 0 (2): 0 157--164, 1963

  33. [33]

    Inconsistency and rank preservation

    Thomas L Saaty. Inconsistency and rank preservation. Journal of Mathematical Psychology, 28 0 (2): 0 205--214, 1984

  34. [34]

    u ber die projektionen konvexer k \

    Rolf Schneider. Zur einem problem von shephard \"u ber die projektionen konvexer k \"o rper. Mathematische Zeitschrift, 101: 0 71--82, 1967

  35. [35]

    BenchScope : dimensionality and reliability of LLM benchmark suites

    Yiyang Sha and Wei Zhao. BenchScope : dimensionality and reliability of LLM benchmark suites. arXiv preprint, 2026

  36. [36]

    Information-Based Complexity

    Joseph F Traub, Grzegorz W Wasilkowski, and Henryk Wo\'zniakowski. Information-Based Complexity. Academic Press, 1988

  37. [37]

    Inverse participation ratio in 2+ dimensions

    Franz Wegner. Inverse participation ratio in 2+ dimensions. Zeitschrift f \"u r Physik B Condensed Matter , 36: 0 209--214, 1980

  38. [38]

    A positive solution to the busemann-petty problem in R ^4

    Gaoyong Zhang. A positive solution to the busemann-petty problem in R ^4 . Annals of Mathematics, 149 0 (2): 0 535--543, 1999

  39. [39]

    General scales unlock ai evaluation

    Lexin Zhou, Lorenzo Pacchiardi, Fernando Martinez-Plumed, Jose Hernandez-Orallo, et al. General scales unlock ai evaluation. Nature, 652: 0 58--67, 2026

This paper was first reviewed by grok-4.5 on July 12, 2026.