Pith. sign in

REVIEW 3 major objections 5 minor 64 references

Under target-blind LLM supervision, every learner faces a sample-size-independent minimax risk floor of at least half the model's admissible-label overlap, certifiable from unlabeled inputs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 05:28 UTC pith:GO2NNX72

load-bearing objection Solid classical minimax applied to LLM-judge channels, with honest scope and a real certification procedure; the theory holds under its assumptions, the empirics are narrow by design. the 3 major comments →

arxiv 2607.08961 v1 pith:GO2NNX72 submitted 2026-07-09 cs.LG cs.AImath.STstat.TH

NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision

classification cs.LG cs.AImath.STstat.TH MSC 68T0568Q3262C20
keywords NL-PACspecification ambiguityminimax risk floortarget-blind supervisionLLM-as-a-judgeadmissible labelsPAC learningfinite-sample certificates
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large language models are increasingly used to supply labels and judgments for tasks written in natural language. When a prompt admits more than one reading and the supervision channel does not reveal which reading is operative, collecting more labels only shrinks sampling error; it cannot identify the true target. This paper introduces NL-PAC, which treats a frozen model's thresholded decoding probabilities as the definition of admissible labels at each input. The mass of inputs where multiple labels clear the threshold equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner's worst-case risk is at least half that diameter at every sample size. The exact randomized minimax value over the admissible core is attained by a data-independent strategy, and both the diameter and the exact value admit finite-sample lower certificates from held-out unlabeled data. An audit of a frozen Qwen 2.5–3B judge produces a positive model-relative certificate for one prespecified prompt and zero for a paraphrase and exact-rule controls; a bridge audit shows that supplied coherent reading clauses fail the admissibility condition needed to transfer the certificate.

Core claim

For a fixed model, prompt, and decoding threshold, the probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class. Under target-blind supervision every learner incurs worst-case risk of at least half this diameter at every sample size; the exact randomized minimax risk over the admissible core equals the expected non-modal admissible mass and is attained by a data-independent strategy. Both quantities can be certified from held-out unlabeled inputs.

What carries the argument

The admissible-overlap mass D*_τ = P(|A_τ(X)| ≥ 2), where A_τ(x) is the set of labels whose model decoding probability meets threshold τ. This mass equals the diameter of the almost-everywhere admissible core; half of it (or the sharper multiplicity-weighted value V*_τ = E[1 − 1/k(X)]) is the blind-channel minimax risk floor, made certifiable by Hoeffding bounds on held-out unlabeled inputs.

Load-bearing premise

The supervision channel never reveals which reading of the prompt is operative, so distinct admissible targets produce identical observation laws; if the channel leaks even partial reading identity, the sample-size-independent floor need not hold.

What would settle it

On the same frozen model, prompt, threshold, and input distribution, find either (a) a positive certified floor for an exact-rule control whose two readings never disagree, or (b) a learner that, under a truly target-blind channel, drives worst-case risk over the admissible core substantially below V*_τ for large sample size.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • More labels from the same ambiguous prompt–model channel cannot erase the certified floor; lowering it requires changing the information structure (clarify the prompt, reveal the reading, or switch model).
  • A positive certificate for a deployed judge quantifies residual worst-case error that remains even with infinite data under that channel.
  • Zero certificates on exact-rule controls show the procedure does not invent floors where the specification leaves no room for ambiguity.
  • Transferring the pointwise model-relative floor to coherent global readings requires a separate coverage-and-admissibility bridge that can fail.
  • The guarantee is configuration-specific: model, prompt, threshold, and input distribution must be re-audited after any change.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Open-weight judges that expose logits support a cheap correction-free audit; sampling-only APIs face a depth barrier that can make certification impractical at scale.
  • Prompt engineering can be reframed as driving admissible-overlap mass below a stated risk tolerance rather than maximizing average accuracy alone.
  • If the same obstruction appears across model families on natural distributions, specification ambiguity may be a first-order bottleneck for LLM-as-judge pipelines that more labeled data cannot remove.
  • The least-favorable cyclic selector construction supplies a template for certifying identification floors in other set-valued supervision settings beyond language models.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces NL-PAC, a framework that treats a frozen model–prompt–threshold triple as defining pointwise admissible label sets A_τ(x) and a candidate target class. It proves that the admissible-overlap mass D*_τ equals the diameter of the almost-everywhere admissible core Sel_τ, that under fixed-description target-blind supervision every learner has worst-case risk at least D*_τ/2 at every sample size, and that the exact randomized minimax risk over Sel_τ equals V*_τ = E[1−1/k(X)], attained by a data-independent uniform draw over admissible labels. Finite-sample Hoeffding certificates for these quantities are given from held-out unlabeled inputs (exposed and sampled-decoding modes). An audit of frozen Qwen 2.5–3B yields a positive model-relative certificate for one prespecified moderation prompt and zero for a paraphrase and exact-rule controls; a held-out bridge audit finds that supplied coherent reading clauses fail the admissibility condition needed to transfer the certificate.

Significance. If the results hold as scoped, the paper supplies a clean, certifiable decision-theoretic account of an identification obstruction that is endogenous to LLM-mediated supervision rather than a classical noise budget. The core geometry (Theorem 2.11), two-point floor (Theorem 3.2), exact minimax identity (Theorem 3.5), and Hoeffding certificates (Theorems 4.1, 4.6, 4.9) are carefully stated with tightness claims and deferred proofs; the exact value V*_τ and its multiplicity sharpening are useful when higher-order overlap has mass. The empirics are honestly delimited (null controls, bridge failure reported, sampled-decoding vacuity flagged), and a reproducibility archive is promised. The contribution is primarily theoretical: it turns model-admissible ambiguity into an auditable minimax floor under an explicit target-blindness assumption, with clear scope limits on transfer to human or coherent readings.

major comments (3)
  1. [Section 5, Theorem 3.9] Section 5 and Theorem 3.9: the held-out bridge reports η̂_U = 0 and ζ̂_U = 0.4904, so |V*_τ − V_blind(C)| ≤ 0.4904 is nearly vacuous. The abstract and introduction frame the problem as specification ambiguity and multiple readings, but the only positive certificate is purely model-relative and does not transfer to the supplied coherent clauses. This is load-bearing for interpretation: either enlarge/validate the reading pool so that held-out admissibility is small, or restructure the abstract and claims so that the primary object is model-admissible overlap under target blindness, with coherent-reading transfer clearly marked as open.
  2. [Section 5, Appendix C.5] Section 5 / C.5: the principal positive certificate uses N = 100 synthetic borderline inputs on a single 3B judge under exposed first-token probabilities; sampled decoding is vacuous at r = 3 (finite-depth term saturates). Exact-rule controls and the controlled pipeline check (C.1) are well designed, but stability across model families and naturally occurring input distributions is untested. For the empirical claim that a real judge can exhibit a positive, non-spurious floor, at least one additional model or a natural-distribution audit (even at modest N) would substantially strengthen the paper; otherwise the empirics should be labeled more explicitly as existence probes rather than validation of practical magnitude.
  3. [Proposition 2.4, Theorems 3.2 and 3.5] Proposition 2.4 / Equation (1) and Theorems 3.2, 3.5: target blindness of the fixed-description channel is the load-bearing assumption for sample-size-independent floors. The paper correctly treats it as a modeling restriction and notes LMaaS opacity, but does not discuss how often deployed judge pipelines leak reading identity (e.g., via returned rationales, multi-turn context, or target-conditioned decoding). A short subsection delimiting when the base channel is approximately blind versus when partial distinguishability would restore a vanishing TV term would make the applicability boundary operational for practitioners.
minor comments (5)
  1. [Figure 1] Figure 1: the positive-certificate region and the comparison point τ = 0.20 are clear, but the y-axis label “estimated overlap mass D” should match the notation D*_τ / D̂* used in the text, and the caption could state the exact ε_N value used for the shaded radius.
  2. [Table 1] Table 1 is helpful; consider adding a one-line pointer from each class to the governing theorem number in the main text as well as in the table, to reduce cross-referencing cost.
  3. [Sections 3.3 and 4.3] Notation: C is overloaded as both |Y| (Theorem 4.5) and a coherent family (Theorem 3.9). The text usually disambiguates by context, but a local rename (e.g., C_Y vs C_read) would help.
  4. [Section 5, C.10] Appendix C.10: the constrained first-token verbalizer is essential to the correction-free audit; a one-sentence reminder in Section 5 that the certificate is relative to this declared-label channel (not unrestricted continuations) would prevent misreading.
  5. [Theorems 2.11, 3.5, 4.6] Several deferred proofs are marked “proof sketch; full proof in Section B” in the main text; ensure every sketch’s key identity (e.g., maximal-spread disagreement set, cyclic selector uniformity) is stated fully enough that a reader can verify without the appendix if desired.

Circularity Check

0 steps flagged

No significant circularity: geometric identities and classical minimax reductions are proved, not fitted or self-defined as predictions.

full rationale

The load-bearing chain is: (i) A_τ from thresholded decoding defines Sel_τ and D*_τ by definition; (ii) diam(Sel_τ)=D*_τ via the maximal-spread pair (Lemma 2.10, Theorem 2.11)—a geometric identity proved by exhibiting a pair that disagrees exactly on the overlap set, not a fitted tautology or self-definitional claim that X derives Y when X is defined as Y; (iii) target blindness (Prop. 2.4) makes observation laws independent of the operative selector, so Le Cam two-point and least-favorable cyclic-selector arguments give the sample-size-independent floors D*_τ/2 and V*_τ=E[1−1/k(X)] (Theorems 3.2, 3.5)—classical decision theory applied to model-induced sets, with achievability by an explicit data-independent kernel; (iv) finite-sample certificates are one-sided Hoeffding bounds on the same observable admissible-set statistics on held-out unlabeled inputs (Theorems 4.1, 4.6, 4.9), i.e. estimation of a theoretically derived quantity, not a fit renamed as prediction. Exact-rule controls certify zero and the coherent-reading bridge fails on held-out admissibility (ζ̂_U=0.4904), so the paper does not smuggle transfer. No load-bearing self-citation uniqueness theorem, no ansatz via author citation, no renaming of a known empirical pattern as the central result. The derivation is self-contained against its stated assumptions.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 3 invented entities

The theory rests on standard PAC/decision-theoretic background plus domain modeling choices that define the LLM channel (target blindness, thresholded decoding admissibility, finite labels). Free parameters are design knobs of the audit (τ, ζ, δ, N, r, ξ), not fitted to force the central identity. Invented entities are the NL-PAC objects themselves; they have operational handles (observable admissible sets, certificates) but no independent human-construct validity without external validation, which the paper states.

free parameters (4)
  • admissibility threshold τ = sweep {0.10,...,0.40}; comparison point 0.20
    Prespecified design choice in (0,1/2); controls which labels enter A_τ and thus D*_τ. Sweep reported; not fitted to maximize the certificate post hoc in the main claim.
  • coverage tolerance ζ
    Tolerance for the open neighborhood F_τ,ζ; enters diameter bracket as +2ζ. Fixed design parameter of the class, not estimated from labels.
  • audit confidence δ and sample size N = δ=0.10, N=100 (main audit)
    Hoeffding radius ε_N = sqrt(log(2/δ)/(2N)); main exposed audit uses δ=0.10, N=100. Standard statistical design knobs.
  • sampled-decoding margin ξ and depth r = r=3, ξ=0.05 (inconclusive mode)
    Plug-in radius terms κ(ξ)+2C exp(-2rξ²). At r=3 the depth term saturates; positive certificates would need r≈877 at ξ=0.05. Design parameters of the access mode.
axioms (5)
  • domain assumption Target blindness of the fixed-description channel: Law(ρ(J,X)|f_NL,X) does not depend on the operative target f (Proposition 2.4).
    Load-bearing for sample-size-independent floors; automatic for description-only oracles but fails if supervision reveals the reading.
  • domain assumption Assumptions (A1)–(A3): measurability of π_LLM, conditional independence of oracle draws, finite nonempty admissible sets a.e.
    Standard technical conditions so selectors and diameters are well-defined measurable objects.
  • standard math Classical two-point / least-favorable-prior minimax identities for finite experiments (Le Cam, Wald, Blackwell–Girshick).
    Used to obtain D/2 floors and exact V*_τ; not re-proved from scratch beyond the model-induced selector family.
  • standard math Hoeffding concentration for bounded i.i.d. indicators on held-out unlabeled inputs.
    Underpins all finite-sample certificates; empirical-Bernstein alternatives noted but not required.
  • ad hoc to paper η-uniform coverage plus ζ-admissibility of a finite coherent-reading pool for transferring V*_τ to V_blind(C) (Theorem 3.9).
    Bridge condition specific to this framework; empirically fails for the supplied two-clause pool on held-out admissibility.
invented entities (3)
  • NL-PAC learning problem and model-admissible labeling class F_τ,ζ / core Sel_τ no independent evidence
    purpose: Make the specification–channel pair the object of PAC analysis via thresholded model decoding.
    New formalization; operational via observable A_τ but model-relative by construction.
  • Admissible-overlap mass D*_τ and exact blind value V*_τ independent evidence
    purpose: Scalar geometry and minimax risk of the admissible core under target blindness.
    Defined from A_τ; equalities and certificates are proved, not fitted. Independent of human ground truth unless externally validated.
  • Coherent-reading subclass F_read and η-coverage bridge no independent evidence
    purpose: Relate pointwise selectors to global natural-language reading clauses.
    Pool-relative construction; paper’s own audit finds the supplied pool fails admissibility transfer.

pith-pipeline@v1.1.0-grok45 · 46574 in / 4072 out tokens · 55754 ms · 2026-07-13T05:28:31.296042+00:00 · methodology

0 comments
read the original abstract

Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.

Figures

Figures reproduced from arXiv: 2607.08961 by Berkay Anahtarci.

Figure 1
Figure 1. Figure 1: Correction-free exposed-probability audit at [PITH_FULL_IMAGE:figures/full_fig_p021_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Cost–precision trade-off of the audit at [PITH_FULL_IMAGE:figures/full_fig_p039_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

64 extracted references · 9 canonical work pages

  1. [1]

    A theory of PAC learnability of partial concept classes

    Noga Alon, Steve Hanneke, Ron Holzman, and Shay Moran. A theory of PAC learnability of partial concept classes. In 2021 IEEE 62nd Annual Symposium on Foundations of Computer Science ( FOCS ) , pages 658--671. IEEE, 2022. doi:10.1109/FOCS52979.2021.00070

  2. [2]

    Angelopoulos and Stephen Bates

    Anastasios N. Angelopoulos and Stephen Bates. Conformal prediction: A gentle introduction. Foundations and Trends in Machine Learning, 16 0 (4): 0 494--591, 2023. doi:10.1561/2200000101

  3. [3]

    Tsybakov

    Jean-Yves Audibert and Alexandre B. Tsybakov. Fast learning rates for plug-in classifiers. The Annals of Statistics, 35 0 (2): 0 608--633, 2007. doi:10.1214/009053606000001217

  4. [4]

    Stop measuring calibration when humans disagree

    Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fern \'a ndez. Stop measuring calibration when humans disagree. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1892--1915, 2022. doi:10.18653/v1/2022.emnlp-main.124

  5. [5]

    Concentration inequalities for sampling without replacement

    R \'e mi Bardenet and Odalric-Ambrym Maillard. Concentration inequalities for sampling without replacement. Bernoulli, 21 0 (3): 0 1361--1385, 2015. doi:10.3150/14-BEJ605

  6. [6]

    Minimax regret of finite partial-monitoring games in stochastic environments

    G \'a bor Bart \'o k, D \'a vid P \'a l, and Csaba Szepesv \'a ri. Minimax regret of finite partial-monitoring games in stochastic environments. In Sham M. Kakade and Ulrike von Luxburg, editors, Proceedings of the 24th Annual Conference on Learning Theory, volume 19 of Proceedings of Machine Learning Research, pages 133--154, Budapest, Hungary, 2011. PML...

  7. [7]

    Valid post-selection inference

    Richard Berk, Lawrence Brown, Andreas Buja, Kai Zhang, and Linda Zhao. Valid post-selection inference. The Annals of Statistics, 41 0 (2): 0 802--837, 2013. doi:10.1214/12-aos1077

  8. [8]

    Comparison of experiments

    David Blackwell. Comparison of experiments. In Jerzy Neyman, editor, Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, pages 93--102, Berkeley, 1951. University of California Press. doi:10.1525/9780520411586-009

  9. [9]

    David Blackwell and M. A. Girshick. Theory of Games and Statistical Decisions. Wiley, New York, 1954

  10. [10]

    Bshouty, Nadav Eiron, and Eyal Kushilevitz

    Nader H. Bshouty, Nadav Eiron, and Eyal Kushilevitz. PAC learning with nasty noise. Theoretical Computer Science, 288 0 (2): 0 255--275, 2002. doi:10.1016/S0304-3975(01)00403-0

  11. [11]

    Learning with bounded instance and label-dependent label noise

    Jiacheng Cheng, Tongliang Liu, Kotagiri Ramamohanarao, and Dacheng Tao. Learning with bounded instance and label-dependent label noise. In Hal Daum \'e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 1789--1799. PMLR, 2020. URL https://proceed...

  12. [12]

    Diagnosing the reliability of LLM -as-a-judge via item response theory, 2026

    Junhyuk Choi, Sohhyung Park, Chanhee Cho, Hyeonchu Park, and Bugeun Kim. Diagnosing the reliability of LLM -as-a-judge via item response theory, 2026

  13. [13]

    Learning from partial labels

    Timothee Cour, Ben Sapp, and Ben Taskar. Learning from partial labels. Journal of Machine Learning Research, 12 0 (42): 0 1501--1536, 2011. URL http://jmlr.org/papers/v12/cour11a.html

  14. [14]

    Underspecification presents challenges for credibility in modern machine learning

    Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, et al. Underspecification presents challenges for credibility in modern machine learning. Journal of Machine Learning Research, 23 0 (226): 0 1--61, 2022. URL https://jmlr.org/papers/v23/20-1335.html

  15. [15]

    A. P. Dawid and A. M. Skene. Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society. Series C (Applied Statistics), 28 0 (1): 0 20--28, 1979. doi:10.2307/2346806

  16. [16]

    Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the data

    Florian Eddie Dorner, Vivian Nastl, and Moritz Hardt. Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the data. In International Conference on Learning Representations, pages 26467--26491, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/hash/4264ee4376776907c0b87ed70b959585-Abstract-Conference.html

  17. [17]

    Adversarial multiclass classification: A risk minimization perspective

    Rizal Fathony, Anqi Liu, Kaiser Asif, and Brian Ziebart. Adversarial multiclass classification: A risk minimization perspective. In D. Lee, M. Sugiyama, U. von Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URL https://proceedings.neurips.cc/paper_files/paper/2016/fi...

  18. [18]

    Ferguson

    Thomas S. Ferguson. Mathematical Statistics: A Decision Theoretic Approach. Academic Press, New York, 1967

  19. [19]

    Towards provably unbiased llm judges via bias-bounded evaluation, 2026

    Benjamin Feuer, Lucas Rosenblatt, and Oussama Elachqar. Towards provably unbiased llm judges via bias-bounded evaluation, 2026

  20. [20]

    Maxmin expected utility with non-unique prior

    Itzhak Gilboa and David Schmeidler. Maxmin expected utility with non-unique prior. Journal of Mathematical Economics, 18 0 (2): 0 141--153, 1989. doi:10.1016/0304-4068(89)90018-9

  21. [21]

    A survey on LLM -as-a-judge

    Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Zhouchi Lin, Bowen Zhang, Lionel Ni, Wen Gao, Yuanzhuo Wang, and Jian Guo. A survey on LLM -as-a-judge. The Innovation, 7 0 (6): 0 101253, 2026. doi:10.1016/j.xinn.2025.101253

  22. [22]

    Validating LLM -as-a-judge systems under rating indeterminacy

    Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Steven Wu, and Alexandra Chouldechova. Validating LLM -as-a-judge systems under rating indeterminacy. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors, Advances in Neural Information Processing Systems, volume 38, pages 112282--112350. Curran Associate...

  23. [23]

    Theory of disagreement-based active learning

    Steve Hanneke. Theory of disagreement-based active learning. Foundations and Trends in Machine Learning, 7 0 (2--3): 0 131--309, 2014. doi:10.1561/2200000037

  24. [24]

    Probability inequalities for sums of bounded random variables

    Wassily Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58 0 (301): 0 13--30, 1963. doi:10.1080/01621459.1963.10500830

  25. [25]

    Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon

    Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics, 49 0 (2): 0 1055--1080, 2021. doi:10.1214/20-aos1991

  26. [26]

    Learning from imprecise and fuzzy observations: Data disambiguation through generalized loss minimization

    Eyke H \"u llermeier. Learning from imprecise and fuzzy observations: Data disambiguation through generalized loss minimization. International Journal of Approximate Reasoning, 55 0 (7): 0 1519--1534, 2014. doi:10.1016/j.ijar.2013.09.003

  27. [27]

    Imbens and Charles F

    Guido W. Imbens and Charles F. Manski. Confidence intervals for partially identified parameters. Econometrica, 72 0 (6): 0 1845--1857, 2004. doi:10.1111/j.1468-0262.2004.00555.x

  28. [28]

    Learning in the presence of malicious errors

    Michael Kearns and Ming Li. Learning in the presence of malicious errors. SIAM Journal on Computing, 22 0 (4): 0 807--837, 1993. doi:10.1137/0222052

  29. [29]

    Kearns and Umesh V

    Michael J. Kearns and Umesh V. Vazirani. An Introduction to Computational Learning Theory. MIT Press, 1994

  30. [30]

    Prometheus: Inducing fine-grained evaluation capability in language models

    Seungone Kim, Jay Shin, yejin cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, S Shin, Ryan, Sungdong Kim, James Thorne, and Minjoon Seo. Prometheus: Inducing fine-grained evaluation capability in language models. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun, editors, International Conference on Learning Representations, vo...

  31. [31]

    Active task disambiguation with LLM s

    Katarzyna Kobalczyk, Nicol \'a s Astorga, Tennison Liu, and Mihaela van der Schaar. Active task disambiguation with LLM s. In Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, editors, International Conference on Learning Representations, volume 2025, pages 37823--37847, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/file/5e07476b6bd2497e1fbd11b8...

  32. [32]

    Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

    Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In International Conference on Learning Representations, 2023. Spotlight

  33. [33]

    Language-models-as-a-service: Overview of a new paradigm and its challenges

    Emanuele La Malfa, Aleksandar Petrov, Simon Frieder, Christoph Weinhuber, Ryan Burnell, Raza Nazar, Anthony Cohn, Nigel Shadbolt, and Michael Wooldridge. Language-models-as-a-service: Overview of a new paradigm and its challenges. Journal of Artificial Intelligence Research, 80: 0 1497--1523, 2024. doi:10.1613/jair.1.15865

  34. [34]

    Eliciting human preferences with language models

    Belinda Li, Alex Tamkin, Noah Goodman, and Jacob Andreas. Eliciting human preferences with language models. In Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, editors, International Conference on Learning Representations, volume 2025, pages 80984--81013, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/file/c9867d5a22653ce98b02595061e40f12-Paper-...

  35. [35]

    Smith, and Yejin Choi

    Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. We're afraid language models aren't modeling ambiguity. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023 a . doi:10.18653/v1/2023.emnlp-main.51

  36. [36]

    Learnability of the superset label learning problem

    Liping Liu and Thomas Dietterich. Learnability of the superset label learning problem. In Eric P. Xing and Tony Jebara, editors, Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, pages 1629--1637, Bejing, China, 2014. PMLR. URL https://proceedings.mlr.press/v32/liug14.html

  37. [37]

    G-Eval : NLG evaluation using GPT -4 with better human alignment

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval : NLG evaluation using GPT -4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2511--2522, 2023 b . doi:10.18653/v1/2023.emnlp-main.153

  38. [38]

    Tsybakov

    Enno Mammen and Alexandre B. Tsybakov. Smooth discrimination analysis. The Annals of Statistics, 27 0 (6): 0 1808--1829, 1999. doi:10.1214/aos/1017939240

  39. [39]

    Charles F. Manski. Partial Identification of Probability Distributions. Springer Series in Statistics. Springer, New York, 2003. doi:10.1007/b97478

  40. [40]

    Empirical Bernstein bounds and sample-variance penalization

    Andreas Maurer and Massimiliano Pontil. Empirical Bernstein bounds and sample-variance penalization. In Proceedings of the 22nd Annual Conference on Learning Theory (COLT), 2009. URL https://www.cs.mcgill.ca/ colt2009/papers/012.pdf

  41. [41]

    Minimax risk classifiers with 0--1 loss

    Santiago Mazuelas, Mauricio Romero, and Peter Gr \"u nwald. Minimax risk classifiers with 0--1 loss. Journal of Machine Learning Research, 24 0 (208): 0 1--48, 2023. URL https://www.jmlr.org/papers/volume24/22-0339/22-0339.pdf

  42. [42]

    Mitchell

    Tom M. Mitchell. Generalization as search. Artificial Intelligence, 18 0 (2): 0 203--226, 1982. doi:10.1016/0004-3702(82)90040-6

  43. [43]

    Random Sets in Econometrics

    Ilya Molchanov and Francesca Molinari. Random Sets in Econometrics. Cambridge University Press, 2018. doi:10.1017/9781316392973

  44. [44]

    Learning with noisy labels

    Nagarajan Natarajan, Inderjit S Dhillon, Pradeep K Ravikumar, and Ambuj Tewari. Learning with noisy labels. In C. J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger, editors, Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013. URL https://proceedings.neurips.cc/paper_files/paper/2013/file/3871bd6401...

  45. [45]

    Yixin Nie, Xiang Zhou, and Mohit Bansal. What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131--9143, 2020. doi:10.18653/v1/2020.emnlp-main.734

  46. [46]

    Norman, Michael U

    Justin D. Norman, Michael U. Rivera, and D. Alex Hughes. Reliability without validity: A systematic, large-scale evaluation of LLM -as-a-judge models across agreement, consistency, and bias, 2026

  47. [47]

    The effects of reward misspecification: Mapping and mitigating misaligned models

    Alexander Pan, Kush Bhatia, and Jacob Steinhardt. The effects of reward misspecification: Mapping and mitigating misaligned models. In International Conference on Learning Representations (ICLR), 2022

  48. [48]

    Inherent disagreements in human textual inferences

    Ellie Pavlick and Tom Kwiatkowski. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7: 0 677--694, 2019. doi:10.1162/tacl_a_00293

  49. [49]

    The ``problem'' of human label variation: On ground truth in data, modeling and evaluation

    Barbara Plank. The ``problem'' of human label variation: On ground truth in data, modeling and evaluation. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022. doi:10.18653/v1/2022.emnlp-main.731

  50. [50]

    Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher R \'e

    Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher R \'e . Snorkel: Rapid training data creation with weak supervision. Proceedings of the VLDB Endowment, 11 0 (3): 0 269--282, 2017. doi:10.14778/3157794.3157797

  51. [51]

    Truth-tracking with non-expert information sources

    Joseph Singleton and Richard Booth. Truth-tracking with non-expert information sources. Journal of Artificial Intelligence Research, 81: 0 619--641, 2024. doi:10.1613/jair.1.15273

  52. [52]

    On general minimax theorems

    Maurice Sion. On general minimax theorems. Pacific Journal of Mathematics, 8 0 (1): 0 171--176, 1958. doi:10.2140/pjm.1958.8.171

  53. [53]

    Minimax regret treatment choice with finite samples

    J \"o rg Stoye. Minimax regret treatment choice with finite samples. Journal of Econometrics, 151 0 (1): 0 70--81, 2009 a . doi:10.1016/j.jeconom.2009.02.013

  54. [54]

    More on confidence intervals for partially identified parameters

    J \"o rg Stoye. More on confidence intervals for partially identified parameters. Econometrica, 77 0 (4): 0 1299--1315, 2009 b . doi:10.3982/ECTA7347

  55. [55]

    Tsybakov

    Alexandre B. Tsybakov. Introduction to Nonparametric Estimation. Springer Series in Statistics. Springer New York, New York, NY, 1 edition, 2009. ISBN 978-0-387-79052-7. doi:10.1007/b13794. URL https://link.springer.com/book/10.1007/b13794

  56. [56]

    Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio

    Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. Learning from disagreement: A survey. Journal of Artificial Intelligence Research, 72: 0 1385--1470, 2021. doi:10.1613/jair.1.12752

  57. [57]

    Leslie G. Valiant. A theory of the learnable. Communications of the ACM , 27 0 (11): 0 1134--1142, 1984. doi:10.1145/1968.1972

  58. [58]

    A new learning paradigm: Learning using privileged information

    Vladimir Vapnik and Akshay Vashist. A new learning paradigm: Learning using privileged information. Neural Networks, 22 0 (5--6): 0 544--557, 2009. doi:10.1016/j.neunet.2009.06.042

  59. [59]

    Algorithmic Learning in a Random World

    Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer New York, New York, NY, 1 edition, 2005. ISBN 978-0-387-25061-8. doi:10.1007/b106715. URL https://link.springer.com/book/10.1007/b106715

  60. [60]

    Statistical Decision Functions

    Abraham Wald. Statistical Decision Functions. Wiley, New York, 1950

  61. [61]

    Estimating means of bounded random variables by betting

    Ian Waudby-Smith and Aaditya Ramdas. Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (1): 0 1--27, 2024. doi:10.1093/jrsssb/qkad009

  62. [62]

    Beyond consensus: Perspectivist modeling and evaluation of annotator disagreement in nlp, 2026

    Yinuo Xu and David Jurgens. Beyond consensus: Perspectivist modeling and evaluation of annotator disagreement in nlp, 2026

  63. [63]

    What prompts don't say: Understanding and managing underspecification in LLM prompts

    Chenyang Yang, Yike Shi, Qianou Ma, Michael Xieyang Liu, Christian K \"a stner, and Tongshuang Wu. What prompts don't say: Understanding and managing underspecification in LLM prompts. In Findings of the Association for Computational Linguistics: ACL 2026, pages 9072--9101. Association for Computational Linguistics, 2026. doi:10.18653/v1/2026.findings-acl.441

  64. [64]

    Information-theoretic distinctions between deception and confusion

    Robin Young. Information-theoretic distinctions between deception and confusion. In Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, and Dhirendra Pratap Singh, editors, Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conferen...