REVIEW 3 major objections 5 minor 64 references
Under target-blind LLM supervision, every learner faces a sample-size-independent minimax risk floor of at least half the model's admissible-label overlap, certifiable from unlabeled inputs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 05:28 UTC pith:GO2NNX72
load-bearing objection Solid classical minimax applied to LLM-judge channels, with honest scope and a real certification procedure; the theory holds under its assumptions, the empirics are narrow by design. the 3 major comments →
NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For a fixed model, prompt, and decoding threshold, the probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class. Under target-blind supervision every learner incurs worst-case risk of at least half this diameter at every sample size; the exact randomized minimax risk over the admissible core equals the expected non-modal admissible mass and is attained by a data-independent strategy. Both quantities can be certified from held-out unlabeled inputs.
What carries the argument
The admissible-overlap mass D*_τ = P(|A_τ(X)| ≥ 2), where A_τ(x) is the set of labels whose model decoding probability meets threshold τ. This mass equals the diameter of the almost-everywhere admissible core; half of it (or the sharper multiplicity-weighted value V*_τ = E[1 − 1/k(X)]) is the blind-channel minimax risk floor, made certifiable by Hoeffding bounds on held-out unlabeled inputs.
Load-bearing premise
The supervision channel never reveals which reading of the prompt is operative, so distinct admissible targets produce identical observation laws; if the channel leaks even partial reading identity, the sample-size-independent floor need not hold.
What would settle it
On the same frozen model, prompt, threshold, and input distribution, find either (a) a positive certified floor for an exact-rule control whose two readings never disagree, or (b) a learner that, under a truly target-blind channel, drives worst-case risk over the admissible core substantially below V*_τ for large sample size.
If this is right
- More labels from the same ambiguous prompt–model channel cannot erase the certified floor; lowering it requires changing the information structure (clarify the prompt, reveal the reading, or switch model).
- A positive certificate for a deployed judge quantifies residual worst-case error that remains even with infinite data under that channel.
- Zero certificates on exact-rule controls show the procedure does not invent floors where the specification leaves no room for ambiguity.
- Transferring the pointwise model-relative floor to coherent global readings requires a separate coverage-and-admissibility bridge that can fail.
- The guarantee is configuration-specific: model, prompt, threshold, and input distribution must be re-audited after any change.
Where Pith is reading between the lines
- Open-weight judges that expose logits support a cheap correction-free audit; sampling-only APIs face a depth barrier that can make certification impractical at scale.
- Prompt engineering can be reframed as driving admissible-overlap mass below a stated risk tolerance rather than maximizing average accuracy alone.
- If the same obstruction appears across model families on natural distributions, specification ambiguity may be a first-order bottleneck for LLM-as-judge pipelines that more labeled data cannot remove.
- The least-favorable cyclic selector construction supplies a template for certifying identification floors in other set-valued supervision settings beyond language models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces NL-PAC, a framework that treats a frozen model–prompt–threshold triple as defining pointwise admissible label sets A_τ(x) and a candidate target class. It proves that the admissible-overlap mass D*_τ equals the diameter of the almost-everywhere admissible core Sel_τ, that under fixed-description target-blind supervision every learner has worst-case risk at least D*_τ/2 at every sample size, and that the exact randomized minimax risk over Sel_τ equals V*_τ = E[1−1/k(X)], attained by a data-independent uniform draw over admissible labels. Finite-sample Hoeffding certificates for these quantities are given from held-out unlabeled inputs (exposed and sampled-decoding modes). An audit of frozen Qwen 2.5–3B yields a positive model-relative certificate for one prespecified moderation prompt and zero for a paraphrase and exact-rule controls; a held-out bridge audit finds that supplied coherent reading clauses fail the admissibility condition needed to transfer the certificate.
Significance. If the results hold as scoped, the paper supplies a clean, certifiable decision-theoretic account of an identification obstruction that is endogenous to LLM-mediated supervision rather than a classical noise budget. The core geometry (Theorem 2.11), two-point floor (Theorem 3.2), exact minimax identity (Theorem 3.5), and Hoeffding certificates (Theorems 4.1, 4.6, 4.9) are carefully stated with tightness claims and deferred proofs; the exact value V*_τ and its multiplicity sharpening are useful when higher-order overlap has mass. The empirics are honestly delimited (null controls, bridge failure reported, sampled-decoding vacuity flagged), and a reproducibility archive is promised. The contribution is primarily theoretical: it turns model-admissible ambiguity into an auditable minimax floor under an explicit target-blindness assumption, with clear scope limits on transfer to human or coherent readings.
major comments (3)
- [Section 5, Theorem 3.9] Section 5 and Theorem 3.9: the held-out bridge reports η̂_U = 0 and ζ̂_U = 0.4904, so |V*_τ − V_blind(C)| ≤ 0.4904 is nearly vacuous. The abstract and introduction frame the problem as specification ambiguity and multiple readings, but the only positive certificate is purely model-relative and does not transfer to the supplied coherent clauses. This is load-bearing for interpretation: either enlarge/validate the reading pool so that held-out admissibility is small, or restructure the abstract and claims so that the primary object is model-admissible overlap under target blindness, with coherent-reading transfer clearly marked as open.
- [Section 5, Appendix C.5] Section 5 / C.5: the principal positive certificate uses N = 100 synthetic borderline inputs on a single 3B judge under exposed first-token probabilities; sampled decoding is vacuous at r = 3 (finite-depth term saturates). Exact-rule controls and the controlled pipeline check (C.1) are well designed, but stability across model families and naturally occurring input distributions is untested. For the empirical claim that a real judge can exhibit a positive, non-spurious floor, at least one additional model or a natural-distribution audit (even at modest N) would substantially strengthen the paper; otherwise the empirics should be labeled more explicitly as existence probes rather than validation of practical magnitude.
- [Proposition 2.4, Theorems 3.2 and 3.5] Proposition 2.4 / Equation (1) and Theorems 3.2, 3.5: target blindness of the fixed-description channel is the load-bearing assumption for sample-size-independent floors. The paper correctly treats it as a modeling restriction and notes LMaaS opacity, but does not discuss how often deployed judge pipelines leak reading identity (e.g., via returned rationales, multi-turn context, or target-conditioned decoding). A short subsection delimiting when the base channel is approximately blind versus when partial distinguishability would restore a vanishing TV term would make the applicability boundary operational for practitioners.
minor comments (5)
- [Figure 1] Figure 1: the positive-certificate region and the comparison point τ = 0.20 are clear, but the y-axis label “estimated overlap mass D” should match the notation D*_τ / D̂* used in the text, and the caption could state the exact ε_N value used for the shaded radius.
- [Table 1] Table 1 is helpful; consider adding a one-line pointer from each class to the governing theorem number in the main text as well as in the table, to reduce cross-referencing cost.
- [Sections 3.3 and 4.3] Notation: C is overloaded as both |Y| (Theorem 4.5) and a coherent family (Theorem 3.9). The text usually disambiguates by context, but a local rename (e.g., C_Y vs C_read) would help.
- [Section 5, C.10] Appendix C.10: the constrained first-token verbalizer is essential to the correction-free audit; a one-sentence reminder in Section 5 that the certificate is relative to this declared-label channel (not unrestricted continuations) would prevent misreading.
- [Theorems 2.11, 3.5, 4.6] Several deferred proofs are marked “proof sketch; full proof in Section B” in the main text; ensure every sketch’s key identity (e.g., maximal-spread disagreement set, cyclic selector uniformity) is stated fully enough that a reader can verify without the appendix if desired.
Circularity Check
No significant circularity: geometric identities and classical minimax reductions are proved, not fitted or self-defined as predictions.
full rationale
The load-bearing chain is: (i) A_τ from thresholded decoding defines Sel_τ and D*_τ by definition; (ii) diam(Sel_τ)=D*_τ via the maximal-spread pair (Lemma 2.10, Theorem 2.11)—a geometric identity proved by exhibiting a pair that disagrees exactly on the overlap set, not a fitted tautology or self-definitional claim that X derives Y when X is defined as Y; (iii) target blindness (Prop. 2.4) makes observation laws independent of the operative selector, so Le Cam two-point and least-favorable cyclic-selector arguments give the sample-size-independent floors D*_τ/2 and V*_τ=E[1−1/k(X)] (Theorems 3.2, 3.5)—classical decision theory applied to model-induced sets, with achievability by an explicit data-independent kernel; (iv) finite-sample certificates are one-sided Hoeffding bounds on the same observable admissible-set statistics on held-out unlabeled inputs (Theorems 4.1, 4.6, 4.9), i.e. estimation of a theoretically derived quantity, not a fit renamed as prediction. Exact-rule controls certify zero and the coherent-reading bridge fails on held-out admissibility (ζ̂_U=0.4904), so the paper does not smuggle transfer. No load-bearing self-citation uniqueness theorem, no ansatz via author citation, no renaming of a known empirical pattern as the central result. The derivation is self-contained against its stated assumptions.
Axiom & Free-Parameter Ledger
free parameters (4)
- admissibility threshold τ =
sweep {0.10,...,0.40}; comparison point 0.20
- coverage tolerance ζ
- audit confidence δ and sample size N =
δ=0.10, N=100 (main audit)
- sampled-decoding margin ξ and depth r =
r=3, ξ=0.05 (inconclusive mode)
axioms (5)
- domain assumption Target blindness of the fixed-description channel: Law(ρ(J,X)|f_NL,X) does not depend on the operative target f (Proposition 2.4).
- domain assumption Assumptions (A1)–(A3): measurability of π_LLM, conditional independence of oracle draws, finite nonempty admissible sets a.e.
- standard math Classical two-point / least-favorable-prior minimax identities for finite experiments (Le Cam, Wald, Blackwell–Girshick).
- standard math Hoeffding concentration for bounded i.i.d. indicators on held-out unlabeled inputs.
- ad hoc to paper η-uniform coverage plus ζ-admissibility of a finite coherent-reading pool for transferring V*_τ to V_blind(C) (Theorem 3.9).
invented entities (3)
-
NL-PAC learning problem and model-admissible labeling class F_τ,ζ / core Sel_τ
no independent evidence
-
Admissible-overlap mass D*_τ and exact blind value V*_τ
independent evidence
-
Coherent-reading subclass F_read and η-coverage bridge
no independent evidence
read the original abstract
Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.
Figures
Reference graph
Works this paper leans on
-
[1]
A theory of PAC learnability of partial concept classes
Noga Alon, Steve Hanneke, Ron Holzman, and Shay Moran. A theory of PAC learnability of partial concept classes. In 2021 IEEE 62nd Annual Symposium on Foundations of Computer Science ( FOCS ) , pages 658--671. IEEE, 2022. doi:10.1109/FOCS52979.2021.00070
-
[2]
Angelopoulos and Stephen Bates
Anastasios N. Angelopoulos and Stephen Bates. Conformal prediction: A gentle introduction. Foundations and Trends in Machine Learning, 16 0 (4): 0 494--591, 2023. doi:10.1561/2200000101
-
[3]
Jean-Yves Audibert and Alexandre B. Tsybakov. Fast learning rates for plug-in classifiers. The Annals of Statistics, 35 0 (2): 0 608--633, 2007. doi:10.1214/009053606000001217
-
[4]
Stop measuring calibration when humans disagree
Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fern \'a ndez. Stop measuring calibration when humans disagree. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1892--1915, 2022. doi:10.18653/v1/2022.emnlp-main.124
-
[5]
Concentration inequalities for sampling without replacement
R \'e mi Bardenet and Odalric-Ambrym Maillard. Concentration inequalities for sampling without replacement. Bernoulli, 21 0 (3): 0 1361--1385, 2015. doi:10.3150/14-BEJ605
-
[6]
Minimax regret of finite partial-monitoring games in stochastic environments
G \'a bor Bart \'o k, D \'a vid P \'a l, and Csaba Szepesv \'a ri. Minimax regret of finite partial-monitoring games in stochastic environments. In Sham M. Kakade and Ulrike von Luxburg, editors, Proceedings of the 24th Annual Conference on Learning Theory, volume 19 of Proceedings of Machine Learning Research, pages 133--154, Budapest, Hungary, 2011. PML...
2011
-
[7]
Valid post-selection inference
Richard Berk, Lawrence Brown, Andreas Buja, Kai Zhang, and Linda Zhao. Valid post-selection inference. The Annals of Statistics, 41 0 (2): 0 802--837, 2013. doi:10.1214/12-aos1077
-
[8]
David Blackwell. Comparison of experiments. In Jerzy Neyman, editor, Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, pages 93--102, Berkeley, 1951. University of California Press. doi:10.1525/9780520411586-009
-
[9]
David Blackwell and M. A. Girshick. Theory of Games and Statistical Decisions. Wiley, New York, 1954
1954
-
[10]
Bshouty, Nadav Eiron, and Eyal Kushilevitz
Nader H. Bshouty, Nadav Eiron, and Eyal Kushilevitz. PAC learning with nasty noise. Theoretical Computer Science, 288 0 (2): 0 255--275, 2002. doi:10.1016/S0304-3975(01)00403-0
-
[11]
Learning with bounded instance and label-dependent label noise
Jiacheng Cheng, Tongliang Liu, Kotagiri Ramamohanarao, and Dacheng Tao. Learning with bounded instance and label-dependent label noise. In Hal Daum \'e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 1789--1799. PMLR, 2020. URL https://proceed...
2020
-
[12]
Diagnosing the reliability of LLM -as-a-judge via item response theory, 2026
Junhyuk Choi, Sohhyung Park, Chanhee Cho, Hyeonchu Park, and Bugeun Kim. Diagnosing the reliability of LLM -as-a-judge via item response theory, 2026
2026
-
[13]
Learning from partial labels
Timothee Cour, Ben Sapp, and Ben Taskar. Learning from partial labels. Journal of Machine Learning Research, 12 0 (42): 0 1501--1536, 2011. URL http://jmlr.org/papers/v12/cour11a.html
2011
-
[14]
Underspecification presents challenges for credibility in modern machine learning
Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, et al. Underspecification presents challenges for credibility in modern machine learning. Journal of Machine Learning Research, 23 0 (226): 0 1--61, 2022. URL https://jmlr.org/papers/v23/20-1335.html
2022
-
[15]
A. P. Dawid and A. M. Skene. Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society. Series C (Applied Statistics), 28 0 (1): 0 20--28, 1979. doi:10.2307/2346806
doi:10.2307/2346806 1979
-
[16]
Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the data
Florian Eddie Dorner, Vivian Nastl, and Moritz Hardt. Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the data. In International Conference on Learning Representations, pages 26467--26491, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/hash/4264ee4376776907c0b87ed70b959585-Abstract-Conference.html
2025
-
[17]
Adversarial multiclass classification: A risk minimization perspective
Rizal Fathony, Anqi Liu, Kaiser Asif, and Brian Ziebart. Adversarial multiclass classification: A risk minimization perspective. In D. Lee, M. Sugiyama, U. von Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URL https://proceedings.neurips.cc/paper_files/paper/2016/fi...
2016
-
[18]
Ferguson
Thomas S. Ferguson. Mathematical Statistics: A Decision Theoretic Approach. Academic Press, New York, 1967
1967
-
[19]
Towards provably unbiased llm judges via bias-bounded evaluation, 2026
Benjamin Feuer, Lucas Rosenblatt, and Oussama Elachqar. Towards provably unbiased llm judges via bias-bounded evaluation, 2026
2026
-
[20]
Maxmin expected utility with non-unique prior
Itzhak Gilboa and David Schmeidler. Maxmin expected utility with non-unique prior. Journal of Mathematical Economics, 18 0 (2): 0 141--153, 1989. doi:10.1016/0304-4068(89)90018-9
-
[21]
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Zhouchi Lin, Bowen Zhang, Lionel Ni, Wen Gao, Yuanzhuo Wang, and Jian Guo. A survey on LLM -as-a-judge. The Innovation, 7 0 (6): 0 101253, 2026. doi:10.1016/j.xinn.2025.101253
-
[22]
Validating LLM -as-a-judge systems under rating indeterminacy
Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Steven Wu, and Alexandra Chouldechova. Validating LLM -as-a-judge systems under rating indeterminacy. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors, Advances in Neural Information Processing Systems, volume 38, pages 112282--112350. Curran Associate...
2025
-
[23]
Theory of disagreement-based active learning
Steve Hanneke. Theory of disagreement-based active learning. Foundations and Trends in Machine Learning, 7 0 (2--3): 0 131--309, 2014. doi:10.1561/2200000037
-
[24]
Probability inequalities for sums of bounded random variables
Wassily Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58 0 (301): 0 13--30, 1963. doi:10.1080/01621459.1963.10500830
-
[25]
Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon
Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics, 49 0 (2): 0 1055--1080, 2021. doi:10.1214/20-aos1991
-
[26]
Eyke H \"u llermeier. Learning from imprecise and fuzzy observations: Data disambiguation through generalized loss minimization. International Journal of Approximate Reasoning, 55 0 (7): 0 1519--1534, 2014. doi:10.1016/j.ijar.2013.09.003
-
[27]
Guido W. Imbens and Charles F. Manski. Confidence intervals for partially identified parameters. Econometrica, 72 0 (6): 0 1845--1857, 2004. doi:10.1111/j.1468-0262.2004.00555.x
-
[28]
Learning in the presence of malicious errors
Michael Kearns and Ming Li. Learning in the presence of malicious errors. SIAM Journal on Computing, 22 0 (4): 0 807--837, 1993. doi:10.1137/0222052
doi:10.1137/0222052 1993
-
[29]
Kearns and Umesh V
Michael J. Kearns and Umesh V. Vazirani. An Introduction to Computational Learning Theory. MIT Press, 1994
1994
-
[30]
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jay Shin, yejin cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, S Shin, Ryan, Sungdong Kim, James Thorne, and Minjoon Seo. Prometheus: Inducing fine-grained evaluation capability in language models. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun, editors, International Conference on Learning Representations, vo...
arXiv 2024
-
[31]
Active task disambiguation with LLM s
Katarzyna Kobalczyk, Nicol \'a s Astorga, Tennison Liu, and Mihaela van der Schaar. Active task disambiguation with LLM s. In Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, editors, International Conference on Learning Representations, volume 2025, pages 37823--37847, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/file/5e07476b6bd2497e1fbd11b8...
2025
-
[32]
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In International Conference on Learning Representations, 2023. Spotlight
2023
-
[33]
Language-models-as-a-service: Overview of a new paradigm and its challenges
Emanuele La Malfa, Aleksandar Petrov, Simon Frieder, Christoph Weinhuber, Ryan Burnell, Raza Nazar, Anthony Cohn, Nigel Shadbolt, and Michael Wooldridge. Language-models-as-a-service: Overview of a new paradigm and its challenges. Journal of Artificial Intelligence Research, 80: 0 1497--1523, 2024. doi:10.1613/jair.1.15865
-
[34]
Eliciting human preferences with language models
Belinda Li, Alex Tamkin, Noah Goodman, and Jacob Andreas. Eliciting human preferences with language models. In Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, editors, International Conference on Learning Representations, volume 2025, pages 80984--81013, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/file/c9867d5a22653ce98b02595061e40f12-Paper-...
2025
-
[35]
Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. We're afraid language models aren't modeling ambiguity. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023 a . doi:10.18653/v1/2023.emnlp-main.51
-
[36]
Learnability of the superset label learning problem
Liping Liu and Thomas Dietterich. Learnability of the superset label learning problem. In Eric P. Xing and Tony Jebara, editors, Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, pages 1629--1637, Bejing, China, 2014. PMLR. URL https://proceedings.mlr.press/v32/liug14.html
2014
-
[37]
G-Eval : NLG evaluation using GPT -4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval : NLG evaluation using GPT -4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2511--2522, 2023 b . doi:10.18653/v1/2023.emnlp-main.153
-
[38]
Enno Mammen and Alexandre B. Tsybakov. Smooth discrimination analysis. The Annals of Statistics, 27 0 (6): 0 1808--1829, 1999. doi:10.1214/aos/1017939240
-
[39]
Charles F. Manski. Partial Identification of Probability Distributions. Springer Series in Statistics. Springer, New York, 2003. doi:10.1007/b97478
-
[40]
Empirical Bernstein bounds and sample-variance penalization
Andreas Maurer and Massimiliano Pontil. Empirical Bernstein bounds and sample-variance penalization. In Proceedings of the 22nd Annual Conference on Learning Theory (COLT), 2009. URL https://www.cs.mcgill.ca/ colt2009/papers/012.pdf
2009
-
[41]
Minimax risk classifiers with 0--1 loss
Santiago Mazuelas, Mauricio Romero, and Peter Gr \"u nwald. Minimax risk classifiers with 0--1 loss. Journal of Machine Learning Research, 24 0 (208): 0 1--48, 2023. URL https://www.jmlr.org/papers/volume24/22-0339/22-0339.pdf
2023
-
[42]
Tom M. Mitchell. Generalization as search. Artificial Intelligence, 18 0 (2): 0 203--226, 1982. doi:10.1016/0004-3702(82)90040-6
-
[43]
Ilya Molchanov and Francesca Molinari. Random Sets in Econometrics. Cambridge University Press, 2018. doi:10.1017/9781316392973
-
[44]
Learning with noisy labels
Nagarajan Natarajan, Inderjit S Dhillon, Pradeep K Ravikumar, and Ambuj Tewari. Learning with noisy labels. In C. J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger, editors, Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013. URL https://proceedings.neurips.cc/paper_files/paper/2013/file/3871bd6401...
2013
-
[45]
Yixin Nie, Xiang Zhou, and Mohit Bansal. What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131--9143, 2020. doi:10.18653/v1/2020.emnlp-main.734
-
[46]
Norman, Michael U
Justin D. Norman, Michael U. Rivera, and D. Alex Hughes. Reliability without validity: A systematic, large-scale evaluation of LLM -as-a-judge models across agreement, consistency, and bias, 2026
2026
-
[47]
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt. The effects of reward misspecification: Mapping and mitigating misaligned models. In International Conference on Learning Representations (ICLR), 2022
2022
-
[48]
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7: 0 677--694, 2019. doi:10.1162/tacl_a_00293
-
[49]
The ``problem'' of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. The ``problem'' of human label variation: On ground truth in data, modeling and evaluation. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022. doi:10.18653/v1/2022.emnlp-main.731
-
[50]
Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher R \'e
Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher R \'e . Snorkel: Rapid training data creation with weak supervision. Proceedings of the VLDB Endowment, 11 0 (3): 0 269--282, 2017. doi:10.14778/3157794.3157797
-
[51]
Truth-tracking with non-expert information sources
Joseph Singleton and Richard Booth. Truth-tracking with non-expert information sources. Journal of Artificial Intelligence Research, 81: 0 619--641, 2024. doi:10.1613/jair.1.15273
-
[52]
Maurice Sion. On general minimax theorems. Pacific Journal of Mathematics, 8 0 (1): 0 171--176, 1958. doi:10.2140/pjm.1958.8.171
-
[53]
Minimax regret treatment choice with finite samples
J \"o rg Stoye. Minimax regret treatment choice with finite samples. Journal of Econometrics, 151 0 (1): 0 70--81, 2009 a . doi:10.1016/j.jeconom.2009.02.013
-
[54]
More on confidence intervals for partially identified parameters
J \"o rg Stoye. More on confidence intervals for partially identified parameters. Econometrica, 77 0 (4): 0 1299--1315, 2009 b . doi:10.3982/ECTA7347
-
[55]
Alexandre B. Tsybakov. Introduction to Nonparametric Estimation. Springer Series in Statistics. Springer New York, New York, NY, 1 edition, 2009. ISBN 978-0-387-79052-7. doi:10.1007/b13794. URL https://link.springer.com/book/10.1007/b13794
doi:10.1007/b13794 2009
-
[56]
Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio
Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. Learning from disagreement: A survey. Journal of Artificial Intelligence Research, 72: 0 1385--1470, 2021. doi:10.1613/jair.1.12752
-
[57]
Leslie G. Valiant. A theory of the learnable. Communications of the ACM , 27 0 (11): 0 1134--1142, 1984. doi:10.1145/1968.1972
-
[58]
A new learning paradigm: Learning using privileged information
Vladimir Vapnik and Akshay Vashist. A new learning paradigm: Learning using privileged information. Neural Networks, 22 0 (5--6): 0 544--557, 2009. doi:10.1016/j.neunet.2009.06.042
-
[59]
Algorithmic Learning in a Random World
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer New York, New York, NY, 1 edition, 2005. ISBN 978-0-387-25061-8. doi:10.1007/b106715. URL https://link.springer.com/book/10.1007/b106715
doi:10.1007/b106715 2005
-
[60]
Statistical Decision Functions
Abraham Wald. Statistical Decision Functions. Wiley, New York, 1950
1950
-
[61]
Estimating means of bounded random variables by betting
Ian Waudby-Smith and Aaditya Ramdas. Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (1): 0 1--27, 2024. doi:10.1093/jrsssb/qkad009
-
[62]
Beyond consensus: Perspectivist modeling and evaluation of annotator disagreement in nlp, 2026
Yinuo Xu and David Jurgens. Beyond consensus: Perspectivist modeling and evaluation of annotator disagreement in nlp, 2026
2026
-
[63]
What prompts don't say: Understanding and managing underspecification in LLM prompts
Chenyang Yang, Yike Shi, Qianou Ma, Michael Xieyang Liu, Christian K \"a stner, and Tongshuang Wu. What prompts don't say: Understanding and managing underspecification in LLM prompts. In Findings of the Association for Computational Linguistics: ACL 2026, pages 9072--9101. Association for Computational Linguistics, 2026. doi:10.18653/v1/2026.findings-acl.441
-
[64]
Information-theoretic distinctions between deception and confusion
Robin Young. Information-theoretic distinctions between deception and confusion. In Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, and Dhirendra Pratap Singh, editors, Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conferen...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.