Pith. sign in

REVIEW 2 major objections 4 minor 42 references

A defect is profitably removable exactly when the detector guarding it can be placed outside the defect—inside, benefit and harm are the same number.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 06:53 UTC pith:ELF6DKER

load-bearing objection The coupling lemma and value formula are genuinely new and the math is mostly clean, but the central iff is false as stated—it drops the L≥L* threshold from Theorem 1, so a confounded detector can earn positive premium at subcritical severities. the 2 major comments →

arxiv 2607.11983 v2 pith:ELF6DKER submitted 2026-07-13 econ.EM cs.AIcs.LGstat.ML

Removable Defects: The Economics and Limits of Deliberate Deficiency

classification econ.EM cs.AIcs.LGstat.ML
keywords removable defectsdeliberate deficiencycoupling lemmadetector placementspecialization premiumselective predictionobservation defectscapacity defects
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper treats a specialist's blind spot as a design variable rather than a cost: a deficiency can be deliberately kept because it concentrates capability where it pays, and removed on demand in the rare fatal case by routing to a compensation channel. It asks when this is economically sound, and answers with an iff: a defect is profitably removable exactly when the detector-relevant distinction survives the restriction (the detector can be placed outside the defect) and an explicit advantage condition holds. The key mechanism is a coupling lemma: inside a deficiency modeled as a coarsening of perception, the detector's rate of capturing specialization gain equals its rate of missing fatal events, so one number prices both and no switch, however perfect, separates benefit from harm. A sympathetic reader should care because this converts a safety-engineering maxim into a decision-theoretic theorem with a computable threshold, and unifies reject-option, selective-prediction, and growth-under-ruin reasoning under a single value formula.

Core claim

The central claim: deliberate blindness is profitably removable exactly when the detector guarding it can be placed outside the defect. A coupling lemma forces that inside a defect (a coarsened perception) the detector's gain-capture rate equals its fatal-miss rate on indistinguishable ensembles, so benefit and harm are one number; a confounded detector therefore earns zero premium even transductively, and under multiplicative dynamics any positive premium destroys long-run growth. A detector outside the defect earns an explicit positive premium under the advantage condition, yielding the iff: removal is profitable iff the detector-relevant distinction survives the restriction and the advant

What carries the argument

The coupling lemma (Lemma 1) is the central mechanism. When a specialist's deficiency is modeled as a coarsening of perception φ, any detector forced to factor through φ has gain-capture rate equal to fatal-miss rate on any safe/fatal pair with equal pushforwards under φ; the quantitative form bounds the wedge by the total-variation distance between pushforwards. This identity carries the necessity direction (zero premium inside the defect) and, with the advantage condition, the sufficiency direction (positive premium outside). A companion value formula—premium equals the support function of the detector class's ROC set at the economic price vector—extends the same mechanism to arbitrary det

Load-bearing premise

The impossibility and necessity results hold when dangerous tasks are mixtures of fixed safe and fatal types that the narrowed view cannot tell apart, with the total dangerous mass capped by a declared budget; they are not established for clustered, cascading, or nonstationary fatal events, which the paper explicitly lists as unmodeled.

What would settle it

In a controlled system with coarsened perception φ and a detector forced to use φ, choose safe and fatal task ensembles with equal pushforwards under φ and measure the detector's gain-capture rate and fatal-miss rate; the coupling lemma predicts they are equal. An experiment showing a within-defect detector with capture rate exceeding miss rate by more than the total-variation distance between pushforwards would falsify the identity. Alternatively, under multiplicative dynamics, hold a confounded detector at fixed miss rate δ and increase the severity L of missed fatal events; the growth form

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A detector that shares the specialist's narrowed perception earns zero specialization premium against adversarial task distributions, even with transductive access to the deployment distribution; under multiplicative dynamics, any positive premium forces long-run log-growth to negative infinity.
  • A detector placed outside the defect earns an explicit positive worst-case premium whenever the specialization gain per period exceeds the insured tail cost (the advantage condition), so keeping a deficiency and escalating the rare fatal case is a computable economic position.
  • Removability is a coefficient, not a yes or no: for any detector class the worst-case premium equals the support function of its ROC set at the economic price vector of captured competence versus leaked fatal exposure, and the removability coefficient is the class's zero-leak capture capacity.
  • Observation defects and capacity defects differ exactly on whether access to the deployment distribution rescues them; the gap decomposes into cross-leak plus a closure deficit, and per-task randomization buys back only the latter.
  • The detector can be learned from stratified samples of declared fatal categories at a one-time training bill roughly linear in loss severity (up to log factors), and the advantage condition survives learning, making the architecture end-to-end feasible.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct consequence for safety auditing: if a monitor's inputs come from the same representation as the actor's, a red-team test measuring the monitor's benefit-capture rate against its fatal-miss rate on indistinguishable task ensembles would either confirm or falsify the placement principle in a real system.
  • The frequency-free form of the advantage condition suggests that insurability of autonomous systems could be certified by auditing per-event miss rates rather than tail-event frequencies; this is testable in insurance and liability schemes.
  • The observation/capacity split hints at a hierarchy of nested defects: if each routing level obeys the same logarithmic detector rent, layered systems might achieve protection that scales exponentially in the number of layers at logarithmic cost; the paper leaves the iterated router open.
  • The linear-in-severity training bill implies that labeled catastrophic examples are a scarce, tradable good; one could test the prediction that a market for such examples would clear at a price growing in proportion to the severity they insure.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper models a specialist that deliberately keeps a competence gap (a 'defect'), guarded by a detector and an escalation/compensation channel. It defines a 'specialization premium' as worst-case payoff above an always-escalate baseline, and calls a defect 'profitably removable' when a within-defect policy earns positive premium. The central theoretical results are: a coupling lemma (Lemma 1) showing that, inside an observation defect, the detector's gain-capture rate equals its fatal-miss rate; a converse (Theorem 1) giving zero premium for confounded detectors once the loss L exceeds an explicit threshold L*; an achievability result (Theorem A) with an explicit positive premium for detectors placed outside the defect; and a unified value formula expressing the premium as the support function of the detector class's ROC set. The paper also gives a taxonomy of observation vs. capacity defects, a decomposition of the joint-realizability deficit into cross-leak and non-closure, randomized-play duality, and end-to-end learning guarantees. It honestly labels Proposition 1 as the Ehrlich-Becker/Townsend insurance margin and lists open problems.

Significance. If the main characterization were true as stated, the paper would be a substantial contribution: it gives a worst-case, decision-theoretic formalization of detector placement, an explicit economic condition for profitable deficiency, a unified pricing formula across detector classes, and a rare-event learning guarantee with a training bill linear in severity. The proofs are transparent and mostly elementary, the paper is unusually explicit about what is proved and what is open, and the core computations (Lemma 1, Theorem 1, Theorem A, Theorem 2') are internally consistent under their stated assumptions. The weakness is the mismatch between the large-L theorems and the unqualified 'iff' stated in the abstract and Corollary A+B; this is a load-bearing issue, not a presentation issue.

major comments (2)
  1. [§5.3, Theorem 1, Corollary A+B; Definition in §5.1] The central iff is false as written. Theorem 1's collapse is proved only for L ≥ L*, but Definition 5.1 of profitable removability has no such threshold, and Corollary A+B/abstract only qualify with 'severity capped or miss rate O(1/L)'. Let φ ≡ const and ν0 = ν1, so φ*ν0 = φ*ν1 and σ̄0 = 0. For any L < L*, the always-act detector is σ(φ)-measurable and has worst-case payoff (1−ε̄)g − ε̄L > −p on F(ν0,ν1), hence earns strictly positive specialization premium despite zero surviving distinction. Since L* > p and (p, L*) is nonempty whenever ε̄ ∈ (0, g/(g+p)), this is not an exotic regime. The statement must either quantify the iff by L ≥ L* or give a separate treatment of the subcritical regime; 'L capped' does not exclude this interval.
  2. [Appendix E, Theorem 1′; §5.3] The claim that removability is 'a coefficient σ0, not a yes/no' is only an L → ∞ statement, but it is presented as the exact content of Theorem 1′ and as part of the paper's central message. For finite L, the theorem's value is a Neyman–Pearson supremum, and the same example as above shows that positive premium can occur with σ0 = 0 when L < L*. The coefficient interpretation in §5.3 should be explicitly qualified as the large-L limit, and the relationship between finite-L profitability and the asymptotic coefficient should be stated precisely.
minor comments (4)
  1. [Corollary A+B / Appendix A] The phrase 'with either L capped or δ = O(1/L)' is ambiguous: 'L capped' needs a definition, and if the cap is below L* the converse does not apply. A sentence defining the cap and its relation to L* would prevent a natural misreading.
  2. [§5.3] When introducing σ0, the text says 'removability is not a yes/no but a number.' This is only true asymptotically; see major comment 2. Please qualify.
  3. [Appendix L / §8] The abbreviation 'NP classification' is used for Neyman–Pearson classification; spell it out, since 'NP' also suggests NP-hardness, which appears later in the same appendix context.
  4. [General] There are several spacing/typographical artifacts in the abstract ('can bekept', 'on demandin'); a final proofreading pass would help.

Circularity Check

0 steps flagged

No significant circularity: the central iff is derived from stated model assumptions, though its L-quantifier is understated (a correctness caveat, not circularity).

full rationale

The derivation chain is self-contained: Lemma 1 (coupling) is proved from Doob–Dynkin factorization and equal pushforwards; Theorem 1/1′ and Theorem 2/2′ compute the premium from the operating-point formulas; Theorem A derives an explicit lower bound and the advantage condition as a threshold; Corollary A+B combines these. No fitted parameter is inserted to force a prediction, and no external result is used as a substitute for proof. The paper explicitly labels Proposition 1 as Ehrlich–Becker/Townsend structure and Theorem A as a re-derivation of Proposition 1, which is honest rather than circular. Self-referential elements (the project's formalization notes, Agent World's shared origin) are declared to be not independent evidence and do not enter the proofs. Citations of classical results (Chow, Goldwasser et al., Wald/Huber–Strassen) are non-self and used as technology, not as the load-bearing source of the claimed impossibility. The main caveat is a scope/correctness issue, not circularity: Theorem 1's converse is proved for L ≥ L*, while the iff in the abstract and Corollary A+B is stated without an explicit L ≥ L* quantifier, so the necessity direction can fail in the subcritical regime L < L*. §9 also honestly lists unmodeled regimes (clusters, cascades, nonstationarity). These are limitations of the stated theorem, not reductions of the conclusion to its inputs.

Axiom & Free-Parameter Ledger

3 free parameters · 9 axioms · 0 invented entities

Theorems rely on standard results (Doob–Dynkin, minimax duality, VC theory) plus domain assumptions about the payoff structure, the adversarial uncertainty class, and the detector's realization. The exogenous quantities are declared bounds, not fitted constants; no hidden free parameter is tuned to produce the results.

free parameters (3)
  • ε̄ (fatal-mass budget)
    Declared upper bound on P(E); appears in the advantage condition, Theorem A, and the critical loss L*. Chosen by the modeler, not estimated from data.
  • c (competence floor)
    Lower bound on P(C) in Theorem A; required for positive premium. Part of the structured class definition, not fitted.
  • detector rate caps δ, α0 and rent c_d
    Hypothesized uniform caps on miss and false-alarm rates and a running cost; the advantage condition and Theorem A depend on them. In a deployment they would need certification (Appendix L), but they are not fitted in the paper.
axioms (9)
  • standard math Doob–Dynkin factorization theorem: any σ(φ)-measurable d can be written d'∘φ
    Used in Lemma 1 (Appendix C) to conclude q0=q1 for observation defects.
  • standard math Chernoff–Stein bound: the minimum sample/rejection cost to achieve miss rate μ scales as O(log(1/μ))
    Used in §4 and Appendix K to argue O(log L) detector rent when δ=O(1/L).
  • standard math Minimax duality (von Neumann / Wald): for finite adversary sets, sup-min equals min-sup over randomized policies
    Used in Theorem M (Appendix I) to collapse the k-direction game to a least-favorable mixture.
  • standard math VC uniform convergence and version-space bounds (Vapnik–Chervonenkis; Blumer et al. 1989)
    Used in Appendix L for Theorems E and E' (empirical rate certification).
  • domain assumption Multiplicative dynamics with i.i.d. multiplicative factors; Kelly growth under logarithmic utility
    Used in Corollary 3 (Appendix J) to translate positive premium into negative expected log-growth at ruin.
  • domain assumption Payoff matrix: act in competence pays +g, act in fatal exposure pays -L, escalate always costs p, with L>p
    Section 2/Appendix A; the whole model is built on this, and L>p is stated as standing assumption.
  • domain assumption Uncertainty class contains mixture families F(ν0,ν1)= {(1-λ)ν0+λν1 : λ∈[0,ε̄]} with equal pushforwards under the coarsening for observation defects
    Appendix A/D; the zero-premium converse is proved only against this structured adversary. Nonstationary or cascading fatal events are explicitly out of scope (§9).
  • domain assumption Declared cover and realizability: fatal side is realizable (∃d*∈D with zero miss on every declared fatal cell)
    Appendix L Theorem E'; the linear training bill depends on it. The paper flags it, and without it the agnostic bill is quadratic (Theorem E).
  • domain assumption Generalist alternative has carrying cost c_gen and gain g_gen<g
    Used in Proposition 1 (§3) to compare specialist-with-detector to the generalist; not needed for the central removability theorems.

pith-pipeline@v1.3.0-alltime-deepseek · 23619 in / 19314 out tokens · 177946 ms · 2026-08-02T06:53:10.946512+00:00 · methodology

0 comments
read the original abstract

A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimized. We treat it as a design variable: a deficiency can be kept because it pays and removed on demand in the rare situation where it would be fatal, by routing to a compensation channel. We give three results. First, an advantage condition under which keeping the deficiency is a computable economic position; structurally it is the Ehrlich-Becker market-vs-self-insurance margin applied to a competence gap, with the detector as a Townsend costly-state-verification technology. Second, a two-sided characterization of removability. A coupling lemma shows that when the deficiency is a coarsening of perception, no switch can separate benefit from harm, yielding a converse (a confounded detector earns zero premium, and any within-defect policy insisting on positive premium is driven, under multiplicative dynamics, to negative long-run growth) and an achievability result (a detector outside the deficiency earns a positive premium). Together, over structured uncertainty classes with severity capped or miss rate O(1/L): a defect is profitably removable iff the detector-relevant distinction survives the restriction and the advantage condition holds; the premium is the support function of the class's ROC set at an economic price vector. Third, observation defects and capacity defects differ exactly on whether access to the deployment distribution rescues them; the gap decomposes as cross-leak plus a closure deficit, and per-task randomization buys back the latter, never the former. The detector can be learned from declared fatal categories at a training bill linear in loss severity (up to a log factor). The results synthesize Chow's reject option, Kelly growth under ruin, and selective prediction.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

42 extracted references · 8 linked inside Pith

  1. [1]

    Bartlett

    Martin Anthony and Peter L. Bartlett. Neural Network Learning: Theoretical Foundations. Cambridge University Press, 1999

  2. [2]

    Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K. Warmuth. Learnability and the V apnik-- C hervonenkis dimension. Journal of the ACM, 36 0 (4): 0 929--965, 1989

  3. [3]

    Optimal gambling systems for favorable games

    Leo Breiman. Optimal gambling systems for favorable games. In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, volume 1, pages 65--78. University of California Press, 1961

  4. [4]

    Learning with the N eyman-- P earson and min-max criteria

    Adam Cannon, James Howse, Don Hush, and Clint Scovel. Learning with the N eyman-- P earson and min-max criteria. Technical Report LA-UR-02-2951, Los Alamos National Laboratory, 2002

  5. [5]

    Selective omniprediction and fair abstention

    S \'i lvia Casacuberta and Varun Kanade. Selective omniprediction and fair abstention. In Advances in Neural Information Processing Systems 38 (NeurIPS 2025), 2025

  6. [6]

    C. K. Chow. An optimum character recognition system using decision functions. IRE Transactions on Electronic Computers, EC-6 0 (4): 0 247--254, 1957

  7. [7]

    C. K. Chow. On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16 0 (1): 0 41--46, 1970

  8. [8]

    Learning with rejection

    Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri. Learning with rejection. In Algorithmic Learning Theory (ALT), volume 9925 of Lecture Notes in Computer Science, pages 67--82. Springer, 2016

  9. [9]

    Cover and Joy A

    Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. Wiley-Interscience, Hoboken, NJ, second edition, 2006

  10. [10]

    Eckhardt and Larry D

    Dave E. Eckhardt and Larry D. Lee. A theoretical basis for the analysis of multiversion software subject to coincident errors. IEEE Transactions on Software Engineering, SE-11 0 (12): 0 1511--1517, 1985

  11. [11]

    A general lower bound on the number of examples needed for learning

    Andrzej Ehrenfeucht, David Haussler, Michael Kearns, and Leslie Valiant. A general lower bound on the number of examples needed for learning. Information and Computation, 82 0 (3): 0 247--261, 1989

  12. [12]

    Isaac Ehrlich and Gary S. Becker. Market insurance, self-insurance, and self-protection. Journal of Political Economy, 80 0 (4): 0 623--648, 1972

  13. [13]

    Is out-of-distribution detection learnable? In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), 2022

    Zhen Fang, Yixuan Li, Jie Lu, Jiahua Dong, Bo Han, and Feng Liu. Is out-of-distribution detection learnable? In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), 2022. arXiv:2210.14707; extended version in Journal of Machine Learning Research 25, 2024

  14. [14]

    Zoubir, and H

    Michael Fau , Abdelhak M. Zoubir, and H. Vincent Poor. Minimax robust detection: Classic results and recent advances. IEEE Transactions on Signal Processing, 69: 0 2252--2283, 2021. arXiv:2105.09836

  15. [15]

    Constructive minimax classification of discrete observations with arbitrary loss function

    Lionel Fillatre. Constructive minimax classification of discrete observations with arbitrary loss function. Signal Processing, 141: 0 322--330, 2017

  16. [16]

    SelectiveNet : A deep neural network with an integrated reject option

    Yonatan Geifman and Ran El-Yaniv. SelectiveNet : A deep neural network with an integrated reject option. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, pages 2151--2159, 2019

  17. [17]

    Beyond perturbations: Learning guarantees with arbitrary adversarial test examples

    Shafi Goldwasser, Adam Tauman Kalai, Yael Tauman Kalai, and Omar Montasser. Beyond perturbations: Learning guarantees with arbitrary adversarial test examples. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 2020. arXiv:2007.05145

  18. [18]

    AI control: Improving safety despite intentional subversion

    Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. AI control: Improving safety despite intentional subversion. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of Proceedings of Machine Learning Research, 2024. arXiv:2312.06784

  19. [19]

    Machine learning with a reject option: A survey

    Kilian Hendrickx, Lorenzo Perini, Dries Van der Plas, Wannes Meert, and Jesse Davis. Machine learning with a reject option: A survey. Machine Learning, 113: 0 3073--3110, 2024. arXiv:2107.11277

  20. [20]

    Peter J. Huber. A robust version of the probability ratio test. Annals of Mathematical Statistics, 36 0 (6): 0 1753--1758, 1965

  21. [21]

    Huber and Volker Strassen

    Peter J. Huber and Volker Strassen. Minimax tests and the N eyman-- P earson lemma for capacities. Annals of Statistics, 1 0 (2): 0 251--263, 1973

  22. [22]

    Reliable agnostic learning

    Adam Tauman Kalai, Varun Kanade, and Yishay Mansour. Reliable agnostic learning. Journal of Computer and System Sciences, 78 0 (5): 0 1481--1495, 2012

  23. [23]

    Kelly, Jr

    John L. Kelly, Jr. A new interpretation of information rate. Bell System Technical Journal, 35 0 (4): 0 917--926, 1956

  24. [24]

    Lehmann and Joseph P

    Erich L. Lehmann and Joseph P. Romano. Testing Statistical Hypotheses. Springer, New York, third edition, 2005

  25. [25]

    Reasoning about the reliability of diverse two-channel systems in which one channel is ``possibly perfect''

    Bev Littlewood and John Rushby. Reasoning about the reliability of diverse two-channel systems in which one channel is ``possibly perfect''. IEEE Transactions on Software Engineering, 38 0 (5): 0 1178--1194, 2012

  26. [26]

    Managoli, K

    Malhar A. Managoli, K. R. Sahasranand, and Vinod M. Prabhakaran. Robust hypothesis testing with abstention, 2025

  27. [27]

    Who should predict? E xact algorithms for learning to defer to humans

    Hussein Mozannar, Hunter Lang, Dennis Wei, Prasanna Sattigeri, Subhro Das, and David Sontag. Who should predict? E xact algorithms for learning to defer to humans. In Proceedings of the 26th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 206 of Proceedings of Machine Learning Research, pages 10520--10545, 2023

  28. [28]

    Jerzy Neyman and Egon S. Pearson. On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society of London, Series A, 231: 0 289--337, 1933

  29. [29]

    Insurance makes wealth grow faster, 2015

    Ole Peters and Alexander Adamou. Insurance makes wealth grow faster, 2015. arXiv:1507.04655

  30. [30]

    Robust classification for imprecise environments

    Foster Provost and Tom Fawcett. Robust classification for imprecise environments. Machine Learning, 42 0 (3): 0 203--231, 2001. arXiv:cs/0009007

  31. [31]

    N eyman-- P earson classification, convexity and stochastic constraints

    Philippe Rigollet and Xin Tong. N eyman-- P earson classification, convexity and stochastic constraints. Journal of Machine Learning Research, 12: 0 2831--2855, 2011

  32. [32]

    A N eyman-- P earson approach to statistical learning

    Clayton Scott and Robert Nowak. A N eyman-- P earson approach to statistical learning. IEEE Transactions on Information Theory, 51 0 (11): 0 3806--3819, 2005

  33. [33]

    Understanding Machine Learning: From Theory to Algorithms

    Shai Shalev-Shwartz and Shai Ben-David. Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, 2014

  34. [34]

    Safe Haven: Investing for Financial Storms

    Mark Spitznagel. Safe Haven: Investing for Financial Storms. Wiley, Hoboken, NJ, 2021

  35. [35]

    Townsend

    Robert M. Townsend. Optimal contracts and competitive markets with costly state verification. Journal of Economic Theory, 21 0 (2): 0 265--293, 1979

  36. [36]

    Know your limits: Uncertainty estimation with ReLU classifiers fails at reliable OOD detection

    Dennis Ulmer and Giovanni Cin \`a . Know your limits: Uncertainty estimation with ReLU classifiers fails at reliable OOD detection. In Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence (UAI), 2021. arXiv:2012.05329

  37. [37]

    Van Mieghem

    Jan A. Van Mieghem. Capacity management, investment, and hedging: Review and recent developments. Manufacturing & Service Operations Management, 5 0 (4): 0 269--302, 2003

  38. [38]

    V. N. Vapnik and A. Ya. Chervonenkis. On the uniform convergence of relative frequencies of events to their probabilities. Theory of Probability and its Applications, 16 0 (2): 0 264--280, 1971

  39. [39]

    Algorithmic Learning in a Random World

    Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer, New York, 2005

  40. [40]

    Statistical decision functions which minimize the maximum risk

    Abraham Wald. Statistical decision functions which minimize the maximum risk. Annals of Mathematics, 46 0 (2): 0 265--280, 1945

  41. [41]

    Statistical Decision Functions

    Abraham Wald. Statistical Decision Functions. Wiley, New York, 1950

  42. [42]

    Williamson

    Oliver E. Williamson. The Economic Institutions of Capitalism: Firms, Markets, Relational Contracting. Free Press, New York, 1985