Pith. sign in

REVIEW 3 major objections 5 minor 31 references

What Makes Treatment Effects Identifiable? Characterizations and Estimators Beyond Unconfoundedness

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper characterizes when the average treatment effect is identifiable from censored observational data: exactly when any two candidate explanations with the same observed distribution also agree on the expected outcome.

desk verdict A genuinely new characterization of when ATE is identifiable beyond unconfoundedness, but the main theorem is stated under a pointwise-density convention that does not match standard measure-theoretic identifiability. read the letter →

arxiv 2506.04194 v2 pith:DKUOUOVM submitted 2025-06-04 math.ST cs.LGecon.EMstat.MEstat.MLstat.TH

classification math.STcs.LGecon.EMstat.MEstat.MLstat.TH MSC 62D2062G0568Q32
keywords averagetreatmenteffectidentifiabilityunconfoundednessoverlapgeneralizedpropensityscoreregressiondiscontinuitystatisticallearningtheorycensoreddata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks when the average treatment effect (ATE) can be recovered from censored observational data, where each unit is seen only with the outcome under the treatment it actually received. It introduces a learning-theoretic setup: the analyst fixes a class of generalized propensity scores and a class of covariate-outcome distributions, and assumes the true study is realizable in these classes. The central result is an if-and-only-if characterization: ATE is identifiable exactly when any two compatible candidate explanations that produce the same censored distribution and the same covariate marginal must also have the same expected outcome. The same condition characterizes identification of the average treatment effect on the treated. The paper then shows the condition covers classical unconfoundedness with overlap, overlap without unconfoundedness (including sensitivity-analysis models), and unconfoundedness without overlap (including regression discontinuity designs), and it gives finite-sample estimators in these settings. This matters because unconfoundedness and overlap are routinely violated in practice, and previously there was no single condition saying when point identification is possible.

What carries the argument

The load-bearing object is Condition 1, a separation condition on the pair of concept classes $(\mathcal{P},\mathcal{D})$, where $\mathcal{P}$ contains generalized propensity scores $p_t(x,y)=\Pr[T=t\mid X=x,Y(t)=y]$ and $\mathcal{D}$ contains joint distributions of covariates and potential outcomes. The censored distribution fixes the covariate marginal and the products $p_t(x,y)P(x,y)$, so Condition 1 demands that any two compatible candidate tuples agreeing on these observables agree on the expected outcome. For finite-sample results, the machinery also uses scale-sensitive complexity (fat-shattering dimension) to control estimation of propensity scores, and extrapolation results from truncated statistics to handle settings where overlap fails.

What would settle it

Take any pair of compatible tuples $(p,P)$, $(q,Q)$ with the same covariate marginal and $p(x,y)P(x,y)=q(x,y)Q(x,y)$ for all $(x,y)$ but with $E_P[y]\neq E_Q[y]$. Following the paper's necessity construction, build the two studies $D^{(1)}$ and $D^{(2)}$; their censored distributions coincide while their ATEs differ, which is the explicit counterexample the theorem predicts. Checking that the constructed censored densities are identical and the ATEs differ settles the necessity direction for that class.

Watch

Extended reading notes

Core claim

The central claim is Theorem 1.1: for any observational study realizable with respect to concept classes $(\mathcal{P},\mathcal{D})$, the average treatment effect $\tau_D$ is identifiable from the censored distribution $C_D$ if and only if Condition 1 holds. Condition 1 says that for any two compatible tuples $(p,P)$, $(q,Q)$ in $\mathcal{P}\times\mathcal{D}$, either the expected outcomes $E_P[y]$ and $E_Q[y]$ coincide, or the covariate marginals differ, or the censored products $p(x,y)P(x,y)$ and $q(x,y)Q(x,y)$ differ at some point. In words, once the observed censored distribution and the covariate distribution are fixed, the expected outcome is pinned down. The same condition characterizes the average treatment effect on the treated, and the paper proves both directions: sufficiency by an explicit construction that reads $E_P[y]$ off $C_D$ when Condition 1 holds, and necessity by building two realizable studies with identical censored distributions but different ATEs whenever Condition 1 fails. The framework then yields new identification results for regression discontinuity designs, for outcome families such as Gaussian, Pareto, and Laplace under overlap without unconfoundedness, and for polynomial log-density and polynomial expectation models under unconfoundedness without overlap.

Load-bearing premise

The characterization assumes the identifier can use exact density functions, so two studies are distinguishable if their densities differ at even a single point, and it assumes the analyst knows the true concept classes and that the study is realizable in them.

Editorial extensions

If this is right

  • In regression discontinuity designs, ATE is identifiable without linearity of the outcome regressions; polynomial expected outcomes suffice, extending sharp RD analysis from local cutoff effects to the global average treatment effect.
  • In overlap-without-unconfoundedness settings covered by sensitivity-analysis models, specifying the outcome family (for example Gaussian, Pareto, or Laplace) turns partial identification intervals into point identification.
  • The same Condition 1 characterizes the average treatment effect on the treated, so identification guarantees for ATE carry over to ATT whenever $\Pr[T=1]>0$.
  • Quantitative refinements of Condition 1 (Conditions 4 and 5) yield finite-sample estimators whose sample complexity is expressed through the fat-shattering dimension of the propensity class and the covering number of the outcome class.
  • A variant of the condition also characterizes identification of the heterogeneous treatment effect function.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not pursue this, but the characterization suggests a practical model-selection rule: choose an outcome family for which censored or truncated observations determine the untruncated mean, and then Condition 1 can be checked without deciding whether unconfoundedness holds.
  • Because Condition 1 is stated pointwise for densities, real implementations must pass to a quantitative analogue; the paper's Conditions 4 and 5 are one such analogue, and minimax lower bounds for these quantitative conditions would be a natural next step.
  • The same template could be applied to fuzzy regression discontinuity designs or non-compliance settings by instantiating appropriate generalized propensity score classes; the paper lists fuzzy RD as future work, and Condition 1 gives a ready criterion for when point identification holds there.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper develops a learning-theoretic framework for identifying and estimating the average treatment effect (ATE) and the average treatment effect on the treated (ATT) from the censored distribution C_D of (X,T,Y(T)). It introduces concept classes P for generalized propensity scores and D for covariate-outcome distributions, and proposes Condition 1: any two compatible tuples (p,P),(q,Q) that yield the same censored product and the same covariate marginal must have the same outcome expectation. Theorems 1.1 and 1.2 characterize ATE/ATT identifiability by Condition 1. The paper then specializes Condition 1 to three scenarios: unconfoundedness with overlap, overlap without unconfoundedness, and unconfoundedness without overlap, and proves identification theorems (Theorems 4.1, 4.3, 4.5) and finite-sample estimation guarantees (Theorems 5.1--5.3) under robust forms of these conditions. Applications include regression discontinuity designs and extreme-propensity-score settings, with connections to probabilistic concept learning and truncated statistics.

Significance. If the main characterization holds, it is a notable conceptual contribution: it gives a single separation condition that is both necessary and sufficient for ATE identification beyond the classical unconfoundedness-plus-overlap regime, and it yields the first sample-complexity guarantees in several settings where point identification was previously unknown. The proofs of the identification claims (Section 3 and Appendix B.1) are self-contained and constructive, with explicit counterexamples in the necessity direction. The link between propensity-score estimation and probabilistic concept learning (Appendix F) is a useful observation that may be of independent interest, and the extrapolation-based treatment of regression discontinuity designs via truncated statistics is an original and productive connection. However, the central theorems are stated for identification from the censored 'distribution' while the proofs rely on pointwise evaluation of densities; this gap in the presentation is load-bearing for the main claim.

major comments (3)
  1. [§1.2, Theorems 1.1--1.2; §3.1, Claim 3.1] The characterization is stated for identifiability from the censored distribution C_D, but Condition 1 and the sufficiency proof (Equation (3) in Claim 3.1) evaluate densities pointwise. Under the standard measure-theoretic meaning of a distribution, the sufficiency direction is false. For a concrete counterexample, let X be uniform on [0,1], P(x,y)=1+α(y-1/2), Q(x,y)=1+β(y-1/2) with α≠β, set p≡1/2, and set q=1/2·P/Q everywhere except at a single point (x0,y0) where q=1/4, with α,β chosen so that all generalized propensity scores lie in (c,1-c) for some c>0. Then P_X=Q_X, E_P[y]≠E_Q[y], and pP=qQ almost everywhere but not at (x0,y0); hence Condition 1 holds pointwise while the two censored distributions coincide as measures but have different ATEs. The paper acknowledges this convention only in a parenthetical remark after Condition 1, not in the abstract or in the theorem statements. The main theorems must be restated to either explicitly adopt the nonstandard convention that the identifying map may use pointwise density values, or, preferably, reformulated in an almost-everywhere version of Condition 1 with matching proofs.
  2. [§5.2, Theorem 5.2 and Condition 4] The estimation theorem for overlap without unconfoundedness is conditional on Condition 4, which requires the existence of a mass function M(·) quantifying how strongly any two distributions with different outcome means separate in density ratio. The paper states in the introduction that Gaussian, Pareto, and Laplace outcome distributions satisfy the identification condition, but it never verifies Condition 4 for any concrete family, nor does it provide any example of a mass function M(·) for the Gaussian case. Since the theorem's sample complexity is stated in terms of M(·), the abstract's claim that ATE 'can be estimated from finite samples' in the scenarios of Scenario II is not instantiated unless such a verification is supplied. The authors should either prove Condition 4 for the Gaussian family (and other claimed families) or explicitly state which families admit a known mass function.
  3. [§1.2 and Definition 5] The definition of Compatibility (Definition 5) writes 'setting (p0,p1,DX,Y(0),DX,Y(1)) = (p,bp,P,bP) (or equivalently the re-ordered assignment ...)' which is ambiguous about which component corresponds to treatment versus control. The subsequent proofs (e.g., Claim 3.2) rely on a specific ordering, and the reader must reconstruct the intended symmetry. A precise statement of compatibility as an unordered pair, or an explicit statement of the two possible orderings, would remove this ambiguity and is needed to fully assess the necessity construction.
minor comments (5)
  1. [§1.2] The abstract and the opening of Section 1.2 state the characterization without any caveat about pointwise density evaluation; the remark after Condition 1 ('we allow the identification algorithms to be a function of the whole density') should be moved into the theorem statements or the abstract so that the domain of the identifying map is unambiguous.
  2. [§5.2, Theorem 5.2] The display for the sample complexity of Theorem 5.2 is typeset as O(1/(ηM(ε/2)^2) · (fat(...)·log(1/(ηcσM(ε/2))) + log(N_{ηcM(ε/2)}(D)/δ))). As written, 'log N/δ' is ambiguous; it should be log(N/δ) with N the covering number, matching the proof sketch in §3.3.
  3. [§5.3, Corollary 5.6] The sample complexity stated in Corollary 5.6, n = O(C^2/(σcε)^2 · (...)), is not the same form as the bound in Theorem 5.3, which gives n = O(C^2/(c^2ε)^4 · (...)). The corollary should either derive its bound from Theorem 5.3 or justify the different exponent; otherwise, it may mislead readers about the rate for regression discontinuity designs.
  4. [Algorithm 1] The pseudo-code of Algorithm 1 uses tolerance 'M(O(ε))' in the description, while the theorem and its proof use M(ε/2). This is a minor notational inconsistency that should be harmonized.
  5. [Appendix B.1.2] The proof of Theorem 4.3 begins with a spurious 'Y' character ('Proof of Theorem 4.3. Y We first prove...'), which should be removed.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; the ATE characterization is a direct equivalence, not a self-fulfilling prediction, and the authors' prior-work citations are external instantiations rather than load-bearing circular support.

full rationale

No significant circularity. Theorem 1.1 is a formal if-and-only-if statement: Condition 1 is shown sufficient in Claim 3.1 by constructing the censored-arm product p0·D_X,Y(0) from C_D and reading off E[Y(0)] from any compatible tuple consistent with that product and the covariate marginal; it is shown necessary in Claim 3.2 by building two valid studies D(1), D(2) with C_D(1)=C_D(2) but different E[Y(0)] whenever Condition 1 fails. This is a direct equivalence, not a fitted-input-called-prediction step, and no parameter is estimated from one part of the data and then 'predicted' on a dependent part. The self-citations (DKTZ21, KTZ19, LMZ24, KMZ24) appear in Lemma 5.4 and Remark 5.5 as external extrapolation and truncated-statistics results used to instantiate Conditions 3 and 5; they are real published theorems, and they do not define Condition 1 or rule out alternative characterizations, so they are not load-bearing in a circular sense. The paper also explicitly flags its pointwise-density convention: 'we allow the identification algorithms to be a function of the whole density. If one does not allow this, then one needs to consider the “almost everywhere” analogue of Condition 1.' That is an acknowledged limitation of the interpretation of C_D and would be a correctness concern under the standard measure-theoretic reading, but it is not a circular derivation. The central claim retains independent mathematical content through the Scenario II and III conditions and their verification for Gaussian, Pareto, Laplace, polynomial, and regression-discontinuity families; the low score reflects only the definitionally close formulation of Condition 1 and the presence of non-load-bearing self-citations.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No free parameters fitted to data appear; the constants (c, sigma, B, M, C) are assumptions rather than fitted values. No new physical entities are introduced. The concept-class framework and conditions are mathematical abstractions, not entities requiring independent evidence.

assumptions (6)
  • domain assumption All distributions are continuous with densities; results extend to discrete by replacing densities with probability mass functions.
    Stateed in Section 2 (Preliminaries) and used throughout the proofs.
  • domain assumption Identification mapping is allowed to depend on the full density, distinguishing models that differ on measure-zero sets.
    Section 1.2 discussion after Condition 1; the paper explicitly notes that without this, an 'almost everywhere' analogue is needed.
  • domain assumption The observational study D is realizable with respect to known concept classes (P,D): p0,p1 in P and D_X,Y(0), D_X,Y(1) in D.
    Definition 4 (Realizability) and used in all theorems; the paper admits this is not testable from C_D.
  • domain assumption Potential outcomes model of Neyman-Rubin with binary treatment and unit-level potential outcomes.
    Section 2 and throughout; this is the standard causal framework.
  • domain assumption For estimation results: finite fat-shattering dimension of P, smoothness and covering of D, and quantitative conditions 4 and 5.
    Theorems 5.2 and 5.3 state these as assumptions; they are not derived from data.
  • standard math Use of DKTZ21 extrapolation result for Lemma 5.4.
    Lemma C.1 cites a published COLT paper (Daskalakis et al. 2021) for the key extrapolation bound.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What Makes Treatment Effects Identifiable? Characterizations and Estimators Beyond Unconfoundedness." pith.science (2026). https://pith.science/paper/DKUOUOVM

@misc{pith2026250604194,
  author       = {Pith},
  title        = {Pith review of: What Makes Treatment Effects Identifiable? Characterizations and Estimators Beyond Unconfoundedness},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DKUOUOVM}},
  note         = {Machine review of arXiv:2506.04194}
}
read the original abstract

Most of the widely used estimators of the average treatment effect (ATE) in causal inference rely on the assumptions of unconfoundedness and overlap. Unconfoundedness requires that the observed covariates account for all correlations between the outcome and treatment. Overlap requires the existence of randomness in treatment decisions for all individuals. Nevertheless, many types of studies frequently violate unconfoundedness or overlap, for instance, observational studies with deterministic treatment decisions - popularly known as Regression Discontinuity designs - violate overlap. In this paper, we initiate the study of general conditions that enable the identification of the average treatment effect, extending beyond unconfoundedness and overlap. In particular, following the paradigm of statistical learning theory, we provide an interpretable condition that is sufficient and necessary for the identification of ATE. Moreover, this condition also characterizes the identification of the average treatment effect on the treated (ATT) and can be used to characterize other treatment effects as well. To illustrate the utility of our condition, we present several well-studied scenarios where our condition is satisfied and, hence, we prove that ATE can be identified in regimes that prior works could not capture. For example, under mild assumptions on the data distributions, this holds for the models proposed by Tan (2006) and Rosenbaum (2002), and the Regression Discontinuity design model introduced by Thistlethwaite and Campbell (1960). For each of these scenarios, we also show that, under natural additional assumptions, ATE can be estimated from finite samples. We believe these findings open new avenues for bridging learning-theoretic insights and causal inference methodologies, particularly in observational studies with complex treatment mechanisms.

Figures

Figures reproduced from arXiv: 2506.04194 by the authors.

Figure 1
Figure 1. Illustration of identifiable and non-identifiable instances in Scenario II: The left plot [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Illustration of Identifiable and Non-Identifiable Instances in Scenario III: Left Plot: If [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. This figure illustrates two regression discontinuity designs. In the first design (left [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

31 extracted references · 31 canonical work pages

  1. [1]

    (Sufficiency) If D satisfies Condition 2, then there is a mapping f : ∆(Rd × {0, 1} ×R) → R with f (CD) =τD for each distribution D realizable with respect to (P O(c), D )

  2. [2]

    Condition 2 (Structure of Class D )

    (Necessity) Otherwise, for any map f : ∆(Rd × {0, 1} ×R) → R, there exists a distribution D realizable with respect to (P O(c), D ) such that f (CD) ̸= τD. Condition 2 (Structure of Class D ). Given a constant c > 0, the class of distributions D over (X, Y) is said to satisfy Condition 2 with constant c if for each P, Q ∈ D with E(x,y)∼P[y] ̸= E(x,y)∼Q[y], either

  3. [3]

    the marginals of P and Q over X are different, i.e., PX ̸= QX or

  4. [4]

    Proof of Theorem 4.3

    there exist some x ∈ supp(PX) and y ∈ R, such that P(x,y)/Q(x,y) /∈ (c/(1−c), (1−c)/c) . Proof of Theorem 4.3. Y We first prove sufficiency and then necessity. Sufficiency. Let D satisfy Condition 2. By Theorem 1.1, it suffices to show that(P O(c), D ) satisfies Condition 1. To this end, consider any (p, P) , (q, Q) ∈ P O(c) × D . If either of the followi...

  5. [5]

    (Sufficiency) If D satisfies Condition 3, then there is a mapping f : ∆(Rd × {0, 1} ×R) → R with f (CD) =τD for each distribution D realizable with respect to (P U(c), D )

  6. [6]

    Condition 3 (Structure of Class D )

    (Necessity) Otherwise, for any map f : ∆(Rd × {0, 1} ×R) → R, there exists a distribution D realizable with respect to (P U(c), D ) such that f (CD) ̸= τD. Condition 3 (Structure of Class D ). Given c ∈ (0, 1/2), a class D is said to satisfy Condition 3 if for each P, Q ∈ D with E(x,y)∼P[y] ̸= E(x,y)∼Q[y] either

  7. [7]

    the marginals of P and Q over X are different, i.e., PX ̸= QX, or

  8. [8]

    Proof of Theorem 4.5

    there is no S ⊆ Rd with vol(S) ≥ c such that P(x, y) =Q(x, y) for each (x, y) ∈ S × R. Proof of Theorem 4.5. We first prove sufficiency and then necessity. Sufficiency. Toward a contradiction, suppose D satisfies Condition 3 and, yet, (P U(c), D ) violate Condition 1. Since Condition 1 is violated, there exist (p, P), (q, Q) ∈ (P U, D ) such that EP[y] ̸=...

Show all 31 references
  1. [9]

    P has a finite fat-shattering dimension (Definition 9) fatγ(P ) < ∞ at scale γ = Θ(c2ε/B)

  2. [10]

    Each P ∈ D has support supp(P) ⊆ [−B, B]. 20Note that (bp, P) witnesses the compatibility of (p, P) (as in the proof of Theorem 1.1) since PX = bPX and Pr[T=0 | X=x] = R y p(x,y)P(x,y)dy/R y P(x,y)dy = 1 − R y p(x,y)P(x,y)dy/R y P(x,y)dy = 1 − Pr[T=1 | X=x]. Similarly for (p′,...

  3. [11]

    D satisfies Condition 4 with mass function M(·)

  4. [12]

    Each P ∈ D is σ-smooth with respect to µ.21

  5. [13]

    P has a finite fat-shattering dimension (Definition 9) fatγ(P ) < ∞ at scale γ = Θ(cσM(ε/2))

  6. [14]

    21Distribution P is said to beσ-smooth with respect toµ if its probability density functionp satisfies p(·) ≤ (1/σ) · µ(·)

    D has a finite covering number with respect to total variation distance NO(cM(ε/2))(D ) < ∞. 21Distribution P is said to beσ-smooth with respect toµ if its probability density functionp satisfies p(·) ≤ (1/σ) · µ(·). 49 There is an algorithm that, given n i.i.d. samples from t...

  7. [15]

    D satisfies Condition 5 with constant C > 0

  8. [16]

    Each P ∈ D is σ-smooth with respect to µ.25

  9. [17]

    P has a finite fat-shattering dimension (Definition 9) fatγ(P ) < ∞ at scale γ = Θ(σc3ε/C)

  10. [18]

    There is an algorithm that, given n i.i.d

    D has a finite covering number with respect to TV distance NO(c3ε/C)(D ) < ∞. There is an algorithm that, given n i.i.d. samples from the censored distribution CD for any D realizable with respect to (P , D ) with 2c < PrD[T=1] < 1 − 2c and ε, δ ∈ (0, 1), outputs an estimate b...

  11. [19]

    Estimate a pair (bp, bP) ∈ P × D such that the product bp(x, y)bP(x, y) is close to the product p(x, y)P(x, y) in the following sense Z Z bp(x, y)bP(x, y) − p(x, y)P(x, y) dxdy < ε . (41)

  12. [20]

    (42) The claim is that E(x,y)∼P′ [y] is O(√ε)-close to ED[Y(1)]

    Estimate a distribution P′ ∈ D such that Z Z x∈ bS1 P′(x, y) − bp(x, y)bP(x, y) be(x) dxdy < O √ε c . (42) The claim is that E(x,y)∼P′ [y] is O(√ε)-close to ED[Y(1)]. However, before proving this, we must verify that the preceding steps can be implemented with finite samples. ...

  13. [21]

    (Polynomial Log-Densities) Each element P of this family can have an arbitrary marginal over X and, for each x, the conditional distribution P(y | x) is parameterized by a polynomial f = fP as P(y|x) ∝ e f (x,y)

  14. [22]

    Proof of Lemma 4.6

    (Polynomial Expectations) Each element P of this family can have an arbitrary marginal over X and, for each x, the conditional distribution P(y | x) satisfies the following for some polynomial f = fP E(x,y)∼P[y|X=x] = f (x) . Proof of Lemma 4.6. We proceed in two parts: one fo...

  15. [23]

    They have the same marginal on X, i.e., PX = QX

  16. [24]

    The densities satisfy: P(x, y1) < Q(x, y1) and P(x, y2) > Q(x, y2)

  17. [25]

    We claim that the tuples (p, P) and (p, Q) witness that (P , D all) violate Condition 1

    For each (x′, y′) ̸∈ S, P(x′, y′) =Q(x′, y′) where S := {(x, y1), (x, y2)} . We claim that the tuples (p, P) and (p, Q) witness that (P , D all) violate Condition 1. To see this, fix any (x′, y′) and from the following cases observe that regardless of the choice of(x′, y′), p(...

  18. [26]

    (Equivalence Outcome Distributions) P = Q

  19. [27]

    (Distinction of Covariate Marginals) PX ̸= QX

  20. [28]

    comparable

    (Distinction under Censoring)∃(x, y) ∈ supp(PX) × R, such that, p(x, y)P(x, y) ̸= q(x, y)Q(x, y) The above condition is sufficient to identify the heterogeneous treatment effect. The reason is similar to why Condition 1 is sufficient to identify ATE: consider two observational...

  21. [29]

    Each P ∈ D is σ-smooth with respect to the distribution µ

  22. [30]

    P has a finite fat-shattering dimension fat(ησε )/16(P ) < ∞ at scale ησε/16

  23. [31]

    Consider any D is realizable by (P , D ) and satisfying Pr[T=1] ∈ (η, 1 − η)

    D has a finite covering number with respect to TV distance N = N(D , dTV, ηε/32) < ∞. Consider any D is realizable by (P , D ) and satisfying Pr[T=1] ∈ (η, 1 − η). Then, there exists an al- gorithm that implements an L 1-approximation oracle of accuracy ε and confidence parame...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.