Pith. sign in

REVIEW 3 major objections 3 minor 33 references

The paper replaces the 'true number of causal subgroups'—a quantity that is model-dependent outside latent-class populations—with a resolution profile, the smallest number of groups explaining a prescribed fraction of causal heterogeneity,

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 18:28 UTC pith:NQ7I7ZV6

load-bearing objection A genuinely new estimand-first framework for subgroup-count uncertainty, with honest threshold theory and set-valued reports; the main caveat is that the advertised calibrated inference needs unique optimal codebooks (Assumption 5(i)), which the paper itself admits via Remark 18 but which narrows the 'every population' claim. the 3 major comments →

arxiv 2607.17280 v1 pith:NQ7I7ZV6 submitted 2026-07-19 stat.ME math.STstat.MLstat.TH

The Resolution of Causal Heterogeneity

classification stat.ME math.STstat.MLstat.TH MSC 62G2062F1562H30
keywords causal heterogeneityresolution profilesubgroup analysisquantizationtreatment effect heterogeneityBayesian bootstrapefficient influence functionthreshold nonregularity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that outside populations generated by a genuine latent-class process, 'the number of causal subgroups' is not a property of the population but an artifact of the model used to cluster, so estimating it with confidence intervals is a category error. The proposed fix is to replace the count with a resolution profile: the smallest number of groups whose best K-group summary explains at least a fraction γ of the variance of the causal feature law. This profile is a well-defined nonparametric estimand for every population without latent structure, and it is estimated through one cross-fitted Bayesian-bootstrap posterior over influence-corrected score evaluations, which the paper shows merges with the efficient Gaussian limit uniformly over the loss class. Because the profile is an integer-valued threshold of a continuous path, it is discontinuous in the law at each knot, and no single-valued rule can select the count with locally uniform consistency there; the matched response is a set-valued report obtained by inverting a simultaneous band, which retains its nominal frequentist coverage over exactly those perturbations. If the central claims hold, subgroup-number uncertainty is resolved as threshold nonregularity, and honest statements about how many groups are needed at any chosen resolution become available even when causal features are unobserved and estimated.

Core claim

The central discovery is that the causal heterogeneity R2 curve ρ(K) = 1 − W(K)/W(1), where W(K) is the population quantization risk of the causal feature law, turns 'how many subgroups' into a family of well-defined estimands: K*(γ) = min{K : ρ(K) ≥ γ}. The paper proves a uniform conditional Bernstein–von Mises theorem for a cross-fitted Bayesian-bootstrap posterior of a single structured moment process with influence-function corrections, showing that posterior draws of the moment process converge to the efficient Gaussian limit uniformly over a loss class containing nonsmooth quantization losses. By composition, paths, profiles, fixed-resolution summaries, and subgroup effects all inherit

What carries the argument

The resolution profile K*(γ) = min{K ∈ [K] : 1 − W(K)/W(1) ≥ γ}, with W(K) = inf_{c∈C^K} P_U(g_c) and g_c(u) = min_h ||u − c_h||^2, is the estimand; it is an integer-valued threshold functional of the continuous quantization path. The inferential engine is a single cross-fitted moment process Ψ_f(P) = E_P[f{Hµ_P(X), µ_P(X)}] whose values are corrected by efficient influence functions ϕ_f, then reweighted by Dirichlet weights to define a feature-law posterior. Theorem 2 shows this posterior converges conditionally to the efficient Gaussian process uniformly over the loss class; Theorems 3–7 transfer that limit to paths, profiles, and subgroup effects, and pair an impossibility result at knots

Load-bearing premise

The causal feature law must be spread out so that no hyperplane carries concentrated mass: every hyperplane must have probability at most a constant times t^α over a small neighborhood, a uniform margin condition that excludes atoms and lower-dimensional concentrations; mixed atomic-continuous laws are explicitly outside the main Gaussian theory.

What would settle it

Simulate a population whose causal feature law is a smooth density plus a small point mass placed on the optimal Voronoi boundary for some K. Because P(dist{U,B} ≤ t) is bounded below by the atom mass for every t, Assumption 4 fails; if the simultaneous band for ρ(K) still covers at nominal level then the margin condition is not needed for coverage, and if it undercovers the paper's stated limitation bites.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Subgroup number is no longer a model parameter to be discovered; it is a coordinate along a resolution path, so reports can honestly state 'two groups suffice for 60% of heterogeneity, five for 90%' with simultaneous uncertainty.
  • One corrected moment process supports all resolutions jointly: quantization risks, R2 curve, profile, knots, and subgroup effects, so partition uncertainty propagates automatically.
  • At every knot of the profile, no single-valued count can be selected with locally uniform consistency; the set-valued report is the attainable summary and is not conservative.
  • The procedure is backward compatible: when the feature law is exactly a finite mixture of well-separated tight components, the profile returns the classical K0 on an interval of resolutions.
  • A noise-floor diagnostic separates resolvable heterogeneity from feature-estimation error; a band crossing zero signals insufficient signal, not zero heterogeneity.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural testable extension is to mixed atomic-continuous feature laws: the paper's own limitation note expects consistency to survive, but the Gaussian process-level guarantees to fail; a simulation varying the atom mass and its distance to optimal Voronoi boundaries would map where set-valued coverage degrades.
  • The same threshold-impossibility logic should apply to any integer-valued feature of a continuous estimated path—for example, level-set or dendrogram cuts from estimated densities—so the set-valued reporting principle likely generalizes beyond quantization.
  • Because the profile depends on the analyst's choice of feature map, metric, covariate population, and ceiling, it is a declared summary rather than an intrinsic property; comparisons across studies are meaningful only when those ingredients are fixed, which is a caveat for replication.
  • The paper's level-shift identity suggests a concrete improvement path: recentering the corrected path with an estimator of the feature-estimation floor could sharpen the level W(1) while leaving the efficiency-bound sampling variance unchanged, an interpretive rather than coverage improvement.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper proposes replacing the question 'how many causal subgroups exist?' with a resolution profile K*(γ), a functional of the law of causal features U(X)=Hμ(X), and develops inference through a single cross-fitted Bayesian-bootstrap posterior for an influence-function-corrected moment process. The main results are: a uniform EIF bias bound with a margin-based exponent (Theorem 1), a uniform conditional Bernstein-von Mises theorem (Theorem 2), delta-method transfer to the quantization path, profile and subgroup effects (Corollary 1, Theorems 3 and 7), a Le Cam impossibility result for single-valued knot selection (Theorem 4), and a set-valued band-inversion report with pointwise (Theorem 5) and locally uniform (Theorem 6) coverage. Simulations in four designs and an application to the MineThatData e-mail experiment support the operating characteristics. The construction is careful and explicit about assumptions, but the advertised universality of the inference is narrower than the abstract suggests, because the coverage theorems require a unique optimal codebook (Assumption 5(i)) and a uniform hyperplane margin (Assumption 4).

Significance. If the theorems hold, this is a substantial contribution. The estimand-first reformulation demotes subgroup number from a model index to a threshold coordinate; the single structured moment process propagates uncertainty to paths, profiles, and subgroup effects; and the pairing of an impossibility theorem with an honest set-valued report is a genuine conceptual advance. The uniform bias bound over the non-smooth quantization class goes beyond fixed-K causal k-means, and the noise-floor diagnostic is practical and well tested. The simulations are extensive and support the claims in the covered regimes; the local-uniformity study is particularly commendable. The main threat is scope: the inference theorems exclude natural margin-regular, non-latent laws with non-unique optimal codebooks, so the central claim as stated in the abstract is broader than what is proved. This is fixable by restricting the claims or adding inference for the non-unique case, and therefore warrants major revision rather than rejection.

major comments (3)
  1. [Assumption 5(i); Cor. 1, Thms. 5-6; §2.2] The advertised coverage is not universal. Corollary 1, Theorem 5(iii), and Theorem 6 all require Assumption 5(i) (unique optimal K-point codebook for each K), while the estimand (3) is value-based and, per §2.2, needs no uniqueness. P_U uniform on the unit circle satisfies Assumption 4 (hyperplane margin, α=1/2) and Assumption 5(ii), but for every K≥2 the optimal codebooks are all rotations of a regular K-gon, so 5(i) fails. The manuscript itself concedes (Theorem 3 discussion; Suppl. Rem. 18) that without uniqueness the envelope map is only directionally differentiable and the posterior is generally inconsistent for the limit. Hence the abstract's set-valued 'locally uniform validity ... over exactly the same perturbations' holds only under 5(i): the estimand is universal, but the advertised inference is not, for a broad class of smooth, margin-regular, non-latent laws. Please restrict
  2. [Assumption 4; Suppl. §S3.4.1; §7] The Gaussian path theory requires Assumption 4 (uniform hyperplane margin), which fails for any law with atoms; the paper itself (Suppl. §S3.4.1 and §7) states that mixed atomic-continuous laws are outside the main theory and would need localized treatment. Since a zero-effect atom embedded in a continuous responder distribution is not 'latent structure,' the abstract's claim of a well-defined estimand and calibrated inference 'for every population without latent structure' is overstated on the inference side. The estimand part is fine; I recommend a one-sentence scope statement in the abstract or Section 1 distinguishing the universal estimand from the margin-regular inference class.
  3. [§4 proofs; Supplementary S6] All proofs of the main theorems (1, 2, 4, 5, 6, 7, and 8) are deferred to Supplementary Section S6, which is not included in the submitted review copy. I could not verify the central derivations, in particular the uniform boundary-term control in Theorem 1(ii) over all codebooks and the contiguity transfer in Theorem 6 on which the local-uniformity claim rests. Please make the supplement available to reviewers, or include proof sketches of these two load-bearing uniformity statements in the main text.
minor comments (3)
  1. [§1.2, §3.1] Several citations have duplicated author names: 'Kim et al. Kim et al. (2026)', 'Yiu et al. Yiu et al. (2025)', 'Shapiro Shapiro (1991)', and 'Dümbgen Dümbgen (1993)'. Please fix.
  2. [Keywords] The keyword 'Causal heterogeneityR2' is missing a space; it should read 'Causal heterogeneity R2'.
  3. [Figure 1 and Table 2 (Study 1)] Figure 1(b) correctly distinguishes the plug-in selector's correct-selection frequency from inclusion probabilities, but Table 2 labels the same quantity 'Coverage of K*(γ) by C-hat(γ)'. Consider adding a sentence in the table footnote clarifying that point selector correctness is not a coverage claim.

Circularity Check

0 steps flagged

No significant circularity: the resolution profile and its inference are derived from the causal feature law via an influence-function-corrected Bayesian bootstrap, with the BvM limit proved rather than assumed.

full rationale

The paper's central estimand, K*(γ)=min{K : 1−W(K)/W(1)≥γ}, is a functional of the causal feature law P_U through the population quantization risk W(K)=inf_c E[g_c(U)]. No step defines the estimand in terms of the estimator or fitted values. The posterior is a Bayesian-bootstrap reweighting of cross-fitted, influence-function-corrected scores, and the uniform conditional Bernstein-von Mises theorem (Theorem 2) is a substantive result proved from Assumptions 1-4 and 7, not an input assumption. Downstream results follow by functional delta method (Theorem 3, Corollary 1), a Le Cam two-point argument (Theorem 4), and contiguity-based local uniformity (Theorem 6), all of which are independent derivations. Acknowledged limitations—Assumption 5(i) uniqueness, the margin condition, and mixed atomic-continuous laws—are scope restrictions explicitly stated by the paper, not circular reductions. The noise-floor diagnostic is explicitly labeled as carrying no coverage claim, so it is not a fitted input disguised as prediction. There are no load-bearing self-citations: the cited prior work on causal k-means, posterior corrections, and quantization is external and used for context or contrast. Thus the derivation chain is self-contained and no circularity is present.

Axiom & Free-Parameter Ledger

3 free parameters · 7 axioms · 0 invented entities

The central theory rests on standard causal identification, smoothness/entropy conditions, a strong uniform margin condition, unique-minimizer assumptions, and rate conditions on nuisances. No parameter is fitted to make the theorems work; the application-level σ is a declared resolution constant. The paper is transparent that mixed atomic-continuous laws are outside the main theory.

free parameters (3)
  • Ceiling K = K=8 in simulations, K=6 in the application
    Analyst-chosen maximum number of groups; the profile and all theorems are stated for fixed K. The paper notes growing ceilings are future work (Remark 16).
  • Feature map H
    Analyst-chosen map from the outcome regression to causal features (e.g., H=(−1,1) for the CATE). The metric and feature map are declared ingredients of the estimand.
  • Soft-projection scale σ = 0.008 (declared, not fitted)
    In the application, the Gaussian mixture's common spherical scale is set at the estimated feature-noise floor. The paper states the criterion has no interior scale, but this choice shapes the reported K=2/3 subgroup descriptions and intervals.
axioms (7)
  • domain assumption Assumption 1: Consistency, no unmeasured confounding, positivity (π_a(X) ≥ ε_π > 0).
    Standard causal identification conditions needed to write µ_a(x) = E[Y^a|X=x] and the causal feature law.
  • domain assumption Assumption 2: Boundedness and truncation (|Y| ≤ B_Y, ||µ_a||∞ ≤ B_µ, compact feature support, W(1)>0).
    Technical bounds used throughout the empirical-process arguments.
  • domain assumption Assumption 3: Function class conditions (C² smooth losses, quantization losses, structured scores; entropy and Lipschitz conditions).
    Ensures Donsker-type control and uniform bias bounds over the moment process.
  • ad hoc to paper Assumption 4: Uniform hyperplane margin, P(dist{U(X),B} ≤ t) ≤ C_M t^{α_M} for every hyperplane B.
    Strong regularity imposed to control the quantization boundary term uniformly over codebooks; fails for atoms and mixed laws, as the paper concedes in S3.4.1.
  • domain assumption Assumption 5: Optimal codebooks are unique as sets, with K distinct interior centers, and W(K)<W(K−1).
    Rules out symmetric codebooks; without uniqueness the infimum map is only directionally differentiable and the posterior is inconsistent (Remark 18).
  • domain assumption Assumption 7: Nuisance rates: √n Rem²(η̂^(−b)) = o_P(1) and score increments δ_n = o_P(1).
    Requires product-rate bias control; for quantization it implies r_µ = o_P(n^{−3/8}) at α_M=1, stronger than standard n^{−1/4}.
  • domain assumption W(1)>0 (positive heterogeneity regime).
    The ρ normalization is undefined at W(1)=0; the paper does not test this and provides a two-tier gate instead.

pith-pipeline@v1.3.0-alltime-deepseek · 55924 in / 16180 out tokens · 158873 ms · 2026-08-01T18:28:46.661230+00:00 · methodology

0 comments
read the original abstract

Causal subgroup analyses often report a small number of groups summarizing treatment effect heterogeneity, as if that number were a well-defined estimand. Outside genuinely latent class populations, however, a ``true'' subgroup count is model dependent rather than a population functional. We replace it with a new population estimand, the resolution profile, a functional of the causal feature law giving the fewest groups explaining a prescribed fraction of causal heterogeneity, defined for every population without latent structure. Inference is organized around one cross-fitted Bayesian-bootstrap posterior for a single structured moment process, its scores corrected with influence functions, so that paths, profiles, fixed-resolution summaries, and subgroup effects follow by composition. A uniform conditional Bernstein--von Mises theorem over a loss class containing the nonsmooth quantization losses shows this posterior merges with the efficient Gaussian limit under stated nuisance-rate and margin conditions. Subgroup-number uncertainty is not model selection but threshold nonregularity, the profile being an integer-valued threshold of a continuous path, discontinuous in the law at each knot. At these knots no single-valued selector is locally uniformly consistent over root-$n$ neighborhoods, and the set-valued report obtained by inverting a simultaneous band retains locally uniform validity over exactly the same perturbations. Simulations support the approximations, and an analysis of the MineThatData e-mail experiment illustrates the resolution-indexed report, in which two to three groups summarize the visit response while finer structure falls below a noise-floor diagnostic.

Figures

Figures reproduced from arXiv: 2607.17280 by Fan Li, Yuki Ohnishi.

Figure 1
Figure 1. Figure 1: Study 1. (a) One replication at n = 4000 in DGP-A, which is the DGP-B family at separation s = 0.8. Point path ρˆ(K), simultaneous 95% band, posterior draws, truth, and reading lines, with corrected posterior draws of ρ occasionally exceeding one in finite samples. (b) Along DGP-B, the population answer at γ = 0.9 changes from 3 to 2 through a knot, the knot being approached as the separation s decreases t… view at source ↗
Figure 2
Figure 2. Figure 2: MineThatData e-mail experiment corrected feature-law resolution summaries for the visit [PITH_FULL_IMAGE:figures/full_fig_p022_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: MineThatData e-mail experiment fine K = 3 soft-projection working summary for the visit outcome. (a) Estimated causal feature law with soft membership coloring. Each point is a cross￾fitted campaign-benefit pair, the men’s and women’s e-mail visit benefits against control, colored by the membership-weighted blend of the three component colors, the weights being that unit’s soft mem￾berships r1, r2, r3, so … view at source ↗
Figure 4
Figure 4. Figure 4: Study 2. Empirical coverage of nominal 95% procedures against n, by nuisance regime. The figure displays four of the six regimes of [PITH_FULL_IMAGE:figures/full_fig_p046_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Study 3. Coverage of each procedure’s own population estimand for the subgroup treatment [PITH_FULL_IMAGE:figures/full_fig_p052_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Supplementary energy-scale study (DGP-C, [PITH_FULL_IMAGE:figures/full_fig_p053_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: MineThatData e-mail experiment coarse K = 2 soft-projection working summary for the visit outcome. (a) Estimated causal feature law with soft membership coloring, each point a cross￾fitted campaign-benefit pair colored by the membership-weighted blend of the two component colors, the weights being that unit’s soft memberships r1, r2, so intermediate hues mark soft membership. Blends are almost absent becau… view at source ↗
Figure 8
Figure 8. Figure 8: MineThatData e-mail experiment control-anchored feature-law resolution summaries for the [PITH_FULL_IMAGE:figures/full_fig_p059_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: MineThatData e-mail experiment penalized profile for the visit outcome under the corrected [PITH_FULL_IMAGE:figures/full_fig_p060_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Project STAR below-floor benchmark under the primary Super Learner propensity route. [PITH_FULL_IMAGE:figures/full_fig_p061_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references

  1. [1]

    Athey, S. and G. Imbens (2016). Recursive partitioning for heterogeneous causal effects.Proceedings of the National Academy of Sciences 113(27), 7353–7360. Bartlett, P. L., T. Linder, and G. Lugosi (1998). The minimax distortion redundancy in empirical quantizer design.IEEE Transactions on Information Theory 44(5), 1802–1813. Biau, G., L. Devroye, and G. ...

  2. [2]

    Consider the affine functions(x) =∥x−ch∥2−∥x−ch⋆∥2 =−2x⊤(ch−ch⋆)+∥ch∥2−∥ch⋆∥2

    35 (b) Expanding∥u−c h⋆∥2 =∥¯u−ch⋆∥2 + 2(¯u−ch⋆)⊤(u−¯u) +∥∆∥ 2 givesrem = { ∥u−c h∥2−∥u− ch⋆∥2} +∥∆∥2. Consider the affine functions(x) =∥x−ch∥2−∥x−ch⋆∥2 =−2x⊤(ch−ch⋆)+∥ch∥2−∥ch⋆∥2. Optimality of the labels givess(u)≤0≤s(¯u), while|s(¯u)−s(u)|≤2∥∆∥d; hence−2∥∆∥d≤s(u)≤0 and0≤s(¯u)≤2∥∆∥d, which yields the bound on|rem|. FinallyB={s= 0}with∥∇s∥= 2d, so dist(...

  3. [3]

    Here we prove the supporting quantization envelope lemma and the path corollary

    is stated in the main text, Section 4.2.1, and proved in Supple- mentary Section S6.4. Here we prove the supporting quantization envelope lemma and the path corollary. S6.6.1 The quantization envelope lemma Lemma 4(Differentiability of the quantization functional).FixK∈[ K]and defineι K :ℓ∞(Fqt)→R byι K(ν) = inf c∈CKν(gc). Thenι K is concave and, atν= Ψ 0...

  4. [4]

    The numerator becomes{W(1)−δ}−{W(1)(1−ρ(K))−δ}=W(1)ρ(K), soρ δ(K) =W(1)ρ(K)/{W(1)−δ}

    Ignoring the boundary terms, the corrected path readsWδ(K) =W(K)−δ, and its normalized curveρ δ(K) = 1−W δ(K)/Wδ(1)satisfies ρδ(K) =ρ(K) W(1) W(1)−δ , ρ(K) = 1− W(K) W(1) .(24) Proof.SubstituteW(K) =W(1){1−ρ(K)}intoρ δ(K) = 1−{W(K)−δ}/{W(1)−δ}. The numerator becomes{W(1)−δ}−{W(1)(1−ρ(K))−δ}=W(1)ρ(K), soρ δ(K) =W(1)ρ(K)/{W(1)−δ}. BecauseW(1)/{W(1)−δ}>1, a ...

  5. [5]

    cluster, then estimate within clusters

    (b)(Simultaneous set-valued coverage.)AssumeVar{G 0(gc⋆(K))}>0for allK≤ K. Let(L W,UW )be the simultaneous posterior band forW(·)constructed as in(15), and define ˆC†(τ) = { K≤ K:L W (K) +τK≤min K′≤K { UW (K′) +τK′}} .(18) 2 Then lim infn→∞ P { S0(τ)⊆ ˆC†(τ)for everyτ∈(0,∞) } ≥1−α. The value process forQ(τ)is regular on compact sets avoiding merge scales,...

  6. [6]

    2:Cross-fit ˆη(−b), formˆUi, ˆRi, the corrected evaluationsˆϕf,i, and the point processˆΨ(f) =n −1∑ iˆϕf,i; apply the estimand functionals toˆΨfor all point estimates

    Algorithm 1Feature-law posterior for causal subgroup analysis 1:Input:data{O i}i≤n; feature mapH; ceilingK; mixture familyk(·;θ); foldsB; drawsS. 2:Cross-fit ˆη(−b), formˆUi, ˆRi, the corrected evaluationsˆϕf,i, and the point processˆΨ(f) =n −1∑ iˆϕf,i; apply the estimand functionals toˆΨfor all point estimates. 3:fors= 1,...,Sdo 4:Draww (s)∼n·Dirichlet(1...

  7. [7]

    What is gained for that price is a calibrated posterior for the corrected moment process

    and the choice ofF. What is gained for that price is a calibrated posterior for the corrected moment process. The object updated is that moment process itself, not a probability law on the feature space, and because a weighted corrected functional can take negative values in finite samples the draws need not be feature-space probability measures. The name...

  8. [8]

    Centering byΨ 0(g)−Ψ 0(g′)gives the displayed influence function

    ={g(U)−g ′(U)}+{∇g(U)−∇g ′(U)}⊤HR. Centering byΨ 0(g)−Ψ 0(g′)gives the displayed influence function. The same subtraction is exact in the empirical one-step process and in the weighted process (12), because both are linear in the corrected evaluations. The resolution pathW(·), the causal heterogeneityR 2 curveρ(·), and the resolution profile are built fro...

  9. [9]

    The first term is a standard exchangeable-bootstrap process; the second is the data-dependent estimated-score increment that must be negligible uniformly overF

       oracle multiplier process + 1√n ∑ i (w(s) i −1) ∆i(f)    increment term ,(28) with∆ i(f) = ˆϕf,i−ϕf(Oi;η 0). The first term is a standard exchangeable-bootstrap process; the second is the data-dependent estimated-score increment that must be negligible uniformly overF. The mode of convergence in (14) is conditional weak convergence in probabili...

  10. [10]

    true number of response classes

    This is why the theory treats the soft projections and why the hard-cell displays of Supplementary Section S5.2 carry no nominal coverage. S3.3 Projection Bernstein–von Mises The following theorem, referenced from Section 4.4 of the main text, supplies the partition-uncertainty component of the subgroup-effect limits. 13 Theorem 8(Projection Bernstein–von...

  11. [11]

    Any atom defeats Assumption 4 as stated, since a hyperplane through it carries mass at every distance

    S3.4.1 A caveat for exact finite response classes Between margin-continuous laws with the full√ntheory and purely atomic laws with assumption-free recovery lies the important mixed case: for example, a zero-effect atom embedded in a continuum of responders. Any atom defeats Assumption 4 as stated, since a hyperplane through it carries mass at every distan...

  12. [12]

    Empirical coverage of the nominal95%pointwise interval forρ(3) and of the nominal95%simultaneous band over2≤K≤8, by nuisance regime and sample size, DGP-A withR= 500replications

    The recentering constantc 1 =E 0[IF1{|IF| ≤M}] = 0.0011, the normc 2 =∥s M∥P0,2 = 0.920, and the achieved driftb=E 0[IFs] = 0.920are precomputed by a fixed-seed Monte Carlo of size10 7, and the 19 Table 3: Study 2 numerical coverage. Empirical coverage of the nominal95%pointwise interval forρ(3) and of the nominal95%simultaneous band over2≤K≤8, by nuisanc...

  13. [13]

    The data are the public MineThatData e-mail experiment released by Hillstrom Hillstrom (2008) for an open analytics challenge and available without restriction

    3.42 (0.05) 4.28 (0.02) 0.93 (1.00) 1.63 (1.00) S5 Additional empirical analyses S5.1 MineThatData e-mail experiment specification and preprocessing This section gives the complete empirical specification. The data are the public MineThatData e-mail experiment released by Hillstrom Hillstrom (2008) for an open analytics challenge and available without res...

  14. [14]

    S5.6 A below-floor benchmark

    and excludes it only at the two smallest displayed prices, a reading consistent with the coarse resolvable structure found on theρscale. S5.6 A below-floor benchmark. Project STAR This section complements the main application with a below-floor benchmark. It exercises the rec- ommended behavior of the pipeline on a trial in which the total causal-feature ...

  15. [16]

    By Lemma 3 the curve t↦→Ut(x)is differentiable at everyxwithsup x∥Ut(x)−U(x)∥≤C|t|

    Forf=g c∈F qt the chain-rule step in (30) is justified as follows. By Lemma 3 the curve t↦→Ut(x)is differentiable at everyxwithsup x∥Ut(x)−U(x)∥≤C|t|. The lossg c is Lipschitz onCand differentiable off the union of Voronoi boundaries, a set that isPU-null by the boundary-null hypothesis of part (i), and automatically so under Assumption 4 as noted above, ...

  16. [17]

    By Theorem 1(ii),sup f|P0∆(b)(f)| ≤CRem 2(ˆη(−b)) = oP(n−1/2)under Assumption 7(i)

    =P n∆·(f) = B∑ b=1 nb n [ P0∆(b)(f) bias + (P(b) nb −P 0)∆(b)(f)    empirical process ] ,(32) withP (b) nb the empirical measure of foldb. By Theorem 1(ii),sup f|P0∆(b)(f)| ≤CRem 2(ˆη(−b)) = oP(n−1/2)under Assumption 7(i). For the empirical-process term, condition on the training dataTb of foldb: the fold-bobservations are i.i.d.P 0 and independen...

  17. [18]

    :f∈ F}isP 0-Donsker with bounded envelope, so√n(Pnϕ·(η0)−Ψ 0)⇝G 0 inℓ∞(F); combining with (13) proves the weak convergence. Coordinatewise, the influence function equals the efficient influence function of Theorem 1(i), and asymptotic linearity with the EIF implies regularity and efficiency (van der Vaart, 1998, Section 25.3).□ S6.3 Proof of Theorem 2, pa...

  18. [19]

    We also record a joint equicontinuity fact used below at weight-dependent random indices. Once the theorem is proved, the conditional weak convergence√n(Ψ(s)− ˆΨ) w⇝G 0 to the tight limitG 0, together with the unconditional convergence√n(ˆΨ−Ψ 0)⇝G 0 of part (i), makes the weighted process√n(Ψ(s)−Ψ0)jointly asymptotically equicontinuous with respect to the...

  19. [20]

    :f∈ F}isP 0-Donsker with bounded envelope (Lemma 9(b)), and by Lemma 12 the Dirichlet weights satisfy the conditions of the exchangeable-bootstrap central limit theorem (Præstgaard and Wellner, 1993, Theorem 2.2) and (van der Vaart and Wellner, 1996, Theorem 3.6.13) with limiting multiplier variancec2 =

  20. [21]

    } f∈F w⇝G 0 inℓ∞(F).(34) (Because∑ i(w(s) i −1) = 0, the display is unchanged ifϕf is replaced byϕf−Pfor any other recentering; this is whyˆΨis the exact natural posterior center.) Increment term.Usingw (s) i =nE i/Sn withE 1,...,E n i.i.d. standard exponential,Sn = ∑ jEj, 38 ¯En =S n/n, algebra gives 1√n ∑ i (w(s) i −1)∆i(f) = n Sn · 1√n ∑ i (Ei−1)∆i(f)−...

  21. [22]

    SinceΦ0 is Donsker,G 0 is tight with paths inCς(F)(van der Vaart and Wellner, 1996, Section 1.5), so the tangentiality hypothesis is met, and both conclusions follow

    In every application in this paper the derivative is a continuous linear evaluation map, such asζ↦→ζ(gc⋆(K))orζ↦→−V−1 β ζ(∇βℓβ⋆), so the extension hypothesis holds automatically. SinceΦ0 is Donsker,G 0 is tight with paths inCς(F)(van der Vaart and Wellner, 1996, Section 1.5), so the tangentiality hypothesis is met, and both conclusions follow. Validity of...

  22. [24]

    The mapν↦→(ι1(ν),...,ι K(ν))is then Hadamard differentiable atΨ 0 (coordinatewise differentiability with linear derivatives implies joint), with derivativeζ↦→{ζ(gc⋆(K))}K; Theorem 3 gives both displayed convergences forW. Forρ(·), the map(x 1,...,x K)↦→(1−xK/x1)K is continuously differentiable at the point(W(1),...,W( K))withW(1)>0, so the chain rule for ...

  23. [25]

    van der Vaart, A. and J. A. Wellner (2011). A local maximal inequality under uniform entropy.Electronic Journal of Statistics 5, 192–203. van der Vaart, A. W. (1998).Asymptotic Statistics. Cambridge: Cambridge University Press. van der Vaart, A. W. and J. A. Wellner (1996).Weak Convergence and Empirical Processes: With Applications to Statistics. New York...

  24. [26]

    } −{1−ρ 0(K0)} { ϕgc⋆(1)(o;η 0)−W(1) }] , 44 mean zero withVar0(IF) =σ 2 K0 >0. For a truncation levelM, setsM = IF1{|IF|≤M}−E 0[IF1{|IF|≤ M}]and˜s=s M/∥sM∥P0,2, so that˜sis bounded, mean zero, withE0˜s2 = 1and b:=E 0[IF ˜s] =E0[IF21{|IF|≤M}] ∥sM∥P0,2 −→σ K0 (M→∞); fixMso large thatb≥σ K0−ε(Cauchy–Schwarz givesb≤σ K0 in any case). DefinedPt = (1+t˜s)dP0 f...

  25. [27]

    declare+

    This proves (36). Next, the minimized values. Uniform convergence (36) and uniqueness of the minimizers (Assump- tion 5(i)) give, by the standard argmin argument, that any minimizerct ofΨ Pt(g·)overC K converges toc ⋆(K)ast→0; andc↦→E 0[{ϕgc−Ψ 0(gc)}˜s]is continuous by Lemma 9(c). Sandwiching now yields the envelope (Danskin) expansion: from above,WPt(K)≤...

  26. [28]

    Then there areε>0, a subsequence(n j), and pairs(u j,sj)∈[−B u,Bu]×Swith coverage probability below1−α−εfor everyj

    Mutual absolute continuity implies that measurable covers under the two laws agree up to null sets, and contiguity givesQn(A∗ n)→0, henceQ ∗ n(|Xn|>ε)→0.□ Step 0 (reduction to sequences).Suppose the display of the theorem fails. Then there areε>0, a subsequence(n j), and pairs(u j,sj)∈[−B u,Bu]×Swith coverage probability below1−α−εfor everyj. Since[−B u,B...

  27. [29]

    Le Cam’s third lemma (van der Vaart, 1998, Example 6.7) gives, underQn, {Gn(IFK)}K≤K ⇝N ( {u∗bK(s∗)}K≤K,Σ H ) , whereΣ H is the covariance matrix ofH

    The vector{Gn(IFK)}K≤K and Λn are jointly asymptotically normal underP⊗n 0 by (38) and the multivariate central limit theorem, with asymptotic covarianceCov0{IFK, u∗s∗}=u ∗bK(s∗)between theKth coordinate andΛn. Le Cam’s third lemma (van der Vaart, 1998, Example 6.7) gives, underQn, {Gn(IFK)}K≤K ⇝N ( {u∗bK(s∗)}K≤K,Σ H ) , whereΣ H is the covariance matrix ...

  28. [30]

    Definev(τ) = min K̸=K†(τ){W(K)+τK}−Q(τ)>0forτ∈T

    For Corollary 4, ties occur only at merge scales: at any otherτ,S0(τ)is a singleton by Proposition 1(i). Definev(τ) = min K̸=K†(τ){W(K)+τK}−Q(τ)>0forτ∈T. SinceTis compact and avoids the kinks, it splitsintofinitelymanycompactpiecesoneachofwhichK †(·)isconstantandviscontinuousandpositive; hencev T = infTv >0. On the event{2 max K|ˆW(K)−W(K)|< v T}, of prob...

  29. [31]

    Suppose, to the contrary, thatXn were asymptotically tight inℓ∞(T′)for a setT ′ containing a right neighborhood ofτ (j). The asymptotic-tightness criterion in (van der Vaart and Wellner, 1996, Theo- rem 1.5.7) supplies a semimetricς′ makingT ′ totally bounded and makingXn asymptotically uniformly ς′-equicontinuous. Applying that criterion with(ε,p 1/2)giv...

  30. [32]

    The common limit is−V−1G0(∇βℓβ⋆)∼N(0,V −1ΣβV−1)with Σβ =Var{ϕ ∇βℓβ⋆(O;η 0)}

    On those events the minimizers are exact zeros, so the delta method operates on the exact-zero functional and the tolerance in the definition ofT β matters only with vanishing probability. The common limit is−V−1G0(∇βℓβ⋆)∼N(0,V −1ΣβV−1)with Σβ =Var{ϕ ∇βℓβ⋆(O;η 0)}. Membership surfaces:β↦→rh(·;β)∈ℓ ∞(C)is Hadamard (indeed Fréchet) 53 differentiable by the ...

  31. [33]

    It does not extend uniformly over the root-nneighborhoods of Theorem 4, on whichP(E n)→1fails, which is why the recommended report when ˆC(γ0)is not a singleton is to give the subgroup effects at every supportedKrather than at a single selected count.□ 54 S6.13 Supporting empirical-process lemmas Lemma 9(Entropy, Donsker property, and index continuity).Le...

  32. [34]

    (d) Withδ n = maxb supf∥∆(b)(f)∥P0,2 andδ ′ n = supf{Pn∆·(f) 2}1/2, (δ′ n)2≤δ 2 n +O P(n−1/2), δ ′ n≤δ n +O P(n−1/4)

    :f∈F}have envelopes 4 ¯F 2 and polynomial entropy with the same type of constants, uniformly inηand conditionally on the training folds. (d) Withδ n = maxb supf∥∆(b)(f)∥P0,2 andδ ′ n = supf{Pn∆·(f) 2}1/2, (δ′ n)2≤δ 2 n +O P(n−1/2), δ ′ n≤δ n +O P(n−1/4). Proof.Part (a) follows from Lemma 9(a):D η is contained in the difference of two uniformly VC-type sco...

  33. [1130]

    Castillo, I. and J. Rousseau (2015). A Bernstein–von Mises theorem for smooth functionals in semipara- metric models.The Annals of Statistics 43(6), 2353–2383. Cheng, G. and J. Z. Huang (2010). Bootstrap consistency for general semiparametric M-estimation.The Annals of Statistics 38(5), 2884–2915. Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C....