{"id":"4a4e44e2-72cf-4bd0-8223-097e6d46c7fd","arxiv_id":"2607.17280","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The resolution profile K*(γ) = min{K: 1−W(K)/W(1) ≥ γ} replaces the model-dependent subgroup count with a population functional, and set-valued band-inversion reports are valid exactly where no single-valued count is.","lead":"Instead of asking 'how many causal subgroups exist?', this paper makes the count a well-defined population target called the resolution profile: the fewest groups that explain a given fraction of treatment-effect variation. It then gives calibrated uncertainty for that target using one influence-corrected Bayesian bootstrap, showing that near ambiguity points the honest answer is a set of possible counts, not a single number.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 5(i) (unique optimal codebook) is the load-bearing restriction: the value-based resolution profile is defined without codebook uniqueness, but the main coverage theorems and set-valued report require it, excluding rotationally symmetric continuous laws that satisfy the margin condition.","rationale":"The reader's Assumption 4 concern is real and partly overlapping: atomic or hyperplane-supported laws violate the margin condition and are excluded from the Gaussian theory. However, the more load-bearing gap for the central claim is Assumption 5(i). The paper itself emphasizes that the resolution profile is a value-based estimand needing no uniqueness, and then requires uniqueness for every inferential statement about that same estimand. Moreover, the unique-codebook condition can fail even in smooth, margin-valid, curved lower-dimensional laws such as the uniform circle, which are not covered by the paper's explicit concession about mixed atomic-continuous laws. This is a structural gap between the conceptual target and the inference, not just a regularity condition on nuisance rates. The reader's verdict of CONDITIONAL remains appropriate: the theory is coherent under Assumptions 4 and 5, but the scope is narrower than 'every population without latent structure,' and the paper should either extend the inference to non-unique codebooks or clearly delimit the set-valued report's validity. I therefore keep the verdict unchanged while identifying a different, more central assumption than the reader's weakest assumption.","tokens_in":56315,"tokens_out":15762,"duration_ms":158857,"concrete_test":"Simulate P_U as the uniform law on the unit circle in R^2 (e.g., U=(cosΘ,sinΘ), Θ∼Unif[0,2π)), with known outcome regressions and propensities so Assumption 7 is satisfied, n=4000, ceiling K=3, S=1000 Dirichlet draws, R=1000 replications. This law satisfies Assumption 4 (α_M=1/2) but violates Assumption 5(i) for K=2 and K=3. Compute the nominal 95% posterior intervals for W(2), ρ(2), and the set-valued report \\hat C(0.5). If coverage is near nominal, the uniqueness assumption may be replaceable; if coverage departs substantially from nominal (or the posterior of W(2) is not correctly centered), Assumption 5(i) is a genuine load-bearing restriction of the main theorems, not a mild technical condition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Assumption 5(i) is the load-bearing restriction for the advertised coverage claims. The resolution profile is deliberately value-based: §2.2 states it is invariant to the choice of optimal codebook when C*(K) is not a singleton and is a well-defined target without any uniqueness condition. Yet Corollary 1, Theorems 5 and 6, and the set-valued report \\hat C(γ) require Assumption 5(i): the optimal codebook must be a unique set for every K. If C*(K) is a continuum, the envelope map ν ↦ inf_c ν(g_c) is only directionally differentiable; the Gaussian delta method of Theorem 3 does not apply to W(K) and ρ(K), and Supplementary Remark 18 concedes the posterior is generally inconsistent for the limit. This is not an exotic or latent-class population. For example, P_U uniform on the unit circle satisfies Assumption 4 (hyperplane margin holds with α_M=1/2: tangent lines give P(dist≤t) ≍ √t) and satisfies Assumption 5(ii), but for every K≥2 the optimal codebooks are all rotations of a regular K-gon, so Assumption 5(i) fails. The paper's stated Gaussian-calibrated set-valued inference therefore has no coverage guarantee for a broad class of smooth, margin-regular, non-latent feature laws. The paper only references a remark for what fails, not a replacement theorem. This makes the central claim 'well-defined estimand and calibrated set-valued reports for every population without latent structure' narrower than stated: the estimand is universal, but the advertised inference is not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes replacing the question 'how many causal subgroups exist?' with a resolution profile K*(γ), a functional of the law of causal features U(X)=Hμ(X), and develops inference through a single cross-fitted Bayesian-bootstrap posterior for an influence-function-corrected moment process. The main results are: a uniform EIF bias bound with a margin-based exponent (Theorem 1), a uniform conditional Bernstein-von Mises theorem (Theorem 2), delta-method transfer to the quantization path, profile and subgroup effects (Corollary 1, Theorems 3 and 7), a Le Cam impossibility result for single-valued knot selection (Theorem 4), and a set-valued band-inversion report with pointwise (Theorem 5) and locally uniform (Theorem 6) coverage. Simulations in four designs and an application to the MineThatData e-mail experiment support the operating characteristics. The construction is careful and explicit about assumptions, but the advertised universality of the inference is narrower than the abstract suggests, because the coverage theorems require a unique optimal codebook (Assumption 5(i)) and a uniform hyperplane margin (Assumption 4).","tokens_in":56758,"tokens_out":8518,"duration_ms":80340,"significance":"If the theorems hold, this is a substantial contribution. The estimand-first reformulation demotes subgroup number from a model index to a threshold coordinate; the single structured moment process propagates uncertainty to paths, profiles, and subgroup effects; and the pairing of an impossibility theorem with an honest set-valued report is a genuine conceptual advance. The uniform bias bound over the non-smooth quantization class goes beyond fixed-K causal k-means, and the noise-floor diagnostic is practical and well tested. The simulations are extensive and support the claims in the covered regimes; the local-uniformity study is particularly commendable. The main threat is scope: the inference theorems exclude natural margin-regular, non-latent laws with non-unique optimal codebooks, so the central claim as stated in the abstract is broader than what is proved. This is fixable by restricting the claims or adding inference for the non-unique case, and therefore warrants major revision rather than rejection.","major_comments":[{"comment":"The advertised coverage is not universal. Corollary 1, Theorem 5(iii), and Theorem 6 all require Assumption 5(i) (unique optimal K-point codebook for each K), while the estimand (3) is value-based and, per §2.2, needs no uniqueness. P_U uniform on the unit circle satisfies Assumption 4 (hyperplane margin, α=1/2) and Assumption 5(ii), but for every K≥2 the optimal codebooks are all rotations of a regular K-gon, so 5(i) fails. The manuscript itself concedes (Theorem 3 discussion; Suppl. Rem. 18) that without uniqueness the envelope map is only directionally differentiable and the posterior is generally inconsistent for the limit. Hence the abstract's set-valued 'locally uniform validity ... over exactly the same perturbations' holds only under 5(i): the estimand is universal, but the advertised inference is not, for a broad class of smooth, margin-regular, non-latent laws. Please restrict","section":"Assumption 5(i); Cor. 1, Thms. 5-6; §2.2"},{"comment":"The Gaussian path theory requires Assumption 4 (uniform hyperplane margin), which fails for any law with atoms; the paper itself (Suppl. §S3.4.1 and §7) states that mixed atomic-continuous laws are outside the main theory and would need localized treatment. Since a zero-effect atom embedded in a continuous responder distribution is not 'latent structure,' the abstract's claim of a well-defined estimand and calibrated inference 'for every population without latent structure' is overstated on the inference side. The estimand part is fine; I recommend a one-sentence scope statement in the abstract or Section 1 distinguishing the universal estimand from the margin-regular inference class.","section":"Assumption 4; Suppl. §S3.4.1; §7"},{"comment":"All proofs of the main theorems (1, 2, 4, 5, 6, 7, and 8) are deferred to Supplementary Section S6, which is not included in the submitted review copy. I could not verify the central derivations, in particular the uniform boundary-term control in Theorem 1(ii) over all codebooks and the contiguity transfer in Theorem 6 on which the local-uniformity claim rests. Please make the supplement available to reviewers, or include proof sketches of these two load-bearing uniformity statements in the main text.","section":"§4 proofs; Supplementary S6"}],"minor_comments":[{"comment":"Several citations have duplicated author names: 'Kim et al. Kim et al. (2026)', 'Yiu et al. Yiu et al. (2025)', 'Shapiro Shapiro (1991)', and 'Dümbgen Dümbgen (1993)'. Please fix.","section":"§1.2, §3.1"},{"comment":"The keyword 'Causal heterogeneityR2' is missing a space; it should read 'Causal heterogeneity R2'.","section":"Keywords"},{"comment":"Figure 1(b) correctly distinguishes the plug-in selector's correct-selection frequency from inclusion probabilities, but Table 2 labels the same quantity 'Coverage of K*(γ) by C-hat(γ)'. Consider adding a sentence in the table footnote clarifying that point selector correctness is not a coverage claim.","section":"Figure 1 and Table 2 (Study 1)"}],"recommendation":"major_revision","confidential_remarks":"The paper is ambitious and mostly internally consistent. The central issue is a genuine scope mismatch between the abstract and the theorems: the estimand is universal, but the calibrated inference is restricted to laws satisfying Assumption 5(i) and Assumption 4. This is fixable within the manuscript's scope by narrowing the headline claims or adding a non-unique-codebook treatment (e.g., via directional differentiability and resampling). Given that all proofs are in S6, I strongly recommend that the supplementary proofs be circulated with the next revision. A referee with expertise in empirical quantization and bootstrap theory should specifically check Theorem 1(ii)'s boundary bound and Theorem 6's contiguity transfer, since these carry the uniformity claims. I did not find evidence of plagiarism or duplicate publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. The paper replaces the ill-defined 'true number of subgroups' with a population functional, the resolution profile K*(γ), and builds inference around one cross-fitted Bayesian-bootstrap posterior over influence-corrected moment processes. The conceptual shift—subgroup count as a threshold coordinate rather than a model parameter—is real and well executed. The process-level Bernstein–von Mises theorem, the impossibility result at knots (Theorem 4), and the matched set-valued report that stays locally uniformly valid exactly where single-valued selection fails are all substantive. Simulations are careful and support the main operating characteristics, and the MineThatData application is honest about what is licensed and what is diagnostic. The paper also flags its own limitations in the Discussion and in remarks, which is good. The main soft spot is precisely the one the stress-test note identifies: the value-based resolution profile is defined without codebook uniqueness, but the advertised coverage guarantees—Corollary 1, Theorems 5 and 6, and the set-valued report—require Assumption 5(i). That rules out rotationally symmetric continuous laws like the uniform on the circle, which satisfy the margin condition but have a continuum of optimal codebooks. The paper concedes this in Remark 18 (posterior generally inconsistent for the limit without uniqueness), but that is a remark, not a replacement theorem. So the central claim 'calibrated set-valued reports for every population without latent structure' is narrower than stated: the estimand is universal, the inference is not. This is a genuine gap in scope, not a sign the main machinery is wrong. Other soft spots: proofs are deferred to the supplement (a practical concern for verification), the uniform hyperplane margin excludes mixed atomic-continuous laws (the paper says so), and the nuisance-rate requirements are strong—though the simulations probe rate violations and the noise-floor diagnostic is a thoughtful buffer. I did not find circularity; the construction is self-contained and the posterior limit is derived from efficient influence functions, not assumed. Who is this for? Statisticians working on causal heterogeneity, clustering, or semiparametric inference. It deserves a serious referee. I would send it out and ask referees to pressure-test the uniqueness assumption and demand that proofs be in the main text or a clearly accessible supplement. I would cite the resolution profile concept with the caveat about Assumption 5(i) stated.","headline":"A genuinely new estimand-first framework for subgroup-count uncertainty, with honest threshold theory and set-valued reports; the main caveat is that the advertised calibrated inference needs unique optimal codebooks (Assumption 5(i)), which the paper itself admits via Remark 18 but which narrows the 'every population' claim.","tokens_in":775,"tokens_out":1627,"would_cite":true,"duration_ms":25148,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G20","62F15","62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper replaces the 'true number of causal subgroups'—a quantity that is model-dependent outside latent-class populations—with a resolution profile, the smallest number of groups explaining a prescribed fraction of causal heterogeneity,","keywords":["causal heterogeneity","resolution profile","subgroup analysis","quantization","treatment effect heterogeneity","Bayesian bootstrap","efficient influence function","threshold nonregularity"],"falsifier":"Simulate a population whose causal feature law is a smooth density plus a small point mass placed on the optimal Voronoi boundary for some K. Because P(dist{U,B} ≤ t) is bounded below by the atom mass for every t, Assumption 4 fails; if the simultaneous band for ρ(K) still covers at nominal level then the margin condition is not needed for coverage, and if it undercovers the paper's stated limitation bites.","tokens_in":56232,"feed_emoji":"📊","tokens_out":6689,"duration_ms":69209,"temperature":0.7,"pith_summary":"The paper argues that outside populations generated by a genuine latent-class process, 'the number of causal subgroups' is not a property of the population but an artifact of the model used to cluster, so estimating it with confidence intervals is a category error. The proposed fix is to replace the count with a resolution profile: the smallest number of groups whose best K-group summary explains at least a fraction γ of the variance of the causal feature law. This profile is a well-defined nonparametric estimand for every population without latent structure, and it is estimated through one cross-fitted Bayesian-bootstrap posterior over influence-corrected score evaluations, which the paper shows merges with the efficient Gaussian limit uniformly over the loss class. Because the profile is an integer-valued threshold of a continuous path, it is discontinuous in the law at each knot, and no single-valued rule can select the count with locally uniform consistency there; the matched response is a set-valued report obtained by inverting a simultaneous band, which retains its nominal frequentist coverage over exactly those perturbations. If the central claims hold, subgroup-number uncertainty is resolved as threshold nonregularity, and honest statements about how many groups are needed at any chosen resolution become available even when causal features are unobserved and estimated.","feed_headline":"Causal subgroup count is a resolution choice, not a truth","feed_subtitle":"One posterior over treatment-effect features yields honest set-valued counts at every heterogeneity level.","key_machinery":"The resolution profile K*(γ) = min{K ∈ [K] : 1 − W(K)/W(1) ≥ γ}, with W(K) = inf_{c∈C^K} P_U(g_c) and g_c(u) = min_h ||u − c_h||^2, is the estimand; it is an integer-valued threshold functional of the continuous quantization path. The inferential engine is a single cross-fitted moment process Ψ_f(P) = E_P[f{Hµ_P(X), µ_P(X)}] whose values are corrected by efficient influence functions ϕ_f, then reweighted by Dirichlet weights to define a feature-law posterior. Theorem 2 shows this posterior converges conditionally to the efficient Gaussian process uniformly over the loss class; Theorems 3–7 transfer that limit to paths, profiles, and subgroup effects, and pair an impossibility result at knots","core_discovery":"The central discovery is that the causal heterogeneity R2 curve ρ(K) = 1 − W(K)/W(1), where W(K) is the population quantization risk of the causal feature law, turns 'how many subgroups' into a family of well-defined estimands: K*(γ) = min{K : ρ(K) ≥ γ}. The paper proves a uniform conditional Bernstein–von Mises theorem for a cross-fitted Bayesian-bootstrap posterior of a single structured moment process with influence-function corrections, showing that posterior draws of the moment process converge to the efficient Gaussian limit uniformly over a loss class containing nonsmooth quantization losses. By composition, paths, profiles, fixed-resolution summaries, and subgroup effects all inherit","pith_inferences":["A natural testable extension is to mixed atomic-continuous feature laws: the paper's own limitation note expects consistency to survive, but the Gaussian process-level guarantees to fail; a simulation varying the atom mass and its distance to optimal Voronoi boundaries would map where set-valued coverage degrades.","The same threshold-impossibility logic should apply to any integer-valued feature of a continuous estimated path—for example, level-set or dendrogram cuts from estimated densities—so the set-valued reporting principle likely generalizes beyond quantization.","Because the profile depends on the analyst's choice of feature map, metric, covariate population, and ceiling, it is a declared summary rather than an intrinsic property; comparisons across studies are meaningful only when those ingredients are fixed, which is a caveat for replication.","The paper's level-shift identity suggests a concrete improvement path: recentering the corrected path with an estimator of the feature-estimation floor could sharpen the level W(1) while leaving the efficiency-bound sampling variance unchanged, an interpretive rather than coverage improvement."],"forward_implications":["Subgroup number is no longer a model parameter to be discovered; it is a coordinate along a resolution path, so reports can honestly state 'two groups suffice for 60% of heterogeneity, five for 90%' with simultaneous uncertainty.","One corrected moment process supports all resolutions jointly: quantization risks, R2 curve, profile, knots, and subgroup effects, so partition uncertainty propagates automatically.","At every knot of the profile, no single-valued count can be selected with locally uniform consistency; the set-valued report is the attainable summary and is not conservative.","The procedure is backward compatible: when the feature law is exactly a finite mixture of well-separated tight components, the profile returns the classical K0 on an interval of resolutions.","A noise-floor diagnostic separates resolvable heterogeneity from feature-estimation error; a band crossing zero signals insufficient signal, not zero heterogeneity."],"fun_headline_variants":["Subgroup count is a resolution dial, not a truth","How many subgroups? It's a resolution choice, not a truth","The true number of subgroups doesn't exist—only a resolution profile","Set-valued subgroup counts: honest answers to an ill-posed question","Causal subgroup count: not a number, but a curve"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The causal feature law must be spread out so that no hyperplane carries concentrated mass: every hyperplane must have probability at most a constant times t^α over a small neighborhood, a uniform margin condition that excludes atoms and lower-dimensional concentrations; mixed atomic-continuous laws are explicitly outside the main Gaussian theory.","fun_headline_variants_meta":{"raw":{"variants":["Subgroup count is a resolution dial, not a truth","How many subgroups? It's a resolution choice, not a truth","The true number of subgroups doesn't exist—only a resolution profile","Set-valued subgroup counts: honest answers to an ill-posed question","Causal subgroup count: not a number, but a curve"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000797,"raw_usage":{"total_tokens":3360,"prompt_tokens":777,"completion_tokens":2583,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":2495}},"tokens_in":521,"tokens_out":2583,"duration_ms":17111,"temperature":1.0,"reasoning_tokens":2495,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T18:28:46.661230+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a population whose causal feature law is a smooth density plus a small point mass placed on the optimal Voronoi boundary for some K. Because P(dist{U,B} ≤ t) is bounded below by the atom mass for every t, Assumption 4 fails; if the simultaneous band for ρ(K) still covers at nominal level then the margin condition is not needed for coverage, and if it undercovers the paper's stated limitation bites.","supporting_citations":[],"review_version":1}