Pith. sign in

REVIEW 3 major objections 3 minor 1 references

Non-Existent Outcomes in Research on Inequality: A Causal Approach

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper argues that routine practice of dropping cases whose outcome doesn't exist—like wages for the unemployed—can understate causal effects, and replaces it with two principal-stratification estimands that separate the effect on whethe

desk verdict The conceptual point is solid and applied researchers will care, but the supplied full text is unreadable and the always-survivor estimand needs an explicit identifying restriction; worth sending to peer review after a clean manuscript is provided. read the letter →

arxiv 2508.14770 v1 pith:CT6MCUO4 submitted 2025-08-20 stat.ME stat.AP

classification stat.MEstat.AP MSC 62D2062P25
keywords non-existentoutcomesprincipalstratificationcausalinferencecomplete-caseanalysisselectiononobservablestreatmenteffectslabormarketinequalityparenthood
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper targets a problem hidden in plain sight in inequality research: many outcomes, such as wages, exist only for part of the population, and researchers routinely drop the people without them. It shows that when a treatment changes both who gets the outcome and what its value is, that dropping can make the true causal effect look smaller than it is—for both beneficial and harmful treatments. The proposed fix is to report two causal effects instead of one: the average effect on whether the outcome exists, and the average effect on its value among the latent subgroup for whom the outcome exists in either treatment condition. A regression-and-simulation implementation extends this approach to the selection-on-observables designs common in observational social science, and the paper illustrates it with parenthood and labor-market outcomes. If the argument holds, researchers can stop silently conditioning on outcome presence and instead separate existence effects from value effects.

What carries the argument

Principal stratification with an 'existence' event: each unit has a potential existence indicator under treatment and under control, placing it in a latent principal stratum (always-reporter, never-reporter, reporter only under treatment, reporter only under control). The two target estimands are the average causal effect on the existence indicator and the average causal effect on the value among the always-reporter stratum—the subgroup whose outcome would exist in either condition. Identification comes from selection-on-observables: regressions for existence and value, conditional on measured confounders, are used to simulate the full potential-outcome distribution and average over the impl

What would settle it

Generate simulated data with known potential outcomes in which a treatment changes both whether the outcome exists and its value among always-reporters, then compare the standard complete-case estimate with the paper's two estimands. The paper predicts the complete-case estimate is biased away from the true always-reporter effect; data or experiments where the two agree across a wide range of such settings would refute the central claim.

Watch

Extended reading notes

Core claim

The central claim is that conditioning on outcome presence does not estimate a causal effect on any well-defined population; it mixes people whose outcome exists under both treatment states with people whose outcome exists only under one state. If treatment shifts existence, this mixture changes, so the complete-case comparison can understate effects of both beneficial and harmful treatments. The paper proposes two principal-stratification estimands—the average effect on outcome existence and the average effect on outcome value among the latent always-reporter subgroup—and provides a regression-simulation implementation for selection-on-observables designs, illustrated on parenthood and labo

Load-bearing premise

The estimates are only valid if, after measured confounders are controlled, treatment is as good as randomly assigned for both outcome existence and outcome value, and the existence effect is monotone; neither condition can be verified from observed data.

Editorial extensions

If this is right

  • Complete-case estimates should no longer be interpreted as the effect of the treatment on outcome values; the paper's decomposition supplies the quantities that answer that question.
  • If a treatment increases outcome existence, the standard estimate will tend to understate the effect among always-reporters; if it decreases existence, understatement can occur in the harmful direction as well.
  • In observational settings with measured confounders, the regression-simulation framework yields point estimates for both existence and always-reporter value effects.
  • In the parenthood and labor-market example, the effect on employment and the effect on wages among always-employed people become separable, so the usual 'motherhood penalty' estimates are interpretable rather than mixed with selection.
  • The proposed first estimand makes the effect on whether an outcome exists an explicit, reportable causal quantity rather than a nuisance to be conditioned away.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same two-estimand logic applies to any outcome defined only conditionally—test scores among test-takers, reports among victims, earnings among taxpayers—so the framework is a general template rather than a labor-market fix.
  • The paper's selection-on-observables implementation could be paired with a formal sensitivity analysis that asks how much unmeasured confounding or non-monotonicity would be needed to erase an estimated always-reporter effect; the estimands give that analysis a well-defined target.
  • With randomized treatment assignment, the always-reporter value effect could be estimated nonparametrically, which suggests the observational framework is one practical special case of a broader identification strategy.
  • Published null or small effects on conditional outcomes may, in some settings, be artifacts of conditioning on outcome presence; re-analyzing existing datasets with the two-estimand decomposition would test that possibility directly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper addresses a common practice in social stratification research: dropping cases for which the outcome is non-existent (e.g., wages for non-employed persons). The authors argue that when a treatment affects both whether the outcome exists and its value, this practice can bias estimated causal effects, and they show that effects of both beneficial and harmful treatments can be underestimated. They propose two principal-stratification estimands: (1) the average effect on outcome existence and (2) the average effect on the outcome among the latent subgroup whose outcome would exist under either treatment condition (the always-survivor stratum). The paper extends this to selection-on-observables settings via a regression-and-simulation framework, and illustrates the approach with an application to the effects of parenthood on labor market outcomes. The conceptual point is plausible and consistent with the truncation-by-death literature, but the supplied full text is heavily corrupted and unreadable, so the identification arguments, simulation details, and empirical results cannot be verified.

Significance. If the methodological claims could be verified, the paper would make a useful contribution to applied inequality research by separating existence effects from value effects and by cautioning against silent conditioning on outcome presence. The proposed approach draws on established principal-stratification ideas, which is a strength. However, the manuscript as submitted is not in readable form: the full text is mojibake, and the only clear statement is the abstract. No proofs, derivations, simulation code, or reproducible results are available to check. The central conceptual claim is not new in the causal inference literature, but the paper's contribution would be in translating it for applied researchers and providing a practical estimator. That contribution cannot be assessed from the current submission.

major comments (3)
  1. [Full text (entire submitted manuscript)] The submitted full text is unreadable: it consists of corrupted character sequences rather than coherent prose, equations, or tables. No identification proof, simulation setup, or empirical analysis can be checked. This is a load-bearing issue because the paper's central claim rests on the regression-simulation framework correctly identifying the always-survivor effect. The authors must provide a complete, readable manuscript before any substantive review can occur.
  2. [Abstract (regression-simulation framework)] The abstract states that the framework 'adjust[s] for measured confounders' and enables principal stratification estimates, but it does not state the identifying restriction needed to recover the always-survivor stratum. Even under ignorability conditional on covariates, the distribution of potential outcomes among the always-survivor group is not identified from observed data without an additional structural assumption such as monotonicity of existence in treatment. The provided text does not show whether such an assumption is made and defended. If the framework simply imputes missing outcomes from a regression fitted on observed survivors, it targets the effect among observed survivors, not among always-survivors. The authors must state and justify the identifying assumptions explicitly.
  3. [Empirical example (parenthood and labor market)] The parenthood application is mentioned in the abstract but the corresponding analysis is not readable. Given the identification concern above, the application is especially vulnerable: monotonicity of employment existence with respect to parenthood is doubtful, since parenthood may lead some people to leave employment and others to enter or intensify work. Without a defense of monotonicity or a sensitivity analysis relaxing it, the empirical estimates are not interpretable as always-survivor effects. The authors should clarify how the application handles this.
minor comments (3)
  1. [Full text] The document contains an unrelated arXiv identifier ('arXiv:2508.14772v2 [cond-mat.str-el]') and other extraneous artifacts. These should be removed in a clean submission.
  2. [Abstract] The phrase 'non-existent outcomes' is intuitive but the paper should more explicitly link it to the established 'truncation-by-death' literature in the abstract, which would help situate the contribution.
  3. [General] Once a readable manuscript is provided, the authors should include a formal definitions section with notation for potential outcomes, existence indicators, and principal strata, since these are central to the proposed estimands.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the bias claim is an analytic decomposition and the principal-stratification estimator is a model-based identification, not a re-labeling of fitted inputs.

full rationale

The abstract's central claim—that dropping cases with non-existent outcomes can obscure causal effects because treatment affects both existence and value—rests on a non-circular decomposition: the observed outcome conditional on existence is a mixture over principal strata, and the commonly used conditional estimand is a weighted combination of the existence effect and the value effect. That is an algebraic identity rather than a conclusion assumed in its own premises. The proposed estimands (average effect on existence; average effect among the latent always-survivor stratum) are standard principal-stratification quantities, and the regression-and-simulation implementation is a model-based identification strategy: under selection-on-observables plus monotonicity or an equivalent identifying restriction, the target is a function of observed-data regressions. The estimates are therefore model-dependent and assumption-dependent, and the always-survivor estimand is not nonparametrically identified without assumptions such as monotonicity—but dependence on untestable identifying assumptions is a correctness/robustness issue, not circularity. No self-citation, uniqueness theorem, or ansatz-via-citation is quoted as load-bearing; nor is any fitted parameter renamed as a prediction. The supplied full text is heavily corrupted, but the readable abstract and fragments do not exhibit any equation that reduces to its own input by construction.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

Everything the central claim rests on. The conceptual claim (dropping non-existent outcomes biases effects) rests on standard counterfactual logic and the known truncation-by-death phenomenon. The two proposed estimands inherit the assumptions of the principal-stratification theory the paper cites. The simulation implementation adds the selection-on-observables assumption. The regression coefficients are fitted from data but only for the empirical illustration, not for the conceptual result. No new entities are introduced: the latent subgroup whose outcome would exist in either treatment condition is the classic always-survivor principal stratum.

free parameters (1)
  • Coefficients of the outcome-existence and outcome-value regression models = Not stated in abstract
    The simulation-based estimator fits models to the observed parenthood sample and simulates counterfactual outcomes; the resulting estimates inherit these fitted coefficient values. This is implementation-level fitting and does not affect the conceptual bias claim.
assumptions (3)
  • domain assumption Conditional ignorability (no unmeasured confounding) given measured confounders
    The abstract's 'selection-on-observables settings... adjust for measured confounders' presumes the regression and simulation framework identifies the principal-stratification estimands only if all confounders of the treatment effects on outcome existence and value are observed.
  • domain assumption Principal-stratification restrictions (e.g., monotonicity of existence in treatment, or an exclusion restriction)
    The always-survivor estimand (effect among those whose outcome would exist in either condition) is not point-identified without such restrictions; the abstract says only that the paper draws on existing principal-stratification approaches.
  • standard math Standard causal assumptions: consistency, no interference, positivity
    Implicit in any counterfactual definition of average causal effects among latent subgroups; not stated in the abstract but required for the estimands to be well-defined.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Non-Existent Outcomes in Research on Inequality: A Causal Approach." pith.science (2026). https://pith.science/paper/CT6MCUO4

@misc{pith2026250814770,
  author       = {Pith},
  title        = {Pith review of: Non-Existent Outcomes in Research on Inequality: A Causal Approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CT6MCUO4}},
  note         = {Machine review of arXiv:2508.14770}
}
read the original abstract

Scholars of social stratification often study exposures that shape life outcomes. But some outcomes (such as wage) only exist for some people (such as those who are employed). We show how a common practice -- dropping cases with non-existent outcomes -- can obscure causal effects when a treatment affects both outcome existence and outcome values. The effects of both beneficial and harmful treatments can be underestimated. Drawing on existing approaches for principal stratification, we show how to study (1) the average effect on whether an outcome exists and (2) the average effect on the outcome among the latent subgroup whose outcome would exist in either treatment condition. To extend our approach to the selection-on-observables settings common in applied research, we develop a framework involving regression and simulation to enable principal stratification estimates that adjust for measured confounders. We illustrate through an empirical example about the effects of parenthood on labor market outcomes.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

1 extracted references · 1 canonical work pages

  1. [1]

    ����� ��� ����������� �� ������ ���������� ����� ���� ������ �� � ������ � ������� ����� ���� � ������ ������ ��� ���� �� ���� �� �� ��� � ������������� ������ ��� ������� ���������� ������ �� �������� ������ ����������� ������� ������� ����� � ��������� ��� ���������� �������� ���� ���� ����������� �������� ������� ����� � ������������� ���������� ������...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.