REVIEW 3 major objections 5 minor 22 references
Which Effect of Race? Causal Inference without Holding All Else Equal
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read The same randomized design can identify a family of causal effects of race, not just the all-else-equal effect.
desk verdict A serious, well-argued framework that separates estimand choice from recovery in race effects; the formal core is sound conditional on a contested product-space premise, and the application needs weight-robustness checks before the empirical claim stands. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The configuration space T = {0,1} × V, a product space in which race and nonracial profile can be paired freely because membership runs through descent, a criterion outside the profile. On this space the potential-outcome function yi defines a family of unit-level race effects, collapsed into estimands by three components: contrast (how many nonracial features are held fixed), weighting, and subset. The AEE and AEWR are the extremes, with mixed effects in between; Proposition 1 shows unit-blindness alone secures unbiased recovery of any member, with the plug-in estimator reweighting cell means to the target weights.
What would settle it
Show a single case in which the criterion for racial membership is not fixed independently of the nonracial profile—for example, a person's race changes solely because their neighborhood, accent, or record changes—and the product-space premise fails. Alternatively, in a perception-based experiment with measured perceived race and nonracial perceptions, demonstrate that the ratio estimator under Assumption 2 does not equal a directly elicited natural race effect among compliers, which would falsify the recovery claim.
Extended reading notes
Core claim
On the paper's own terms, a contrast in which nonracial attributes move with racial membership is a well-defined causal effect in the potential-outcomes framework. The central result (Proposition 1) is that a unit-blind assignment—identical assignment probabilities across decision-maker units—recovers every member of the estimand family without bias, provided every profile with positive weight has positive assignment probability. Profile-blindness, the condition that race is assigned independently of nonracial profile, is neither necessary nor sufficient for inference; it selects the all-else-equal member. The same decomposition holds in the perception-based regime, where a weaker condition
Load-bearing premise
The framework requires that racial membership be fixed by descent (actual or presumed) independently of the nonracial profile; if a strong constructivism is right that race is constituted wholly by phenotype and indexed attributes, then atypical configurations become ill-defined and both the AEE and the AEWR lose their status as well-defined effects.
Editorial extensions
If this is right
- Studies that currently justify all-else-equal designs on credibility grounds can instead report a family of estimands from the same experiment.
- A within-race effect can be causal and estimable without manipulating race.
- In perception-based audits, researchers need only assume no isolated nonracial shift for non-compliers to estimate the NREC; cue stability for everyone is stronger than needed.
- Researchers can choose weights—for example, race-conditional campaign distributions—to represent how categories actually exist, making typicality and external validity design choices.
Reading between the lines
- If unit-blindness suffices, many natural experiments and observational designs that achieve assignment independent of units—but not profile balance—can be repurposed to estimate within-race effects, expanding the design space beyond randomized audits.
- The framework implies that 'manipulation checks' in race-cue experiments should measure joint movement of racial and nonracial perceptions, rather than aiming for isolated perceived-race shifts; a cue that moves several perceptions is informative, not contaminated.
- A testable extension is to reanalyze published audit studies where within-group distributional data exist: report both AEE and AEWR to map how conclusions depend on the conception of the category.
- The same logic should apply to other identity categories that index associated traits—religion, ethnicity, language—suggesting that all-else-equal designs in those fields also embed an unacknowledged estimand choice.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper separates two roles of randomization in discrimination experiments: selecting an estimand and recovering it. It defines a family of causal race estimands over configurations T={0,1}×V—the all-else-equal effect (AEE), the all-else-within-race effect (AEWR), and mixed contrasts—and, in the perception-based regime, the natural race effect among compliers (NREC). Proposition 1 states that a unit-blind assignment with fixed cell counts recovers every member of the family by reweighting cell means, while profile-blindness is neither necessary nor sufficient for recovery and instead selects the AEE. Propositions 2 and 3 show that identifying the NREC requires a no-isolated-nonracial-shift condition on non-compliers, a weaker condition than full cue stability. The application reanalyzes the Zárate et al. candidate-evaluation experiment and reports AEWR=0.10 (95% CI [0.04,0.15]) versus AEE=−0.01 (95% CI [−0.05,0.03]).
Significance. If the formal results stand, the contribution is substantial: it gives a precise language for a debate that has largely been conducted in prose, proves that profile-blindness is not required for unbiased recovery, and weakens the excludability condition for perception-based designs from cue stability for all units to a condition on non-compliers only. The supplement contains careful proofs of unbiasedness, exact variance, conservative variance bounds, and consistency for the plug-in estimator. The empirical illustration is transparent and shows that the estimand choice can reverse a substantive conclusion. The main risk is not in the algebra but in the philosophical premise required for the potential outcomes to be well defined, and in the strength of Assumption 2; both need explicit attention before the claims are accepted at full generality.
major comments (3)
- [Section 3.1, Eq. (1), Proposition 1] The entire estimand family presupposes that T={0,1}×V is a product space with potential outcomes y_i(z,v) defined at every pair. The only justification offered is that racial membership is fixed by descent, 'a criterion outside the nonracial profile.' If one accepts the strong constructivist reading of Hu and Kohler-Hausmann (2025) that racial and nonracial features 'co-constitute the category of race,' then a configuration such as a White arrestee in Brownsville is not merely atypical but ill-defined: y_i(0,v_Brownsville) in Eq. (1) would have no referent, and Proposition 1 would have no domain. Section 2 cites descent-based theories, but it does not show that those theories defeat the co-constitution thesis; it asserts the premise. This is load-bearing: the formal claims should either be explicitly restricted to a stated descent-based membership assumption, or the paper should argue wh
- [Section 6.2, Assumption 2, Propositions 2–3] Assumption 2 (no isolated nonracial shift) is necessary for the Wald ratio in Eq. (7) to equal the NREC. It is weaker than full cue stability, but it is still strong: non-compliers must have perfectly cue-invariant nonracial perceptions. The paper's own Figure 7 shows that 'Josué' signals foreignness and class relative to White-coded names. A decision-maker who does not perceive the candidate as Hispanic under either name—a non-complier—may nevertheless read the two names differently on these nonracial dimensions, violating Assumption 2. In the Spanish conditions the first stage is weak (text after Figure 6), so non-compliers may be numerous and the bias from a violation potentially large. Please provide a sensitivity analysis over plausible violations of Assumption 2, or at least a substantive discussion of the likely sign and magnitude of the resulting bias in the NREC estimates.
- [Section 8.1, Table 1] The AEWR point estimate and confidence interval treat q_0 and q_1 as fixed weights, but these weights are estimated from 126 campaign candidacies (49 Anglo, 77 Hispanic). Section 4 states that the weights are researcher-specified and not estimated from data, yet the application derives them from empirical frequencies. If they are estimates, the reported 95% CI [0.04,0.15] ignores their sampling uncertainty; if they are fixed design weights, the basis for fixing them at the observed frequencies needs justification. Please clarify the status of these weights and, if they are estimated, either propagate the uncertainty or provide a sensitivity analysis.
minor comments (5)
- [References / Supplementary Material] Several works cited in the supplement are missing from the reference list, including Kaufman et al. (2026), Neyman (1923), Li and Ding (2017), Aronow et al. (2014), Cochran (1977), Gaddis (2017), Nosofsky (1986), and Shepard (1987). The bibliography should be completed.
- [Abstract and Conclusion] The abstract states that the reanalysis finds 'coethnic preference among Hispanic voters under the within-race effect.' This is an estimand-dependent conclusion: the AEWR lets language vary with race, so under a classical conception of race it is better described as a combined race-language effect. Section 8.3 is careful about this, but the abstract and conclusion state it more categorically.
- [Section 7.1 and Proposition S4] The text says the variance of the plug-in estimator 'grows as the weights g_z(v) depart from the assigned shares.' This is only true holding within-cell variances and cross-cell covariances constant; the exact variance in Proposition S4 also depends on S²(z,v). Suggest rephrasing to avoid implying a purely monotone relationship.
- [Notation] The population mean of potential outcomes is written both as \bar y(z,v) (Proposition S1) and as μ(z,v) (Section S1.8). Unify the notation to avoid confusion.
- [Figure 2 caption] The caption refers to contours enclosing 'at least a 1−ε share' of a group's mass, but ε is not defined in the main text and is first introduced in Proposition S2 in the supplement. Please define it in the caption or in Section 5.
Circularity Check
No load-bearing circularity; identification theorems are self-contained given the paper's stated assumptions.
full rationale
The paper's central formal claims are derived from stated definitions and the randomization benchmark, not from fitted inputs or self-citations. The estimands (AEE, AEWR, NREC) are defined a priori as linear functionals of potential outcomes (Eqs. 4-6), and Proposition 1's unbiasedness result is proved in Section S1.8 by linearity of expectation under complete random assignment; the proof never assumes the estimand it targets. Part (ii) of Proposition 1 is a Bayes-law equivalence between profile-blindness and a common observed weighting, so it is a structural fact about the assignment rather than a prediction from data. Proposition 2's NREC identification follows from Assumption 2 plus standard IV logic, and Proposition 3 is an algebraic decomposition of cue stability. The application's AEE-vs-AEWR contrast uses external campaign-language weights (Zarate et al. 2024) applied to the same experimental cell means; the divergence is an empirical comparison, not manufactured by estimation. The self-citations to Leavitt and Rivera-Burgos (2024, 2026) appear only as sensitivity-analysis analogies in Section 6.2 and Supplement S1.4, and are not premises of the main theorems. The load-bearing product-space premise (race membership fixed by descent outside the non-racial profile) is a substantive philosophical assumption supported by external citations, not a self-referential derivation; rejecting it would undermine the estimands' interpretation, but that is assumption-dependence, not circularity. The paper's own limitations—e.g., "This paper does not answer these questions" and developing normative grounds least—are scope statements, not circular moves. No specific equation reduces to its input by construction.
Assumptions & free parameters
free parameters (6)
- q_0 (within-Anglo language weights, application) =
Non-Native Spanish 0.96, Native Spanish 0.04
- q_1 (within-Hispanic language weights, application) =
Non-Native 0.17, Native 0.83
- ρ (common language weight, application) =
Non-Native 0.48, Native 0.52
- λ_z (typicality concentration parameters) =
λ_0 ≈ 3.16, λ_1 ≈ 1.59
- prototypes p_z and salience weights a_z
- complier share |C|/N
assumptions (10)
- domain assumption SUTVA (Assumption S1): potential outcomes depend only on the unit's own configuration
- domain assumption Cue-level SUTVA (Assumption S2)
- domain assumption Perception sufficiency (Assumption 1): outcome depends on the cue only through the perceived configuration
- ad hoc to paper Assumption 2 (No isolated nonracial shift): non-compliers' nonracial perceptions are cue-invariant
- domain assumption No defiers (Assumption S3)
- domain assumption Existence of at least one complier (Assumption S4)
- domain assumption Complete random assignment of configurations/cues (Assumptions 3 and 4)
- domain assumption Racial membership is fixed by descent, actual or presumed, independently of the nonracial profile
- domain assumption Researcher-specified weights q_z and ρ define the estimand and carry no hat
- standard math Finiteness of V and standard probability/normalization facts
invented entities (3)
-
AEWR — All-Else-Within-Race Effect
independent evidence
-
NREC — Natural Race Effect among Compliers
independent evidence
-
Mixed-contrast effects
independent evidence
Cite this review
Pith. "Pith review of Which Effect of Race? Causal Inference without Holding All Else Equal." pith.science (2026). https://pith.science/paper/RX7M7ITC
@misc{pith2026260716371,
author = {Pith},
title = {Pith review of: Which Effect of Race? Causal Inference without Holding All Else Equal},
year = {2026},
howpublished = {\url{https://pith.science/paper/RX7M7ITC}},
note = {Machine review of arXiv:2607.16371}
}
read the original abstract
Empirical studies of racial discrimination vary race while holding nonracial traits fixed, a design the literature defends as what credible inference requires. This defense bundles two claims: which effect of race a study should target, and whether its design can recover that effect. I separate them. The same randomization that secures credible estimation and inference recovers a family of race estimands, from the all-else-equal effect to a within-race effect that lets associated traits vary with race. Every member of that family is causal rather than descriptive, and the choice among members is a claim about what a racial category is -- a claim extending to ethnicity, religion, and other identity categories that index associated traits. I derive conditions, weaker than the literature's, for credible estimation and inference. Reanalyzing a Spanish-language campaign experiment, I find coethnic preference among Hispanic voters under the within-race effect and none under the all-else-equal effect.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Identification of Causal Effects Using Instrumental Variables
Angrist, Joshua D., Guido W. Imbens, and Donald B. Rubin (1996). “Identification of Causal Effects Using Instrumental Variables”. In:Journal of the American Statistical Association91.434, pp. 444–455
1996
-
[2]
and Jörn-Steffen Pischke (2008).Mostly Harmless Econometrics: An Empiricist’s Companion
Angrist, Joshua D. and Jörn-Steffen Pischke (2008).Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton, NJ: Princeton University Press
2008
-
[3]
Randomization-based Con- fidence Sets for the Local Average Treatment Effect
Aronow, Peter M., Haoge Chang, and Patrick Lopatto (2026). “Randomization-based Con- fidence Sets for the Local Average Treatment Effect”. In:Biometrika113.2, asag010
2026
-
[4]
Sharp Bounds on the Variance in Randomized Experiments
Aronow, Peter M., Donald P. Green, and Donald K. K. Lee (2014). “Sharp Bounds on the Variance in Randomized Experiments”. In:The Annals of Statistics42.3, pp. 850–871
2014
-
[5]
Improving the External Validity of Conjoint Analysis: The Essential Role of Profile Distribution
Cochran, William G. (1977).Sampling Techniques. 3rd. Hoboken, NJ: John Wiley & Sons. de la Cuesta, Brandon, Naoki Egami, and Kosuke Imai (2022). “Improving the External Validity of Conjoint Analysis: The Essential Role of Profile Distribution”. In:Political Analysis30.1, pp. 19–45
1977
-
[6]
Signaling Race, Ethnicity, and Gen- der with Names: Challenges and Recommendations
Elder, Elizabeth Mitchell and Matthew Hayes (2023). “Signaling Race, Ethnicity, and Gen- der with Names: Challenges and Recommendations”. In:The Journal of Politics85.2, pp. 764–770
2023
-
[7]
How Black Are Lakisha and Jamal? Racial Perceptions from Names Used in Correspondence Audit Studies
Gaddis, S. Michael (2017). “How Black Are Lakisha and Jamal? Racial Perceptions from Names Used in Correspondence Audit Studies”. In:Sociological Science4, pp. 469–489. 115 Gärdenfors, Peter (2000).Conceptual Spaces: The Geometry of Thought. Cambridge, MA: The MIT Press. — (2014).The Geometry of Meaning: Semantics Based on Conceptual Spaces. Cambridge, MA...
2017
-
[8]
and Donald P
Gerber, Alan S. and Donald P. Green (2012).Field Experiments: Design, Analysis, and Interpretation. New York, NY: W.W. Norton
2012
Show all 22 references
-
[9]
Causal Inference in Conjoint Analysis: Understanding Multidimensional Choices via Stated Preference Experiments
Hainmueller, Jens, Daniel J. Hopkins, and Teppei Yamamoto (2014). “Causal Inference in Conjoint Analysis: Understanding Multidimensional Choices via Stated Preference Experiments”. In:Political Analysis22.1, pp. 1–30
2014
-
[10]
Statistics and Causal Inference
Holland, Paul W. (1986). “Statistics and Causal Inference”. In:Journal of the American Statistical Association81.396, pp. 945–960
1986
-
[11]
Identification and Estimation of Local Average Treatment Effects
Imbens, Guido W. and Joshua D. Angrist (1994). “Identification and Estimation of Local Average Treatment Effects”. In:Econometrica62.2, pp. 467–475
1994
-
[12]
Robust, Accurate Confidence Intervals with a Weak Instrument: Quarter of Birth and Education
Imbens, Guido W. and Paul R. Rosenbaum (2005). “Robust, Accurate Confidence Intervals with a Weak Instrument: Quarter of Birth and Education”. In:Journal of the Royal Statistical Society: Series A (Statistics in Society)168.1, pp. 109–126
2005
-
[13]
and Donald B
Imbens, Guido W. and Donald B. Rubin (2015).Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. New York, NY: Cambridge University Press
2015
-
[14]
Inference for Instrumental Vari- ables: A Randomization Inference Approach
Kang, Hyunseung, Laura Peck, and Luke Keele (2018). “Inference for Instrumental Vari- ables: A Randomization Inference Approach”. In:Journal of the Royal Statistical Soci- ety. Series A: Statistics in Society181.4, pp. 1231–1254
2018
-
[15]
Improving Compliance in Experimental Studies of Discrimination
Kaufman, Aaron R., Jacob M. Grumbach, and Christopher Celaya (2026). “Improving Compliance in Experimental Studies of Discrimination”. In:The Journal of Politics 88.1
2026
-
[16]
Do Name-Based Treatments Vio- late Information Equivalence? Evidence from a Correspondence Audit Experiment
Landgrave, Michelangelo and Nicholas Weller (2022). “Do Name-Based Treatments Vio- late Information Equivalence? Evidence from a Correspondence Audit Experiment”. In: Political Analysis30.1, pp. 142–148. 116
2022
-
[17]
Audit Experiments of Racial Discrim- ination and the Importance of Symmetry in Exposure to Cues
Leavitt, Thomas and Viviana Rivera-Burgos (2024). “Audit Experiments of Racial Discrim- ination and the Importance of Symmetry in Exposure to Cues”. In:Political Analysis 32.4, pp. 445–462. — (2026). “Navigating the Mismeasurement of Intermediary Variables in Message-Based Exp...
2024
-
[18]
General Forms of Finite Population Central Limit The- orems with Applications to Causal Inference
Li, Xinran and Peng Ding (2017). “General Forms of Finite Population Central Limit The- orems with Applications to Causal Inference”. In:Journal of the American Statistical Association112.520, pp. 1759–1769
2017
-
[19]
Sur les applications de la théorie des probabilités aux experiences agricoles: Essai des principes
Neyman, Jersey (1923). “Sur les applications de la théorie des probabilités aux experiences agricoles: Essai des principes”. In:Roczniki Nauk Rolniczych10, pp. 1–51
1923
-
[20]
Attention, Similarity, andthe Identification-Categorization Relationship
Nosofsky, Robert Mark (1986). “Attention, Similarity, andthe Identification-Categorization Relationship”. In:Journal of Experimental Psychology: General115.1, pp. 39–61
1986
-
[21]
Note on the Delta Method for Finite Population Inference with Applications to Causal Inference
Pashley, Nicole E. (2022). “Note on the Delta Method for Finite Population Inference with Applications to Causal Inference”. In:Statistics & Probability Letters188, p. 109540
2022
-
[22]
Toward a Universal Law of Generalization for Psychological Science
Shepard, Roger N. (1987). “Toward a Universal Law of Generalization for Psychological Science”. In:Science237.4820, pp. 1317–1323. Zárate, Marques G., Enrique Quezada-Llanes, and Angel D. Armenta (2024). “Se Habla Español: Spanish-Language Appeals and Candidate Evaluations in ...
1987
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.