Pith. sign in

REVIEW 3 major objections 4 minor 81 references

The Role of Confounders and Linearity in Ecological Inference: A Reassessment

T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read This paper argues that ecological inference is a confounded regression problem, and that aggregation imposes a partially linear structure that unifies all existing methods.

desk verdict A genuinely useful reformulation of ecological inference as confounded regression, with solid formal results and valuable ground-truth validations, but the paper overclaims by stating that all EI methods fail when CAR is violated. read the letter →

arxiv 2601.07668 v2 pith:3XXFSXRS submitted 2026-01-12 stat.AP

classification stat.AP MSC 62P2562J05
keywords ecologicalinferencecoarseningatrandomconfoundingpartiallylinearmodelracialpolarizationticketsplittingaggregatedatacausal
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the ecological inference problem—estimating individual-level conditional means from aggregate data—should be understood as a problem of confounding and linearity. It formalizes a coarsening-at-random assumption analogous to selection on observables, under which the global estimand is identified and the aggregate outcome is partially linear in the group shares. This reformulation shows that widely used methods are essentially the same linear model, differing only in distributional assumptions. The paper also uses voter-file and cast-vote-record data to show that common EI methods systematically overestimate racial polarization and underestimate ticket splitting, and that adding covariates can reduce the bias.

What carries the argument

The accounting identity Y_g = X_g^T B_g—the aggregate outcome is exactly a weighted average of unobserved group means with weights equal to group shares—combined with the coarsening-at-random assumption. This forces the conditional expectation of the aggregate outcome to be partially linear in group shares (a varying-coefficient model), which is the linchpin of both identification and the unifying regression framework.

What would settle it

Simulate data with a known individual-level truth, aggregate it in a way that satisfies coarsening at random by construction (e.g., drawing local group means independently of group shares), then apply the proposed plug-in regression; if the estimate deviates from the true global mean by more than sampling error, the identification argument is wrong. Conversely, in the North Carolina voter-file validation, test whether the residual association between local group means and group shares survives conditioning on the paper's covariates; if it does, CAR is violated and the paper's estimates would b

Watch

Extended reading notes

Core claim

The central discovery is that aggregation imposes a strong functional-form restriction on the conditional expectation function: under coarsening at random, E[Y_g | Z_g, X_g] = f(Z_g)^T X_g, where X_g is the vector of group shares. This partial linearity means that the analyst only needs to interact covariates with group categories, and that OLS is a natural estimator. The paper further shows that the identification condition is equivalent to requiring that the unobserved local group means be independent of the group composition after conditioning on covariates, and that violations of this condition produce exactly the ecological fallacies documented in the literature. Finally, the paper demo

Load-bearing premise

The accounting identity Y_g = X_g^T B_g holds exactly—aggregate outcomes are error-free, perfectly aligned linear combinations of unobserved group means and observed group shares; any measurement error or misalignment breaks the linearity and the identification argument.

Editorial extensions

If this is right

  • Under coarsening at random, the global mean for each group is identified by a weighted average of predicted outcomes from a regression of the aggregate outcome on group shares interacted with covariates, evaluated at that group's share equal to one.
  • All point-identification methods for ecological inference—from the classic regression approach to the random-coefficient models and count models—are special cases of the same partially linear regression; differences reduce to error distribution, bounds, and whether covariates are included.
  • If coarsening at random fails, every EI method is biased; the direction of the bias can be anticipated from regression diagnostics such as extrapolation, influence, and collinearity.
  • Because the aggregate outcome is exactly linear in group shares, there is no functional-form ambiguity about how covariates and shares enter the model: interacting them is sufficient.
  • The empirical validations imply that published EI results on racial polarization and ticket splitting may be systematically off, and that including context covariates can partially correct them.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's formalization suggests that sensitivity analysis for unobserved confounders—like that used in causal inference—should become standard for ecological inference; one can adapt partial-identification bounds to report how large the omitted-confounder effect must be to change conclusions.
  • The same partial-linearity argument applies beyond political science to any setting where outcomes are exact aggregates of group-specific means (e.g., public-health incidence by race within census areas, market shares by consumer type); the bias mechanisms identified here should generalize.
  • A testable extension of the paper's empirical finding: if finer geographies reduce extrapolation, then precinct-level estimates should be more accurate than county-level estimates even when identification conditions are equally plausible; this could be verified with the same ground-truth data.
  • The strong claim that 'all methods fail' when CAR is violated means that the only defensible route for practitioners is to collect covariates that plausibly satisfy CAR, or to use sensitivity analysis; the paper does not, however, provide a formal test of CAR itself.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reassesses ecological inference (EI) by framing it as a missing-data/coarsening problem. It defines local group means B_gk and the global estimand beta, states an accounting identity Y_g = X_g^T B_g, and introduces coarsening completely at random (CCAR) and coarsening at random (CAR) identification conditions. Proposition 3.1 shows CCAR identifies Goodman regression; Proposition 3.2 shows CAR identifies beta via a plug-in formula. Section 3.3 derives partial linearity of the aggregate CEF under CAR and argues that fully interacted linear regression is the natural model. Section 4 reinterprets King's 2x2 model, R x C count models, and a semiparametric/double-machine-learning estimator (seine) within this regression framework. Section 5 validates methods on North Carolina voter-file data and cast vote records, finding that all tested methods overestimate racial polarization and underestimate ticket splitting, with covariates sometimes helping. The main identification propositions are short and essentially correct under CAR, but the paper overstates the necessity of CAR and the absence of functional-form concerns.

Significance. If the identification results are read as sufficient conditions and the overclaims are corrected, the paper is a valuable synthesis. It connects EI to causal-inference selection-on-observables, clarifies the role of covariates, and shows that aggregation imposes partial linearity in X. The comparison of King, Goodman, and R x C models under one regression umbrella is useful, and the two empirical validations with observed ground truth (NC voter file, cast vote records) are a strength. The paper is frank that all tested methods are biased in the applications. The self-cited semiparametric estimator is used in the empirical section, so the demonstration that covariates help is partly a proof-of-concept for the authors' own method; this should be disclosed but is not circular because the identification theory is independent of that estimator.

major comments (3)
  1. [Section 1, p.2; Section 3.2] The sentence "All methods of ecological inference fail to consistently estimate the quantity of interest when this condition does not hold" is false as stated. CAR is proved sufficient (Prop. 3.2), but no necessity theorem is supplied. A concrete counterexample: with K=2, X_g ~ U(0,1), B_g1 = beta1 + gamma X_g, B_g2 = beta2, CAR fails but Y_g = beta2 + (beta1 - beta2) X_g + gamma X_g^2, so OLS of Y on (1, X, X^2) consistently estimates beta = (beta1 + gamma E[X], beta2). The correct claim is that standard EI estimators that impose CCAR/CAR fail, or that CAR is sufficient. Please revise the universal claim and related statements that imply CAR is necessary.
  2. [Section 3.3, Eqs. (7)-(8)] The text states that if Z satisfies CAR and is fully interacted with X, "there is no concern for functional form misspecification." This is not implied by Eq. (7), where f(Z_g) is an arbitrary function of Z_g; Eq. (8) requires f_k(Z_g) to be linear in Z_g. The paper's own Section 4.3 allows nonlinear f via basis expansions. The "no functional form" claim should be restricted to the linearity of the CEF in X (the varying-coefficient structure), not the covariate response.
  3. [Proposition 3.2; Appendix B.3] Proposition 3.2 states beta_k = E[ E[Y | Z, X_k=1] N_k / E[N_k] ], but Section 3.1 defines the global mean as B_k = sum_g N_gk B_gk / sum_g N_gk and beta = E[B]. These are not the same object: the formula identifies E[N_gk f_k(Z)] / E[N_gk] (the probability limit of the N_gk-weighted average), whereas E[ sum_g N_gk B_gk / sum_g N_gk ] is a ratio of sums. The last step of B.3, "E[N_gk B_gk] = E[N_gk] beta_k," requires an additional exchangeability or asymptotic assumption. Please state the target estimand and the sampling model explicitly.
minor comments (4)
  1. [Section 3.1 and Proposition 3.2] The notation N_k is used both for the global count in Section 3.1 and for the geography-specific count inside the expectation in Proposition 3.2. Use N_gk consistently in the proposition and in the plug-in algorithm.
  2. [Section 3.6] Typo: "senstivity" should be "sensitivity."
  3. [Figure 2] In the discussion of the influence point, "remaining20variables" should read "remaining 20 observations."
  4. [Figure 7 and Section 4.3] The estimator is sometimes "seine" and sometimes "Seine"; please make the capitalization consistent.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: identification results are derived from stated assumptions, and the self-cited estimator is assessed against external ground-truth data.

full rationale

The paper's central identification chain is self-contained and does not reduce to its inputs. Proposition 3.2 defines the estimand as beta = E[B_g] and proves, under the explicit CAR assumption in Eq. 5, that beta_k = E[ E[Y | Z, X_k=1] N_k / E[N_k] ]; the proof is a direct law-of-total-expectation argument from the accounting identity Eq. 4 and the CAR condition, not a restatement of the estimand. Proposition 3.1 follows as the no-covariate special case, and the partial-linearity result in Eq. 7 is derived from CAR plus Eq. 4 rather than assumed. The semiparametric estimator seine is attributed to the authors' own McCartan and Kuriwaki (2025a), which is a self-citation, but it is used as an evaluation tool and benchmarked against external ground truth (North Carolina voter file and cast vote records). Those comparisons are externally falsifiable, so the self-citation does not carry the load of the empirical conclusions. The only substantive concern is the unsupported overclaim that all ecological inference methods fail when CAR does not hold; that is a correctness and scope issue, not a circularity, because no theorem in the paper forces it and the paper does not define EI methods so as to make the statement true by construction. No fitted parameter is relabeled as a prediction, and no uniqueness or ansatz is imported solely through self-citation.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The theoretical contribution is lean: it adds CAR/CCAR formalization and the partial-linearity observation on top of the accounting identity. The heavy domain assumptions are exact aggregation, CAR, positivity, plus proxy/sample assumptions in the empirical validations. No new ontological entities are introduced.

free parameters (2)
  • Biden voteshare bins = 5
    In the ticket-splitting seine model, precinct-level Biden voteshare is coarsened into five equally sized bins and entered as a covariate; the bin count is a hand-chosen discretization that affects the estimates.
  • Ridge penalty λ = selected by LOO-CV
    In the semiparametric estimator, the ridge regularization strength is chosen automatically by leave-one-out cross-validation; it is a tuning input rather than a parameter fitted to the target estimand.
assumptions (6)
  • standard math Accounting identity Y_g = X_g^T B_g holds exactly (Eq. 3-4).
    Definitional identity when aggregate means are exact; the entire linearity argument depends on it.
  • standard math Regularity conditions: E[XX^T] invertible and finite second moments (Prop 3.1).
    Standard population-regression regularity needed for identification of beta.
  • domain assumption CAR: E[B_g | Z_g, X_g, N_g] = E[B_g | Z_g] (Eq. 5).
    The central identification condition; the paper grants it may fail and provides sensitivity-analysis references.
  • domain assumption Positivity/overlap: sufficient residual variation in X after conditioning on Z.
    Discussed in Section 3.4 as a tradeoff but not formally stated as a required condition.
  • domain assumption Party registration is a valid proxy for vote choice in the North Carolina validation.
    Authors acknowledge surveys show the two are heavily correlated, but registration is not vote choice; biases in registration may not transfer exactly to vote choice.
  • domain assumption Cast vote records accurately represent ballots in the selected districts.
    CVR coverage varies by county and the analysis is restricted to districts with contested House races and at least 3 counties; nonignorable missingness could bias the ground truth.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Role of Confounders and Linearity in Ecological Inference: A Reassessment." pith.science (2026). https://pith.science/paper/3XXFSXRS

@misc{pith2026260107668,
  author       = {Pith},
  title        = {Pith review of: The Role of Confounders and Linearity in Ecological Inference: A Reassessment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3XXFSXRS}},
  note         = {Machine review of arXiv:2601.07668}
}
read the original abstract

Estimating conditional means using only the marginal means available from aggregate data is known as the ecological inference problem. We reassess this literature, arguing that it has understudied two issues: how practitioners should control for confounding, and how methodologists can leverage the linearity inherent in the structure of the problem. On the former, we formalize ignorability conditions like those in causal inference and outline consistent plug-in estimators: These are credible when covariates make the ignorability condition plausible. On the latter, we show that aggregation restricts the target function to be partially linear. Such linearity clarifies the connections between King's (1997) methodology, its predecessors, and subsequent developments. That motivates a recent doubly-robust technique that enters covariates flexibly while leveraging linearity. Finally, we test these methods in datasets where the ground truth is fortuitously observed. In these common applications, all methods tested were prone to overestimating racial polarization and underestimating split-ticket voting.

Figures

Figures reproduced from arXiv: 2601.07668 by the authors.

Figure 1
Figure 1. Examples of the Ecological Fallacy. X-axis shows the Black population as a share of the overall voting-age population (VAP). The y-axis shows the total number of votes cast for George Wallace as a share of VAP. The estimated VAP turnout in the two states was 54%, which was in turn split across Nixon (21%), Wallace (17%), and Humphrey (16%). Observers of these elections knew that these results could not be meant to i… view at source ↗
Figure 2
Figure 2. Intuition for EI as linear regression. A simulated example where the quantity of interest is 𝛽1 = 0.5. In panel (a), 𝑛 = 20 data points are simulated from the model of King (1997). A simple OLS fit evaluated at 𝑋1 = 1 provides an estimate that agrees with both King’s EI algorithm and the ground truth. In panel (b), an outlier with high leverage is added to the dataset, shown in the top-right part of the figure. The … view at source ↗
Figure 3
Figure 3. The Distribution of Party and Race Registration in North Carolina. [PITH_FULL_IMAGE:figures/full_fig_p022_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Accuracy of EI Methods in Uncovering Partisanship among Racial Groups. [PITH_FULL_IMAGE:figures/full_fig_p023_4.png]
Figure 5
Figure 5. Figure 5: Potential Confounders for Racially Polarized Voting. [PITH_FULL_IMAGE:figures/full_fig_p024_5.png]
Figure 6
Figure 6. Figure 6: Typical Data Aggregation Problem in Ticket Splitting in Wisconsin’s 8th Congres [PITH_FULL_IMAGE:figures/full_fig_p026_6.png]
Figure 7
Figure 7. Figure 7: Predictive Performance of Ticket Splitting. [PITH_FULL_IMAGE:figures/full_fig_p026_7.png]
Figure 8
Figure 8. Figure 8: Variation in Ticket Splitting Rates. The relationship between 𝑋 and quantities of interest across a random sample of 500 precincts with more than 100 voters for visual clarity. Each point is a sampled precinct, sorted on the horizontal axis by the percentage of total v…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 5 linked inside Pith

  1. [1]

    Achen, C. H. and Shively, W. P. (1995). Cross-level Inference . University of Chicago Press

  2. [2]

    and Cho, W

    Anselin, L. and Cho, W. K. T. (2002). Spatial effects and ecological inference. Political analysis , 10(3):276--297

  3. [3]

    and Rivers, D

    Ansolabehere, S. and Rivers, D. (1995). Bias in ecological regression. Massachusetts Institute of Technology

  4. [4]

    and Weymouth, S

    Baccini, L. and Weymouth, S. (2021). Gone for good: Deindustrialization, white voter backlash, and us presidential voting. American Political Science Review , 115(2):550--567

  5. [5]

    Barreto, M., Collingwood, L., Garcia-Rios, S., and Oskooii, K. A. (2022). Estimating candidate support in voting rights act cases: Comparing iterative ei and ei - R x C methods. Sociological Methods & Research , 51(1):271--304

  6. [6]

    and Haile, P

    Berry, S. and Haile, P. (2024). Nonparametric identification of differentiated products demand using micro data. Econometrica , 92(4):1135--1162

  7. [7]

    Berry, S., Levinsohn, J., and Pakes, A. (2004). Differentiated products demand systems from a combination of micro and macro data: The new car market. Journal of political Economy , 112(1):68--105

  8. [8]

    Blackwell, M. (2025). A User's Guide to Statistical Inference and Regression . CRC Press

Show all 81 references
  1. [9]

    Blalock, H. M. (1984). Contextual effects models: theoretical and methodological issues. Annual review of sociology , pages 353--372

  2. [10]

    Brown, P. J. and Payne, C. D. (1986). Aggregate data, ecological regression, and voting transitions. Journal of the American Statistical Association , 81(394):452--460

  3. [11]

    Burden, B. C. and Kimball, D. C. (1998). A new approach to the study of ticket splitting. American Political Science Review , 92(3):533--544

  4. [12]

    Burden, B. C. and Kimball, D. C. (2009). Why Americans split their tickets: Campaigns, competition, and divided government . University of Michigan Press

  5. [13]

    and Escolar, M

    Calvo, E. and Escolar, M. (2003). The local voter: A geographically weighted approach to ecological inference. American Journal of Political Science , 47(1):189--204

  6. [14]

    D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., and Riddell, A

    Carpenter, B., Gelman, A., Hoffman, M. D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., and Riddell, A. (2017). Stan: A probabilistic programming language. Journal of statistical software , 76:1--32

  7. [15]

    Chambers, R. L. and Steel, D. G. (2001). Simple methods for ecological inference in 2 x 2 tables. Journal of the Royal Statistical Society: Series A (Statistics in Society) , 164(1):175--192

  8. [16]

    and Hadi, A

    Chatterjee, S. and Hadi, A. S. (1986). Influential observations, high leverage points, and outliers in linear regression. Statistical science , pages 379--393

  9. [17]

    Chen, Q., Syrgkanis, V., and Austern, M. (2022). Debiased machine learning without sample-splitting for stable estimators. Advances in Neural Information Processing Systems , 35:3096--3109

  10. [18]

    Chernozhukov, V., Cinelli, C., Newey, W., Sharma, A., and Syrgkanis, V. (2022). Long story short: Omitted variable bias in causal machine learning. Technical report, National Bureau of Economic Research

  11. [19]

    K., and Robins, J

    Chernozhukov, V., Newey, W. K., and Robins, J. (2018). Double/de-biased machine learning using regularized riesz representers. Technical report, cemmap working paper

  12. [20]

    Cho, W. K. T. and Gaines, B. J. (2004). The limits of ecological inference: The case of split-ticket voting. American Journal of Political Science , 48(1):152--171

  13. [21]

    and Hazlett, C

    Cinelli, C. and Hazlett, C. (2020). Making sense of sensitivity: Extending omitted variable bias. Journal of the Royal Statistical Society Series B: Statistical Methodology , 82(1):39--67

  14. [22]

    and Stanig, P

    Colantone, I. and Stanig, P. (2018). Global competition and brexit. American political science review , 112(2):201--218

  15. [23]

    Cross, P. J. and Manski, C. F. (2002). Regressions, short and long. Econometrica , 70(1):357--368

  16. [24]

    P., Hennig, P., and Lacoste-Julien, S

    Cunningham, J. P., Hennig, P., and Lacoste-Julien, S. (2011). Gaussian probabilities and expectation propagation. arXiv preprint arXiv:1111.6832

  17. [25]

    P., Hitt, M

    Darr, J. P., Hitt, M. P., and Dunaway, J. L. (2018). Newspaper closures polarize voting behavior. Journal of Communication , 68(6):1007--1028

  18. [26]

    de Benedictis-Kessner , J. (2015). Evidence in voting rights act litigation: Producing accurate estimates of racial voting patterns. Election Law Journal , 14(4):361--381

  19. [27]

    Ding, P., Imbens, G., Qu, Z., and Ye, Y. (2024). Computationally efficient estimation of large probit models. arXiv preprint arXiv:2407.09371

  20. [28]

    Duncan, O. D. and Davis, B. (1953). An alternative to ecological correlation. American Sociological Review , 18(6)

  21. [29]

    E., and Morton, C

    Elzayn, H., Goldin, J., Guage, C., Ho, D. E., and Morton, C. (2025). Monotone ecological inference. National Bureau of Economic Research

  22. [30]

    and Zhang, W

    Fan, J. and Zhang, W. (1999). Statistical estimation in varying coefficient models. The annals of Statistics , 27(5):1491--1518

  23. [31]

    Fan, Y., Sherman, R., and Shum, M. (2016). Estimation and inference in an ecological inference model. Journal of Econometric Methods , 5(1):17--48

  24. [32]

    and Rosenman, E

    Fishman, N. and Rosenman, E. (2024). Estimating vote choice in us elections with approximate poisson-binomial logistic regression. In OPT 2024: Optimization for Machine Learning

  25. [33]

    R., Wang, Y.-X., and Smola, A

    Flaxman, S. R., Wang, Y.-X., and Smola, A. J. (2015). Who supported obama in 2012? ecological inference through distribution regression. In Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , pages 289--298

  26. [34]

    A., Klein, S

    Freedman, D. A., Klein, S. P., Sacks, J., Smyth, C. A., and Everett, C. G. (1991). Ecological regression and voting rights. Evaluation review , 15(6):673--711

  27. [35]

    K., Ansolabehere, S., Price, P

    Gelman, A., Park, D. K., Ansolabehere, S., Price, P. N., and Minnite, L. C. (2001). Models, assumptions and model checking in ecological regressions. Journal of the Royal Statistical Society Series A: Statistics in Society , 164(1):101--118

  28. [36]

    N., Wakefield, J., Handcock, M

    Glynn, A. N., Wakefield, J., Handcock, M. S., and Richardson, T. S. (2008). Alleviating linear ecological bias and optimal design with subsample data. Journal of the Royal Statistical Society Series A: Statistics in Society , 171(1):179--202

  29. [37]

    Goodman, L. A. (1953). Ecological regressions and behavior of individuals. American Sociological Review , 18(6)

  30. [38]

    Greiner, D. J. and Quinn, K. M. (2010). Exit polling and racial bloc voting: Combining individual-level and R x C ecological data. The Annals of Applied Statistics , pages 1774--1796

  31. [39]

    Greiner, J. D. and Quinn, K. M. (2009). R x C ecological inference: bounds, correlations, flexibility and transparency of assumptions. Journal of the Royal Statistical Society Series A: Statistics in Society , 172(1):67--81

  32. [40]

    Griffiths, W. E. (1972). Estimation of actual response coefficients in the hildreth-houck random coefficient model. Journal of the American Statistical Association , 67(339):633--635

  33. [41]

    and Merrill, S

    Grofman, B. and Merrill, S. (2004). Ecological regression and ecological inference. Ecological Inference: New Methodological Strategies , page Ch. 5

  34. [42]

    A., Jackson, J

    Hanushek, E. A., Jackson, J. E., and Kain, J. F. (1974). Model specification, use of aggregate data, and the ecological correlation fallacy. Political Methodology , pages 89--107

  35. [43]

    and Tibshirani, R

    Hastie, T. and Tibshirani, R. (1993). Varying-coefficient models. Journal of the Royal Statistical Society Series B: Statistical Methodology , 55(4):757--779

  36. [44]

    Heitjan, D. F. and Rubin, D. B. (1991). Ignorability and coarse data. The Annals of Statistics , pages 2244--2253

  37. [45]

    Herron, M. C. and Shotts, K. W. (2003). Using ecological inference point estimates as dependent variables in second-stage linear regressions. Political Analysis , 11(1):44--64

  38. [46]

    Imai, K., Lu, Y., and Strauss, A. (2008). Bayesian and likelihood inference for 2x2 ecological tables: an incomplete-data approach. Political Analysis , 16(1):41--69

  39. [47]

    Jiang, W., King, G., Schmaltz, A., and Tanner, M. A. (2020). Ecological regression with partial identification. Political Analysis , 28(1):65--86

  40. [48]

    King, G. (1997). A Solution to the Ecological Inference Problem: Reconstructing Individual Behavior from Aggregate Data . Princeton University Press

  41. [49]

    King, G., Rosen, O., and Tanner, M. A. (1999). Binomial-beta hierarchical models for ecological inference. Sociological Methods & Research , 28(1):61--90

  42. [50]

    W., Molnar, C., Schlesinger, T., and K \"u chenhoff, H

    Klima, A., Thurner, P. W., Molnar, C., Schlesinger, T., and K \"u chenhoff, H. (2016). Estimation of voter transitions based on ecological inference: An empirical assessment of different approaches. AStA Advances in Statistical Analysis , 100(2):133--159

  43. [51]

    Kuriwaki, S. (2025). Ticket splitting in a nationalized era. The Journal of Politics

  44. [52]

    Kuriwaki, S., Ansolabehere, S., Dagonel, A., and Yamauchi, S. (2024). The geography of racially polarized voting: Calibrating surveys at the district level. American Political Science Review , 118(2):922--939

  45. [53]

    T., and Kellermann, M

    Lau, O., Moore, R. T., and Kellermann, M. (2007). eiPack : R x C ecological inference and higher-dimension data management. New Functions for Multivariate Analysis , 7(1):43

  46. [54]

    Lewis, J. B. (2001). Understanding king's ecological inference model: A method-of-moments approach. Historical Methods: A Journal of Quantitative and Interdisciplinary History , 34(4):170--188

  47. [55]

    Lewis, J. B. (2004). Extending king’s ecological inference model to multiple elections using markov chain monte carlo. Ecological Inference: New Methodological Strategies , page Ch. 4

  48. [56]

    Manski, C. F. (2018). Credible ecological inference for medical decisions with personalized risk assessment. Quantitative Economics , 9(2):541--569

  49. [57]

    McCartan, C. (2025). bases: Basis Expansions for Regression Modeling . R package version 0.1.2

  50. [58]

    and Kuriwaki, S

    McCartan, C. and Kuriwaki, S. (2025a). Identification and semiparametric estimation of conditional means from aggregate data. arXiv preprint arXiv:2509.20194

  51. [59]

    and Kuriwaki, S

    McCartan, C. and Kuriwaki, S. (2025b). seine: Semiparametric ecological inference

  52. [60]

    McCue, K. F. (2001). The statistical foundations of the ei method. The American Statistician , 55(2):106--110

  53. [61]

    Moskowitz, D. J. (2021). Local news, information, and the nationalization of us elections. American Political Science Review , 115(1):114--129

  54. [62]

    Park, W.-h. (2008). Ecological Inference and Aggregate Analysis of Elections. PhD thesis

  55. [63]

    J., and Biggers, D

    Park, W.-h., Hanmer, M. J., and Biggers, D. R. (2014). Ecological inference under unfavorable conditions: Straight and split-ticket voting in diverse settings and small samples. Electoral Studies , 36:192--203

  56. [64]

    Pav \'i a, J. M. and Romero, R. (2024). Improving estimates accuracy of voter transitions. two new algorithms for ecological inference based on linear programming. Sociological Methods & Research , 53(3):1491--1533

  57. [65]

    Pav \'i a, J. M. and Thomsen, S. R. (2024). ecolRxC : Ecological inference estimation of R x C tables using latent structure approaches. Political Science Research and Methods , pages 1--19

  58. [66]

    Phillips, K. P. (2014). The emerging republican majority: updated edition . Princeton University Press

  59. [67]

    and Ginebra, J

    Puig, X. and Ginebra, J. (2015). Ecological inference and spatial variation of individual behavior: National divide and elections in catalonia. Geographical Analysis , 47(3):262--283

  60. [68]

    Quinn, K. (2004). Ecological inference in the presence of temporal dependence. Ecological Inference: New Methodological Strategies , page Ch. 9

  61. [69]

    Rivers, D. (1998). Review of: A solution to the ecological inference problem by gary king. American Political Science Review , 92(2):442--443

  62. [70]

    M., Rotnitzky, A., and Zhao, L

    Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1995). Analysis of semiparametric regression models for repeated outcomes in the presence of missing data. Journal of the american statistical association , 90(429):106--121

  63. [71]

    Robinson, W. (1950). Ecological correlations and the behavior of individuals. American Sociological Review , 15(3):351--357

  64. [72]

    Rodden, J. A. (2019). Why cities lose: The deep roots of the urban-rural political divide . Basic Books

  65. [73]

    Rosen, O., Jiang, W., King, G., and Tanner, M. A. (2001). Bayesian and frequentist inference for ecological inference: The R x C case. Statistica Neerlandica , 55(2):134--156

  66. [74]

    and Viswanathan, N

    Rosenman, E. and Viswanathan, N. (2018). Using poisson binomial GLMs to reveal voter preferences. arXiv preprint arXiv:1802.01053

  67. [75]

    Schoenberger, R. A. and Segal, D. R. (1971). The ecology of dissent: the southern wallace vote in 1968. Midwest Journal of Political Science , 15(3):583--586

  68. [76]

    and Wong, W

    Shen, X. and Wong, W. H. (1994). Convergence rate of sieve estimates. The Annals of Statistics , pages 580--615

  69. [77]

    Teele, D. L. (2024). The political geography of the gender gap. The Journal of Politics , 86(2):428--442

  70. [78]

    Thomsen, S. R. (1987). Danish elections 1920-79. a logit approach to ecological analysis and inference. Politica

  71. [79]

    Wakefield, J. (2004). Ecological inference for 2 x 2 tables (with discussion). Journal of the Royal Statistical Society Series A: Statistics in Society , 167(3):385--445

  72. [80]

    Wright, G. C. (1977). Contextual models of electoral behavior: The southern wallace vote. American Political Science Review , 71(2):497--508

  73. [81]

    and Gardner, J

    Wu, K. and Gardner, J. R. (2024). A fast, robust elliptical slice sampling implementation for linearly truncated multivariate normal distributions. arXiv preprint arXiv:2407.10449

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.