Pith. sign in

REVIEW 3 major objections 5 minor 15 references

Parameter Effects in ReCom Ensembles

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read ReCom ensemble statistics depend on algorithm variant and county-preservation parameters, while population tolerance has no material effect.

desk verdict A substantial multi-state study of ReCom parameter sensitivity with a plausible average-margin consistency claim, held back by placeholder data links and significance testing that ignores effective sample sizes. read the letter →

arxiv 2505.21326 v1 pith:UZ46MJSA submitted 2025-05-27 physics.soc-ph cs.CYstat.AP

classification physics.soc-phcs.CYstat.AP PACS 89.65.-s
keywords ReComredistrictingensemblesMarkovchainMonteCarloeffectivesamplesizecountypreservationpartisanbiascompetitivenessminority-majoritydistricts
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper systematically tests whether the statistics of ReCom redistricting ensembles depend on the algorithm's tuning parameters, using 315 ensembles across seven states and three legislative chambers. It establishes that population tolerance—the allowed deviation from equal population—has a negligible effect on every score examined, so practitioners can set tolerance for computational or legal reasons without changing substantive conclusions. It also establishes that the choice of ReCom variant and the strength of county preservation can shift partisan bias, minority-opportunity, and competitiveness statistics, with some effects inconsistent across jurisdictions and others surprisingly consistent. In particular, the average district margin rises steadily as maps become more compact and county-preserving, in every state and chamber studied. The paper also introduces effective-sample-size and redundancy diagnostics intended to verify that the chains are long enough for reliable statistics.

What carries the argument

The machinery is the ReCom (recombination) Markov chain and its parameterized variants: ReCom-A through D, which differ in how they choose the adjacent district pair and how they draw spanning trees, plus county-aware versions R25–R100 that add a surcharge to edges crossing county lines, and RevReCom, the reversible variant. The load-bearing statistical tools are the AR(1)-based effective sample size approximation $n_{\text{eff}} \approx n \cdot (1-\gamma_1)/(1+\gamma_1)$ and two redundancy measures, $\phi_{\text{avg}}$ and $\phi_{\text{max}}$, which together support the convergence claims. The explanatory identity for the most consistent parameter effect is $AM = (V_D - 0.5) + 2S_R$, decomposing the average margin into the statewide Democratic share plus twice the Republican surplus-vote fraction, which makes the average margin a monotone function of compactness under the paper's assumptions.

What would settle it

Compute the lag-2 and lag-3 autocorrelations of a score along a ReCom chain and compare them with the squares and cubes of the lag-1 correlation; if the measured values are systematically larger, the geometric-decay model fails and the reported effective sample sizes would shrink. A second check would rerun the full 315-ensemble experiment with chains several times longer and see whether any of the significant parameter effects in the paper's main tables change sign or lose significance.

Watch

Extended reading notes

Core claim

The central discovery is that there is no single 'ReCom ensemble' for a state and chamber: the distribution of maps, and hence the statistics used in redistricting litigation, depends on modeling choices. Across seven states and three chambers, the paper finds that altering the population tolerance barely moves any score, whereas switching among ReCom variants (A–D) and increasing the county-aware surcharge from 0.25 to 1.00 frequently produces statistically significant and sometimes material changes in Democratic seat counts, minority-majority districts, and competitive-seat counts. These effects are directionally inconsistent across states and chambers for partisan and minority-opportunity scores, but the average margin score increases consistently with compactness and county preservation. The paper argues this consistency follows from an identity: the average margin equals the surplus-vote fraction, which is a monotone function of compactness given a fixed statewide Democratic vote share. In addition, the paper claims that its autocorrelation-based effective sample sizes and redundancy measurements show the non-reversible chains are well mixed, with RevReCom the notable exception in some jurisdictions.

Load-bearing premise

The chain's scores are assumed to lose correlation with each other at a fixed rate as they get further apart in the chain; if the correlation actually lingers longer, the effective sample sizes are overestimated and the significance levels are too optimistic.

Editorial extensions

If this is right

  • Practitioners using ReCom ensembles in court or policy work should report the algorithm variant and county-preservation surcharge alongside the resulting statistics, because these choices can move partisan and minority-opportunity scores by half a seat or more.
  • Population tolerance can be varied freely within the studied ranges without altering substantive conclusions, so it can be chosen to satisfy legal or computational constraints.
  • Because the average margin score responds consistently and monotonically to compactness and county preservation, comparisons of competitiveness across ensembles should control for these parameters or use the same settings.
  • The convergence diagnostics (effective sample size and redundancy) give a practical template for checking whether a chain is long enough before interpreting ensemble statistics.
  • RevReCom ensembles are not reliably proxied by the cheaper ReCom variants, so users who need the reversible algorithm's target distribution should generate RevReCom chains directly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extending the paper's logic, sensitivity analysis across parameter settings could become a standard robustness check in redistricting litigation, since a single ensemble from one parameter setting may understate the uncertainty in the statistics.
  • The monotone relationship between compactness and average margin might generalize beyond ReCom: any algorithm that increases compactness by reducing long district boundaries would tend to increase surplus votes and thus average margins, a testable claim for other sampling algorithms.
  • The effective-sample-size method could be applied to judge chain length in other MCMC sampling contexts where only a single run is available, with the caveat that the AR(1) assumption should be checked against higher-lag autocorrelations.
  • Future work could test whether the inconsistency of partisan-score effects across states is itself predictable from state political geography, such as the distribution of Democratic strongholds relative to county lines.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a systematic study of how ReCom algorithm parameters affect redistricting ensemble statistics. For seven states and three legislative chambers, the authors generate 315 ensembles of 20,000 plans, varying the population tolerance, the ReCom variant (A, B, C, D, and reversible ReCom), and a county-preservation surcharge (R25 through R100). They compare ensembles using compactness, partisan bias, competitiveness, minority-opportunity, and county-split scores. They introduce two convergence diagnostics: an effective sample size estimate based on an AR(1) autocorrelation model, and district-level redundancy measures. Their main empirical findings are that population tolerance has negligible effects, that county surcharge and algorithm variant affect partisan and minority-opportunity scores in directions that are inconsistent across states and chambers, and that the average-margin score increases consistently with compactness and county preservation, with the ordering D < B < C/RevReCom < A0 < R25 < R50 < R75 < R100. The paper also proposes an explanation for the average-margin pattern via the identity AM = SD + SR.

Significance. If the convergence and significance claims hold, this is a valuable large-scale contribution: it is the first multi-state, multi-chamber systematic parameter sweep for ReCom, and the convergence diagnostics (autocorrelation-based effective sample size and redundancy measures) are useful additions to the redistricting toolbox. The paper shares code and data links, and the tables report results for all 315 ensembles across many scores. The most novel and actionable claim is the surprisingly consistent effect of compactness and county preservation on average margin, which, if confirmed, would matter for practitioners choosing ReCom parameters in litigation and policy analyses. The main uncertainty is statistical: the reported p-values and convergence statements depend on effective sample size estimates and an AR(1) autocorrelation assumption that are not fully validated for the county-preserving and RevReCom ensembles.

major comments (3)
  1. [Section 5.2, Table 3; Section 6.5, Table 8] The paper's own convergence benchmark is neff >= 15,094 (Section 5.2), but several R75/R100 ensembles fall below it: for example, FL 120 R100 has neff = 5,776, MI 110 R75/R100 have 8,259/6,519, NC 120 R75/R100 have 7,411/5,922, NY 150 R75/R100 have 9,706/7,323, OH 99 R75/R100 have 10,249/7,752, and WI 99 R75/R100 have 10,299/6,299. The text asserts that these low values are due to the county-splits score alone, but Table 3 reports only the maximum autocorrelation over all non-MMD scores and does not identify which score produced the minimum neff. Since the central consistency claim concerns the average-margin score specifically (Table 8), the convergence evidence for that score is not documented. Please provide per-score autocorrelations and neff values for average margin (and for the other headline scores in Tables 6-11), or qualify the convergence claim accordingly.
  2. [Section 5.2, Eq. (5)] Equation (5) assumes the lag-k autocorrelation decays geometrically, gamma_i = gamma_1^i. The paper does not present any empirical lag-k autocorrelations or any test of this geometric-decay assumption. ReCom chains with county surcharges can exhibit persistent county-block structures whose autocorrelation plausibly decays more slowly than geometric; if so, the neff values in Table 3 would be overestimates. The consistency checks mentioned in Section 5.2 (multi-start dKS for A0 and redundancy measurements) do not test the decay model. Please add a concrete validation, such as plots or summaries of the empirical lag-k autocorrelation for representative chains, and consider a more robust neff estimator (e.g., initial-monotone-sequence or batch means). The headline convergence claims for R75/R100 and RevReCom currently rest on an unvalidated model.
  3. [Section 4.6, Tables 6-11] The p-values reported in Tables 6-11 appear to treat each ensemble as n = 20,000 independent draws. The paper introduces neff for convergence assessment but does not state that the significance tests in Tables 6-11 use neff-corrected sample sizes or standard errors. For autocorrelated ensembles, including R75/R100 and especially RevReCom (which has neff values as low as 46 for NY 150 and 77 for NC 120 in Table 3), these p-values are anti-conservative. Moreover, Table 5 explicitly omits RevReCom because its effective sample sizes are insufficient for accurate dKS comparisons, yet RevReCom columns are included with bold 'p < .001' entries in Tables 6-11. This is internally inconsistent. Please either recompute the significance tests using neff-adjusted effective sample sizes, or clearly flag comparisons with insufficient effective sample size, and reconcile the treatment of RevReCom.
minor comments (5)
  1. [Tables 5-11] The first column of Tables 5-11 is labeled 'A1' in the headers but the text and captions refer to this ensemble as A0; please make the labels consistent.
  2. [Table 10 caption] The caption contains 'hisanic' instead of 'Hispanic'; please fix the typo.
  3. [Section 6.2 heading] The heading reads 'dramatically effects compactness' but should read 'dramatically affects compactness'; please correct the verb.
  4. [Section 6.5] The sentence 'the formula AM = .05 + 2SD shows that AM would increase monotonically' does not follow from the preceding observation that the Dem-share-related quantity behaves in a complicated, not-necessarily-monotonic way; because AM is affine in SD, the monotonicity claim needs a separate argument or a conditional wording.
  5. [Supplementary Materials] The data and code links currently contain placeholders such as 'https://TBD-url' and 'github.com/proebsting/TBD'; please provide the actual URLs in the final version.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the ReCom parameter-sensitivity results are empirical measurements, not derivations from fitted inputs or self-citations.

full rationale

The paper is an empirical parameter-sensitivity study of ReCom ensembles, not a derivation. The central claims—population tolerance has negligible effects, while algorithm variant and county surcharge affect partisan, minority-opportunity, and competitiveness scores—are supported by comparing ensembles generated under different parameter settings and reporting mean differences, dKS distances, and p-values. No parameter is fitted to a subset of data and then renamed as a prediction. The AM = SD + SR identity in Section 6.5 is an interpretive aid used to rationalize why the average margin correlates with compactness; it does not by construction force the observed ordering, and the paper does not claim that it does. The effective-sample-size and redundancy analyses in Sections 5.2–5.3 are convergence heuristics; the AR(1) autocorrelation assumption of Equation 5 is an approximation whose fragility could weaken the convergence evidence, but that is a correctness or validity concern, not circular reasoning. The only self-citation is [Tap25], used in Section 4.4 to cite a partial explanation of why RMST algorithms produce more compact plans than UST algorithms; this is an aside and is not load-bearing for any of the paper's principal findings. Consequently, no circular step can be exhibited, and the appropriate finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical or conceptual entities. The free parameters are experimental design choices and a heuristic threshold, not fitted constants. The key load-bearing assumptions are the AR(1) autocorrelation model and the mixing of the MCMC chains.

free parameters (4)
  • Population tolerance settings = 0.005/0.01/0.015 (congressional); 0.025/0.05/0.075 (legislative)
    Chosen to bracket the default setting. The paper concludes these have negligible effects, but this conclusion is conditional on this narrow range.
  • County surcharge values = 0.25, 0.50, 0.75, 1.00
    Chosen as a grid for region-awareness strength. The effects of these values are the central object of study, not fitted parameters.
  • Chain length and subsampling interval = 50 million steps / every 2500 for ReCom variants; 1 billion / 50000 for RevReCom
    Chosen for manageability and to offset repetition. Convergence diagnostics are used to justify sufficiency, but these settings are hand-selected.
  • dKS threshold for visible difference = 0.1
    Subjective threshold equated with the 'eyeball test' in Section 4.6; used to interpret the magnitude of ensemble differences.
assumptions (5)
  • ad hoc to paper The AR(1) model approximates the autocorrelation of score sequences in Equation 5.
    Section 5.2 assumes geometric decay of autocorrelation to estimate effective sample sizes. This is a modeling assumption not validated against the full autocorrelation spectrum.
  • domain assumption ReCom chains mix to the stationary distribution within the chosen chain lengths.
    The entire ensemble analysis presumes mixing. The paper provides multi-start and autocorrelation heuristics but cannot prove convergence for all variants.
  • domain assumption Input data from Dave's Redistricting (precincts, population, VAP/CVAP, election returns) are accurate.
    All scores depend on the correctness of DRA's data, including the blended election index. Cited in Section 4.2.
  • domain assumption The blended election index is an appropriate summary of partisan voting behavior.
    Partisan scores use a weighted blend of presidential, senate, governor, and attorney general races; different blends could change the results.
  • domain assumption The chosen scores (Reock, Polsby-Popper, efficiency gap, etc.) are valid measures for their purposes.
    The paper uses standard metrics from the literature but does not independently justify their validity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Parameter Effects in ReCom Ensembles." pith.science (2026). https://pith.science/paper/UZ46MJSA

@misc{pith2026250521326,
  author       = {Pith},
  title        = {Pith review of: Parameter Effects in ReCom Ensembles},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UZ46MJSA}},
  note         = {Machine review of arXiv:2505.21326}
}
read the original abstract

Ensemble analysis has become central to redistricting litigation, but parameter effects remain understudied. We analyze 315 ReCom ensembles across the three legislative chambers in 7 states, systematically varying the population tolerance, county preservation strength, and algorithm variant. To validate convergence, we introduce new methods to approximate effective sample size and measure redundancy. We find that varying the population tolerance has a negligible effect on all scores, whereas the algorithm and county-preservation parameters can significantly affect some metrics, inconsistently in some cases but surprisingly consistently in others across jurisdictions. These findings suggest parameter choices should be thoughtfully considered when using ReCom ensembles.

Figures

Figures reproduced from arXiv: 2505.21326 by the authors.

Figure 1
Figure 1. kde plots for multiple ReCom variants (to be described later) with [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Ordered seats plots for two ensembles of FL congressional plans. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. The figure is specific to the lower legislative chamber of Florida, [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3: County splits vs cut edges for FL lower (with [PITH_FULL_IMAGE:figures/full_fig_p014_3.png]
Figure 4
Figure 4. Figure 4: Box and whisker plots for WI congress (with [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Score orders: The position of an ensemble represents how many standard deviations is lies from A0 with respect to that score. Each cyan box shows the range [−.033, .033]; positions outside of the box are significantly dif￾ferent from A0 at p-value ≤ 0.001. 20 [PITH_FU…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 12 canonical work pages

  1. [1]

    Herschlag, Zach Hunter, and Jonathan C

    Eric Autry, Daniel Carter, Gregory J. Herschlag, Zach Hunter, and Jonathan C. Mattingly. Metropolized forest recombination for monte carlo sampling of graph partitions. SIAM Journal on Applied Mathematics , 83(4):1366–1391, August 2023

  2. [2]

    Colorado in Context: Congressional Redistricting and Competing Fairness Criteria in Colorado

    Jeanne Clelland, Haley Colgate, Daryl DeFord, Beth Malmskog, and Flavia Sancier-Barbosa. Colorado in context: Congressional redistricting and competing fairness criteria in colorado, March 2021. arXiv:2011.06049

  3. [3]

    Repetition effects in a Sequential Monte Carlo sampler

    Sarah Cannon, Daryl DeFord, and Moon Duchin. Repetition effects in a sequential monte carlo sampler, September 2024. arXiv:2409.19017

  4. [4]

    Spanning tree methods for sampling graph partitions, October 2022

    Sarah Cannon, Moon Duchin, Dana Randall, and Parker Rule. Spanning tree methods for sampling graph partitions, October 2022. arXiv:2210.01401

  5. [5]

    Recombination: A family of markov chains for redistricting

    Daryl DeFord, Moon Duchin, and Justin Solomon. Recombination: A family of markov chains for redistricting. Harvard Data Science Review , 3(1), December 2020

  6. [6]

    Discrete geometry for electoral geography, August 2023

    Moon Duchin and Bridget Eileen Tenner. Discrete geometry for electoral geography, August 2023. arXiv:1808.05860

  7. [7]

    Benjamin Fifield, Kosuke Imai, Jun Kawahara, and Christopher T. Kenny. The essential role of empirical validation in legislative redistricting simulation. Statistics and Public Policy , 7(1):52–68, January 2020

  8. [8]

    Carlin, Hal S

    Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, and Donald B. Rubin. Bayesian Data Analysis . CRC Press, November 2013. Google-Books-ID: eSHSBQAAQBAJ

Show all 15 references
  1. [9]

    Charles J. Geyer. Practical markov chain monte carlo. Statistical Science , 7(4):473–483, November 1992

  2. [10]

    Charles J. Geyer. Introduction to markov chain monte carlo. In Handbook of Markov Chain Monte Carlo , pages 3--48. Chapman and Hall/CRC, 1st edition, 2011

  3. [11]

    Katz, Gary King, and Elizabeth Rosenblatt

    Jonathan N. Katz, Gary King, and Elizabeth Rosenblatt. Theoretical foundations and empirical evaluations of partisan fairness in district-based democracies. American Political Science Review , 114:1: 164--178, 2020

  4. [12]

    Sequential monte carlo for sampling balanced and compact redistricting plans, February 2023

    Cory McCartan and Kosuke Imai. Sequential monte carlo for sampling balanced and compact redistricting plans, February 2023. arXiv:2008.06131

  5. [13]

    Yossi Rubner, Carlo Tomasi, and Leonidas J. Guibas. The earth mover’s distance as a metric for image retrieval. International Journal of Computer Vision , 40(2):99–121, November 2000

  6. [14]

    Partisan gerrymandering and the efficiency gap

    Nicholas Stephanopoulos and Eric McGhee. Partisan gerrymandering and the efficiency gap. University of Chicago Law Review , 82(2), March 2015

  7. [15]

    On the minimum spanning tree distribution in grids, January 2025

    Kristopher Tapp. On the minimum spanning tree distribution in grids, January 2025. arXiv:2401.17947

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.