{"id":"c9e8a1b1-22f1-48fc-9fb5-45fd91ab79ad","arxiv_id":"2509.07328","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An ensemble analysis of redistricting in New Hampshire finds the enacted State Senate and Executive Council maps are Republican-leaning outliers relative to typical nonpartisan plans.","lead":"This paper generates thousands of plausible New Hampshire legislative maps and compares the enacted Senate and Executive Council plans against them. It finds the enacted maps consistently lean more Republican than typical maps in close elections, and shows that the choice of election data can change fairness conclusions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Outlier claim depends on an unvalidated 4:1 town-weight prior: the baseline is anchored to the enacted plan's split-town count, so tail status may be a modeling artifact rather than evidence of partisan intent.","rationale":"The reader's weakest assumption and my concern coincide: the ensemble's status as a neutral baseline rests on the town-weighting ratio and population tolerances chosen in §2.2. The paper does not justify these choices from statutory criteria and performs no sensitivity analysis. This matters because the headline conclusion is purely comparative: the enacted plan is an outlier relative to this particular ensemble. If a different but equally defensible baseline moves the plan into the bulk of the distribution, the central claim collapses. The concern is concrete and testable: rerun the ensembles under a small grid of weight ratios and population tolerances and recompute tail ranks. Unlike missing tail probabilities or absent code, this is a direct test of the inference's foundation. I do not think the paper should be rejected; the directional pattern across many metrics and elections is credible, and the authors already provide mixing diagnostics and reproducibility-oriented details. But the current manuscript does not establish that the outlier status is robust to the prior, so the conditional verdict is appropriate. The proposed sensitivity check would either strengthen the conclusion to near-acceptance or reveal that the finding is a modeling artifact.","tokens_in":11684,"tokens_out":8839,"duration_ms":124594,"concrete_test":"Re-run the Executive Council and State Senate ensembles under at least four baselines: (i) current 4:1 town weight with 5%/1% population deviations; (ii) 1:1 town weight; (iii) 10:1 town weight; (iv) a hard no-town-split constraint where feasible; also vary population deviations to 2% and 10%. For each baseline, compute the empirical tail probability/rank of the enacted plan for mean-median score, partisan bias, and efficiency gap in PRES16, GOV16, and PRES20, and for Republican vote share in the Hanover/Concord/Keene districts. If the enacted plan remains outside the 5–95% band in all settings, the §4 conclusion is robust; if its rank moves into the bulk under any plausible setting, the ensemble is not a neutral baseline and the partisan-intent claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference—that the enacted plans are partisan outliers—requires the RECOM ensemble to be a neutral baseline of legally plausible maps. That baseline is not validated. In §2.2 the paper states that New Hampshire law requires districts to be made of contiguous towns/cities/wards without splitting them, yet the ensemble only discourages splits via a soft 4:1 edge-weight ratio. The upper bound of 4 is chosen to 'balance mixing time' and 'reduction of split towns,' and Figure 2 uses the enacted plan's three split towns as the reference point. Thus the baseline is partially anchored to the very plan it is meant to judge. The 5% Senate / 1% Executive Council population tolerances are similarly arbitrary. The paper even concedes in §4 that the RECOM target distribution is unknown. If a hard no-split constraint, or a different town-weight ratio, shifts the ensemble's partisan distribution so that the enacted plan is no longer in the tail, then the 'outlier' finding is an artifact of the prior, not evidence of intent. This is the load-bearing weak point: the conclusion in §4 follows only if the baseline is both legally faithful and insensitive to reasonable modeling choices.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the RECOM ensemble method to the New Hampshire State Senate (24 districts) and Executive Council (5 districts). Districts are built from precinct dual graphs with soft town weighting and population-deviation tolerances of 5% (Senate) and 1% (Executive Council). The enacted maps are compared to the ensemble across eight statewide elections (2012–2020) using compactness measures, seat counts, sorted district vote shares, mean-median score, efficiency gap, and partisan bias. The authors report that the enacted plans are consistently more Republican-leaning than typical nonpartisan plans, especially in close elections, and that district-level patterns suggest packing of Democratic voters in Executive Council District 2. They conclude that the analysis supports claims made in Brown v. Scanlan that the maps may have been drawn with partisan interests, while noting that the exact target distribution of the ensemble is unknown and that town-weighting was chosen to balance mixing and split-town reduction.","tokens_in":11968,"tokens_out":5512,"duration_ms":66531,"significance":"If the outlier findings are robust, the paper provides a useful state-specific evidentiary analysis for ongoing litigation and a clear demonstration that election-data choice affects partisan-symmetry conclusions. The study uses standard, well-established methods; reports mixing diagnostics; examines multiple metrics and elections; and includes the enacted plan as a seed. These are real strengths. The main value is empirical rather than methodological: the paper does not introduce new theory, but it operationalizes redistricting rules for New Hampshire and offers a template for similar small-state analyses. Its broader significance hinges on whether the enacted plans are indeed outliers under a defensible neutral baseline—a point that the current manuscript supports only qualitatively.","major_comments":[{"comment":"The ensemble is presented as a neutral baseline, but the legal requirement that districts be composed of contiguous towns/cities/wards without splitting is implemented only as a soft 4:1 edge-weight prior. §2.2 says the upper bound was 'chosen to balance mixing time ... and reduction of split towns', and Fig. 2 anchors the plot to the enacted plan's three split towns. Since §4 admits the target distribution is unknown, the outlier conclusion is only meaningful if tail status is insensitive to reasonable modeling choices. Please report sensitivity analyses: hard no-split constraints, town-weight ratios such as 2:1 and 8:1, and population deviations of 0.5%, 2%, and 10%, with tail probabilities for each.","section":"§2.2, §4"},{"comment":"Central claims that the enacted plan is 'often' an outlier or 'outside the typical range' are supported only by visual inspection of Figs. 4–7 and 10–12. No empirical quantiles, tail probabilities, or p-values are reported for any metric or election. For the Executive Council, distributions are discrete with five seats, so exact tail probabilities are straightforward; for the State Senate, report the fraction of ensemble plans more extreme than the enacted plan for each metric and election. State how multiple comparisons across eight elections and several metrics are handled.","section":"§3.1, §3.2"},{"comment":"The 'firewall' claim—that the enacted plan favors Democrats in landslides but Republicans in close elections—is inferred from seat-count plots (Figs. 4a and 10a) without any statistical test. With only 5 and 24 seats, seat counts are coarse and a single district can move the result. Quantify the enacted plan's seat count as a tail event for each election, and test whether the interaction between election margin and enacted-vs-ensemble difference is significant (e.g., logistic regression or permutation test).","section":"§3.1.2, §3.2.2"},{"comment":"The mixing diagnostics in Tables A1–A4 are computed for sorted district Republican vote totals, but the paper's conclusions rely on nonlinear summary statistics (mean-median, efficiency gap, partisan bias, cut edges). Reporting KS distances for vote totals does not by itself establish that the chains have mixed for these metrics. Please also compute pairwise KS distances and autocorrelation lags for the actual metrics used in §3, or justify why vote-total mixing suffices for all reported comparisons.","section":"Appendix A"}],"minor_comments":[{"comment":"Typos: 'have been been' is repeated in the Introduction; 'preformed' should be 'performed' in §2; Fig. 9 caption spells 'Polsby-Popper' as 'Polsby-Poppy'; ref. [33] spells 'Hillsborough' as 'Hilsborough'.","section":"§1, §2"},{"comment":"The manuscript provides no data or code availability statement. Since exact preprocessing of MGGG shapefiles and the GerryChain version are described, a repository link or appendix with replication code would strengthen reproducibility.","section":"General"},{"comment":"The sentence 'The candidate up for election appears to have a greater effect than in other states' is vague; specify the comparison set or remove it.","section":"§3.1.2"},{"comment":"The conclusion uses 'may have been drawn with partisan interests,' which is appropriately hedged, but then states 'Our result validates some of the claims made in Brown v. Scanlan.' 'Validates' overstates what an ensemble-based outlier analysis can establish; rephrase to 'is consistent with' or 'supports.'","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"This is a competent applied ensemble study, but the central claim—that the enacted maps are partisan outliers—currently rests on an unvalidated baseline and on visual inspection of figures. The requested sensitivity analyses and quantitative tail probabilities are feasible within the scope of the manuscript and would materially strengthen the paper. The empirical contribution is valuable for the state, but novelty is modest and the fit to a general CS/SI audience would be improved by a more rigorous robustness section."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent ensemble case study that does something new for New Hampshire and something useful for the literature on election-data sensitivity. The central claim—enacted maps lean Republican relative to the ensemble, especially in close elections—is credible, but the paper doesn't quantify how far into the tail the enacted plans sit, and the town-weight prior is a free parameter that needs robustness work.\n\nThe genuinely new pieces: first ensemble analysis of the NH Executive Council and State Senate; and a systematic look at eight statewide elections in a state where presidential and gubernatorial results diverge sharply. That second piece is the real contribution. It shows that partisan symmetry conclusions flip depending on which election you use, which matters for anyone doing this kind of analysis anywhere. The authors also tie their findings to Brown v. Scanlan, and their district-level vote-total plots (Hanover, Concord, Dover, Portsmouth) give a concrete picture of the packing argument. Mixing diagnostics are in the appendix, which is more than many papers do.\n\nThe soft spots are addressable but real. Outlier claims are supported by visual inspection only. There are no quantiles or tail probabilities for any of the metrics, so \"often outliers\" is doing too much work. The firewall pattern in close elections looks post hoc, though it is consistent in direction across two chambers and several metrics. The bigger issue is the town-weight prior. The 4:1 ratio for same-town vs cross-town edges is chosen to balance mixing and split-town reduction, but there's no sensitivity analysis. The stress-test worry about the baseline being anchored to the enacted plan is overstated—the paper doesn't tune to match the enacted plan's three split towns—but the choice still matters. If a different ratio, or a hard no-split constraint, moves the ensemble's partisan distribution, the tail claim changes. The paper itself admits the target distribution is unknown, so a robustness section would go a long way.\n\nAlso, no code or data release, which is a minus for reproducibility, though the underlying data are public.\n\nWho this is for: people working on ensemble redistricting, state court cases, and anyone thinking about which elections to use in these analyses. It deserves peer review. A referee should ask for tail probabilities, robustness to the town-weight parameter, and ideally artifacts. With those, this becomes a strong case study.\n\nYes, I'd take it if I were editing.","headline":"Solid, useful case study; main claim credible but under-quantified, and the town-weight prior needs robustness checks.","tokens_in":12399,"tokens_out":2074,"would_cite":true,"duration_ms":24016,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"New Hampshire's enacted State Senate and Executive Council maps are Republican-leaning outliers in a neutral ensemble of plans.","keywords":["Computational Redistricting","Markov Chain Monte Carlo","Ensemble Analysis","Spanning Trees","New Hampshire","partisan symmetry","gerrymandering","RECOM"],"falsifier":"Generate new ensembles under different but still legally compliant parameter choices—for example, town-edge weights drawn from [0,2] instead of [0,4], or a Senate population tolerance of 2% instead of 5%—and see whether the enacted plan's mean-median and partisan-bias values remain outside the ensemble's central range. The paper's claim predicts they do; a neutral-baseline explanation predicts they would fall inside.","tokens_in":11604,"feed_emoji":"🗳️","tokens_out":8111,"duration_ms":84103,"temperature":0.7,"pith_summary":"The paper tries to establish that the State Senate and Executive Council maps New Hampshire enacted after the 2020 census are unlikely to be neutral products of the state's political geography. By generating millions of legally plausible districting plans with a recombination Markov chain and comparing the enacted maps against this ensemble, the authors find the enacted plans sit consistently at the Republican-favoring tail of partisan symmetry measures, even though New Hampshire's geography by itself tends to give a slight Democratic edge. The pattern is strongest in close elections, where the enacted maps would turn narrow Republican wins into seat majorities while narrow Democratic wins would not. The paper presents this as quantitative context for pending state litigation claiming Democratic-leaning towns were packed together. A careful reader should care because the result shows that New Hampshire's enacted maps are measurable outliers on standard fairness metrics across eight different statewide elections.","feed_headline":"New Hampshire's enacted maps are Republican-leaning outliers","feed_subtitle":"A baseline of millions of legally plausible plans shows the tilt is real, not geography.","key_machinery":"The recombination (RECOM) Markov chain algorithm is the central instrument: it merges two districts, builds a weighted spanning tree, and cuts an edge to produce a new pair of districts, with edge weights favoring same-town connections so that towns are rarely split, matching New Hampshire's legal requirements. From this chain the authors draw an ensemble of millions of plans that defines the 'typical' range of compactness and partisan behavior for the state. The comparison metric that carries the argument is the position of the enacted plan's value relative to the ensemble distribution on vote-share, seats-won, mean-median, efficiency-gap, and partisan-bias measures, computed separately for","core_discovery":"The central discovery is that relative to an ensemble of districts sampled under New Hampshire's own rules—towns kept whole, population deviations capped at 5% for the Senate and 1% for the Executive Council—the enacted plans are Republican-leaning outliers on mean-median, efficiency gap, and partisan bias. The ensemble's typical plan carries a slight Democratic advantage from the state's political geography, so the Republican edge cannot be explained by geography alone. The specific mechanism shows in the Executive Council district containing Hanover, Keene, and Concord, which packs Democratic voters and pushes other districts' Republican shares above 50%, and in the State Senate's less-Dem","pith_inferences":["The paper's outlier verdict is conditional on its ensemble being a legitimate neutral baseline; a skeptic could re-run the same analysis with different but still legal parameter choices (town-edge weights, population tolerances, or a metropolized sampler with an explicit target distribution) and check whether the enacted plans remain outliers.","A natural next test is to apply the same ensemble construction to New Hampshire's multi-member and floterial House districts, which the paper explicitly leaves for future work, to see whether the Republican advantage appears there as well.","The split-ticket volatility in New Hampshire makes it a useful test bed for the general question of which election data courts should use when evaluating fairness; the paper implies that single-election analyses can be misleading, and a cross-state comparison could sharpen that lesson.","Because the paper only evaluates the enacted plan against its own ensemble, a direct out-of-sample check—comparing the enacted maps' performance in the 2022 or 2024 elections with ensemble predictions—would test whether the observed bias is stable over time."],"forward_implications":["If the analysis is right, the enacted Senate and Executive Council maps are not ordinary reflections of New Hampshire's geography; they are statistical outliers with a consistent Republican tilt.","In close statewide elections, the enacted maps would give Republicans more seats than typical neutral plans, while in landslide elections the Democratic advantage in packed districts shows up instead—a pattern the paper describes as a Republican firewall.","The ensemble provides a neutral benchmark that courts and map-drawers could use to pre-check a plan for partisanship before adoption, not just after a challenge.","The specific packing of the Executive Council district containing Hanover, Keene, and Concord, and the weakened Democratic middle districts in the Senate, are the mechanisms that produce the outlier values.","The result holds across eight different elections, so it does not depend on which one statewide race a critic chooses to measure partisanship."],"supporting_citations":[{"why":"Defines the recombination (RECOM) Markov chain algorithm that generates the ensemble; the paper's core method.","marker":"[34]"},{"why":"Introduces the county/town-weighted variant of RECOM and the Kolmogorov-Smirnov multi-start mixing heuristic used to tune chain length.","marker":"[26]"},{"why":"Establishes the ensemble-outlier framework and the North Carolina precedent for quantifying gerrymandering with cut edges and partisan measures.","marker":"[8]"},{"why":"The state court challenge whose packing claims the analysis is built to evaluate; supplies the specific districts and allegations.","marker":"[33]"},{"why":"Motivates examining many elections and checking for 'firewall' patterns that hide asymmetric bias in close races.","marker":"[38]"},{"why":"Defines the efficiency gap, one of the three partisan symmetry measures used to compare enacted and ensemble plans.","marker":"[12]"}],"fun_headline_variants":["NH's GOP tilt in maps isn't geography—ensemble proves it","NH redistricting: under state rules, GOP edge is real","Packing Dems in NH districts tilts maps GOP—math shows","NH maps: Republican outliers despite Democratic geography"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the ensemble generated with the chosen parameters is a fair and neutral baseline for legally compliant New Hampshire maps; if the town-weighting, population tolerances, or election set are not representative, then showing the enacted plan is an outlier does not by itself show partisan intent.","fun_headline_variants_meta":{"raw":{"variants":["NH's GOP tilt in maps isn't geography—ensemble proves it","NH redistricting: under state rules, GOP edge is real","Packing Dems in NH districts tilts maps GOP—math shows","NH maps: Republican outliers despite Democratic geography"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1254,"prompt_tokens":617,"completion_tokens":637,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":361,"completion_tokens_details":{"reasoning_tokens":567}},"tokens_in":361,"tokens_out":637,"duration_ms":9129,"temperature":1.0,"reasoning_tokens":567,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:22:25.561412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate new ensembles under different but still legally compliant parameter choices—for example, town-edge weights drawn from [0,2] instead of [0,4], or a Senate population tolerance of 2% instead of 5%—and see whether the enacted plan's mean-median and partisan-bias values remain outside the ensemble's central range. The paper's claim predicts they do; a neutral-baseline explanation predicts they would fall inside.","supporting_citations":[{"cited_title":"Harvard Data Science Review3(1) (2021) https://hdsr.mitpress.mit.edu/pub/1ds8ptxu","cited_arxiv_id":null,"evidence_quote":"Defines the recombination (RECOM) Markov chain algorithm that generates the ensemble; the paper's core method."},{"cited_title":"Journal of Computational Social Science5, 180–226 (2021) https://doi.org/10.1007/s42001-021-00119-7","cited_arxiv_id":null,"evidence_quote":"Introduces the county/town-weighted variant of RECOM and the Kolmogorov-Smirnov multi-start mixing heuristic used to tune chain length."},{"cited_title":"Statistics and Public Policy7(1), 30–38 (2020) https://doi.org/10.1080/2330443X.2020.1796400","cited_arxiv_id":null,"evidence_quote":"Establishes the ensemble-outlier framework and the North Carolina precedent for quantifying gerrymandering with cut edges and partisan measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The state court challenge whose packing claims the analysis is built to evaluate; supplies the specific districts and allegations."},{"cited_title":"Accessed 10-31-2023 (2018)","cited_arxiv_id":null,"evidence_quote":"Motivates examining many elections and checking for 'firewall' patterns that hide asymmetric bias in close races."},{"cited_title":"University of Chicago Law Review82(2), 831–900 (2014) 17","cited_arxiv_id":null,"evidence_quote":"Defines the efficiency gap, one of the three partisan symmetry measures used to compare enacted and ensemble plans."}],"review_version":1}