{"id":"d3d88ea1-1eab-488f-bfc3-844a2d30bb55","arxiv_id":"2507.16722","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using a natural experiment in North Carolina judicial races, the paper estimates that candidate-order flips relative to the presidential race cause 11.8% of Democratic and 15.4% of Republican partisan voters to cast votes contrary to their intent.","lead":"A statistical study of North Carolina judicial ballots finds that 12 to 15 percent of partisan voters may vote for the wrong candidate when a race omits party labels and the candidate order is flipped relative to the presidential race. The finding suggests that mixing labeled and unlabeled races on one ballot can systematically distort election outcomes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 11.8%/15.4% mistake rates are identified by extrapolating a global polynomial flip-effect curve to x=1, a boundary with no direct empirical support, and interpreting |f̂(1)| as a structural mistake rate without a formal derivation.","rationale":"I agree with the reader that the extrapolation to x=1 is the most load-bearing assumption. The central claim—that mixed-label ballots cause 11.8% and 15.4% of partisan voters to cast incorrect votes—is a direct transformation of f̂(1), and the paper provides no structural derivation or sensitivity analysis for that boundary. The placebo and heterogeneity tests are supportive but do not validate the magnitude claim. The concern is not fatal: the paper could be strengthened by a robustness check using registration-based conditioning or local estimation, and the verdict CONDITIONAL is appropriate. I find no additional internal inconsistency beyond the abstract's 12.0% vs 11.8% discrepancy, which is minor.","tokens_in":18736,"tokens_out":8264,"duration_ms":85647,"concrete_test":"Re-estimate Eq. (1) using the precinct-level share of registered Democrats (available in the same NC voter dataset, Section 2.2) as the conditioning variable X, and evaluate |f̂(0)| and |f̂(1)| at 0% and 100% registration. Compare these to |f̂(1)| from the presidential-share model; a material difference indicates that the x=1 interpretation conflates presidential choice with party composition. As a second check, fit a local linear flip effect using only precincts with X>0.9 and compare the boundary estimate to the cubic extrapolation; divergence beyond the reported SE indicates polynomial misspecification at the boundary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline estimates are |f̂_d(1)|=11.8% and |f̂_r(1)|=15.4% (Section 4.1), where f̂ is the cubic fit from Eq. (1)-(2) estimated on 13 contests. The mapping from a conditional treatment effect at x=1 to 'share of partisan-voting mistakes' is asserted in Section 3.1: 'In these precincts, there are no voters from the opposing party.' This requires that presidential vote share is a perfect proxy for partisan composition, that no crossover or independent voting occurs at the boundary, and that the only behavioral response to a flip is the party-cue mistake. None of these conditions is derived or tested. The data contain precinct-level party registration (Section 2.2), which could directly measure partisan composition, but the analysis conditions on presidential vote share instead. Moreover, f̂(1) is an extrapolation: the polynomial is global, so the boundary value is a linear combination of all coefficients; support near 0 or 1 may be thin, and the placebo test with party labels (Section 4.2) does not validate the boundary structural assumption. Specification sensitivity in Table 2 (linear R: 11.0%, cubic R: 15.4%, SE 3.9%) shows the boundary estimate is not stable across plausible polynomials. The abstract also reports 12.0% for Democrats while the full text reports 11.8%, a minor inconsistency that does not affect the main concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes North Carolina statewide judicial races without party labels (2004–2016) to estimate the effect of a party-order \"flip\" relative to the presidential race on judicial vote shares. Using precinct-level data from 13 nonpartisan contests (6 flipped, 7 not) and a double machine learning estimator with a polynomial flip-effect function, the authors report a heterogeneous conditional flip effect and interpret the boundary value |f(1)| as the share of partisan voters who cast votes contrary to their intent, yielding 11.8% for Democratic and 15.4% for Republican voters. A placebo test using 13 judicial races with party labels finds no flip effect, and robustness checks compare cubic, linear, and constant specifications.","tokens_in":19061,"tokens_out":4361,"duration_ms":50017,"significance":"If the headline estimates were structurally identified, the paper would make a consequential contribution to the study of ballot design: it would show that mixing labeled and unlabeled contests can systematically distort down-ballot outcomes, with 10–15% of partisan voters potentially voting against their intent. The manuscript has real strengths: the double machine learning procedure is described in algorithmic detail, inference uses cluster-robust standard errors and uniform confidence bands, the placebo design is appropriate for testing the party-label mechanism, and the data are from public sources. However, the central interpretation of |f(1)| as a mistake rate is not derived from the identifying assumptions, and the boundary estimate is extrapolated from a global polynomial estimated on only 13 contests. Those issues are load-bearing for the main claim, so the paper needs substantive revision rather than minor polishing.","major_comments":[{"comment":"The identification of the share of partisan-voting mistakes as |f(1)| is asserted, not derived. The flip effect f(x) is a conditional average treatment effect on judicial vote share given presidential vote share; converting f(1) into a mistake rate requires that, in a precinct where one party receives 100% of the presidential vote, all voters are partisans of that party, that they all vote in the judicial race (no roll-off), that no independent or cross-party voters exist, and that the only behavioral response to a flip is the party-cue mistake. None of these conditions is stated as an assumption or tested. The dataset includes precinct-level party registration (Section 2.2), which could be used to measure partisan composition directly or to validate the presidential-vote-share proxy; without such validation, the headline numbers should be presented as boundary extrapolations of the flip effect, not as structural mistake rates.","section":"Section 3.1, Eq. (1)–(2)"},{"comment":"The headline estimate is an extrapolation of a global polynomial to x=1, and the estimate is not stable across plausible polynomial degrees. For Republicans, the linear specification gives 11.0% while the cubic gives 15.4%; for Democrats, the linear gives 12.4% and the cubic 11.8%. The paper does not report the empirical support of the presidential vote share near 0 or 1, the number of precincts in that region, or local sensitivity analyses at the boundary. Because |f^(1)| is a linear combination of all polynomial coefficients, the paper should report support diagnostics and assess robustness using local estimates, restricted samples, or registration-based measures of partisan composition.","section":"Section 3.2 and Section 4.1/Table 2"},{"comment":"Identification relies on contest-level treatment assignment with only 13 clusters, 6 of which are flipped. Cluster-robust standard errors with 13 clusters can understate uncertainty, and the paper reports no balance tests across flipped and non-flipped contests for contest-level characteristics such as incumbency, candidate gender, year, or race composition, even though these covariates are described as available in Section 2.2. The paper should present covariate balance and consider small-cluster corrections or randomization inference for the contest-level assignment, especially because Assumption 2 (known random assignment) is justified heuristically rather than verified for the realized set of contests.","section":"Section 2.2, Section 3.1, Table 1"}],"minor_comments":[{"comment":"The abstract reports 12.0% (SE 3.6%) of Democratic voters casting incorrect votes, while the main text and Table 1 report 11.8% (SE 4.0%). These numbers should be reconciled.","section":"Abstract and Section 4.1"},{"comment":"There are typos in the text: \"stablish\" in Section 3 and \"hypotesis\" in Section 3.2 should be corrected.","section":"Section 3, Section 3.2"},{"comment":"The claim that the mistake shares under the linear and cubic specifications are \"not statistically different from each other\" is stated without reporting the test statistic or p-value; please provide the formal comparison.","section":"Section 4.3"},{"comment":"The conclusion describes the design as a \"state-level treatment assignment,\" whereas Section 3.1 defines treatment at the contest level; this wording should be aligned for precision.","section":"Section 5 and Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The empirical design is interesting and the placebo test is a genuine strength, but the headline interpretation of |f(1)| as a mistake rate requires either a formal derivation with explicit assumptions or a substantial reframing of the estimates as boundary treatment effects. The small number of clusters and the absence of balance tests also need to be addressed. I see no sign of misconduct; the abstract discrepancy appears to be a reporting error that should be corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The flip-effect design is genuinely new, and the placebo test is the strongest part of the paper. Using party-order flips relative to the presidential race in nonpartisan judicial contests, and then decomposing the net effect into offsetting mistake rates via the x=1 boundary, is a real angle I haven't seen in the ballot-order literature. The institutional detail on NC election law is careful, and the DML implementation with cluster-robust standard errors and uniform confidence bands is professionally done. The placebo test — showing the flip effect disappears when party labels are present — is exactly the right check, and it passes cleanly. That is real evidence for the mechanism.\n\nThe soft spots are where the stress-test note points, and I largely agree. The 11.8% and 15.4% numbers are |f(1)|, i.e., the fitted cubic evaluated at the boundary. That is a transformation of the same coefficients, not an independent measurement. The mapping from f(1) to 'share of partisan-voting mistakes' requires that presidential vote share is a perfect proxy for partisan composition and that only the party-cue mistake operates at the boundary. Those assumptions are asserted, not derived. The data contain precinct-level party registration, which could directly speak to composition, but the analysis conditions on presidential vote share instead. The specification sensitivity is real: the Republican estimate moves from 11.0% (linear) to 15.4% (cubic), which is a wide range for a headline number.\n\nThe cluster issue is actually worse than the stress-test note says. With 13 contests and only 6 flipped, cluster-robust standard errors are known to understate variance; the point estimates are essentially identified off a handful of contests. No balance tests are reported for contest-level covariates, so unobserved contest-level confounders (candidate quality, campaign spending) cannot be ruled out. These are not 'fatal' objections — the heterogeneity pattern across x is as predicted, and the placebo is strong — but they make the specific magnitudes fragile.\n\nMinor: the abstract reports 12.0% (SE 3.6%) for Democrats while the text reports 11.8% (SE 4.0%). That should be fixed.\n\nWho is this for? Political scientists and election-administration researchers will get value from the design and the placebo. Operations researchers interested in causal inference with clustered treatments will find the DML application instructive, though the small-cluster inference is a cautionary example. It deserves a serious referee. A referee should ask for balance tables, a sensitivity analysis across polynomial degrees and estimation methods, a derivation or at least a formal statement of the boundary assumptions, and ideally an analysis using party registration to validate the x=1 interpretation. I'd engage with it after those revisions.","headline":"A clever natural experiment on ballot order with a strong placebo, but the headline mistake rates are boundary extrapolations of a fitted polynomial on 13 contests and should be read as model-dependent estimates, not measured facts.","tokens_in":19586,"tokens_out":2410,"would_cite":false,"duration_ms":29854,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Party-order flips misled about 12–15% of NC judicial voters","keywords":["ballot design","causal inference","double machine learning","conditional average treatment effect","party-order flip","partisan and nonpartisan elections","candidate order","voter intent"],"falsifier":"Re-estimate $|f(1)|$ using only precincts with presidential vote share above 0.9, or with candidate-name-order fixed effects added; if the fitted curve flattens or reverses at the boundary, the mistake-rate estimates are artifacts of extrapolation rather than measured voter behavior.","tokens_in":18503,"feed_emoji":"🗳️","tokens_out":6784,"duration_ms":64790,"temperature":0.7,"pith_summary":"The paper tries to establish that when a ballot mixes races with and without party labels, partisan voters use the party order in the labeled races as a cue in the unlabeled races, and that when the two orders differ they cast votes against their own intent. Using 13 North Carolina statewide judicial races without party labels from 2004–2016, it estimates that 11.8% of Democratic and 15.4% of Republican partisan voters cast incorrect judicial votes because of such a party-order flip. The estimated flip effect is heterogeneous: it helps a party's judicial candidate in precincts where that party's presidential candidate is weak and hurts them where the party is strong, so the net average effect is near zero even though mistakes are large. A placebo test on races with party labels shows the flip effect vanishes when affiliation is visible. If right, mixed-label ballots systematically misrepresent voter intent in down-ballot races.","feed_headline":"Order flips on NC ballots misled 12–15% of partisan voters","feed_subtitle":"Analyzing 13 unlabeled races shows order mismatch, not voter preference, drove judicial vote shifts.","key_machinery":"The carrying object is the flip effect $f(x)$, the conditional average treatment effect of a party-order flip on a party's judicial vote share given the presidential vote share $x$ in the same precinct. It is modeled in a partially linear outcome equation $Y = f(X)T + g(X,W,Z) + \\epsilon$, with $f$ a degree-$q$ polynomial (cubic in the main results), estimated by double machine learning with cluster-robust standard errors and bootstrap uniform confidence bands. Identification of the mistake share uses the boundary value $x=1$: in a hypothetical precinct where one party receives all presidential votes, opposing-party mistakes cannot exist, so $|f(1)|$ measures the fraction of that party's partisan voters misled by the flip. The contest-level treatment assignment is argued to be quasi-randomized because the NC General Statute's ballot-order rules give each candidate roughly a 50% chance of being listed first, independent of precinct characteristics.","core_discovery":"The central discovery is that party-order flips are a real causal force in nonpartisan judicial races, not noise. Defining a flip as a difference between the party order of a judicial contest and the party order of the presidential race on the same ballot, the paper estimates the conditional average flip effect $f(x)$ on a party's judicial vote share given presidential vote share $x$. Using contest-level treatment assignment generated by the North Carolina statute and double machine learning on 37,690 precinct-level observations, it finds $f$ is significantly nonzero and heterogeneous, with tests rejecting both a zero flip effect and a homogeneous flip effect for each party. The share of partisan-voting mistakes, identified as $|f(1)|$, is 11.8% for Democrats and 15.4% for Republicans. In the placebo sample of races with party labels, the flip effect is indistinguishable from zero. The paper concludes that ballots mixing labeled and unlabeled contests mislead many voters and should be avoided.","pith_inferences":["Beyond the paper, any state with mixed-label ballots and random or rotating candidate order could exhibit the same cue-taking; a natural extension would be re-running the same design on other states' judicial elections where party endorsements can be recovered.","The boundary extrapolation is the fragile link: if name-order effects or nonpartisan crossover voting operate at extreme presidential-vote-share precincts, $|f(1)|$ is not a clean mistake rate, though the qualitative conclusion that flip effects are real likely survives.","A sharper test of the mechanism would use individual-level ballot images or voter files to verify that voters select the candidate in the same position as their party's presidential candidate, rather than the candidate affiliated with their party.","The method's logic—estimating offsetting mistake shares from a heterogeneous treatment-effect curve—could transfer to other settings where agents use a visible ordering cue to infer an unlabeled attribute, such as ranked lists in retail or information displays."],"forward_implications":["A party-order flip in one unlabeled race can shift judicial vote shares by enough to matter in close elections, even when the statewide average effect is zero.","Analyses that only estimate an average treatment effect will miss the distortion entirely; heterogeneity across presidential vote share is necessary to see it.","Including party designations on all contests removes the effect, as the placebo test shows.","The 11.8% and 15.4% figures imply that roughly one in eight Democrats and one in seven Republicans voting in these races did not vote their intent purely because of candidate order.","Ballots that mix partisan and nonpartisan contests should be redesigned, since they misrepresent voter intent rather than merely shift margins."],"supporting_citations":[{"why":"Supplies the double/debiased machine learning estimation used to recover $f(x)$ with valid inference.","marker":"Chernozhukov et al. 2018a"},{"why":"Provides the cluster-robust variance formulas used because precinct outcomes are correlated within contests.","marker":"Cameron and Miller 2015"},{"why":"Establishes that party identification drives vote choice in both partisan and nonpartisan elections, motivating the cue-taking mechanism.","marker":"Bonneau and Cann 2015"},{"why":"Provides the natural-experiment approach to ballot effects and the no-interference and random-assignment assumptions adapted here.","marker":"Ho and Imai 2006"},{"why":"Documents ballot-position and roll-off effects that the paper's design must separate from party-order flips.","marker":"Augenblick and Nicholson 2015"},{"why":"Provides the butterfly-ballot precedent justifying the claim that ballot design can misrecord voter intent at scale.","marker":"Wand et al. 2001"},{"why":"Formalizes the ignorability and stability assumptions used for identification.","marker":"VanderWeele 2008"}],"fun_headline_variants":["Party-order flips mislead 12-15% of NC voters","Ballot order shifts cause wrong votes in judicial races","NC ballot flips: 12.0% D, 15.4% R voted wrong","Mixing labeled and unlabeled races confuses partisan voters","Judicial ballot order flips explain 12-15% voter mistakes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The estimate that 11.8% and 15.4% of voters were misled rests on extrapolating the fitted flip-effect curve to a hypothetical precinct where one party wins 100% of presidential votes; at that boundary every lost judicial vote is attributed to that party's own voters being fooled by the order change.","fun_headline_variants_meta":{"raw":{"variants":["Party-order flips mislead 12-15% of NC voters","Ballot order shifts cause wrong votes in judicial races","NC ballot flips: 12.0% D, 15.4% R voted wrong","Mixing labeled and unlabeled races confuses partisan voters","Judicial ballot order flips explain 12-15% voter mistakes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000371,"raw_usage":{"total_tokens":1966,"prompt_tokens":905,"completion_tokens":1061,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":966}},"tokens_in":521,"tokens_out":1061,"duration_ms":10187,"temperature":1.0,"reasoning_tokens":966,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:03:56.360082+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate $|f(1)|$ using only precincts with presidential vote share above 0.9, or with candidate-name-order fixed effects added; if the fitted curve flattens or reverses at the boundary, the mistake-rate estimates are artifacts of extrapolation rather than measured voter behavior.","supporting_citations":[{"cited_title":"Journal of human resources 50(2):317--372","cited_arxiv_id":null,"evidence_quote":"Provides the cluster-robust variance formulas used because precinct outcomes are correlated within contests."},{"cited_title":"Political Behavior 37(1):43--66","cited_arxiv_id":null,"evidence_quote":"Establishes that party identification drives vote choice in both partisan and nonpartisan elections, motivating the cue-taking mechanism."},{"cited_title":"Journal of the American Statistical Association 101(475):888--900","cited_arxiv_id":null,"evidence_quote":"Provides the natural-experiment approach to ballot effects and the no-interference and random-assignment assumptions adapted here."},{"cited_title":"The Review of Economic Studies 83(2):460--480","cited_arxiv_id":null,"evidence_quote":"Documents ballot-position and roll-off effects that the paper's design must separate from party-order flips."},{"cited_title":"American Political Science Review 95(4):793--810","cited_arxiv_id":null,"evidence_quote":"Provides the butterfly-ballot precedent justifying the claim that ballot design can misrecord voter intent at scale."},{"cited_title":"Statistics in Medicine 27(11):1934--1943","cited_arxiv_id":null,"evidence_quote":"Formalizes the ignorability and stability assumptions used for identification."}],"review_version":1}