Pith. sign in

REVIEW 1 major objections 5 minor 17 references

A parliamentary majority can be certified by pooling seat-level evidence processes rather than certifying every seat, and the resulting audit reduces ballots inspected by orders of magnitude.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 08:31 UTC pith:IUJDBOGQ

load-bearing objection Solid, useful RLA paper for parliamentary majorities; the main theorem is right under an unstated independence assumption, and Section 2.3 overclaims robustness to cross-seat dependence. the 1 major comments →

arxiv 2607.21082 v1 pith:IUJDBOGQ submitted 2026-07-23 stat.AP cs.CRcs.CYstat.ME

Risk-Limiting Audits for Parliamentary Majorities

classification stat.AP cs.CRcs.CYstat.ME
keywords risk-limiting auditparliamentary majoritypartial conjunction hypothesise-processanytime-valid sequential testadaptive samplingelection auditing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper's central claim is that certifying a parliamentary majority can be reduced to testing the partial conjunction hypothesis that at least r of the reported seats were truly won — a 'partial conjunction' meaning some but not necessarily all of the seats must be confirmed. The authors construct an e-process (a sequence of nonnegative statistics that grows as evidence accumulates) for each seat, and combine the weakest ones by multiplication. They prove that this combined process controls the risk limit, so once it crosses 1/α the majority is certified. Simulations on synthetic data and the 2014 Indian Lok Sabha election show the majority audit can require orders of magnitude fewer ballots than auditing every reported winning seat, sometimes by almost a thousandfold.

Core claim

Formally, the reported winning party is seen as having won a set W of seats, with r the threshold for a majority. The null hypothesis that the party did not truly win a majority is equivalent to the union, over all subsets C of W of size |W|-r+1, of the intersection nulls 'all seats in C were falsely reported'. For each seat, the authors define an e-process E_{s,t} that is the minimum over a set of per-assertion betting processes. Their majority statistic E_{r/W,t} is the product of the |W|-r+1 smallest of these seat-level processes at time t. Theorem 1 states that this product is anytime-valid for testing the majority null: under that null, the probability that E_{r/W,t} ever reaches 1/α is

What carries the argument

The central object is the seat-level e-process E_{s,t}, built as the minimum over per-assertion betting processes for that seat. The majority statistic E_{r/W,t} multiplies the |W|-r+1 smallest such values at each time. This product grows whenever any sufficiently large group of seats accumulates evidence, and its anytime-validity follows because each E_{s,t} is a test supermartingale and because, conditional on the audit history, the next-drawn ballot values are independent across seats. The adaptive sampling schemes determine which seats' ballots to draw next; since the decisions depend only on the history, they preserve the anytime-validity property.

Load-bearing premise

The risk-limit proof hinges on the assumption that, given the audit history, the next ballot drawn from any seat is independent of the next ballot drawn from any other seat — that is, each seat's ballot order is an independent random permutation.

What would settle it

Simulate a small parliament (e.g., five seats, majority threshold three) where ballots across seats are drawn from a single shared random stream so that future draws are dependent across seats; set the truth so that only two seats are truly won; run the audit repeatedly and count how often the product statistic E_{r/W} reaches 1/α before a full recount. If that frequency exceeds the risk limit α, the anytime-validity claim fails under cross-seat dependence.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Election officials can certify a parliamentary majority after hand-inspecting a small fraction of ballots, often orders of magnitude fewer than certifying every reported seat.
  • The majority audit can run alongside a full seat-by-seat audit; the majority becomes certified as soon as the product statistic crosses 1/α, and the seat-by-seat process can continue or escalate to recounts.
  • The method applies to any 'at least r of the reported winners' hypothesis, not just majorities, and to any contest expressible as assertions about mean ballot scores.
  • The direct product of e-processes is uniformly more powerful than the previously proposed combination that uses a chi-square threshold, so implementations gain efficiency by using the product directly.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the method is adopted, the dominant cost factor becomes the number of 'spare' seats beyond the majority threshold, since the product statistic is driven by the smallest |W|-r+1 seat-level processes; this suggests audits of coalitions or supermajorities could be tailored to the size of the cushion.
  • The Bayesian filter used to exclude likely-false seats introduces a prior whose influence on sampling is not covered by the anytime-validity guarantee; a fully frequentist or e-value-based filtering rule could offer the same robustness without a prior.
  • The independence assumption across seats may be violated in implementations using a single shared random stream; the paper does not provide a valid combination rule for that case, so a practical protocol should either enforce independent per-seat permutations or develop a dependence-robust variant.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper proposes a risk-limiting audit (RLA) for parliamentary majority outcomes. It formalizes the audit as a partial conjunction test: rather than certifying every reported winning seat, it suffices to certify that at least r of the reported winning seats W were truly won, requiring testing H_{r/W} (fewer than r true wins). The authors construct seat-level e-processes E_{s,t} using SHANGRLA assertions and ALPHA-style betting, then combine them into a parliament-level e-process E_{r/W,t} as the product of the |W|-r+1 smallest seat-level processes. They prove anytime-validity of E_{r/W} via Ville's inequality (Theorem 1). They introduce adaptive sampling rules (non-adaptive, greedy, filtered, and combinations) that allocate sampling effort across seats, and evaluate them in simulations on synthetic two-candidate contests and the 2014 Indian Lok Sabha election, reporting large reductions in ballots sampled compared with certifying all reported seats.

Significance. If Theorem 1 holds under the intended assumptions, the paper makes a substantial contribution to election auditing. The partial conjunction formulation is natural and appears new for RLA in this form, and the e-process combination is more efficient than the earlier Fisher-based proposal. The adaptive sampling schemes are practical and well-motivated, and the simulation study is extensive and reproducible, with code provided. The theoretical development is based on standard supermartingale and Ville arguments and is mostly careful. However, as discussed below, one key assumption in the proof of Theorem 1 is implicit and the paper overclaims robustness to cross-seat dependence; this must be corrected before the result can be accepted as stated.

major comments (1)
  1. [Section 2.3 / Appendix B, step (a)] The proof of the anytime-validity of E_{1/C} (and hence E_{r/W}) requires that, conditional on F_{t-1}, the next-drawn assorter values X_{s,j_s,n_s(t)} are independent across seats s∈C. This is used at step (a) of the Appendix B derivation. The manuscript does not state this assumption in Section 2.1 or in Theorem 1; moreover, Section 2.3 explicitly claims that the construction 'stays valid even when sampling induces cross-seat dependence, extending Vovk and Wang [17]'. That claim is not supported: if the per-seat random permutations are coupled (e.g., a common random stream selects ballot positions in multiple seats), the product of nonnegative supermartingales need not be a supermartingale, so the risk limit can be violated. This is a load-bearing issue for the main theorem. The fix is straightforward: explicitly assume that the seat-level sampling permutations are independent across s
minor comments (5)
  1. [Section 2.3] The sentence 'As shown in Appendix B, it stays valid even when sampling induces cross-seat dependence' is in tension with the later sentence 'their e-processes are conditionally independent and can be multiplied.' Please clarify that the adaptive decisions D_{s,t} may be cross-seat dependent (because they depend on F_{t-1}), while the ballot draws themselves are independent across seats.
  2. [Section 3, Greedy scheme] The definition of δ_t in (12) would be easier to follow if the constant k = |W|−r+1 were introduced explicitly, and if the text explained why the set of i satisfying the condition is an initial segment of {1,…,r}.
  3. [Appendix C] The statement that Mohanty et al. [9] 'justified the theoretical validity of their Fisher combination proposal incorrectly' is categorical. Please provide a precise citation to the step in [9] that relies on independence, or soften the wording.
  4. [Section 2.1] Typo: 'satisfiy' should be 'satisfy'. Also, the phrase 'for seatsisn't' should read 'for seat s isn't'.
  5. [Figures 2 and 3] The legends are dense; consider marking the benchmark methods ('All seats', 'Reported top-r seats') with a distinct symbol or label so the comparison to the proposed methods is easier to read.

Circularity Check

0 steps flagged

No circular derivation; the only self-citation is a non-load-bearing ALPHA tuning choice.

full rationale

The derivation chain is self-contained. Seat-level M_{s,j} is proved to be a nonnegative test supermartingale under H_{s,j} (Appendix A, eqs. (5)-(6)); E_s inherits anytime-validity via the min construction and Ville's inequality. E_{1/C} is a product of seat-level e-processes, with the proof explicitly relying on conditional independence across seats in Appendix B step (a). E_{r/W} is the product of the |W|-r+1 smallest seat-level processes, and the union representation (10) supplies the subset C_0 whose intersection null holds under H_{r/W}. No parameter is fitted to a subset of the data and then presented as a prediction; the simulation results are benchmark evaluations, not fitted outputs. The ALPHA tuning constants d=200 and eta0=0.51 are taken from Ek et al. [4], a self-citation for one co-author, but Theorem 1 holds for any admissible lambda sequence, so this citation is not load-bearing for the anytime-validity guarantee; it only affects simulated efficiency. One non-circular caveat: Section 2.3 claims the construction 'stays valid even when sampling induces cross-seat dependence, extending Vovk and Wang [17]', while Appendix B step (a) requires that, conditional on F_{t-1}, the next-drawn assorter values 'are independent across s in C'. This is an assumption/overclaim issue rather than circularity. Score 2 reflects only the minor non-load-bearing self-citation; no circular step is present.

Axiom & Free-Parameter Ledger

6 free parameters · 7 axioms · 0 invented entities

All hand-set tuning parameters (d, η0, κ0, τ, a, ε) affect the growth rate and allocation of sampling, but not the anytime-validity guarantee. The central theorem depends on standard e-process theory plus the domain assumption that seat-level ballot permutations are independent. No new entities are introduced; the 'filter' is a Bayesian posterior used only for sampling decisions.

free parameters (6)
  • ALPHA truncated-shrinkage d = 200
    Tuning parameter for λ_s,j,i in the seat-level betting process, taken from Ek et al. (2024); affects growth rate and therefore sample sizes, not validity.
  • ALPHA truncated-shrinkage η0 = 0.51
    Second ALPHA tuning parameter, same source; setting slightly above 0.5 keeps the prior mean just over the null boundary.
  • Dirichlet concentration κ_s,0 = 200
    Prior concentration for the filtering posterior, chosen to mimic η0 and d; only affects which seats are sampled, not the risk limit.
  • posterior exclusion threshold τ = 0.01
    Hand-chosen threshold below which a seat is judged unlikely to have been truly won; a liberal filter.
  • greedy aggressiveness a = 0 or 3
    Constant in the d_t formula; a=0 is most aggressive, a=3 hedges by sampling extra seats; simulations use both.
  • filter margin epsilon = 0
    Fixed at 0 so the filter only deprioritizes seats that appear falsely reported; positive epsilon is left to future work.
axioms (7)
  • domain assumption SHANGRLA: each seat's social choice function decomposes into finitely many half-average assertions with non-negative assorter functions
    Assumed throughout Section 2.1 as the basis for seat-level e-processes; SHANGRLA provides this decomposition.
  • domain assumption Ballots in each seat are a random permutation; permutations are independent across seats
    Required for the conditional-independence step (a) in Appendix B that makes the product E_{1/C} a supermartingale.
  • domain assumption Sampling without replacement within each seat, with D_{s,t} measurable w.r.t. F_{t-1}
    The process is defined over rounds where each seat may or may not be sampled; the auditor's decisions must not look ahead.
  • standard math Ville's inequality and e-process/e-value theory
    Used in Theorem 1 to convert test supermartingales into anytime-valid tests.
  • domain assumption Ballot-polling with no invalid ballots in the simulations
    Section 4 assumes no invalid ballots; the text notes invalid ballots could be scored as 1/2 but this is not simulated.
  • domain assumption India 2014 dataset from community-scraped ECI data is accurate
    Section 4.2 uses candidate-level vote totals from a GitHub dataset; errors would change the simulation numbers, not the method's validity.
  • ad hoc to paper Dirichlet-multinomial posterior approximation for filtering
    The multinomial model approximates sampling without replacement; the authors note it is only used to choose sampling decisions and does not affect the risk limit, but it is an ad hoc modeling choice.

pith-pipeline@v1.3.0-alltime-deepseek · 14304 in / 21773 out tokens · 211957 ms · 2026-08-01T08:31:34.122248+00:00 · methodology

0 comments
read the original abstract

Existing methods for risk-limiting audits typically focus on certifying individual contests. In parliamentary elections, however, the politically relevant outcome is often whether a party has won enough seats to form government, not whether every reported seat outcome is correct. Extending on the work of Mohanty et al. (2019), we formulate the certification of a parliamentary majority as a partial conjunction testing problem: it is enough to verify that the reported winning party truly won at least a majority of its reported seats. Building on the SHANGRLA auditing framework, we construct a sequential audit statistic for the majority outcome by combining seat-level statistics. We then propose adaptive sampling strategies that allocate auditing effort across seats, including variants that learn to avoid spending excessive effort on seats that appear unlikely to have been truly won. Using simulations based on synthetic and real data, from the 2014 Indian Lok Sabha election, we show that auditing the parliamentary majority can substantially reduce the number of ballots inspected (by almost a thousand-fold) compared to certifying every reported winning seat.

Figures

Figures reproduced from arXiv: 2607.21082 by Damjan Vukcevic, Dennis Leung, Jack Freestone.

Figure 1
Figure 1. Figure 1: Illustration of the four sampling schemes at a single round, with |W| = 5 seats and r = 3. Bars show the current values of Es,t; the dashed red line is 1/α. Blue bars indicate seats sampled in the next round; grey bars indicate seats above the active-set cutoff; orange bars indicate seats excluded by the posterior filter (πs ⩽ τ ). Seat s1 is falsely reported. (a) Non-adaptive: every seat is sampled. (b) G… view at source ↗
Figure 2
Figure 2. Figure 2: Simulated two-candidate plurality contests. The total number of ballots sam￾pled to certification: each point shows the mean across 100 replicates and error bars denote ±2 standard deviations. Columns correspond to the number of falsely reported seats, nfalse. Rows correspond to different ways to set the margins: top row shows Scenario 1, bottom row shows Scenario 2. Results for the All seats method are on… view at source ↗
Figure 3
Figure 3. Figure 3: India 2014. Box plots of the total number of ballots sampled to certification across 100 replicates, with one panel per number of falsely reported seats nfalse. Results for the All seats method are omitted for nfalse > 0, like in [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Extended simulation results for the two-candidate plurality contests of Sce￾nario 1 from Section 4.1. The plot matches the first row of [PITH_FULL_IMAGE:figures/full_fig_p019_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Extended simulation results for the two-candidate plurality contests of Sce￾nario 2 from Section 4.1. The plot matches the second row of [PITH_FULL_IMAGE:figures/full_fig_p020_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Extended simulation results for the two-candidate plurality contests of Sce￾nario 2 from Section 4.1. The plot matches the second row of [PITH_FULL_IMAGE:figures/full_fig_p021_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Extended simulation results for the two-candidate plurality contests of Sce￾nario 2 from Section 4.1. The plot matches the second row of [PITH_FULL_IMAGE:figures/full_fig_p022_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

17 extracted references · 12 linked inside Pith

  1. [1]

    Bio- metrics 64(4), 1215–1222 (2008)

    Benjamini, Y., Heller, R.: Screening for partial conjunction hypotheses. Bio- metrics 64(4), 1215–1222 (2008)

  2. [2]

    Statistical Science 38(4), 602–620 (2023), Preprint: arXiv:2210.00522

    Bogomolov, M., Heller, R.: Replicability across multiple studies. Statistical Science 38(4), 602–620 (2023), Preprint: arXiv:2210.00522

  3. [3]

    In: E-Vote-ID 2023

    Ek, A., Stark, P.B., Stuckey, P.J., Vukcevic, D.: Adaptively weighted audits of instant-runoff voting elections: A W AIRE. In: E-Vote-ID 2023. LNCS, vol. 14230, pp. 35–51. Springer (2023), Preprint: arXiv:2307.10972

  4. [4]

    In: Financial Cryptography and Data Security

    Ek, A., Stark, P.B., Stuckey, P.J., Vukcevic, D.: Efficient weighting schemes for auditing instant-runoff voting elections. In: Financial Cryptography and Data Security. FC 2024. LNCS, vol. 14746, pp. 18–32. Springer (2025), Preprint: arXiv:2403.15400

  5. [5]

    Hafner Publishing Co., New York, fourteenth edn

    Fisher, R.A.: Statistical methods for research workers. Hafner Publishing Co., New York, fourteenth edn. (1973)

  6. [6]

    Journal of the Royal Statistical Society Series B: Statistical Methodology 87(1), 56–73 (2025), Preprint: arXiv:2306.09976 16 Freestone J, Leung D, Vukcevic D

    Gablenz, P., Sabatti, C.: Catch me if you can: signal localization with knock- off e-values. Journal of the Royal Statistical Society Series B: Statistical Methodology 87(1), 56–73 (2025), Preprint: arXiv:2306.09976 16 Freestone J, Leung D, Vukcevic D

  7. [7]

    Journal of Statistical Computation and Simulation 92(10), 2184–2204 (2022), Preprint: arXiv:2104.13081

    Hoang, A.T., Dickhaus, T.: Combining independent p-values in replicability analysis: A comparative study. Journal of Statistical Computation and Simulation 92(10), 2184–2204 (2022), Preprint: arXiv:2104.13081

  8. [8]

    IEEE Security & Privacy 10(5), 42–49 (2012)

    Lindeman, M., Stark, P.B.: A gentle introduction to risk-limiting audits. IEEE Security & Privacy 10(5), 42–49 (2012)

  9. [10]

    In: E- Vote-ID 2018

    Ottoboni, K., Stark, P.B., Lindeman, M., McBurnett, N.: Risk-limiting audits by stratified union-intersection tests of elections (SUITE). In: E- Vote-ID 2018. LNCS, vol. 11143, pp. 174–188. Springer (2018), Preprint: arXiv:1809.04235

  10. [11]

    Statistical Science 38(4), 576–601 (2023), Preprint: arXiv:2210.01948

    Ramdas, A., Gr¨ unwald, P., Vovk, V., Shafer, G.: Game-theoretic statistics and safe anytime-valid inference. Statistical Science 38(4), 576–601 (2023), Preprint: arXiv:2210.01948

  11. [12]

    arXiv:2409.06680 (2026)

    Spertus, J.V., Sridhar, M., Stark, P.B.: Sequential stratified inference for the mean. arXiv:2409.06680 (2026)

  12. [13]

    In: Financial Cryptography and Data Security

    Stark, P.B.: Sets of half-average nulls generate risk-limiting audits: SHANGRLA. In: Financial Cryptography and Data Security. FC 2020. LNCS, vol. 12063, pp. 319–336. Springer (2020), Preprint: arXiv:1911.10035

  13. [14]

    The Annals of Applied Statistics 17(1), 641–679 (2023), Preprint: arXiv:2201.02707

    Stark, P.B.: ALPHA: Audit that learns from previously hand-audited bal- lots. The Annals of Applied Statistics 17(1), 641–679 (2023), Preprint: arXiv:2201.02707

  14. [15]

    Ville, J.: Etude critique de la notion de collectif, Monographies des Proba- bilites, vol. 3. Gauthier-Villars Paris (1939)

  15. [16]

    The Annals of Statistics 49(3), 1736–1754 (2021), Preprint: arXiv:1912.06116

    Vovk, V., Wang, R.: E-values: Calibration, combination and applica- tions. The Annals of Statistics 49(3), 1736–1754 (2021), Preprint: arXiv:1912.06116

  16. [17]

    Electronic Journal of Statistics 18(1), 1185–1205 (2024), Preprint: arXiv:2007.06382

    Vovk, V., Wang, R.: Merging sequential e-values via martingales. Electronic Journal of Statistics 18(1), 1185–1205 (2024), Preprint: arXiv:2007.06382

  17. [18]

    Waudby-Smith, I., Ramdas, A.: Estimating means of bounded random vari- ables by betting. Journal of the Royal Statistical Society Series B: Statistical Methodology 86(1), 1–27 (2024), Preprint: arXiv:2010.09686 A Proof thatM s,j is a test supermartingale underHs,j Non-negativity holds sinceX s,j,i ⩾0 andλ s,j,i ⩽µ −1 s,j,i in (4), andM s,j,0 = 1 by defini...