{"id":"b4444368-c852-4ae5-a86d-32652ecabb06","arxiv_id":"2506.12961","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new quantitative framework measures the degree to which voting rules violate Arrow's independence and unanimity axioms, and an empirical study finds Borda performs best on Scottish and synthetic elections.","lead":"This paper introduces two numeric scores that measure how much voting rules violate Arrow's fairness axioms on real election data. It tests five voting rules on Scottish ranked-choice ballots and finds Borda is most stable and most majoritarian.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'Borda consistently receives the highest sigma_U' is contradicted by the paper's own bootstrap grouping of Borda with 3-Approval; the empirical central claim needs qualification.","rationale":"I read the paper as making two connected claims: a theoretical equivalence (dictatorship iff sigma_IIA=1 and sigma_U>0 for all profiles) and an empirical finding that Borda consistently scores highest on both metrics. The theoretical equivalence is sound: sigma_IIA=1 recovers IIA by Proposition 2, sigma_U>0 recovers unanimity by Definition 6, and the proof via Arrow's theorem is valid. The quantitative framing is modest, since it is a binary recovery rather than a trade-off bound, but that is a contribution-scope issue, not a correctness error. The empirical finding is where the central claim is least secure. The reader's weakest_assumption concerned STV output conversion; that is a legitimate validity threat, and the authors acknowledge it in Section 5.1. However, a more direct problem is internal: the paper's own bootstrap grouping in Appendix B places Borda and 3-Approval in the same sigma_U group, which contradicts the abstract's 'consistently highest' phrasing. No significance tests or paired comparisons are reported, so the 'consistently' claim is not established. This does not require changing the CONDITIONAL verdict; it requires qualified empirical statements and appropriate statistical support in revision. My concrete test would settle whether Borda actually separates from 3-Approval on sigma_U and sigma_IIA, or whether the headline should be weakened.","tokens_in":12542,"tokens_out":9839,"duration_ms":104143,"concrete_test":"Using the released code and data (github.com/Suvadip2776/Quantitative_fairness), compute paired per-profile differences delta_i = sigma_U(Borda, P_i) - sigma_U(3-Approval, P_i) for all Scottish profiles and for each Bradley-Terry scenario, and likewise for sigma_IIA. Report the fraction of profiles with delta_i > 0 and a paired bootstrap 95% confidence interval for the mean (or median) delta. If the interval includes 0, or 3-Approval has a non-negligible win rate, replace 'consistently highest' with 'statistically tied with 3-Approval on sigma_U' or 'highest in a majority of profiles but not uniformly.' This single check settles whether the empirical headline is accurate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim, stated in the abstract, is that 'the Borda rule consistently receives the highest sigma_IIA and sigma_U scores across observed and synthetic elections.' The body is more cautious: Section 4.1 says Borda 'most frequently receives the highest' scores, and Appendix B reports that, for sigma_U, the bootstrap analysis produces 'three groupings: {Borda, 3-Approval}, {2-Approval}, and {Plurality, STV}.' That grouping means the bootstrapped confidence intervals for Borda and 3-Approval overlap, so Borda is not statistically distinguishable from 3-Approval on sigma_U. The abstract's 'consistently highest' overstates the evidence. This matters because the paper's applied takeaway is Borda's empirical superiority; if 3-Approval is tied on sigma_U, the headline comparison changes. The theoretical Quantitative Arrow theorem is unaffected; it is a correct restatement of Arrow's theorem, though not a quantitative trade-off bound. The empirical claim is the load-bearing part of the paper's contribution, and it is currently stronger than the reported statistics support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two real-valued metrics, σ_IIA and σ_U, that measure, for a given preference profile, how stable a voting rule is under candidate deletion and how well its output respects pairwise majority preferences. It shows that σ_IIA ≡ 1 is equivalent to classical IIA and that σ_U > 0 for all profiles is equivalent to classical Unanimity, yielding a quantitative restatement of Arrow's theorem: a rule is dictatorial if and only if both conditions hold. It also defines a greedy rule that maximizes σ_U and reports empirical comparisons of Plurality, 2-Approval, 3-Approval, Borda, and STV on Scottish local-election data and Bradley-Terry synthetic profiles.","tokens_in":12729,"tokens_out":8692,"duration_ms":96804,"significance":"If the empirical claims held as stated, the paper would provide an interpretable, profile-level framework for measuring axiom compliance in real elections and a useful comparison of common voting rules. Strengths include transparent definitions, a clean equivalence with Arrow's theorem, public code and data, and bootstrap uncertainty analysis. However, the headline empirical claim is stronger than the reported statistics: Table 2 shows ties, Appendix B groups Borda with 3-Approval for σ_U, and no pairwise significance tests are reported. The STV ranking conversion is also acknowledged to be misaligned with the rule's actual output. The theoretical result is a correct but direct corollary of Arrow's theorem rather than a quantitative trade-off bound, so the paper's incremental contribution rests mainly on the metrics and the empirical comparison, which currently need tightening.","major_comments":[{"comment":"The abstract's claim that the Borda rule 'consistently receives the highest σ_IIA and σ_U scores' is not supported by the reported statistics. Section 4.1 uses the weaker wording 'most frequently receives the highest,' Table 2 shows Borda tied with 2-Approval on both metrics for the example profile, and Appendix B reports bootstrap groupings for σ_U of {Borda, 3-Approval}, {2-Approval}, and {Plurality, STV}, meaning Borda and 3-Approval are not statistically separated. No pairwise significance tests are reported. Please either add formal comparisons or qualify the conclusion to state that Borda belongs to the top group rather than being consistently and uniquely highest.","section":"Abstract; §4.1; Appendix B"},{"comment":"The empirical comparison treats STV as a complete-ranking rule by filling in winners and eliminated candidates, a conversion the authors themselves say in Section 5.1 is 'not fully aligned with the nature of the elections where it was used.' This conversion can change σ_IIA and σ_U values for STV and therefore the relative ordering of rules. Because the abstract's headline includes STV among the compared rules, please provide a sensitivity analysis for alternative ways of ranking STV outputs or remove STV from the headline comparison until the conversion is justified.","section":"§2.2; §5.1"},{"comment":"The proof states that g is increasing on [0,1), h is decreasing, and hence g∘h is increasing on [0,n]; the composition of an increasing function with a decreasing function is decreasing, not increasing. The subsequent minimization formula g(h(max δ)) is what would follow from a decreasing composition, so the result appears salvageable, but the monotonicity claim must be corrected and the proof rechecked.","section":"Appendix A, proof of Lemma 4"}],"minor_comments":[{"comment":"The Quantitative Arrow's Theorem should explicitly state the standing assumption m ≥ 3, since Arrow's theorem does not apply to fewer than three candidates.","section":"§3.4"},{"comment":"The Bradley-Terry generation procedure is not fully specified; please provide the exact generative model, including how the Dirichlet parameter α maps to candidate strengths and how partial ballots are produced.","section":"§4.2"},{"comment":"The text notes that Borda and 2-Approval give the same outcome on the example profile; the implication for the 'highest scores' wording should be acknowledged in the main text.","section":"Table 2"},{"comment":"The phrase 'compliment the descriptive box plots' should be 'complement the descriptive box plots.'","section":"Appendix B"},{"comment":"The sentence describing how candidates left when seats are filled are placed 'in between, in order of first-place votes when the process terminates' is vague; please clarify the tie-breaking and termination conventions for STV.","section":"§2.2"},{"comment":"The comparison with Zhao et al. [2024b] would be clearer if the relationship between σ_IIA and their edit-distance-based similarity score were stated formally rather than only informally.","section":"§1.1"}],"recommendation":"major_revision","confidential_remarks":"The theoretical result is a direct corollary of Arrow's theorem, so the paper's novelty rests on the metrics and the empirical comparison. The current empirical overstatement and the acknowledged STV conversion issue should be resolved before publication. No concerns about novelty disclosure or citation behavior beyond the minor issues noted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The theoretical core is in good shape. The two metrics, sigma_IIA built on swap distance and sigma_U built on majority alignment, are natural and the equivalence results connecting them to classical IIA and Unanimity are correct. The greedy rule that provably maximizes sigma_U is a genuine contribution, and the proof of its optimality is clean. The quantitative Arrow theorem is, as the authors more or less admit, a direct corollary of Arrow's theorem via those equivalences, not a new impossibility result. That is fine, but it should not be oversold.\n\nThe empirical story is where the paper gets shaky. The abstract says Borda 'consistently receives the highest' sigma_IIA and sigma_U scores across observed and synthetic elections. The body is more careful, saying Borda 'most frequently' receives the highest scores, and the bootstrap analysis in Appendix B shows that for sigma_U, Borda and 3-Approval fall into the same statistical grouping. That means the confidence intervals overlap, so on the paper's own numbers, Borda is not distinguishable from 3-Approval on sigma_U. The 'consistently highest' claim in the abstract is stronger than the evidence, and this is the load-bearing empirical takeaway. No pairwise significance tests are reported, and the STV-to-ranking conversion, while acknowledged in Section 5.1 as 'not fully aligned' with the actual winner-set elections, could still distort the relative scores. These are fixable issues, but they need to be fixed before the empirical conclusions are stated so baldly.\n\nWhat is genuinely new and useful: the swap-distance formulation of sigma_IIA (an improvement over the edit-distance version in Zhao et al.), the sigma_U metric, the optimization result, and the empirical sweep over Scottish and synthetic data. The metrics give people a concrete, profile-level way to talk about axiom compliance, and the code and data are public. The literature engagement is honest, including the acknowledgment of the prior edit-distance metric.\n\nWho should read this: anyone working in computational social choice, and to a lesser extent AI alignment researchers who want graded measures of voting-rule behavior. It deserves a serious referee. I would send it to peer review, but with a clear request to the authors: align the abstract with the bootstrap results, report uncertainty in the rule comparisons, and either handle STV more carefully or qualify the claims about it.\n\nMy verdict: conditional acceptance, pending the empirical claims being scaled back to what the statistics actually support.","headline":"A clean metric framework for quantifying IIA and Unanimity violations, but the abstract overclaims Borda's empirical edge in a way that the paper's own bootstrap numbers do not support.","tokens_in":13272,"tokens_out":1809,"would_cite":true,"duration_ms":21648,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91B14","91B12"],"pacs":[],"model":"deepseek-v4-flash","headline":"A voting rule passes both relaxed axioms exactly when it is a dictatorship.","keywords":["voting rules","social choice axioms","impossibility theorem","independence of irrelevant alternatives","unanimity","Borda rule","Kendall tau distance","Bradley-Terry model"],"falsifier":"For the empirical claim: re-run the Scottish and Bradley-Terry comparisons treating STV output as a winner set rather than a converted ranking, and check whether Borda's top scores persist. For the theoretical claim: exhibit any single non-dictatorial rule with $\\sigma_{IIA}=1$ and $\\sigma_U>0$ on every profile, which would contradict the quantitative impossibility theorem.","tokens_in":12313,"feed_emoji":"🗳️","tokens_out":11976,"duration_ms":107030,"temperature":0.7,"pith_summary":"This paper converts two classical axioms of social choice, independence of irrelevant alternatives and unanimity, from binary guarantees into profile-level scores. $\\sigma_{IIA}(f,P)$ measures how much a voting rule's output ranking changes when a candidate is removed, and $\\sigma_U(f,P)$ measures how well the output respects majority pairwise preferences. The paper proves that a rule scores $\\sigma_{IIA}=1$ and $\\sigma_U>0$ on every profile exactly when it is a dictatorship, so the classical impossibility theorem survives verbatim in the metric setting. Used as tools, these scores let five common rules be compared on thousands of real and simulated elections, and the Borda rule comes out with the best scores on both dimensions. This turns a long-standing critique of binary axiomatics into a practical way to grade election rules by their actual behavior.","feed_headline":"Only voting rules passing both relaxed-axiom tests are dictatorships","feed_subtitle":"New continuous scores measure stability and majority agreement; Borda ranks first on real and simulated elections.","key_machinery":"The machinery is a pair of continuous scores computed from a voting rule $f$ and a profile $P$. $\\sigma_{IIA}$ uses Kendall-tau swap distance between the ranking $f(P)$ and the ranking produced after deleting each candidate one at a time; it equals 1 when deletion never changes the relative order among remaining candidates and 0 when every deletion reverses the order completely. $\\sigma_U$ uses the pairwise comparison graph of majority margins: it is 1 when the output ranking is a topological sort of that graph, and it decreases toward 0 as the worst-aligned pair of candidates approaches a unanimous preference that the rule reverses. These definitions make the classical axioms exact limits of the metrics, which is what allows the quantitative impossibility theorem to be proved by direct appeal to the classical one.","core_discovery":"The central claim is the quantitative impossibility theorem: a voting rule $f$ is a dictatorship if and only if $\\sigma_{IIA}(f,P)=1$ and $\\sigma_U(f,P)>0$ for all preference profiles $P$. The proof runs through a proposition showing that $\\sigma_{IIA}\\equiv 1$ is equivalent to classical IIA, and through the definition of $\\sigma_U$, whose positivity rules out unanimity violations. Thus the theorem is a faithful quantitative restatement of the classical impossibility result, not a new weakening of it. The empirical discovery is that on 1,070 Scottish local-government elections and on Bradley-Terry synthetic profiles, Borda consistently records the highest average $\\sigma_{IIA}$ and $\\sigma_U$ among Plurality, 2-Approval, 3-Approval, Borda, and STV, a pattern the authors connect to a recent result that weakening IIA to allow preference intensity uniquely selects Borda.","pith_inferences":["Editorial inference: because $\\sigma_U$ depends only on $f(P)$ and the profile, it can be computed from a single output ranking, making it a cheap audit tool for deployed voting algorithms in multi-agent or LLM settings.","Editorial inference: the same template could be applied to other binary axioms such as monotonicity, producing a family of continuous axiom scores and a multi-dimensional axiom profile for each rule.","Editorial inference: if STV were scored as a winner set instead of the converted complete ranking used here, the relative standing of the rules could shift, so the Borda-dominance conclusion is tied to the ranking representation."],"forward_implications":["For any fixed election profile, axiom compliance becomes a number in $[0,1]$ instead of a yes/no answer, so rules can be ranked by how closely they satisfy the axioms in that election.","The quantitative impossibility theorem implies that no non-dictatorial rule can maintain perfect candidate-removal stability and full unanimity respect across all possible profiles.","The paper's greedy topological-sort algorithm shows that a rule can be constructed in polynomial time to maximize $\\sigma_U$ on any profile, so majoritarian alignment is an optimizable objective.","On both observed Scottish elections and synthetic Bradley-Terry profiles, Borda has the highest $\\sigma_{IIA}$ and $\\sigma_U$ among the five rules tested, supporting the view that Borda is a good compromise between stability and majority responsiveness."],"supporting_citations":[{"why":"The classical impossibility theorem whose quantitative restatement is the paper's main theoretical result.","marker":"[Arrow, 1950]"},{"why":"Recent result that weakening IIA to include preference intensity uniquely selects Borda, which frames the empirical finding.","marker":"[Maskin, 2025]"},{"why":"VoteKit software package used to implement the voting rules and compute STV outcomes in the experiments.","marker":"[Data and Democracy Lab, 2024]"},{"why":"Serves as the justification that uniform synthetic models are unrealistic and motivates the Bradley-Terry simulation design.","marker":"[Tideman and Plassmann, 2010]"},{"why":"Closely related IIA-stability metric based on edit distance that the paper compares against and argues swap distance improves upon.","marker":"[Zhao et al., 2024b]"},{"why":"Documents the Scottish local-government STV elections that provide the real-world preference profiles used in the empirical tests.","marker":"[McCune and Graham-Squire, 2024]"}],"fun_headline_variants":["Perfection is dictatorship: quantitative Arrow axioms","Borda wins when Arrow's axioms go continuous","New scores turn Arrow's theorem into a fine-grained test","Dictatorship is the only perfect match to relaxed Arrow axioms","Quantifying fairness: Borda tops relaxed Arrow tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The empirical comparison assumes that converting STV winner sets into a complete ranking for the metrics does not distort the comparison; the authors themselves note that this representation is 'not fully aligned' with how the elections were run.","fun_headline_variants_meta":{"raw":{"variants":["Perfection is dictatorship: quantitative Arrow axioms","Borda wins when Arrow's axioms go continuous","New scores turn Arrow's theorem into a fine-grained test","Dictatorship is the only perfect match to relaxed Arrow axioms","Quantifying fairness: Borda tops relaxed Arrow tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000413,"raw_usage":{"total_tokens":2214,"prompt_tokens":1100,"completion_tokens":1114,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":716,"completion_tokens_details":{"reasoning_tokens":1038}},"tokens_in":716,"tokens_out":1114,"duration_ms":8935,"temperature":1.0,"reasoning_tokens":1038,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:37:15.562295+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For the empirical claim: re-run the Scottish and Bradley-Terry comparisons treating STV output as a winner set rather than a converted ranking, and check whether Borda's top scores persist. For the theoretical claim: exhibit any single non-dictatorial rule with $\\sigma_{IIA}=1$ and $\\sigma_U>0$ on every profile, which would contradict the quantitative impossibility theorem.","supporting_citations":[{"cited_title":"A difficulty in the concept of social welfare","cited_arxiv_id":null,"evidence_quote":"The classical impossibility theorem whose quantitative restatement is the paper's main theoretical result."},{"cited_title":"Borda’s rule and A rrow’s independence condition","cited_arxiv_id":null,"evidence_quote":"Recent result that weakening IIA to include preference intensity uniquely selects Borda, which frames the empirical finding."},{"cited_title":"VoteKit : Python package","cited_arxiv_id":null,"evidence_quote":"VoteKit software package used to implement the voting rules and compute STV outcomes in the experiments."},{"cited_title":"The structure of the election-generating universe","cited_arxiv_id":null,"evidence_quote":"Serves as the justification that uniform synthetic models are unrealistic and motivates the Bradley-Terry simulation design."},{"cited_title":"Monotonicity anomalies in S cottish local government elections","cited_arxiv_id":null,"evidence_quote":"Documents the Scottish local-government STV elections that provide the real-world preference profiles used in the empirical tests."}],"review_version":1}