{"id":"adf14b02-38ec-4356-9868-0ee5d104d596","arxiv_id":"2412.01845","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper claims computer simulations show that roughly 175,000 votes in Georgia's 2024 parliamentary election were manipulated, with a plausible range of 90,000 to 245,000.","lead":"This paper uses official precinct-level results from Georgia's 2024 parliamentary election to argue that between 140,000 and 245,000 votes were manipulated, with 175,000 as the most likely figure. It is a focused statistical forensics claim about a contested national election, relevant to readers tracking election integrity and to researchers testing fraud-detection methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 175,000-vote estimate rests on an unvalidated Gaussian-uniformity null model; the paper's own per-district sigma results contradict the model, so the histogram mismatch is not a measure of manipulation.","rationale":"The reader's weakest assumption identifies the same load-bearing point: the simulated null distribution is assumed to be Gaussian and tightly clustered around district means, and the paper provides no external validation for that assumption. My stress-test confirms this is the critical hinge. The central numerical claim of approximately 175,000 manipulated votes depends entirely on the shape of the simulated histogram, since the manipulation count is literally the difference between simulated and official histograms in Phases II and III. If the null model is misspecified, that difference is not a count of manipulated votes but a measure of model error. The paper even demonstrates the misspecification in its own Section 3: using per-district sigma estimated from official data, the model cannot reproduce the official 54% result, producing only 51-52%. The authors treat this as evidence of fraud, but it is equally (and more parsimoniously) evidence that within-district heterogeneity in real elections exceeds the model's Gaussian small-sigma assumption. Urban-rural divides, ethnic geography, and diaspora voting patterns can produce exactly the kind of spread that Phase II/III capture. Thus the load-bearing assumption is not merely unproven; it is contradicted by the paper's own analysis. A nonparametric or covariate-adjusted counterfactual would settle whether any discrepancy remains after accounting for legitimate heterogeneity. I therefore agree with the reader's rejection, without changing the verdict. The paper does provide some real anomalies, such as precincts with over 100% turnout, but these are not quantitatively linked to the 175,000 figure and would not survive as statistical evidence of manipulation without a valid counterfactual model.","tokens_in":5547,"tokens_out":3912,"duration_ms":37623,"concrete_test":"Recompute the counterfactual histogram without the Gaussian-uniformity assumption: for each district, draw precinct-level Georgian Dream shares by nonparametric bootstrap from the district's actual empirical distribution of official precinct shares, weighting draws so the overall Georgian Dream share matches the official 54%; then compute the Phase II/III vote deficit exactly as in Section 4. If the deficit falls outside the paper's 90,000-245,000 range or changes sign, the 175,000 estimate is an artifact of the assumed null model; if the deficit persists with similar magnitude, the concern is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central 175,000-vote estimate (Section 4) is obtained by subtracting the official precinct-level histogram from a simulated counterfactual. The counterfactual is generated by a model (Section 3) that assumes that, absent manipulation, every precinct's Georgian Dream vote share is a draw from a Gaussian with a small, common standard deviation sigma (0 < sigma < 0.1) around the district mean. This assumption is never validated against any fair-election baseline. More importantly, the paper itself provides evidence against it: when sigma is instead estimated per district from the official data, the simulation yields an overall Georgian Dream share of only 51-52%, not the official 54% (Section 3). The authors interpret this as 'a strong indication of manipulation,' but the correct reading is that the official data contain more within-district heterogeneity than the Gaussian-uniformity model allows. Legitimate urban-rural, ethnic, and diaspora compositional differences can generate such heterogeneity. Consequently, the histogram deficit in 'Phase II' and surplus in 'Phase III' are a measure of model misspecification, not of manipulated votes. The 'most probable' 175,000 figure is therefore not a mathematically grounded estimate of falsification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes official precinct-level returns from Georgia's October 2024 parliamentary election and claims to find statistical evidence that the result was rigged. The authors construct a simulation in which precinct-level Georgian Dream vote shares are drawn from a Gaussian distribution around each district mean, with a common parameter sigma restricted to (0, 0.1); simulations are retained only if the overall Georgian Dream share falls between 53% and 55%. Comparing the simulated histogram of precinct-level shares to the official histogram, the authors identify a deficit of precincts in an intermediate range (Phase II) and a surplus at higher shares (Phase III), and convert the histogram differences into an estimate of 140,000–200,000 manipulated votes (5% bins) or 90,000–245,000 votes (1% bins), with a claimed most probable value of 175,000. The paper also notes anomalies such as precincts with turnout exceeding 100% and precincts with missing registered-voter data, and concludes that the election was rigged.","tokens_in":5702,"tokens_out":5644,"duration_ms":50694,"significance":"If the 175,000-vote estimate were methodologically sound, the paper would be a significant contribution to election forensics and to public understanding of the Georgian election. The authors deserve credit for making their code and data available on GitHub and for highlighting objective anomalies (turnout above 100%, missing voter rolls) that independently warrant investigation. However, the central estimate rests on an unvalidated and internally contradicted null model; the paper does not supply a transparent derivation of the vote count from histogram differences, and it does not compare its method against any known fair election. The significance of the paper as evidence for manipulation is therefore not established.","major_comments":[{"comment":"The paper's own second approach, in which sigma is estimated per electoral district from the official data, yields an overall Georgian Dream share of only 51–52% rather than the official 54% (Section 3). The authors interpret this as 'a strong indication of manipulation' (Section 6), but the more direct reading is that the official precinct-level data contain more within-district heterogeneity than the Gaussian-uniformity model permits. Legitimate compositional effects such as urban–rural divides, ethnic segregation, and differential diaspora turnout generate exactly this kind of spread. The per-district sigma result therefore undercuts rather than supports the model, and the Phase II/Phase III histogram difference is not identified as manipulated votes.","section":"Section 3, second approach"},{"comment":"The simulation is built to reproduce the official 54% total: 'ensuring that the party achieved the same 54% result overall' (Section 3), and sigma is restricted to (0, 0.1) by that requirement. In addition, only simulations in which the overall result falls between 53% and 55% are retained, which excludes approximately 70% of the runs. The counterfactual is therefore conditioned on a summary of the very official data it is compared with. Consequently, the histogram difference measures only the shape of the distribution conditional on the official total; it cannot speak to whether the official total itself is the product of manipulation, which is the load-bearing claim of the paper.","section":"Section 3, calibration and filtering"},{"comment":"The central estimate of 175,000 manipulated votes is not derived transparently. The text states only that the number is obtained 'by calculating the difference between our computational results and real data in Phase II' (Section 4), but it does not give the formula that converts histogram bin counts into a number of votes, nor does it explain how the 803 simulated elections are aggregated to produce a 'most probable' value (mean, median, mode, or something else). The ranges differ substantially between the 5%-bin analysis (140,000–200,000) and the 1%-bin analysis (90,000–245,000), and no explanation is given for why the most probable value is identical across the two, nor is any uncertainty or confidence interval attached to the estimate. Without the conversion formula, the headline number is not reproducible.","section":"Section 4, vote-count conversion"},{"comment":"The division into Phase I (about 0–45%), Phase II, and Phase III is introduced only after inspecting the data, and the percentage boundaries of Phase II and Phase III are never defined precisely. The paper states that 'Phases II and III correspond to rural areas,' but it does not specify the numerical ranges or the algorithm by which bins are assigned to phases. Because the estimated vote count is the sum of histogram differences over the Phase II bins, the result is directly sensitive to the arbitrary placement of the phase boundary. A principled, pre-specified definition of the phases is required before the histogram mismatch can be interpreted as a quantitative estimate of manipulation.","section":"Section 4, phase boundaries"},{"comment":"The paper itself provides evidence against the Gaussian-uniformity assumption. In Section 4 it reports that 'a large majority of the precincts that fall within this Phase are urban precincts. Phases II and III correspond to rural areas,' and it notes that diaspora precincts produce a sharp peak in the simulated distribution because they were 'treated as a single electoral district in the code,' an approach that 'does not perfectly reflect reality.' These admissions indicate that known covariates—urbanicity and diaspora status—systematically affect precinct-level vote shares. A model that ignores these covariates cannot serve as a valid counterfactual for the absence of manipulation; the observed discrepancy may simply reflect the omitted covariates rather than fraud.","section":"Section 4, model misspecification"}],"minor_comments":[{"comment":"The manuscript contains many grammatical errors and typographical issues that impede readability; for example, 'The official data provided by Central Election Commission was analyzed' should be 'The official data ... were analyzed,' and 'the elections' is often used where 'the election' is meant. A thorough language edit is needed.","section":"General"},{"comment":"The description of the histograms is ambiguous: 'received from 0%-20% in a small number of precincts, about 40% in 60 precincts, 80% in 25 precincts, and 100% in zero precincts' does not clarify whether the percentages are bin centers or exact values, and the phrase '100% in zero precincts' is confusing. The figure caption should specify the bin labels and the interpretation.","section":"Section 2, Figure 1"},{"comment":"The definition of the standard deviation uses N instead of N-1, so it is a population standard deviation; the text should state this explicitly. Also, the sentence 'we did not need to choose anything manually because the requirement that the party must achieve 54% overall fixed the parameter sigma within a very small range' is not logically transparent, since the range 0<sigma<0.1 is an assumption rather than a derived consequence.","section":"Section 3"},{"comment":"The phrase 'stolen votes damaged the opposition twice' (Section 5) is unclear; the paper does not explain the double-counting mechanism. In the Conclusion, the statement that the result 'would have changed the election results dramatically' is asserted without defining the counterfactual seat allocation or margin.","section":"Section 5 and 6"},{"comment":"The reference to a probability textbook [1] is not used to justify the assertion that a Gaussian is 'the most common model to describe statistical processes in real life,' and the paper does not engage with the substantial literature on election forensics (e.g., distribution-based methods, digit tests, or previous validation studies). Adding such references would help situate the method.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper addresses a politically consequential topic, and the authors have made a good-faith effort to provide open code and data. However, the central quantitative claim is not supported by the method as presented: the null model is unvalidated, the paper's own per-district sigma analysis contradicts it, and the conversion from histogram differences to a vote count is not shown. These are load-bearing problems that cannot be repaired by local revision; the analysis would need to be redone with a credible counterfactual (e.g., validated on past elections or incorporating known covariates). I therefore recommend rejection, while encouraging the authors to pursue a more rigorous framework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: the 175,000-vote estimate is not supported by the paper's own analysis. The Gaussian-uniformity assumption is asserted, not validated, and the per-district sigma result that the authors cite as evidence of manipulation actually shows the model cannot reproduce the official data without assuming a global sigma. So the histogram deficit is a measure of misspecification, not falsification.\n\nWhat the paper does well: it uses official CEC data, shares code and data, documents real anomalies (four precincts with >100% turnout, several with no registered-voter counts, all strongly pro-Georgian Dream), and it is transparent about the 51–52% result under per-district sigmas. Those anomaly observations are worth a look.\n\nWhere it breaks down: the counterfactual is built to reproduce the official 54% by selecting a global sigma in (0,0.1). That makes the simulation a rescaled version of the data, not an independent fair-election baseline. The phase boundaries (Phase II vs III) are chosen after seeing the histogram, so the \"deficit\" is partly a function of the authors' eyeballing. And the conversion from histogram differences to a vote count appears nowhere in the paper. The claim that 'both methods identify 175,000' is misleading: both methods use the same simulation and same filtering, so the agreement is not independent confirmation.\n\nThe most serious problem is the model's premise: precinct vote shares within a district need not be tightly clustered. Urban-rural divides, ethnic composition, and diaspora precincts (the paper itself treats all precincts outside Georgia as one district) can produce exactly the spread the model calls abnormal. The per-district sigma result (51–52%) is the model saying the official data has more heterogeneity than the Gaussian assumption allows; interpreting that as manipulation is circular.\n\nWho is this for? Someone working on election forensics might use the anomaly list and the GitHub code as a starting point, but the paper cannot serve as a reliable estimate of manipulated votes. It deserves a serious referee only if the question is whether the method can be repaired; as submitted, I would not send it to a statistics journal. My recommendation: do not publish as is. Suggest the authors validate their null model on past Georgian elections or on precincts with known clean results, and show the full arithmetic from histogram difference to vote count.\n\nSerious thinker: yes — they are honest about the sigma discrepancy and present their reasoning clearly, even though the conclusion overreaches.","headline":"The 175k estimate is an artifact of an unvalidated Gaussian null model; the paper's own per-district sigma result undercuts it, though the anomaly documentation and shared code are useful.","tokens_in":6303,"tokens_out":2121,"would_cite":false,"duration_ms":19345,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that roughly 175,000 votes were manipulated in Georgia's 2024 parliamentary election, enough to have changed the result if they had been cast normally.","keywords":["computer simulation","statistics","election forensics","vote manipulation","Georgia 2024 parliamentary election","precinct-level data","counterfactual baseline","election integrity"],"falsifier":"Apply the same simulation to a Georgian parliamentary election that all observers accept as free and fair (for example, 2016 or 2020): if the method also finds a discrepancy corresponding to over 100,000 manipulated votes, the baseline is invalid. More directly, regress the official precinct-level Georgian Dream share on urban/rural status, diaspora location, and demographic composition; if the excess mass in the high-share precincts disappears once these covariates are controlled, the anomaly is explained without invoking manipulation.","tokens_in":5282,"feed_emoji":"🗳️","tokens_out":6897,"duration_ms":54771,"temperature":0.7,"pith_summary":"The paper sets out to prove that the official results of Georgia's October 2024 parliamentary election were rigged, with roughly 175,000 votes manipulated in favor of the ruling Georgian Dream party. The authors build a computer simulation that reproduces the party's official 54% overall result under normal voting conditions, then compare the simulated precinct-level vote-share distribution with the official Central Election Commission data. Two bin-width analyses converge on the same most probable number of manipulated votes: 175,000, within ranges of 140,000–200,000 (5% bins) and 90,000–245,000 (1% bins). If the claim is correct, the official outcome did not reflect voter intent, and the election result would have been dramatically different without the manipulation.","feed_headline":"Simulation puts manipulated votes in Georgia at 175,000","feed_subtitle":"Two methods agree on the number of votes that would flip the result.","key_machinery":"The load-bearing object is the simulated counterfactual histogram of precinct-level vote shares for Georgian Dream. Each voter is assigned a probability of voting for the party drawn from a Gaussian distribution whose mean is the official party share in that voter's district; the standard deviation $\\sigma$ is taken as a single global value and is pinned to the interval $0<\\sigma<0.1$ by the condition that the simulated party total comes out at approximately 54%. The experiment runs 2,500 simulated elections, keeps the 803 that give 53–55% for the party, and uses their average histogram as the fair-election baseline. The manipulated-vote estimate is then obtained by measuring the difference between the official histogram and this baseline in the Phase II/III ranges: votes that should appear in precincts around 60% are found instead in precincts with much higher shares.","core_discovery":"The central discovery is a quantitative mismatch between the official precinct-level vote-share distribution and a simulated fair-election distribution with the same overall result. The authors assume that in a fair election voters are assigned to precincts roughly randomly, so a party's share in each precinct should be tightly clustered near its district average; they therefore model each voter's probability of voting for Georgian Dream with a Gaussian distribution and require the party to reach 54% overall, which fixes the standard deviation to a small range $0<\\sigma<0.1$. From 2,500 simulated elections they retain 803 that yield 53–55% for the party, and compare the resulting histogram of precinct shares to official data. They find a deficit of official precincts around 60% (Phase II) and an excess of precincts with very high shares (Phase III), and compute the number of shifted votes as most probably 175,000, with ranges 140,000–200,000 for 5% bins and 90,000–245,000 for 1% bins. They also report that using district-specific standard deviations, the simulation yields only 51–52% for the party, which they interpret as further evidence that precinct-level variability in official data is inconsistent with an unmanipulated election. From this they conclude the election was rigged.","pith_inferences":["A testable next step is to replace the Gaussian baseline with a demographic model that controls for urban/rural and diaspora status; if the excess mass in the high-share precincts is fully explained by those covariates, the 175,000 estimate would measure legitimate heterogeneity rather than fraud.","The same simulation could be run on the 2012, 2016, and 2020 Georgian parliamentary elections cited in the paper; if those accepted-clean elections also flag a large number of manipulated votes, the baseline is invalid, while clean results there would strongly reinforce the 2024 conclusion.","The paper's observation that district-specific standard deviations produce only 51–52% could be converted into a formal statistical significance statement: under fair aggregation the probability of observing the official 54% with such precinct-level scatter is the quantity a referee would want reported as a p-value."],"forward_implications":["If the central claim is right, the official 54% / 46% split is not a valid expression of voter intent: removing roughly 175,000 manipulated votes would change the winner and the overall result.","The convergence of the 5% and 1% bin analyses on the same most probable count (175,000) indicates the anomaly is not an artifact of bin width.","Because stolen votes damage the opposition twice (a vote not cast for the opposition, plus a vote added to the winner), the real effect on the race exceeds the raw 175,000 figure, and the opposition's true support would be higher still.","The method offers a repeatable forensics template: any election with published precinct-level results can be tested against this sort of simulation, provided the no-manipulation baseline is credible."],"supporting_citations":[{"why":"Supplies the probability background that justifies using the Gaussian distribution as the model for precinct-level vote shares.","marker":"[1]"},{"why":"Provides the Edison Research exit-poll baseline, showing that a 13-point gap between exit polls and final results is unprecedented in previous Georgian elections and motivating the investigation.","marker":"[2]"}],"fun_headline_variants":["Simulation finds 175k likely manipulated votes in Georgia","Two methods converge on 175k rigged votes in Georgia","Mathematical evidence points to 175k Georgian votes faked","Statistically, 175k votes in Georgia were manipulated","Data suggests 175k votes flipped in Georgia's election"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that in a fair election voters are distributed across precincts randomly enough that each party's share in a precinct is tightly clustered around the district average; if legitimate urban-rural, demographic, or diaspora differences spread precinct shares widely, the simulated baseline does not represent a fair election.","fun_headline_variants_meta":{"raw":{"variants":["Simulation finds 175k likely manipulated votes in Georgia","Two methods converge on 175k rigged votes in Georgia","Mathematical evidence points to 175k Georgian votes faked","Statistically, 175k votes in Georgia were manipulated","Data suggests 175k votes flipped in Georgia's election"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1434,"prompt_tokens":911,"completion_tokens":523,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":450}},"tokens_in":527,"tokens_out":523,"duration_ms":4685,"temperature":1.0,"reasoning_tokens":450,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:10:41.088382+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same simulation to a Georgian parliamentary election that all observers accept as free and fair (for example, 2016 or 2020): if the method also finds a discrepancy corresponding to over 100,000 manipulated votes, the baseline is invalid. More directly, regress the official precinct-level Georgian Dream share on urban/rural status, diaspora location, and demographic composition; if the excess mass in the high-share precincts disappears once these covariates are controlled, the anomaly is explained without invoking manipulation.","supporting_citations":[{"cited_title":"Tsitsiklis","cited_arxiv_id":null,"evidence_quote":"Supplies the probability background that justifies using the Gaussian distribution as the model for precinct-level vote shares."},{"cited_title":"Edison Research 2024 Republic of Georgia Exit Poll - Edison Research","cited_arxiv_id":null,"evidence_quote":"Provides the Edison Research exit-poll baseline, showing that a 13-point gap between exit polls and final results is unprecedented in previous Georgian elections and motivating the investigation."}],"review_version":1}