REVIEW 4 major objections 5 minor 3 references
Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that Google's text and image search results during the 2023 Swiss federal election campaign were gendered, with men receiving more news media links and women receiving more stereotypically pleasant images, and that these…
desk verdict The descriptive gender-bias audit is solid and worth publishing; the electoral-performance 'prediction' is an in-sample correlation and should be toned down. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two instruments carry the argument. The first is media source prominence, a rank-weighted share of news media domains defined as $\sum_{i\in\text{media}} 1/r_i$ over $\sum_{j=1}^N 1/r_j$, where $r_i$ is the rank of the $i$-th organic result; it converts Google's source mix and ordering into a single candidate-level visibility score. The second is a set of visual-stereotyping measures produced by a commercial computer-vision API from over a million returned images: share of women depicted, average smile probability, average positive affect (happiness plus calmness), and average negative affect (fear, anger, sadness, disgust). These measures are aggregated per candidate per wave and entered into mixed-effects regressions with by-canton random intercepts, so the audit infrastructure, which used virtual machines with Swiss IPs and cookie-cleared browser agents querying google.ch in two waves, makes the comparison controlled and the link to official election data possible.
What would settle it
Re-run the hierarchical regressions with the candidates' personal vote share from the 2019 Swiss federal election, or an equivalent prior-popularity measure, included as a control; if the 6 to 8 percent variance attributed to text and image search measures collapses toward zero, the predictive claim is an artifact of successful candidates generating more search results.
Extended reading notes
Core claim
The paper's central claim is that Google's selection and ranking of candidate information during the 2023 Swiss Federal Elections was systematically different for men and women, and that the resulting search-representation measures were associated with actual voting outcomes. In text results, media source prominence was higher for men candidates (b = -3.94 in wave 1; b = -2.72 in wave 2), meaning men received more news media links and better placement. In image results, women made up about 38.7% to 39.1% of depicted persons versus 42.5% of candidates on official lists, and queries for women returned images with more smiling (about 5.4 percentage points more), more positive affect (b = 1.82 and 2.22), and less negative affect (b = -1.05 and -1.08). Hierarchical regressions with by-canton random intercepts found that text and image search blocks together added 6 to 8 percent explained variance to models of log-transformed personal votes, with media source prominence positively predicting votes and, in wave 1, a significant interaction showing women benefited more from media prominence than men. The paper does not claim that negative affect is punished for women; that hypothesis was not supported.
Load-bearing premise
The load-bearing premise is that the observed correlation between Google search representation and personal votes is not driven by reverse causality, meaning that popular candidates are not simply generating more and more positive search results; the models do not control for prior vote share, offline media coverage, or campaign spending.
Editorial extensions
If this is right
- Women candidates start from a lower base of rank-weighted news media visibility, and because media source prominence is positively associated with personal votes, the gender gap in Google text results implies a corresponding electoral disadvantage.
- The consistent gender-by-party pattern in image output means that women from conservative parties are most exposed to stereotypically positive and smiling portrayals, making the interaction between gender and party part of the algorithmic-curation story.
- Search-based measures explain 6 to 8 percent of variance in personal votes with the included controls, a magnitude the paper argues is politically consequential in elections decided by a few percentage points.
- The null result for the negative-affect penalty indicates that not every gendered stereotype translates into an electoral punishment, so the double bind for women candidates is more conditional than the visual stereotype literature might suggest.
- The two-wave design shows the main gendered patterns in text and image output were stable across the final month of the campaign, which points to a persistent, not transient, algorithmic environment.
Reading between the lines
- The most serious rival explanation, not settled by the paper, is reverse causality: candidates who are expected to win attract more media coverage and more image material, so the 6 to 8 percent predictive variance may partly be popularity rather than search-engine influence. A natural extension is to add prior vote share or campaign spending to the models and see whether the search coefficients su
- Because the image results may mirror candidates' own campaign photography and party visual strategy, an editorial inference is that comparing Google image output with the candidates' official portraits or social-media feeds would separate algorithmic stereotyping from self-presentation.
- A randomized or natural-experiment perturbation of search rankings, for example a temporary change in a candidate's media prominence that is unrelated to their popularity, could convert the predictive association into a causal test; the paper's design is observational and does not itself identify causal direction.
- If the reverse-causality concern is borne out, regulators would need to focus on whether Google amplifies existing inequality rather than creates it; the paper leaves that distinction explicitly open in its discussion.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a large-scale algorithm audit of Google text and image search results for all 5,883 candidates in the 2023 Swiss federal elections, with data collected three weeks and one week before election day. It finds that text search results give men candidates higher media source prominence than women candidates, and that image search results for women show more smiling, more positive affect, and less negative affect, especially for right-leaning candidates. The paper further claims that these search representation measures are predictive of electoral performance, with search-derived variables adding 6–8% of explained variance to models of candidates' log-transformed personal votes.
Significance. If the descriptive findings hold, the study provides valuable large-scale evidence on algorithmic gender bias in political search output, extending prior work on media visibility and visual stereotyping to a full-candidate audit in a proportional electoral context. The audit design is a strength: it covers all candidates, uses two temporally separated waves, employs virtual agents to reduce personalization, and combines manual source coding with computer-vision affect measurement. The gender differences in media source prominence (b = -3.94 and -2.72) and in positive/negative affect (b = 1.82/2.22 and -1.05/-1.08) are consistent across waves and reported with confidence intervals. The weakest part is the electoral-performance analysis: the cross-sectional regressions cannot distinguish search-engine influence from reverse causality or omitted confounders, so the 'predictive' claim in the abstract and H2 is not currently supported.
major comments (4)
- [Search engines and electoral results; Discussion] The claim that search-output measures are 'predictive of electoral performance' (Abstract; Table 2) is not identified as a causal or predictive effect. The hierarchical regressions control gender, party, incumbency, list position, and canton random intercepts, but not prior vote share, offline media coverage volume, or campaign spending. Since the paper itself states that 'the presence in traditional media is a necessary condition for media source prominence' (Section 'Text searches: Algorithmic curation and media source prominence'), the news-media-prominence coefficients (Table 2; wave 1 b = 0.27) plausibly reflect pre-existing candidate popularity and news coverage rather than search-engine influence on voters. The Discussion acknowledges that distinguishing amplification from distortion is 'essential but challenging', but the Abstract's 'predictive' wording and the H2 formulation go beyond what the cross-sectional design can establish. Please either add controls or robustness checks (e.g., prior election results, offline coverage volume) or reframe the claim as an association.
- [Table 2 vs. Results text] There is a direct inconsistency between the text and Table 2 for the wave-2 media source prominence effect. The text reports b = 1.34, 95% CI [1.001, 1.58] for media prominence in wave 2, but Table 2 lists 'News media' as b = 0.37 (SE = 0.11) and 'Social media' as b = 1.34 (SE = 0.18). Please correct the reported coefficient/CI or the table, and ensure the interpretation (news media vs. social media) is consistent.
- [Methodology/Data analysis; Table 2] The sample size is inconsistent. The methods state that each wave contains n = 5,883 candidates (Table 1, k = 5,883), but Table 2 reports 5,952 and 5,948 observations for the two waves. Please clarify whether some candidates are represented by multiple agent-level records after aggregation, or whether the N in Table 2 should be 5,883; this affects degrees of freedom and the reported fit statistics.
- [Search engines and electoral results] The 'predictive' language is not supported by out-of-sample evaluation. The likelihood-ratio tests and Δ marginal R² values are computed on the same data used to fit the models; they quantify in-sample explanatory power, not predictive performance. If the authors intend 'predictive' in the forecasting sense, cross-validation or a temporal holdout is needed; otherwise, terms like 'associated with' or 'explained variance' should be used consistently.
minor comments (5)
- [Introduction] The introduction states that 'text search output included more and higher rank links to media sources for queries of women candidates', which contradicts the Results, Abstract, and Discussion, where women have lower media source prominence; please correct this sentence.
- [Image search results] The positive-affect finding is attributed to H4a, but H4a is about smiling; H4b is about more positive affect and H4c about less negative affect. Please fix the hypothesis labels in this section.
- [Methodology/Data analysis strategy] The models predicting log-transformed personal votes are described as 'generalised linear mixed-effects regression models'; unless a non-Gaussian family and link function are specified, these should be described as linear mixed-effects models.
- [Online Appendix, Table S4] Table S4 contains a duplicated column header ('affect positive'); please clean up the table formatting.
- [Data availability] No data or code availability statement is included; please add one or state that materials are available upon request, as this is standard for algorithm audits.
Circularity Check
No significant circularity: the gender/party differences are measured audit outputs, and the 'predictive' performance claim is an in-sample regression fit, not a derivation that reduces to its own inputs.
full rationale
This paper is an empirical algorithm audit rather than a first-principles derivation. The gender and party differences in Google text and image results are directly measured from the collected search outputs (e.g., media-source prominence differences of b = -3.94 and b = -2.72; image affect differences), so they are observations, not fitted targets. The central 'predictive of electoral performance' claim (Abstract; 'Search engines and electoral results', Table 2) is based on hierarchical mixed-effects regressions fitted to the same Swiss 2023 election outcomes; the reported 6-8% explained variance is an in-sample model-fit increment, not an out-of-sample forecast. This is a presentation limitation, but it is not circular by construction: the coefficients and likelihood-ratio tests could have come out null, and no parameter is defined in terms of the outcome it is said to predict. The Discussion explicitly acknowledges the harder identification issue: it is 'unclear whether these algorithms merely amplify existing social biases or create distortions in social reality' and that 'Establishing a robust baseline for search engine curation each politician... is essential but challenging.' The cited prior work by the same authors (Rohrbach et al., 2024; Makhortykh et al., 2025a/b) supplies background and earlier audit evidence, but it is not invoked as a load-bearing uniqueness theorem or an ansatz that forces the present results. A separate textual inconsistency exists (the Introduction says women received more media links while the Results say men did), but that is an editing error, not a circular step. Overall, the audit's measured outputs are self-contained, and the predictive language, while not a genuine forecast, does not make the derivation circular.
Assumptions & free parameters
free parameters (3)
- Inverse rank weighting in media source prominence =
1/rank
- Affect valence groupings =
positive = happiness + calmness; negative = fear + anger + sadness + disgust
- Most prominent face smile score =
probability score for the single most prominent face
assumptions (5)
- domain assumption Google search results are a relevant source of political information for voters in Switzerland.
- domain assumption Virtual agents with cleared cookies and static IPs produce search results representative of ordinary Swiss voters.
- domain assumption Amazon Rekognition's emotion labels accurately measure displayed emotions in political images.
- domain assumption The single student assistant's manual coding of unknown domains is reliable.
- domain assumption The regression controls (incumbency, list position, party, canton, gender) are sufficient to prevent reverse causality from explaining the search-vote association.
Cite this review
Pith. "Pith review of Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023." pith.science (2026). https://pith.science/paper/MWPAQ3LX
@misc{pith2026250706018,
author = {Pith},
title = {Pith review of: Campaigning through the lens of Google: A large-scale algorithm audit of Google searches in the run-up to the Swiss Federal Elections 2023},
year = {2026},
howpublished = {\url{https://pith.science/paper/MWPAQ3LX}},
note = {Machine review of arXiv:2507.06018}
}
read the original abstract
Search engines like Google have become major sources of information for voters during election campaigns. To assess potential biases across candidates' gender and partisan identities in the algorithmic curation of candidate information, we conducted a large-scale algorithm audit analyzing Google's selection and ranking of information about candidates for the 2023 Swiss Federal Elections, three and one week before the election day. Results indicate that text searches prioritize media sources in search output but less so for women politicians. Image searches revealed a tendency to reinforce stereotypes about women candidates, marked by a disproportionate focus on stereotypically pleasant emotions for women, particularly among right-leaning candidates. Crucially, we find that patterns of candidates' representation in Google text and image searches are predictive of their electoral performance.
Reference graph
Works this paper leans on
-
[1]
Aaldering, L., Van der Meer, T., & Van der Brug, W. (2018). Mediated Leader Effects: The Impact of Newspapers’ Portrayal of Party Leadership on Electoral Support. The International Journal of Press/Politics , 23 (1), 70–94. https://doi.org/10.1177/1940161217740696 Bandy, J., & Diakopoulos, N. (2020). Auditing News Curation Systems: A Case Study Examining ...
arXiv 2018
-
[12]
Foreign beauties want to meet you
Newman, N., Fletcher, R., Eddy, K., Robertson, C. T., & Nielsen, R. K. (2023). Reuters Institute digital news report 2023 . Reuters Institute for the study of Journalism. https://reutersinstitute.politics.ox.ac.uk/sites/default/files/2023-06/Digital_News_Repo rt_2023.pdf Norocel, O. C., & Lewandowski, D. (2023). Google, data voids, and the dynamics of the...
arXiv 2023
-
[33]
https://doi.org/10.1007/s42001-025-00361-3 Makhortykh, M., Rohrbach, T., Sydorova, M., & Kuznetsova, E. (2025b). Search engines in polarized media environment: Auditing political information curation on Google and Bing prior to 2024 US elections. arXiv preprint arXiv:2501.04763. Mittelstadt, B. (2016). Automation, algorithms, and politics| auditing for tr...
work page Pith review arXiv 2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.