Pith. sign in

REVIEW 3 major objections 5 minor 71 references

Learn, Explore and Reflect by Chatting: Understanding the Value of an LLM-Based Voting Advice Application Chatbot

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read An LLM-based chatbot can make voting advice applications clearer, more flexible, and more reflective, a 331-user field study suggests.

desk verdict A genuinely first field deployment of a conversational LLM-based VAA chatbot; the usability and reflection findings are solid, but the knowledge-gain claim rests on self-report and no one checked whether the chatbot's facts were right. read the letter →

arxiv 2505.09806 v1 pith:VO4E5QKU submitted 2025-05-14 cs.HC cs.CY

classification cs.HCcs.CY
keywords votingadviceapplicationsLLMchatbotciviceducationdeliberationtrustinAIpoliticalknowledgeconversationaluserinterfacesEuropeanParliamentelection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a chatbot powered by a large language model can fix two central usability problems of voting advice applications (VAAs): their reliance on complex political jargon and their rigid click-through format. In a field deployment with 331 German voters before the 2024 European Parliament election, participants described the chatbot as intuitive and informative, valuing its plain-language explanations and the freedom to ask follow-up questions at any point. Beyond accessibility, the paper claims the conversational format acted as a catalyst for reflection and rationalization: users were prompted to ask themselves why they held a position, not just which party matched it. The significance, if the finding holds, is that the same tool that helps voters compare parties can also exercise their political judgment, and that design choices about transparency will decide whether they trust it.

What carries the argument

The load-bearing mechanism is the hybrid interaction design of the prototype chatbot, which combines an unstructured question-and-answer phase with a structured, VAA-style opinion survey. In the unstructured phase, users ask any election-related question and can follow up until a topic is clear; in the structured phase, they choose parties and topics, then answer ten policy statements scored $+1$ for agreement, $-1$ for disagreement, and $0$ for neutrality against each party's platform, ending in a ranked list. That design is implemented with a prompt-engineered commercial large language model, deliberately without retrieval or fine-tuning, so the conversational affordances—plain-language explanations, two-way comparisons, invitations to go deeper—carry the observed effects. The mixed-method apparatus of surveys, chat logs, and ten follow-up interviews is the evidentiary machinery that connects those affordances to the claims about accessibility, reflection, and trust.

What would settle it

A randomized controlled trial with an objective knowledge test, such as factual questions about party positions before and after use, comparing the LLM chatbot with a traditional click-based VAA: if the chatbot group shows no greater objective knowledge gain than the click-based group, the paper's claim that conversation enhances political knowledge collapses. A second, cheaper check is to fact-check a random sample of the recorded chat logs against party manifestos and official election sources; a substantial share of wrong statements would show that the perceived accuracy users reported is not justified trust.

Watch

Extended reading notes

Core claim

The central discovery is that an LLM-based VAA chatbot does not merely repackage the same information in friendlier prose; it changes what users do with the information. Participants who had found traditional VAAs overwhelming engaged with the chatbot through open questions, on-demand party comparisons, and a turn-by-turn opinion survey that asked them to justify stances. The reported effects split into two layers: an accessibility gain, with lower-education and moderately low-self-efficacy users more likely to report learning something, and a cognitive-engagement gain, with users describing curiosity-driven exploration and a felt obligation to reason about their own answers. The paper also documents a trust paradox: users were broadly aware that LLMs can fabricate and slant information, yet most trusted this chatbot's plausible, neutral-sounding output, and nearly all wanted source citations and third-party validation before relying on it in an election.

Load-bearing premise

The central claim rests on participants' self-reported political knowledge gain, measured with a single seven-point survey question, with no objective knowledge test, no control group, and no verification that the chatbot's answers were factually correct; if perceived understanding diverges from real understanding, the claim that the chatbot enhances political knowledge is unsupported.

Editorial extensions

If this is right

  • If the accessibility finding holds, a conversational interface could extend the reach of VAAs to voters who are put off by jargon and long text walls, including people with less formal education.
  • If the reflection finding holds, VAA chatbots could be designed not just to match voters to parties but to elicit reasons, making the tool serve deliberative rather than purely preference-aggregating models of democracy.
  • The trust findings imply that a public-facing VAA chatbot needs source citations, disclosure of training data and developers, and third-party certification; without these, even satisfied users will withhold full trust.
  • The sycophancy observation implies a design trade-off: a chatbot that flatters users' existing views risks creating an echo chamber, so future systems need to introduce counterarguments without being perceived as biased.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because political knowledge gain was measured with one self-report item, a fair test of the paper's strongest claim would be a randomized comparison with an objective pre/post factual quiz; the paper's results do not by themselves establish that understanding actually improved.
  • The reflection and rationalization finding suggests a concrete hypothesis for follow-up work: conversational VAAs should increase the deliberative quality of a vote choice relative to click-based VAAs, which is testable with think-aloud protocols or post-decision justification tasks.
  • The pattern that less-educated and moderately low-efficacy users reported greater gains may partly reflect the sample's demographic skew, so the accessibility benefit for the general electorate remains an open question rather than a measured fact.
  • A useful engineering extension implied by the trust findings is a retrieval-augmented version of the chatbot that grounds every claim in a primary source, allowing a direct test of whether citation affordances increase justified trust without reducing perceived neutrality.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports an exploratory mixed-method deployment of an LLM-based voting advice application (VAA) chatbot before the 2024 European Parliament election in Germany. A total of 331 Prolific participants interacted with a custom GPT-4o chatbot and completed survey questionnaires, and 10 participants were interviewed afterwards. The paper argues that the chatbot improves accessibility compared with traditional VAAs through simpler language and on-demand clarification, that the conversational format fosters curiosity-driven exploration, reflection, and rationalization, and that users' trust is tempered by concerns about truthfulness, transparency, and bias. It concludes with design recommendations for future VAA chatbots. The central contribution is the claim that such a chatbot can 'enhance political knowledge' and serve as a 'catalyst for reflection and rationalization.'

Significance. Strengths: the work is a timely, real-world deployment with a relatively large sample for a CUI study; it combines surveys, interaction logs, and interviews; and it includes the full survey instrument, system prompts, and classification prompt in appendices, which supports replication and reuse. The qualitative findings on reflection, rationalization, and trust are coherent and well supported by interview excerpts, and the design recommendations are actionable. If the knowledge-enhancement claim were supported by objective evidence, the paper would make a strong contribution to civic education and conversational AI. As it stands, the self-report nature of the knowledge-gain measure and the absence of any audit of the chatbot's factual accuracy substantially narrow the significance: the paper convincingly shows that users felt informed and engaged, but it does not yet show that the chatbot actually improved their political knowledge or that its information was reliable.

major comments (3)
  1. [Section 4.1.4 and Appendix A] The paper's headline claim that the chatbot can 'enhance political knowledge' (abstract, contributions, Section 5) is measured with a single 7-point Likert item ('By using the chatbot, I have gained more understanding of the political landscape') and no objective pre/post knowledge test or control condition. Self-reported understanding can diverge substantially from actual understanding, especially in a task where users are motivated to report a positive experience. This item cannot support the causal and objective wording used in the contributions. Please reframe all knowledge-related conclusions as 'perceived knowledge gain' or supplement them with objective evidence such as a factual-knowledge quiz or a verification of the information delivered in the conversation logs.
  2. [Sections 3.2.1, 4.3.3, and Appendix D] The 'informative' claim presupposes that the chatbot's answers were factually accurate, but the paper never audits the recorded conversation logs against ground truth. The system prompt only instructs the model to 'suggest websites' when unsure about facts; there is no retrieval grounding, no answer verification, and no reported fact-check. The interview evidence that users found the answers plausible and recalled no errors is not sufficient, because the paper itself shows that users can be unable to detect inaccuracies (P3 accepted outputs at face value until reminded). If a nontrivial fraction of the GPT-4o answers were wrong, the informational contribution would be unsupported or misleading. I recommend either adding a factuality audit of a sample of the logged answers or explicitly limiting the claims to perceived informativeness.
  3. [Section 3.3 and Section 5.3] The external validity of the quantitative findings is constrained by a sample that is younger, more left-leaning, and much more experienced with both VAAs (88% had used Wahl-O-Mat) and LLM chatbots than the general electorate. Although the limitations section acknowledges the age issue, the abstract and discussion present the regression results on lower-educated and less politically efficacious users as if they characterize those population segments. These findings should be framed as hypothesis-generating observations from a convenience sample, and the corresponding claims in the discussion should be tempered accordingly.
minor comments (5)
  1. [Section 5.2.1] The word 'receommendation' should be 'recommendation'.
  2. [Appendix B] The category label 'IRRELEV ANT' should be 'IRRELEVANT'.
  3. [Section 4.3.2] In the sentence beginning 'As participants were curious about the people behind the chatbot (Section 4.3.2),' the following clause starts with a capital 'This may translate to...' and should be lowercase.
  4. [Section 3.2.2 and Figure 1] The chatbot's name 'ChatEP2024' appears only in Appendix D; please introduce it when the prototype is first described in the main text.
  5. [Section 4.1.4] The ordinal regression section reports odds ratios for four models selected by AIC; it would help to state explicitly how the Benjamini-Hochberg correction was applied across the multiple models and outcome variables.

Circularity Check

0 steps flagged · score 0.0 of 10

No derivational circularity: findings are empirical self-reports, not fit-forced or self-referential claims.

full rationale

This paper is an empirical mixed-method user study, not a derivation. The central claims—that participants found the LLM-based VAA chatbot intuitive, informative, and reflective—rest on survey responses, conversation logs, and follow-up interviews. There are no fitted constants, normalization choices, uniqueness theorems, or self-citation chains that force the conclusions. The perceived political knowledge gain is measured by a single 7-point Likert self-report item (Section 4.1.4, Appendix A) and is never validated against an objective knowledge test, and the paper itself acknowledges that chatbot accuracy was not audited; however, this is a validity or correctness limitation, not circularity. Using GPT-4o both as the chatbot and as the zero-shot question classifier (Section 3.6.1) is a methodological choice that does not make the user-experience findings equivalent to their inputs by construction. The authors cite no prior work of their own as load-bearing evidence, and the design recommendations are grounded in participants' stated desires rather than imported from an unverified self-citation. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities; this is an observational study. The load-bearing premises are the validity of self-reported outcomes, the representativeness of the Prolific sample, and the unverified accuracy of GPT-4o answers during the study.

assumptions (4)
  • domain assumption Self-reported Likert items measure actual political knowledge and motivation gains.
    The central quantitative outcome is a single self-report item, with no objective knowledge test or behavioral outcome. This enters at Section 4.1.4 and Appendix A.
  • domain assumption GPT-4o's answers during the field study were sufficiently accurate that user perceptions reflect the LLM's capabilities rather than an unusually error-prone prototype.
    No expert verification of chatbot outputs was reported; interviewees recalled no errors, but logs were not systematically fact-checked. This underlies all findings about perceived usefulness and trust.
  • domain assumption Prolific panelists with 88% prior Wahl-O-Mat use provide a useful signal about general voter experience.
    The sample skews young, left, and experienced with VAAs; generalizability to less-experienced voters is asserted rather than demonstrated. Acknowledged in Section 5.3.
  • domain assumption Thematic analysis by two researchers yields stable themes.
    Standard qualitative method, but inter-coder reliability is not reported. This underlies the qualitative findings in Sections 4.1 through 4.3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learn, Explore and Reflect by Chatting: Understanding the Value of an LLM-Based Voting Advice Application Chatbot." pith.science (2026). https://pith.science/paper/VO4E5QKU

@misc{pith2026250509806,
  author       = {Pith},
  title        = {Pith review of: Learn, Explore and Reflect by Chatting: Understanding the Value of an LLM-Based Voting Advice Application Chatbot},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VO4E5QKU}},
  note         = {Machine review of arXiv:2505.09806}
}
read the original abstract

Voting advice applications (VAAs), which have become increasingly prominent in European elections, are seen as a successful tool for boosting electorates' political knowledge and engagement. However, VAAs' complex language and rigid presentation constrain their utility to less-sophisticated voters. While previous work enhanced VAAs' click-based interaction with scripted explanations, a conversational chatbot's potential for tailored discussion and deliberate political decision-making remains untapped. Our exploratory mixed-method study investigates how LLM-based chatbots can support voting preparation. We deployed a VAA chatbot to 331 users before Germany's 2024 European Parliament election, gathering insights from surveys, conversation logs, and 10 follow-up interviews. Participants found the VAA chatbot intuitive and informative, citing its simple language and flexible interaction. We further uncovered VAA chatbots' role as a catalyst for reflection and rationalization. Expanding on participants' desire for transparency, we provide design recommendations for building interactive and trustworthy VAA chatbots.

Figures

Figures reproduced from arXiv: 2505.09806 by the authors.

Figure 1
Figure 1. Examples of guidance from the chatbot in English. The actual user study was performed in German. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Conceptual walk-through of the interaction with the chatbot (as explained in section 3.2.2). [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Overall flow of the study (as explained in Section 3.4). [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 31 canonical work pages

  1. [1]

    Julia Angwin, Alondra Nelson, and Rina Palta. 2024. Seeking Re- liable Election Information? Don’t Trust AI . Technical Report. The AI Democracy Projects. https://www.ias.edu/sites/default/files/AIDP_ SeekingReliableElectionInformation-DontTrustAI_2024.pdf

  2. [2]

    Anthropic. 2024. Testing and mitigating elections-related risks. https://www. anthropic.com/news/testing-and-mitigating-elections-related-risks

  3. [3]

    Voelkel, Johannes Christopher Eichstaedt, and Robb Willer

    Hui Bai, Jan G. Voelkel, Johannes Christopher Eichstaedt, and Robb Willer

  4. [4]

    Yoav Benjamini and Yosef Hochberg. 1995. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society. Series B (Methodological) 57, 1 (1995), 289–300. https://www. jstor.org/stable/2346101 Publisher: Royal Statistical Society, Oxford University Press

  5. [5]

    Maximilian Bleick, Nils Feldhus, Aljoscha Burchardt, and Sebastian Möller. 2024. German Voter Personas Can Radicalize LLM Chatbots via the Echo Chamber Effect. In Proceedings of the 17th International Natural Language Generation Conference, Saad Mahamood, Nguyen Le Minh, and Daphne Ippolito (Eds.). Association for Computational Linguistics, Tokyo, Japan, ...

  6. [6]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology 3, 2 (Jan. 2006), 77–101. https://doi.org/10. 1191/1478088706qp063oa

  7. [7]

    John Brooke. 1996. SUS: A ’Quick and Dirty’ Usability Scale. In Usability Evaluation In Industry. CRC Press. Num Pages: 6

  8. [8]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proc. ACM Hum.-Comput. Interact. 5, CSCW1 (April 2021), 188:1–188:21. https://doi.org/10.1145/3449287

Show all 71 references
  1. [9]

    Ilias Chalkidis and Stephanie Brandl. 2024. Llama meets EU: Investigating the European political spectrum through the lens of LLMs. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies...

  2. [11]

    Lu Ding, Tong Li, Shiyan Jiang, and Albert Gapud. 2023. Students’ perceptions of using ChatGPT in a physics class as a virtual tutor. International Journal of Educational Technology in Higher Education 20, 1 (Dec. 2023), 63. https: //doi.org/10.1186/s41239-023-00434-1

  3. [12]

    Esin Durmus, Liane Lovitt, Alex Tamkin, Stuart Ritchie, Jack Clark, and Deep Ganguli. 2024. Measuring the Persuasiveness of Language Models. https: //www.anthropic.com/news/measuring-model-persuasiveness

  4. [13]

    Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. In Proceedings of the 61st Annual Meeting of the Association for Computation...

  5. [14]

    Trech- sel, and Diego Garzia

    Frederico Ferreira Da Silva, Andres Reiljan, Lorenzo Cicchi, Alexander H. Trech- sel, and Diego Garzia. 2023. Three sides of the same coin? comparing party positions in VAAs, expert surveys and manifesto data.Journal of European Public Policy 30, 1 (Jan. 2023), 150–173. https:...

  6. [15]

    Thomas Fossen and Joel Anderson. 2014. What’s the point of voting advice applications? Competing perspectives on democracy and citizenship. Electoral Studies 36 (Dec. 2014), 244–251. https://doi.org/10.1016/j.electstud.2014.04.001

  7. [16]

    Diego Garzia and Stefan Marschall. 2019. Voting Advice Applications. In Oxford Research Encyclopedia of Politics . Oxford University Press. https://doi.org/10. 1093/acrefore/9780190228637.013.620

  8. [17]

    Micha Germann, Fernando Mendez, and Kostas Gemenis. 2023. Do Voting Advice Applications Affect Party Preferences? Evidence from Field Experiments in Five European Countries. Political Communication 40, 5 (Sept. 2023), 596–614. https://doi.org/10.1080/10584609.2023.2181896 Publ...

  9. [18]

    Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz

  10. [19]

    Kobi Hackenburg and Helen Margetts. 2024. Evaluating the persuasive influence of political microtargeting with large language models.Proceedings of the National Academy of Sciences 121, 24 (June 2024), e2403116121. https://doi.org/10.1073/ pnas.2403116121 Publisher: Proceeding...

  11. [20]

    Rafik Hadfi and Takayuki Ito. 2022. Augmented Democratic Deliberation: Can Conversational Agents Boost Deliberation in Social Media?. In Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems (AAMAS ’22). International Foundation for Auton...

  12. [21]

    Stef Hankel, Christine Liebrecht, and Naomi Kamoen. 2024. ‘Hi Chatbot, let’s Talk about Politics!’ Examining the Impact of Verbal Anthropomorphism in Conversational Agent Voting Advice Applications (CAVAAs) on Higher and Lower Politically Sophisticated Users. Interacting with ...

  13. [22]

    Sandra G. Hart. 1986. NASA Task Load Index (TLX). https://ntrs.nasa.gov/ citations/20000021488 NTRS Author Affiliations: NASA Ames Research Center NTRS Document ID: 20000021488 NTRS Research Center: Ames Research Center (ARC)

  14. [23]

    Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. The po- litical ideology of conversational AI: Converging evidence on ChatGPT’s pro- environmental, left-libertarian orientation. https://doi.org/10.48550/arXiv.2301. 01768 arXiv:2301.01768 [cs]

  15. [24]

    Clara Helming, Angela Müller, Matthias Spielkamp, Anna Lena Schiller, and Waldemar Kesler. 2023. Generative AI and elections: Are chat- bots a reliable source of information for voters? Technical Report. https://algorithmwatch.org/en/wp-content/uploads/2023/12/AlgorithmWatch_ ...

  16. [25]

    Samuel Holmes, Anne Moorhead, Raymond Bond, Huiru Zheng, Vivien Coates, and Michael Mctear. 2019. Usability testing of a healthcare chatbot: Can we use conventional methods to assess conversational user interfaces?. InProceedings of the 31st European Conference on Cognitive Er...

  17. [26]

    Yiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu, Munmun De Choudhury, and Srijan Kumar. 2024. Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries. In Proceedings of the ACM Web Conference 2024 (WWW ’24) . Association for Computin...

  18. [27]

    Seth Jolly, Ryan Bakker, Liesbet Hooghe, Gary Marks, Jonathan Polk, Jan Rovny, Marco Steenbergen, and Milada Anna Vachudova. 2022. Chapel Hill Expert Survey trend file, 1999–2019. Electoral Studies 75 (Feb. 2022), 102420. https: //doi.org/10.1016/j.electstud.2021.102420

  19. [28]

    Bernhard Jordan, Laura Koesten, and Torsten Möller. 2023. Chatting About Data - Interacting with Voice Interfaces to Engage with Election Panel Data. In Proceedings of the 5th International Conference on Conversational User Interfaces (CUI ’23) . Association for Computing Mach...

  20. [29]

    Naomi Kamoen and Bregje Holleman. 2017. I don’t get it. Response difficulties in answering political attitude statements in Voting Advice Applications. Survey Research Methods Vol 11 (Aug. 2017), 125–140 Pages. https://doi.org/10.18148/ SRM/2017.V11I2.6728 Artwork Size: 125-14...

  21. [30]

    Naomi Kamoen and Christine Liebrecht. 2022. I Need a CAVAA: How Con- versational Agent Voting Advice Applications (CAVAAs) Affect Users’ Political Knowledge and Tool Experience. Frontiers in Artificial Intelligence 5 (2022). https://www.frontiersin.org/articles/10.3389/frai.20...

  22. [31]

    Yoonsu Kim, Jueon Lee, Seoyoung Kim, Jaehyuk Park, and Juho Kim. 2024. Un- derstanding Users’ Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level. InProceedings of the 29th International Conference on Intelligent User Interfaces ...

  23. [32]

    Johann Laux, Sandra Wachter, and Brent Mittelstadt. 2024. Trustwor- thy artificial intelligence and the European Union AI act: On the con- flation of trustworthiness and acceptability of risk. Regulation & Gov- ernance 18, 1 (2024), 3–32. https://doi.org/10.1111/rego.12512 _ep...

  24. [33]

    Hao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson. 2025. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. In...

  25. [34]

    Lewis and Jeff Sauro

    James R. Lewis and Jeff Sauro. 2018. Item benchmarks for the system usability scale. J. Usability Studies 13, 3 (May 2018), 158–167

  26. [35]

    Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Zhu et al. Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval- Augmented Generation for Knowledge-Intensive NLP Tasks. In ...

  27. [36]

    Shyam Sundar

    Q.Vera Liao and S. Shyam Sundar. 2022. Designing for Responsible Trust in AI Systems: A Communication Perspective. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22) . Association for Computing Machinery, New York, NY, USA, 1257...

  28. [37]

    Zhuoran Lu, Dakuo Wang, and Ming Yin. 2024. Does More Advice Help? The Effects of Second Opinions in AI-Assisted Decision Making. Proceedings of the ACM on Human-Computer Interaction 8, CSCW1 (April 2024), 1–31. https: //doi.org/10.1145/3653708

  29. [38]

    Claudia López, Alexandra Davidoff, Francisca Luco, Mónica Humeres, and Teresa Correa. 2024. Users’ Experiences of Algorithm-Mediated Public Ser- vices: Folk Theories, Trust, and Strategies in the Global South. Interna- tional Journal of Human–Computer Interaction 0, 0 (2024), ...

  30. [39]

    José Alberto Mancera Andrade and Luis Terán. 2024. From GenAI to Political Profiling Avatars: A Data-Driven Approach to Crafting Virtual Experts for Voting Advice Applications. In Proceedings of the 25th Annual International Conference on Digital Government Research (dg.o ’24)...

  31. [40]

    McKee, Verena Rieser, and Iason Gabriel

    Arianna Manzini, Geoff Keeling, Nahema Marchal, Kevin R. McKee, Verena Rieser, and Iason Gabriel. 2024. Should Users Trust Advanced AI Assistants? Justified Trust As a Function of Competence and Alignment. In The 2024 ACM Conference on Fairness, Accountability, and Transparenc...

  32. [41]

    Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2023. More human than human: measuring ChatGPT political bias. Public Choice (Aug. 2023). https: //doi.org/10.1007/s11127-023-01097-2

  33. [42]

    Simon Munzert, Pablo BarberÁ, Andrew Guess, and JungHwan Yang. 2021. Do Online Voter Guides Empower Citizens? Public Opinion Quarterly 84, 3 (Jan. 2021), 675–698. https://doi.org/10.1093/poq/nfaa037

  34. [43]

    Simon Munzert and Sebastian Ramirez-Ruiz. 2021. Meta-Analysis of the Effects of Voting Advice Applications. Political Communication 38, 6 (Nov. 2021), 691–706. https://doi.org/10.1080/10584609.2020.1843572

  35. [44]

    Diana C. Mutz. 2006. Hearing the Other Side: Deliberative versus Participa- tory Democracy (1 ed.). Cambridge University Press. https://doi.org/10.1017/ CBO9780511617201

  36. [45]

    OpenAI. 2024. GPT-4o System Card. Technical Report. https://cdn.openai.com/ gpt-4o-system-card.pdf

  37. [46]

    Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red Teaming Language Models with Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing...

  38. [47]

    Leon Reicherts, Gun Woo Park, and Yvonne Rogers. 2022. Extending Chatbots to Probe Users: Enhancing Complex Decision-Making Through Probing Con- versations. In Proceedings of the 4th Conference on Conversational User Interfaces (CUI ’22) . Association for Computing Machinery, ...

  39. [48]

    Leon Reicherts and Yvonne Rogers. 2020. Do Make me Think! How CUIs Can Support Cognitive Processes. In Proceedings of the 2nd Conference on Conversa- tional User Interfaces (CUI ’20) . Association for Computing Machinery, New York, NY, USA, 1–4. https://doi.org/10.1145/3405755.3406157

  40. [49]

    Luca Rettenberger, Markus Reischl, and Mark Schutera. 2024. Assessing Political Bias in Large Language Models. https://doi.org/10.48550/arXiv.2405.13041 arXiv:2405.13041 [cs]

  41. [50]

    Pedro Riera and Francisco Cantú. 2022. Electoral systems and ideological voting. European Political Science Review 14, 4 (Nov. 2022), 463–481. https://doi.org/10. 1017/S1755773922000248

  42. [51]

    Salvatore Romano, Riccardo Angius, Natalie Kerby, Paul Bouchaud, Jacopo Amidei, and Andreas Kaltenbrunner. 2024. A Dataset to Assess Microsoft Copilot Answers in the Context of Swiss, Bavarian and Hessian Elections. Proceedings of the International AAAI Conference on Web and S...

  43. [52]

    Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger

  44. [53]

    Martin Schiele, Yannick Gittmann, Stefan Ilchmann, Ante Gojsalić, Dominik Jurinčić, and Phylis Klempt. 2024. Voting Advice Applications: Implementation of RAG-supported LLMs. https://doi.org/10.36227/techrxiv.172115156.64500701/v1

  45. [54]

    Martin Schultze. 2014. Effects of Voting Advice Applications (VAAs) on Political Knowledge About Party Positions. Policy & Internet 6, 1 (2014), 46–68. https://doi.org/10.1002/1944-2866.POI352 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/1944-2866.POI352

  46. [56]

    Isabelle Stadelmann-Steffen, Hannah Rajski, and Sophie Ruprecht. 2023. The role of vote advice application in direct-democratic opinion formation: an experiment from Switzerland. Acta Politica 58, 4 (Oct. 2023), 792–818. https://doi.org/10. 1057/s41269-022-00264-5

  47. [57]

    Thitaree Tanprasert, Sidney S Fels, Luanne Sinnamon, and Dongwook Yoon. 2024. Debate Chatbots to Facilitate Critical Thinking on YouTube: Social Identity and Conversational Style Make A Difference. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI...

  48. [58]

    Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, a...

  49. [59]

    Mathias Wessel Tromborg and Andreas Albertsen. 2023. Candidates, voters, and voting advice applications. European Political Science Review (April 2023), 1–18. https://doi.org/10.1017/S1755773923000103 Publisher: Cambridge University Press

  50. [60]

    Jasper van de Pol, Bregje Holleman, Naomi Kamoen, André Krouwel, and Claes de Vreese. 2014. Beyond Young, Highly Educated Males: A Typology of VAA Users. Journal of Information Technology & Politics 11, 4 (Oct. 2014), 397–

  51. [61]

    Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Goldstein. 2024. System- atic Biases in LLM Simulations of Debates. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Assoc...

  52. [62]

    Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West. 2023. Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks. https://arxiv.org/abs/2306.07899v1

  53. [63]

    Markus Wagner and Outi Ruusuvirta. 2012. Matching voters to parties: Voting advice applications and models of party choice. Acta Politica 47, 4 (Oct. 2012), 400–422. https://doi.org/10.1057/ap.2011.29

  54. [64]

    Maxime Walder, Jan Fivaz, Daniel Schwarz, and Nathalie Giger. 2024. Explaining the (non-) Use of Voting Advice Applications. Digit. Gov.: Res. Pract. 5, 3 (Oct. 2024), 33:1–33:29. https://doi.org/10.1145/3689214

  55. [65]

    Nina van Zanten and Roel Boumans. 2024. Voting Assistant Chatbot for Increas- ing Voter Turnout at Local Elections: An Exploratory Study. In Chatbot Research and Design, Asbjørn Følstad, Theo Araujo, Symeon Papadopoulos, Effie L.-C. Law, Ewa Luger, Morten Goodwin, Sebastian Ho...

  56. [66]

    Magdalena Wischnewski, Nicole Krämer, Christian Janiesch, Emmanuel Müller, Theodor Schnitzler, and Carina Newen. 2024. In Seal We Trust? Investigating the Effect of Certifications on Perceived Trustworthiness of AI Systems. Human- Machine Communication 8 (2024), 141–162. https...

  57. [67]

    Vera Liao, Michelle Zhou, Tyrone Grandison, and Yunyao Li

    Ziang Xiao, Q. Vera Liao, Michelle Zhou, Tyrone Grandison, and Yunyao Li. 2023. Powering an AI Chatbot with Expert Sourcing to Support Credible Health Infor- mation Access. In Proceedings of the 28th International Conference on Intelligent User Interfaces (IUI ’23) . Associati...

  58. [68]

    Wei Jie Yeo, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024. How Inter- pretable are Reasoning Explanations from Prompting Large Language Models?. In Findings of the Association for Computational Linguistics: NAACL 2024, Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). ...

  59. [69]

    Thomas Waldvogel, Monika Oberle, and Johanna Leunig. 2023. What Do Pupils Learn from Voting Advice Applications in Civic Education Classes? Effects of a Digital Intervention Using Voting Advice Applications on Students’ Political Dispositions. Social Sciences 12, 11 (Nov. 2023...

  60. [73]

    Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. Can Large Language Models Transform Computational So- cial Science? Computational Linguistics 50, 1 (March 2024), 237–291. https: //doi.org/10.1162/coli_a_00502 Place: Cambridge, MA Publishe...

  61. [411]

    https://doi.org/10.1080/19331681.2014.958794 Publisher: Routledge _eprint: https://doi.org/10.1080/19331681.2014.958794

  62. [2023]

    https: //doi.org/10.31219/osf.io/stakv

    Artificial Intelligence Can Persuade Humans on Political Issues. https: //doi.org/10.31219/osf.io/stakv

  63. [2024]

    2024), pgae034

    How persuasive is AI-generated propaganda? PNAS Nexus 3, 2 (Feb. 2024), pgae034. https://doi.org/10.1093/pnasnexus/pgae034

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.