Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Exploring LLMs for Automated Generation and Adaptation of Questionnaires

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that an LLM pipeline can generate clear, specific survey questions and adapt existing questionnaires for new audiences, with human raters finding LLM-adapted items slightly clearer and less biased than the original U.S.

desk verdict A transparent early-stage empirical study of LLM questionnaire generation and adaptation; the generation part holds up, but the pretesting claim rests on a self-referential loop that needs external validation. read the letter →

arxiv 2501.05985 v2 pith:KUEUHITH submitted 2025-01-10 cs.HC cs.CY

classification cs.HCcs.CY
keywords LLMquestionnairegenerationsurveypretestingpersonasimulationcross-culturaladaptationcognitiveinterviewingretrieval-augmentedGPT-4omethodology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether large language models can take over two labor-intensive parts of survey design: writing new questionnaires and adapting validated ones for a different culture or country. It builds a pipeline that generates questions with retrieved context, invents fictional target-audience personas, runs a simulated pilot interview, and has a second LLM review the transcripts to suggest rewrites. In two online studies, participants rated LLM-generated items as clear and specific, but pretesting through simulated personas made items less clear in the creation task; in the adaptation task, the LLM's rewrites of a U.S. climate-change survey were seen as marginally clearer and less biased by South African raters. The authors conclude that LLMs are useful for generating and adapting questionnaires, while automated pretesting needs human oversight.

What carries the argument

The pipeline's central object is the simulated pilot study. The system first generates a questionnaire using a research question, ten retrieved questions from a standardized-survey database, and a summary of recent news articles; it also invents a list of personas representing the target audience. An interviewer LLM then administers the questionnaire to persona-conditioned participant LLMs, asking cognitive-interviewing follow-up questions. A reviewer LLM reads the transcripts and returns suggested rewrites for questions where personas showed confusion, ambiguity, or bias. The rewrites are the output that gets compared against the originals in the human evaluation.

What would settle it

Compare LLM-suggested rewrites against rewrites from human cognitive interviews on the same original questionnaire: if a substantial share of the LLM's flagged problems are not raised by any human respondent, or if the LLM's rewrites are rated lower on clarity by the target audience, the pretesting component's value is falsified.

Watch

Extended reading notes

Core claim

The central claim is that a pipeline combining retrieval-augmented generation with an LLM-simulated pilot study can produce questionnaires that are relevant and specific, and can adapt a standardized questionnaire to a new target audience. On the creation side, the paper reports that LLM-generated statements were rated clearer than LLM-pretested statements by both lay raters and experts, while pretesting improved specificity for lay raters; on the adaptation side, the LLM-adapted version of a U.S. climate-change questionnaire was rated marginally clearer and less biased than the original by South African participants. The authors interpret this as evidence that LLM pretesting adds little when generating from scratch but helps when recontextualizing existing instruments.

Load-bearing premise

The load-bearing premise is that the LLM-simulated pilot interviews, conducted with fictional personas, produce feedback about how a real target population will understand and react to the questions; if those simulated reactions diverge from real respondents, the pretesting step can make questions worse.

Editorial extensions

If this is right

  • If the pipeline works as claimed, researchers can quickly generate draft questionnaires with current-event context and standard-question grounding, cutting the time to a first testable draft.
  • For cross-cultural adaptation, LLM pretesting could flag questions whose assumptions (e.g., 'tax rebate', 'Congress') do not transfer, and produce rewrites that real raters find at least as clear.
  • The mixed creation results imply that automated pretesting should be applied selectively; generated questions may already be clear enough that further LLM revision reduces clarity.
  • The evaluation design, with paired clarity/bias/relevance/specificity ratings, gives survey researchers a reusable template for judging LLM-assisted instruments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to run the same pipeline with human cognitive interviews instead of LLM personas; if the simulated pilot does not flag the same problems human respondents do, the pretesting step should be skipped for new questionnaires.
  • The clarity/specificity tradeoff the paper observes suggests that LLM pretesting might be tuned for particular question types: rewrites that add concrete examples may help adaptation but hurt brevity-sensitive items.
  • The pipeline's dependence on GPT-4o raises the question of whether cheaper open models would show the same pattern, since the 'pretesting made it less clear' result may be a property of the specific LLM's revision style.
  • One could build a cost-aware variant that only asks for LLM pretesting when a human pilot is infeasible, since the paper's adaptation gains are small and its creation gains are negative.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents an LLM-based pipeline (GPT-4o) for questionnaire generation and pretesting. The pipeline retrieves context from news summaries and SQP questions, generates personas of target respondents, simulates pilot interviews, and uses a reviewer LLM to propose question revisions. The authors then evaluate the pipeline in two studies: RQ1 uses a political-trust questionnaire rated by 238 Prolific US participants and 13 experts; RQ2 adapts a U.S. climate-change survey for a South African audience and evaluates the five items the LLM chose to modify with 118 Prolific participants. The reported findings are that LLM-generated items were rated clearer and LLM-pretested items more specific in RQ1 (on Prolific), and that in RQ2 the adapted items were rated marginally clearer and less biased than the original items. The paper concludes that LLMs are promising for questionnaire generation and adaptation, but that pretesting did not improve the political-trust questionnaire and that human oversight remains necessary.

Significance. If the findings hold, this is a useful exploratory contribution to a small but growing literature on LLM-assisted survey methodology. The paper provides public prompts and code, uses human raters including domain experts, and evaluates items on multiple criteria with both direct and reversed statements. The two Prolific studies have reasonable sample sizes for exploratory work. However, the central validity of the simulated-pilot pretesting step is untested, the headline effects are small and mostly marginal, and RQ1 lacks a human-written baseline. The significance is therefore real but bounded; the paper is best read as an existence proof of the pipeline's feasibility rather than as evidence that LLM-pretesting generally improves questionnaires.

major comments (3)
  1. [§3.1.2, §3.2, Appendix A.5–A.7] The pretesting module is a closed GPT-4o loop: the same model family generates the personas, simulates the respondent interviews, and writes the reviewer feedback that produces the revised items. Section 3.2's formative check only reports that the pipeline 'detected issues in 5 out of the 13' GESIS questions, with no precision, recall, or comparison to the human pretest's issue list, so it does not validate the module. Because the RQ2 result (Section 5.2, Figure 8) is built from the five items the LLM selected to modify, the observed small gains in clarity and bias could be generic rewriting effects rather than evidence that persona-based pretesting identifies real comprehension problems. The manuscript needs an external anchor—e.g., a human cognitive-interview issue list or a randomized control where rewrites of unproblematic items are also evaluated.
  2. [Table 8, §4.2, §5.2] Table 8 reports 12 paired t-tests without any multiple-comparison correction, and the two headline RQ2 effects (clarity and bias) are only marginally significant (p<0.1). Under a standard correction such as Benjamini-Hochberg at α=0.05, these p-values would not survive. The abstract and Section 5.2 state these as findings; the manuscript should either apply a correction, report effect sizes and confidence intervals, or explicitly label the results as exploratory. The same issue affects the RQ1 clarity and specificity claims in Sections 4.2.1 and 4.2.2.
  3. [§4.1–§4.2, §6.1.1] RQ1 has no human-written baseline. The paper claims in Section 6.1.1 that LLMs 'have the potential to generate relevant and specific questions,' but the comparison is only between LLM-generated and LLM-pretested versions, both produced by the same pipeline. Without a comparison against a human-authored questionnaire (e.g., the WVS module that inspired the items), the generation claim is unsupported. The abstract's statement that 'LLM-generated text clearer' is a relative comparison to the LLM-pretested version, not an absolute quality assessment.
minor comments (6)
  1. [§4.2.1–§4.2.2, Abstract] The abstract and Section 6.1.1 state that LLM-pretested text was 'more specific,' but this held only for Prolific participants; the 13 experts rated the pretested text as less specific (Figure 6). The claim should be qualified to the Prolific sample.
  2. [§5.1.1] The number of personas generated for the South African pilot is not reported, whereas RQ1 explicitly states 50 personas. This matters for reproducibility and for assessing the diversity of the simulation.
  3. [Table 3] Table 3 is hard to read: the layout does not clearly separate the direct and reversed evaluation statements for each criterion. Consider using one row per criterion with two columns for the two statements.
  4. [§7, §4.2.2] The expert sample (N=13) is acknowledged as small in Section 7, but Section 4.2.2 states the expert findings without hedging. Add a cautionary note in the results section itself, not only in the limitations.
  5. [§8] The conclusion states that LLMs can make questionnaires 'more explicit, unbiased, and relevant,' but relevance was not significantly different in RQ2 (p>0.1). This overstates the results; revise to reflect the actual findings.
  6. [Throughout] The paper uses 'pretesting' and 'pre-testing' interchangeably. Standardize to one form.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the main claims rest on independent human ratings, not on LLM outputs alone.

full rationale

The paper's load-bearing findings (RQ1 and RQ2) are evaluated by 356 Prolific participants and 13 experts using Likert ratings and forced-choice comparisons of original versus LLM-pretested or adapted items, not by the same LLM that produced the revisions. The simulated-pilot module (Section 3.1.2) is a generation step whose output could be and was contradicted by human raters: RQ1 experts rated pretested statements as less clear and more biased, so the empirical claims are falsifiable and not forced by construction. The formative feasibility study (Section 3.2) reports that the pipeline 'detected issues in 5 out of the 13 tested questions' without reporting precision or recall against the GESIS human pretest; this is a validation gap, not a circular reduction. No load-bearing self-citations or imported uniqueness theorems appear. Accordingly, no step reduces by definition to its inputs.

Assumptions & free parameters 4 free parameters · 5 assumptions · 1 invented entities

The paper introduces no fitted parameters in the mathematical sense, but several pipeline hyperparameters (news count, SQP count, persona count, sampling settings) are hand-chosen and affect results. The main epistemic burden is the assumption that LLM-simulated personas and crowdsourced raters provide valid evidence about real questionnaire quality.

free parameters (4)
  • Number of retrieved news articles = 5
    Chosen by hand in the generation prompt (Section 3.1.1); changes generation context and specificity.
  • Number of SQP questions retrieved = 10
    Chosen by hand for FAISS retrieval (Section 3.1.1); affects grounding and specificity.
  • Number of personas in political trust pretest = 50
    Chosen by hand for RQ1 simulated pilot (Section 4.1.1); the persona set defines the pretest sample.
  • Temperature and random seed = not reported for main studies
    Mentioned as varied in formative feasibility studies (Section 3.2) but not fixed or reported for the main Prolific studies, so generation variance is uncontrolled.
assumptions (5)
  • domain assumption LLM-simulated personas can represent target-audience interpretation patterns
    The pretesting step in Section 3.1.2 relies on persona interviews to surface real questionnaire problems; no validation against actual target-population behavior is provided.
  • domain assumption Prolific participants with domain degrees are valid judges of questionnaire clarity, bias, relevance, and specificity
    The evaluation in Sections 4.1.2 and 5.1.2 treats these ratings as ground truth for questionnaire quality.
  • standard math Likert ratings can be treated as interval data for paired t-tests
    Sections 3.5.1 and Table 8 use mean ratings and t-tests without addressing ordinality.
  • domain assumption The source instruments (WVS political trust module and Yale climate change survey) are appropriate standardized benchmarks
    RQ1 and RQ2 use these as the basis for generation and adaptation (Sections 4.1.1, 5.1.1).
  • domain assumption GPT-4o thematic analysis of open-ended responses reliably categorizes participant reasoning
    Section 3.5.2 uses GPT-4o to categorize free-text responses; no inter-rater reliability with human coding is reported.
invented entities (1)
  • Fictional personas for simulated pilot participants
    purpose: Act as stand-ins for target-audience respondents in pretesting interviews
    These are generated fictional profiles with no external validation that their responses match real demographic groups; the paper cites prior work but does not validate this pipeline's personas.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring LLMs for Automated Generation and Adaptation of Questionnaires." pith.science (2026). https://pith.science/paper/KUEUHITH

@misc{pith2026250105985,
  author       = {Pith},
  title        = {Pith review of: Exploring LLMs for Automated Generation and Adaptation of Questionnaires},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KUEUHITH}},
  note         = {Machine review of arXiv:2501.05985}
}
read the original abstract

Effective questionnaire design improves the validity of the results, but creating and adapting questionnaires across contexts is challenging due to resource constraints and limited expert access. Recently, the emergence of LLMs has led researchers to explore their potential in survey research. In this work, we focus on the suitability of LLMs in assisting the generation and adaptation of questionnaires. We introduce a novel pipeline that leverages LLMs to create new questionnaires, pretest with a target audience to determine potential issues and adapt existing standardized questionnaires for different contexts. We evaluated our pipeline for creation and adaptation through two studies on Prolific, involving 238 participants from the US and 118 participants from South Africa. Our findings show that participants found LLM-generated text clearer, LLM-pretested text more specific, and LLM-adapted questions slightly clearer and less biased than traditional ones. Our work opens new opportunities for LLM-driven questionnaire support in survey research.

Figures

Figures reproduced from arXiv: 2501.05985 by the authors.

Figure 1
Figure 1. We explore the role of LLMs in questionnaire design by understanding their capabilities in creating and pretesting [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Our complete questionnaire creation and pretesting pipeline involve: (1) Using the research question and relevant [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. (1) Based on the research question, the domain and [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: (Simplified) Summary of the tasks in the evaluation studies. (1) In the question rating task, the participants were [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Mean ratings for the LLM-generated and LLM [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 7
Figure 7. Figure 7: Comparison of question preferences between par [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Mean ratings (with standard errors) for the original [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: The results of the questionnaire comparison task [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Example question from rating task of the study where participants were asked to evaluate the given statement on a [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: Example question from comparison task of the study where participants were asked to select the clearer question. [PITH_FULL_IMAGE:figures/full_fig_p020_11.png]
Figure 12
Figure 12. Figure 12: Manually coded reasons from the free text responses of the experts when selecting the clearer question in the [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research

    cs.DL 2025-09 conditional novelty 4.0 of 10

    A work-in-progress systematic review finds LLM use in survey research clusters in instrument development, synthetic respondent modeling, and automated text classification, leaving interviewing and cross-lingual work thin.

Reference graph

Works this paper leans on

93 extracted references · 49 canonical work pages · cited by 1 Pith paper

  1. [1]

    Noorhan Abbas, Thomas Pickard, Eric Atwell, and Aisha Walker. 2021. University Student Surveys Using Chatbots: Artificial Intelligence Conversational Agents. InLearning and Collaboration Technologies: Games and Virtual Environments for Learning, Panayiotis Zaphiris and Andri Ioannou (Eds.). Springer International Publishing, Cham, 155–169. doi:10.1007/978...

  2. [2]

    Argyle, Ethan C

    Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples.Political Analysis31, 3 (July 2023), 337–351. doi:10. 1017/pan.2023.2

  3. [3]

    Beaton, Claire Bombardier, Francis Guillemin, and Marcos Bosi Ferraz

    Dorcas E. Beaton, Claire Bombardier, Francis Guillemin, and Marcos Bosi Ferraz

  4. [4]

    Beaton, Claire Bombardier, Francis Guillemin, and Mar- cos Bosi Ferraz

    Dorcas E. Beaton, Claire Bombardier, Francis Guillemin, and Mar- cos Bosi Ferraz. 2000. Guidelines for the Process of Cross-Cultural Adaptation of Self-Report Measures.Spine25, 24 (Dec. 2000), 3186. https://journals.lww.com/spinejournal/pages/articleviewer.aspx?year= 2000&issue=12150&article=00014&type=Fulltext

  5. [5]

    Dorothée Behr, Michael Braun, and Luisa Aiglstorfer. 2024. Showcasing the usefulness of web probing: Do subtle variations in questionnaire translation lead to different survey responding?Quality & Quantity(March 2024). doi:10.1007/ s11135-024-01843-8

  6. [6]

    Dorothée Behr, Katharina Meitinger, Michael Braun, and Lars Kaczmirek. 2017. Web probing – implementing probing techniques from cognitive, interviewing Exploring LLMs for Automated Generation and Adaptation of Questionnaires in web surveys with the goal to assess the validity of survey questions (GESIS Survey Guidelines)Web probing – implementing probing ...

  7. [7]

    Johnny Blair, Allison Ackermann, Linda Piccinino, and Rachel Levenstein. 2007. Using Behavior Coding to Validate Cognitive Interview Findings. (Jan. 2007)

  8. [8]

    Johnny Blair and Frederick G. Conrad. 2011. Sample Size for Cognitive Interview Pretesting.The Public Opinion Quarterly75, 4 (2011), 636–658. https://www. jstor.org/stable/41288411 Publisher: American Association for Public Opinion Research

Show all 93 references
  1. [9]

    Johnny Blair and Linda Piccinino. 2005. The development and testing of instru- ments for cross-cultural and multi-cultural surveys. InMethodological aspects in cross-national research, Jürgen H. P. Hoffmeyer-Zlotnik and Janet Harkness (Eds.). ZUMA-Nachrichten Spezial, Vol. 11....

  2. [10]

    Petra M Boynton and Trisha Greenhalgh. 2004. Selecting, designing, and devel- oping your questionnaire.BMJ : British Medical Journal328, 7451 (May 2004), 1312–1315. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC420179/

  3. [11]

    Zi Chai and Xiaojun Wan. 2020. Learning to Ask More: Semi-Autoregressive Sequential Question Generation under Dual-Graph Interaction. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel...

  4. [12]

    Roberts Dar ‘gis, Guntis B ¯arzdin, š, Inguna Skadi n, a, Normunds Gr ¯uz¯ıtis, and Baiba Saul¯ıte. 2024. Evaluating Open-Source LLMs in Low-Resource Lan- guages: Insights from Latvian High School Exams. InProceedings of the 4th International Conference on Natural Language Pro...

  5. [13]

    Elizabeth Dean, Brian Head, and Jodi Swicegood. 2013. Virtual Cognitive In- terviewing Using Skype and Second Life. InSocial Media, Sociality, and Survey Research. 107–132. doi:10.1002/9781118751534.ch5 Journal Abbreviation: Social Media, Sociality, and Survey Research

  6. [14]

    Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023. Can AI language models replace human participants?Trends in Cognitive Sciences27, 7 (July 2023), 597–600. doi:10.1016/j.tics.2023.04.008 Publisher: Elsevier

  7. [15]

    Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. 2024. Questioning the Survey Responses of Large Language Models. doi:10.48550/ arXiv.2306.07951 arXiv:2306.07951 [cs]

  8. [16]

    Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The Faiss library. (2024). arXiv:2401.08281 [cs.LG]

  9. [17]

    Jonathan Drennan. 2003. Cognitive interviewing: verbal data in the design and pretesting of questionnaires.Journal of Advanced Nursing42, 1 (2003), 57–63. doi:10.1046/j.1365-2648.2003.02579.x arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1046/j.1365-2648.2003.02579.x

  10. [18]

    Osborne, Gerald R

    Jonathan Epstein, Richard H. Osborne, Gerald R. Elsworth, Dorcas E. Beaton, and Francis Guillemin. 2015. Cross-cultural adaptation of the Health Education Impact Questionnaire: experimental study showed expert committee, not back- translation, added value.Journal of Clinical E...

  11. [19]

    Jonathan Epstein, Ruth Miyuki Santo, and Francis Guillemin. 2015. A review of guidelines for cross-cultural adaptation of questionnaires could not bring out a consensus.Journal of Clinical Epidemiology68, 4 (April 2015), 435–441. doi:10.1016/j.jclinepi.2014.11.021

  12. [20]

    Weinstein, Vasil I

    Jennifer Fei, Jessica Wolff, Michael Hotard, Hannah Ingham, Saurabh Khanna, Duncan Lawrence, Beza Tesfaye, Jeremy M. Weinstein, Vasil I. Yasenov, and Jens Hainmueller. 2022. Automated Chat Application Surveys Using Whatsapp: Evidence from Panel Surveys and a Mode Experiment. d...

  13. [21]

    Yifan Gao, Piji Li, Irwin King, and Michael R. Lyu. 2019. Interconnected Question Generation with Coreference Alignment and Conversation Flow Modeling. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and L...

  14. [22]

    Yubin Ge, Ziang Xiao, Jana Diesner, Heng Ji, Karrie Karahalios, and Hari Sun- daram. 2023. What should I Ask: A Knowledge-driven Approach for Follow-up Questions Generation in Conversational Surveys. doi:10.48550/arXiv.2205.10977

  15. [23]

    Linn Gjersing, John RM Caplehorn, and Thomas Clausen. 2010. Cross-cultural adaptation of research instruments: language, setting, time and statistical considerations.BMC Medical Research Methodology10, 1 (10 Feb 2010), 13. doi:10.1186/1471-2288-10-13

  16. [24]

    Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2024. Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs. doi:10.48550/arXiv.2311.04892 arXiv:2311.04892 [cs]

  17. [25]

    Melissa Guyre, Liz Holland, Nirva Shah, and Rahul R. Divekar. 2024. Prompt Engineering an LLM into Roleplaying a Management Coach: a Short Guide by and for Non-NLP Experts. InProceedings of the 6th ACM Conference on Conversational User Interfaces (CUI ’24). Association for Com...

  18. [26]

    Otso Haavisto and Robin Welsch. 2024. Questionnaires for Everyone: Stream- lining Cross-Cultural Questionnaire Adaptation with GPT-Based Translation Quality Evaluation. doi:10.48550/arXiv.2407.20608 arXiv:2407.20608 [cs]

  19. [27]

    Christian Haerpfer, Ronald Inglehart, Alejandro Moreno, Christian Welzel, Kseniya Kizilova, Jaime Diez-Medrano, Marta Lagos, Pippa Norris, Eduard Ponarin, and Bi Puranen. 2020. World Values Survey wave 7 (2017-2020) cross- national data-set

  20. [28]

    Anne-Wil Harzing. 2005. Does the Use of English-language Questionnaires in Cross-national Research Obscure National Differences?International Jour- nal of Cross Cultural Management5, 2 (Aug. 2005), 213–224. doi:10.1177/ 1470595805054494 Publisher: SAGE Publications

  21. [29]

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI generates covertly racist decisions about people based on their dialect.Nature 633, 8028 (01 Sep 2024), 147–154. doi:10.1038/s41586-024-07856-5

  22. [30]

    John J. Horton. 2023. Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? doi:10.48550/arXiv.2301.07543 arXiv:2301.07543 [econ, q-fin]

  23. [31]

    Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case Study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, N...

  24. [32]

    Sviatlana Höhn, Jauwairia Nasir, Ali Paikan, Pouyan Ziafati, and Elisabeth André

  25. [33]

    Nashwa Ismail, Gary Kinchin, and Julie-Ann Edwards. 2018. Pilot study, Does it really matter? Learning lessons from conducting a pilot study for a qualitative PhD thesis.International Journal of Social Science Research6, 1 (March 2018), 1–17. doi:10.5296/ijssr.v6i1.11720 Num P...

  26. [34]

    Jansen, Soon-gyo Jung, and Joni Salminen

    Bernard J. Jansen, Soon-gyo Jung, and Joni Salminen. 2023. Employing large language models in survey research.Natural Language Processing Journal4 (Sept. 2023), 100020. doi:10.1016/j.nlp.2023.100020

  27. [35]

    Stephen Jenkins and Tony Solomonides. 2000. Automating Questionnaire Design and Construction.International Journal of Market Research42, 1 (Jan. 2000), 1–13. doi:10.1177/147078530004200106

  28. [36]

    Ng Chirk Jenn. 2006. Designing A Questionnaire.Malaysian Family Physician : the Official Journal of the Academy of Family Physicians of Malaysia1, 1 (April 2006), 32–35. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4797036/

  29. [37]

    Johanson and Gordon P

    George A. Johanson and Gordon P. Brooks. 2010. Initial Scale Development: Sam- ple Size for Pilot Studies.Educational and Psychological Measurement70, 3 (June 2010), 394–400. doi:10.1177/0013164409355692 Publisher: SAGE Publications Inc

  30. [38]

    Kendra Kamp, Gwen Wyatt, Sharon Dudley-Brown, Kelly Brittain, and Barbara Given. 2018. Using cognitive interviewing to improve questionnaires: An exem- plar study focusing on individual and condition-specific factors.Applied Nursing Research43 (Oct. 2018), 121–125. doi:10.1016...

  31. [39]

    Junsol Kim and Byungkyu Lee. 2023. AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction. https://arxiv.org/abs/ 2305.09620v2

  32. [40]

    Sunwoong Kim, Jongho Jeong, Jin Soo Han, and Donghyuk Shin. 2024. LLM-Mirror: A Generated-Persona Approach for Survey Pre-Testing. arXiv:2412.03162 [cs.CY] https://arxiv.org/abs/2412.03162

  33. [41]

    Soomin Kim, Joonhwan Lee, and Gahgene Gweon. 2019. Comparing Data from Chatbot and Web Surveys: Effects of Platform and Conversational Style on Survey Response Quality. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19). Association for Co...

  34. [42]

    Hill, and A Herzog

    Bärbel Knäuper, Robert Belli, D. Hill, and A Herzog. 1997. Question Difficulty and Respondents’ Cognitive Ability: The Effect on Data Quality.Journal of Official Statistics13 (Jan. 1997)

  35. [43]

    Lucrezia Laraspata, Fabio Cardilli, Giovanna Castellano, and Gennaro Vessio. [n. d.]. Enhancing Human Capital Management through GPT-driven Question- naire Generation. ([n. d.])

  36. [44]

    Yan Lei, Liang Pang, Yuanzhuo Wang, Huawei Shen, and Xueqi Cheng. 2024. Qs- nail: A Questionnaire Dataset for Sequential Question Generation. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2...

  37. [45]

    Timo Lenzer, Patricia Hadler, and Cornelia Neuert. 2024. Cognitive Pretesting. Adhikari et al

  38. [46]

    Timo Lenzner. 2012. Effects of Survey Question Comprehensibility on Response Quality.Field Methods24, 4 (Nov. 2012), 409–428. doi:10.1177/1525822X12448166 Publisher: SAGE Publications Inc

  39. [47]

    Timo Lenzner and Cornelia E. Neuert. 2017. Pretesting Survey Questions Via Web Probing – Does it Produce Similar Results to Face-to-Face Cognitive Inter- viewing?Survey Practice10, 4 (Oct. 2017). doi:10.29115/SP-2017-0020

  40. [48]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Gen- eration for Knowledge-Intensive NLP Tasks. InAdvances in...

  41. [49]

    Xun Liang, Hanyu Wang, Yezhaohui Wang, Shichao Song, Jiawei Yang, Simin Niu, Jie Hu, Dan Liu, Shunyu Yao, Feiyu Xiong, and Zhiyu Li. 2024. Controllable Text Generation for Large Language Models: A Survey. arXiv:2408.12599 [cs.CL] https://arxiv.org/abs/2408.12599

  42. [50]

    Jieli Liu and Haining Wang. 2024. Assessing Gender and Racial Bias in Large Language Model-Powered Virtual Reference.Proceedings of the Association for Information Science and Technology61, 1 (2024), 576–580. doi:10.1002/pra2.1061 arXiv:https://asistdl.onlinelibrary.wiley.com/...

  43. [51]

    Antonio Maiorino, Zoe Padgett, Chun Wang, Misha Yakubovskiy, and Peng Jiang

  44. [52]

    Marlon, Xinran Wang, Parrish Bergquist, Peter D

    Jennifer R. Marlon, Xinran Wang, Parrish Bergquist, Peter D. Howe, Anthony Leiserowitz, Edward Maibach, Matto Mildenberger, and Seth Rosenthal. 2022. Change in US state-level public opinion about climate change: 2008–2020.Envi- ronmental Research Letters17, 12 (2022). doi:10.1...

  45. [53]

    Catriona Rachel Mayland, Christina Gerlach, Katrin Sigurdardottir, Marit Irene Tuen Hansen, Wojciech Leppert, Andrzej Stachowiak, Maria Krajew- ska, Eduardo Garcia-Yanneo, Vilma Adriana Tripodoro, Gabriel Goldraij, Mar- tin Weber, Lair Zambon, Juliana Nalin Passarini, Ivete Br...

  46. [54]

    K Meitinger, C Neuert, C Beitz, and N Menold. 2016. Pretesting of special module on ICT at work, working conditions & learning digital skills

  47. [55]

    Felix Ndashimye, Oumarou Hebie, and Jasper Tjaden. 2024. Effectiveness of WhatsApp for Measuring Migration in Follow-Up Phone Surveys. Lessons from a Mode Experiment in Two Low-Income Countries during COVID Contact Restrictions.Social Science Computer Review42, 2 (April 2024),...

  48. [56]

    Cornelia Neuert. 2023. Design of multiple open-ended probes in cognitive online pretests using web probing. (2023). doi:10.13094/SMIF-2023-00005 Publisher: FORS/PUMA/GESIS

  49. [57]

    Francisco Olivos and Minhui Liu. 2024. ChatGPTest: opportunities and cautionary tales of utilizing AI for questionnaire pretesting. doi:10.48550/arXiv.2405.06329

  50. [59]

    Kristen Olson. 2010. An Examination of Questionnaire Evaluation by Expert Re- viewers.Field Methods22, 4 (Nov. 2010), 295–318. doi:10.1177/1525822X10379795 Publisher: SAGE Publications Inc

  51. [60]

    Jessica Parker, Veronica Richard, and Kimberly Becker. 2023. Flexibility & It- eration: Exploring the Potential of Large Language Models in Developing and Refining Interview Protocols.The Qualitative Report28, 9 (Sept. 2023), 2772–2791. doi:10.46743/2160-3715/2023.6695

  52. [61]

    Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Iryna Gurevych and Yusuke Miyao (Eds.). Associati...

  53. [62]

    Pooja Rao S B, Manish Agnihotri, and Dinesh Babu Jayagopi. 2021. Improv- ing Asynchronous Interview Interaction with Follow-up Question Generation. (March 2021). doi:10.9781/ijimai.2021.02.010 Accepted: 2022-04-25T09:02:55Z Publisher: International Journal of Interactive Multi...

  54. [63]

    Nina Reynolds, Adamantios Diamantopoulos, and Bodo Schlegelmilch. 1993. Pre-Testing in Questionnaire Design: A Review of the Literature and Suggestions for Further Research.Market Research Society. Journal.35, 2 (March 1993), 1–11. doi:10.1177/147078539303500202 Publisher: SAG...

  55. [64]

    Rothschild, James Brand, Hope Schroeder, and Jenny Wang

    David M. Rothschild, James Brand, Hope Schroeder, and Jenny Wang. 2024. Opportunities and risks of LLMs in survey research. doi:10.2139/ssrn.5001645

  56. [65]

    David Rozado. 2024. The political preferences of LLMs.PLOS ONE19, 7 (07 2024), 1–15. doi:10.1371/journal.pone.0306621

  57. [66]

    Adler, Lea Rau, and Bernd Schmitt

    Marko Sarstedt, Susanne J. Adler, Lea Rau, and Bernd Schmitt. 2024. Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines.Psychology & Marketing41, 6 (2024), 1254–1270. doi:10.1002/mar.21982

  58. [67]

    Rothkopf, and Kristian Kersting

    Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A. Rothkopf, and Kristian Kersting. 2022. Large pre-trained language models contain human- like biases of what is right and wrong to do.Nature Machine Intelligence4, 3 (March 2022), 258–268. doi:10.1038/s42256-022-00...

  59. [68]

    Kerry Scott, Dipanwita Gharai, Manjula Sharma, Namrata Choudhury, Bibha Mishra, Sara Chamberlain, and Amnesty LeFevre. 2020. Yes, no, maybe so: the importance of cognitive interviewing to enhance structured surveys on respectful maternity care in northern India.Health Policy a...

  60. [69]

    Josh Seltzer, Jiahua Pan, Kathy Cheng, Yuxiao Sun, Santosh Kolagati, Jimmy Lin, and Shi Zong. 2023. $SmartProbe$: A Virtual Moderator for Market Research Surveys. doi:10.48550/arXiv.2305.08271 arXiv:2305.08271 [cs]

  61. [70]

    Vanessa E. C. Sousa, Jeffrey Matson, and Karen Dunn Lopez. 2017. Questionnaire Adapting: Little Changes Mean a Lot.Western Journal of Nursing Research 39, 9 (Sept. 2017), 1289–1300. doi:10.1177/0193945916678212 Publisher: SAGE Publications Inc

  62. [71]

    John M. Stahura. 2005. Methods for Testing and Evaluating Survey Ques- tionnaires.Contemporary Sociology34, 4 (July 2005), 427–428. doi:10.1177/ 009430610503400457 Publisher: SAGE Publications Inc

  63. [72]

    Ming-Hsiang Su, Chung-Hsien Wu, Kun-Yi Huang, Qian-Bei Hong, and Huai- Hung Huang. 2018. Follow-up Question Generation Using Pattern-based Seq2seq with a Small Corpus for Interview Coaching. 1006–1010. doi:10.21437/ Interspeech.2018-1007

  64. [73]

    Guangzhi Sun, Xiao Zhan, and Jose Such. 2024. Building Better AI Agents: A Provocation on the Utilisation of Persona in LLM-based Conversational Agents. InProceedings of the 6th ACM Conference on Conversational User Interfaces (CUI ’24). Association for Computing Machinery, Ne...

  65. [74]

    Synodinos

    Nicolaos E. Synodinos. 2003. The “art” of questionnaire construction: some important considerations for manufacturing studies.Integrated Manufacturing Systems14, 3 (Jan. 2003). doi:10.1108/09576060310463172

  66. [75]

    Hamed Taherdoost. 2022. Designing a Questionnaire for a Research Paper: A Comprehensive Guide to Design and Develop an Effective Questionnaire. Asian Journal of Managerial Science11, 1 (March 2022), 8–16. doi:10.51983/ajms- 2022.11.1.3087 Number: 1

  67. [76]

    Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. 2024. Do LLMs exhibit human-like response biases? A case study in survey design. doi:10.48550/arXiv.2311.04076 arXiv:2311.04076 [cs]

  68. [77]

    Brown, Jennifer V

    Alejandro Cuevas Villalba, Eva M. Brown, Jennifer V. Scurrell, Jason Entenmann, and Madeleine I. G. Daepp. 2023. Automated Interviewer or Augmented Survey? Collecting Social Data with Large Language Models. doi:10.48550/arXiv.2309. 10187

  69. [78]

    Dickerson

    Angelina Wang, Jamie Morgenstern, and John P. Dickerson. 2024. Large language models cannot replace human participants because they cannot portray identity groups. doi:10.48550/arXiv.2402.01908 arXiv:2402.01908 [cs]

  70. [79]

    W.E. Saris. 2022.SQP 3.0. Universitat Pompeu Fabra, Barcelona, GESIS – Leibniz- Institut für Sozialwissenschaften e.V., Mannheim. https://sqp.gesis.org

  71. [80]

    Lang, Quirin Würschinger, and Frauke Kreuter

    Alexander Wuttke, Matthias Aßenmacher, Christopher Klamm, Max M. Lang, Quirin Würschinger, and Frauke Kreuter. 2024. AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers. arXiv:2410.01824 [cs.HC] https://arxiv.org/abs/2410.01824

  72. [81]

    Dongling Xiao, Han Zhang, Yukun Li, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020. ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation. doi:10.48550/arXiv.2001.11314 arXiv:2001.11314 [cs] version: 3

  73. [82]

    Yifan Yang, Xiaoyu Liu, Qiao Jin, Furong Huang, and Zhiyong Lu. 2024. Un- masking and quantifying racial bias of large language models in medical report generation.Communications Medicine4, 1 (10 Sep 2024), 176. doi:10.1038/s43856- 024-00601-z

  74. [83]

    Griffith, and Timothy Bickmore

    Hye Sun Yun, Mehdi Arjmand, Phillip Raymond Sherlock, Michael Paasche- Orlow, James W. Griffith, and Timothy Bickmore. 2023. Keeping Users Engaged During Repeated Administration of the Same Questionnaire: Using Large Lan- guage Models to Reliably Diversify Questions. doi:10.48...

  75. [84]

    Zhao Zou, Omar Mubin, Fady Alnajjar, and Luqman Ali. 2024. A pilot study of measuring emotional response and perception of LLM-generated questionnaire and human-generated questionnaires.Scientific Reports14, 1 (Feb. 2024), 2781. doi:10.1038/s41598-024-53255-1 Publisher: Nature...

  76. [88]

    Summary of the first article </relevant_articles> <relevant_questions> * First relevant question </relevant_questions> Questions:

  77. [89]

    Why do you like this product? Type: closed-ended Options: Ease of use, Quality, Price, Brand

  78. [90]

    language

    What suggestions do you have for the product? Type: open-ended Options: - Adhikari et al. A.2 Persona generation system prompt You are a researcher working on Survey research. As a part of your research, you need to determine the correct samples to pick from the intended audie...

  79. [91]

    How you feel about the national congress in your country?

    Generate 6 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) on the question "How you feel about the national congress in your country?". Number it from P1 to P6

  80. [92]

    How do you feel about the government in your country?

    Generate 6 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) on the question "How do you feel about the government in your country?". Number it from G1 to G6

  81. [93]

    How do you feel about the United Nations?

    Generate 6 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) on the question "How do you feel about the United Nations?". Number it from UN1 to UN6

  82. [94]

    Number it from A to O

    Generate 15 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) related to trust on Politicians and the government. Number it from A to O

  83. [95]

    Disagree

    Generate a question asking about their trust on the <head of state>in their country. Replace <head of state>with the respective head of state in the country. Make sure to use a likert scale of 0-10, 0 meaning no trust at all and 10 meaning a great deal of trust. Adhikari et al...

  84. [2000]

    https://journals.lww.com/spinejournal/fulltext/ 2000/12150/guidelines_for_the_process_of_cross_cultural.14.aspx

    Guidelines for the Process of Cross-Cultural Adaptation of Self-Report Measures.Spine25, 24 (2000). https://journals.lww.com/spinejournal/fulltext/ 2000/12150/guidelines_for_the_process_of_cross_cultural.14.aspx

  85. [2023]

    InProceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM ’23)

    Application and Evaluation of Large Language Models for the Generation of Survey Questions. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM ’23). Association for Computing Machinery, New York, NY, USA, 5244–5245. doi:10.1145...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.