REVIEW 3 major objections 6 minor 1 cited by
Exploring LLMs for Automated Generation and Adaptation of Questionnaires
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that an LLM pipeline can generate clear, specific survey questions and adapt existing questionnaires for new audiences, with human raters finding LLM-adapted items slightly clearer and less biased than the original U.S.
desk verdict A transparent early-stage empirical study of LLM questionnaire generation and adaptation; the generation part holds up, but the pretesting claim rests on a self-referential loop that needs external validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The pipeline's central object is the simulated pilot study. The system first generates a questionnaire using a research question, ten retrieved questions from a standardized-survey database, and a summary of recent news articles; it also invents a list of personas representing the target audience. An interviewer LLM then administers the questionnaire to persona-conditioned participant LLMs, asking cognitive-interviewing follow-up questions. A reviewer LLM reads the transcripts and returns suggested rewrites for questions where personas showed confusion, ambiguity, or bias. The rewrites are the output that gets compared against the originals in the human evaluation.
What would settle it
Compare LLM-suggested rewrites against rewrites from human cognitive interviews on the same original questionnaire: if a substantial share of the LLM's flagged problems are not raised by any human respondent, or if the LLM's rewrites are rated lower on clarity by the target audience, the pretesting component's value is falsified.
Extended reading notes
Core claim
The central claim is that a pipeline combining retrieval-augmented generation with an LLM-simulated pilot study can produce questionnaires that are relevant and specific, and can adapt a standardized questionnaire to a new target audience. On the creation side, the paper reports that LLM-generated statements were rated clearer than LLM-pretested statements by both lay raters and experts, while pretesting improved specificity for lay raters; on the adaptation side, the LLM-adapted version of a U.S. climate-change questionnaire was rated marginally clearer and less biased than the original by South African participants. The authors interpret this as evidence that LLM pretesting adds little when generating from scratch but helps when recontextualizing existing instruments.
Load-bearing premise
The load-bearing premise is that the LLM-simulated pilot interviews, conducted with fictional personas, produce feedback about how a real target population will understand and react to the questions; if those simulated reactions diverge from real respondents, the pretesting step can make questions worse.
Editorial extensions
If this is right
- If the pipeline works as claimed, researchers can quickly generate draft questionnaires with current-event context and standard-question grounding, cutting the time to a first testable draft.
- For cross-cultural adaptation, LLM pretesting could flag questions whose assumptions (e.g., 'tax rebate', 'Congress') do not transfer, and produce rewrites that real raters find at least as clear.
- The mixed creation results imply that automated pretesting should be applied selectively; generated questions may already be clear enough that further LLM revision reduces clarity.
- The evaluation design, with paired clarity/bias/relevance/specificity ratings, gives survey researchers a reusable template for judging LLM-assisted instruments.
Reading between the lines
- A testable extension is to run the same pipeline with human cognitive interviews instead of LLM personas; if the simulated pilot does not flag the same problems human respondents do, the pretesting step should be skipped for new questionnaires.
- The clarity/specificity tradeoff the paper observes suggests that LLM pretesting might be tuned for particular question types: rewrites that add concrete examples may help adaptation but hurt brevity-sensitive items.
- The pipeline's dependence on GPT-4o raises the question of whether cheaper open models would show the same pattern, since the 'pretesting made it less clear' result may be a property of the specific LLM's revision style.
- One could build a cost-aware variant that only asks for LLM pretesting when a human pilot is infeasible, since the paper's adaptation gains are small and its creation gains are negative.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an LLM-based pipeline (GPT-4o) for questionnaire generation and pretesting. The pipeline retrieves context from news summaries and SQP questions, generates personas of target respondents, simulates pilot interviews, and uses a reviewer LLM to propose question revisions. The authors then evaluate the pipeline in two studies: RQ1 uses a political-trust questionnaire rated by 238 Prolific US participants and 13 experts; RQ2 adapts a U.S. climate-change survey for a South African audience and evaluates the five items the LLM chose to modify with 118 Prolific participants. The reported findings are that LLM-generated items were rated clearer and LLM-pretested items more specific in RQ1 (on Prolific), and that in RQ2 the adapted items were rated marginally clearer and less biased than the original items. The paper concludes that LLMs are promising for questionnaire generation and adaptation, but that pretesting did not improve the political-trust questionnaire and that human oversight remains necessary.
Significance. If the findings hold, this is a useful exploratory contribution to a small but growing literature on LLM-assisted survey methodology. The paper provides public prompts and code, uses human raters including domain experts, and evaluates items on multiple criteria with both direct and reversed statements. The two Prolific studies have reasonable sample sizes for exploratory work. However, the central validity of the simulated-pilot pretesting step is untested, the headline effects are small and mostly marginal, and RQ1 lacks a human-written baseline. The significance is therefore real but bounded; the paper is best read as an existence proof of the pipeline's feasibility rather than as evidence that LLM-pretesting generally improves questionnaires.
major comments (3)
- [§3.1.2, §3.2, Appendix A.5–A.7] The pretesting module is a closed GPT-4o loop: the same model family generates the personas, simulates the respondent interviews, and writes the reviewer feedback that produces the revised items. Section 3.2's formative check only reports that the pipeline 'detected issues in 5 out of the 13' GESIS questions, with no precision, recall, or comparison to the human pretest's issue list, so it does not validate the module. Because the RQ2 result (Section 5.2, Figure 8) is built from the five items the LLM selected to modify, the observed small gains in clarity and bias could be generic rewriting effects rather than evidence that persona-based pretesting identifies real comprehension problems. The manuscript needs an external anchor—e.g., a human cognitive-interview issue list or a randomized control where rewrites of unproblematic items are also evaluated.
- [Table 8, §4.2, §5.2] Table 8 reports 12 paired t-tests without any multiple-comparison correction, and the two headline RQ2 effects (clarity and bias) are only marginally significant (p<0.1). Under a standard correction such as Benjamini-Hochberg at α=0.05, these p-values would not survive. The abstract and Section 5.2 state these as findings; the manuscript should either apply a correction, report effect sizes and confidence intervals, or explicitly label the results as exploratory. The same issue affects the RQ1 clarity and specificity claims in Sections 4.2.1 and 4.2.2.
- [§4.1–§4.2, §6.1.1] RQ1 has no human-written baseline. The paper claims in Section 6.1.1 that LLMs 'have the potential to generate relevant and specific questions,' but the comparison is only between LLM-generated and LLM-pretested versions, both produced by the same pipeline. Without a comparison against a human-authored questionnaire (e.g., the WVS module that inspired the items), the generation claim is unsupported. The abstract's statement that 'LLM-generated text clearer' is a relative comparison to the LLM-pretested version, not an absolute quality assessment.
minor comments (6)
- [§4.2.1–§4.2.2, Abstract] The abstract and Section 6.1.1 state that LLM-pretested text was 'more specific,' but this held only for Prolific participants; the 13 experts rated the pretested text as less specific (Figure 6). The claim should be qualified to the Prolific sample.
- [§5.1.1] The number of personas generated for the South African pilot is not reported, whereas RQ1 explicitly states 50 personas. This matters for reproducibility and for assessing the diversity of the simulation.
- [Table 3] Table 3 is hard to read: the layout does not clearly separate the direct and reversed evaluation statements for each criterion. Consider using one row per criterion with two columns for the two statements.
- [§7, §4.2.2] The expert sample (N=13) is acknowledged as small in Section 7, but Section 4.2.2 states the expert findings without hedging. Add a cautionary note in the results section itself, not only in the limitations.
- [§8] The conclusion states that LLMs can make questionnaires 'more explicit, unbiased, and relevant,' but relevance was not significantly different in RQ2 (p>0.1). This overstates the results; revise to reflect the actual findings.
- [Throughout] The paper uses 'pretesting' and 'pre-testing' interchangeably. Standardize to one form.
Circularity Check
No significant circularity: the main claims rest on independent human ratings, not on LLM outputs alone.
full rationale
The paper's load-bearing findings (RQ1 and RQ2) are evaluated by 356 Prolific participants and 13 experts using Likert ratings and forced-choice comparisons of original versus LLM-pretested or adapted items, not by the same LLM that produced the revisions. The simulated-pilot module (Section 3.1.2) is a generation step whose output could be and was contradicted by human raters: RQ1 experts rated pretested statements as less clear and more biased, so the empirical claims are falsifiable and not forced by construction. The formative feasibility study (Section 3.2) reports that the pipeline 'detected issues in 5 out of the 13 tested questions' without reporting precision or recall against the GESIS human pretest; this is a validation gap, not a circular reduction. No load-bearing self-citations or imported uniqueness theorems appear. Accordingly, no step reduces by definition to its inputs.
Assumptions & free parameters
free parameters (4)
- Number of retrieved news articles =
5
- Number of SQP questions retrieved =
10
- Number of personas in political trust pretest =
50
- Temperature and random seed =
not reported for main studies
assumptions (5)
- domain assumption LLM-simulated personas can represent target-audience interpretation patterns
- domain assumption Prolific participants with domain degrees are valid judges of questionnaire clarity, bias, relevance, and specificity
- standard math Likert ratings can be treated as interval data for paired t-tests
- domain assumption The source instruments (WVS political trust module and Yale climate change survey) are appropriate standardized benchmarks
- domain assumption GPT-4o thematic analysis of open-ended responses reliably categorizes participant reasoning
invented entities (1)
-
Fictional personas for simulated pilot participants
Cite this review
Pith. "Pith review of Exploring LLMs for Automated Generation and Adaptation of Questionnaires." pith.science (2026). https://pith.science/paper/KUEUHITH
@misc{pith2026250105985,
author = {Pith},
title = {Pith review of: Exploring LLMs for Automated Generation and Adaptation of Questionnaires},
year = {2026},
howpublished = {\url{https://pith.science/paper/KUEUHITH}},
note = {Machine review of arXiv:2501.05985}
}
read the original abstract
Effective questionnaire design improves the validity of the results, but creating and adapting questionnaires across contexts is challenging due to resource constraints and limited expert access. Recently, the emergence of LLMs has led researchers to explore their potential in survey research. In this work, we focus on the suitability of LLMs in assisting the generation and adaptation of questionnaires. We introduce a novel pipeline that leverages LLMs to create new questionnaires, pretest with a target audience to determine potential issues and adapt existing standardized questionnaires for different contexts. We evaluated our pipeline for creation and adaptation through two studies on Prolific, involving 238 participants from the US and 118 participants from South Africa. Our findings show that participants found LLM-generated text clearer, LLM-pretested text more specific, and LLM-adapted questions slightly clearer and less biased than traditional ones. Our work opens new opportunities for LLM-driven questionnaire support in survey research.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research
A work-in-progress systematic review finds LLM use in survey research clusters in instrument development, synthetic respondent modeling, and automated text classification, leaving interviewing and cross-lingual work thin.
Reference graph
Works this paper leans on
-
[1]
Noorhan Abbas, Thomas Pickard, Eric Atwell, and Aisha Walker. 2021. University Student Surveys Using Chatbots: Artificial Intelligence Conversational Agents. InLearning and Collaboration Technologies: Games and Virtual Environments for Learning, Panayiotis Zaphiris and Andri Ioannou (Eds.). Springer International Publishing, Cham, 155–169. doi:10.1007/978...
-
[2]
Argyle, Ethan C
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples.Political Analysis31, 3 (July 2023), 337–351. doi:10. 1017/pan.2023.2
2023
-
[3]
Beaton, Claire Bombardier, Francis Guillemin, and Marcos Bosi Ferraz
Dorcas E. Beaton, Claire Bombardier, Francis Guillemin, and Marcos Bosi Ferraz
-
[4]
Beaton, Claire Bombardier, Francis Guillemin, and Mar- cos Bosi Ferraz
Dorcas E. Beaton, Claire Bombardier, Francis Guillemin, and Mar- cos Bosi Ferraz. 2000. Guidelines for the Process of Cross-Cultural Adaptation of Self-Report Measures.Spine25, 24 (Dec. 2000), 3186. https://journals.lww.com/spinejournal/pages/articleviewer.aspx?year= 2000&issue=12150&article=00014&type=Fulltext
2000
-
[5]
Dorothée Behr, Michael Braun, and Luisa Aiglstorfer. 2024. Showcasing the usefulness of web probing: Do subtle variations in questionnaire translation lead to different survey responding?Quality & Quantity(March 2024). doi:10.1007/ s11135-024-01843-8
2024
-
[6]
Dorothée Behr, Katharina Meitinger, Michael Braun, and Lars Kaczmirek. 2017. Web probing – implementing probing techniques from cognitive, interviewing Exploring LLMs for Automated Generation and Adaptation of Questionnaires in web surveys with the goal to assess the validity of survey questions (GESIS Survey Guidelines)Web probing – implementing probing ...
2017
-
[7]
Johnny Blair, Allison Ackermann, Linda Piccinino, and Rachel Levenstein. 2007. Using Behavior Coding to Validate Cognitive Interview Findings. (Jan. 2007)
2007
- [8]
Show all 93 references
-
[9]
Johnny Blair and Linda Piccinino. 2005. The development and testing of instru- ments for cross-cultural and multi-cultural surveys. InMethodological aspects in cross-national research, Jürgen H. P. Hoffmeyer-Zlotnik and Janet Harkness (Eds.). ZUMA-Nachrichten Spezial, Vol. 11....
2005
-
[10]
Petra M Boynton and Trisha Greenhalgh. 2004. Selecting, designing, and devel- oping your questionnaire.BMJ : British Medical Journal328, 7451 (May 2004), 1312–1315. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC420179/
2004
-
[11]
Zi Chai and Xiaojun Wan. 2020. Learning to Ask More: Semi-Autoregressive Sequential Question Generation under Dual-Graph Interaction. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel...
2020 doi
-
[12]
Roberts Dar ‘gis, Guntis B ¯arzdin, š, Inguna Skadi n, a, Normunds Gr ¯uz¯ıtis, and Baiba Saul¯ıte. 2024. Evaluating Open-Source LLMs in Low-Resource Lan- guages: Insights from Latvian High School Exams. InProceedings of the 4th International Conference on Natural Language Pro...
2024 doi
-
[13]
Elizabeth Dean, Brian Head, and Jodi Swicegood. 2013. Virtual Cognitive In- terviewing Using Skype and Second Life. InSocial Media, Sociality, and Survey Research. 107–132. doi:10.1002/9781118751534.ch5 Journal Abbreviation: Social Media, Sociality, and Survey Research
2013 doi
-
[14]
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023. Can AI language models replace human participants?Trends in Cognitive Sciences27, 7 (July 2023), 597–600. doi:10.1016/j.tics.2023.04.008 Publisher: Elsevier
2023 doi
- [15]
-
[16]
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The Faiss library. (2024). arXiv:2401.08281 [cs.LG]
2024 arXiv
-
[17]
Jonathan Drennan. 2003. Cognitive interviewing: verbal data in the design and pretesting of questionnaires.Journal of Advanced Nursing42, 1 (2003), 57–63. doi:10.1046/j.1365-2648.2003.02579.x arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1046/j.1365-2648.2003.02579.x
2003
-
[18]
Osborne, Gerald R
Jonathan Epstein, Richard H. Osborne, Gerald R. Elsworth, Dorcas E. Beaton, and Francis Guillemin. 2015. Cross-cultural adaptation of the Health Education Impact Questionnaire: experimental study showed expert committee, not back- translation, added value.Journal of Clinical E...
2015 doi
-
[19]
Jonathan Epstein, Ruth Miyuki Santo, and Francis Guillemin. 2015. A review of guidelines for cross-cultural adaptation of questionnaires could not bring out a consensus.Journal of Clinical Epidemiology68, 4 (April 2015), 435–441. doi:10.1016/j.jclinepi.2014.11.021
2015 doi
-
[20]
Weinstein, Vasil I
Jennifer Fei, Jessica Wolff, Michael Hotard, Hannah Ingham, Saurabh Khanna, Duncan Lawrence, Beza Tesfaye, Jeremy M. Weinstein, Vasil I. Yasenov, and Jens Hainmueller. 2022. Automated Chat Application Surveys Using Whatsapp: Evidence from Panel Surveys and a Mode Experiment. d...
2022 doi
-
[21]
Yifan Gao, Piji Li, Irwin King, and Michael R. Lyu. 2019. Interconnected Question Generation with Coreference Alignment and Conversation Flow Modeling. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and L...
2019 doi
- [22]
-
[23]
Linn Gjersing, John RM Caplehorn, and Thomas Clausen. 2010. Cross-cultural adaptation of research instruments: language, setting, time and statistical considerations.BMC Medical Research Methodology10, 1 (10 Feb 2010), 13. doi:10.1186/1471-2288-10-13
2010 doi
- [24]
-
[25]
Melissa Guyre, Liz Holland, Nirva Shah, and Rahul R. Divekar. 2024. Prompt Engineering an LLM into Roleplaying a Management Coach: a Short Guide by and for Non-NLP Experts. InProceedings of the 6th ACM Conference on Conversational User Interfaces (CUI ’24). Association for Com...
2024
- [26]
-
[27]
Christian Haerpfer, Ronald Inglehart, Alejandro Moreno, Christian Welzel, Kseniya Kizilova, Jaime Diez-Medrano, Marta Lagos, Pippa Norris, Eduard Ponarin, and Bi Puranen. 2020. World Values Survey wave 7 (2017-2020) cross- national data-set
2020
-
[28]
Anne-Wil Harzing. 2005. Does the Use of English-language Questionnaires in Cross-national Research Obscure National Differences?International Jour- nal of Cross Cultural Management5, 2 (Aug. 2005), 213–224. doi:10.1177/ 1470595805054494 Publisher: SAGE Publications
2005
-
[29]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI generates covertly racist decisions about people based on their dialect.Nature 633, 8028 (01 Sep 2024), 147–154. doi:10.1038/s41586-024-07856-5
2024 doi
-
[30]
John J. Horton. 2023. Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? doi:10.48550/arXiv.2301.07543 arXiv:2301.07543 [econ, q-fin]
2023 doi
-
[31]
Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case Study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, N...
2023
-
[32]
Sviatlana Höhn, Jauwairia Nasir, Ali Paikan, Pouyan Ziafati, and Elisabeth André
-
[33]
Nashwa Ismail, Gary Kinchin, and Julie-Ann Edwards. 2018. Pilot study, Does it really matter? Learning lessons from conducting a pilot study for a qualitative PhD thesis.International Journal of Social Science Research6, 1 (March 2018), 1–17. doi:10.5296/ijssr.v6i1.11720 Num P...
2018 doi
-
[34]
Jansen, Soon-gyo Jung, and Joni Salminen
Bernard J. Jansen, Soon-gyo Jung, and Joni Salminen. 2023. Employing large language models in survey research.Natural Language Processing Journal4 (Sept. 2023), 100020. doi:10.1016/j.nlp.2023.100020
2023
-
[35]
Stephen Jenkins and Tony Solomonides. 2000. Automating Questionnaire Design and Construction.International Journal of Market Research42, 1 (Jan. 2000), 1–13. doi:10.1177/147078530004200106
2000 doi
-
[36]
Ng Chirk Jenn. 2006. Designing A Questionnaire.Malaysian Family Physician : the Official Journal of the Academy of Family Physicians of Malaysia1, 1 (April 2006), 32–35. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4797036/
2006
-
[37]
Johanson and Gordon P
George A. Johanson and Gordon P. Brooks. 2010. Initial Scale Development: Sam- ple Size for Pilot Studies.Educational and Psychological Measurement70, 3 (June 2010), 394–400. doi:10.1177/0013164409355692 Publisher: SAGE Publications Inc
2010 doi
-
[38]
Kendra Kamp, Gwen Wyatt, Sharon Dudley-Brown, Kelly Brittain, and Barbara Given. 2018. Using cognitive interviewing to improve questionnaires: An exem- plar study focusing on individual and condition-specific factors.Applied Nursing Research43 (Oct. 2018), 121–125. doi:10.1016...
2018 doi
-
[39]
Junsol Kim and Byungkyu Lee. 2023. AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction. https://arxiv.org/abs/ 2305.09620v2
2023 arXiv
-
[40]
Sunwoong Kim, Jongho Jeong, Jin Soo Han, and Donghyuk Shin. 2024. LLM-Mirror: A Generated-Persona Approach for Survey Pre-Testing. arXiv:2412.03162 [cs.CY] https://arxiv.org/abs/2412.03162
2024 arXiv
-
[41]
Soomin Kim, Joonhwan Lee, and Gahgene Gweon. 2019. Comparing Data from Chatbot and Web Surveys: Effects of Platform and Conversational Style on Survey Response Quality. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19). Association for Co...
2019
-
[42]
Hill, and A Herzog
Bärbel Knäuper, Robert Belli, D. Hill, and A Herzog. 1997. Question Difficulty and Respondents’ Cognitive Ability: The Effect on Data Quality.Journal of Official Statistics13 (Jan. 1997)
1997
-
[43]
Lucrezia Laraspata, Fabio Cardilli, Giovanna Castellano, and Gennaro Vessio. [n. d.]. Enhancing Human Capital Management through GPT-driven Question- naire Generation. ([n. d.])
-
[44]
Yan Lei, Liang Pang, Yuanzhuo Wang, Huawei Shen, and Xueqi Cheng. 2024. Qs- nail: A Questionnaire Dataset for Sequential Question Generation. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2...
2024
-
[45]
Timo Lenzer, Patricia Hadler, and Cornelia Neuert. 2024. Cognitive Pretesting. Adhikari et al
2024
-
[46]
Timo Lenzner. 2012. Effects of Survey Question Comprehensibility on Response Quality.Field Methods24, 4 (Nov. 2012), 409–428. doi:10.1177/1525822X12448166 Publisher: SAGE Publications Inc
2012 doi
-
[47]
Timo Lenzner and Cornelia E. Neuert. 2017. Pretesting Survey Questions Via Web Probing – Does it Produce Similar Results to Face-to-Face Cognitive Inter- viewing?Survey Practice10, 4 (Oct. 2017). doi:10.29115/SP-2017-0020
2017 doi
-
[48]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Gen- eration for Knowledge-Intensive NLP Tasks. InAdvances in...
2020
-
[49]
Xun Liang, Hanyu Wang, Yezhaohui Wang, Shichao Song, Jiawei Yang, Simin Niu, Jie Hu, Dan Liu, Shunyu Yao, Feiyu Xiong, and Zhiyu Li. 2024. Controllable Text Generation for Large Language Models: A Survey. arXiv:2408.12599 [cs.CL] https://arxiv.org/abs/2408.12599
2024 arXiv
-
[50]
Jieli Liu and Haining Wang. 2024. Assessing Gender and Racial Bias in Large Language Model-Powered Virtual Reference.Proceedings of the Association for Information Science and Technology61, 1 (2024), 576–580. doi:10.1002/pra2.1061 arXiv:https://asistdl.onlinelibrary.wiley.com/...
2024 doi
-
[51]
Antonio Maiorino, Zoe Padgett, Chun Wang, Misha Yakubovskiy, and Peng Jiang
-
[52]
Marlon, Xinran Wang, Parrish Bergquist, Peter D
Jennifer R. Marlon, Xinran Wang, Parrish Bergquist, Peter D. Howe, Anthony Leiserowitz, Edward Maibach, Matto Mildenberger, and Seth Rosenthal. 2022. Change in US state-level public opinion about climate change: 2008–2020.Envi- ronmental Research Letters17, 12 (2022). doi:10.1...
2022 doi
-
[53]
Catriona Rachel Mayland, Christina Gerlach, Katrin Sigurdardottir, Marit Irene Tuen Hansen, Wojciech Leppert, Andrzej Stachowiak, Maria Krajew- ska, Eduardo Garcia-Yanneo, Vilma Adriana Tripodoro, Gabriel Goldraij, Mar- tin Weber, Lair Zambon, Juliana Nalin Passarini, Ivete Br...
2019
-
[54]
K Meitinger, C Neuert, C Beitz, and N Menold. 2016. Pretesting of special module on ICT at work, working conditions & learning digital skills
2016
-
[55]
Felix Ndashimye, Oumarou Hebie, and Jasper Tjaden. 2024. Effectiveness of WhatsApp for Measuring Migration in Follow-Up Phone Surveys. Lessons from a Mode Experiment in Two Low-Income Countries during COVID Contact Restrictions.Social Science Computer Review42, 2 (April 2024),...
2024
-
[56]
Cornelia Neuert. 2023. Design of multiple open-ended probes in cognitive online pretests using web probing. (2023). doi:10.13094/SMIF-2023-00005 Publisher: FORS/PUMA/GESIS
2023 doi
- [57]
-
[59]
Kristen Olson. 2010. An Examination of Questionnaire Evaluation by Expert Re- viewers.Field Methods22, 4 (Nov. 2010), 295–318. doi:10.1177/1525822X10379795 Publisher: SAGE Publications Inc
2010 doi
-
[60]
Jessica Parker, Veronica Richard, and Kimberly Becker. 2023. Flexibility & It- eration: Exploring the Potential of Large Language Models in Developing and Refining Interview Protocols.The Qualitative Report28, 9 (Sept. 2023), 2772–2791. doi:10.46743/2160-3715/2023.6695
2023
-
[61]
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Iryna Gurevych and Yusuke Miyao (Eds.). Associati...
2018 doi
-
[62]
Pooja Rao S B, Manish Agnihotri, and Dinesh Babu Jayagopi. 2021. Improv- ing Asynchronous Interview Interaction with Follow-up Question Generation. (March 2021). doi:10.9781/ijimai.2021.02.010 Accepted: 2022-04-25T09:02:55Z Publisher: International Journal of Interactive Multi...
2021 doi
-
[63]
Nina Reynolds, Adamantios Diamantopoulos, and Bodo Schlegelmilch. 1993. Pre-Testing in Questionnaire Design: A Review of the Literature and Suggestions for Further Research.Market Research Society. Journal.35, 2 (March 1993), 1–11. doi:10.1177/147078539303500202 Publisher: SAG...
1993 doi
-
[64]
Rothschild, James Brand, Hope Schroeder, and Jenny Wang
David M. Rothschild, James Brand, Hope Schroeder, and Jenny Wang. 2024. Opportunities and risks of LLMs in survey research. doi:10.2139/ssrn.5001645
2024 doi
-
[65]
David Rozado. 2024. The political preferences of LLMs.PLOS ONE19, 7 (07 2024), 1–15. doi:10.1371/journal.pone.0306621
2024 doi
-
[66]
Adler, Lea Rau, and Bernd Schmitt
Marko Sarstedt, Susanne J. Adler, Lea Rau, and Bernd Schmitt. 2024. Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines.Psychology & Marketing41, 6 (2024), 1254–1270. doi:10.1002/mar.21982
2024 doi
-
[67]
Rothkopf, and Kristian Kersting
Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A. Rothkopf, and Kristian Kersting. 2022. Large pre-trained language models contain human- like biases of what is right and wrong to do.Nature Machine Intelligence4, 3 (March 2022), 258–268. doi:10.1038/s42256-022-00...
2022 doi
-
[68]
Kerry Scott, Dipanwita Gharai, Manjula Sharma, Namrata Choudhury, Bibha Mishra, Sara Chamberlain, and Amnesty LeFevre. 2020. Yes, no, maybe so: the importance of cognitive interviewing to enhance structured surveys on respectful maternity care in northern India.Health Policy a...
2020 doi
- [69]
-
[70]
Vanessa E. C. Sousa, Jeffrey Matson, and Karen Dunn Lopez. 2017. Questionnaire Adapting: Little Changes Mean a Lot.Western Journal of Nursing Research 39, 9 (Sept. 2017), 1289–1300. doi:10.1177/0193945916678212 Publisher: SAGE Publications Inc
2017 doi
-
[71]
John M. Stahura. 2005. Methods for Testing and Evaluating Survey Ques- tionnaires.Contemporary Sociology34, 4 (July 2005), 427–428. doi:10.1177/ 009430610503400457 Publisher: SAGE Publications Inc
2005
-
[72]
Ming-Hsiang Su, Chung-Hsien Wu, Kun-Yi Huang, Qian-Bei Hong, and Huai- Hung Huang. 2018. Follow-up Question Generation Using Pattern-based Seq2seq with a Small Corpus for Interview Coaching. 1006–1010. doi:10.21437/ Interspeech.2018-1007
2018
-
[73]
Guangzhi Sun, Xiao Zhan, and Jose Such. 2024. Building Better AI Agents: A Provocation on the Utilisation of Persona in LLM-based Conversational Agents. InProceedings of the 6th ACM Conference on Conversational User Interfaces (CUI ’24). Association for Computing Machinery, Ne...
2024
-
[74]
Synodinos
Nicolaos E. Synodinos. 2003. The “art” of questionnaire construction: some important considerations for manufacturing studies.Integrated Manufacturing Systems14, 3 (Jan. 2003). doi:10.1108/09576060310463172
2003 doi
-
[75]
Hamed Taherdoost. 2022. Designing a Questionnaire for a Research Paper: A Comprehensive Guide to Design and Develop an Effective Questionnaire. Asian Journal of Managerial Science11, 1 (March 2022), 8–16. doi:10.51983/ajms- 2022.11.1.3087 Number: 1
2022 doi
- [76]
-
[77]
Brown, Jennifer V
Alejandro Cuevas Villalba, Eva M. Brown, Jennifer V. Scurrell, Jason Entenmann, and Madeleine I. G. Daepp. 2023. Automated Interviewer or Augmented Survey? Collecting Social Data with Large Language Models. doi:10.48550/arXiv.2309. 10187
2023 doi
- [78]
-
[79]
W.E. Saris. 2022.SQP 3.0. Universitat Pompeu Fabra, Barcelona, GESIS – Leibniz- Institut für Sozialwissenschaften e.V., Mannheim. https://sqp.gesis.org
2022
-
[80]
Lang, Quirin Würschinger, and Frauke Kreuter
Alexander Wuttke, Matthias Aßenmacher, Christopher Klamm, Max M. Lang, Quirin Würschinger, and Frauke Kreuter. 2024. AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers. arXiv:2410.01824 [cs.HC] https://arxiv.org/abs/2410.01824
2024 arXiv
- [81]
-
[82]
Yifan Yang, Xiaoyu Liu, Qiao Jin, Furong Huang, and Zhiyong Lu. 2024. Un- masking and quantifying racial bias of large language models in medical report generation.Communications Medicine4, 1 (10 Sep 2024), 176. doi:10.1038/s43856- 024-00601-z
2024 doi
-
[83]
Griffith, and Timothy Bickmore
Hye Sun Yun, Mehdi Arjmand, Phillip Raymond Sherlock, Michael Paasche- Orlow, James W. Griffith, and Timothy Bickmore. 2023. Keeping Users Engaged During Repeated Administration of the Same Questionnaire: Using Large Lan- guage Models to Reliably Diversify Questions. doi:10.48...
-
[84]
Zhao Zou, Omar Mubin, Fady Alnajjar, and Luqman Ali. 2024. A pilot study of measuring emotional response and perception of LLM-generated questionnaire and human-generated questionnaires.Scientific Reports14, 1 (Feb. 2024), 2781. doi:10.1038/s41598-024-53255-1 Publisher: Nature...
2024 doi
-
[88]
Summary of the first article </relevant_articles> <relevant_questions> * First relevant question </relevant_questions> Questions:
-
[89]
Why do you like this product? Type: closed-ended Options: Ease of use, Quality, Price, Brand
-
[90]
language
What suggestions do you have for the product? Type: open-ended Options: - Adhikari et al. A.2 Persona generation system prompt You are a researcher working on Survey research. As a part of your research, you need to determine the correct samples to pick from the intended audie...
-
[91]
How you feel about the national congress in your country?
Generate 6 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) on the question "How you feel about the national congress in your country?". Number it from P1 to P6
-
[92]
How do you feel about the government in your country?
Generate 6 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) on the question "How do you feel about the government in your country?". Number it from G1 to G6
-
[93]
How do you feel about the United Nations?
Generate 6 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) on the question "How do you feel about the United Nations?". Number it from UN1 to UN6
-
[94]
Number it from A to O
Generate 15 statements to rate on a likert scale (1-5, 1 meaning Agree strongly and 5 meaning disagree strongly) related to trust on Politicians and the government. Number it from A to O
-
[95]
Disagree
Generate a question asking about their trust on the <head of state>in their country. Replace <head of state>with the respective head of state in the country. Make sure to use a likert scale of 0-10, 0 meaning no trust at all and 10 meaning a great deal of trust. Adhikari et al...
-
[2000]
https://journals.lww.com/spinejournal/fulltext/ 2000/12150/guidelines_for_the_process_of_cross_cultural.14.aspx
Guidelines for the Process of Cross-Cultural Adaptation of Self-Report Measures.Spine25, 24 (2000). https://journals.lww.com/spinejournal/fulltext/ 2000/12150/guidelines_for_the_process_of_cross_cultural.14.aspx
2000
-
[2023]
InProceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM ’23)
Application and Evaluation of Large Language Models for the Generation of Survey Questions. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM ’23). Association for Computing Machinery, New York, NY, USA, 5244–5245. doi:10.1145...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.