{"id":"d99d1189-5175-4074-b448-40737a914971","arxiv_id":"2501.05985","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLM-generated survey questions were rated as clear and specific, while LLM-based pretesting improved some adapted questions but often made original questions wordier and less clear.","lead":"Researchers tested whether large language models can write and adapt survey questionnaires by building a pipeline that generates questions, simulates pilot interviews with fictional personas, and revises the text. In two crowdsourced studies with 356 participants, people rated the AI-generated wording as generally clear and the AI-adapted climate questions as slightly clearer and less biased than the original U.S. version.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pretesting validity is untested: the persona-simulated pilot is a closed GPT-4o loop, and RQ2's only positive evidence compares five self-selected rewritten items, so the central adaptation claim needs an external check.","rationale":"The paper's central contribution is not simply text generation; it is the claim that an LLM-simulated pilot can pretest questionnaires and that this pretesting improves adaptation. For that claim to hold, the LLM's simulated personas must produce feedback that is at least roughly aligned with how the actual target population interprets the items. The manuscript offers only one piece of direct evidence for this, the §3.2 feasibility study, and it is reported in a single sentence with no quantitative agreement metric. The main studies do not close the loop because they evaluate the pretesting pipeline's outputs, not its diagnostic validity. In RQ2, only the five questions the LLM flagged as problematic were compared; this makes the evaluation hinge on the very assumption at issue. If the LLM had flagged questions at random or for stylistic reasons, the same evaluation could show \"improvements\" on the rewritten items. The fact that RQ1 shows pretesting reducing clarity for both experts and Prolific participants is additional evidence that the feedback from simulated personas is not reliably aligned with human judgments. I therefore do not see grounds for accepting the central claim as established; the conditional acceptance recommended by the reader is appropriate, with the external validation check as the condition.\n\nThe paper has independent strengths: the code and prompts are open, the two Prolific studies are substantial for this area, and the paper candidly reports the negative RQ1 result and lists limitations including the small expert sample and LLM biases. These do not resolve the validity gap in the pretesting module. No formal verification exists, so the empirical check is the appropriate route.","tokens_in":25572,"tokens_out":6843,"duration_ms":70761,"concrete_test":"Conduct standard cognitive interviews with 10–15 South African respondents on the original 30-item climate questionnaire before revealing any LLM output; code the interview transcripts for comprehension problems using the same issue taxonomy as the reviewer prompt (A.7). Compare the human-flagged problem questions and stated reasons against the LLM's five flagged items and its rationale, computing precision, recall, and Cohen's kappa at the question and issue-type level. If agreement is low (e.g., kappa < 0.4 or fewer than half of the LLM's flagged issues are independently confirmed), the simulated pilot does not capture real respondent interpretation, and the RQ2 clarity/bias improvements must be re-attributed to LLM rewriting rather than valid pretesting. If agreement is high, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing component is the simulated pilot (§3.1.2, Appendix A.5–A.7): GPT-4o generates the personas, plays the respondents, and reviews the transcripts, with no external criterion tying its flagged problems to how real South African or US respondents interpret the items. The formative check in §3.2 only says the pipeline \"detected issues in 5 out of 13\" GESIS questions; it reports no precision, recall, or comparison with the human pretest's issue list, so it does not validate the module.\n\nThis matters because the RQ2 result, the main support for the adaptation claim, is built from the five questions the LLM chose to modify (Tables 5 and 7, Figure 9). Those items were not independently established to be problematic for the target audience, so the observed small gains in clarity and bias (Figure 8, p<0.1) could be generic rewriting effects rather than evidence that persona-based pretesting identifies real comprehension problems. RQ1's data are consistent with this worry: pretesting decreased rated clarity for both Prolific participants and experts (§4.2.1–4.2.2), and experts rated pretested items as more biased. The paper's own limitations (§7) note LLM biases and misrepresentation of identity groups, but the main pipeline is deployed without quantifying that risk against a human pretest benchmark.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an LLM-based pipeline (GPT-4o) for questionnaire generation and pretesting. The pipeline retrieves context from news summaries and SQP questions, generates personas of target respondents, simulates pilot interviews, and uses a reviewer LLM to propose question revisions. The authors then evaluate the pipeline in two studies: RQ1 uses a political-trust questionnaire rated by 238 Prolific US participants and 13 experts; RQ2 adapts a U.S. climate-change survey for a South African audience and evaluates the five items the LLM chose to modify with 118 Prolific participants. The reported findings are that LLM-generated items were rated clearer and LLM-pretested items more specific in RQ1 (on Prolific), and that in RQ2 the adapted items were rated marginally clearer and less biased than the original items. The paper concludes that LLMs are promising for questionnaire generation and adaptation, but that pretesting did not improve the political-trust questionnaire and that human oversight remains necessary.","tokens_in":25880,"tokens_out":6389,"duration_ms":59034,"significance":"If the findings hold, this is a useful exploratory contribution to a small but growing literature on LLM-assisted survey methodology. The paper provides public prompts and code, uses human raters including domain experts, and evaluates items on multiple criteria with both direct and reversed statements. The two Prolific studies have reasonable sample sizes for exploratory work. However, the central validity of the simulated-pilot pretesting step is untested, the headline effects are small and mostly marginal, and RQ1 lacks a human-written baseline. The significance is therefore real but bounded; the paper is best read as an existence proof of the pipeline's feasibility rather than as evidence that LLM-pretesting generally improves questionnaires.","major_comments":[{"comment":"The pretesting module is a closed GPT-4o loop: the same model family generates the personas, simulates the respondent interviews, and writes the reviewer feedback that produces the revised items. Section 3.2's formative check only reports that the pipeline 'detected issues in 5 out of the 13' GESIS questions, with no precision, recall, or comparison to the human pretest's issue list, so it does not validate the module. Because the RQ2 result (Section 5.2, Figure 8) is built from the five items the LLM selected to modify, the observed small gains in clarity and bias could be generic rewriting effects rather than evidence that persona-based pretesting identifies real comprehension problems. The manuscript needs an external anchor—e.g., a human cognitive-interview issue list or a randomized control where rewrites of unproblematic items are also evaluated.","section":"§3.1.2, §3.2, Appendix A.5–A.7"},{"comment":"Table 8 reports 12 paired t-tests without any multiple-comparison correction, and the two headline RQ2 effects (clarity and bias) are only marginally significant (p<0.1). Under a standard correction such as Benjamini-Hochberg at α=0.05, these p-values would not survive. The abstract and Section 5.2 state these as findings; the manuscript should either apply a correction, report effect sizes and confidence intervals, or explicitly label the results as exploratory. The same issue affects the RQ1 clarity and specificity claims in Sections 4.2.1 and 4.2.2.","section":"Table 8, §4.2, §5.2"},{"comment":"RQ1 has no human-written baseline. The paper claims in Section 6.1.1 that LLMs 'have the potential to generate relevant and specific questions,' but the comparison is only between LLM-generated and LLM-pretested versions, both produced by the same pipeline. Without a comparison against a human-authored questionnaire (e.g., the WVS module that inspired the items), the generation claim is unsupported. The abstract's statement that 'LLM-generated text clearer' is a relative comparison to the LLM-pretested version, not an absolute quality assessment.","section":"§4.1–§4.2, §6.1.1"}],"minor_comments":[{"comment":"The abstract and Section 6.1.1 state that LLM-pretested text was 'more specific,' but this held only for Prolific participants; the 13 experts rated the pretested text as less specific (Figure 6). The claim should be qualified to the Prolific sample.","section":"§4.2.1–§4.2.2, Abstract"},{"comment":"The number of personas generated for the South African pilot is not reported, whereas RQ1 explicitly states 50 personas. This matters for reproducibility and for assessing the diversity of the simulation.","section":"§5.1.1"},{"comment":"Table 3 is hard to read: the layout does not clearly separate the direct and reversed evaluation statements for each criterion. Consider using one row per criterion with two columns for the two statements.","section":"Table 3"},{"comment":"The expert sample (N=13) is acknowledged as small in Section 7, but Section 4.2.2 states the expert findings without hedging. Add a cautionary note in the results section itself, not only in the limitations.","section":"§7, §4.2.2"},{"comment":"The conclusion states that LLMs can make questionnaires 'more explicit, unbiased, and relevant,' but relevance was not significantly different in RQ2 (p>0.1). This overstates the results; revise to reflect the actual findings.","section":"§8"},{"comment":"The paper uses 'pretesting' and 'pre-testing' interchangeably. Standardize to one form.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable exploratory contribution, and the public release of prompts and code is commendable. The main barrier is the missing external validation of the simulated-pilot pretesting, which the RQ2 claims depend on. The marginal p-values and lack of multiple-comparison correction should be addressed. If the authors can add a human-baseline comparison for RQ1 and a validity check for the pretesting module (or at least substantially soften the claims), the paper would be publishable. The fit with a CUI journal is acceptable given the conversational nature of the simulated pilot."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, this is a useful empirical snapshot, but the central pretesting claim is not yet established. The paper does something real: it runs two Prolific studies with human raters (238 US, 118 South Africa), reports the numbers, and releases prompts, code, and data. The RAG-based generation component is a sensible extension of prior work on specificity. And the RQ1 result, where pretesting actually reduced rated clarity, is a genuinely useful cautionary finding about simulated pilots.\n\nThe problem is the pretesting module itself. GPT-4o generates the personas, plays the respondents, and reviews the transcripts—a closed loop. The formative check against the GESIS pretest dataset only says it detected issues in 5 of 13 questions; there is no precision or recall against the human pretest's issue list, so it does not validate the module. In RQ2, the only positive evidence for adaptation comes from five questions that the LLM chose to modify. Those items were not independently shown to be problematic for a South African audience, so the small gains in clarity and bias (both p<0.1) could just be generic rewriting effects rather than evidence that the simulated pilot identified real comprehension problems. This worry is consistent with RQ1, where pretesting made questions less clear to both Prolific participants and experts.\n\nThere are also minor statistical issues: no multiple-comparison correction across the eight t-tests, the expert sample is N=13, and RQ2 has no non-pretested LLM-adapted arm. The paper is honest about most of this in Section 7, which I appreciate, but the limitations don't change the fact that the load-bearing claim—that LLM-simulated pilots improve questionnaires—needs an external check.\n\nWho gets value from this: survey methodologists who want to see an actual attempt at LLM-based pretesting, and HCI/practitioners building LLM tools for questionnaire design. I would not cite it as evidence that persona-based pretesting works, but I would cite it as an early empirical effort with transparent methods.\n\nRecommendation: send it to peer review. It is a legitimate empirical study with clear methods, reproduced artifacts, and modest claims. The right review will ask for an external benchmark of the pretesting module, a human-generated baseline, and better statistical hygiene. Those are revision-level asks, not a desk reject.","headline":"A transparent early-stage empirical study of LLM questionnaire generation and adaptation; the generation part holds up, but the pretesting claim rests on a self-referential loop that needs external validation.","tokens_in":26400,"tokens_out":2551,"would_cite":true,"duration_ms":24768,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an LLM pipeline can generate clear, specific survey questions and adapt existing questionnaires for new audiences, with human raters finding LLM-adapted items slightly clearer and less biased than the original U.S.","keywords":["LLM questionnaire generation","survey pretesting","persona simulation","cross-cultural questionnaire adaptation","cognitive interviewing","retrieval-augmented generation","GPT-4o","survey methodology"],"falsifier":"Compare LLM-suggested rewrites against rewrites from human cognitive interviews on the same original questionnaire: if a substantial share of the LLM's flagged problems are not raised by any human respondent, or if the LLM's rewrites are rated lower on clarity by the target audience, the pretesting component's value is falsified.","tokens_in":25392,"feed_emoji":"📋","tokens_out":4713,"duration_ms":41765,"temperature":0.7,"pith_summary":"The paper asks whether large language models can take over two labor-intensive parts of survey design: writing new questionnaires and adapting validated ones for a different culture or country. It builds a pipeline that generates questions with retrieved context, invents fictional target-audience personas, runs a simulated pilot interview, and has a second LLM review the transcripts to suggest rewrites. In two online studies, participants rated LLM-generated items as clear and specific, but pretesting through simulated personas made items less clear in the creation task; in the adaptation task, the LLM's rewrites of a U.S. climate-change survey were seen as marginally clearer and less biased by South African raters. The authors conclude that LLMs are useful for generating and adapting questionnaires, while automated pretesting needs human oversight.","feed_headline":"LLM pipeline drafts survey questions and adapts them across cultures","feed_subtitle":"Two online studies find LLM-generated items clear and specific, and LLM-adapted items slightly clearer than the originals.","key_machinery":"The pipeline's central object is the simulated pilot study. The system first generates a questionnaire using a research question, ten retrieved questions from a standardized-survey database, and a summary of recent news articles; it also invents a list of personas representing the target audience. An interviewer LLM then administers the questionnaire to persona-conditioned participant LLMs, asking cognitive-interviewing follow-up questions. A reviewer LLM reads the transcripts and returns suggested rewrites for questions where personas showed confusion, ambiguity, or bias. The rewrites are the output that gets compared against the originals in the human evaluation.","core_discovery":"The central claim is that a pipeline combining retrieval-augmented generation with an LLM-simulated pilot study can produce questionnaires that are relevant and specific, and can adapt a standardized questionnaire to a new target audience. On the creation side, the paper reports that LLM-generated statements were rated clearer than LLM-pretested statements by both lay raters and experts, while pretesting improved specificity for lay raters; on the adaptation side, the LLM-adapted version of a U.S. climate-change questionnaire was rated marginally clearer and less biased than the original by South African participants. The authors interpret this as evidence that LLM pretesting adds little when generating from scratch but helps when recontextualizing existing instruments.","pith_inferences":["A testable extension is to run the same pipeline with human cognitive interviews instead of LLM personas; if the simulated pilot does not flag the same problems human respondents do, the pretesting step should be skipped for new questionnaires.","The clarity/specificity tradeoff the paper observes suggests that LLM pretesting might be tuned for particular question types: rewrites that add concrete examples may help adaptation but hurt brevity-sensitive items.","The pipeline's dependence on GPT-4o raises the question of whether cheaper open models would show the same pattern, since the 'pretesting made it less clear' result may be a property of the specific LLM's revision style.","One could build a cost-aware variant that only asks for LLM pretesting when a human pilot is infeasible, since the paper's adaptation gains are small and its creation gains are negative."],"forward_implications":["If the pipeline works as claimed, researchers can quickly generate draft questionnaires with current-event context and standard-question grounding, cutting the time to a first testable draft.","For cross-cultural adaptation, LLM pretesting could flag questions whose assumptions (e.g., 'tax rebate', 'Congress') do not transfer, and produce rewrites that real raters find at least as clear.","The mixed creation results imply that automated pretesting should be applied selectively; generated questions may already be clear enough that further LLM revision reduces clarity.","The evaluation design, with paired clarity/bias/relevance/specificity ratings, gives survey researchers a reusable template for judging LLM-assisted instruments."],"supporting_citations":[{"why":"Supplies the relevance and specificity evaluation metrics and the sequential question generation framework that the pipeline extends.","marker":"[44]"},{"why":"Provides the standardized U.S. climate-change public opinion survey used as the baseline questionnaire for the adaptation study.","marker":"[52]"},{"why":"Supplies the World Values Survey political trust module that the generation study is inspired by and bases its question structure on.","marker":"[27]"},{"why":"Demonstrates that persona-conditioned LLMs can simulate responses from demographic subgroups, justifying the simulated pilot approach.","marker":"[2]"},{"why":"Proposes using LLMs to generate 'silicon samples' for pretesting, the direct precedent for the pipeline's simulated participants.","marker":"[66]"},{"why":"Argues that LLMs can support the pretesting of survey instruments, motivating the automated pretesting step.","marker":"[34]"},{"why":"Shows that LLMs can review survey questions to identify potential issues, the basis for the reviewer LLM in the pipeline.","marker":"[57]"},{"why":"Provides evidence that LLMs can generate high-quality questions and discusses anchoring risks, used in the paper's explanation of why generation may not need pretesting.","marker":"[64]"}],"fun_headline_variants":["LLM pipeline writes and adapts surveys across cultures","AI drafts survey questions, adapts them for new audiences","LLM-generated surveys rated clearer and more specific","Automated questionnaire design with LLMs shows promise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the LLM-simulated pilot interviews, conducted with fictional personas, produce feedback about how a real target population will understand and react to the questions; if those simulated reactions diverge from real respondents, the pretesting step can make questions worse.","fun_headline_variants_meta":{"raw":{"variants":["LLM pipeline writes and adapts surveys across cultures","AI drafts survey questions, adapts them for new audiences","LLM-generated surveys rated clearer and more specific","Automated questionnaire design with LLMs shows promise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1159,"prompt_tokens":839,"completion_tokens":320,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":258}},"tokens_in":455,"tokens_out":320,"duration_ms":3290,"temperature":1.0,"reasoning_tokens":258,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:05:25.197090+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare LLM-suggested rewrites against rewrites from human cognitive interviews on the same original questionnaire: if a substantial share of the LLM's flagged problems are not raised by any human respondent, or if the LLM's rewrites are rated lower on clarity by the target audience, the pretesting component's value is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the relevance and specificity evaluation metrics and the sequential question generation framework that the pipeline extends."},{"cited_title":"ChatGPTest: opportunities and cautionary tales of utilizing AI for questionnaire pretesting","cited_arxiv_id":"2405.06329","evidence_quote":"Shows that LLMs can review survey questions to identify potential issues, the basis for the reviewer LLM in the pipeline."}],"review_version":1}