{"id":"bbc29df8-4b96-4b3c-bc9a-d92b3600df9b","arxiv_id":"2509.08702","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A preregistered 2x3 factorial experiment on 262 real survey questions shows GPT-4.0 and a survey-expert persona make ChatGPT flag more, and different, survey question problems than GPT-3.5 or no persona.","lead":"This paper ran a controlled experiment asking ChatGPT (two versions, three 'personas') to critique 262 real survey questions, then coded the feedback with a survey methodology taxonomy. It found that model version and persona choices systematically change which problems the AI flags, and argues ChatGPT can serve as a low-cost quality check for survey design.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GPT-4 effect sizes inherit an untested codebook and a single-author human benchmark; ~80% agreement is likely base-rate-driven.","rationale":"The paper's central contribution is that ChatGPT feedback can be mapped onto established survey-problem categories and that model/persona choice changes that feedback. The preregistered factorial design, Holm-corrected multilevel models, and source-based convergent checks (e.g., WVS questions being more likely to elicit Code 10) are genuine strengths. But the quantities carrying the conclusion are code indicators produced with a codebook refined while reading AI output, with only 70% full agreement, and the human benchmark is a single author's coding. Overall agreement of about 80% is uninformative for rare codes without kappa or positive agreement. The largest GPT-4 effects are on the most subjective codes, and the paper itself acknowledges more 'AI-only' flags from GPT-4. So the load-bearing premise is that the code indicators are reliable and the human benchmark is a valid reference; neither is currently demonstrated. This is exactly the weakest assumption the reader identified. If independent coding reproduces the effects, the conditional verdict could move toward acceptance; if not, the abstract's 'valuable tool' claim would need substantial qualification. The treatment-order/time-confounding possibility is a separate secondary concern, but the measurement-validity issue is more fundamental because it affects every code-based outcome regardless of design.","tokens_in":24468,"tokens_out":8783,"duration_ms":104184,"concrete_test":"Using the posted Dataverse replication data, draw a stratified random sample of ~150 treatment-question outputs and have two independent survey methodologists, blinded to treatment, apply the Appendix B codebook; also have them independently code the 262 questions as the human reference. Compute per-code Cohen's kappa and positive agreement between independent coders and between independent codes and the authors' codes, then re-estimate the Section 3.1 multilevel models on the independent codes. If kappa is below 0.6 for Code 9/Code 10, or if the GPT-4 coefficients (+0.14, +0.15) shift by more than ~0.05 or lose significance, the model effects are not robust to coding subjectivity and the Section 3.4.3 validity check is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline estimates in Section 3.1 (GPT-4 produces 0.55 more codes; Code 9 +0.14, Code 10 +0.15, both p<0.001) are only interpretable if the 11-code instrument and the 'human expert' reference are valid. Appendix B shows the codebook was refined abductively on AI output, with categories merged, deleted, and redefined during coding, and the only reported reliability is 70% full agreement; low inter-coder reliability led the authors to double-check all codes. The human benchmark in Section 3.4.3 (footnote 6) is the first author's own coding of all 262 questions using that same codebook. This is not an independent reference standard: agreement between AI-generated codes and this benchmark cannot establish that the AI feedback maps onto real survey problems rather than onto the coding team's expectations. The reported ~80% agreement is also likely dominated by 'Neither' for rare codes (e.g., Code 8 prevalence ~0, Code 7 ~0.03); no per-code kappa or positive agreement is reported. Because the largest model effects are on subjective codes (9 and 10), and the paper itself notes GPT-4 produced more 'AI-only' flags that the human did not find, the central claim rests on the least secure part of the measurement chain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a preregistered 2×3 factorial experiment (OSF, Nov 2023) in which 262 survey questions from Gallup Q12, WVS, and LGPI were each prompted through six ChatGPT treatments: GPT-3.5 vs GPT-4.0 crossed with no persona, survey-design-expert persona, or linguist persona. The zero-shot prompts asked for up to five features that could cause respondents to interpret the question differently. Each output statement was qualitatively coded with an 11-category scheme adapted from QAS-99 and Rothgeb et al., plus NOTA and an emergent SysVar subcode. Multilevel linear probability and logistic models with question random intercepts, with Holm-corrected p-values, estimate the effects of model and persona on the total number of codes and on each code's presence. The main results are that GPT-4 produces 0.55 more codes on average, is more likely to flag Code 3 (syntax), Code 5 (double-barrelled), Code 8 (complex estimation), Code 9 (sensitivity), and Code 10 (leading), and is 17 pp less likely to produce NOTA statements; personas shift output in smaller but plausible ways. The authors also present exploratory analyses of question source, statement ordering, and agreement between AI codes and a human coder, concluding that generative AI is a valuable 'safety net' for survey question refinement.","tokens_in":24679,"tokens_out":5674,"duration_ms":56457,"significance":"If the measurement chain is accepted, the paper provides one of the first systematic, pre-registered estimates of how model choice and persona change ChatGPT's survey-feedback content, with effect sizes at a meaningful magnitude (e.g., 0.14–0.15 for Codes 9 and 10) and robustness across linear and logistic specifications. The design has real strengths: preregistration, factorial manipulation, random-intercepts modeling, Holm correction, publicly posted data/code, and coding categories anchored (initially) in established instruments. The paper also honestly reports limitations such as the unusable Code 1 and the risk of spurious AI-only flags. However, the headline claim that the feedback 'maps onto established survey problem categories' depends critically on the validity of the codebook and the independence of the human benchmark, both of which are currently problematic.","major_comments":[{"comment":"The 'human expert' benchmark is not independent: the comparison in Figure 3 and the surrounding text is based on first author Metheney's own coding of all 262 questions using the same codebook that was refined while coding AI output (Appendix B). Agreement with this benchmark cannot establish that AI feedback corresponds to real survey problems rather than to the coding team's expectations. The Discussion claim that AI 'provides feedback similar to that of a human expert' is therefore unsupported. Please obtain independent expert coding (at least on a subset) or explicitly reframe §3.4.3 as descriptive concordance between AI and the authors' coding, not as validity evidence.","section":"§3.4.3, footnote 6"},{"comment":"The codebook was revised while coding AI output: categories were merged, terms were changed, and emergent codes (Answer set, SysVar) were added during the coding process. Reliability is reported only as 'above 80% agreement for all codes and 70% full agreement' (Appendix B), with no per-code kappa or positive agreement for most codes. Given the low prevalence of some codes (e.g., Code 8, Code 7 at ~0.03), overall agreement may be dominated by 'Neither' codings. This weakens the interpretability of every code-level effect in Tables 5–10. Add per-code reliability statistics (e.g., Gwet's AC1, positive agreement) and, if reliability is low, restrict the main analysis to codes with acceptable reliability.","section":"Appendix B, Table 2"},{"comment":"Code 1 (vague term) was coded in 98% of question-treatment sets, and the authors state that 'the results cannot be meaningfully investigated' for it. Yet H3 was preregistered as testing whether the linguist persona increases Code 1 and Code 3. The paper reports 'partial support for all three preregistered hypotheses,' but for H3 only Code 3 can actually be evaluated. Please explicitly exclude Code 1 from hypothesis testing and state that H3 is supported only for Code 3.","section":"§3.1 and H3"}],"minor_comments":[{"comment":"Typographical issues: 'wholistically' (Introduction), 'estimatation' (§3.1), 'Holmes' instead of 'Holm' in Table titles (Tables 5–7 and 8–10), 'M1 = GPT-3.4' in Figure 2 legend (should be GPT-3.5), and 'Air on the side' in Appendix B (should be 'Err on the side').","section":"Throughout"},{"comment":"The citation 'Olivos and Liu 0' in the Introduction has an incomplete publication year; please update with the full reference.","section":"References"},{"comment":"The statement-order analysis reports average placements but no measures of dispersion; adding standard deviations or boxplots would aid interpretation.","section":"§3.4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely topic and has a sound experimental skeleton. The main revisions are feasible: strengthen or reinterpret the human-comparison analysis, supply per-code reliability metrics, and handle Code 1 explicitly in hypothesis testing. If independent expert coding is not available, removing the 'similar to human expert' claim and limiting conclusions to 'agreement with the authors' coding' would make the paper's claims match the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth a serious look. It runs a clean 2x3 factorial experiment — two GPT versions, three personas — across 262 real survey questions, codes the output into literature-grounded categories, and finds real differences: GPT-4 produces 0.55 more codes per question, is 17 points less likely to produce off-task NOTA output, and shifts the mix toward sensitivity and leading-question codes. The survey-expert persona adds focused codes like double-barreled and answer-set problems; the linguist persona mostly adds syntax. That is genuinely new. Prior work was example-based or single-case. The design is careful: preregistered, multilevel models with question random intercepts, Holm correction, both linear and logistic specifications, and the authors honestly report the ugly parts — Code 1 was unusable at 98% prevalence, and the AI flagged many things the human did not.\n\nThe weak spot is the measurement chain, exactly where the reader and stress-test put it. The 'human expert' benchmark is the first author applying a codebook that was refined while coding AI output, and the agreement statistics are overall percent agreement, which is going to be dominated by 'Neither' for rare codes like Code 8. No per-code kappa or positive agreement is reported. So the ~80% agreement cannot carry the weight of 'AI feedback maps onto real survey problems.' The internal treatment comparisons — GPT-4 vs GPT-3.5, persona vs no persona — are statements about ChatGPT output, and those are trustworthy. The external interpretation, that the flagged codes are genuine problems, needs an independent coding pass with the final codebook.\n\nThe source-based validity checks are a nice counterweight: the AI flags more issues in the less-refined WVS and LGPI questions than in Gallup, which is the kind of pattern that would be hard to fake. But the largest model effects land on the most subjective codes (sensitivity, leading), so the measurement concern is not minor. Also, the abstract's 'valuable tool today, even for an average AI user' oversells it. The tested conditions are 2023 models, a web interface, short structured prompts, and three English-language surveys. That is a proof of concept, not an endorsement for the average user.\n\nFor survey methodologists, this is a useful, honest baseline. It deserves peer review, with referees who will push for independent coding, per-code reliability, and a more calibrated abstract. My verdict is skeptical only about the external claims, not about the core experimental result.","headline":"A solid, preregistered experiment showing GPT version and persona steer survey feedback, but the paper's 'valuable for the average user' claim outruns its measurement chain.","tokens_in":25239,"tokens_out":2450,"would_cite":true,"duration_ms":29962,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An ordinary ChatGPT user, writing short prompts with no training examples, gets feedback that aligns with known survey-question problems—and both the version of the model and the assigned persona change what gets flagged.","keywords":["survey question design","generative AI","ChatGPT","prompt engineering","questionnaire pretesting","qualitative coding","total survey error","model comparison"],"falsifier":"Have independent survey-methodologist coders, blind to treatment, re-code the full set of ChatGPT outputs; then re-estimate the model and persona effects. If inter-coder agreement with the original codes is low, or if the GPT-4.0 advantages in sensitivity and leading flags shrink or reverse under the independent codes, the paper's central claim would fail.","tokens_in":24276,"feed_emoji":"🤖","tokens_out":7186,"duration_ms":68454,"temperature":0.7,"pith_summary":"This paper aims to establish that generative AI, used in the simplest way an ordinary user would use it, is already a workable safety net for survey question refinement—catching problems that match known survey methodology categories before a questionnaire is fielded. The authors ran a preregistered 2-by-3 experiment in which two ChatGPT versions (3.5 and 4.0) and three persona conditions (none, survey design expert, linguist) were applied to 262 questions from three established surveys, and all output was coded with an 11-category scheme of survey-question problems. They report that GPT-4.0 flags more problems than GPT-3.5, especially respondent-centered issues such as sensitivity and leading/biasing questions, and produces less irrelevant output; assigning a persona shifts the kinds of problems flagged. If correct, this means AI feedback can supplement expensive pretesting methods when time, money, or expertise is limited, and that users should expect model and prompt choices to matter.","feed_headline":"GPT-4 finds more survey flaws than GPT-3.5 from plain prompts","feed_subtitle":"Model version and persona steer ChatGPT feedback toward known survey-problem categories.","key_machinery":"The load-bearing machinery is a combination of a zero-shot prompt experiment and a qualitative coding scheme. Each of 262 survey questions was asked six times—once per combination of model version (GPT-3.5 or GPT-4.0) and persona (none, survey design expert, linguist)—with prompts that asked for up to five features causing respondents to interpret the question differently. Every statement was then coded into an 11-category scheme of survey-question problems (vague terms, specialized knowledge, syntax, unfair presumption, double-barreled, reference-period issues, recall difficulty, complex estimation, sensitivity, leading/biasing, answer set), plus a 'none of the above' category. The scheme,","core_discovery":"The central claim is that an average ChatGPT user—someone who writes a short, single-request prompt with no model training—can expect feedback aligned with the survey-methodology literature on question problems. Concretely, GPT-4.0 produced on average 0.55 more codes per question-treatment pair than GPT-3.5 (p < 0.001), was 17 percentage points less likely to produce 'none of the above' statements, and was more likely to produce codes for syntax problems, double-barreled questions, complex estimation, sensitivity, and leading/biasing questions, with the latter two showing the largest effects. Persona also mattered: a survey-design-expert persona increased the number of codes and shifted outp","pith_inferences":["A natural next experiment would test whether asking for fewer than five features reduces the late-arriving off-task statements; the order analysis suggests the five-item cap creates filler.","The systematic-variation output could be repurposed as a feature: asking the model to suggest control variables or subgroups that answer differently would turn off-task comments into analysis-planning input.","Because the human benchmark was a single coder, a multi-coder panel re-coding the same output could quantify false-positive rates per code and reveal which AI-only flags are novel vs spurious."],"forward_implications":["A researcher without special AI expertise can obtain useful, codeable feedback on draft questions using short prompts, so the quality-check layer is accessible beyond pretesting specialists.","Model version is not neutral: upgrading from GPT-3.5 to GPT-4.0 changes both the quantity and the focus of feedback, so survey teams should re-test whatever AI procedure they use as models are updated.","Persona choice can be used deliberately: asking for a survey-design expert increases flags on double-barreled questions and answer-set issues, while asking for a linguist increases syntax flags.","Irrelevant or off-task output (NOTA) is common—about 17% of question-treatment pairs—but some of it, labeled systematic variation, is useful for analysis planning rather than question wording.","The human comparison suggests AI output overlaps with expert judgment most of the time, but also that AI flags things a human did not, so it should be treated as a complement to expert review rather than a replacement."],"supporting_citations":[{"why":"Supplies the QAS-99 question appraisal checklist that the paper truncates into its coding scheme.","marker":"Willis and Lessler 1999"},{"why":"Provides the four-type problem scheme (content, structure, retrieval, judgment) that the codebook adapts and extends with emergent codes.","marker":"Rothgeb, Willis, and Forsyth 2007"},{"why":"Supplies the total survey error framework that frames AI as a quality safety net under resource constraints.","marker":"Biemer 2010"},{"why":"Catalogues prompt patterns and motivates the expectation that persona specification changes output.","marker":"White et al. 2023"},{"why":"Provides the simulation-based power analysis used to size the multilevel models.","marker":"Green and MacLeod 2016"},{"why":"Establishes that different pretesting methods surface different problem types, the premise behind treating AI as one more appraisal method.","marker":"Presser and Blair 1994"}],"fun_headline_variants":["GPT-4 spots more survey flaws than GPT-3.5","Simple ChatGPT prompts sharpen survey question design","Persona and model version steer ChatGPT's survey feedback","ChatGPT's survey critique hinges on prompt persona"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The results stand on the assumption that the 11-code scheme and the single expert's coding of all 262 questions are a valid reference for what counts as a survey-question problem; if that coding is biased or unreliable, the measured effects tell us about the codebook as much as about the AI.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 spots more survey flaws than GPT-3.5","Simple ChatGPT prompts sharpen survey question design","Persona and model version steer ChatGPT's survey feedback","ChatGPT's survey critique hinges on prompt persona"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1483,"prompt_tokens":769,"completion_tokens":714,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":652}},"tokens_in":513,"tokens_out":714,"duration_ms":8692,"temperature":1.0,"reasoning_tokens":652,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T20:17:28.127788+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent survey-methodologist coders, blind to treatment, re-code the full set of ChatGPT outputs; then re-estimate the model and persona effects. If inter-coder agreement with the original codes is low, or if the GPT-4.0 advantages in sensitivity and leading flags shrink or reverse under the independent codes, the paper's central claim would fail.","supporting_citations":[],"review_version":1}