Pith. sign in

REVIEW 3 major objections 3 minor 16 references

An ordinary ChatGPT user, writing short prompts with no training examples, gets feedback that aligns with known survey-question problems—and both the version of the model and the assigned persona change what gets flagged.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A preregistered 2x3 factorial experiment on 262 real survey questions shows GPT-4.0 and a survey-expert persona make ChatGPT flag more, and different, survey question problems than GPT-3.5 or no persona.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection A solid, preregistered experiment showing GPT version and persona steer survey feedback, but the paper's 'valuable for the average user' claim outruns its measurement chain. the 3 major comments →

arxiv 2509.08702 v1 pith:QZANLPMT submitted 2025-09-10 stat.ME cs.CYstat.AP

Generative AI as a Safety Net for Survey Question Refinement

classification stat.ME cs.CYstat.AP
keywords survey question designgenerative AIChatGPTprompt engineeringquestionnaire pretestingqualitative codingtotal survey errormodel comparison
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to establish that generative AI, used in the simplest way an ordinary user would use it, is already a workable safety net for survey question refinement—catching problems that match known survey methodology categories before a questionnaire is fielded. The authors ran a preregistered 2-by-3 experiment in which two ChatGPT versions (3.5 and 4.0) and three persona conditions (none, survey design expert, linguist) were applied to 262 questions from three established surveys, and all output was coded with an 11-category scheme of survey-question problems. They report that GPT-4.0 flags more problems than GPT-3.5, especially respondent-centered issues such as sensitivity and leading/biasing questions, and produces less irrelevant output; assigning a persona shifts the kinds of problems flagged. If correct, this means AI feedback can supplement expensive pretesting methods when time, money, or expertise is limited, and that users should expect model and prompt choices to matter.

Core claim

The central claim is that an average ChatGPT user—someone who writes a short, single-request prompt with no model training—can expect feedback aligned with the survey-methodology literature on question problems. Concretely, GPT-4.0 produced on average 0.55 more codes per question-treatment pair than GPT-3.5 (p < 0.001), was 17 percentage points less likely to produce 'none of the above' statements, and was more likely to produce codes for syntax problems, double-barreled questions, complex estimation, sensitivity, and leading/biasing questions, with the latter two showing the largest effects. Persona also mattered: a survey-design-expert persona increased the number of codes and shifted outp

What carries the argument

The load-bearing machinery is a combination of a zero-shot prompt experiment and a qualitative coding scheme. Each of 262 survey questions was asked six times—once per combination of model version (GPT-3.5 or GPT-4.0) and persona (none, survey design expert, linguist)—with prompts that asked for up to five features causing respondents to interpret the question differently. Every statement was then coded into an 11-category scheme of survey-question problems (vague terms, specialized knowledge, syntax, unfair presumption, double-barreled, reference-period issues, recall difficulty, complex estimation, sensitivity, leading/biasing, answer set), plus a 'none of the above' category. The scheme,

Load-bearing premise

The results stand on the assumption that the 11-code scheme and the single expert's coding of all 262 questions are a valid reference for what counts as a survey-question problem; if that coding is biased or unreliable, the measured effects tell us about the codebook as much as about the AI.

What would settle it

Have independent survey-methodologist coders, blind to treatment, re-code the full set of ChatGPT outputs; then re-estimate the model and persona effects. If inter-coder agreement with the original codes is low, or if the GPT-4.0 advantages in sensitivity and leading flags shrink or reverse under the independent codes, the paper's central claim would fail.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A researcher without special AI expertise can obtain useful, codeable feedback on draft questions using short prompts, so the quality-check layer is accessible beyond pretesting specialists.
  • Model version is not neutral: upgrading from GPT-3.5 to GPT-4.0 changes both the quantity and the focus of feedback, so survey teams should re-test whatever AI procedure they use as models are updated.
  • Persona choice can be used deliberately: asking for a survey-design expert increases flags on double-barreled questions and answer-set issues, while asking for a linguist increases syntax flags.
  • Irrelevant or off-task output (NOTA) is common—about 17% of question-treatment pairs—but some of it, labeled systematic variation, is useful for analysis planning rather than question wording.
  • The human comparison suggests AI output overlaps with expert judgment most of the time, but also that AI flags things a human did not, so it should be treated as a complement to expert review rather than a replacement.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural next experiment would test whether asking for fewer than five features reduces the late-arriving off-task statements; the order analysis suggests the five-item cap creates filler.
  • The systematic-variation output could be repurposed as a feature: asking the model to suggest control variables or subgroups that answer differently would turn off-task comments into analysis-planning input.
  • Because the human benchmark was a single coder, a multi-coder panel re-coding the same output could quantify false-positive rates per code and reveal which AI-only flags are novel vs spurious.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. This paper reports a preregistered 2×3 factorial experiment (OSF, Nov 2023) in which 262 survey questions from Gallup Q12, WVS, and LGPI were each prompted through six ChatGPT treatments: GPT-3.5 vs GPT-4.0 crossed with no persona, survey-design-expert persona, or linguist persona. The zero-shot prompts asked for up to five features that could cause respondents to interpret the question differently. Each output statement was qualitatively coded with an 11-category scheme adapted from QAS-99 and Rothgeb et al., plus NOTA and an emergent SysVar subcode. Multilevel linear probability and logistic models with question random intercepts, with Holm-corrected p-values, estimate the effects of model and persona on the total number of codes and on each code's presence. The main results are that GPT-4 produces 0.55 more codes on average, is more likely to flag Code 3 (syntax), Code 5 (double-barrelled), Code 8 (complex estimation), Code 9 (sensitivity), and Code 10 (leading), and is 17 pp less likely to produce NOTA statements; personas shift output in smaller but plausible ways. The authors also present exploratory analyses of question source, statement ordering, and agreement between AI codes and a human coder, concluding that generative AI is a valuable 'safety net' for survey question refinement.

Significance. If the measurement chain is accepted, the paper provides one of the first systematic, pre-registered estimates of how model choice and persona change ChatGPT's survey-feedback content, with effect sizes at a meaningful magnitude (e.g., 0.14–0.15 for Codes 9 and 10) and robustness across linear and logistic specifications. The design has real strengths: preregistration, factorial manipulation, random-intercepts modeling, Holm correction, publicly posted data/code, and coding categories anchored (initially) in established instruments. The paper also honestly reports limitations such as the unusable Code 1 and the risk of spurious AI-only flags. However, the headline claim that the feedback 'maps onto established survey problem categories' depends critically on the validity of the codebook and the independence of the human benchmark, both of which are currently problematic.

major comments (3)
  1. [§3.4.3, footnote 6] The 'human expert' benchmark is not independent: the comparison in Figure 3 and the surrounding text is based on first author Metheney's own coding of all 262 questions using the same codebook that was refined while coding AI output (Appendix B). Agreement with this benchmark cannot establish that AI feedback corresponds to real survey problems rather than to the coding team's expectations. The Discussion claim that AI 'provides feedback similar to that of a human expert' is therefore unsupported. Please obtain independent expert coding (at least on a subset) or explicitly reframe §3.4.3 as descriptive concordance between AI and the authors' coding, not as validity evidence.
  2. [Appendix B, Table 2] The codebook was revised while coding AI output: categories were merged, terms were changed, and emergent codes (Answer set, SysVar) were added during the coding process. Reliability is reported only as 'above 80% agreement for all codes and 70% full agreement' (Appendix B), with no per-code kappa or positive agreement for most codes. Given the low prevalence of some codes (e.g., Code 8, Code 7 at ~0.03), overall agreement may be dominated by 'Neither' codings. This weakens the interpretability of every code-level effect in Tables 5–10. Add per-code reliability statistics (e.g., Gwet's AC1, positive agreement) and, if reliability is low, restrict the main analysis to codes with acceptable reliability.
  3. [§3.1 and H3] Code 1 (vague term) was coded in 98% of question-treatment sets, and the authors state that 'the results cannot be meaningfully investigated' for it. Yet H3 was preregistered as testing whether the linguist persona increases Code 1 and Code 3. The paper reports 'partial support for all three preregistered hypotheses,' but for H3 only Code 3 can actually be evaluated. Please explicitly exclude Code 1 from hypothesis testing and state that H3 is supported only for Code 3.
minor comments (3)
  1. [Throughout] Typographical issues: 'wholistically' (Introduction), 'estimatation' (§3.1), 'Holmes' instead of 'Holm' in Table titles (Tables 5–7 and 8–10), 'M1 = GPT-3.4' in Figure 2 legend (should be GPT-3.5), and 'Air on the side' in Appendix B (should be 'Err on the side').
  2. [References] The citation 'Olivos and Liu 0' in the Introduction has an incomplete publication year; please update with the full reference.
  3. [§3.4.2] The statement-order analysis reports average placements but no measures of dispersion; adding standard deviations or boxplots would aid interpretation.

Circularity Check

0 steps flagged

No significant circularity: the GPT-4 effects are direct empirical comparisons of coded AI output, not predictions derived from the codebook or from self-citations.

full rationale

The paper's central estimates (Section 3.1: GPT-4 produces 0.55 more codes; Code 9 +0.14, Code 10 +0.15) are empirical comparisons of qualitatively coded ChatGPT output under randomized treatments. The coding scheme is explicitly adapted from external sources (QAS-99; Rothgeb et al. 2007) and the emergent codes were added during abductive refinement, but this does not make the model-effect regressions circular: the outcome variables are human codings of model output, and no parameter is fitted to the quantity being 'predicted.' The human-comparison validity check (Section 3.4.3, footnote 6) uses the first author's coding as a benchmark and the paper discloses low inter-coder reliability and double-checking; these are validity/reliability limitations, not instances where a result reduces by construction to its inputs. The only self-citation is the replication-data deposit (Metheney and Yehle 2025), which is not load-bearing. Concerns about codebook validity, base-rate-driven agreement, and the single-author benchmark should be assessed as measurement risk, not circularity.

Axiom & Free-Parameter Ledger

1 free parameters · 5 axioms · 0 invented entities

The measurement pipeline is the main source of unverified input: a hand-built 11-code scheme refined while coding AI output, a 5-statement cap chosen arbitrarily, and a single-author human benchmark. The statistical models are standard; the treatments are unpinned black-box LLM outputs. These items, rather than fitted constants, are what the reader pays for upstream.

free parameters (1)
  • requested number of feedback features (output cap) = 5
    The prompt asks for up to 5 features; the authors state this was chosen arbitrarily (Section 2.1). It caps the number of statements per question-treatment and thus the observed code count, and the statement-order analysis (Section 3.4.2) suggests the AI reaches to fill the 5 slots, generating filler content.
axioms (5)
  • domain assumption The 11-code scheme adapted from QAS-99 and Rothgeb et al. (2007) captures the survey question problems that matter for data quality.
    Underlies the entire outcome measurement (Section 2.5, Appendix B); the scheme was redefined and extended while coding AI output.
  • domain assumption The ChatGPT web interface with default parameters, accessed September to December 2023, provides a stable and comparable treatment for the GPT-3.5 vs GPT-4.0 contrast.
    Section 2.4 and Algorithm 1: model snapshots and sampling settings are unpinned, so version comparisons include uncontrolled drift and stochasticity.
  • domain assumption The prompt template is representative of what an average AI user would write, supporting the abstract's average-user generalization.
    Section 2.1: prompts are structured and domain-specific; no evidence shows typical users write such prompts.
  • domain assumption First author Metheney's coding of the 262 questions is an adequate human-expert reference for the validity comparisons.
    Section 3.4.3 and footnote 6: a single author coded with the same codebook applied to AI output, and that codebook was refined in part on AI output (Appendix B).
  • standard math The random-intercepts multilevel logistic and linear probability models are correctly specified for the factorial within-question design.
    Section 2.6, Equations (1) and (2).

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Generative AI as a Safety Net for Survey Question Refinement." pith.science (2026). https://pith.science/paper/QZANLPMT

@misc{pith2026250908702,
  author       = {Pith},
  title        = {Pith review of: Generative AI as a Safety Net for Survey Question Refinement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QZANLPMT}},
  note         = {Machine review of arXiv:2509.08702}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Writing survey questions that easily and accurately convey their intent to a variety of respondents is a demanding and high-stakes task. Despite the extensive literature on best practices, the number of considerations to keep in mind is vast and even small errors can render collected data unusable for its intended purpose. The process of drafting initial questions, checking for known sources of error, and developing solutions to those problems requires considerable time, expertise, and financial resources. Given the rising costs of survey implementation and the critical role that polls play in media, policymaking, and research, it is vital that we utilize all available tools to protect the integrity of survey data and the financial investments made to obtain it. Since its launch in 2022, ChatGPT and other generative AI model platforms have been integrated into everyday life processes and workflows, particularly pertaining to text revision. While many researchers have begun exploring how generative AI may assist with questionnaire design, we have implemented a prompt experiment to systematically test what kind of feedback on survey questions an average ChatGPT user can expect. Results from our zero--shot prompt experiment, which randomized the version of ChatGPT and the persona given to the model, shows that generative AI is a valuable tool today, even for an average AI user, and suggests that AI will play an increasingly prominent role in the evolution of survey development best practices as precise tools are developed.

Figures

Figures reproduced from arXiv: 2509.08702 by Erica Ann Metheney, Lauren Yehle.

Figure 1
Figure 1. Figure 1: Illustration depicting how we can conceptualize AI as a safety net for survey [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Marginal Effects Plot Showing the Impact of Each Experimental Attribute on [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Comparison of Codes between Human Assessment and Experimental Output. [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Power analysis results for detecting a Model effect of sizes 0.1, 0.41, and 0.7 [PITH_FULL_IMAGE:figures/full_fig_p035_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Power analysis results for detecting a Model effect of sizes 0.1, 0.41, and 0.7 [PITH_FULL_IMAGE:figures/full_fig_p036_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Power analysis results for detecting a Persona effect of sizes 0.1, 0.41, and 0.7 [PITH_FULL_IMAGE:figures/full_fig_p037_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Power analysis results for detecting a Persona effect of sizes 0.1, 0.41, and 0.7 [PITH_FULL_IMAGE:figures/full_fig_p038_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Power analysis results for detecting an interaction effect between Model and [PITH_FULL_IMAGE:figures/full_fig_p039_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Power analysis results for detecting an interaction effect between Model and [PITH_FULL_IMAGE:figures/full_fig_p040_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Comparison of Human and AI Use of Each Code [PITH_FULL_IMAGE:figures/full_fig_p057_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

16 extracted references · 4 canonical work pages

  1. [3]

    https://www.pnas.org/doi/10.1073/pnas

    https : //doi.org/10.1073/pnas.2314021121. https://www.pnas.org/doi/10.1073/pnas. 2314021121. Bais, Frank, Barry Schouten, Peter Lugtig, Vera Toepoel, Judit Arends-T` oth, Salima Douhou, Natalia Kieruj, Mattijn Morren, and Corrie Vis

  2. [5]

    https: //doi.org/10.1007/978-94-007-0753-5

  3. [6]

    Converse, Jean M., and Stanley Presser

    2025.ESRA 2025 Program.Online: https://www.europeansurveyresearch.org/conf2025/prog.php. Converse, Jean M., and Stanley Presser. 1986.Survey Questions: Handcrafting the Stan- dardized Questionnaire[in en]. Google-Books-ID: AXRZbfHM 94C. SAGE, Septem- ber.isbn: 978-0-8039-2743-8. 58 Fichtner, Urs Alexander, Jochen Knaus, Erika Graf, Georg Koch, J¨ org Sahl...

  4. [9]

    Summary of ChatGPT- Related Research and Perspective Towards the Future of Large Language Models

    “Summary of ChatGPT- Related Research and Perspective Towards the Future of Large Language Models.” Meta-Radiology1:100017. Metheney, Erica, and Lauren Yehle. 2025.Replication Data for: AI Shows Immense Promise to Revolutionize How We Design Surveys.V. V1. https : / / doi . org / 10 . 7910/DVN/AN3U9F. https://doi.org/10.7910/DVN/AN3U9F. 59 Moineddin, Rahi...

  5. [10]

    ChatGPTest: Opportunities and Cautionary Tales of Utilizing AI for Questionnaire Pretesting

    “ChatGPTest: Opportunities and Cautionary Tales of Utilizing AI for Questionnaire Pretesting.”Field Methods0 (0): 1525822X241280574. https://doi.org/10.1177/1525822X241280574. eprint: https://doi.org/10.1177/ 1525822X241280574. https://doi.org/10.1177/1525822X241280574. Ongena, Yfke P, and Wil Dijkstra

  6. [15]

    https://doi.org/10.48550/arXiv.2302. 11382. http://arxiv.org/abs/2302.11382. Willis, Gordon B, and Judith T Lessler

  7. [1994]

    Survey Pretesting: Do Different Methods Produce Different Results?

    “Survey Pretesting: Do Different Methods Produce Different Results?”Sociological Methodology24:73–104. R Core Team. 2023.R: A Language and Environment for Statistical Computing.Vienna, Austria: R Foundation for Statistical Computing. https://www.R-project.org/. Rothgeb, Jennifer, Gordon Willis, and Barbara Forsyth

  8. [1999]

    Question appraisal system QAS-99

    “Question appraisal system QAS-99.” National Cancer Institute-. World Values Survey Association. 2020.Home>What we do>Questionnaire Develop- ment.Available at https : / / www . worldvaluessurvey . org / whoweareContents . jsp Accessed: 2023/08/11. . 2023.Home - Who we are.Available at https://www.worldvaluessurvey.org/who we areContents.jsp Accessed: 2023...

  9. [2007]

    Questionnaire Pretesting Methods: Do Different Techniques and Different Organizations Produce Similar Results?

    “Questionnaire Pretesting Methods: Do Different Techniques and Different Organizations Produce Similar Results?”BMS: Bulletin of Sociological Methodology / Bulletin de Methodologie Sociologique96 (96): 5–31. Sturgis, Patrick, Thomas S Robinson, Laura Fung, and Caroline Roberts. 2025.SOCbot: Using Large Language Models to measure and classify occupations i...

  10. [2011]

    Coding the Behavior of Interviewers and Respondents to Evaluate Sur- vey Questions

    “Coding the Behavior of Interviewers and Respondents to Evaluate Sur- vey Questions.” InQuestion evaluation methods: contributing to the science of data quality,edited by Jennifer Madans, Kristen Miller, Aaron Maitland, and Gordon B Willis, 7–21. John Wiley & Sons. Fowler, Floyd J. 1995.Improving Survey Questions: Design and Evaluation[in en]. Google-Book...

  11. [2016]

    simr: an R package for power analysis of generalised linear mixed models by simulation

    “simr: an R package for power analysis of generalised linear mixed models by simulation.”Methods in Ecology and Evolution 7 (4): 493–498. https://doi.org /10.1111/2041- 210X.12504. https://CRAN.R- project.org/package=simr. Holbrook, Allyson, Young Ik Cho, and Timothy Johnson

  12. [2019]

    Sample size issues in multilevel logistic regression models

    “Sample size issues in multilevel logistic regression models” [in en]. Publisher: Public Library of Science, PLOS ONE14, no. 11 (November): e0225427.issn: 1932-6203, accessed April 15,

  13. [2020]

    Using Time Series Mod- els to Understand Survey Costs

    “Using Time Series Mod- els to Understand Survey Costs.” Eprint: https://academic.oup.com/jssam/article- pdf/9/5/943/41727211/smaa024.pdf,Journal of Survey Statistics and Methodology 9, no. 5 (November): 943–960.issn: 2325-0984. https://doi.org/10.1093/jssam/ smaa024. https://doi.org/10.1093/jssam/smaa024. 60 White, Jules, Quchen Fu, Sam Hays, Michael San...

  14. [2023]

    https: //doi.org/10.1177/0049124117729692

    https://doi.org/10.1177/0049124117729692. https: //doi.org/10.1177/0049124117729692. Biemer, Paul P

  15. [2024]

    Adaptive Questionnaire Design Using AI Agents for People Profiling

    “Adaptive Questionnaire Design Using AI Agents for People Profiling.” InICAART (3),633–640. Payne, Stanley Le Baron. 1951.The art of asking questions: Studies in public opinion, 3.Princeton University Press. Presser, Stanley, and Johnny Blair

  16. [2025]

    https://journals.plos.org/ plosone/article?id=10.1371/journal.pone.0225427

    https://doi.org/10.1371/journal.pone.0225427. https://journals.plos.org/ plosone/article?id=10.1371/journal.pone.0225427. Alvesson, Mats, J¨ orgen Sandberg, and Katja Einola

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.