REVIEW 2 major objections 2 minor 78 references
Standard hypothesis tests are invalid under models of generative surveying that include prompt perturbations.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Standard hypothesis tests fail for generative surveys under realistic prompt perturbations; a permutation test is valid and practical guidance on budget allocation is given.
T0 review reviewed 2026-06-29 challenge →
load-bearing objection The paper shows standard tests fail under a generative survey model with prompt perturbations and offers a permutation test plus budget advice, but the model's match to real LLMs is untested. the 2 major comments →
When prompt perturbations break your A/B test: A valid statistical test for generative surveying
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Under a statistical model for generative surveying that includes realistic perturbation structure, standard hypothesis tests including the sign test and Wilcoxon signed-rank test are invalid, whereas a permutation test is valid. The paper formally characterizes the conditions under which the standard tests fail and demonstrates that both the magnitude and direction of estimated effects are sensitive to the choice of model.
What carries the argument
A permutation test that respects the perturbation structure in the generative surveying model.
Load-bearing premise
The statistical model for generative surveying with realistic perturbation structure accurately describes how LLMs respond to semantically equivalent prompt variations.
What would settle it
Collect data from an LLM survey with multiple perturbations per persona and compare the p-values from the standard sign test and the proposed permutation test; if they frequently disagree on significance, the invalidity claim is supported.
If this is right
- The sign test and Wilcoxon signed-rank test fail to maintain correct type I error rates when perturbations are present.
- Effect estimates in generative surveys depend on the specific LLM used, even within the same family.
- Budget should be allocated across personas, perturbations, and replicates to achieve desired power.
- Conditions for failure of standard tests can be characterized in terms of the perturbation model parameters.
Where Pith is reading between the lines
- Existing generative survey results that used standard tests may need to be reanalyzed with the permutation test.
- The approach could be extended to other settings where input variations affect statistical conclusions, such as in robustness testing for machine learning models.
- Power calculations under this model can inform minimal sample sizes for future generative surveys.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that standard hypothesis tests (sign test, Wilcoxon signed-rank) are invalid for A/B testing in generative surveying because LLM responses exhibit dependence induced by prompt perturbations. It derives a permutation test that is valid under the proposed generative model, formally characterizes the conditions under which standard tests fail, applies the framework to estimate model parameters in a simple surveying example, analyzes the power of the permutation test, and supplies guidance on allocating budget across personas, perturbations, and replicates. It also shows that effect magnitude and direction can be sensitive to model choice even within the same LLM family.
Significance. If the perturbation model accurately captures the dependence structure of real LLM responses to semantically equivalent prompts, the work supplies a practically important correction to inference in generative surveying. The valid permutation test and budget-allocation results would directly improve the reliability of LLM-based market research. The formal characterization of test failure conditions is a methodological contribution; the parameter estimation and power analysis add concrete guidance. These strengths are tempered by the absence of external validation that the assumed perturbation structure matches observed LLM behavior.
major comments (2)
- [§2] §2 (generative model): the invalidity result for the sign and Wilcoxon tests is derived under a specific perturbation dependence structure; without an explicit equation or proof sketch showing how the covariance between perturbed responses violates the exchangeability or independence assumptions of those tests, the central claim cannot be verified.
- [§4–5] §4–5 (application and power): parameters are estimated and power is characterized under the model, yet no comparison of the fitted perturbation distribution to empirical LLM output distributions (e.g., response similarity across prompt variants) is reported; this leaves the practical guidance conditional on an untested modeling assumption.
minor comments (2)
- [Abstract] Abstract: the phrase 'realistic perturbation structure' is used without a one-sentence gloss of the key dependence assumption.
- [Model section] Notation: the distinction between persona-level, perturbation-level, and replicate-level random effects is not always visually clear in the displayed equations.
Simulated Author's Rebuttal
We thank the referee for the careful reading and for identifying points that will improve the clarity and transparency of the manuscript. We address each major comment below.
read point-by-point responses
-
Referee: [§2] §2 (generative model): the invalidity result for the sign and Wilcoxon tests is derived under a specific perturbation dependence structure; without an explicit equation or proof sketch showing how the covariance between perturbed responses violates the exchangeability or independence assumptions of those tests, the central claim cannot be verified.
Authors: Section 2 introduces the generative model in which responses to semantically equivalent prompts share a common perturbation factor, inducing dependence. The text states that this structure violates the assumptions of the sign and Wilcoxon tests, but we agree that an explicit covariance equation and short proof sketch would make the violation transparent. In the revised manuscript we will insert a short derivation showing that Cov(Y_{ij}, Y_{ik}) > 0 for j ≠ k under shared perturbation, which directly breaks the exchangeability required by those tests. revision: yes
-
Referee: [§4–5] §4–5 (application and power): parameters are estimated and power is characterized under the model, yet no comparison of the fitted perturbation distribution to empirical LLM output distributions (e.g., response similarity across prompt variants) is reported; this leaves the practical guidance conditional on an untested modeling assumption.
Authors: We acknowledge that the paper contains no direct empirical comparison of the fitted perturbation distribution to observed LLM response similarities. The power and budget-allocation results in §§4–5 are therefore conditional on the modeling assumption. In revision we will add an explicit limitations subsection that states this assumption, motivates it from existing literature on prompt sensitivity, and outlines the empirical checks that would be needed to validate it. We cannot supply the missing comparison without new data collection outside the scope of the present work. revision: partial
- Empirical validation of the assumed perturbation dependence structure against observed LLM output distributions.
Circularity Check
No circularity detected; derivation is model-internal but self-contained
full rationale
The paper defines a generative surveying model that incorporates prompt perturbations and derives the invalidity of sign/Wilcoxon tests (plus validity of a permutation test) strictly under that model. No equations, fitted parameters, or self-citations appear in the abstract or provided text that would reduce any claim to its inputs by construction. The load-bearing assumption that the model captures 'realistic' structure is an external modeling choice, not a definitional loop or renamed fit. This is the normal case of a paper working inside its stated assumptions without circular reduction.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of When prompt perturbations break your A/B test: A valid statistical test for generative surveying." pith.science (2026). https://pith.science/paper/AZZ4Q7OZ
@misc{pith2026260527463,
author = {Pith},
title = {Pith review of: When prompt perturbations break your A/B test: A valid statistical test for generative surveying},
year = {2026},
howpublished = {\url{https://pith.science/paper/AZZ4Q7OZ}},
note = {Machine review of arXiv:2605.27463}
}
read the original abstract
Generative surveying -- where collections of LLM-based personas provide feedback on messages -- has emerged as a cheap and scalable alternative to traditional market research. However, LLMs are sensitive to small variations in prompt design and conclusions drawn from generative surveys may depend on arbitrary phrasing choices. Controlling for this sensitivity requires including semantically equivalent perturbations in the analysis. In this paper, we show that standard hypothesis tests, including the sign test and Wilcoxon signed-rank test, are invalid under a statistical model for generative surveying that includes realistic perturbation structure. We propose a permutation test that is valid under this model and formally characterize the conditions under which standard tests fail. Applying our framework to a simple generative surveying problem, we estimate relevant parameters, characterize the power of the permutation test under realistic conditions, and provide practical guidance on budget allocation across personas, perturbations, and replicates. Finally, we show that both the magnitude and direction of the estimated effect are sensitive to the choice of model, even within the same model family.
Figures
Reference graph
Works this paper leans on
-
[1]
In Proceedings of the 40th International Conference on Machine Learning, pages 337–371
Using large language models to simulate mul- tiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning, pages 337–371. John Arbuthnott. 1710. An argument for divine prov- idence, taken from the constant regularity observ’d in the births of both sexes. by dr. john arbuthnott, physitian in ordi...
-
[2]
Out of one, many: Using language mod- els to simulate human samples.Political Analysis, 31(3):337–351. Peter J. Bickel and Kjell A. Doksum. 1977.Mathe- matical Statistics: Basic Ideas and Selected Topics. Holden-Day, San Francisco. Paul P. Biemer and Lars E. Lyberg. 2003.Introduction to Survey Quality. John Wiley & Sons, Hoboken, NJ. James Bisbee, Joshua ...
work page internal anchor Pith review Pith/arXiv arXiv 1977
-
[3]
In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–19
Evaluating large language models in generat- ing synthetic HCI research data: A case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–19. Hayden Helm, Tianyi Chen, Harvey McGuinness, Paige Lee, Brandon Duderstadt, and Carey E. Priebe. 2025. Toward a digital twin of U.S. congress.arXiv preprint arXiv:2505.0000...
-
[4]
A statistical turing test for generative mod- els.Preprint, arXiv:2309.08913. John J. Horton, Apostolos Filippas, and Benjamin S. Manning. 2023. Large language models as simulated economic agents: What can we learn from Homo Silicus? Technical Report 31122, National Bureau of Economic Research. Jon A. Krosnick and Stanley Presser. 2010. Question and quest...
-
[5]
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109. Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, pages 29971–30004. Howard Schuman and Stanley Presser. 1981....
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[6]
I’d like to buy a pair of sneakers
-
[7]
I’m looking to purchase some new trainers
-
[8]
I want to shop for a pair of athletic casual shoes
-
[9]
I’m considering buying sneakers
-
[10]
I’m in the market for a new pair of sneakers
-
[11]
I’m thinking about purchasing some trainers
-
[12]
I’d like to get myself a pair of sneakers
-
[13]
I’m interested in buying athletic casual shoes. 11
-
[14]
I want to buy some new sneakers
-
[15]
I’m planning to purchase a pair of trainers
-
[16]
I’m looking to get a pair of sneakers
-
[17]
I’d like to invest in some new athletic casual shoes
-
[18]
I’m considering buying a pair of sneakers
-
[19]
I want to shop for trainers
-
[20]
I’m interested in purchasing sneakers
-
[21]
I’d like to grab a pair of athletic casual shoes
-
[22]
I’m thinking about getting some new sneakers
-
[23]
I’m in the mood to buy trainers
-
[24]
I’m looking to acquire a pair of sneakers
-
[25]
I want to pick up some new athletic casual shoes
-
[26]
I’m considering shopping for sneakers
-
[27]
I’d like to obtain a pair of trainers
-
[28]
I’m planning to get myself some new sneakers
-
[29]
I want to purchase a pair of trainers
-
[30]
I’m looking to purchase some athletic shoes
-
[31]
Can you show me some sneakers to buy?
-
[32]
I want to shop for a new pair of trainers
-
[33]
I’m interested in buying some casual athletic shoes
-
[34]
Do you have sneakers available for purchase?
-
[35]
I’m thinking about getting a pair of sneakers
-
[36]
I’d like to explore options for buying sneakers
-
[37]
I want to find a pair of sneakers to buy
-
[38]
I’m in the market for some new trainers
-
[39]
I’m considering purchasing a pair of sneakers
-
[40]
I’d like to buy some comfortable athletic shoes
-
[41]
I’m looking to add sneakers to my collection
-
[42]
I want to check out sneakers to purchase
-
[43]
I’m interested in buying a pair of running shoes
-
[44]
I’d like to shop for some stylish sneakers
-
[45]
I’m hoping to find sneakers to buy soon
-
[46]
I want to purchase a pair of casual sneakers
-
[47]
I’m considering buying some athletic footwear
-
[48]
I’d like to explore sneaker options for purchase
-
[49]
I want to buy some new trainers
-
[50]
I’m interested in purchasing some sneakers
-
[51]
I’d like to add a pair of sneakers to my wardrobe
-
[52]
I’m considering buying some casual athletic shoes
-
[53]
I’d like to buy a new pair of sneakers
-
[54]
Could you show me some trainers that are available for purchase?
-
[55]
A.4 Boot Perturbations (M= 25) All 25 semantically equivalent paraphrases of the boot purchase-intent query, generated and validated via mistral-small-latest:
I’m looking to shop for a pair of athletic casual shoes. A.4 Boot Perturbations (M= 25) All 25 semantically equivalent paraphrases of the boot purchase-intent query, generated and validated via mistral-small-latest:
-
[56]
I’d like to buy a pair of boots
-
[57]
I’m looking to purchase some boots
-
[58]
I’m in the market for a new pair of boots
-
[59]
I’d like to shop for boots
-
[60]
I’m interested in getting boots
-
[61]
I’m considering purchasing boots
-
[62]
I’m planning to buy a pair of boots
-
[63]
I’m thinking about buying boots
-
[64]
I’m hoping to purchase boots soon
-
[65]
I’m looking for boots to buy
-
[66]
I’m ready to buy boots
-
[67]
I’d like to find boots to purchase
-
[68]
I’m shopping for boots
-
[69]
I’m aiming to buy boots
-
[70]
I’m set on purchasing boots. 12
-
[71]
I’m eager to buy a pair of boots
-
[72]
I’m on the hunt for boots to buy
-
[73]
I’m keen to purchase boots
-
[74]
I’d like to buy some boots
-
[75]
I’m looking for a pair of boots to purchase
-
[76]
I’m planning to shop for boots soon
-
[77]
I want to get myself a pair of boots
-
[78]
I’m thinking about buying boots, specifically a pair. B Individual Model Validity Table 5 reports the null rejection rate at α= 0.05 for the sign test and permutation test for each individual model under the null condition (sneakers vs. sneakers split), across six (N, M, R) configurations. Each cell is based on 200 random sub-samples. The sign test is ove...
This paper was first reviewed by grok-4.3 on June 29, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.