Pith. sign in

REVIEW 2 major objections 2 minor 78 references

Standard hypothesis tests are invalid under models of generative surveying that include prompt perturbations.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-29 16:23 UTC pith:AZZ4Q7OZ

load-bearing objection The paper shows standard tests fail under a generative survey model with prompt perturbations and offers a permutation test plus budget advice, but the model's match to real LLMs is untested. the 2 major comments →

arxiv 2605.27463 v1 pith:AZZ4Q7OZ submitted 2026-05-26 stat.ME cs.AIstat.AP

When prompt perturbations break your A/B test: A valid statistical test for generative surveying

classification stat.ME cs.AIstat.AP
keywords generative surveyingprompt perturbationshypothesis testingpermutation testA/B testingLLM feedbackstatistical validity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper shows that standard tests such as the sign test and Wilcoxon signed-rank test do not control type I error when applied to generative surveys that account for sensitivity to prompt phrasing. It introduces a permutation test that is valid under a model incorporating realistic perturbation structure. The work also examines how effect estimates vary with model choice and offers guidance on allocating experimental resources across personas, perturbations, and replicates. A sympathetic reader would care because many current applications of LLMs for market research rely on these invalid tests.

Core claim

Under a statistical model for generative surveying that includes realistic perturbation structure, standard hypothesis tests including the sign test and Wilcoxon signed-rank test are invalid, whereas a permutation test is valid. The paper formally characterizes the conditions under which the standard tests fail and demonstrates that both the magnitude and direction of estimated effects are sensitive to the choice of model.

What carries the argument

A permutation test that respects the perturbation structure in the generative surveying model.

Load-bearing premise

The statistical model for generative surveying with realistic perturbation structure accurately describes how LLMs respond to semantically equivalent prompt variations.

What would settle it

Collect data from an LLM survey with multiple perturbations per persona and compare the p-values from the standard sign test and the proposed permutation test; if they frequently disagree on significance, the invalidity claim is supported.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • The sign test and Wilcoxon signed-rank test fail to maintain correct type I error rates when perturbations are present.
  • Effect estimates in generative surveys depend on the specific LLM used, even within the same family.
  • Budget should be allocated across personas, perturbations, and replicates to achieve desired power.
  • Conditions for failure of standard tests can be characterized in terms of the perturbation model parameters.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Existing generative survey results that used standard tests may need to be reanalyzed with the permutation test.
  • The approach could be extended to other settings where input variations affect statistical conclusions, such as in robustness testing for machine learning models.
  • Power calculations under this model can inform minimal sample sizes for future generative surveys.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper claims that standard hypothesis tests (sign test, Wilcoxon signed-rank) are invalid for A/B testing in generative surveying because LLM responses exhibit dependence induced by prompt perturbations. It derives a permutation test that is valid under the proposed generative model, formally characterizes the conditions under which standard tests fail, applies the framework to estimate model parameters in a simple surveying example, analyzes the power of the permutation test, and supplies guidance on allocating budget across personas, perturbations, and replicates. It also shows that effect magnitude and direction can be sensitive to model choice even within the same LLM family.

Significance. If the perturbation model accurately captures the dependence structure of real LLM responses to semantically equivalent prompts, the work supplies a practically important correction to inference in generative surveying. The valid permutation test and budget-allocation results would directly improve the reliability of LLM-based market research. The formal characterization of test failure conditions is a methodological contribution; the parameter estimation and power analysis add concrete guidance. These strengths are tempered by the absence of external validation that the assumed perturbation structure matches observed LLM behavior.

major comments (2)
  1. [§2] §2 (generative model): the invalidity result for the sign and Wilcoxon tests is derived under a specific perturbation dependence structure; without an explicit equation or proof sketch showing how the covariance between perturbed responses violates the exchangeability or independence assumptions of those tests, the central claim cannot be verified.
  2. [§4–5] §4–5 (application and power): parameters are estimated and power is characterized under the model, yet no comparison of the fitted perturbation distribution to empirical LLM output distributions (e.g., response similarity across prompt variants) is reported; this leaves the practical guidance conditional on an untested modeling assumption.
minor comments (2)
  1. [Abstract] Abstract: the phrase 'realistic perturbation structure' is used without a one-sentence gloss of the key dependence assumption.
  2. [Model section] Notation: the distinction between persona-level, perturbation-level, and replicate-level random effects is not always visually clear in the displayed equations.

Simulated Author's Rebuttal

2 responses · 1 unresolved

We thank the referee for the careful reading and for identifying points that will improve the clarity and transparency of the manuscript. We address each major comment below.

read point-by-point responses
  1. Referee: [§2] §2 (generative model): the invalidity result for the sign and Wilcoxon tests is derived under a specific perturbation dependence structure; without an explicit equation or proof sketch showing how the covariance between perturbed responses violates the exchangeability or independence assumptions of those tests, the central claim cannot be verified.

    Authors: Section 2 introduces the generative model in which responses to semantically equivalent prompts share a common perturbation factor, inducing dependence. The text states that this structure violates the assumptions of the sign and Wilcoxon tests, but we agree that an explicit covariance equation and short proof sketch would make the violation transparent. In the revised manuscript we will insert a short derivation showing that Cov(Y_{ij}, Y_{ik}) > 0 for j ≠ k under shared perturbation, which directly breaks the exchangeability required by those tests. revision: yes

  2. Referee: [§4–5] §4–5 (application and power): parameters are estimated and power is characterized under the model, yet no comparison of the fitted perturbation distribution to empirical LLM output distributions (e.g., response similarity across prompt variants) is reported; this leaves the practical guidance conditional on an untested modeling assumption.

    Authors: We acknowledge that the paper contains no direct empirical comparison of the fitted perturbation distribution to observed LLM response similarities. The power and budget-allocation results in §§4–5 are therefore conditional on the modeling assumption. In revision we will add an explicit limitations subsection that states this assumption, motivates it from existing literature on prompt sensitivity, and outlines the empirical checks that would be needed to validate it. We cannot supply the missing comparison without new data collection outside the scope of the present work. revision: partial

standing simulated objections not resolved
  • Empirical validation of the assumed perturbation dependence structure against observed LLM output distributions.

Circularity Check

0 steps flagged

No circularity detected; derivation is model-internal but self-contained

full rationale

The paper defines a generative surveying model that incorporates prompt perturbations and derives the invalidity of sign/Wilcoxon tests (plus validity of a permutation test) strictly under that model. No equations, fitted parameters, or self-citations appear in the abstract or provided text that would reduce any claim to its inputs by construction. The load-bearing assumption that the model captures 'realistic' structure is an external modeling choice, not a definitional loop or renamed fit. This is the normal case of a paper working inside its stated assumptions without circular reduction.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract-only review supplies no explicit free parameters, axioms, or invented entities; the central model is described only at the level of 'realistic perturbation structure' without further decomposition.

pith-pipeline@v0.9.1-grok · 5700 in / 1058 out tokens · 31930 ms · 2026-06-29T16:23:20.840178+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of When prompt perturbations break your A/B test: A valid statistical test for generative surveying." pith.science (2026). https://pith.science/paper/AZZ4Q7OZ

@misc{pith2026260527463,
  author       = {Pith},
  title        = {Pith review of: When prompt perturbations break your A/B test: A valid statistical test for generative surveying},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AZZ4Q7OZ}},
  note         = {Machine review of arXiv:2605.27463}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Generative surveying -- where collections of LLM-based personas provide feedback on messages -- has emerged as a cheap and scalable alternative to traditional market research. However, LLMs are sensitive to small variations in prompt design and conclusions drawn from generative surveys may depend on arbitrary phrasing choices. Controlling for this sensitivity requires including semantically equivalent perturbations in the analysis. In this paper, we show that standard hypothesis tests, including the sign test and Wilcoxon signed-rank test, are invalid under a statistical model for generative surveying that includes realistic perturbation structure. We propose a permutation test that is valid under this model and formally characterize the conditions under which standard tests fail. Applying our framework to a simple generative surveying problem, we estimate relevant parameters, characterize the power of the permutation test under realistic conditions, and provide practical guidance on budget allocation across personas, perturbations, and replicates. Finally, we show that both the magnitude and direction of the estimated effect are sensitive to the choice of model, even within the same model family.

Figures

Figures reproduced from arXiv: 2605.27463 by Carey Priebe, Hayden Helm.

Figure 1
Figure 1. Figure 1: (Left) Illustration of persona-based generative message testing for a single persona. The illustrated [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (Left) Simple statistical model for persona-based generative surveying where outcomes are binary. Light [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Empirical CDF of p-values under the null [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Power of the permutation test as a function of total query budget [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Effect size (βˆ 1) versus number of active pa￾rameters for models in the Mistral-3 and Qwen-3 fami￾lies. Effect size is non-monotonic in model scale within both families. Same scale models across families can produce estimates of opposite sign (e.g., Mistral-8B and Qwen-3-8B estimate effects in opposite directions) meaning that model choice alone can determine the con￾clusion of a generative survey. should… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

78 extracted references · 4 canonical work pages · 2 internal anchors

  1. [1]

    In Proceedings of the 40th International Conference on Machine Learning, pages 337–371

    Using large language models to simulate mul- tiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning, pages 337–371. John Arbuthnott. 1710. An argument for divine prov- idence, taken from the constant regularity observ’d in the births of both sexes. by dr. john arbuthnott, physitian in ordi...

  2. [2]

    Out of one, many: Using language mod- els to simulate human samples.Political Analysis, 31(3):337–351. Peter J. Bickel and Kjell A. Doksum. 1977.Mathe- matical Statistics: Basic Ideas and Selected Topics. Holden-Day, San Francisco. Paul P. Biemer and Lars E. Lyberg. 2003.Introduction to Survey Quality. John Wiley & Sons, Hoboken, NJ. James Bisbee, Joshua ...

  3. [3]

    In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–19

    Evaluating large language models in generat- ing synthetic HCI research data: A case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–19. Hayden Helm, Tianyi Chen, Harvey McGuinness, Paige Lee, Brandon Duderstadt, and Carey E. Priebe. 2025. Toward a digital twin of U.S. congress.arXiv preprint arXiv:2505.0000...

  4. [4]

    A statistical turing test for generative mod- els.Preprint, arXiv:2309.08913. John J. Horton, Apostolos Filippas, and Benjamin S. Manning. 2023. Large language models as simulated economic agents: What can we learn from Homo Silicus? Technical Report 31122, National Bureau of Economic Research. Jon A. Krosnick and Stanley Presser. 2010. Question and quest...

  5. [5]

    LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

    Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109. Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, pages 29971–30004. Howard Schuman and Stanley Presser. 1981....

  6. [6]

    I’d like to buy a pair of sneakers

  7. [7]

    I’m looking to purchase some new trainers

  8. [8]

    I want to shop for a pair of athletic casual shoes

  9. [9]

    I’m considering buying sneakers

  10. [10]

    I’m in the market for a new pair of sneakers

  11. [11]

    I’m thinking about purchasing some trainers

  12. [12]

    I’d like to get myself a pair of sneakers

  13. [13]

    I’m interested in buying athletic casual shoes. 11

  14. [14]

    I want to buy some new sneakers

  15. [15]

    I’m planning to purchase a pair of trainers

  16. [16]

    I’m looking to get a pair of sneakers

  17. [17]

    I’d like to invest in some new athletic casual shoes

  18. [18]

    I’m considering buying a pair of sneakers

  19. [19]

    I want to shop for trainers

  20. [20]

    I’m interested in purchasing sneakers

  21. [21]

    I’d like to grab a pair of athletic casual shoes

  22. [22]

    I’m thinking about getting some new sneakers

  23. [23]

    I’m in the mood to buy trainers

  24. [24]

    I’m looking to acquire a pair of sneakers

  25. [25]

    I want to pick up some new athletic casual shoes

  26. [26]

    I’m considering shopping for sneakers

  27. [27]

    I’d like to obtain a pair of trainers

  28. [28]

    I’m planning to get myself some new sneakers

  29. [29]

    I want to purchase a pair of trainers

  30. [30]

    I’m looking to purchase some athletic shoes

  31. [31]

    Can you show me some sneakers to buy?

  32. [32]

    I want to shop for a new pair of trainers

  33. [33]

    I’m interested in buying some casual athletic shoes

  34. [34]

    Do you have sneakers available for purchase?

  35. [35]

    I’m thinking about getting a pair of sneakers

  36. [36]

    I’d like to explore options for buying sneakers

  37. [37]

    I want to find a pair of sneakers to buy

  38. [38]

    I’m in the market for some new trainers

  39. [39]

    I’m considering purchasing a pair of sneakers

  40. [40]

    I’d like to buy some comfortable athletic shoes

  41. [41]

    I’m looking to add sneakers to my collection

  42. [42]

    I want to check out sneakers to purchase

  43. [43]

    I’m interested in buying a pair of running shoes

  44. [44]

    I’d like to shop for some stylish sneakers

  45. [45]

    I’m hoping to find sneakers to buy soon

  46. [46]

    I want to purchase a pair of casual sneakers

  47. [47]

    I’m considering buying some athletic footwear

  48. [48]

    I’d like to explore sneaker options for purchase

  49. [49]

    I want to buy some new trainers

  50. [50]

    I’m interested in purchasing some sneakers

  51. [51]

    I’d like to add a pair of sneakers to my wardrobe

  52. [52]

    I’m considering buying some casual athletic shoes

  53. [53]

    I’d like to buy a new pair of sneakers

  54. [54]

    Could you show me some trainers that are available for purchase?

  55. [55]

    A.4 Boot Perturbations (M= 25) All 25 semantically equivalent paraphrases of the boot purchase-intent query, generated and validated via mistral-small-latest:

    I’m looking to shop for a pair of athletic casual shoes. A.4 Boot Perturbations (M= 25) All 25 semantically equivalent paraphrases of the boot purchase-intent query, generated and validated via mistral-small-latest:

  56. [56]

    I’d like to buy a pair of boots

  57. [57]

    I’m looking to purchase some boots

  58. [58]

    I’m in the market for a new pair of boots

  59. [59]

    I’d like to shop for boots

  60. [60]

    I’m interested in getting boots

  61. [61]

    I’m considering purchasing boots

  62. [62]

    I’m planning to buy a pair of boots

  63. [63]

    I’m thinking about buying boots

  64. [64]

    I’m hoping to purchase boots soon

  65. [65]

    I’m looking for boots to buy

  66. [66]

    I’m ready to buy boots

  67. [67]

    I’d like to find boots to purchase

  68. [68]

    I’m shopping for boots

  69. [69]

    I’m aiming to buy boots

  70. [70]

    I’m set on purchasing boots. 12

  71. [71]

    I’m eager to buy a pair of boots

  72. [72]

    I’m on the hunt for boots to buy

  73. [73]

    I’m keen to purchase boots

  74. [74]

    I’d like to buy some boots

  75. [75]

    I’m looking for a pair of boots to purchase

  76. [76]

    I’m planning to shop for boots soon

  77. [77]

    I want to get myself a pair of boots

  78. [78]

    I’m thinking about buying boots, specifically a pair. B Individual Model Validity Table 5 reports the null rejection rate at α= 0.05 for the sign test and permutation test for each individual model under the null condition (sneakers vs. sneakers split), across six (N, M, R) configurations. Each cell is based on 200 random sub-samples. The sign test is ove...