{"id":"7499d79e-7f02-4059-82ba-d1db530c00b0","arxiv_id":"2605.27463","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Standard hypothesis tests fail for generative surveys under realistic prompt perturbations; a permutation test is valid and practical guidance on budget allocation is given.","lead":"The paper models generative surveying with LLMs and shows that standard A/B tests like the sign test become invalid once prompt perturbations are included. It proposes a permutation test that remains valid under this perturbation structure and notes that effect estimates vary by model choice.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Central claim of test invalidity holds only under the paper's specific perturbation model, whose form is not shown to match LLM behavior.","rationale":"The reader's weakest assumption directly identifies the same point. Because the full manuscript was referenced but the reader's verdict was formed on the abstract, the model-validity gap remains the single load-bearing concern; no internal inconsistency or derivation error is visible from the given material.","tokens_in":1658,"tokens_out":311,"duration_ms":21819,"concrete_test":"Fit the paper's perturbation model parameters to a new dataset of 50 personas each queried with 10 semantically equivalent prompt variants on the same A/B messages; recompute the sign-test type-I error under the fitted model versus the empirical response distribution. If the empirical error rate stays near nominal while the model predicts inflation, the invalidity claim does not apply to observed LLM behavior.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper defines a generative surveying model that includes prompt perturbations and derives that sign/Wilcoxon tests are invalid while a permutation test is valid. This derivation is internal to the model. The load-bearing step is the claim that the model captures 'realistic perturbation structure'; without that, the invalidity result does not transfer to actual LLM surveying. The abstract and title emphasize this model, yet the provided text supplies no external check (e.g., fitted parameters vs. real LLM output distributions) that would confirm the dependence structure between personas, perturbations, and responses.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that standard hypothesis tests (sign test, Wilcoxon signed-rank) are invalid for A/B testing in generative surveying because LLM responses exhibit dependence induced by prompt perturbations. It derives a permutation test that is valid under the proposed generative model, formally characterizes the conditions under which standard tests fail, applies the framework to estimate model parameters in a simple surveying example, analyzes the power of the permutation test, and supplies guidance on allocating budget across personas, perturbations, and replicates. It also shows that effect magnitude and direction can be sensitive to model choice even within the same LLM family.","tokens_in":1768,"tokens_out":490,"duration_ms":38210,"significance":"If the perturbation model accurately captures the dependence structure of real LLM responses to semantically equivalent prompts, the work supplies a practically important correction to inference in generative surveying. The valid permutation test and budget-allocation results would directly improve the reliability of LLM-based market research. The formal characterization of test failure conditions is a methodological contribution; the parameter estimation and power analysis add concrete guidance. These strengths are tempered by the absence of external validation that the assumed perturbation structure matches observed LLM behavior.","major_comments":[{"comment":"§2 (generative model): the invalidity result for the sign and Wilcoxon tests is derived under a specific perturbation dependence structure; without an explicit equation or proof sketch showing how the covariance between perturbed responses violates the exchangeability or independence assumptions of those tests, the central claim cannot be verified.","section":"§2"},{"comment":"§4–5 (application and power): parameters are estimated and power is characterized under the model, yet no comparison of the fitted perturbation distribution to empirical LLM output distributions (e.g., response similarity across prompt variants) is reported; this leaves the practical guidance conditional on an untested modeling assumption.","section":"§4–5"}],"minor_comments":[{"comment":"Abstract: the phrase 'realistic perturbation structure' is used without a one-sentence gloss of the key dependence assumption.","section":"Abstract"},{"comment":"Notation: the distinction between persona-level, perturbation-level, and replicate-level random effects is not always visually clear in the displayed equations.","section":"Model section"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and for identifying points that will improve the clarity and transparency of the manuscript. We address each major comment below.","responses":[{"response":"Section 2 introduces the generative model in which responses to semantically equivalent prompts share a common perturbation factor, inducing dependence. The text states that this structure violates the assumptions of the sign and Wilcoxon tests, but we agree that an explicit covariance equation and short proof sketch would make the violation transparent. In the revised manuscript we will insert a short derivation showing that Cov(Y_{ij}, Y_{ik}) > 0 for j ≠ k under shared perturbation, which directly breaks the exchangeability required by those tests.","revision_made":"yes","referee_comment":"[§2] §2 (generative model): the invalidity result for the sign and Wilcoxon tests is derived under a specific perturbation dependence structure; without an explicit equation or proof sketch showing how the covariance between perturbed responses violates the exchangeability or independence assumptions of those tests, the central claim cannot be verified."},{"response":"We acknowledge that the paper contains no direct empirical comparison of the fitted perturbation distribution to observed LLM response similarities. The power and budget-allocation results in §§4–5 are therefore conditional on the modeling assumption. In revision we will add an explicit limitations subsection that states this assumption, motivates it from existing literature on prompt sensitivity, and outlines the empirical checks that would be needed to validate it. We cannot supply the missing comparison without new data collection outside the scope of the present work.","revision_made":"partial","referee_comment":"[§4–5] §4–5 (application and power): parameters are estimated and power is characterized under the model, yet no comparison of the fitted perturbation distribution to empirical LLM output distributions (e.g., response similarity across prompt variants) is reported; this leaves the practical guidance conditional on an untested modeling assumption."}],"tokens_in":1341,"tokens_out":440,"duration_ms":33477,"standing_objections":["Empirical validation of the assumed perturbation dependence structure against observed LLM output distributions."]},"desk_editor":{"model":"grok-4.3","letter":"The main things to know are that under their model of LLM personas plus prompt perturbations, the sign test and Wilcoxon signed-rank test become invalid due to induced dependence, while a permutation test stays valid, and they supply power calculations and sample allocation guidance for a simple case. They also note that effect size and direction can shift with model choice even inside the same family.\n\nWhat is new is the framing of prompt perturbations as a source of statistical dependence that breaks common nonparametric tests in this generative surveying setup. The paper does a clean job deriving the conditions for test failure and showing the permutation test works inside the model. The budget allocation section and the caution about model sensitivity are the parts that would actually help someone running these surveys.\n\nThe soft spot is the lack of any check that the assumed perturbation structure matches how LLMs actually behave. The invalidity result is internal to the model; without evidence that the dependence between personas, perturbations, and responses looks like real LLM output, the practical claim does not carry over. The abstract mentions estimating parameters but gives no detail on data, fitting method, or external validation, so that step remains unverified.\n\nThis is for applied statisticians or product teams using LLMs for cheap surveys. A reader who needs a dependence-aware test for this exact setting would get value from the permutation test and the allocation tips. The internal logic looks like standard permutation test application, so the paper deserves a serious referee to verify the derivations and push for some external grounding of the model.","headline":"The paper shows standard tests fail under a generative survey model with prompt perturbations and offers a permutation test plus budget advice, but the model's match to real LLMs is untested.","tokens_in":2255,"tokens_out":383,"would_cite":false,"duration_ms":30048,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Standard hypothesis tests are invalid under models of generative surveying that include prompt perturbations.","keywords":["generative surveying","prompt perturbations","hypothesis testing","permutation test","A/B testing","LLM feedback","statistical validity"],"falsifier":"Collect data from an LLM survey with multiple perturbations per persona and compare the p-values from the standard sign test and the proposed permutation test; if they frequently disagree on significance, the invalidity claim is supported.","tokens_in":2541,"feed_emoji":"📊","tokens_out":574,"duration_ms":31281,"temperature":0.7,"pith_summary":"The paper shows that standard tests such as the sign test and Wilcoxon signed-rank test do not control type I error when applied to generative surveys that account for sensitivity to prompt phrasing. It introduces a permutation test that is valid under a model incorporating realistic perturbation structure. The work also examines how effect estimates vary with model choice and offers guidance on allocating experimental resources across personas, perturbations, and replicates. A sympathetic reader would care because many current applications of LLMs for market research rely on these invalid tests.","feed_headline":"Perturbations break standard A/B tests for LLM surveys","feed_subtitle":"A permutation test stays valid when including prompt variations in generative surveying of personas.","key_machinery":"A permutation test that respects the perturbation structure in the generative surveying model.","core_discovery":"Under a statistical model for generative surveying that includes realistic perturbation structure, standard hypothesis tests including the sign test and Wilcoxon signed-rank test are invalid, whereas a permutation test is valid. The paper formally characterizes the conditions under which the standard tests fail and demonstrates that both the magnitude and direction of estimated effects are sensitive to the choice of model.","pith_inferences":["Existing generative survey results that used standard tests may need to be reanalyzed with the permutation test.","The approach could be extended to other settings where input variations affect statistical conclusions, such as in robustness testing for machine learning models.","Power calculations under this model can inform minimal sample sizes for future generative surveys."],"forward_implications":["The sign test and Wilcoxon signed-rank test fail to maintain correct type I error rates when perturbations are present.","Effect estimates in generative surveys depend on the specific LLM used, even within the same family.","Budget should be allocated across personas, perturbations, and replicates to achieve desired power.","Conditions for failure of standard tests can be characterized in terms of the perturbation model parameters."],"fun_headline_variants":["Prompt perturbations invalidate standard tests for LLM surveys","Permutation test valid for perturbed generative surveying of personas","Standard tests fail with prompt variations in generative surveys","Effect magnitude depends on model in perturbed surveys","Sign and Wilcoxon tests fail under generative survey perturbations"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The statistical model for generative surveying with realistic perturbation structure accurately describes how LLMs respond to semantically equivalent prompt variations.","fun_headline_variants_meta":{"raw":{"variants":["Prompt perturbations invalidate standard tests for LLM surveys","Permutation test valid for perturbed generative surveying of personas","Standard tests fail with prompt variations in generative surveys","Effect magnitude depends on model in perturbed surveys","Sign and Wilcoxon tests fail under generative survey perturbations"]},"model":"grok-4.3","cost_usd":0.004341,"raw_usage":{"total_tokens":2146,"prompt_tokens":604,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":43412000,"prompt_tokens_details":{"text_tokens":604,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1474,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":604,"tokens_out":68,"duration_ms":16610,"temperature":1.0,"reasoning_tokens":1474,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T16:23:20.840178+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Collect data from an LLM survey with multiple perturbations per persona and compare the p-values from the standard sign test and the proposed permutation test; if they frequently disagree on significance, the invalidity claim is supported.","supporting_citations":[],"review_version":1}