{"id":"853b41e4-3719-4e48-92e1-06105f78b548","arxiv_id":"2607.05554","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.5,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Prompt robustness in LLMs is systematically lower for subjective survey items than for objective questions, with the largest gap under option-order perturbations and strong model–dataset–prompt interactions.","lead":"LLM answers to value and belief surveys are less stable under prompt changes than answers to objective multiple-choice questions, especially when option order changes. The finding warns that single-prompt political or value scores should not be read as fixed model beliefs.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged meaning-preservation assumption.","rationale":"The strongest claim is an empirical interaction (dataset type × prompt category) under a shared forced-choice protocol, supported by consistent heatmaps (Fig. 1), marginal gaps (Table 2), and large Wald tests (Tables 3–5). The only assumption that could re-interpret rather than merely qualify that claim is equal meaning preservation across objective and subjective items—the exact point the reader already flags. Other limitations (family-level naming, no artifacts, forced-choice only) affect reproducibility and generalizability but not the internal validity of the reported consistency differences. Because the concern is already correctly identified and does not overturn the descriptive result, the CONDITIONAL verdict and its rationale remain appropriate; no adjustment is warranted.","tokens_in":10105,"tokens_out":495,"duration_ms":4622,"concrete_test":"Have 3–5 independent annotators rate, for a stratified sample of ~50 Type-II items, whether each paraphrase and logical-equivalent variant preserves the original survey intent (binary + free-text). If >15–20% of variants are judged non-preserving, recompute Table 2 and the Type×prompt GEE after excluding those variants; a material shrinkage of the 0.157 option-order gap or loss of the χ²=175 interaction would confirm the concern lands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly isolates the central interpretive hinge: that the shared perturbation taxonomy (§3.2–3.3, Table 1) preserves intended task meaning equally for Type-I and Type-II items, so that lower consistency on subjective datasets can be read as robustness failure rather than legitimate re-interpretation of ambiguous survey wording. The paper states the criterion explicitly and designs the taxonomy to separate semantic, surface, and answer-presentation sources, but supplies no independent human validation that paraphrase/logical-equivalent variants leave survey-item meaning fixed. That is a real soft spot for causal attribution of the Type-I vs Type-II gap (especially under option-order, Table 2). However, it does not undermine the descriptive claim that consistency differs by dataset type and interacts with prompt category; the GEE results and heatmaps still stand as evidence of differential sensitivity. No stronger internal inconsistency or hidden assumption appears load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that prompt robustness in LLMs is task-dependent: subjective/belief-style survey items are less answer-consistent under meaning-preserving prompt variants than objective multiple-choice items. Four instruction-tuned model families are evaluated on three Type-I datasets (MMLU, ARC, CulturalBench-Easy) and three Type-II datasets (Political Compass Test, ValueBench, WVS). A seven-category perturbation taxonomy (paraphrase, spelling noise, lexical substitution, logical equivalence, label substitution, format variation, option shuffling) is applied under temperature-0 forced-choice decoding with answer normalization. Consistency is defined per item as C_i = max_y n_i(y)/N_i. Binomial GEE models with item clustering yield significant main effects of model, dataset, dataset type, and prompt category, plus large dataset-type × prompt-category and model × prompt-category interactions. Option order produces the largest Type-I vs Type-II gap (0.485 vs 0.328).","tokens_in":10301,"tokens_out":1072,"duration_ms":8080,"significance":"If the result holds, it supplies a concrete, statistically supported caution against treating single-prompt survey responses as stable measures of LLM values or beliefs. The unified design across objective and subjective tasks, the explicit separation of semantic vs surface vs answer-presentation perturbations, and the clustered GEE analysis are genuine strengths relative to prior single-prompt or single-domain sensitivity studies. The work is useful for evaluation methodology and for any paper that reports political-compass or value-survey scores for LLMs. It does not claim a new training method or a closed-form theory; its contribution is empirical and methodological.","major_comments":[{"comment":"§3.2–3.3 and Table 1: The central causal reading of the Type-I/Type-II gap rests on the claim that every perturbation “preserves the intended task” equally for factual and survey items. No independent human validation (e.g., annotator agreement that paraphrase/logical-equivalent variants leave survey-item meaning fixed) is reported. Without that check, lower consistency on Type-II items under option-order or format changes (Table 2) can be read either as robustness failure or as legitimate re-interpretation of ambiguous survey wording. A short validation study or explicit limitation that the gap is descriptive rather than purely causal would make the claim load-bearing-safe.","section":null},{"comment":"§3.1 and Limitations: Models are reported only at family level (Gemma, Llama, Mistral, Qwen) with no checkpoint IDs, parameter sizes, or instruction-tuning recipes. Because the model × prompt interaction is large (Tables 3–4, Appendix Table 6), family-level aggregation leaves open whether the robustness ranking is driven by size, post-training, or particular checkpoints. Exact model identifiers are needed for reproducibility and for interpreting the model-level claims in RQ3.","section":null},{"comment":"§3.4 Eq. (1) and GEE setup: Consistency is defined relative to the majority answer among the observed variants for that item. When N_i is small or the answer distribution is multimodal (plausible for subjective items under option shuffling), the majority label itself is noisy and can inflate or deflate C_i. The paper should report the distribution of N_i per category, sensitivity of results to alternative anchors (e.g., a fixed canonical prompt), and the working correlation structure used in the GEE.","section":null}],"minor_comments":[{"comment":"Figure 1 heatmaps are informative but the color scale and numeric overlays become hard to read at small size; consider a supplementary table of the same means.","section":null},{"comment":"Table 6 (Appendix) appears to reuse the same χ² values for Gemma as the within-dataset Table 5; clarify whether Gemma is the reference or whether the numbers are intentionally identical.","section":null},{"comment":"§6 Limitations correctly notes forced-choice and deterministic decoding; a one-sentence pointer to how open-ended or temperature > 0 settings might change the gap would help readers.","section":null},{"comment":"Minor wording: “Ismithdeen et al.” appears consistently; verify the author spelling against the cited Promptception paper.","section":null},{"comment":"CulturalBench-Easy is justified as Type-I because of fixed keys; a brief note on how many items were retained after any filtering would aid replication.","section":null}],"recommendation":"major_revision","confidential_remarks":"The descriptive claim (consistency differs by dataset type and interacts with prompt category) is solid and the GEE design is appropriate. The main risk is over-claiming a causal “robustness failure” interpretation without meaning-preservation validation. With the three major points addressed, the paper is a clear contribution to evaluation methodology; without them it remains useful but overstated. Scope fit for a CL/ML evaluation venue is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: under a shared forced-choice setup and temperature 0, subjective value/belief items (PCT, ValueBench, WVS) show lower answer consistency than objective MCQs (MMLU, ARC, CulturalBench-Easy), and the gap is largest for option-order and format changes rather than pure lexical paraphrase. The GEE results (dataset-type main effect and especially type × prompt-category interaction) line up with the heatmaps, so the descriptive claim is real.\n\nWhat is new is the controlled head-to-head under one perturbation taxonomy and one consistency metric (majority-based C_i). Prior work already showed prompt sensitivity and survey instability separately; this paper puts both under the same experimental roof and shows the interaction is large. The taxonomy itself is cleanly motivated (semantic vs surface vs answer-presentation), answer normalization is careful, and they correctly treat subjective items as stability questions rather than accuracy questions. The stats are appropriate for clustered item-level data.\n\nSoft spots are real but not load-bearing for the main claim. Models are named only at family level, no code/data/checkpoints are released, and everything is forced-choice + deterministic decoding, so open-ended survey behavior is out of scope. The interpretive hinge the reader flags is fair: we have to trust that paraphrase/logical-equivalent variants preserve intended meaning equally for survey items; there is no human validation of that. That weakens causal talk about “robustness failure” more than it weakens the descriptive pattern that consistency differs by type and interacts with category. Option-order effects are already known; the paper’s contribution is showing they hit subjective items harder under this design.\n\nThis is for people who run or critique LLM value/political surveys and for evaluation methodologists. It is not a new theory of prompting. I would send it to peer review: the design is coherent, the numbers are clear, and the practical warning is worth having on the record even if referees demand exact checkpoints and public artifacts. Worth reading and citing for the comparison, not as a definitive benchmark.","headline":"Solid empirical comparison showing subjective survey items are less prompt-stable than objective MCQs, with large type × perturbation interactions; useful methodological warning, moderate novelty.","tokens_in":11007,"tokens_out":523,"would_cite":true,"duration_ms":4853,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Prompt robustness in LLMs is task-dependent: subjective belief and value questions are less stable under prompt changes than objective questions.","keywords":["prompt robustness","large language models","survey evaluation","objective vs subjective questions","answer consistency","prompt sensitivity","value measurement","option order"],"falsifier":"A replication on the same or expanded datasets that finds no significant dataset-type main effect and no dataset-type by prompt-category interaction in the binomial GEE models would falsify the central claim.","tokens_in":10964,"feed_emoji":"🔄","tokens_out":794,"duration_ms":15772,"temperature":0.7,"pith_summary":"Survey-style evaluations of large language models often treat a single prompted answer as evidence of the model's values, politics, or beliefs. This paper tests whether that practice is reliable by comparing answer consistency on objective questions that have fixed correct answers against subjective questions that ask for opinions or values. Across four instruction-tuned model families and six datasets, subjective items produce lower consistency under meaning-preserving prompt variants such as rewording, formatting, label changes, and especially option reordering. Statistical models show significant effects of model, dataset type, and prompt category, plus a large interaction between dataset type and prompt category. The finding matters because a single survey score can no longer be treated as a stable model trait once the prompt form is allowed to vary.","feed_headline":"Belief questions flip more under prompt tweaks than facts","feed_subtitle":"Option order hits value surveys hardest; single-prompt LLM scores are not stable beliefs.","key_machinery":"Answer consistency ratio for each item (the share of prompt variants that match the majority answer), analyzed with binomial generalized estimating equations that treat repeated variants of the same item as clustered data.","core_discovery":"Prompt robustness depends jointly on question type, the kind of prompt change, and the model. Subjective survey-style questions yield systematically lower answer consistency than objective multiple-choice questions under the same families of meaning-preserving perturbations, and the gap is largest for answer-presentation changes such as option order.","pith_inferences":["Political and value benchmarks may need default protocols that randomize or balance option order and labels before any score is published.","The larger instability on subjective items may partly reflect models treating surface cues (order, labels, framing) as legitimate stance signals rather than pure noise.","Open-ended, non-forced-choice formats could either shrink or enlarge the objective-subjective gap and should be tested as a natural extension.","Model cards that claim value alignment should include multi-prompt consistency ranges rather than single-prompt point estimates."],"forward_implications":["A single prompted answer to a political or value question should not be treated as a stable model belief unless it survives controlled prompt variation.","Survey-style LLM evaluations need to report consistency across multiple prompt families, not only a final score.","Option-order and other answer-presentation perturbations should be treated as higher-priority robustness checks for subjective items than pure lexical paraphrases.","Robustness scores must be reported as a joint function of model, dataset, and prompt category rather than as a single global model property.","Objective and subjective tasks cannot be assumed to share the same sensitivity profile when designing or interpreting evaluations."],"fun_headline_variants":["Belief questions flip more than facts under prompt tweaks","Subjective LLM answers less stable than objective ones","Prompt robustness hinges on question type and change","Option order hits value surveys harder than fact quizzes","Same tweaks break beliefs more than fixed-answer tasks"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The curated prompt changes are assumed to keep the intended task meaning the same for both factual and survey items, so any answer flip can be counted as a robustness failure rather than a legitimate reinterpretation.","fun_headline_variants_meta":{"raw":{"variants":["Belief questions flip more than facts under prompt tweaks","Subjective LLM answers less stable than objective ones","Prompt robustness hinges on question type and change","Option order hits value surveys harder than fact quizzes","Same tweaks break beliefs more than fixed-answer tasks"]},"model":"grok-4.5","effort":"low","cost_usd":0.002866,"raw_usage":{"total_tokens":985,"prompt_tokens":713,"num_sources_used":0,"completion_tokens":56,"cost_in_usd_ticks":28660000,"prompt_tokens_details":{"text_tokens":713,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":216,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":713,"tokens_out":56,"duration_ms":2513,"temperature":1.0,"reasoning_tokens":216,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T05:51:58.889214+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A replication on the same or expanded datasets that finds no significant dataset-type main effect and no dataset-type by prompt-category interaction in the binomial GEE models would falsify the central claim.","supporting_citations":[],"review_version":1}