{"id":"46819067-1972-497e-a3b1-6ecdc44632a3","arxiv_id":"2605.29027","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Tonal variations in prompts cause systematic but model-dependent accuracy changes in LLMs on objective multiple-choice questions.","lead":"This paper tests whether different tones in prompts change how accurately LLMs answer multiple-choice questions across two datasets and four models. Results indicate tone effects are systematic yet vary sharply by model, advising against assuming consistent LLM reliability.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Tone variants may alter question semantics or difficulty, preventing clean isolation of tone as the sole variable.","rationale":"The reader's weakest_assumption directly identifies the same isolation requirement. Because the full text was not available to the reader, the current assessment treats the concern as still open rather than refuted by unreported methods details. No other internal inconsistency (e.g., statistical reporting or model selection) rises to the same load-bearing level for the central claim.","tokens_in":1653,"tokens_out":351,"duration_ms":15538,"concrete_test":"Take the 50 base questions and their five tone variants; have three independent annotators (blind to condition) rate each variant on a 1-5 scale for (a) factual equivalence to the base and (b) unchanged difficulty; compute mean score per variant. If any variant averages below 4.0 or shows >15% inter-annotator disagreement on equivalence, re-run the main accuracy experiments on only the verified-equivalent subset and check whether the reported tonal swings remain statistically significant.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the five and seven tone variants preserve factual content, difficulty, and intended meaning exactly. The 50-question and 570-question MMLU datasets are transformed into tonal versions, but the paper provides no reported verification (human equivalence ratings, semantic similarity thresholds, or side-by-side difficulty comparisons) that the transformations are meaning-preserving. If any tone variant introduces subtle rephrasing that changes clarity or cognitive load, observed accuracy differences cannot be attributed to tone alone. This assumption is load-bearing because the headline result (systematic but model-dependent tonal effects) collapses without it.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper investigates whether and how tonal variations in prompts affect LLM accuracy on objective multiple-choice questions. Using a 50-question dataset with five tone variants and a 570-question MMLU subset with seven tone variants, it evaluates four models (ChatGPT-4o, ChatGPT-5-nano, Gemini 2.5 Flash, Gemini 2.5 Flash Lite). The central claim is that tonal effects are systematic but highly model-dependent, with some models exhibiting small yet statistically significant accuracy shifts and others showing large swings; the work also reports subject-level differences in tone sensitivity and proposes a routing framework for how tones influence internal reasoning modes.","tokens_in":1773,"tokens_out":458,"duration_ms":21279,"significance":"If the tone variants are shown to preserve question content and difficulty, the empirical results across multiple models and datasets would usefully demonstrate prompt sensitivity in LLMs and caution against assuming tone-robust performance. The multi-model, multi-subject design and identification of model-specific patterns constitute a concrete contribution to prompt engineering literature. The routing framework, if better substantiated, could provide explanatory value beyond the accuracy measurements.","major_comments":[{"comment":"Methods (tone variant generation and dataset construction): No verification is reported (human equivalence ratings, semantic similarity thresholds, or side-by-side difficulty comparisons) that the five- and seven-tone transformations preserve factual content, difficulty, and intended meaning. This assumption is load-bearing; without it, accuracy differences cannot be cleanly attributed to tone rather than unintended changes in clarity or cognitive load.","section":"Methods (tone variant generation and dataset construction)"},{"comment":"Abstract and Results: The claim of statistically significant shifts and model differences is stated without visible details on the statistical tests employed, handling of multiple comparisons, per-condition sample sizes, or raw data availability, preventing assessment of whether the data support the reported significance levels.","section":"Abstract and Results"}],"minor_comments":[{"comment":"The routing framework is introduced in the discussion but lacks a dedicated methods subsection or quantitative validation against the accuracy results.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment below and will revise the manuscript to strengthen the methods and statistical reporting sections.","responses":[{"response":"We agree this verification is important for attributing effects to tone. The variants were produced by prompting an LLM to rephrase while explicitly preserving meaning and facts, but the manuscript omitted explicit checks. In revision we will add a methods subsection with the exact generation prompts, report embedding-based semantic similarity (cosine > 0.95 threshold) between base and toned questions, and include a small human equivalence study (n=20 raters) confirming no change in perceived difficulty or content.","revision_made":"yes","referee_comment":"Methods (tone variant generation and dataset construction): No verification is reported (human equivalence ratings, semantic similarity thresholds, or side-by-side difficulty comparisons) that the five- and seven-tone transformations preserve factual content, difficulty, and intended meaning. This assumption is load-bearing; without it, accuracy differences cannot be cleanly attributed to tone rather than unintended changes in clarity or cognitive load."},{"response":"We acknowledge the need for fuller statistical transparency. Significance was assessed via McNemar’s tests on paired per-question outcomes (n=50 and n=570 per tone), with Bonferroni correction across the 4 models × 7 tones. We will expand the results section to state the exact tests, report all p-values and effect sizes, confirm sample sizes, and add a data-availability statement with a link to the anonymized per-question accuracy matrix and analysis scripts.","revision_made":"yes","referee_comment":"Abstract and Results: The claim of statistically significant shifts and model differences is stated without visible details on the statistical tests employed, handling of multiple comparisons, per-condition sample sizes, or raw data availability, preventing assessment of whether the data support the reported significance levels."}],"tokens_in":1351,"tokens_out":416,"duration_ms":19757,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The headline result is that some LLMs show clear accuracy changes across tone variants on multiple-choice questions while others stay mostly flat, and the size of the effect also varies by subject. They ran this on four recent cost-efficient models using a 50-question set with five tones and a 570-question MMLU slice with seven tones, then added a routing framework to tie tones to different reasoning modes.\n\nThe concrete numbers on model and subject differences are the useful part. Running the same base questions through tonal rewrites and reporting statistical significance gives practitioners something to look at when they worry about prompt sensitivity in evaluations. The subject-level breakdown adds a bit more texture than a single aggregate score.\n\nThe load-bearing problem is the missing check that the tone variants left factual content, difficulty, and intended meaning untouched. Rephrasing for tone almost always alters wording, and nothing in the work shows human equivalence ratings, semantic similarity thresholds, or side-by-side difficulty comparisons. Without that step the accuracy gaps cannot be cleanly attributed to tone rather than to unintended changes in clarity or cognitive load.\n\nThe datasets are modest but workable for this style of experiment, and the models are the right practical ones to test. The routing framework stays at the level of explanation rather than a tested mechanism. Citation patterns look standard for prompt-sensitivity work.\n\nThis is for people who run LLM evaluations or build routing systems and want to see whether tone is worth controlling. It is worth sending to referees so they can ask for the equivalence verification and more detail on how the tones were generated. The empirical observation has enough practical bite to justify the time even if the current version needs tightening.","headline":"Tone shifts accuracy on MMLU subsets in model-dependent ways, but the variants likely change question content enough to blur what is actually being measured.","tokens_in":2223,"tokens_out":409,"would_cite":false,"duration_ms":20665,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Tonal variations in prompts cause systematic but model-dependent shifts in LLM accuracy on multiple-choice questions.","keywords":["LLM performance","prompt tone","multiple-choice questions","model-dependent effects","MMLU","prompt engineering","accuracy variation","reasoning modes"],"falsifier":"A larger replication experiment on the same MMLU subset that finds no statistically significant accuracy differences across any of the seven tones for all four models would falsify the claim of systematic tonal effects.","tokens_in":2555,"feed_emoji":"🤖","tokens_out":465,"duration_ms":22410,"temperature":0.7,"pith_summary":"The paper tests whether changing the tone of prompts alters how accurately large language models answer objective multiple-choice questions. It applies five to seven tone variants to two fixed datasets—one small set of 50 questions and a 570-question slice of MMLU covering 57 subjects—while holding factual content constant. Four cost-efficient models are evaluated, revealing that tonal effects appear consistently within each model yet differ sharply in size across models. Some models register only small statistically significant changes while others show large accuracy swings, and the study notes that sensitivity also varies by subject. A routing framework is offered to account for how tones may steer internal reasoning modes, leading to the practical caution that users cannot assume tone-independent reliability.","feed_headline":"Tone shifts LLM accuracy in model-specific patterns","feed_subtitle":"Tests across four models and MMLU questions show modest changes for some and large swings for others when prompt tone varies.","key_machinery":"A routing framework that explains how different prompt tones may attune or switch among internal reasoning modes within an LLM.","core_discovery":"Across models, tonal effects are systematic but highly model-dependent. Some models show small, yet statistically significant, shifts, while others exhibit large accuracy swings across tones. Further, we identify subject-level differences in tone sensitivity and present a routing framework to explain how tones may attune internal reasoning modes. Our findings caution users against assuming tone-robust reliability in LLM deployments.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Tone alters LLM accuracy by model","LLM performance varies with prompt tone per model","Tones cause model-dependent accuracy changes in LLMs","Prompt tone leads to different LLM results across models"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The tone variants can be applied without changing the factual content, difficulty level, or intended meaning of the questions.","fun_headline_variants_meta":{"raw":{"variants":["Tone alters LLM accuracy by model","LLM performance varies with prompt tone per model","Tones cause model-dependent accuracy changes in LLMs","Prompt tone leads to different LLM results across models"]},"model":"grok-4.3","cost_usd":0.005301,"raw_usage":{"total_tokens":2537,"prompt_tokens":618,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":53012000,"prompt_tokens_details":{"text_tokens":618,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1864,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":618,"tokens_out":55,"duration_ms":15070,"temperature":1.0,"reasoning_tokens":1864,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T12:17:49.626536+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A larger replication experiment on the same MMLU subset that finds no statistically significant accuracy differences across any of the seven tones for all four models would falsify the claim of systematic tonal effects.","supporting_citations":[],"review_version":1}