{"id":"ea44b745-0cce-4f01-93bf-9f001a4c2c95","arxiv_id":"2607.21498","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"LLMs overuse the 'not X, but Y' self-correction pattern in persuasive registers and underuse it in informal Q&A; a prompt or a detachable LoRA dial adjusts it to human levels.","lead":"Large language models systematically over-use the ancient rhetorical figure of self-correction ('This is not a course; it is a journey') in persuasive genres and under-use it in informal writing. The paper introduces an Epanorthosis Index and shows that a one-line prompt or a small LoRA adapter can calibrate the habit to human rates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Detector's 0.17 precision on human text inflates human baseline densities, so the EI ratios and the two-sided miscalibration claim are not established.","rationale":"The reader's weakest assumption identifies precisely the detector's poor precision on human text. This is the most load-bearing concern because the paper's central quantitative claim is calibration relative to human rates: every EI value in Table 1 divides by H_g. If H_g is inflated, both the overshoot and undershoot findings are suspect, and the paper's headline 'mis-calibration by register in both directions' collapses. The concern is not merely external criticism; the paper itself reports precision 0.17 on human text and acknowledges the asymmetry, yet still computes ratios and p-values from those baselines. This is an internal validity issue, not a disagreement with consensus.\n\nOther concerns—uncorrected p-values, the single model family, and the untested RLHF attribution—are real but secondary. The p-value issue is explicitly disclosed and the results labeled suggestive; the RLHF attribution is explicitly called inferential. The detector validity problem is different: it affects the very numbers that are presented as the empirical foundation. Without a correction or bound on false positives in human windows, the Table 1 effects cannot be interpreted as rhetorical miscalibration.\n\nThe recommended verdict remains CONDITIONAL/UNCHANGED because the paper is transparent about its pilot status and the qualitative phenomenon is plausibly real. The condition is that the authors must re-analyze the human baselines with a validated detector or manual annotation. This is exactly the reader's condition, so no verdict change is needed.","tokens_in":16936,"tokens_out":4268,"duration_ms":44677,"concrete_test":"Manually annotate all human windows used in Table 1, or a substantial random sample (at minimum the 19 oratory and 75 Q&A windows), using the same adjudicated protocol as Appendix B, and recompute H_g for each genre from true-positive corrections only. Then recompute EI_g and the permutation p-values. If the Q&A undershoot loses significance or the oratory EI drops below ~1.8, the two-sided miscalibration claim is not supported. Complementary check: validate the English detector on a human gold standard (not just Italian) and report precision by genre; if precision on human text is anywhere near 0.17, the Table 1 baselines must be adjusted before any EI interpretation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Epanorthosis Index is defined as EI_g = D_g/H_g (Section 4.1), so the human baseline H_g is the denominator of every calibration claim. The only validation of the detector reports precision 0.45 overall, but 0.17 on human windows versus 0.82 on model windows (Appendix B, Table B.2; Section 4). Since H_g is counted by the same lexical channel, it is inflated by ordinary adversatives misread as corrections. For informal Q&A (H=8.2, EI=0.2, p=0.05), even a modest overcount of H would materially reduce or erase the undershoot; for oratory (H=14.9), inflation is absorbed into the denominator, so the reported 'twofold' overshoot and the p=0.03 are not trustworthy. Moreover, the permutation test on per-window densities is biased: with precision 0.82 on model windows and 0.17 on human windows, the test compares a clean signal to a noisy one, so it can return significance for detector behavior rather than for rhetorical behavior. The English detector is not separately validated (Limitations), so the Italian 0.17 figure is the only evidence we have. The paper explicitly acknowledges this asymmetry but does not correct or bound its effect on Table 1, even though EI ratios are the paper's central quantitative output.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that LLMs systematically overuse epanorthosis ('Not X. Y' self-correction) because of a training distribution rich in promotional prose and because RLHF rewards confident phrasing, with left-to-right generation only an amplifier. It proposes an Epanorthosis Index (density relative to a genre-specific human baseline), reports an English pilot study claiming bidirectional miscalibration (overshoot in oratory, undershoot in informal Q&A), and presents mitigation results: a LoRA adapter with an α-scaling 'dial', a one-line prompting intervention in Italian, and training-free best-of-n/rewrite baselines. The paper is candid about limitations, including per-genre p-values that do not survive multiple-comparison correction, detector precision asymmetry, small samples, and the absence of a base-model-versus-aligned comparison.","tokens_in":17276,"tokens_out":3889,"duration_ms":43335,"significance":"If the measurement were reliable, the paper would make a useful conceptual contribution: it connects a classical rhetorical figure to a measurable stylistic property of LLMs, proposes a register-relative calibration target rather than simple suppression, and provides reproducible artifacts (LoRA recipe, adapter weights, evaluation scripts, a negative DPO result). The framing of mitigation as calibration to human rates rather than elimination is thoughtful. However, the central quantitative claims currently rest on pilot-scale evidence with a detector that is much noisier on human text than on model text, uncorrected multiple testing, and no significance tests for the mitigation tables. The significance is therefore conditional: the paper is a promising research programme, not yet an established finding.","major_comments":[{"comment":"The detector's precision asymmetry is load-bearing for the central EI ratios. Table B.2 reports precision 0.17 on human windows versus 0.82 on model windows. Since EI_g = D_g/H_g divides by the human baseline H_g, an inflated H_g (due to ordinary adversatives misclassified as epanorthosis) directly deflates every EI value in Table 1. This could erase or reverse the reported overshoot/undershoot pattern. The English detector is not separately validated, so the Italian figures are the only evidence. The manuscript acknowledges the asymmetry but does not bound, correct, or sensitivity-test its effect on Table 1. I request a precision/recall-adjusted estimate of H_g, a sensitivity analysis, or a re-annotation of the human windows before the EI claims can be accepted.","section":"§4.1, Eq. (EI), Appendix B, Table B.2"},{"comment":"The abstract and Section 4.1 claim 'mis-calibration by register in both directions' with oratory p=0.03 and informal Q&A p=0.05. The Limitations state that under strict multiple-comparison correction neither result survives. As reported, the two headline effects are at best suggestive. The abstract and conclusion should be tempered to explicitly say these are exploratory, and Table 1 should report corrected p-values or clearly mark the uncorrected values. This is not a fatal flaw given the paper's transparency, but the central claim of bidirectional miscalibration is not currently established at conventional significance levels.","section":"§4.1, Limitations (Statistical power and multiplicity)"},{"comment":"The mitigation claims are based on point estimates without significance tests: six generations per prompt in Table 2 and per-genre baselines resting on two prompts in Table 5. The 'dial' claim—specifically that argument reaches the human rate at α=0.75—rests on a single point estimate (12.3) with no confidence interval. The promotional spike at α=0.25 (96.0, above the α=0 baseline of 61.4) could be noise or a real non-monotonic effect; without uncertainty bounds it cannot be interpreted. Additionally, the content-fidelity evaluation prescribed in §7.7 is not reported, so the adapter's principal risk remains unmeasured. Please provide confidence intervals, more generations, or a significance test, and report the content-fidelity results before presenting the adapter as a validated calibration tool.","section":"§7.2, Table 2; §7.9, Table 5"},{"comment":"The abstract states that overuse is 'a trained disposition, driven mainly by training distribution and preference tuning (RLHF)', but the paper's own measurements do not test this causal claim: there is no base-model-versus-aligned comparison, and all three models come from one alignment pipeline. The section correctly labels the RLHF attribution as inferential, but this caveat is absent from the abstract and conclusion. Either soften the causal attribution to an explicitly stated hypothesis, or add the base/aligned comparison that would test it. This is load-bearing because the paper's framing ('trained disposition') is presented as an explanation, not merely a conjecture.","section":"§3, Limitations (Causal attribution, Single model family)"}],"minor_comments":[{"comment":"The claim that models use fewer neutral connectives (1.5 vs. 7.5 in oratory) supports the interpretation that the spike is not marker-heaviness, but no significance test or window-count is reported. Please provide the underlying comparison or mark it as descriptive.","section":"§4.1, 'specificity check'"},{"comment":"There is a stray space in 'T wo pilots'. Throughout the manuscript there are similar ligature/formatting artifacts (e.g., 'diﬀicult', 'eﬀiciently', 'suﬀicient', 'T able'). These should be cleaned in the production version.","section":"§7.2, 'T wo pilots'"},{"comment":"The model mismatch is acknowledged, but Section 7.2's adapter is trained on Qwen2.5-7B-Instruct while the measurement and prompting demonstrations use Claude Haiku/Sonnet/Opus. This limits the direct comparability of the α-dial calibration in Table 2 to the human baselines in Table 1. Please state this explicitly in the main text rather than only in Limitations.","section":"§7.2 vs. §4.1"},{"comment":"The per-genre precision/recall rows for Speech, Narrative, and Social rest on very few true instances (seven or fewer). This is acknowledged in the text, but the table itself would benefit from an explicit 'n' column showing the number of true instances per row.","section":"Appendix B, Table B.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more of an essay with pilot experiments than a definitive empirical study. Its main risk is the detector precision asymmetry, which directly affects the EI ratios that the abstract highlights. The author's transparency is commendable, but the abstract overstates what the evidence supports. I would encourage the editor to send the paper back with the requested sensitivity analyses and significance tests, and to consider whether the journal's scope welcomes this format after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the headline: the paper identifies a real, recognizable LLM tic—the 'not X, but Y' upgrade frame—and treats it seriously as a rhetorical figure with a classical pedigree. It defines a clean metric (Epanorthosis Index = density over human baseline), ships reproducible artifacts (scripts, a LoRA adapter, a training recipe), and reports its own negative result (DPO failed) and its validation numbers. That transparency is the paper's best feature.\n\nThe problem is the load-bearing measurement. The central claim—models overshoot in oratory and undershoot in informal Q&A—rests on human baseline densities counted by a detector with precision 0.17 on human text versus 0.82 on model text. So human baselines are probably inflated, which deflates the EI ratios and could cancel the undershoot. The paper acknowledges this asymmetry but doesn't bound its effect. The headline p-values (0.03, 0.05) don't survive multiple-comparison correction. The three models are one family sharing an alignment pipeline, and the RLHF attribution is untested. The English detector isn't separately validated. The paper says all of this; the limitations section is exemplary. But the result remains a first signal, not a settled measurement.\n\nWhat's genuinely new and useful: the naming and framing of epanorthosis as a genre-miscalibrated stylistic bias, the EI as a calibration target, and the demonstration that a one-line instruction cuts the figure by 70% in Italian. The LoRA α-dial tables are suggestive, not definitive, but they establish a mechanism. The paper is honest that content fidelity is not yet measured.\n\nWho is this for? Computational stylistics, controllable generation, AI-text detection audiences. They'll get a good essay with pilot results and a clear research program. I'd bring it to reading group.\n\nRecommendation: send it to peer review. The phenomenon is important, the artifacts are reproducible, and the weaknesses are fixable—better detector validation on human text, cross-family and base-vs-aligned comparisons, proper multiplicity correction. I'd expect the referee to ask for exactly that.","headline":"A transparent, well-written pilot that names a real LLM tic and ships a LoRA dial, but the headline miscalibration numbers rest on a detector that is not valid on human text.","tokens_in":17739,"tokens_out":2771,"would_cite":true,"duration_ms":26706,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models overuse the rhetorical figure of epanorthosis — \"not X, but Y\" self-correction — because training data and preference tuning reward emphatic phrasing, and the excess can be measured and dialed back toward human rates.","keywords":["epanorthosis","rhetorical figures","large language models","RLHF","LoRA","controllable text generation","stylistic calibration","AI-text detection"],"falsifier":"Hand-annotate a human oratory sample with the same criteria used for model text and recompute the Epanorthosis Index with the corrected human density; if the corrected human rate is not below the model rate, the paper's signature overshoot result is a detector artifact.","tokens_in":16829,"feed_emoji":"🤖","tokens_out":7234,"duration_ms":66378,"temperature":0.7,"pith_summary":"Epanorthosis — the rhetorical move of negating a word to replace it with a stronger one, as in \"This is not a course. It is a journey of transformation\" — has become a signature tic of AI-generated writing. The paper argues that this overuse is a trained disposition: the training distribution and preference tuning reward confident, emphatic phrasing, while left-to-right generation merely surfaces the pattern unedited. It introduces a genre-relative Epanorthosis Index (density divided by the human rate) and reports two-sided miscalibration — overshoot in oratory, undershoot in informal Q&A. It then shows the excess is correctable: a one-line instruction cuts the figure by 70–72 percent in Italian, and a lightweight LoRA adapter can nearly eliminate it, with a scaling dial that lands on the human rate. The stakes are that we may begin to write like the machines, so calibrating style to human registers matters beyond aesthetics.","feed_headline":"AI overuses 'not X, but Y' because training rewards it","feed_subtitle":"A one-line prompt cuts the rhetorical overuse by up to 72 percent; a small adapter can dial it back to human rates.","key_machinery":"The load-bearing machinery is the Epanorthosis Index: for a genre g, EI_g = D_g/H_g, the model's emphatic-epanorthosis density (per ten thousand words, counting the \"not X, but Y\" family) divided by the human baseline density. An index of 1 is human-like calibration, above 1 overshoot, below 1 undershoot. The paper's theoretical bridge is Fontanier's classification of epanorthosis as a figure of thought, which licenses measuring the figure in a statistical model regardless of intention. On the mitigation side, the central mechanism is a low-rank adaptation (LoRA) adapter trained by supervised fine-tuning on de-emphasised paraphrases; scaling its contribution by a coefficient alpha at inferen","core_discovery":"The paper argues that when a model writes \"This is not a course. It is a journey of transformation,\" it is reproducing a figure Cicero and Quintilian catalogued, not inventing a tic: epanorthosis, the self-correction that negates a term to replace it with a stronger one. The central claim is that LLMs overuse this figure because it is a trained disposition — training corpora are rich in promotional prose and preference tuning rewards confident, emphatic phrasing — while left-to-right generation is only an amplifier. Using an Epanorthosis Index (model density divided by human density per ten thousand words), the paper measures two-sided miscalibration: models overshoot in oratory (33.5 vs 14.","pith_inferences":["Editorial inference: Because the detector's precision on human text is only 0.17, the human densities in the baseline table are likely inflated by false positives; recomputing the Epanorthosis Index with human-validated labels could reduce or even reverse the oratory overshoot.","Editorial inference: The attribution of the effect to RLHF is extrapolated, not directly tested; a base-model-versus-aligned comparison across several developer model families would confirm or refute it.","Editorial inference: If epanorthosis is a figure of thought that models rediscover statistically, other named figures (e.g., chiasmus, epiphora) should also show register-dependent, detectable frequencies, giving AI-text detection a whole-rhetoric fingerprint.","Editorial inference: The supervised adapter's near-total elimination of the figure at full strength may suppress legitimate correction in genres like argument; content-fidelity and appropriateness evaluation, prescribed but not yet reported, is needed before deployment."],"forward_implications":["The same training stages meant to make models helpful — instruction tuning and preference optimization — are what bake in the rhetorical excess, so fixing the data and reward signal would address the cause rather than the symptom.","A cheap one-line instruction cuts emphatic self-correction by 70–72 percent in Italian oratory and argument, so practical mitigation is available even without retraining.","A supervised LoRA adapter nearly removes the figure, and its scaling coefficient can settle output at the human rate rather than at zero — the paper's definition of proper calibration.","Measuring style genre-by-genre reveals two-sided errors that a raw frequency count hides, so evaluations of AI writing should use register-relative baselines.","If left unchecked, the miscalibrated style can transfer to human writing, which is the paper's stated reason to care."],"fun_headline_variants":["Why LLMs overuse 'not X, but Y' — it's trained, not innate","One-line prompt slashes AI's 'not X, but Y' by up to 72%","AI's 'not X, but Y' habit is a trained rhetoric flourish","Epanorthosis in LLMs: overuse stems from RLHF, not generation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The human baseline rates that define \"overuse\" rest on the paper's own detector, which it reports as far less reliable on human text than on model text; if those baselines are inflated, the central overshoot finding shrinks or disappears.","fun_headline_variants_meta":{"raw":{"variants":["Why LLMs overuse 'not X, but Y' — it's trained, not innate","One-line prompt slashes AI's 'not X, but Y' by up to 72%","AI's 'not X, but Y' habit is a trained rhetoric flourish","Epanorthosis in LLMs: overuse stems from RLHF, not generation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000342,"raw_usage":{"total_tokens":1785,"prompt_tokens":879,"completion_tokens":906,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":812}},"tokens_in":623,"tokens_out":906,"duration_ms":8659,"temperature":1.0,"reasoning_tokens":812,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:14:39.621668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hand-annotate a human oratory sample with the same criteria used for model text and recompute the Epanorthosis Index with the corrected human density; if the corrected human rate is not below the model rate, the paper's signature overshoot result is a detector artifact.","supporting_citations":[],"review_version":1}