{"id":"70a3e184-e032-4c9f-835a-3dea16fbf975","arxiv_id":"2605.28826","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Stylistic collapse in LLMs is largely pre-alignment and scale-independent; strong entropy regularization reduces it far more than weak smoothing or RLHF.","lead":"Modern LLMs massively overuse headers, lists, and stock phrases while nearly erasing complex punctuation, and this “AI voice” shows up even before RLHF. A small model trained with strong entropy control can look more human on these probes than frontier APIs, with implications for detection and future training data.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"The λ=5.0 “beats frontier” claim rests on a high-perplexity small model whose divergence metric is not shown to preserve generation quality.","rationale":"The multi-model frequency survey (selective amplification of headers/lists/discourse markers and suppression of complex punctuation) is solid and reproducible from the reported protocol; the reader correctly treats that as the durable contribution. The softest load-bearing step for the stronger causal story—that objectives matter more than scale and that only strong entropy regularization fixes the collapse—is exactly the one the reader flags: treating a 410M from-scratch run at perplexity ~786 as a fair comparator that “beats” frontier and larger open models on the authors’ divergence metric. Non-significant base–instruct p-values and proxy baselines for closed models are secondary weaknesses; they do not by themselves overturn the measurement paper. Because the quality of the λ=5.0 generations is unvalidated, the verdict should stay CONDITIONAL (valuable if claims are narrowed to open-model frequency shifts and exploratory regularization), not move to ACCEPT or REJECT. The proposed preference/fluency check is the single experiment that would decide whether the scale-disadvantage claim survives.","tokens_in":22607,"tokens_out":722,"duration_ms":9072,"concrete_test":"Generate 200 matched samples from pythia-410m-λ=5.0 and from an unregularized or lightly regularized checkpoint (and, if possible, a same-prompt frontier baseline) under the paper’s exact protocol; run blind pairwise preference or fluency ratings plus a simple task check (e.g., topical coherence / instruction following). If λ=5.0 loses on quality while only winning on mean AR, the “outperforms frontier despite scale” claim does not hold as stated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that training objectives dominate scale, and that strong entropy control is the fix, depends on Tables 4–6 and A.6: pythia-410m trained from scratch with λ=5.0 reaches distance-from-1.0 of 0.22 and is said to be 96.7–98.2% better than frontier APIs (and better than 1–8B open models) despite 200–1000× scale disadvantage. That comparison is load-bearing for the title and abstract. Yet the same table reports perplexity 786.5 (vs 48.4 at λ=0), while feature-level Table 8 shows non-monotonic restoration: some suppressed features reappear but several structural/discourse markers still rise, and many probes remain at zero in both models. Without evidence that the λ=5.0 outputs remain fluent, on-topic, and preference-competitive under the same 15 formal prompts, the low mean AR can be an artifact of undertraining or quality collapse rather than successful distributional control. The paper asserts perplexity is “decoupled from generation quality” but supplies no human ratings, win-rates, or task metrics to support that assertion for the model that underpins the scale-vs-objectives conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that modern LLMs systematically reallocate stylistic probability mass relative to human training corpora: across 17 models and 24 string/regex probes, discourse and structural features are amplified by mean factors of roughly 1,949–16,853% (peaks to ~209,675%) while complex punctuation is suppressed to ~3.2–23.2% of baseline. Matched base vs instruction-tuned pairs show statistically indistinguishable divergence (p > 0.25), so the authors conclude the effect is not caused or worsened by RLHF and is instead driven by deployment context shift plus low-entropy “absorbing stylistic states.” They further claim that only strong entropy regularization (λ=5.0) during from-scratch Pythia-410M pretraining reduces divergence (40.5% improvement; 96.7–98.2% closer to AR=1 than frontier APIs) while weak λ=1.0 exacerbates collapse, establishing a “control strength principle” that training objectives matter more than scale.","tokens_in":22953,"tokens_out":1334,"duration_ms":18326,"significance":"If the survey results hold, the paper supplies a concrete, reproducible stylometric characterization of the “AI voice” across open and commercial systems, with clear implications for AI detection, recursive training-data contamination, and long-term linguistic drift. The base–instruct comparison is a useful corrective to narratives that pin formulaic style solely on RLHF. The entropy-regularization ablations are a genuine attempt at a training-time fix and report useful diversity metrics (distinct-n, repetition, vocab diversity). Strengths include a fully specified 24-feature taxonomy, large generation samples (1,000 per model), open experimental protocol in the appendix, and explicit limitations. The work is significant as an empirical audit even if the mechanistic story and the “beats frontier despite scale” claim require tightening.","major_comments":[{"comment":"Table 3 and the abstract claim that alignment “does not exacerbate” stylistic divergence because all four base–instruct pairs have p > 0.25. Three of four pairs show large numerical mean-AR increases (+1,194%, +174%, +138%); non-significance is not evidence of equivalence, especially with n=4 pairs and high cross-feature variance. The load-bearing claim that the AI voice is “upstream of alignment” and “alignment-independent” needs equivalence tests (e.g., TOST), confidence intervals on the change, or a clearer statement that the study is underpowered to detect moderate exacerbation rather than that exacerbation is ruled out.","section":null},{"comment":"Tables 4–6 and A.6 underpin the title claim that training objectives dominate scale: pythia-410m-λ=5.0 reports distance-from-1.0 of 0.22 and is said to be 96.7–98.2% better than frontier APIs. The same tables report perplexity 786.5 (vs 48.4 at λ=0). Table 8 shows non-monotonic feature restoration and many probes still at zero in both models. Without human ratings, preference win-rates, or task metrics under the same 15 prompts, the low mean AR may reflect undertraining or quality collapse rather than successful distributional control. The assertion that “perplexity is decoupled from generation quality” is currently unsupported for the model that carries the scale-vs-objectives conclusion; either add quality evidence or substantially qualify the frontier comparison.","section":null},{"comment":"Sections 4.1–4.3 and A.2–A.4: amplification ratios for commercial APIs (and for models without open corpora) use Pile/Dolma human baselines and 1,000 generations from 15 exclusively formal expository English prompts at temperature 0.7. That prompt set itself selects the formal-expository slice the theory calls “context shift,” so measured AR may partly be an artifact of the evaluation distribution rather than a pure property of the models. For closed models the true PC(f) is unknown. The survey claim remains directionally credible for open models with matched corpora, but the universality and magnitude claims for frontier systems need either multi-register prompts or explicit sensitivity analysis to baseline choice.","section":null}],"minor_comments":[{"comment":"Abstract vs body: abstract says “17 models”; main tables and A.4 list 13 evaluated models plus four trained Pythia variants—clarify the count consistently.","section":null},{"comment":"Mean AR aggregation treats all 24 features equally; Appendix A.7.1 shows top-10 features drive rank correlation. State whether mean AR is unweighted and whether results are robust to top-k or category-weighted aggregation.","section":null},{"comment":"Figure 1 caption mentions “OLMo-2-Instruct” while tables use “OLMo-1B-Instruct”; align naming.","section":null},{"comment":"Section 3.2 presents “absorbing stylistic states” as mechanistic explanation; Limitations correctly notes lack of causal verification—consider moving stronger causal language to future work throughout the Theory section.","section":null},{"comment":"Table 2 lists “It’s worth noting” and sentence-initial “Certainly”/“Absolutely” at 0.0 AR; clarify whether these are true zeros or below detection, and how zeros enter the mean AR.","section":null},{"comment":"δ=0.1 and λ∈{0,0.1,1,5} are free parameters; a short sensitivity note (or pointer to ablations) would help readers assess robustness of the control-strength principle.","section":null}],"recommendation":"major_revision","confidential_remarks":"The empirical audit of open models is publishable and useful; the overclaim is concentrated in (i) interpreting p>0.25 as “does not exacerbate” and (ii) the λ=5.0 vs frontier quality-free comparison that drives the title. If the authors reframe those two points and add even modest quality checks or multi-register prompts, this becomes a solid empirical contribution. Fit for a methods/empirical track is good; novelty relative to prior stylometry and diversity-collapse work should be checked but is not a reject reason on its own."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The durable piece here is the 24-probe survey across open and frontier models. Selective amplification of headers, lists, and discourse markers (\"delve into\", \"in conclusion\") and suppression of semicolons/em-dashes is tabulated carefully, with baselines from Pile/Dolma and fixed generation settings. That map is new enough relative to prior AI–human stylometry and DetectGPT-style work, and it will be useful to people who care about detectors and data contamination.\n\nThe base–instruct comparison is the second real contribution. Four matched pairs show no statistically significant change in mean AR (all p>0.25). That undercuts the common story that RLHF is the main source of formulaic style. The paper is right that the effect is already present in base models; the non-significance is not equivalence, and some pairs show large point estimates, but the direction is clear enough that the claim “does not worsen under RLHF” is fair if worded carefully.\n\nSoft spots are real but concentrated. The load-bearing title claim—that training objectives dominate scale and that strong entropy control is the fix—rests on from-scratch Pythia-410M runs. λ=5.0 reaches distance-from-1.0 of 0.22 and is said to beat frontier APIs by 96–98%, yet perplexity jumps to 786. Feature-level results are non-monotonic; many probes stay at zero. The paper asserts perplexity is decoupled from quality but gives no human ratings, win rates, or task metrics under the same prompts. Closed-model baselines are proxies, and the 15 formal English prompts make “context shift” partly by construction. Mechanisms (absorbing stylistic states) are post-hoc. None of this sinks the measurement paper; it does mean the scale-vs-objectives conclusion needs heavy narrowing.\n\nMath and citation pattern look ordinary and honest: amplification ratios, Bonferroni notes, standard diversity metrics, and the right priors (Holtzman, Mitchell, Kirk, Pereyra, Zhang). Code is promised. For readers who want a concrete feature inventory of the AI voice and evidence that alignment is not the root cause, this is worth time. I would send it to peer review with instructions to force the claims back to what the open-model frequencies and the non-monotonic λ finding actually support.","headline":"Solid multi-model stylometry of the AI voice; the base–instruct non-effect is useful; the λ=5.0 “beats frontier” claim is oversold on a high-perplexity 410M model.","tokens_in":23556,"tokens_out":578,"would_cite":true,"duration_ms":7406,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Training objectives, not model scale or RLHF, drive the extreme stylistic redistribution that makes LLM text sound like AI.","keywords":["stylistic divergence","entropy regularization","context shift","absorbing stylistic states","RLHF independence","amplification ratio","AI text detection","mode collapse"],"falsifier":"If matched base and instruction-tuned pairs of the same architecture, evaluated on the same 24 probes and the same prompt set, produced statistically significant differences in mean amplification (p≪0.25), or if strong entropy regularization at larger scale failed to reduce divergence relative to unregularized controls, the central claim would be falsified.","tokens_in":23438,"feed_emoji":"📉","tokens_out":711,"duration_ms":6314,"temperature":0.7,"pith_summary":"This paper argues that modern language models systematically reallocate probability mass over linguistic features, amplifying discourse markers and structural scaffolding by thousands of percent relative to human training corpora while suppressing complex punctuation. Across seventeen models spanning 410M to frontier scale and twenty-four probes, the same selective pattern appears. Matched base and instruction-tuned pairs show statistically indistinguishable divergence, so RLHF and instruction tuning are not the primary drivers. The author attributes the effect to deployment context shift into formal expository regimes plus self-reinforcing low-entropy “absorbing” stylistic states during autoregressive generation. Weak entropy regularization makes collapse worse; only strong regularization (λ=5.0) reduces divergence substantially and can outperform far larger frontier models on distributional naturalness. The result matters because the redistribution is invisible to ordinary quality metrics yet detectable by simple probes, with consequences for AI detection, future training data, and how human writing norms may evolve.","feed_headline":"Scale and RLHF do not fix AI writing style; training does","feed_subtitle":"Strong entropy regularization beats frontier models on stylistic naturalness despite 200–1000× size gap","key_machinery":"Amplification ratio AR_M(f) = P_M(f)/P_C(f) across a 24-feature taxonomy, together with the control-strength principle that entropy regularization L_CE − λH(P_θ) only mitigates collapse when λ is large enough (λ=5.0 works; λ=1.0 worsens it).","core_discovery":"Instruction-tuned and frontier LLMs systematically reallocate stylistic probability mass—amplifying discourse and structural features by mean factors of roughly 1,949–16,853 percent (peaks to ~209,675 percent) while suppressing complex punctuation to 3.2–23.2 percent of corpus baselines—and this divergence is statistically indistinguishable across matched base versus instruction-tuned pairs (p>0.25). Therefore the “AI voice” is not primarily created or worsened by RLHF; only sufficiently strong entropy regularization, not weak smoothing or scale, substantially reduces it.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Training objectives collapse LLM style far more than scale or RLHF","Strong entropy control beats frontier scale on stylistic naturalness","Instruction tuning reallocates stylistic mass; RLHF adds little","Why alignment objectives reshape language distributions more than size","Weak regularization worsens collapse; strong control outperforms giants"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That amplification ratios against Pile/Dolma baselines, measured on a thousand generations from fifteen formal expository English prompts, are a valid proxy for real deployment context shift—and that a from-scratch 410M model with high perplexity is still a fair test of whether strong regularization beats scale.","fun_headline_variants_meta":{"raw":{"variants":["Training objectives collapse LLM style far more than scale or RLHF","Strong entropy control beats frontier scale on stylistic naturalness","Instruction tuning reallocates stylistic mass; RLHF adds little","Why alignment objectives reshape language distributions more than size","Weak regularization worsens collapse; strong control outperforms giants"]},"model":"grok-4.5","effort":"low","cost_usd":0.006486,"raw_usage":{"total_tokens":1765,"prompt_tokens":930,"num_sources_used":0,"completion_tokens":81,"cost_in_usd_ticks":64860000,"prompt_tokens_details":{"text_tokens":930,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":754,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":930,"tokens_out":81,"duration_ms":6330,"temperature":1.0,"reasoning_tokens":754,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T08:56:12.600853+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If matched base and instruction-tuned pairs of the same architecture, evaluated on the same 24 probes and the same prompt set, produced statistically significant differences in mean amplification (p≪0.25), or if strong entropy regularization at larger scale failed to reduce divergence relative to unregularized controls, the central claim would be falsified.","supporting_citations":[],"review_version":1}