{"id":"db2a4ce6-94b4-41d5-8862-89f1f15d5bcd","arxiv_id":"2506.03902","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using time-scaled harmonic regression on six RST discourse corpora, the paper reports that surprisal contours show periodic structure aligned with elementary discourse units, with first-order EDU-scaled sinusoids carrying the highest amplitude.","lead":"Across six languages, this paper finds that word-level surprisal (information content) rises and falls in patterns that track discourse units like clauses and paragraphs, and it introduces a time-scaled harmonic regression method to detect such periodic structure. The finding offers a more precise alternative to the uniform information density hypothesis, with a reusable statistical framework for testing how linguistic structure shapes information flow.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"EDU-scaled sinusoids can absorb a non-periodic within-EDU surprisal trend; without a relative-position baseline, the periodicity claim is indistinguishable from a boundary-ramp effect.","rationale":"The reader's weakest assumption is that surprisal estimates from one LLM per language may drive the fitted sinusoids via model artifacts. This is a real concern, but it is partly acknowledged in Sec. 2.2 and the boundary predictability captured by any reasonable LM is linguistically meaningful; even a perfect human language model would show lower surprisal before and higher surprisal after discourse boundaries. The more decisive weakness is internal to the argument: the paper does not compare time-scaled sinusoids against non-periodic within-unit position features. First-order EDU-scaled sinusoids are smooth functions of position within the EDU, so they can perfectly mimic a gradual boundary-to-boundary trend. Since the baseline only includes short boundary windows and document-level relative position, the reported advantage of EDU-scaled harmonics could be a re-description of the boundary effect rather than evidence for oscillation. This does not invalidate the boundary-surprisal findings, which are solid and replicated across languages, nor the harmonic-regression framework as a descriptive tool; but it leaves Hypothesis 2.3 underdetermined. The proposed polynomial baseline test would settle whether periodic components explain variance beyond any smooth within-unit trend. The reader's verdict of CONDITIONAL is therefore appropriate, and the concern here reinforces it rather than changing it.","tokens_in":26393,"tokens_out":4373,"duration_ms":47436,"concrete_test":"Refit the maximal model with the EDU-scaled sinusoidal family replaced by (or augmented with) linear, quadratic, and cubic terms of the token's relative position within its EDU, keeping all other baseline features and the same 10-fold CV and ANOVA procedure. If the polynomial within-EDU model matches the EDU-scaled model's MSE and ANOVA significance, the periodic interpretation is unsupported; if the sinusoids still yield a significant improvement over the polynomial model, the periodicity claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is treating one-cycle-per-EDU (k=1) sinusoids as evidence of periodicity. Under time scaling (Eq. 8), sin(2πt/U_t) and cos(2πt/U_t) are smooth functions anchored to EDU start and end; with fitted coefficients they can represent any smooth within-unit trend that starts high and ends low (or vice versa). The baseline in App. E controls only for boundary windows of 1, 2, and 4 tokens and document-level relative position, not for a gradual within-EDU positional trend. So the large amplitudes and significance of EDU-scaled first harmonics (Tab. 2) may simply re-express the documented boundary asymmetry (low surprisal before, high after; Sec. 6) as a smooth ramp over the EDU, rather than demonstrating oscillation. Hypothesis 2.3 asserts periodicity, not merely boundary-local modulation; without a non-periodic within-unit position baseline, the central claim cannot be distinguished from a much weaker 'surprisal follows relative position within discourse units' account.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes the Harmonic Surprisal (HS) hypothesis, according to which surprisal contours vary periodically and the dominant periods align with the boundaries of structural discourse units (EDUs, sentences, paragraphs). To test this, the authors introduce time-scaled harmonic regression, in which sine/cosine predictors for each token are scaled by the length of the token's containing structural unit, and compare such models against a linear baseline that includes boundary-window flags, previous surprisal, token length, and document-level relative position. Using LLM-derived surprisal estimates on RST-annotated corpora in six languages, they report that EDU-scaled sinusoids yield the lowest cross-validated MSE in five of six languages and that low-order EDU harmonics have the largest amplitudes. The paper also documents a systematic boundary profile (low surprisal before, high after boundaries) and includes a permutation control.","tokens_in":26514,"tokens_out":8715,"duration_ms":85265,"significance":"If established, the HS hypothesis would give a concrete, falsifiable functional form to the link between global information contours and discourse structure, and the time-scaling regression would be a reusable tool for probing periodicity at other linguistic levels. The paper has clear strengths: six languages, cross-validated model comparison with L1 feature selection, Holm-corrected significance tests, a permutation control in App. G, and publicly available code. However, as detailed below, the current evidence does not yet distinguish the periodicity claim from a much weaker non-periodic within-unit position effect, so the central empirical contribution remains under-supported. The paper is well written and the computational framework is a useful methodological contribution regardless of the outcome of that control.","major_comments":[{"comment":"EDU-scaled sinusoids can absorb a non-periodic within-EDU trend; the missing baseline makes the periodicity claim indeterminate. For a token t in an EDU of length U_t, sin(2πt/U_t) and cos(2πt/U_t) are smooth functions of the relative position t/U_t, so a fitted linear combination is any smooth within-EDU ramp of the kind documented for boundaries in Sec. 6 and Table 12. The baseline in App. E includes boundary-window flags of width 1, 2, and 4 tokens, document-level relative position, token length, and previous surprisal, but not a continuous or discretized relative-position feature within the EDU, sentence, or paragraph. Therefore the significant MSE reductions in Table 1 and the dominant low-order EDU amplitudes in Table 2 do not discriminate between Hypothesis 2.3 and the weaker structured-context account that surprisal follows relative position within the unit. In addition, Eq. (8) imposes the period rather than estimating it, so the analysis tests the predictive value of a particular set of unit-anchored functions, not whether the contour is periodic at those periods. I request a control condition with non-periodic within-unit positional features (linear, quadratic, or spline terms in relative position) and, if possible, a free-period condition in which the sinusoid periods are estimated from the data; the main claim should be re-evaluated after those controls.","section":"§3.2, Eq. (8); §5, Tables 1–2; App. E"},{"comment":"The cross-linguistic evidence rests on a single LLM per language, and several of the selected models are instruction-tuned; without estimator-level controls, the documented boundary effects could be artifacts of those models' tokenization or positional biases. The permutation control in App. G destroys the temporal structure of the surprisal values and therefore cannot address this concern; it shows only that the harmonic fit is not an artifact of arbitrary permutations. Adding a second, architecturally different surprisal estimator (e.g., an LSTM, a non-instruction-tuned causal LM, or an n-gram model) for at least one corpus per language would substantially strengthen the claim that the pattern is a property of linguistic information contours rather than of a particular model family. If the paper is intended only as a claim about LLM-computed surprisal, this concern is less severe, but the abstract and Hypothesis 2.3 are framed in language-general terms.","section":"§2.2; §4.2; Tab. 6"}],"minor_comments":[{"comment":"The sentence 'To account for the multiple comparisons, we we the Holm–Bonferroni correction' contains a typo and should read 'we used the Holm–Bonferroni correction.'","section":"App. F"},{"comment":"The column header 'Variance' appears to contain standard deviations (e.g., mean 12.87 tokens per English EDU with value 9.13); if so, the header should be corrected to 'Standard deviation' to avoid misreading.","section":"Tables 4–5"},{"comment":"Amplitudes are reported as means over cross-validation folds without a measure of variability; since the table is used to support the claim that EDU harmonics are strongest, standard errors or per-fold ranges would be informative.","section":"Table 2"},{"comment":"The text says all EDU-scaled sinusoids that remain after feature selection are consistently significant across folds; this is true for the reported EDU rows, but the sentence immediately follows a table that also contains non-EDU rows with fold subscripts lower than 10, so a reader could misread the claim as applying to the whole table.","section":"§5.1, Table 2"},{"comment":"The phrase 'Many dominant frequencies align with discourse structure' overstates the design: the periods are imposed by time scaling rather than discovered, so a more precise formulation would be 'EDU-anchored sinusoids are predictive of surprisal contours.'","section":"Abstract and §5"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically solid and the missing within-unit position baseline is a fixable but load-bearing gap. If the authors can show that EDU-scaled sinusoids survive after conditioning on non-periodic within-unit position features, the paper would be a strong contribution; otherwise the conclusions should be weakened to claim only smooth within-EDU modulation, not periodicity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a careful empirical paper with a real methodological contribution, but its headline claim—periodicity aligned with discourse units—is not actually distinguished from a much weaker claim about within-unit positional trends. The boundary effects are real and well-documented; the interpretation of them as harmonic is the soft spot.\n\nThe genuinely new thing here is time-scaled harmonic regression. Scaling the sinusoid periods by structural unit length is a simple, reusable idea, and the paper applies it across six languages with cross-validated MSE comparisons, permutation controls, and Holm–Bonferroni corrections. The boundary-surprisal statistics (low before, high after) are solid and consistent across languages and boundary types. They also ship code and data, which I value. The reader's skepticism about single LLM surprisal proxies is a legitimate concern but the paper acknowledges it, and it doesn't undermine the boundary effects.\n\nThe load-bearing weakness is the missing baseline. The stress-test note is right: sin and cos scaled to the EDU can absorb any smooth within-unit trend, and the baseline only includes boundary windows and document-level relative position, not relative position within the EDU. So the large first-harmonic amplitudes and their significance may simply re-express the boundary ramp as a smooth function across the unit. That is a much weaker claim than the HS hypothesis as stated in Hypothesis 2.3. The paper needs a non-periodic within-unit positional baseline (e.g., linear or quadratic relative position, or a spline) to show that sinusoids add something beyond a ramp. Without that, the periodicity claim is indistinguishable from \"surprisal follows relative position within discourse units,\" which is already the structured context hypothesis.\n\nA second, minor statistical issue: the significance tests treat tokens as independent, but surprisal is autocorrelated. Residual autocorrelation could inflate the ANOVA significance. This is fixable with cluster-robust errors or block permutation, and I'd mention it in review.\n\nOverall: this is a serious paper worth engaging with. The method is useful, the boundary effects are solid, and the hypothesis is interesting. But the central claim needs a stronger test before I'd believe it. I'd send it to peer review, with the expectation that the authors add the missing baseline and temper the periodicity language if the evidence doesn't hold up.","headline":"Solid boundary-effects paper with a novel method, but the periodicity claim needs a within-unit position baseline.","tokens_in":622,"tokens_out":748,"would_cite":true,"duration_ms":28271,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that word-level surprisal in texts oscillates periodically, with the dominant periods matching the spans of elementary discourse units (EDUs).","keywords":["harmonic surprisal","uniform information density","surprisal contours","discourse structure","rhetorical structure theory","harmonic regression","time scaling","information rate"],"falsifier":"Fit the same EDU-scaled harmonic regression to surprisal contours of texts whose clause order has been shuffled or whose EDU boundaries have been replaced by same-length random spans; if the first-order EDU sinusoid keeps its amplitude and significance, the periodic structure is not specific to real discourse segmentation. Equivalently, recompute the analysis with several different LLM families and check whether the EDU harmonic persists across models with different architectures and training data.","tokens_in":1495,"feed_emoji":"📈","tokens_out":1830,"duration_ms":56020,"temperature":0.7,"pith_summary":"The paper claims that the rate at which words convey information, measured as surprisal, does not merely drift around a mean but oscillates on regular cycles whose dominant lengths are set by discourse structure. It introduces the Harmonic Surprisal hypothesis: surprisal within a document varies periodically, with periods corresponding to the boundaries of structural units, most strongly elementary discourse units. Using harmonic regression with a new time-scaling step, the authors fit sinusoids scaled to document, paragraph, sentence, and EDU lengths in six languages and find that first-order EDU-scaled sinusoids are consistently significant predictors with the largest amplitudes. The result matters because it refines the uniform information density hypothesis: rather than flattening information globally, speakers modulate it in a structure-aligned rhythm, connecting information theory to discourse organization.","feed_headline":"Surprisal oscillates with discourse-unit spans in six languages","feed_subtitle":"A harmonic-regression analysis finds that word-level information rises and falls on the span of elementary discourse units.","key_machinery":"Harmonic regression is a linear regression whose predictors are sine and cosine components at integer multiples of a fundamental frequency, allowing a signal to be approximated as a sum of periodic trends. Time scaling modifies the period of those sinusoids by replacing absolute time with position normalized by the length of the structural unit containing the token, so that a first-order sinusoid completes exactly one oscillation per unit span. This machinery lets the authors test whether observed periodicity aligns with specific annotated structures: an EDU-scaled first-order sinusoid models one full cycle per elementary discourse unit, and its amplitude measures how strongly that particular period shapes surprisal. Feature selection uses L1 regularization, with significance assessed through one-way ANOVA and paired t-tests against a boundary-feature baseline.","core_discovery":"The central claim is Hypothesis 2.3: the components of a document's surprisal contour vary periodically, with periods that correspond to the boundaries of structural units within that contour. Operationally, the authors fit harmonic regression models whose sine and cosine predictors are time-scaled by the token length of the containing document, paragraph, sentence, or EDU, and evaluate them with ten-fold cross-validation and ANOVA significance tests. In the maximal model, low-order EDU-scaled sinusoids show the highest amplitudes across languages, they persist through feature selection in all folds, and they are the only structure type whose retained sinusoids are consistently significant everywhere, while surprisal reliably decreases before discourse boundaries and increases after them. The authors take this as evidence that discourse constituents, especially elementary discourse units, are the periods that organize information flow in text.","pith_inferences":["The framework predicts a directly testable extension: applying time-scaled harmonic regression to syntactic constituents, multi-word expressions, or words should yield high-amplitude sinusoids at those finer periods, revealing a hierarchy of nested information cycles.","The method could be applied to dialogue turns, prosodic units in speech, or explicitly structured text such as code, and if analogous periodicity appears there, the harmonic-surprisal claim would generalize beyond written expository discourse.","A null test with randomly assigned same-length unit boundaries would isolate whether the EDU-aligned periodicity is genuine discourse alignment or merely an artifact of unit-length scaling; the paper's permutation checks rearrange surprisal values but not the structural segmentation itself.","Because surprisal estimates come from one LLM per language, comparing the fitted EDU harmonic across models with different architectures, context lengths, and training objectives would measure how much of the signal is a property of language versus a property of the estimator."],"forward_implications":["If the Harmonic Surprisal hypothesis holds, uniform information density is not contradicted but refined: surprisal regresses toward a mean while oscillating periodically around it, so global uniformity and structure-aligned modulation coexist.","Position within an EDU, not just distance from a boundary, carries predictive information about surprisal, making discourse segmentation a useful feature for surprisal and reading-time modeling.","The time-scaled harmonic regression framework is a general tool: any annotated structural unit, at any granularity, can be tested as a candidate period for information-rate fluctuations.","The cross-linguistic consistency across English, Spanish, German, Dutch, Basque, and Brazilian Portuguese suggests the periodic pattern is a general property of written discourse rather than an artifact of one language or corpus.","Surprisal's reliable dip before and rise after boundaries implies that speakers modulate information rate to align with processing costs that accumulate toward unit ends and reset at unit starts."],"supporting_citations":[{"why":"Introduces the Structured Context hypothesis and the discourse-level positional predictors that this work refines into periodic sinusoids.","marker":"Tsipidi et al. (2024)"},{"why":"Defines Rhetorical Structure Theory and elementary discourse units, the structural periods at the center of the analysis.","marker":"Mann and Thompson (1988)"},{"why":"Establishes global entropy-rate constancy as a discourse-level uniform information density claim, the baseline view this paper refines.","marker":"Genzel and Charniak (2002)"},{"why":"Formulates uniform information density as speakers' pressure to balance information density, the hypothesis that Harmonic Surprisal elaborates.","marker":"Levy and Jaeger (2006)"},{"why":"Provides evidence that UID operates as regression toward a global mean, the backdrop against which periodic fluctuation is detected.","marker":"Meister et al. (2021)"},{"why":"Supplies the English RST Discourse Treebank, one of the six corpora used to estimate surprisal and EDU spans.","marker":"Carlson et al. (2001)"}],"fun_headline_variants":["Language's information flow pulses with discourse boundaries","Surprisal rhythm aligns with elementary discourse units","Information rate oscillates on discourse-unit spans","Six languages show periodic surprisal tied to discourse spans","Discourse units set the beat for word-surprisal waves"],"cache_read_input_tokens":29312,"weakest_assumption_plain":"The load-bearing premise is that the transformer language models used to compute surprisal are faithful proxies for the unknown human language model; if their position-dependent predictions or boundary artifacts create the fitted sinusoids, then the periodic alignment is a property of those models rather than of language itself.","fun_headline_variants_meta":{"raw":{"variants":["Language's information flow pulses with discourse boundaries","Surprisal rhythm aligns with elementary discourse units","Information rate oscillates on discourse-unit spans","Six languages show periodic surprisal tied to discourse spans","Discourse units set the beat for word-surprisal waves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000334,"raw_usage":{"total_tokens":1823,"prompt_tokens":882,"completion_tokens":941,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":875}},"tokens_in":498,"tokens_out":941,"duration_ms":7757,"temperature":1.0,"reasoning_tokens":875,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:53:37.087209+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the same EDU-scaled harmonic regression to surprisal contours of texts whose clause order has been shuffled or whose EDU boundaries have been replaced by same-length random spans; if the first-order EDU sinusoid keeps its amplitude and significance, the periodic structure is not specific to real discourse segmentation. Equivalently, recompute the analysis with several different LLM families and check whether the EDU harmonic persists across models with different architectures and training data.","supporting_citations":[{"cited_title":"Florian Jaeger","cited_arxiv_id":null,"evidence_quote":"Formulates uniform information density as speakers' pressure to balance information density, the hypothesis that Harmonic Surprisal elaborates."}],"review_version":1}