{"id":"7aad9a53-df8e-46ff-a49f-da4d9de584ae","arxiv_id":"2608.03507","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"ChronoLens uses feature-aligned crosscoders to show that historical language change has comparable magnitude across linguistic levels within a language, but divergent timing and direction across five parliamentary languages.","lead":"This paper introduces ChronoLens, a framework pairing multilingual language models with sparse crosscoder features to measure how morphology, syntax, semantics, and pragmatics change across five parliamentary languages from 1803 to 2026. It finds that within a language these levels change by similar amounts, while languages differ markedly in when, how far, and in which direction they change.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pragmatics trajectories are unsupported by the paper's own feature-count criterion, contradicting the abstract's four-level claim.","rationale":"The reader's weakest assumption (corpus/register comparability) is a genuine external threat to all empirical results, and I partially agree. But the most load-bearing, checkable weakness is internal: the supplementary material explicitly states that pragmatics never reaches the minimum feature count and that filtered trajectories cover only morphology, syntax, and semantics, while the main text and abstract nonetheless report pragmatics trajectories and include pragmatics in the central comparable-magnitude finding. This is not a disagreement with consensus; it is a contradiction between the paper's own limitations appendix and its headline claim. A single computational check — recomputing Figures 2–4 under the stated minimum-feature criterion — would settle whether the four-level claim survives. The paper still has substantial value in the corpus, the crosscoder methodology, and the morphology/syntax/semantics results, so I would keep the reader's conditional verdict rather than reject it outright.","tokens_in":27261,"tokens_out":7076,"duration_ms":76917,"concrete_test":"Re-run the magnitude and direction analyses (Figures 2–4) using the paper's own minimum-feature rule (C.6: at least 8 assigned features per language–level–decade cell; also report cells with fewer). If no pragmatics cell meets the criterion, remove pragmatics from the trajectory claims and revise the abstract and Figure 3; if some cells do meet it, report per-cell feature counts and trajectory stability for pragmatics to show the values are not driven by a handful of features.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central four-level claim ('morphology, syntax, semantics, and pragmatics generally change by comparable amounts') depends on crosscoder trajectories for pragmatics, but the paper's own threshold analysis says those trajectories do not exist under its stated criterion. Appendix C.4 reports that pragmatics features are only 0.4–0.6% of assigned features at every threshold setting and 'never reaches that count' — the minimum-feature criterion — 'so the filtered trajectories cover morphology, syntax and semantics only.' Yet Figure 3 and Section 5.2 report pragmatics magnitudes (0.29–0.50) and the abstract includes pragmatics in the comparable-amounts finding. Either the Figure 3 pragmatics values are computed from cells with fewer than the C.6 minimum of 8 assigned features, making them unstable and not comparable to the other levels, or the appendix refers to a different analysis and the main text must reconcile the two. In either case, the headline empirical finding as stated is not supported by the released methodology.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ChronoLens, a framework combining frozen multilingual language models with feature-aligned crosscoders to measure historical language change across five parliamentary traditions (English, German, Italian, Polish, Turkish) and four linguistic levels (morphology, syntax, semantics, pragmatics). The corpus spans 1803–2026 (44.98M documents, ~17.2B tokens). The core claims are: (i) the crosscoder's sparse representations agree more strongly with linguistic statistics than dense embeddings or a pooled SAE (ρ=0.72 vs 0.29/0.28; Table 2); (ii) within a language, the four linguistic levels change by comparable magnitudes, while languages differ in timing, magnitude, and direction (Figures 2–4); and (iii) magnitude and direction of language change are separable properties. The framework trains crosscoders without linguistic supervision, then assigns features to linguistic levels via post-hoc probe ablations, and computes magnitude/direction metrics in the shared feature space.","tokens_in":27517,"tokens_out":4942,"duration_ms":55649,"significance":"If the claims hold, ChronoLens offers a genuinely unifying representation for cross-linguistic, cross-level diachronic comparison, addressing a recognized gap in computational historical linguistics. The paper's strengths include its very large harmonized corpus, the use of four diverse multilingual backbones, explicit condition-shuffled nulls and permuted-label controls, and a transparent threshold-sensitivity analysis. The empirical finding that magnitude and direction are separable is important and falsifiable. However, the central validation and the four-level empirical claim are currently undermined by two problems: the 'independent' linguistic indicators are derived from the same annotations used to train the assignment probes and select layers, and the paper's own Appendix C.4 states that pragmatics trajectories never meet the minimum feature-count criterion, directly contradicting the abstract's four-level claim. The contribution is potentially significant, but the evidence as presented does not yet support the headline conclusions.","major_comments":[{"comment":"The threshold-sensitivity analysis reports that pragmatics features are only 0.4–0.6% of assigned features at every threshold setting and 'never reaches that count', so 'the filtered trajectories cover morphology, syntax and semantics only.' This directly contradicts the abstract's statement that morphology, syntax, semantics, and pragmatics 'generally change by comparable amounts,' and Figure 3's pragmatics magnitudes (e.g., 0.29–0.50) are computed from cells that, by C.6's own rule, fall below the minimum of eight available features. Either the figure reports unstable low-coverage cells, or the appendix refers to a different analysis; the manuscript must reconcile this and either remove pragmatics from the four-level claim or demonstrate stable trajectories with adequate feature coverage.","section":"Appendix C.4; §5.2, Fig. 3; Abstract"},{"comment":"The headline validation metric 'linguistic agreement' is described as correlating representational displacement with 'direct changes in independently measured linguistic indicators.' However, Appendix D states that the observable measures are 'derived from the same sentence-level annotations used for feature attribution,' and §4.1 and Appendix C.2 show that both layer selection and feature-to-level assignment use probes trained on those same annotations. Features are therefore selected because they predict the labels against which agreement is measured, so ρ=0.72 is partly by construction. The comparison with embeddings/SAE remains informative, but the claim of independent linguistic grounding requires a validation source held out from probe training and layer selection, or an explicit demonstration that the agreement survives such a split.","section":"§4.2, Table 2; Appendix D; §4.1/C.2"},{"comment":"The empirical trajectories are interpreted as language change, but corpus comparability is not established. Coverage differs markedly by language (English begins 1803, Turkish 1950), German additionally includes German Digital Library newspapers, all languages contain manifesto records in different proportions, and OCR noise is concentrated in the earliest material. Length- and policy-frame matching controls for some sentence-level confounds, but not for source composition, register, or OCR degradation. The Limitations admit that 'residual OCR errors... may still resemble linguistic change despite our filtering.' Since the central findings concern magnitudes and directions over time, the paper should include per-source or per-register stability analyses (e.g., parliamentary speeches only, common periods) and show that the trajectories are not driven by OCR or source shifts.","section":"§3, §4.1; Limitations"},{"comment":"Even setting aside the feature-count issue, the pragmatics attribution rests on weak probes: speech-act selectivity ranges 0.05–0.21 and the majority class covers 0.84–0.96 of held-out examples, and the paper itself cautions that lower-selectivity tasks should be interpreted with caution. Reporting pragmatics magnitudes as comparable to other levels (Figure 3) is not supported by the released probe diagnostics. At minimum, the pragmatics rows should be flagged as low-confidence and the cross-level 'comparable amounts' claim should be restricted to the levels with reliable attribution.","section":"Appendix C.5, Table 9; §5.2, Fig. 3"}],"minor_comments":[{"comment":"The main text says 'We examine 23 measures grouped under morphology, syntax, semantics, and pragmatics' and '37 of the 90 fitted trends remain significant.' Appendix D reports 18 observable measures, and 18×5=90; 23 measures would give 115 tests. The mismatch should be corrected.","section":"§5.2 vs Appendix D"},{"comment":"The variable name 'parend' appears to be a typo; the surrounding text and Eq. (15) define a cosine between displacement vectors. Please rename for clarity.","section":"Appendix C.6"},{"comment":"The row 'Linguistic agreement, ρ↑' does not state whether ρ is Spearman or Pearson. The text calls it 'mean Spearman correlation' in §4.2 but Table 2 is ambiguous. Please state the estimator in the table caption.","section":"Table 2"},{"comment":"Cells with fewer than the minimum number of features are retained but marked as low coverage according to C.6, yet Figure 3 does not mark any low-coverage cells. Adding a visual indicator for low-coverage cells would make the reliability of the pragmatics rows transparent.","section":"Fig. 3; Appendix C.6"},{"comment":"The text refers to 'the pre-registered target × decade convergence contrast' but no pre-registration identifier or repository link is provided. Please add the registration reference or remove the term.","section":"Appendix C.4"},{"comment":"Typo: 'ross-language agreement' should be 'cross-language agreement.'","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"This is a high-risk, potentially high-reward submission. The unifying framework and corpus are valuable, and the authors have included unusually transparent diagnostics, but the current validation is circular and the pragmatics claim is contradicted by their own Appendix C.4. I believe major revision is appropriate: if the authors can (a) supply a genuinely independent validation of the linguistic-agreement metric, and (b) either remove pragmatics from the four-level claim or produce stable pragmatics trajectories with sufficient feature coverage, the paper could become suitable for publication. If these points are not addressable within the manuscript's scope, the central empirical claims would need to be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe one thing to know: this paper is worth reading, but do not trust the abstract's four-level claim until the pragmatics issue is resolved. The framework—feature-aligned crosscoders across five languages and 200 years, with magnitude separated from direction—is genuinely new, and the corpus effort is substantial. The rho=0.72 validation is a real result but weaker than it looks because the same annotations drive feature attribution and the comparison target, and layer selection uses the same probes.\n\nWhat is actually good: the corpus construction is careful (22 sources, quality control, sampling caps, matched strata). The use of crosscoders for diachronic multilingual comparison is a real step beyond pooled SAEs. The separation of magnitude and direction, and the finding that similar magnitudes can hide different directions, is a useful contribution. The sensitivity analyses (thresholds, condition-shuffled null, backbone agreement) are more thorough than typical for this area.\n\nThe soft spots: first, the internal contradiction. Appendix C.4 explicitly says pragmatics features are only 0.4–0.6% of assigned features and 'never reaches that count' (the minimum-feature criterion), so the filtered trajectories cover morphology, syntax and semantics only. Yet Figure 3 and Section 5.2 report pragmatics magnitudes and the abstract includes pragmatics in the comparable-amounts finding. That is a load-bearing inconsistency: either Figure 3 is computed from cells with fewer than the minimum features, or the appendix refers to a different analysis. Either way, the paper's central empirical claim as stated is not supported by its own methodology. This needs to be fixed before publication.\n\nSecond, the validation circularity: the 'independent' linguistic indicators in Appendix D are derived from the same annotations used to train the probes that assign features to levels. The rho=0.72 is partly by construction. The paper does not provide an external benchmark or a proper held-out validation. Also, the 'comparable amounts' claim lacks formal testing—no significance test or effect size is reported for the between-level comparison.\n\nThird, corpus comparability: unequal historical coverage, different source mixes (German includes newspapers and manifestos), and OCR noise concentrated in early material. The matching on length and policy frame helps but does not fully control. The Limitations admit residual OCR errors may resemble language change. That is a genuine threat to all empirical results.\n\nWho for: computational historical linguists and interpretability researchers. It deserves a serious referee, but not acceptance as is. The pragmatics contradiction and the circularity need to be addressed, ideally with an external validation benchmark.\n\nRecommendation: send to peer review, with the understanding that the pragmatics claims must be reconciled with the appendix and the validation needs external grounding.","headline":"A genuinely useful framework for comparing diachronic change across languages and levels, but the pragmatics results contradict the paper's own feature-count criterion and the validation is partly circular.","tokens_in":27995,"tokens_out":1646,"would_cite":true,"duration_ms":17143,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single feature-aligned space measures language change across five languages and four linguistic levels at once, revealing that the levels shift by comparable amounts within a language while languages differ in timing, magnitude, and direc","keywords":["language change","diachronic linguistics","crosscoders","sparse autoencoders","multilingual language models","parliamentary corpora","cross-linguistic comparison","linguistic levels"],"falsifier":"Train the pipeline on a condition-shuffled null — tuples whose decade or language identities are randomly exchanged, a control the paper already runs at feature level — and apply the full magnitude-and-direction analysis to it. If the null reproduces the ρ ≈ 0.72 linguistic agreement or the German-versus-Turkish magnitude split, the signal is a method artifact, not a fact about time. A complementary check: apply the pipeline to a decades-spanning corpus of stable, repeatedly reprinted texts (legal or liturgical boilerplate) with a similar OCR profile; if it reports substantial change, the traj","tokens_in":27152,"feed_emoji":"🗣️","tokens_out":15984,"duration_ms":157301,"temperature":0.7,"pith_summary":"ChronoLens asks a question earlier computational work could not answer: do morphology, syntax, semantics, and pragmatics change together over time, or do they follow separate historical paths? The paper's claim is that they can be measured together — by encoding sentences from five parliamentary traditions with frozen multilingual language models and projecting them into one feature-aligned sparse space learned by crosscoders — and that when measured together, the four levels change by roughly comparable amounts within a language, while languages differ sharply in when, how far, and in which direction they change. The framework's representations agree with direct linguistic statistics far better than dense embeddings or a pooled sparse autoencoder (Spearman ρ = 0.72 versus 0.29 and 0.28), which the paper reads as evidence that its features track real linguistic variation rather than reconstruction artifacts. If the magnitude–direction separation is real, then single-language, single-level studies have been under-reporting the structure of historical change, because similar magnitudes can conceal very different trajectories.","feed_headline":"How far a language changes doesn't tell which way it drifts","feed_subtitle":"A 17-billion-token study of five parliaments separates change size from change direction.","key_machinery":"The load-bearing object is the crosscoder: a sparse autoencoder with one shared feature index across all conditions — languages in the cross-lingual setting, or historical periods within a language — but separate encoders and decoders per condition. A single sparse feature vector f is inferred jointly from a tuple of matched sentences (ReLU plus BatchTopK), so feature j carries the same identity in every condition, while each condition's decoder lets that feature contribute with different strength there. This shared-index design is what makes cross-lingual and cross-temporal comparison possible, since independently trained sparse dictionaries have no guaranteed feature correspondence. Lingui","core_discovery":"On the paper's own terms, the discovery is that a crosscoder — a sparse autoencoder whose feature indices are shared across conditions but which gives each language or period its own decoder — yields representations whose decade-to-decade movement tracks independently measured linguistic statistics more closely than dense embeddings or a pooled sparse autoencoder do (mean Spearman ρ = 0.72 versus 0.29 and 0.28; linguistic specificity 0.89 versus 0.74 and 0.73). Applied to 44.98 million parliamentary documents and about 17.2 billion tokens across English, German, Italian, Polish, and Turkish (1803–2026), this representation shows that within a language the four linguistic levels change by com","pith_inferences":["The paper's own threshold-sensitivity analysis reports that features assigned to pragmatics are scarce (0.4–0.6% of assigned features at every threshold setting), so the claim that pragmatics changes in step with the other levels may rest on a thinner feature base; a dedicated pragmatic feature bank would test whether the high cross-language pragmatics agreement (cosine 0.92) survives a denser inv","The hard one-level-per-feature assignment is a measurement choice, not an established fact about language: genuinely cross-level features, which the paper acknowledges exist, are forced into a dominant level, so phenomena like grammaticalization that plausibly run through morphology, syntax, and semantics at once may be under-counted in the reported trajectories.","The headline direction metric compares net displacement between first and last decades, so pairs like German and Turkish could reach similar bearings through different intermediate paths; the paper's secondary stepwise-alignment measure is the stricter test of whether two languages truly follow shared trajectories.","Because directionally similar pairs (English–Italian, German–Turkish) are not genealogically close, the framework offers a quantitative handle on contact and institutional-convergence hypotheses — for instance, whether shared European political discourse is steering parliamentary registers toward a common pragmatic style, which the pragmatics results make plausible but do not causally establish."],"forward_implications":["If the ρ = 0.72 agreement is real, feature-aligned sparse representations are the appropriate substrate for diachronic comparison, and the pooled autoencoder's failure shows that a shared sparse dictionary alone, without condition-specific decoders, is not enough.","Within each language, morphology, syntax, semantics, and pragmatics change by comparable amounts, so single-level histories (the common practice) systematically miss that the other levels are moving just as much.","Magnitude does not determine direction: Italian and Turkish share net magnitude 0.53 but move in different directions, while German (0.89) and Turkish (0.53) move different distances in similar directions, so cross-linguistic comparison must report both quantities.","Pragmatics is the most shared dimension of change across the five languages (mean pairwise cosine 0.92, with personal deixis rising in all five), while syntax shows no common direction (0.00), suggesting shared institutional or cultural pressures shape parliamentary language beyond genealogical relatedness.","Rankings of which language changed most depend on the historical window: German is largest over its full record, Turkish over the common 1950–2020 interval, so claims about relative change are interval-relative."],"supporting_citations":[{"why":"Supplies the sparse crosscoder architecture — a shared feature index with separate decoders per checkpoint — that ChronoLens adapts from model checkpoints to languages and historical periods.","marker":"Lindsey et al., 2024"},{"why":"Motivates the BatchTopK choice over L1 sparsity and supplies the latent-scaling check used to keep shared features from being mislabeled condition-specific.","marker":"Minder et al., 2026"},{"why":"Supplies the ParlaMint multilingual parliamentary corpora that anchor the comparable data across all five languages.","marker":"Erjavec et al., 2023, 2024"},{"why":"Supplies the UK Hansard corpus (1803–2004), the main historical English source.","marker":"Coole et al., 2020"},{"why":"Supplies GermaParl, the principal German parliamentary protocol corpus.","marker":"Blätte and Blessing, 2018"},{"why":"Supplies the Polish Parliamentary Corpus, the main Polish source since 1919.","marker":"Ogrodniczuk and Nitoń, 2020"},{"why":"Supplies ItaParlCorpus of Italian parliamentary speech turns since 1948.","marker":"Cova, 2025"},{"why":"Supplies the TBMM corpus of Turkish Grand National Assembly transcripts since 1950.","marker":"Güngör, 2018"},{"why":"Supplies the permuted-label probe-control methodology used to validate that level assignments are decodable beyond chance.","marker":"Hewitt and Liang, 2019"},{"why":"Defines the diachronic-embedding approach and evaluation style that the embedding baseline and the paper's comparisons build on.","marker":"Hamilton et al., 2016b"}],"fun_headline_variants":["Change size doesn't predict drift direction in 5 languages","17B tokens show language change has two dimensions: size & direction","How far vs which way: 44M documents from 5 parliaments","Crosscoder: same change size, different drift direction","5 languages, one space: size vs direction of change"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The measured displacements are read as language change rather than corpus change: the five parliamentary traditions have unequal historical coverage, different source mixes, and OCR noise concentrated in the earliest material, so if length- and topic-matching do not fully control those differences, the magnitude and direction results would describe corpus artifacts, not language change.","fun_headline_variants_meta":{"raw":{"variants":["Change size doesn't predict drift direction in 5 languages","17B tokens show language change has two dimensions: size & direction","How far vs which way: 44M documents from 5 parliaments","Crosscoder: same change size, different drift direction","5 languages, one space: size vs direction of change"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000821,"raw_usage":{"total_tokens":3430,"prompt_tokens":745,"completion_tokens":2685,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":2598}},"tokens_in":489,"tokens_out":2685,"duration_ms":21727,"temperature":1.0,"reasoning_tokens":2598,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:41:19.007881+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the pipeline on a condition-shuffled null — tuples whose decade or language identities are randomly exchanged, a control the paper already runs at feature level — and apply the full magnitude-and-direction analysis to it. If the null reproduces the ρ ≈ 0.72 linguistic agreement or the German-versus-Turkish magnitude split, the signal is a method artifact, not a fact about time. A complementary check: apply the pipeline to a decades-spanning corpus of stable, repeatedly reprinted texts (legal or liturgical boilerplate) with a similar OCR profile; if it reports substantial change, the traj","supporting_citations":[],"review_version":1}