{"id":"ee1d0dcd-8a43-479c-a6cc-6e74bcd8d989","arxiv_id":"1906.11483","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 28 languages, an information-theoretic irregularity score derived from neural transduction models correlates positively with frequency, with stronger effects when aggregated over paradigms.","lead":"The paper measures morphological irregularity across 28 languages using a neural model of form predictability and finds that irregularity correlates positively with word frequency. A smart generalist might read it to see how computational methods can test classic ideas from linguistics about why some words resist regular patterns.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Neural transduction model's irregularity proxy may embed frequency-dependent learning artifacts rather than pure linguistic irregularity","rationale":"Reader's weakest assumption directly identifies the same load-bearing point. Full text availability does not remove the need for external validation of the proxy; the concern therefore leaves the verdict at UNVERDICTED.","tokens_in":1642,"tokens_out":279,"duration_ms":15436,"concrete_test":"For one language with existing hand-coded regular/irregular labels (e.g., English verbs), recompute the frequency-irregularity correlation using the traditional binary labels instead of the neural scores; if the correlation vanishes or reverses while the neural version retains it, the proxy is the source of the signal.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on an information-theoretic irregularity score derived from a neural transducer's predictive performance. Because the model is trained on the same frequency distribution whose correlation is being measured, any frequency-dependent optimization (e.g., better memorization or lower cross-entropy on high-count items) can directly modulate the estimated irregularity, creating a mechanical rather than linguistic link. The abstract and reader's weakest assumption correctly flag this; without explicit frequency-balanced training, held-out evaluation that isolates count effects, or external validation against hand-labeled regular/irregular classes, the reported correlation cannot be cleanly attributed to morphology.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper defines an information-theoretic measure of morphological irregularity based on predictability from a neural transduction model, estimates it for forms across 28 languages, performs validatory and exploratory analyses, and reports a correlation with frequency: higher-frequency items tend to be more irregular and irregular items tend to be more frequent. The correlation strengthens when aggregated over paradigms, which the authors interpret as support for abstract stem/lexeme representations. Code is released.","tokens_in":1752,"tokens_out":434,"duration_ms":18955,"significance":"If the central correlation survives controls for frequency-dependent artifacts in the neural model, the result would supply the broadest cross-linguistic empirical support yet for classic linguistic claims linking irregularity and frequency. The paradigm-level finding and the public code release are clear strengths that aid reproducibility and allow direct testing of the measure.","major_comments":[{"comment":"Methods (neural transduction model and irregularity estimation): because the model is trained on the same frequency distribution later used for the correlation, any frequency-dependent optimization (better memorization or lower cross-entropy on high-count items) can directly affect the estimated irregularity score. Without frequency-balanced training, frequency-stratified held-out evaluation, or external validation against hand-labeled regular/irregular classes, the reported correlation risks being partly mechanical rather than linguistic.","section":"Methods (neural model training and irregularity estimation)"},{"comment":"Results (paradigm-level aggregation): the claim that the correlation is 'more robust when aggregated at the level of whole paradigms' is central to the linguistic interpretation, yet the manuscript provides no explicit definition of how paradigms are constructed, no statistical comparison of form-level vs. paradigm-level effect sizes, and no error analysis showing that the improvement is not driven by a few high-frequency irregular paradigms.","section":"Results (correlation analyses)"}],"minor_comments":[{"comment":"Abstract: 'irregular items are more likely be highly frequent' contains a missing 'to'.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments. We address each major point below, providing the strongest honest defense of the manuscript while acknowledging where revisions are warranted.","responses":[{"response":"We acknowledge the potential for frequency-dependent effects in model training. However, any such bias would improve predictability (reduce estimated irregularity) for high-frequency items, which works directly against the observed positive correlation between frequency and irregularity. The reported result is therefore conservative with respect to this artifact. The manuscript already includes validatory analyses comparing the measure against known morphological patterns in several languages; we will add an explicit discussion of this point in the revision.","revision_made":"partial","referee_comment":"Methods (neural transduction model and irregularity estimation): because the model is trained on the same frequency distribution later used for the correlation, any frequency-dependent optimization (better memorization or lower cross-entropy on high-count items) can directly affect the estimated irregularity score. Without frequency-balanced training, frequency-stratified held-out evaluation, or external validation against hand-labeled regular/irregular classes, the reported correlation risks being partly mechanical rather than linguistic."},{"response":"We agree these details should be clarified. In the revised manuscript we will (i) explicitly define paradigm construction (forms grouped by shared lemma), (ii) report a direct statistical comparison of effect sizes between form-level and paradigm-level analyses, and (iii) include a robustness check or error analysis confirming the stronger paradigm-level correlation is not driven by a small number of high-frequency paradigms. These additions will be straightforward to implement from the existing data and code.","revision_made":"yes","referee_comment":"Results (paradigm-level aggregation): the claim that the correlation is 'more robust when aggregated at the level of whole paradigms' is central to the linguistic interpretation, yet the manuscript provides no explicit definition of how paradigms are constructed, no statistical comparison of form-level vs. paradigm-level effect sizes, and no error analysis showing that the improvement is not driven by a few high-frequency irregular paradigms."}],"tokens_in":1302,"tokens_out":442,"duration_ms":15190,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main result is that irregularity, scored as poor predictability from a neural transducer, correlates with frequency across 28 languages, and the link is stronger when forms are grouped into paradigms. This is the widest test of the classic hypothesis so far, and the paradigm-level finding gives some support to lexeme-based models of morphology. Releasing the code is useful for anyone who wants to rerun or extend it. The validatory analyses mentioned in the abstract are a step in the right direction. The soft spot is the potential artifact in the irregularity measure itself. The transducer is trained on the same frequency distribution it is later correlated against, so high-count items can simply receive more effective learning and appear more predictable. That creates a mechanical link rather than a purely linguistic one. The abstract does not spell out frequency-balanced training, held-out checks that isolate count effects, or direct comparison to hand-labeled regular/irregular classes, so it is hard to tell how much of the reported correlation survives those controls. The stress-test concern lands here. This work is aimed at computational morphologists and people modeling language change who need large-scale data points. A reader can extract the paradigm result and the language coverage even if they later replace the neural proxy with something else. It is worth sending to peer review because the scale is new and the question is substantive; the methods section will decide whether the central claim holds up.","headline":"Broad multi-language correlation between frequency and neural-measured irregularity, but the proxy may inherit frequency biases from training.","tokens_in":2255,"tokens_out":346,"would_cite":false,"duration_ms":15632,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Linguistic morphology correlation study using neural transduction; no RS-shaped machinery","alignment":"orthogonal","rationale":"Paper defines information-theoretic irregularity via neural model predictability P(w|ℓ,σ,L−ℓ) and correlates with frequency across 28 languages. Central machinery is entirely within computational linguistics (UniMorph, wug-testing, lexeme-level aggregation). RS framework (reality_from_one_distinction, J-cost functional equation, φ-ladder, 8-tick periodicity, AlexanderDuality D=3) has no opinion on morphological irregularity or frequency distributions.","tokens_in":49323,"confidence":"high","tokens_out":135,"duration_ms":5133,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Analyses of 28 languages show higher-frequency words are more likely to be morphologically irregular.","keywords":["morphological irregularity","word frequency","neural transduction model","information-theoretic measure","linguistic paradigms","cross-linguistic analysis","morphology"],"falsifier":"A new cross-linguistic dataset in which linguist-assigned irregularity ratings show no correlation with the model's scores, or in which frequency and irregularity are uncorrelated.","tokens_in":2519,"feed_emoji":"","tokens_out":407,"duration_ms":17399,"temperature":0.7,"pith_summary":"The paper defines morphological irregularity through an information-theoretic measure of how predictable a word form is given others in its paradigm. Using a neural transduction model, it estimates this quantity across forms in 28 languages and conducts analyses that reveal a correlation with frequency. Higher-frequency items tend to be irregular, and irregular items tend to be high-frequency. The pattern strengthens when measured at the level of entire paradigms rather than isolated forms. A sympathetic reader would care because the result supplies broad empirical backing for longstanding linguistic ideas about how usage shapes structure and how abstract stems organize inflected words.","feed_headline":"Higher-frequency words tend to be morphologically irregular","feed_subtitle":"Study of 28 languages finds the pattern strengthens when forms are grouped into whole paradigms.","key_machinery":"An information-theoretic measure of irregularity based on the predictability of forms, estimated by a neural transduction model.","core_discovery":"The central claim is that morphological irregularity correlates with frequency: higher frequency items are more likely to be irregular and irregular items are more likely to be highly frequent. This holds across the 28 languages examined, is more robust when forms are grouped into whole paradigms, and supplies the first confirmation of this breadth for proposals from the linguistics literature.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Irregularity correlates with frequency","Morphological irregularity links to higher frequency","Frequency and irregularity correlate in 28 languages","Whole paradigms confirm frequency irregularity tie","High frequency words show more irregularity"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The neural model's estimate of irregularity matches the linguistic notion of irregularity instead of reflecting model-specific artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Irregularity correlates with frequency","Morphological irregularity links to higher frequency","Frequency and irregularity correlate in 28 languages","Whole paradigms confirm frequency irregularity tie","High frequency words show more irregularity"]},"model":"grok-4.3","cost_usd":0.006859,"raw_usage":{"total_tokens":3137,"prompt_tokens":572,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":68587000,"prompt_tokens_details":{"text_tokens":572,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2508,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":572,"tokens_out":57,"duration_ms":17932,"temperature":1.0,"reasoning_tokens":2508,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T15:15:59.110411+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new cross-linguistic dataset in which linguist-assigned irregularity ratings show no correlation with the model's scores, or in which frequency and irregularity are uncorrelated.","supporting_citations":[],"review_version":1}