{"id":"ba7c797e-8cb7-4e7b-a212-7e9b940f6e8a","arxiv_id":"1908.03676","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GLM maximum likelihood estimators have almost sure error of order sqrt(log log n / n), and penalties growing faster than log log n make model selection strongly consistent.","lead":"This paper proves a law of the iterated logarithm for maximum likelihood estimates in generalized linear models, covering both independent and weakly dependent responses. The sharp almost sure rate is then used to show that BIC-type model selection criteria select the simplest correct model almost surely.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LIL lower bound for dependent responses is not proved: Lemma H.1 requires strict stationarity, while the weighted score sums in §H have time-varying coefficients, and the proof after (70) only yields an O(1) upper bound.","rationale":"The reader's weakest assumption correctly identifies the dependent-case proof's central gap: Lemma H.1 is a strict-stationarity LIL, and it is applied to nonstationary weighted score sums. The proof of Theorem 3.5 itself only derives an O(1) upper-bound normalizer for the score after (70), which cannot yield the positive lower bound b>0 claimed in the second display. This is the single most load-bearing concern because Theorem 3.5 is presented as the dependent analogue of Theorem 3.1, and the paper's title and abstract advertise the LIL as the main contribution. The upper-bound half of the argument, O(√(n^{-1}log log n)) almost surely, appears coherent and is what actually drives the BIC/SCC consistency proofs in Theorems 3.4 and 3.6; those consequences would survive even if the sharp lower bound is weakened or deferred. Thus the appropriate disposition remains CONDITIONAL: the manuscript needs a proof of a nonstationary LIL for weighted ρ-mixing score sums (or an explicit weakening of the theorem to the upper bound), not a rejection of the model selection results. I agree with the reader's choice of weakest assumption; my rationale additionally notes that the independent-case Corollary A.1 has a related normalization gap, but the dependent-case stationarity failure is the clearest and most consequential flaw. The proposed simulation is a direct check: with an alternating design, the strict-stationarity hypothesis is visibly violated, and the limsup constant should be tested rather than assumed.","tokens_in":22540,"tokens_out":11860,"duration_ms":138205,"concrete_test":"Take p=1, x_i=(-1)^i, β0=1, and an AR(1) response y_i=1+ε_i with ε_i=0.5ε_{i-1}+e_i, e_i iid N(0,1). Simulate 10^4 paths of length n=10^6 and compute the ratio |S_n|/√(2 E[S_n²] log log E[S_n²]) for S_n=Σ_{i=1}^n x_i ε_i, where E[S_n²] is the exact variance of the alternating weighted sum. If the running limsup of this ratio depends on the phase of x_i (e.g., start at i=1 versus i=2) or drifts with n instead of converging to one fixed constant, the strict-stationarity premise of Lemma H.1 fails exactly as in §H, and the lower bound b>0 in Theorem 3.5 cannot be derived from the submitted argument.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central dependent-case claim, Theorem 3.5, needs both the upper bound (15) and the positive lower bound b>0 in the second display. The proof in Section H invokes Lemma H.1, a LIL for strictly stationary ρ-mixing sequences, for the score components S_n(β0)_j = Σ_{k=1}^n w_k x_{kj} u'(x_k^Tβ0)[y_k − ˙b(u(x_k^Tβ0))]. These summands have coefficients that depend on k through the fixed design x_k and weights w_k, and the response means E(y_k) = ˙b(u(x_k^Tβ0)) also vary with x_k. Thus, unless x_k^Tβ0 is constant, the summands are not identically distributed and the process is not strictly stationary, so Lemma H.1 does not apply. This inconsistency is visible already in (12), where y_i = E(y_i)+ε_i with fixed covariates, side by side with the restriction to strictly stationary sequences in Section 3.4. Even if Lemma H.1 were applicable, the paragraph around (69)-(70) only verifies Var(S_n(β0)_j)=O(n) and concludes limsup |S_n(β0)_j|/√(2n log log n)=O(1). That is enough for the almost-sure rate (15), but it cannot establish the positive lower bound b>0 claimed in Theorem 3.5. The same structural gap appears in the independent proof: Corollary A.1 applies Lemma A.2 to non-identically distributed weighted sums, and condition (i) of Lemma A.2 is checked only as an O(1) bound in (24), not as a limit. The upper-bound route to model selection consistency, Theorems 3.4 and 3.6, is not affected, but the sharp LIL lower bound stated as the headline result is unsupported for both independent and dependent responses as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the law of the iterated logarithm (LIL) for maximum likelihood estimation in fixed-dimensional generalized linear models with a weighted log-likelihood, for independent responses and for weakly dependent (rho-mixing or m-dependent) responses. The main theoretical claims are Theorem 3.1 and Theorem 3.5: for any correct sub-model, the Euclidean distance between the weighted MLE and the true parameter is O(sqrt(n^{-1} log log n)) almost surely, with a positive limsup constant b when normalized by sqrt(n^{-1} log log n). Theorems 3.2 and 3.3 give corresponding bounds on log-likelihood discrepancies for correct and incorrect sub-models, and Theorems 3.4 and 3.6 conclude that BIC and SCC are strongly consistent for model selection while AIC is not. A simulation study in Section 4 compares BIC and AIC for negative-binomial, probit, and dependent linear models.","tokens_in":22863,"tokens_out":11346,"duration_ms":125228,"significance":"If the LIL results were valid, they would provide the sharpest strong-convergence rate for GLM estimators in fixed dimension, and they would give a clean proof of strong consistency for BIC and SCC, including non-canonical links and weakly dependent responses. That would be a genuinely useful contribution. The paper has a sensible high-level strategy: use strong convexity of the negative log-likelihood (Lemma A.1) to localize the estimator, bound the linear and quadratic error terms by score LILs, and then separate correct from incorrect sub-models. The simulation results for BIC versus AIC are also in the expected direction, and the authors honestly state that no real-data analysis is included. However, the proof of the LIL statements has load-bearing gaps in both the independent and dependent cases: the variance normalization in Corollary A.1 is not the correct normalization for general weights, and the dependent-case proof applies a stationary LIL to nonstationary weighted sums. These issues are central, not cosmetic, and they affect the headline theorems.","major_comments":[{"comment":"Petrov's LIL, Lemma A.2, requires condition (i) as a limit: lim_{n->infty} n^{-1} Var(S_n) = sigma^2 < infinity. The proof of Corollary A.1 verifies only n^{-1} Var(S_{nj}) <= W n^{-1} I_n(beta_0)(j,j) = O(1), which is not the same as convergence. In addition, Var(S_{nj}) = sum_k w_k^2 dot{u}(x_k^t beta_0)^2 ddot{b}(u(x_k^t beta_0)) x_{kj}^2, whereas I_n(beta_0)(j,j) uses w_k rather than w_k^2. For a constant weight c, the limsup in Eq. (20) is c^{1/2} times the claimed value, so Eq. (20) as stated is false for general weights. Consequently Corollary A.1 is not established, and the positive lower-bound construction in the proof of Theorem 3.1 after Eq. (41) is unsupported.","section":"Appendix B, Corollary A.1, Eqs. (20)-(24)"},{"comment":"Lemma H.1 is a LIL for strictly stationary rho-mixing sequences, but the score components S_n(beta_0)_j = sum_k w_k x_{kj} dot{u}(x_k^t beta_0)[y_k - dot{b}(u(x_k^t beta_0))] have coefficients and means that depend on k through the fixed design x_k. Under (H.1), the linear predictors x_k^t beta_0 cannot be constant, so the summands are non-identically distributed and the process is not strictly stationary. The restriction to strictly stationary sequences in Section 3.4 is therefore incompatible with the model (12) and with (H.1). The same issue applies to the m-dependent case via Lemma H.2. Even if a nonstationary LIL were supplied, the calculation around (69)-(70) establishes only I_n(beta_0)(k,k) = O(n) and gives limsup |S_n|/sqrt(2n log log n) = O(1), which cannot yield the positive lower bound b > 0 claimed in Theorem 3.5.","section":"Section H, Lemma H.1 and Eq. (12)"},{"comment":"The model-selection claims inherit the same gaps. For example, Eq. (57) in the proof of Theorem 3.2 uses the score bound from Corollary A.1, and Eq. (58) uses Corollary A.2, both of which rest on the unproved normalization in Corollary A.1. In the dependent case, Section H is the only score-LIL argument, and it is not valid for the nonstationary weighted sums. Thus Theorems 3.4 and 3.6 are not established as written. This is not to say the upper-bound rate or BIC/SCC consistency are false; they may be salvageable with additional assumptions such as two-sided bounded weights and an appropriate nonstationary LIL, but the present manuscript does not provide such a proof.","section":"Propagation to Theorems 3.2-3.4 and 3.6"}],"minor_comments":[{"comment":"The statement of (H.1) says 'Fisher information of beta given by (6)', but Eq. (6) defines S_n(alpha); the Fisher information I_n(beta) is defined just before Eq. (6). This cross-reference should be corrected.","section":"Section 3.2, condition (H.1)"},{"comment":"The notation ||x_k^t||_2^1 is garbled; it should likely be ||x_k||_1 or a clearly defined norm. The same notation recurs in later bounds and should be fixed for readability.","section":"Eq. (30) and surrounding text"},{"comment":"The displayed statement says limsup |S_n|/sqrt(2n log log n) = sigma^2, but under the displayed definition sigma^2 is a variance; the right-hand side should be sigma, not sigma^2, unless the definition is changed.","section":"Lemma H.2"},{"comment":"There is a typo 'Simualtion' in the table caption, and the AIC rows for the dependent linear model at n=300 are identical across MR(2) and MR(3), which suggests a copy-paste error that should be corrected.","section":"Table 1"},{"comment":"The citation 'Nelder (1972)' should be 'Nelder and Wedderburn (1972)', since the cited work is the joint paper introducing generalized linear models.","section":"References"}],"recommendation":"reject","confidential_remarks":"The decision is based on the technical gaps in Corollary A.1 and Section H, not on any concern about circularity or attribution. I do not see evidence of prior-art suppression; the cited literature is appropriate. The main problem is that the headline LIL theorems, and the dependent-case results in particular, are not proven as stated, and the normalization error in Corollary A.1 is not a local typo but a structural mismatch with the stated theorem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: This paper extends the GLM strong-limit literature from binomial responses to general exponential families with non-canonical links and weakly dependent responses. The upper-bound rate and the BIC/SCC consistency application look sound and are worth having. The exact LIL lower bound, however, is not proved as stated, in either the independent or the dependent case.\n\nThe paper's core move is standard and correctly executed for the upper bound: convexity of the negative log-likelihood gives quadratic control, the score is bounded by O(sqrt(n loglog n)) a.s., and the third-order remainder is lower order. That yields ||beta_hat - beta0|| = O(sqrt(n^{-1} loglog n)) a.s., which is the rate needed for the model-selection consistency argument. The BIC/SCC consistency proof is the usual penalty comparison, and the simulations, while not extensive, support the claimed distinction between BIC and AIC. The citation pattern is honest: the convexity-plus-LIL technique comes from Wu and Zen and Qian and Wu, and the paper credits them.\n\nThe soft spots are in the lower-bound half. Corollary A.1 asserts a limsup = 1 law for the weighted score, but its proof verifies only Var(S_nj)/n = O(1). Petrov's LIL needs a proper limit for the normalizer; an O(1) bound cannot establish a positive limsup. The dependent case is worse: Lemma H.1 is a LIL for strictly stationary rho-mixing sequences, while the score sums have non-random weights and design vectors that vary with k, so the summands are not identically distributed and the process is not stationary. The paper even states in Section 3.4 that it restricts to strictly stationary sequences, which sits awkwardly with the fixed-covariate model in (12). The calculation after (70) only produces limsup |S_n| / sqrt(2n log log n) = O(1), which cannot give the positive constant b in Theorem 3.5. There is also a small but real error in (13), where the estimating equation is written as S_n(beta0)=0 rather than S_n(beta)=0.\n\nNone of this undermines the upper-bound rate or the model-selection consistency theorems if restated with the upper bound only. So the paper is a conditional contribution: it identifies the right problems and gets the easy half right, but the sharp LIL needs a genuine proof for nonstationary weighted sums, or the claims need to be downgraded.\n\nI'd send it to a serious referee, with the clear expectation that the dependent-case LIL and the normalization in Corollary A.1 must be fixed or removed. Someone working on strong limit theorems for GLMs will want to read it; the model-selection application makes it worth engaging.","headline":"Useful upper-bound and model-selection results wrapped in an unproved LIL lower bound; referee it, but expect the sharp claims to be repaired or dropped.","tokens_in":23489,"tokens_out":3303,"would_cite":false,"duration_ms":31946,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","60F15","62J12","62M10"],"pacs":[],"model":"deepseek-v4-flash","headline":"For generalized linear models, the MLE obeys the law of the iterated logarithm, which makes BIC and SCC select the simplest correct model almost surely while AIC does not.","keywords":["law of the iterated logarithm","generalized linear models","model selection consistency","BIC","AIC","SCC","ρ-mixing","m-dependent"],"falsifier":"For a probit GLM with $n=10{,}000$, fixed bounded design, and geometric $\\rho$-mixing Bernoulli responses, compute the weighted score coordinate $S_n(\\beta_0)_j$ and its empirical variance $\\sigma_n^2$, and trace the ratio $|S_n|/\\sqrt{2\\sigma_n^2\\log\\log\\sigma_n^2}$ along the LIL subsequence $n_k = \\lfloor e^{k}\\rfloor$; if the ratio does not approach 1 almost surely (e.g., it converges to 0 or oscillates with the weight pattern), Theorem 3.5's positive lower bound fails.","tokens_in":22251,"feed_emoji":"📈","tokens_out":13201,"duration_ms":118854,"temperature":0.7,"pith_summary":"This paper sets out to prove that maximum likelihood estimation in fixed-dimensional generalized linear models reaches the sharpest almost-sure convergence rate known for such problems: under mild conditions, the Euclidean error $\\|\\hat\\beta-\\beta_0\\|$ is $O(\\sqrt{n^{-1}\\log\\log n})$ almost surely, and the limsup of the normalized error is a strictly positive constant. That rate is exactly what is needed to settle model selection: with it, any penalized-likelihood criterion whose penalty grows faster than $\\log\\log n$ but slower than $n$ selects the simplest correct submodel with probability one. The paper shows BIC and the stochastic complexity criterion have precisely this property, while AIC does not. The same conclusions are claimed when the responses are weakly dependent — $\\rho$-mixing with geometric decay or $m$-dependent — extending the theory to time-series GLMs. Simulations for negative-binomial, probit, and dependent linear models confirm the BIC/AIC contrast that the theory predicts.","feed_headline":"BIC beats AIC in GLMs: the law of the iterated logarithm proves it","feed_subtitle":"Under mild conditions, the MLE obeys the sharp strong rate, so BIC and SCC select the true model almost surely.","key_machinery":"The argument's engine is the weighted score function evaluated at the truth, $S_n(\\beta_0) = \\sum_{k=1}^n w_k x_k \\dot{u}(x_k^{\\scriptscriptstyle T}\\beta_0)[y_k - \\dot{b}(u(x_k^{\\scriptscriptstyle T}\\beta_0))]$, a sum of centered random variables. A classical LIL for independent sums — Lemma A.2 in the paper — fixes the almost-sure limsup of each coordinate of $S_n(\\beta_0)$ at the scale $\\sqrt{n\\log\\log n}$. The strongly convex log-likelihood (Lemma A.1) then turns this score bound into a quadratic lower bound on the log-likelihood ratio $K_n(\\beta,\\beta_0)$ on spheres $\\partial B_n$ of radius $\\tau_n\\sqrt{n^{-1}\\log\\log n}$; because the quadratic term dominates the linear and cross terms on that boundary, the minimizer must lie inside $B_n$, giving the $O(\\sqrt{n^{-1}\\log\\log n})$ error rate. The positive limsup $b>0$ is forced by a subsequence argument: one coordinate of the score reaches its LIL limit along $n_i$, and if the estimator's error were $o(\\sqrt{n^{-1}\\log\\log n})$, the likelihood ratio at a cleverly chosen competitor would have to be both $o(\\log\\log n)$ and $\\le -O(\\log\\log n)$. For dependent responses the same machinery is run with the score LIL replaced by mixing-sequence LILs (Lemmas H.1 and H.2).","core_discovery":"The central discovery is that the MLE in a GLM satisfies a law of the iterated logarithm. Theorem 3.1 states that under (H.1)–(H.4), for every correct submodel $\\alpha$, $\\|\\hat\\beta(\\alpha)-\\beta_0(\\alpha)\\| = O(\\sqrt{n^{-1}\\log\\log n})$ almost surely and $\\limsup_n \\|\\hat\\beta(\\alpha)-\\beta_0(\\alpha)\\|/\\sqrt{n^{-1}\\log\\log n} = b > 0$ almost surely. The proof goes through the weighted score at $\\beta_0$, whose coordinates obey a classical LIL for independent sums, and a strong-convexity/local-quadratic argument that converts score bounds into estimator bounds. From this, Theorem 3.2 bounds the maximized log-likelihood gap of any correct model by $O(\\log\\log n)$ almost surely, and Theorem 3.3 shows any wrong model loses to the true likelihood by a positive multiple of $n$; Theorem 3.4 then concludes that penalties of order strictly between $\\log\\log n$ and $n$ — in particular BIC and SCC but not AIC — select the smallest correct model almost surely. Theorem 3.5 extends the LIL to $\\rho$-mixing (geometric decay) and $m$-dependent responses, and Theorem 3.6 carries the same model-selection conclusion to those settings.","pith_inferences":["The proof structure — strong convexity plus a LIL for the score — should carry over to penalized convex M-estimators outside exponential families, since the GLM-specific inputs are only the moment bound and the bounded-Hessian conditions.","The strictly positive constant $b$ is an existence result: the paper neither identifies it nor suggests how to estimate it, so the LIL is best used as a rate guarantee rather than as a basis for confidence sets.","A natural stress test is to run the dependent-case simulation with deterministic oscillating weights $w_k = 1 + 0.5\\sin(k)$ (allowed by (H.4)) and check whether the normalized score still has limsup 1; this checks whether the stationarity assumption in the cited mixing LIL matters for the conclusion."],"forward_implications":["Any penalty term whose order lies strictly between $O(\\log\\log n)$ and $O(n)$ yields strong model-selection consistency; the $\\log n$ penalties of BIC and SCC are the standard instances, and AIC's $O(1)$ penalty is not.","Under (H.1)–(H.4), every fixed-dimensional correct GLM submodel has MLE error $O(\\sqrt{n^{-1}\\log\\log n})$ almost surely, so the estimators are strongly consistent at the iterated-logarithm rate.","With geometric $\\rho$-mixing or $m$-dependence, the same LIL and hence the same BIC/SCC consistency holds for time-series GLM responses.","Since the conditions cover non-natural links, probit regression and negative-binomial regression fall within the theory, not just canonical-link GLMs."],"supporting_citations":[{"why":"Supplies Petrov's LIL (Lemma A.2) for centered independent sums, the backbone of the independent score bound.","marker":"Stout (1974)"},{"why":"Provides Lemma H.1, the stationary ρ-mixing LIL invoked in the dependent-response extension.","marker":"Lin and Lu (1997)"},{"why":"Provides Lemma H.2, the m-dependent LIL used for the second dependent case.","marker":"Chen (1997)"},{"why":"The binomial-GLM strong model-selection LIL that this paper generalizes to arbitrary GLM links; also the template for the correct/wrong-model penalty-order comparison.","marker":"Qian and Wu (2006)"},{"why":"Initiates the LIL-based strong consistency argument for model selection in regression that the penalty-order theorem follows.","marker":"Rao and Wu (1989)"},{"why":"Establishes strongly consistent information criteria for linear M-estimation; supplies the regularity-condition spirit behind (H.1).","marker":"Wu and Zen (1999)"},{"why":"Lemma A.1 supplies the strong-convexity sandwich bounds that turn score LILs into estimator error bounds.","marker":"Wright (2017)"},{"why":"Provides the GLM MLE consistency/asymptotic-normality framework and the non-natural-link setup the paper starts from.","marker":"Fahrmeir and Kaufmann (1985)"},{"why":"Its Lemma 6.1 gives the exponential-family third-moment bound used in Remark 1 to verify the moment condition of the LIL.","marker":"Rigollet (2012)"}],"fun_headline_variants":["LIL proves BIC selects true GLM model a.s.","Sharp law: BIC picks correct GLM model almost surely","GLM model selection: BIC provably best via LIL","Law of iterated logarithm settles GLM selection: BIC wins","BIC's edge in GLMs: LIL guarantees true model choice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the weighted score sums obeying a law of the iterated logarithm with a strictly positive limsup constant; for dependent responses the proof invokes an LIL that assumes stationarity, while the score sums themselves are not stationary because of the fixed weights and covariates.","fun_headline_variants_meta":{"raw":{"variants":["LIL proves BIC selects true GLM model a.s.","Sharp law: BIC picks correct GLM model almost surely","GLM model selection: BIC provably best via LIL","Law of iterated logarithm settles GLM selection: BIC wins","BIC's edge in GLMs: LIL guarantees true model choice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2892,"prompt_tokens":1002,"completion_tokens":1890,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1800}},"tokens_in":618,"tokens_out":1890,"duration_ms":13004,"temperature":1.0,"reasoning_tokens":1800,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:07:08.751116+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a probit GLM with $n=10{,}000$, fixed bounded design, and geometric $\\rho$-mixing Bernoulli responses, compute the weighted score coordinate $S_n(\\beta_0)_j$ and its empirical variance $\\sigma_n^2$, and trace the ratio $|S_n|/\\sqrt{2\\sigma_n^2\\log\\log\\sigma_n^2}$ along the LIL subsequence $n_k = \\lfloor e^{k}\\rfloor$; if the ratio does not approach 1 almost surely (e.g., it converges to 0 or oscillates with the weight pattern), Theorem 3.5's positive lower bound fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Petrov's LIL (Lemma A.2) for centered independent sums, the backbone of the independent score bound."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Lemma H.1, the stationary ρ-mixing LIL invoked in the dependent-response extension."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Lemma H.2, the m-dependent LIL used for the second dependent case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The binomial-GLM strong model-selection LIL that this paper generalizes to arbitrary GLM links; also the template for the correct/wrong-model penalty-order comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Initiates the LIL-based strong consistency argument for model selection in regression that the penalty-order theorem follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Lemma A.1 supplies the strong-convexity sandwich bounds that turn score LILs into estimator error bounds."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the GLM MLE consistency/asymptotic-normality framework and the non-natural-link setup the paper starts from."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Its Lemma 6.1 gives the exponential-family third-moment bound used in Remark 1 to verify the moment condition of the LIL."}],"review_version":1}