{"id":"a0b26d1c-2c22-4a9d-a5f1-14143496c379","arxiv_id":"1908.03329","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A tutorial showing that OLS, ridge, and LASSO estimators arise as Bayesian posterior modes under different priors, with a derivation of KIC for model selection.","lead":"This note derives the formulas for ordinary least squares, ridge regression, LASSO, and the Kashyap information criterion inside a single Bayesian framework. It is a tutorial rather than a research contribution, and the derivations are mostly standard textbook material.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 2.4's fixed-sign Gaussian derivation is not the LASSO posterior; a one-dimensional sparse example gives no sign-consistent solution, and the KIC in §3.1 does not apply.","rationale":"The reader's conditional verdict already rests on the LASSO sign/Gaussian assumption. My independent check of §2.4 confirms that the concern is real and concrete, not a style preference: in a one-parameter problem with |y|<λ, the true LASSO posterior mode is zero, while the paper's fixed-sign equation has no admissible solution. This directly falsifies Eq. (16) as the LASSO MAP and invalidates the subsequent KIC use for sparse models. I do not see a reason to escalate to rejection: the paper is explicitly a tutorial, the OLS and ridge derivations (modulo the uniform-prior truncation issue) are standard, and the flaw is localized to §2.4/§3.1. A corrected treatment of the Laplace prior — e.g., presenting the MAP as the solution of the l1-regularized least-squares problem and omitting the Gaussian evidence formula for that case — would be enough to make the notes usable. Hence the verdict remains conditional.","tokens_in":6631,"tokens_out":8451,"duration_ms":90652,"concrete_test":"Run the scalar example N=P=1, Ψ=1, y=0.2, C_M=1, a0=0, Λ=1. The exact posterior is p(a|y) ∝ exp(-0.5(0.2-a)^2 - |a|); its mode is a*=0. Apply the §2.4 fixed-sign iteration: s=0 gives â=0.2, s=+1 gives â=-0.8 (sign mismatch), s=-1 gives â=1.2 (sign mismatch). No sign-consistent fixed point exists, so Eq. (16) is not the MAP. Then compute the exact posterior normalization and compare with the Gaussian N(â,1) implied by Eq. (16); the densities and the resulting log-evidence differ, confirming the KIC in §3.1 is not applicable to LASSO.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The tutorial's central claim is that OLS, ridge, LASSO, and KIC all follow from one Bayesian framework. The LASSO leg is the least secure. In §2.4 the Laplace-prior log posterior is linearized by fixing s = sgn(a - a0), after which the posterior is declared Gaussian, Eq. (16), with mean a0 + H^{-1}(Ψ^T C_M^{-1}(y - Ψa0) - Λ s) and covariance H^{-1}. This is not the LASSO posterior. The exact posterior is a mixture of orthant-truncated Gaussians, one for each sign pattern; it is not a single global Gaussian, so the normalizing constant and covariance in Eq. (16) are wrong. Worse, when the MAP has a coefficient exactly at the prior mean — the typical sparse LASSO situation — sgn is undefined, and treating it as zero silently removes the l1 penalty for that coordinate. The fixed-point formula cannot represent such zero coefficients unless the data term happens to vanish. Consequently Eq. (16) is not generally the LASSO MAP, and the model-evidence/KIC calculation in §3.1, which assumes a Gaussian posterior and a Hessian at the MAP, is invalid for the LASSO case: at a kink the Hessian does not exist. The OLS and ridge sections are not affected, but the paper's own N.B. admits the sign is unknown, which shows the Gaussian conclusion in Eq. (16) is doing unacknowledged work.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"These notes are a pedagogical derivation of Bayesian linear regression. Starting from the linear model y = Ψa + ε with a Gaussian likelihood, the paper derives the posterior distribution of the coefficients under a uniform prior (interpreted as ordinary/generalized least squares), a Gaussian prior (ridge regression), and a Laplace prior (LASSO). It then derives the Kashyap information criterion (KIC) by Laplace approximation of the Bayesian model evidence, proposing KIC for model selection among different subsets A of basis functions. The abstract states that the article presents the author's understanding of how these methods are \"better unified in the Bayesian framework,\" with explicit formulas derived for each estimator.","tokens_in":6907,"tokens_out":5554,"duration_ms":62991,"significance":"If correct, the paper would be a useful tutorial for practitioners who want to see common linear-regression estimators derived from a single Bayesian perspective. The OLS and ridge derivations are standard, the algebra is shown in detail, and the text cites the relevant historical literature. The treatment of the LASSO, however, is not the standard LASSO solution and is mathematically incorrect as written; the KIC derivation is also only valid for smooth posterior densities and cannot be applied to the LASSO case. Since the central claim of the paper is that OLS, ridge, LASSO, and KIC all fit within one framework, this is not a local blemish: the LASSO and KIC legs of the unification need substantial correction. The paper does not report numerical experiments or new statistical results, so its value is essentially tutorial, and tutorial value is only realized if the mathematics is sound.","major_comments":[{"comment":"The derivation leading to Eq. (16) is not a derivation of the LASSO posterior. The sign of b = a - a0 is a function of b, and fixing s = sgn(b) before completing the square and then declaring the posterior to be Gaussian is not valid. The exact posterior under a Laplace prior is a mixture of orthant-truncated Gaussians, one for each sign pattern, with different normalizing constants on each orthant; it is not a single Gaussian. Consequently the mean and covariance in Eq. (16) are not the LASSO MAP and its posterior covariance. A concrete failure occurs in the common sparse situation where the MAP has a coefficient exactly equal to a0: there sgn(b) is undefined, and if one treats it as zero, the l1 penalty for that coordinate is silently removed. The paper's own N.B. after Eq. (16) states that one cannot guess sgn(a - a0) before computing a-hat, which indicates that the Gaussian conclusion in Eq. (16) rests on an unverified assumption, not on a theorem.","section":"Section 2.4, Eqs. (14)-(16)"},{"comment":"The KIC derivation assumes twice differentiability of log p(y, a | A) at the MAP, since Eq. (18) uses a Taylor expansion and Eq. (19) defines Σ^{-1} as the negative Hessian at a-hat. For the Laplace prior used in Section 2.4, the log-joint density is not differentiable at any coefficient value equal to the prior location a0, which is precisely the situation that makes LASSO solutions sparse. Therefore Eq. (22) is not a valid approximation of the Bayesian model evidence for LASSO models. The claim in Section 3.1 that the KIC at the MAP applies generally to the models considered in the paper is too broad; it should be restricted to likelihood/prior combinations for which the posterior density is smooth at the mode, such as the Gaussian-prior ridge case. The closing remark in Section 4 that the KIC in the BSPCE application is implemented with a Gaussian prior is consistent with this limitation, but the main text does not flag it.","section":"Section 3.1, Eqs. (18)-(23)"},{"comment":"The treatment of the uniform prior is incorrect as written. If the prior is uniform on a bounded rectangular domain Ω, the posterior is proportional to the Gaussian likelihood restricted to Ω, which is a truncated Gaussian. It is not the Gaussian density N(a-tilde, C-tilde_aa) even when the unconstrained mean a-tilde lies inside Ω, because truncation changes the normalizing constant and the covariance; and when a-tilde lies outside Ω, the posterior is not zero. The N.B. after Eq. (7) asserts that the posterior is N(a-tilde, C-tilde_aa) if a-tilde ∈ Ω and zero otherwise, which is false. The standard OLS result can be recovered by taking an improper uniform prior over the whole coefficient space, but the manuscript as written bounds Ω first and then ignores the truncation. This needs to be corrected because the OLS section depends on it.","section":"Section 2.1, Eq. (7) and N.B."}],"minor_comments":[{"comment":"There is a typographical error: \"m¯C^{-1}_{aa}\" should read \"C^{-1}_{aa}\" in the expression for the MAP estimate.","section":"Section 2.3, Eq. (13)"},{"comment":"The citation \"G. Schwartz, Estimating the dimensions of a model\" should be to Gideon Schwarz (1978); the name is spelled Schwarz.","section":"References"},{"comment":"The notation in the transition from the likelihood to the Gamma distribution is confusing: the text writes p(σ^2 | ·) and then p(σ^{-2} | ·), and states that the mode is σ-tilde^2. Since the Gamma distribution is stated for σ^{-2}, the mode of that Gamma distribution is 1/σ-tilde^2, not σ-tilde^2. The parameterization should be made explicit.","section":"Section 2.2, Eqs. (9)-(10)"}],"recommendation":"major_revision","confidential_remarks":"This is a tutorial manuscript rather than a research contribution, so its value depends almost entirely on mathematical correctness. The OLS and ridge parts are standard and usable, and I see no reason to reject the manuscript outright: the LASSO section can be corrected or explicitly reduced to a discussion of the standard LASSO optimization problem, and the KIC section can be restricted to smooth posterior cases. However, the current version should not be accepted as is, because the abstract and introduction promise a unified derivation that presently contains load-bearing errors in the LASSO and KIC parts. I would ask the author to either repair Section 2.4 and Section 3.1 or clearly state the limited conditions under which the derived formulas hold."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you need a clean derivation of Bayesian OLS and ridge. The paper is a tutorial, not a research contribution: it re-derives known formulas. The first two sections and the KIC derivation for the Gaussian case are fine. The algebra is clear and the notation is consistent. For someone who wants to see how a uniform prior leads to the MLE, and a Gaussian prior to ridge, these notes work.\n\nThe soft spot is Section 2.4. The Laplace-prior posterior is not a single Gaussian. The author fixes sgn(a - a0), completes the square, and declares p(a|y,A) Gaussian with mean a0 + H^{-1}(...) and covariance H^{-1}. That is the posterior only if the sign pattern is known in advance, and even then it is an orthant-truncated Gaussian, not a global Gaussian. In the sparse cases LASSO is supposed to produce, some coefficients sit at the prior mean, where sgn is undefined; setting it to zero silently drops the l1 penalty for that coordinate. So Eq. (16) is not the LASSO MAP. The N.B. admits the sign is unknown, which is exactly the problem: the Gaussian conclusion is doing unacknowledged work.\n\nThat error propagates. Section 3.1 derives KIC from a Laplace approximation around the MAP. For a Gaussian posterior this is standard and correct. For the LASSO case, at a kink the Hessian does not exist and the posterior is not Gaussian, so the KIC formula does not apply. The author's own conclusion section says BSPCE uses KIC with a Gaussian prior, so this flaw does not affect that application.\n\nOne minor point: in Section 2.1 the uniform prior is on a bounded rectangle; the posterior is a truncated Gaussian, not a full Gaussian. The N.B. acknowledges the support restriction, but Eq. (7) drops the indicator.\n\nThe citation pattern is fine. The self-citation to Shao et al. is only in the concluding remark and is not load-bearing.\n\nWho is this for? Someone who wants a quick, self-contained derivation of OLS and ridge from a Bayesian viewpoint. Do not use it as a reference for the Bayesian interpretation of LASSO; Park and Casella (2008) is the right source. I would not cite it, but I would consider it for a reading group as a cautionary example of what goes wrong when you linearize a nonsmooth prior. If a teaching-oriented or arXiv-focused venue asks me, yes, referee it: it is short, readable, and the LASSO mistake is worth catching. But it should not be published as research, and as written the LASSO section needs major revision before the tutorial is trustworthy.","headline":"A clear but non-novel tutorial whose OLS and ridge sections are reliable, whose LASSO derivation is mathematically wrong, and whose KIC application to LASSO does not survive scrutiny.","tokens_in":7437,"tokens_out":3916,"would_cite":false,"duration_ms":38879,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J05","62F15","62J07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that ordinary least squares, weighted least squares, ridge regression, and LASSO are all the same Bayesian calculation run with different priors on the coefficients.","keywords":["linear regression","Bayesian inference","generalized least squares","ridge regression","LASSO","Kashyap information criterion","model selection","maximum a posteriori estimation"],"falsifier":"For a small dataset, set $a_0=0$ and $\\Lambda_{aa}=\\lambda I$, then compare the closed-form $\\hat{a}$ from Eq. (16) with the mode obtained by directly maximizing the posterior density in Eq. (14) on a fine grid for one or two coefficients; if the sign of $\\hat{a}-a_0$ does not match the sign of the posterior mode, the sign-constancy assumption fails and Eq. (16) is not the LASSO MAP.","tokens_in":6424,"feed_emoji":"📐","tokens_out":9189,"duration_ms":85559,"temperature":0.7,"pith_summary":"These notes aim to show that least-squares, ridge, and LASSO regression are not separate algorithms but the same Bayesian calculation performed with different priors on the coefficients. Under a Gaussian error model, each choice of prior—flat, Gaussian, or Laplace—turns the log-posterior into a quadratic form, so completing the square yields the estimator as the posterior mean and its covariance in closed form. The paper also derives the Kashyap information criterion from the same setup, giving a rule for choosing which coefficient subset to keep. If the derivation is right, a practitioner gets one framework that reproduces several textbook methods and reveals what each implicitly assumes before seeing data. The LASSO part rests on an extra sign-constancy assumption that the author flags and says must be checked afterward.","feed_headline":"One Bayesian formula unifies least squares, ridge, and LASSO","feed_subtitle":"Each common regression estimator is a different prior in disguise, and the same derivation yields the KIC model selector.","key_machinery":"The load-bearing object is completing the square in the exponent of a Gaussian likelihood times a prior: rewriting $(y-\\Psi a)^T C_M^{-1}(y-\\Psi a)+(a-a_0)^T C_{aa}^{-1}(a-a_0)$ as $(a-\\hat{a})^T \\hat{C}_{aa}^{-1}(a-\\hat{a})+\\text{const}$ turns the posterior into $N(\\hat{a},\\hat{C}_{aa})$. This one algebraic manipulation produces the MLE/MAP formulas for every method in the paper. For the Laplace prior the same manipulation is applied to $b=a-a_0$ under the condition that $\\operatorname{sgn}(b)$ stays fixed, and the posterior covariance is taken from the least-squares term alone.","core_discovery":"The paper's claim is that the usual linear-regression toolbox is a single Bayesian inference: choose a Gaussian likelihood for the error and a prior for the coefficients, and the posterior is Gaussian because the log-posterior is a quadratic form in $a$. Completing the square gives the posterior mean and covariance in closed form. A uniform prior on a rectangular domain yields the maximum-likelihood (weighted or ordinary least-squares) estimate; a Gaussian prior $N(a_0,C_{aa})$ yields the ridge-regression MAP; a Laplace prior yields LASSO, provided the sign of $a-a_0$ is constant across the posterior. The same quadratic view supports a Laplace-approximation derivation of the Kashyap information criterion for comparing candidate coefficient subsets, with the best model being the one with lowest KIC.","pith_inferences":["If the unification holds, the practical choice between OLS, ridge, and LASSO becomes a choice of prior belief about coefficients, which gives a principled way to set penalties through prior widths rather than by cross-validation alone.","The same completing-the-square machinery extends to any prior whose log-density is quadratic in the coefficients, suggesting a general recipe for building new shrinkage estimators; the paper does not pursue this generalization.","For the Laplace prior, a fully Bayesian treatment would keep the non-Gaussian posterior, whose mode is the usual LASSO estimator; a corrected version would replace the Gaussian approximation with direct optimization and would need the true posterior curvature for KIC.","A testable extension is to compare KIC computed from the Gaussian approximation against numerically integrated model evidence for small models; the paper does not include such a check."],"forward_implications":["Ordinary and weighted least squares emerge as maximum-likelihood or MAP estimates under a flat prior, so no separate theory is needed to justify them.","Ridge regression with penalty $\\lambda$ is the posterior mean of a Gaussian prior with covariance $\\lambda I$, making the regularization strength a prior variance.","Under the paper's sign-constancy assumption, LASSO is the MAP estimate under a Laplace prior, given as a shifted least-squares solution with a sign-dependent correction.","The Kashyap information criterion, derived by Laplace approximation around the MAP, gives a consistent rule for selecting among linear Gaussian candidate models: choose the lowest KIC.","Because all estimators come from the same Bayesian recipe, a user can move between methods by changing only the prior, not the inference machinery."],"supporting_citations":[{"why":"defines ridge regression, which the paper derives as the MAP estimate under a Gaussian prior with covariance $\\lambda I$","marker":"Hoerl and Kennard [1970]"},{"why":"introduces LASSO, which the paper derives as the MAP estimate under a Laplace prior under the sign-constancy assumption","marker":"Tibshirani [1996]"},{"why":"introduces the Kashyap information criterion that the paper derives from a Laplace approximation to Bayesian model evidence","marker":"Kashyap [1982]"},{"why":"supplies the comparison showing KIC typically outperforms BIC and AIC for linear Gaussian models, motivating the model-selection section","marker":"Schöniger et al. [2014]"},{"why":"uses the derived formulas in a sparse polynomial chaos expansion algorithm, showing the derivations support an actual estimation procedure","marker":"Shao et al. [2017]"}],"fun_headline_variants":["Every regression estimator is a different prior","Bayesian unification: OLS, ridge, LASSO as prior choices","OLS, ridge, LASSO, and KIC all from one Bayesian derivation","KIC falls out of the same quadratic derivation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The LASSO derivation assumes that the sign of each shifted coefficient $a-a_0$ stays fixed across the posterior and that the posterior is then Gaussian, an assumption the author says must be checked after computing the estimate; if it fails, the closed-form LASSO solution and the KIC built on it are not the true Bayesian MAP.","fun_headline_variants_meta":{"raw":{"variants":["Every regression estimator is a different prior","Bayesian unification: OLS, ridge, LASSO as prior choices","OLS, ridge, LASSO, and KIC all from one Bayesian derivation","KIC falls out of the same quadratic derivation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001238,"raw_usage":{"total_tokens":4990,"prompt_tokens":758,"completion_tokens":4232,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":374,"completion_tokens_details":{"reasoning_tokens":4164}},"tokens_in":374,"tokens_out":4232,"duration_ms":27956,"temperature":1.0,"reasoning_tokens":4164,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:17:16.242005+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a small dataset, set $a_0=0$ and $\\Lambda_{aa}=\\lambda I$, then compare the closed-form $\\hat{a}$ from Eq. (16) with the mode obtained by directly maximizing the posterior density in Eq. (14) on a fine grid for one or two coefficients; if the sign of $\\hat{a}-a_0$ does not match the sign of the posterior mode, the sign-constancy assumption fails and Eq. (16) is not the LASSO MAP.","supporting_citations":[],"review_version":1}