{"id":"14e6272d-6890-42b5-999d-b6540b96b940","arxiv_id":"2505.22211","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A sparse Bayesian Beta regression method is proposed, but its Gibbs sampler does not target the Beta model and its theoretical results are not proven.","lead":"This paper proposes a high-dimensional Bayesian Beta regression model using a tempered posterior and a Horseshoe prior, and adds a Polya-Gamma Gibbs sampler and claimed posterior concentration guarantees. The core sampler is invalid: it targets a quasi-binomial likelihood rather than the Beta likelihood, and the theory is stated without proof.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §3 Gibbs sampler misapplies the Pólya–Gamma identity to the Beta likelihood; the true Beta density is not proportional to (e^η)^{φ y}/(1+e^η)^φ, so the Gaussian full conditional does not target the stated posterior.","rationale":"The paper's central computational claim is an exact Pólya–Gamma Gibbs sampler for high-dimensional Beta regression. That claim depends entirely on rewriting the Beta log-likelihood as κ_i η_i − ω_i η_i²/2. This is the same step the reader identified as the weakest assumption, and it fails: the Beta density includes the beta normalizing constant and both y_i^{μ_i φ−1} and (1−y_i)^{(1−μ_i)φ−1}; substituting μ_i=logit^{-1}(η_i) does not turn it into the binomial logistic kernel. The identity stated in §3 is correct in isolation but irrelevant. Because the sampler does not sample from the stated fractional posterior, the reported empirical results do not test the proposed method. The theoretical section is also incomplete, with Appendix A ending in a placeholder reference instead of proofs. Both issues independently justify rejecting the paper in its current form; the sampler mismatch is the more fundamental one.","tokens_in":18444,"tokens_out":8116,"duration_ms":85510,"concrete_test":"Verify the proportionality used at the start of §3: fix y_i=0.3, φ=10, and compute the true density f(y_i|η) at two values of η (e.g., η=0 and η=1, with μ=logit^{-1}(η)). The claimed PG augmentation requires f(y_i|η)(1+e^η)^φ/e^{φ y_i η} to be constant in η; the true Beta density gives different values because the Γ(μφ) and Γ((1−μ)φ) terms vary with η. Equivalently, compute ∂² log f(y_i|η)/∂η² for y_i∈(0,1); any dependence on η contradicts the constant −ω implied by the claimed full-conditional update. This calculation is one line and settles whether the sampler targets the stated model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Step 1 of Section 3 uses the identity (e^η)^{φ y_i}/(1+e^η)^φ ∝ exp(φ(y_i−1/2)η) ∫ exp(−ωη²/2)p(ω)dω with ω∼PG(φ,η). This identity is valid for a binomial logistic kernel, but it is not the Beta likelihood in Eq. (1). The true log-density is log f(y_i|η_i)=log Γ(φ)−log Γ(μ_i φ)−log Γ((1−μ_i)φ)+(μ_i φ−1)log y_i+((1−μ_i)φ−1)log(1−y_i), with μ_i=logit^{-1}(η_i). The β-dependent part is not quadratic in η_i: the digamma terms from the normalizing constant and the η-dependent coefficients multiplying log y_i and log(1−y_i) preclude a representation as κ_i η_i−ω η_i²/2+const. For example, ∂² log f(y_i|η_i)/∂η_i² is not a constant independent of η_i, as the claimed quadratic augmentation would require. Hence Step 2's multivariate-normal conditional for β is not the full conditional under the stated posterior; the sampler targets a different 'binomialized' model. The simulation studies and trace plots cannot validate the claimed Beta-regression posterior because they test the wrong target. Separately, Appendix A does not contain proofs: after Lemma 3 the text reads 'Proofs of Lemma 1, 2, and 3 are given at the end of the proof section, ??', so the theoretical claims are also currently unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a sparse Bayesian Beta regression model for high-dimensional bounded responses, using a fractional posterior with a Horseshoe prior. The authors claim two main contributions: a novel Gibbs sampling algorithm based on Polya-Gamma augmentation for posterior inference, and the first theoretical guarantees (posterior consistency and convergence rates) for Bayesian Beta regression. The manuscript includes simulation studies in low and high dimensions, a sensitivity analysis for the precision parameter, and a real-data application. The theoretical results are stated in Section 4 and rely on appendices for proofs.","tokens_in":18787,"tokens_out":2268,"duration_ms":24113,"significance":"The problem is practically relevant, and a correct sparse Bayesian Beta regression method with computable inference and theoretical guarantees would be a valuable addition. The paper, however, does not deliver these contributions as written. The central computational claim is based on an incorrect likelihood identity, and the theoretical results are not supported because the required lemmas are unproved and the appendix contains a placeholder instead of proofs. If the computational error were isolated, the method could be salvaged by using a different sampler, but that would change the paper's central novelty. The paper is therefore not publishable in its current form.","major_comments":[{"comment":"The Polya-Gamma identity used in Step 1, (e^{η})^{φ y}/(1+e^{η})^φ ∝ exp(κ_i η) ∫ exp(−ω η²/2)p(ω)dω with κ_i = φ(y_i − 1/2), is valid for a binomial logistic kernel but is not the Beta likelihood in Eq. (1). The Beta log-density contains the log-normalizer log Γ(φ) − log Γ(μ_i φ) − log Γ((1−μ_i)φ) and the terms (μ_i φ −1) log y_i + ((1−μ_i)φ−1) log(1−y_i), where μ_i depends on η_i through the logit link. The quadratic augmentation would require the log-density to be linear-quadratic in η_i, which it is not; for instance, the second derivative of the log-density with respect to η_i is not a constant independent of η_i. Consequently, the multivariate normal full conditional for β in Step 2 is not the full conditional under the stated posterior. The sampler targets a different, binomialized model, so the simulation results and trace plots cannot be interpreted as evidence about the Beta regression posterior the paper claims to fit.","section":"Section 3"},{"comment":"The proofs of Lemmas 1, 2, and 3 are missing; the text reads 'Proofs of Lemma 1, 2, and 3 are given at the end of the proof section, ??', with a literal placeholder. These lemmas are the technical core needed to verify the KL conditions in Theorem 2 of [1] that underpin Theorem 1. Lemma 1 in particular must control the KL divergence between two Beta distributions with different means, which requires careful handling of the Beta normalizing constant; it is not proved, so the main concentration result is currently unsupported.","section":"Appendix A"},{"comment":"Even granting Lemmas 1–3, the proof of Theorem 1 is not supplied. The paper cites Theorems 2 and 3 from Alquier and Ridgway [1] and Lemma 4 from the author's prior work [24], but it does not show that the required conditions hold for the Beta regression model. Specifically, it does not verify the existence of a distribution ρ_n satisfying ∫ KL(P_{β0}, P_β) ρ_n(dβ) ≤ ε_n and KL(ρ_n, π) ≤ n ε_n for the Beta likelihood with horseshoe prior, nor does it verify the second-moment condition needed for the high-probability bound in inequality (5). The unproved lemmas are not merely technical; they are exactly the steps that connect the prior mass bound of Lemma 4 to the KL divergence of the Beta model. The theoretical claims are therefore not established in the manuscript.","section":"Section 4"}],"minor_comments":[{"comment":"The manuscript contains numerous typos and grammatical errors, including 'hierachical', 'conducte', 'P´olya' spacing, and inconsistent use of 'Cau+' for the half-Cauchy distribution. A careful copyedit is needed.","section":"Throughout"},{"comment":"In the n = 500, ρ_X = 0.5 block, the FDR for the Horseshoe method is reported as 0.99 (0.03), which appears to be an error; this value is inconsistent with the reported Precision and Specificity and should be corrected.","section":"Table 1"},{"comment":"The test set size is fixed at n_test = 30 regardless of training sample size; for n = 100 this is a sizable fraction, and for n = 1000 it is very small. Reporting results for a fixed test size makes comparisons across settings less informative.","section":"Section 5.1"},{"comment":"The claim that the Horseshoe method 'consistently' outperforms the Lasso is too strong for the s* = 20, n = 80 cases, where the Horseshoe recall is low (0.43 and 0.21) and the test prediction error is comparable to or worse than Lasso in some settings.","section":"Section 5.2"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an incomplete draft: Appendix A contains a literal '??' placeholder instead of proofs, and the central sampler is based on a likelihood identity that does not hold for the Beta distribution. The paper also relies heavily on the author's own prior work and on external theorems without verification. These issues go beyond what a normal revision could fix within the paper's stated scope, so I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's central selling points—a Polya–Gamma Gibbs sampler for Beta regression and the first concentration rates—do not hold up. The sampler step is wrong. The Beta likelihood in Eq. (1) is not proportional to (e^η)^{φy}/(1+e^η)^φ; the gamma-function normalizing constant and the µ-dependent coefficients on log y and log(1−y) are not quadratic in η. So the posterior the sampler explores is not the stated Beta-regression posterior. That is a load-bearing flaw: every simulation and trace plot in Section 5 tests a different model. I don't think there is any way to reinterpret the numerical results as evidence for the claimed method.\n\nThe paper does some things well. The gap it targets—high-dimensional sparse Bayesian Beta regression without guarantees—is real, and the combination of a fractional posterior and a Horseshoe prior is a reasonable place to start. The writing is clear and the literature review is competent. If the theory were actually proven, it would be a useful addition.\n\nBut the theory is not proven. Theorem 1 is a direct appeal to Alquier and Ridgway, yet the KL conditions required to apply their theorem are never verified. The lemmas in Appendix A are stated without proofs, and after Lemma 3 the text literally reads 'Proofs of Lemma 1, 2, and 3 are given at the end of the proof section, ??'. That placeholder is a strong signal the manuscript is incomplete. Lemma 4 comes from the author's earlier paper, which is fine as a citation, but the rest of the derivation is absent.\n\nThere are smaller issues: φ is treated as known, and the sensitivity analysis is the only workaround; the real-data example is a bit thin. Those would be minor if the core were sound.\n\nBottom line: this needs to be rejected. It is not ready for referees to spend time on, because the main algorithm is invalid and the theory is a sketch. If the author replaces the sampler with a correct one targeting the true Beta posterior and completes the proofs, the underlying idea might be worth a fresh look. But the current version does not deserve a full referee cycle.","headline":"The paper's PG-augmented Gibbs sampler is built on a false identity, so it samples a different model, and the theory is an unproven sketch; the work should be rejected.","tokens_in":19271,"tokens_out":3489,"would_cite":false,"duration_ms":35326,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J07","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to build the first Gibbs sampling algorithm for high-dimensional sparse Beta regression, using Polya-Gamma augmentation and a Horseshoe prior, together with the first posterior consistency and concentration-rate results…","keywords":["Beta regression","Horseshoe prior","Polya-Gamma augmentation","Gibbs sampler","high-dimensional sparsity","fractional posterior","posterior concentration","bounded response"],"falsifier":"Choose one observation with fixed $\\varphi$, a grid of $\\eta$ values, and a fixed $y\\in(0,1)$; compute the true Beta likelihood $f(y;\\mu(\\eta),\\varphi)$ and compare it with the marginal $\\int \\operatorname{PG}(\\omega;\\varphi,\\eta)\\exp(\\kappa\\eta-\\omega\\eta^2/2)\\,d\\omega$ over the grid. If the ratio of the two quantities changes with $\\eta$, the augmentation does not reproduce the Beta likelihood and the sampler does not target the stated posterior.","tokens_in":18189,"feed_emoji":"📊","tokens_out":10529,"duration_ms":106006,"temperature":0.7,"pith_summary":"The paper aims to make Bayesian Beta regression work in high dimensions, where the number of predictors exceeds the sample size. It proposes a fractional posterior combined with a Horseshoe prior, and claims the first Gibbs sampling algorithm for Beta regression by invoking Polya-Gamma augmentation to rewrite the log-likelihood in a quadratic form. It also claims the first theoretical guarantees for Bayesian Beta regression: posterior consistency and a concentration rate $\\varepsilon_n = K s^*\\log(p/s^*)/n$ in $\\alpha$-Rényi divergence, with the bound remaining meaningful when $p$ grows faster than $n$. If these claims are right, proportion-type outcomes could be analyzed with a conjugate sampler in sparse settings where classical Beta regression cannot be fitted at all.","feed_headline":"First Gibbs sampler for sparse Beta regression","feed_subtitle":"Polya-Gamma augmentation and a Horseshoe prior target p > n proportion data with concentration guarantees.","key_machinery":"The load-bearing object is the Polya-Gamma augmentation identity for a logistic kernel: $\\frac{(e^{\\eta})^{\\varphi y}}{(1+e^{\\eta})^{\\varphi}} \\propto e^{\\kappa\\eta}\\int_0^\\infty e^{-\\omega\\eta^2/2}p(\\omega)\\,d\\omega$ with $\\kappa = \\varphi(y-\\tfrac12)$. The paper applies this identity to the Beta regression likelihood to obtain a conditionally Gaussian form for each observation, which in turn yields a multivariate normal full conditional for $\\beta$. The Horseshoe prior enters through its hierarchical inverse-gamma representation, giving conjugate updates for all shrinkage variances, and the fractional posterior $\\pi_{n,\\alpha}(\\beta)\\propto L_n(\\beta)^\\alpha\\pi(\\beta)$ supplies the theoretical framework in which the concentration proof is carried out.","core_discovery":"The paper's central claim is that a Beta regression likelihood, written with $\\mu_i = (1+e^{-\\eta_i})^{-1}$ and precision $\\varphi$, can be transformed into a Gaussian form in $\\eta_i$ through the Polya-Gamma identity: $\\log p(y_i \\mid \\eta_i) \\propto \\kappa_i\\eta_i - \\omega_i\\eta_i^2/2$ with $\\kappa_i = \\varphi(y_i - 1/2)$. This makes the full conditional of the regression vector $\\beta$ multivariate normal, so inference can proceed by a closed-form Gibbs sampler alternating between Polya-Gamma draws, Gaussian draws for $\\beta$, and inverse-gamma draws for the Horseshoe scale parameters. The paper further claims that, with the tempered likelihood $L_n(\\beta)^\\alpha$ and the Horseshoe prior, the fractional posterior concentrates around $\\beta_0$ at rate $\\varepsilon_n = K s^*\\log(p/s^*)/n$, measured in $\\alpha$-Rényi divergence, Hellinger distance, and total variation; the posterior mean estimator satisfies the same rate in prediction loss $\\|X^\\top(\\hat\\beta-\\beta_0)\\|_2^2$.","pith_inferences":["If the Polya-Gamma rewrite is exact, the same augmentation strategy would plausibly extend to other unit-interval response models whose mean is linked through the logit function, not just the Beta family.","The proof structure, which bounds divergence by squared linear-predictor distance, suggests the concentration result may carry over to neighboring bounded-response families with smooth densities, conditional on a similar prior-mass lemma.","A natural testable extension is a zero-one inflated version of the sampler, where a point-mass mixture component is added while the Polya-Gamma augmentation handles the continuous part of the response."],"forward_implications":["Beta regression becomes usable in $p>n$ regimes: the method estimates coefficients and selects variables when the number of covariates exceeds the sample size, a setting where classical maximum-likelihood Beta regression cannot even be fitted.","If the sampler is exact, it replaces Metropolis-Hastings updates for Beta regression with closed-form conditional draws, making posterior inference for bounded responses substantially faster.","The claimed concentration bound says the fractional posterior learns the true coefficient vector at near-minimax rate $s^*\\log(p/s^*)/n$ with probability at least $1-2/(n\\varepsilon_n)$.","The numerical comparisons claim the Horseshoe-Beta method reduces estimation error and improves variable-selection precision relative to transformed Lasso in both low- and high-dimensional simulations."],"supporting_citations":[{"why":"Supplies the Polya-Gamma identity that the sampler's Gaussian full-conditional step relies on.","marker":"[38]"},{"why":"Introduces the Horseshoe prior used for shrinkage and variable selection.","marker":"[8]"},{"why":"Provides the hierarchical inverse-gamma sampler for the Horseshoe prior that the update steps implement.","marker":"[30]"},{"why":"Defines fractional posteriors and gives the general concentration framework the theory builds on.","marker":"[6]"},{"why":"Gives the tempered-posterior concentration theorems used to prove the Renyi-divergence rates.","marker":"[1]"},{"why":"Supplies the prior-mass lower bound on a sparse neighborhood of the true coefficient vector used in the proof.","marker":"[24]"},{"why":"Establishes optimal posterior concentration rates for the Horseshoe, motivating its use as a sparse prior.","marker":"[40]"},{"why":"Provides the maximum-likelihood Beta regression baseline used in low-dimensional numerical comparisons.","marker":"[12]"}],"fun_headline_variants":["High-dim Beta regression: Gibbs sampler with Horseshoe prior","Sparse Beta regression gets fast Gibbs sampler","Polya-Gamma Gibbs for high-dim Beta regression","New Gibbs sampler for sparse Beta regression in p>n","Beta regression: first Gibbs sampler with concentration guarantees"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The sampler is exact only if each Beta observation's likelihood equals, up to a constant independent of the linear predictor, the logistic form $(e^{\\eta})^{\\varphi y}/(1+e^{\\eta})^{\\varphi}$; the actual Beta density depends on $\\eta$ through $\\mu=(1+e^{-\\eta})^{-1}$ in the shape parameters and through the $\\beta$ normalizing constant, so this equivalence is the premise on which the whole Gibbs scheme rests.","fun_headline_variants_meta":{"raw":{"variants":["High-dim Beta regression: Gibbs sampler with Horseshoe prior","Sparse Beta regression gets fast Gibbs sampler","Polya-Gamma Gibbs for high-dim Beta regression","New Gibbs sampler for sparse Beta regression in p>n","Beta regression: first Gibbs sampler with concentration guarantees"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000305,"raw_usage":{"total_tokens":1759,"prompt_tokens":964,"completion_tokens":795,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":720}},"tokens_in":580,"tokens_out":795,"duration_ms":6351,"temperature":1.0,"reasoning_tokens":720,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:13:14.225471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Choose one observation with fixed $\\varphi$, a grid of $\\eta$ values, and a fixed $y\\in(0,1)$; compute the true Beta likelihood $f(y;\\mu(\\eta),\\varphi)$ and compare it with the marginal $\\int \\operatorname{PG}(\\omega;\\varphi,\\eta)\\exp(\\kappa\\eta-\\omega\\eta^2/2)\\,d\\omega$ over the grid. If the ratio of the two quantities changes with $\\eta$, the augmentation does not reproduce the Beta likelihood and the sampler does not target the stated posterior.","supporting_citations":[{"cited_title":"G., Scott, J","cited_arxiv_id":null,"evidence_quote":"Supplies the Polya-Gamma identity that the sampler's Gaussian full-conditional step relies on."},{"cited_title":"M., Polson, N","cited_arxiv_id":null,"evidence_quote":"Introduces the Horseshoe prior used for shrinkage and variable selection."},{"cited_title":"and Schmidt, D","cited_arxiv_id":null,"evidence_quote":"Provides the hierarchical inverse-gamma sampler for the Horseshoe prior that the update steps implement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines fractional posteriors and gives the general concentration framework the theory builds on."},{"cited_title":"and Ridgway, J","cited_arxiv_id":null,"evidence_quote":"Gives the tempered-posterior concentration theorems used to prove the Renyi-divergence rates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the prior-mass lower bound on a sparse neighborhood of the true coefficient vector used in the proof."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes optimal posterior concentration rates for the Horseshoe, motivating its use as a sparse prior."},{"cited_title":"and Zeileis, A","cited_arxiv_id":null,"evidence_quote":"Provides the maximum-likelihood Beta regression baseline used in low-dimensional numerical comparisons."}],"review_version":1}