{"id":"216aaf5d-a730-4427-864a-f52ce129faa6","arxiv_id":"2606.20480","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Priors with p-exponential tails (p<1) yield improved posterior contraction rates and near-adaptive smoothness recovery in white noise and random design regression, including overparameterized ReLU networks up to beta=2.","lead":"The paper shows that Bayesian posteriors contract faster when priors on function coefficients have heavier tails (smaller p in p-exponential distributions), achieving near-full adaptation to unknown smoothness in nonparametric regression and shallow ReLU networks. A smart generalist might read it because it suggests a simple prior tweak can make Bayesian methods work better for flexible models like neural nets without manual tuning.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"NN application relies on modeling as independent coefficients over a fixed dictionary, but overparametrized ReLU networks use continuous parameter priors instead","rationale":"Reader correctly flags the fixed-dictionary modeling choice as the weakest link; the NN claim is the place where that assumption is most strained because the dictionary is not a priori fixed. The concern is internal to the argument rather than external consensus, so a targeted check on the reduction step would settle it without rejecting the series-prior results.","tokens_in":1639,"tokens_out":389,"duration_ms":15879,"concrete_test":"Extract the precise prior definition and proof sketch for the ReLU case (likely §4 or §5); check whether the argument invokes a fixed finite dictionary obtained by gridding the hidden weights or directly analyzes the continuous-parameter posterior; recompute the entropy integral or prior mass lower bound under the continuous parameterization and verify it matches the p-exponential tail assumption used for the series case.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The core contraction results (Theorems on posterior rates for p-exponential tails) are proved under independent coefficient priors on a fixed, countable dictionary with controlled metric entropy. The shallow ReLU claim requires representing the network as a linear combination over a dictionary of ReLU atoms; however, in the overparametrized regime the atoms are indexed by continuous hidden weights, so the prior is placed on the weight vectors rather than on fixed coefficients. This introduces an extra layer of approximation (discretization or covering of the parameter space) whose entropy and prior mass must be controlled uniformly in the unknown smoothness β. The abstract states the result holds for 0 ≤ β ≤ 2, but without an explicit reduction showing that the continuous-parameter prior satisfies the same tail and entropy conditions used in the fixed-dictionary theorems, the adaptation statement for networks rests on an unverified modeling equivalence.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.3","summary":"The manuscript studies Bayesian posterior contraction in nonparametric regression using priors with p-exponential tails (including Laplace and heavier tails for p<1) placed independently on coefficients of a fixed basis or dictionary. It claims that contraction rates improve as p decreases and that an appropriate p→0 regime yields full adaptation to unknown smoothness up to logarithmic factors. Applications are given to series priors in white noise regression and to shallow ReLU networks in random design regression, with the specific claim that overparametrized shallow ReLU networks adapt to any regularity 0≤β≤2. A simulation study is included to support the theory.","tokens_in":1828,"tokens_out":351,"duration_ms":15573,"significance":"If the results hold, the work would contribute to Bayesian nonparametrics by showing how tail choice alone can deliver adaptation without hyperpriors on smoothness. The neural-network application, if rigorously justified, would be notable for linking tail-based priors to overparametrized ReLU models. The empirical component provides concrete support for the predicted behavior.","major_comments":[{"comment":"The core contraction theorems are proved under independent coefficient priors on a fixed countable dictionary with controlled metric entropy. The claim that overparametrized shallow ReLU networks adapt to 0≤β≤2 requires an explicit reduction showing that the continuous-parameter prior on hidden weights satisfies the same tail and entropy conditions uniformly in the unknown β; without this reduction the adaptation statement for networks rests on an unverified modeling equivalence between the fixed-dictionary setting and the continuous-atom setting.","section":"Application to shallow ReLU networks (abstract and corresponding theorems)"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their careful reading and for identifying a key point regarding the rigor of the ReLU network application. We address the concern directly below and commit to a targeted revision that supplies the missing explicit reduction.","responses":[{"response":"We agree that the general theorems are stated for a fixed countable dictionary and that the network claim requires a separate verification step. In the revised manuscript we will insert a new subsection (and supporting appendix) that explicitly reduces the continuous-weight prior to the dictionary setting. Concretely, we will (i) derive the marginal prior on the effective coefficients obtained by integrating the p-exponential prior over the hidden weights, (ii) verify that this marginal satisfies the required p-exponential tail bound uniformly in β, and (iii) establish a uniform bound on the metric entropy of the resulting function class for β ∈ [0,2]. With these steps the adaptation statement will rest on a verified reduction rather than an implicit modeling equivalence.","revision_made":"yes","referee_comment":"[Application to shallow ReLU networks (abstract and corresponding theorems)] The core contraction theorems are proved under independent coefficient priors on a fixed countable dictionary with controlled metric entropy. The claim that overparametrized shallow ReLU networks adapt to 0≤β≤2 requires an explicit reduction showing that the continuous-parameter prior on hidden weights satisfies the same tail and entropy conditions uniformly in the unknown β; without this reduction the adaptation statement for networks rests on an unverified modeling equivalence between the fixed-dictionary setting and the continuous-atom setting."}],"tokens_in":1265,"tokens_out":337,"duration_ms":12149,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper's main point is that placing p-exponential tail priors on coefficients, with p less than 1, tightens posterior contraction rates compared to the p=1 Laplace case, and the p approaching 0 regime recovers full adaptation to unknown smoothness up to log factors. They work this out for series priors in white noise and then apply it to shallow ReLU networks in random design regression.\n\nThe new piece is the explicit improvement for heavier tails and the resulting adaptation statement for ReLU networks across 0 to 2 in regularity. The simulation study is a straightforward check that the predicted behavior shows up in finite samples. The arguments rest on standard contraction theory but extend the tail family in a direct way.\n\nOne spot to watch is the neural net application. The core theorems use independent coefficient priors on a fixed countable dictionary with controlled entropy. Overparametrized ReLU networks instead put priors on continuous weight vectors, so the reduction needs to control covering numbers and prior mass uniformly in the unknown beta. The abstract states the result holds, but the details of that step matter for whether the adaptation claim goes through cleanly.\n\nThis is for readers already working on Bayesian nonparametric contraction rates or adaptive priors for neural nets. Someone tracking that literature would find the rate calculations and the ReLU extension worth seeing. The thinking is clear and engages the existing results without obvious internal gaps.\n\nI would send it to a serious referee.","headline":"The paper shows p-exponential tails with p<1 improve contraction rates over Laplace and yield adaptation in the p to 0 limit, including a claim for overparametrized shallow ReLU nets up to beta=2.","tokens_in":2312,"tokens_out":379,"would_cite":false,"duration_ms":14025,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Priors with p-exponential tails on coefficients achieve improving contraction rates and full adaptation to smoothness as p goes to zero.","keywords":["Bayesian nonparametrics","posterior contraction","adaptation","heavy-tailed priors","ReLU neural networks","white noise model","Sobolev balls"],"falsifier":"A concrete counterexample would be a function in a Sobolev ball whose posterior fails to contract at the predicted adaptive rate when p is taken small, or where an overparametrized shallow ReLU network posterior does not adapt beyond regularity level 2.","tokens_in":2543,"feed_emoji":"","tokens_out":643,"duration_ms":33577,"temperature":0.7,"pith_summary":"This paper examines Bayesian posterior contraction in nonparametric settings where independent p-exponential tail priors are placed on coefficients of a fixed basis or dictionary. It establishes that contraction rates improve as the tail parameter p decreases, with full adaptation to unknown smoothness obtained up to logarithmic factors in the regime where p tends to zero. The results cover series priors in white noise regression and extend to overparametrized shallow ReLU networks in random design regression, where adaptation holds for any regularity level between 0 and 2. A reader would care because the prior tail choice directly controls the ability to adapt without prior knowledge of smoothness.","feed_headline":"Heavier tails improve Bayesian adaptation to unknown smoothness","feed_subtitle":"p-exponential tail priors on coefficients achieve full adaptation as p approaches zero, including for overparametrized ReLU networks.","key_machinery":"p-exponential tail priors placed independently on the coefficients of a fixed basis or dictionary","core_discovery":"Placing independent priors with p-exponential tails on the coefficients of a fixed basis or dictionary produces posterior contraction rates that improve as p decreases. In an appropriate regime as p tends to zero, this setup yields full adaptation to the unknown smoothness parameter, up to logarithmic factors. The same mechanism is applied to shallow ReLU neural networks, showing that overparametrized versions adapt to any regularity 0 ≤ β ≤ 2 in random design regression.","pith_inferences":["The tail-based mechanism could be examined in related nonparametric problems such as density estimation.","One might test whether the adaptation extends when the basis or dictionary is allowed to vary with sample size.","Practitioners could explore very small p values in software implementations to check empirical adaptation in moderate sample sizes."],"forward_implications":["Posterior contraction rates improve as the tail parameter p is decreased.","Full adaptation to unknown smoothness holds up to logarithmic factors when p approaches zero.","Overparametrized shallow ReLU networks adapt to any regularity 0 ≤ β ≤ 2.","Simulation studies show agreement between observed behavior and the predicted rates."],"fun_headline_variants":["p-exponential tails contract posteriors faster as p drops","As p approaches zero posteriors adapt to smoothness","Shallow ReLU nets adapt via p-exponential coefficient tails","Bayesian adaptation to smoothness via decreasing tail exponent p","Overparametrized ReLU adapts to beta via p tails"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The true function belongs to a Sobolev or Besov-type ball of unknown radius, and the priors are placed independently on coefficients of a fixed basis or dictionary.","fun_headline_variants_meta":{"raw":{"variants":["p-exponential tails contract posteriors faster as p drops","As p approaches zero posteriors adapt to smoothness","Shallow ReLU nets adapt via p-exponential coefficient tails","Bayesian adaptation to smoothness via decreasing tail exponent p","Overparametrized ReLU adapts to beta via p tails"]},"model":"grok-4.3","cost_usd":0.004494,"raw_usage":{"total_tokens":2197,"prompt_tokens":584,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":44937000,"prompt_tokens_details":{"text_tokens":584,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1535,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":584,"tokens_out":78,"duration_ms":12389,"temperature":1.0,"reasoning_tokens":1535,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T15:09:19.350332+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A concrete counterexample would be a function in a Sobolev ball whose posterior fails to contract at the predicted adaptive rate when p is taken small, or where an overparametrized shallow ReLU network posterior does not adapt beyond regularity level 2.","supporting_citations":[],"review_version":1}