{"id":"66fe4043-0ee6-429e-b851-f5b1d4756fba","arxiv_id":"2607.14064","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Deep ReLU maximum-likelihood estimators are near-minimax optimal for speckle regression, achieving the same rates as additive-noise regression up to log factors.","lead":"This paper proves that likelihood-trained deep networks can estimate signals from multiplicative speckle noise at near-optimal rates, matching ordinary additive-noise regression. It supplies finite-sample guarantees and matching lower bounds for both low- and sparse high-dimensional problems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Upper bounds are proven only for σ_τ=0; the asserted transfer to σ_τ>0 is an unproved bridge for the paper's main model.","rationale":"The reader's conditional verdict is supported: the single most load-bearing gap is the σ_τ transfer. The paper's abstract and main theorems claim results for the model (1) with both multiplicative and additive Gaussian noise, but the entire upper-bound proof is carried out for σ_τ=0. The text itself flags this as an asserted rather than demonstrated transfer. The concern is a genuine proof gap, not a demonstrated falsehood: the pieces look fixable because σ_τ is constant and f* is uniformly bounded away from zero, so the variance model remains well-conditioned and the square-root map is Lipschitz on the relevant range. But a mathematical theorem requires the proof to be supplied; as it stands, Theorems 1 and 2 are not fully proven for the stated model. I therefore see no reason to change the reader's CONDITIONAL verdict. I also credit the lower-bound construction and the extensive σ_τ=0 analysis as real evidence that the claimed rates are plausible; the objection is specifically that the general-σ transfer is unproved.","tokens_in":32732,"tokens_out":18773,"duration_ms":176149,"concrete_test":"Re-derive Theorem 1 with y_i ∼ N(0, (f*(x_i))² + σ_τ²), σ_τ>0 fixed, keeping loss (2) (and separately loss (8)). The decisive check is whether Lemma 1 can approximate h(x)=sqrt((f*(x))²+σ_τ²) at rate N^{−2γ*} and whether inequality (39) holds with h replacing f*. If both hold, the transfer is valid; if not, Theorems 1–2 for σ_τ>0 are unproven.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV explicitly reduces the analysis to the pure speckle model (9) and states that for general σ_τ the proof 'remains essentially the same,' with f̂² consistent for f*²+σ_τ² and f* recovered via sqrt(f̂²−σ_τ²). However, Appendices A and B and all supporting lemmas are written for model (9) with loss (8); there is no displayed argument for σ_τ>0. The gap is not merely cosmetic: for σ>0 the full likelihood (2) and the simplified loss (8) are different objectives; the target of (8) becomes h=sqrt(f*²+σ_τ²), so Lemma 1's approximation guarantee must be transferred from f* to h, and the expected-likelihood lower bound (39) must be re-derived with h in place of f*. The paper does not show h belongs to H(d,l,P) or has the same DNN approximation rate N^{−2γ*}, nor does it prove the square-root map preserves the ∥·∥_n rate. Since Theorems 1 and 2 are stated for the general model (1), failure of any of these steps would leave the upper bounds unsupported; the lower bound alone would not establish the claimed minimax equivalence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies nonparametric regression under a model with multiplicative Gaussian speckle noise and additive Gaussian noise: y_i = f*(x_i) ξ_i + τ_i, with ξ_i, τ_i independent zero-mean Gaussian noises. Because E[y_i|x_i]=0, ordinary least-squares regression is inapplicable; the authors instead minimize a negative log-likelihood over deep ReLU networks, in both low-dimensional and sparse high-dimensional settings. The main claims are finite-sample upper bounds of order Õ(n^{-2γ*/(2γ*+1)}) on the squared estimation error, where γ* is a dimension-adjusted smoothness parameter of a hierarchical composition class H(d,l,P), and a matching lower bound of order Ω(n^{-2γ*/(2γ*+1)}), so that the method is near-minimax and the difficulty is comparable to additive-noise nonparametric regression. Numerical experiments illustrate the behavior of the estimator and the ability of the sparse version to detect active coordinates.","tokens_in":33044,"tokens_out":17532,"duration_ms":151125,"significance":"If correct, this is the first minimax theoretical analysis of likelihood-based DNN estimators for speckle regression in multivariate and sparse settings, extending the one-dimensional results of Malekian et al. (2025). The paper contains a substantial σ_τ=0 upper-bound proof based on empirical process theory and Hanson–Wright inequalities, and a detailed multidimensional lower-bound construction in Appendix D that appears carefully executed. It is also transparent in Remark 5 that the general-σ_τ case is not fully treated. However, the central theorems as stated cover the general model (1), and the proofs do not; the asserted transfer from σ_τ=0 to σ_τ>0 is a load-bearing gap rather than a cosmetic omission.","major_comments":[{"comment":"Theorems 1–2 are stated for model (1) with known σ_τ=O(1) and for the full likelihood (2), but all displayed proofs are for the pure speckle model (9) with loss (8). Section IV asserts that for general σ_τ the proof 'remains essentially the same' because the optimizer of (8) estimates h=sqrt(f*^2+σ_τ^2), from which f* is recovered by sqrt. This transfer is not proved and is nontrivial: (i) the theorem's estimator minimizes (2), not (8), and the relation between the two optimizers is not established; (ii) h is not shown to lie in H(d,l,P) or to admit the same N^{-2γ*} DNN approximation rate; (iii) the expected-likelihood lower bound (39) is derived for (8)/(9), not for (2)/(1); (iv) the square-root map is not shown to preserve the ∥·∥_n rate, and pointwise feasibility f̂^2≥σ_τ^2 is not guaranteed. Since Theorems 1–2 are the main upper-bound claims for the general model, this gap is load-b","section":"Section IV and Appendices A–B"},{"comment":"The proof of Theorem 2 drops the term λ|J| without a condition on |J|. From (63), the bound contains λ|J|; with λ≍log(ndN)/n, this is of order |J|logn/n. The displayed rate (64) requires |J| ≲ n^{1/(2γ*+1)} up to logarithms. The theorem statement only says |J| << d, which does not ensure this when d≤n^{c0}. The statement should either assume |J| is bounded (or specify the required growth condition) or include a |J|-dependent term in the rate. As written, the proof does not close for diverging |J|.","section":"Appendix B, Eqs. (63)–(64)"}],"minor_comments":[{"comment":"The notation ψ_τ(x)=|x|^τ ∧1 reuses τ, which also denotes the additive-noise variance σ_τ, and the displayed formula for λ appears to be missing parentheses: it should presumably read λ=c_2(log(ndN)+Llog(BN))/n. Please clarify.","section":"Theorem 2 statement"},{"comment":"The reduction to the class of (β*,C)-smooth d*-variate functions uses, without proof, the inclusion of this class in H(|J|,l,P) for l≥2. The inclusion is true by representing coordinate projections as elements of H(|J|,1,P), but it should be stated and proved.","section":"Appendix D, start of proof"},{"comment":"The phrase 'satisfying (3)' in Definition 4 appears to refer to Definition 3, not to equation (3). The cross-reference should be corrected.","section":"Definition 4, p. 6"},{"comment":"The figure captions use 'Test MSE' while the text describes the mean squared error relative to the true generating function. Please align the terminology so that it is clear whether the plotted quantity is an oracle MSE or a held-out-data MSE.","section":"Section III-B, Fig. 1"}],"recommendation":"major_revision","confidential_remarks":"The main load-bearing issue is the unproved transfer from σ_τ=0 to σ_τ>0. The lower bound and the σ_τ=0 upper bound are substantial and appear worth publishing; however, Theorems 1–2 as stated are not yet supported. The sparse theorem also needs a precise condition on |J|. I recommend major revision, with attention to supplying the transfer argument or restating the theorems for σ_τ=0 and moving the general-σ_τ claim to a conjecture."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know the main thing before anything else: this paper has a real result, but its headline theorems outrun the proof. The minimax rate for likelihood-based DNN despeckling—matching additive-noise rates up to log factors—is a legitimate contribution, and the multidimensional lower bound is genuinely new. But Theorem 1 and Theorem 2 are stated for the full model with multiplicative plus additive noise, and the proof is carried out only for the pure speckle case σ_τ = 0. The transfer to σ_τ > 0 is asserted in Section IV, not proved. I checked the stress-test note about this and it holds up: the full likelihood (2) and the simplified loss (8) are different objectives, and the argument would need h = sqrt(f*^2 + σ_τ^2) to inherit the same smoothness class and the same DNN approximation rate, plus a square-root stability result for the empirical L2 norm. None of that is shown. As written, the upper bounds are unsupported for the main model.\n\nWhat is actually new and good: first multidimensional minimax lower bound for Gaussian speckle regression, first DNN likelihood upper bounds for this model, and a sparse high-dimensional extension with an l1-penalized projection layer. The machinery is substantial and mostly credible—the likelihood expansion, the Hanson-Wright based empirical process control, the covering number bounds, and the peeling argument for the sparse case are all laid out in detail. The lower bound uses the standard hypercube construction and Berry-Esseen; it looks sound, and the σ_τ dependence is handled inside the likelihood ratio. The Fan–Gu approximation lemma is imported cleanly, and the self-citations are background, not load-bearing.\n\nSoft spots, in proportion: the σ_τ transfer is the one that matters. I think it is fixable—if f* is bounded away from zero, h should be at least as smooth as f*^2 and the square-root step is likely a Lipschitz perturbation—but it has to be written down, and the likelihood lower bound (39) must be re-derived with h. Minor points: the numerical experiments show consistency but no baselines, no code, and no comparison to existing despeckling methods; the paper's own Remarks 5 and 8 admit the σ_τ restriction, which is honest. The log factors in the upper bounds are standard DNN artifacts.\n\nBottom line: this is a solid contribution to the minimax theory of deep learning under multiplicative noise, and the rate identity—speckle is no harder than additive noise up to logs—is well supported for σ_τ = 0. It deserves a serious referee. I would send it to peer review with the expectation that the σ_τ transfer be made rigorous or the theorems explicitly restricted.","headline":"The paper's central rate claims are probably right, but Theorems 1–2 are only proved for σ_τ = 0; the asserted transfer to σ_τ > 0 is a genuine gap that a referee should insist on closing.","tokens_in":33461,"tokens_out":2159,"would_cite":true,"duration_ms":24511,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G20","62C20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that likelihood-based deep neural networks estimate signals corrupted by multiplicative speckle noise at the minimax-optimal rate, matching additive-noise regression up to logarithmic factors.","keywords":["speckle noise","nonparametric regression","deep neural networks","minimax optimality","likelihood-based estimation","sparse high-dimensional regression","ReLU networks"],"falsifier":"Run the likelihood estimator on data generated from y = f*(x)ξ + τ with a fixed σ_τ>0 and a smooth f*, and measure the empirical MSE as n grows. If the error does not decay as (log n)^{2γ*/(2γ*+1)} n^{−2γ*/(2γ*+1)}—e.g., it plateaus or decays slower—the σ_τ-transfer claim fails. Alternatively, a direct mathematical check of whether the proof's lower bound on ℓ̄(f)−ℓ̄(f*) holds with (f*)²+σ_τ² in the denominator would settle it.","tokens_in":32622,"feed_emoji":"📡","tokens_out":6149,"duration_ms":67664,"temperature":0.7,"pith_summary":"Speckle noise multiplies the signal, so the regression function is invisible to the usual least-squares loss: E[y|x]=0. This paper proposes a likelihood-based deep neural network estimator, using the negative log-likelihood with f^2+σ_τ² in the denominator, and proves that its mean squared error is, up to logarithmic factors, n^{-2γ*/(2γ*+1)} for smooth hierarchical functions in both low-dimensional and sparse high-dimensional settings. A matching lower bound shows the estimator is minimax optimal up to logs. The rates coincide with additive-noise nonparametric regression, so multiplicative speckle does not add statistical difficulty—only the estimation strategy must change.","feed_headline":"DNN despeckling hits the optimal statistical rate","feed_subtitle":"Likelihood-based deep learning recovers signals hidden in speckle as efficiently as under additive Gaussian noise.","key_machinery":"The load-bearing object is the negative log-likelihood ℓ(f) = (1/n) Σ_i [ y_i²/(f(x_i)²+σ_τ²) + log(f(x_i)²+σ_τ²) ]. Its expectation difference ℓ̄(f)−ℓ̄(f*) equals (1/n) Σ_i [(f*/f)² − 1 − log((f*/f)²)], which is lower-bounded by a constant times Σ_i (f*−f)² via the inequality x−1−log x ≥ a(x−1)² for bounded x. This converts the MLE's likelihood inequality into a squared-error bound. The random fluctuation is controlled with the Hanson-Wright inequality on the quadratic form y^T(Σ* − Σ_f)y, where Σ_f = diag(1/f(x_i)²); a covering-number bound for the ReLU network class (width N ≈ (n/log n)^{1/(4γ*+2)}) lets the argument go through without sample splitting. For sparse high-dimensional feature","core_discovery":"The paper's central claim is that estimating a signal f* from observations y_i = f*(x_i)ξ_i + τ_i — speckle noise (Gaussian multiplier) plus additive Gaussian noise — via maximum likelihood with a deep ReLU network is minimax optimal. For f* in a hierarchical smooth function class whose hardest composition has dimension-adjusted smoothness γ* = β*/d* > 1/2, the likelihood-based estimator satisfies, with high probability, ||f̂−f*||_n² ≤ c (log n)^{2γ*/(2γ*+1)} n^{−2γ*/(2γ*+1)} (Theorem 1). The same rate holds in sparse high-dimensional settings through an ℓ1-penalized projection layer (Theorem 2), and a matching lower bound Ω(n^{−2γ*/(2γ*+1)}) (Theorem 3) shows the estimator is within logarit","pith_inferences":["If σ_τ is unknown, the model identifies only f*²+σ_τ²; one could add a smoothness prior on f* or a bounded-variance assumption to estimate σ_τ and then recover f* by the same square-root step.","The proof's σ_τ=0-to-σ_τ>0 transfer is asserted rather than demonstrated; a dedicated numerical or theoretical check with σ_τ>0 would either confirm the rate or reveal where the argument needs strengthening.","Because the lower bound uses bump functions with disjoint supports, the result likely extends to any fixed design satisfying a Riemann-sum condition, not just grids; this could be verified directly with random designs via the paper's Proposition 1."],"forward_implications":["Despeckling with deep networks is provably near-optimal: the DNN estimator matches the minimax lower bound up to log factors, so practitioners can use likelihood-based DNNs without paying a statistical penalty.","In sparse high-dimensional settings, the same rate is achieved automatically: the ℓ1-projection recovers active coordinates, so the curse of dimensionality is avoided when the signal depends on few features.","Multiplicative speckle noise does not make function recovery fundamentally harder than additive Gaussian noise: the rates agree, so only the loss function needs to change, not the sample complexity.","The logarithmic slack is linked to the DNN analysis; the author suggests classical nonparametric estimators may remove the logs, which would pin down the exact minimax rate."],"fun_headline_variants":["Deep likelihood nets hit minimax rate for speckle regression","Likelihood-based deep nets: optimal rates despite multiplicative noise","Speckle regression: deep likelihood attains optimal error rates","Minimax-optimal despeckling via likelihood-based deep nets"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper depends on knowing the additive noise level σ_τ and on the assertion—stated but not fully proved in Section IV—that the theory for pure speckle (σ_τ=0) carries over to σ_τ>0 by estimating f* as the square root of f̂²−σ_τ².","fun_headline_variants_meta":{"raw":{"variants":["Deep likelihood nets hit minimax rate for speckle regression","Likelihood-based deep nets: optimal rates despite multiplicative noise","Speckle regression: deep likelihood attains optimal error rates","Minimax-optimal despeckling via likelihood-based deep nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000542,"raw_usage":{"total_tokens":2465,"prompt_tokens":807,"completion_tokens":1658,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":1587}},"tokens_in":551,"tokens_out":1658,"duration_ms":13530,"temperature":1.0,"reasoning_tokens":1587,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:52:46.405882+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the likelihood estimator on data generated from y = f*(x)ξ + τ with a fixed σ_τ>0 and a smooth f*, and measure the empirical MSE as n grows. If the error does not decay as (log n)^{2γ*/(2γ*+1)} n^{−2γ*/(2γ*+1)}—e.g., it plateaus or decays slower—the σ_τ-transfer claim fails. Alternatively, a direct mathematical check of whether the proof's lower bound on ℓ̄(f)−ℓ̄(f*) holds with (f*)²+σ_τ² in the denominator would settle it.","supporting_citations":[],"review_version":1}