REVIEW 3 major objections 7 minor 47 references
Sparse recovery rates for nonlinear inverse problems proven minimax-optimal
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-09 09:55 UTC pith:ANTEEP4Y
load-bearing objection Minimax-optimal rates for ℓ¹-regularized nonlinear statistical inverse learning; proofs are sound and the k_t-to-VSC bridge holds for the minimax claim as stated. the 3 major comments →
Statistical inverse learning and ell¹-regularization
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The minimax-optimal convergence rate for ℓ¹-regularized sparse statistical inverse learning with nonlinear forward operators is n^{-r/(1+b-br)} in ℓ¹ reconstruction norm, where r encodes the solution's sparsity-driven smoothness (via a variational source condition) and b encodes the effective dimension of the learning problem (via spectral decay of the covariance operator). This rate is achieved by the regularized estimator with parameter λ* = n^{-(1-r)/(1+b-br)} and cannot be improved by any estimator over the prior class P_{r,b}. A key structural finding is that the variational source condition — the analytical engine for the rates — is equivalent to membership in an approximation space kₜ
What carries the argument
The variational source condition (Assumption 5) with index function φ(t) = t^r, which quantifies how well the true solution can be approximated relative to the forward operator's sensitivity; the effective dimension N(λ) with polynomial bound N(λ) ≤ Cλ^{-b}, capturing statistical complexity via covariance spectral decay; the weighted bi-Lipschitz property (Assumption 6), which provides two-sided stability of the forward operator between weighted sequence space and data space; the approximation space kₘ
Load-bearing premise
The weighted bi-Lipschitz property (Assumption 6) requires the forward operator to satisfy a lower Lipschitz bound — meaning small changes in the data must reflect proportionally small changes in the solution. This fails for severely ill-posed problems like the backward heat equation or electrical impedance tomography, restricting the theory to finitely smoothing operators.
What would settle it
Construct a probability distribution in P_{r,b} for which the ℓ¹-regularized estimator with λ* = n^{-(1-r)/(1+b-br)} converges slower than n^{-r/(1+b-br)}, or exhibit any learning algorithm achieving a faster rate over the full class P_{r,b}.
If this is right
- For filtered Radon transforms with a boxcar filter, the effective dimension exponent is b = 2/3; with a Gaussian filter, eigenvalues decay super-polynomially so any b > 0 is admissible, yielding near-parametric rates n^{-r} for sufficiently sparse signals.
- For cartoon-like images represented in shearlet frames (best n-term approximation rate n^{-1}), the theory predicts r = 1/4, giving concrete convergence rates of n^{-1/6} (boxcar) or n^{-1/4} (Gaussian) in ℓ¹ reconstruction norm.
- The framework recovers classical kernel ridge regression results as a degenerate case when the forward operator is the synthesis identity, with the weighted bi-Lipschitz property holding as an equality and the same rate structure governing both sparse and non-sparse regimes.
- The equivalence chain σ_n(f) = O(n^{1/2-1/t}) ⟺ f ∈ k_t ⟹ VSC with φ(s) = s^{(1-t)/(2-t)} provides a practical diagnostic: measuring best n-term approximation decay for a given signal class directly determines the achievable statistical convergence rate.
- The parameter choice λ* = n^{-(1-r)/(1+b-r)} depends on unknown r and b, but the dual-function characterization in Corollary 4.4 suggests a data-driven balancing principle analogous to the discrepancy principle, potentially enabling adaptive selection without prior knowledge of smoothness.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning. The unknown is modeled as an element of ℓ¹, and observations are generated through a possibly nonlinear forward operator A: ℓ¹ → H, where H is a vector-valued reproducing kernel Hilbert space (vv-RKHS). The authors propose an ℓ¹-regularized empirical risk minimizer and establish almost-sure consistency, non-asymptotic high-probability convergence rates in both prediction and ℓ¹ reconstruction norms, and matching minimax lower bounds. The rates depend on a source smoothness parameter r (characterized by a variational source condition) and an effective dimension exponent b (polynomial spectral decay of the covariance operator). The theory is connected to practical sparsity models via approximation spaces k_t, and the assumptions are verified for two representative inverse problems: reaction coefficient identification in elliptic PDEs and sparse computed tomography.
Significance. The paper makes a substantial contribution by extending the statistical inverse learning framework to ℓ¹-regularization in a Banach-space setting with possibly nonlinear forward operators. The minimax optimality of the derived rates n^{-r/(1+b-br)} (in ℓ¹ reconstruction norm) is a central and significant claim, established through matching upper and lower bounds over the prior class P_{r,b}. The connection between approximation spaces k_t, variational source conditions, and best n-term approximation errors (Theorem 5.7, Lemma 5.9) provides a concrete and verifiable bridge between sparse approximation theory and statistical convergence rates. The application to filtered Radon transforms, including explicit effective-dimension asymptotics for boxcar and Gaussian filters (Proposition 6.5), yields concrete convergence rates for standard image models and sparsifying systems, adding practical value to the theoretical contributions.
major comments (3)
- §5.2, Theorem 5.7 and the text following it: The converse direction of the k_t-to-VSC bridge is asserted but not proved. The text states: 'Conversely, again under Assumption 6(i), if f_ρ satisfies the variational source condition of Assumption 5 with φ(s)=s^r, then f_ρ belongs to the smoothness space k_t with t=(2r-1)/(r-1).' However, Theorem 5.7 only establishes the forward direction (k_t → VSC). While the skeptic's analysis confirms that the minimax claim in Corollary 5.8 (restricted to r ∈ (0,1/2)) does not logically depend on the unproved converse—since the lower bound constructs specific ρ* ∈ P_{r,b} via the proved forward direction—the asserted converse should either be proved, stated as a conjecture, or removed. As written, it could mislead readers into thinking the equivalence is fully established within the paper.
- §5.2, Corollary 5.8 vs. §4, Corollary 4.7: There is a mismatch in the range of r for which minimax optimality is claimed. The upper bound (Corollary 4.7) is stated for r ∈ (0,1), while the lower bound (Theorem 5.4) and the minimax optimality claim (Corollary 5.8) are restricted to r ∈ (0,1/2). Remark 5.10 acknowledges this restriction, but the abstract and introduction state that 'matching minimax lower bounds' are established without clearly communicating this restriction on r. The abstract should be amended to accurately reflect that minimax optimality is established for r ∈ (0,1/2), or the scope of the lower bound should be extended.
- §2.5, Assumption 6: The weighted bi-Lipschitz property (both parts (i) and (ii) with the same weight w) is load-bearing for the minimax optimality claim (Corollary 5.8), as the authors note. However, the verification in §6.1 for finitely smoothing operators A = G ∘ S shows that parts (i) and (ii) hold with different weights in the shearlet case (w_λ = 2^{-|λ|(a+1/2)} for (i) and w_λ = 2^{-|λ|(a-1/2)} for (ii)). The minimax optimality result requires the same weight w for both parts. The paper should clarify whether the minimax claim applies to the shearlet case, or whether it is limited to cases (like the wavelet case or the direct synthesis operator in §6.4) where the weights coincide.
minor comments (7)
- §1.1, Table 1: The 'Rate' column for the present work lists 'n^{-r/(1+b-br)}', which corresponds to the ℓ¹ reconstruction norm rate (p=1). The table would benefit from also listing the general interpolation norm rate n^{-(2r-pr+p-1)/(p(1+b-br))} for completeness, or noting that the displayed rate is the special case p=1.
- §4, Corollary 4.6, Eq. (28): The exponent in the log factor is written as log^{2/p}(4/η), but the derivation from Theorem 4.3 (which has log²(4/η)) should make this log^{2/p}(4/η). This appears correct but could be stated more explicitly for the reader.
- §6.3, Proposition 6.5: The boxcar filter case yields b = 2/3, and the Gaussian filter case yields any b ∈ (0,1). It would be helpful to explicitly state the resulting convergence rates (as done at the end of §6.3 for specific examples) in terms of the general formula n^{-r/(1+b-br)} for these two filter choices, to make the practical implications more immediately visible.
- §3, proof of Theorem 3.1: The application of Scheffé's lemma to conclude strong convergence from weak-* convergence and norm convergence in ℓ¹ is correct but could benefit from a brief justification or citation, as this is a less commonly used tool in this context.
- §2.7, Definition 2.4: The class P_{r,b} is defined with 0 ≤ r ≤ 1, but the main results (e.g., Corollary 4.5) require 0 < r < 1. The boundary cases r = 0 and r = 1 should be discussed or excluded from the definition for consistency.
- §5.1, Theorem 5.3: The condition on the weight sequence (49) involves constants c_0, ε_0, q. Remark 5.5 verifies this for polynomially decaying weights w_m = m^{-a}, but the relationship between the constant a and the parameters r, b is not fully explicit. Stating the constraint on a in terms of r and b would improve clarity.
- Typographical: §6.1, the sentence 'In the shearlet case, using H^{-a} ↪ S^{-a-1/2}_{2,2}, thus and Assumption 6(ii) is verified with w_λ = 2^{-|λ|(a+1/2)}' contains a grammatical error ('thus and').
Simulated Author's Rebuttal
We thank the referee for a careful reading and three substantive comments, all of which are correct. We address each below and describe the revisions we will make.
read point-by-point responses
-
Referee: §5.2, Theorem 5.7 and the text following it: The converse direction of the k_t-to-VSC bridge is asserted but not proved. The text states the converse but Theorem 5.7 only establishes the forward direction. The asserted converse should either be proved, stated as a conjecture, or removed.
Authors: The referee is correct. Theorem 5.7 proves only the forward implication: membership in k_t (together with Assumption 6(ii)) implies the variational source condition with r = (1-t)/(2-t). The converse statement — that a variational source condition with φ(s) = s^r implies membership in k_t with t = (2r-1)/(r-1) under Assumption 6(i) — appears in the text following Theorem 5.7 without a proof. We do not currently have a complete proof of this converse direction. As the referee notes, the minimax claim in Corollary 5.8 does not logically depend on the converse: the lower bound construction in Theorem 5.4 uses the proved forward direction to exhibit specific ρ* ∈ P_{r,b}. Nevertheless, the unqualified assertion of the converse is misleading as written. We will revise the text to state the converse explicitly as a conjecture (Conjecture 5.8 or similar), clearly marking it as unproved, and will adjust the surrounding discussion to avoid any implication that the equivalence is fully established within the paper. revision: yes
-
Referee: §5.2, Corollary 5.8 vs. §4, Corollary 4.7: There is a mismatch in the range of r for which minimax optimality is claimed. The upper bound is stated for r ∈ (0,1), while the lower bound and minimax optimality claim are restricted to r ∈ (0,1/2). The abstract and introduction state that 'matching minimax lower bounds' are established without clearly communicating this restriction on r.
Authors: The referee is correct. The upper convergence rate (Corollary 4.7) is established for r ∈ (0,1), while the minimax lower bound (Theorem 5.4) and the minimax optimality statement (Corollary 5.8) are restricted to r ∈ (0,1/2). This restriction arises because the mapping t ↦ r = (1-t)/(2-t) from the approximation space k_t to the source condition index r has range (0,1/2), and the lower bound construction relies on this connection. Remark 5.10 acknowledges this, but the abstract and introduction do not. We will amend the abstract to read: 'We further prove matching minimax lower bounds for r ∈ (0,1/2), showing that the obtained convergence rates are optimal in this regime.' We will make a corresponding adjustment in the introduction (contribution (iii)) and in the statement of Corollary 5.8 to make the restriction on r explicit and prominent. revision: yes
-
Referee: §2.5, Assumption 6: The weighted bi-Lipschitz property requires the same weight w for both parts (i) and (ii), but the verification in §6.1 for finitely smoothing operators A = G ∘ S shows that parts (i) and (ii) hold with different weights in the shearlet case. The minimax optimality result requires the same weight w for both parts. The paper should clarify whether the minimax claim applies to the shearlet case, or whether it is limited to cases where the weights coincide.
Authors: The referee is correct. The minimax optimality result in Corollary 5.8 requires both parts of Assumption 6 to hold with the same weight w. In the wavelet case (Section 6.1), both parts (i) and (ii) are verified with w_λ = 2^{-|λ|a}, so the minimax claim applies. In the shearlet case, part (i) holds with w_λ = 2^{-|λ|(a+1/2)} and part (ii) holds with w_λ = 2^{-|λ|(a-1/2)}, which are different. Consequently, the minimax optimality result of Corollary 5.8 does not directly apply to the shearlet case as currently stated. The upper convergence rates (Corollary 4.7) still hold for the shearlet case, since they require only Assumption 6(ii), but the matching lower bound is not established in this setting. We will add a clarifying remark in Section 6.1 explicitly stating that: (a) the minimax optimality claim applies to the wavelet case and to the direct synthesis operator case (Section 6.4), where the weights coincide; (b) for the shearlet case, the upper rates remain valid but the minimax lower bound is not established, because the two parts of Assumption 6 hold with different weights. We view extending the lower bound to the shearlet case — either by refining the construction to accommodate distinct weights or by identifying a common weight under which both bounds hold — as an interesting open problem, which we will mention. revision: yes
Circularity Check
No significant circularity found; the minimax argument is self-contained with a minor self-citation for a standard concentration inequality.
full rationale
The paper's central minimax claim (Corollary 5.8) is derived through genuinely independent upper and lower bound arguments. The upper bound (Corollary 4.7) applies the variational source condition (Assumption 5, φ(t)=t^r) and polynomial spectral decay (Assumption 7) directly to the ℓ¹-regularized estimator, yielding rate n^{-(2r-pr+p-1)/(p(1+b-br))}. The lower bound (Theorem 5.4) constructs hard instances f_i ∈ k_t (approximation space), then uses Theorem 5.7 — which is PROVED in the forward direction (k_t → VSC under Assumption 6(ii)) — to certify that these instances belong to P_{r,b}. Fano's inequality then yields the matching lower rate. The forward direction of Theorem 5.7, which is the only direction needed for the lower bound (constructing specific ρ* ∈ P_{r,b}), is rigorously proved via Lemma 5.6 (itself a variant of Lemma 13 in [33], an external reference with no author overlap) combined with Assumption 6(ii). The converse direction (VSC → k_t) is asserted but not proved; however, it is NOT load-bearing for the minimax claim, since minimax optimality only requires that the hard instances lie in P_{r,b} (forward direction), not that all of P_{r,b} lies in k_t. The self-citation to [37] (Rastogi, Blanchard, Mathé) for Proposition A.1 is a standard Bernstein-type concentration inequality, not a premise that defines the conclusion. The regularization parameter λ* = n^{-(1-r)/(1+b-br)} is an a priori choice balancing bias and variance, not a data-driven fit. No step in the derivation chain reduces to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (7)
- r (source smoothness)
- b (effective dimension exponent)
- λ* (regularization parameter) =
n^{-(1-r)/(1+b-br)}
- C_β (spectral decay constant)
- M, Σ (noise parameters)
- L_A (Lipschitz constant of A)
- L (weighted bi-Lipschitz constant)
axioms (6)
- domain assumption Sub-exponential noise (Assumption 2): the noise satisfies a Bernstein-type condition with constants M, Σ.
- domain assumption Variational source condition (Assumption 5): ∥f−f_ρ∥_{ℓ¹} ≤ ∥f∥_{ℓ¹} − ∥f_ρ∥_{ℓ¹} + ϕ(∥A(f)−A(f_ρ)∥²_{H_μ}).
- domain assumption Polynomial spectral decay (Assumption 7): N(λ) ≤ C_β λ^{-b} for 0 < b < 1.
- domain assumption Weighted bi-Lipschitz property (Assumption 6): A is Lipschitz and lower-Lipschitz with respect to ∥·∥_{w,2}.
- domain assumption Weak-to-weak sequential continuity of S_μ ∘ A (Theorem 3.1).
- domain assumption Eigenvalue decay s_j ≤ β j^{-1/b} of the covariance operator (Section 5.1, before Theorem 5.3).
read the original abstract
We study the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning. The unknown is modeled as an element of $\ell^1$, and observations are generated through a possibly nonlinear forward operator $A:\ell^1\to H$, where $H$ is a vector-valued reproducing kernel Hilbert space. We propose an $\ell^1$-regularized empirical risk minimizer and develop a theoretical analysis of its statistical properties. Under mild assumptions, we establish almost-sure consistency and derive non-asymptotic high-probability convergence rates in both the prediction and $\ell^1$ reconstruction norms. The rates depend on the source smoothness parameter $r$, characterized by a variational source condition, and the effective dimension exponent $b$, describing the polynomial spectral decay of the covariance operator. We further prove matching minimax lower bounds, showing that the obtained convergence rates are optimal. To relate the theory to practical sparsity models, we consider finitely smoothing operators of the form $A=G\circ S$, where $S$ is a synthesis operator, and show that approximation-space assumptions imply the required variational source conditions. In particular, we prove that membership in the approximation space $k_t$ is equivalent to polynomial decay of the best $n$-term approximation error. Finally, we verify the assumptions for two representative inverse problems: reaction coefficient identification in elliptic PDEs and sparse computed tomography. For filtered Radon transforms, we derive explicit effective-dimension asymptotics, yielding concrete convergence rates for standard image models and sparsifying systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Laplace priors and spatial inhomogeneity in Bayesian inverse problems.Bernoulli, 30(2):878–910, 2024
Sergios Agapiou and Sven Wang. Laplace priors and spatial inhomogeneity in Bayesian inverse problems.Bernoulli, 30(2):878–910, 2024
work page 2024
-
[2]
Compressed sensing for inverse problems II: applications to deconvolution, source recovery, and MRI
Giovanni S Alberti, Alessandro Felisi, Matteo Santacesaria, and S Ivan Trapasso. Compressed sensing for inverse problems ii: applications to deconvolution, source recovery, and mri.arXiv preprint arXiv:2501.01929, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[3]
Giovanni S Alberti, Alessandro Felisi, Matteo Santacesaria, and Salvatore Ivan Trapasso. Compressed sensing for inverse problems and the sample complexity of the sparse radon transform.Journal of the European Mathematical Society, pages 1–56, 2025
work page 2025
-
[4]
´Alvarez, Lorenzo Rosasco, and Neil D
Mauricio A. ´Alvarez, Lorenzo Rosasco, and Neil D. Lawrence.Kernels for vector-valued functions: A review, volume 4. Now Foundations and Trends, 2012
work page 2012
-
[5]
Theory of reproducing kernels.Transactions of the American Mathe- matical Society, 68:337–404, 1950
Nachman Aronszajn. Theory of reproducing kernels.Transactions of the American Mathe- matical Society, 68:337–404, 1950
work page 1950
-
[6]
On regularization algorithms in learn- ing theory.Journal of Complexity, 23(1):52 – 72, 2007
Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco. On regularization algorithms in learn- ing theory.Journal of Complexity, 23(1):52 – 72, 2007
work page 2007
-
[7]
Bickel, Ya’acov Ritov, and Alexandre B
Peter J. Bickel, Ya’acov Ritov, and Alexandre B. Tsybakov. Simultaneous analysis of Lasso and Dantzig selector.Annals of Statistics, 37(4):1705–1732, 2009. 45
work page 2009
-
[8]
Gilles Blanchard and Nicole M¨ ucke. Optimal rates for regularization of statistical inverse learning problems.Foundations of Computational Mathematics, 18(4):971–1013, 2018
work page 2018
-
[9]
Gilles Blanchard and Nicole M¨ ucke. Kernel regression, minimax rates and effective dimen- sionality: Beyond the regular case.Analysis and Applications, 18(04):683–696, 2020
work page 2020
-
[10]
Bubba, Martin Burger, Tapio Helin, and Luca Ratti
Tatiana A. Bubba, Martin Burger, Tapio Helin, and Luca Ratti. Convex regularization in statistical inverse learning problems.Inverse Problems and Imaging, 17(6):1193–1225, 2023
work page 2023
-
[11]
Tatiana A Bubba and Luca Ratti. Shearlet-based regularization in statistical inverse learning with an application to x-ray tomography.Inverse Problems, 38(5):054001, 2022
work page 2022
-
[12]
Emmanuel J Cand` es and David L Donoho. New tight frames of curvelets and optimal rep- resentations of objects with piecewise c2 singularities.Communications on Pure and Applied Mathematics: A Journal Issued by the Courant Institute of Mathematical Sciences, 57(2):219– 266, 2004
work page 2004
-
[13]
Emmanuel J. Cand` es and Terence Tao. The dantzig selector: Statistical estimation whenp is much larger thann.The Annals of Statistics, 35(6):2313–2351, 2007
work page 2007
-
[14]
Andrea Caponnetto and Ernesto De Vito. Optimal rates for the regularized least-squares algorithm.Foundations of Computational Mathematics, 7(3):331–368, 2007
work page 2007
-
[15]
Claudio Carmeli, Ernesto De Vito, and Alessandro Toigo. Vector valued reproducing ker- nel Hilbert spaces of integrable functions and Mercer theorem.Analysis and Applications, 4(04):377–408, 2006
work page 2006
-
[16]
Claudio Carmeli, Ernesto De Vito, Alessandro Toigo, and Veronica Umanit´ a. Vector valued reproducing kernel Hilbert spaces and universality.Analysis and Applications, 8(01):19–61, 2010
work page 2010
-
[17]
Nonlinear approximation and the space bv(r2).American Journal of Mathematics, 121(3):587–628, 1999
Albert Cohen, Ronald DeVore, Pencho Petrushev, and Hong Xu. Nonlinear approximation and the space bv(r2).American Journal of Mathematics, 121(3):587–628, 1999
work page 1999
-
[18]
Ronald DeVore, Gerard Kerkyacharian, Dominique Picard, and Vladimir Temlyakov. Ap- proximation methods for supervised learning.Foundations of Computational Mathematics, 6(1):3–58, 2006
work page 2006
-
[19]
Applications of Mathematics, Springer, New York, 1996
Luc Devroye, L´ aszl´ o Gy¨ orfi, and G´ abor Lugosi.A probabilistic theory of pattern recognition, volume 31. Applications of Mathematics, Springer, New York, 1996
work page 1996
-
[20]
Jens Flemming and Daniel Gerth. Injectivity and weak*-to-weak continuity suffice for con- vergence rates inℓ 1-regularization.Journal of Inverse and Ill-posed Problems, 26(1):85–94, 2018
work page 2018
-
[21]
David Gilbarg, Neil S Trudinger, David Gilbarg, and NS Trudinger.Elliptic partial differential equations of second order, volume 2. Springer, 1998
work page 1998
-
[22]
Kanghui Guo and Demetrio Labate. Optimally sparse multidimensional representation using shearlets.SIAM journal on mathematical analysis, 39(1):298–318, 2007
work page 2007
-
[23]
Niklas Hartung, Martin Wahl, Abhishake Rastogi, and Wilhelm Huisinga. Nonparametric goodness-of-fit testing for parametric covariate models in pharmacometric analyses.CPT: Pharmacometrics & Systems Pharmacology, 10(6):564–576, 2021
work page 2021
-
[24]
Thorsten Hohage and Philip Miller. Optimal convergence rates for sparsity promoting wavelet- regularization in besov spaces.Inverse Problems, 35(6):065005, 2019
work page 2019
-
[25]
Springer- Verlag, Berlin Heidelberg, 1985
Lars H¨ ormander.The Analysis of Linear Partial Differential Operators, volume IV. Springer- Verlag, Berlin Heidelberg, 1985. 46
work page 1985
-
[26]
Inverse problems involving pdes with applications to imaging
Taufiquar Khan. Inverse problems involving pdes with applications to imaging. In Pammy Manchanda, Ren´ e Pierre Lozi, and Abul Hasan Siddiqi, editors,Mathematical Modelling, Optimization, Analytic and Numerical Solutions, pages 181–195. Springer, Singapore, 2020
work page 2020
-
[27]
Shearlet smoothness spaces.Journal of Fourier Analysis and Applications, 19(3):577–611, 2013
Demetrio Labate, Lucia Mantovani, and Pooran Negi. Shearlet smoothness spaces.Journal of Fourier Analysis and Applications, 19(3):577–611, 2013
work page 2013
-
[28]
Junhong Lin, Alessandro Rudi, Lorenzo Rosasco, and Volkan Cevher. Optimal rates for spec- tral algorithms with least-squares regression over Hilbert spaces.Applied and Computational Harmonic Analysis, 48(3):868–890, 2020
work page 2020
-
[29]
Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri
L. Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri. Spectral algorithms for supervised learning.Neural Computation, 20(7):1873–1897, 2008
work page 2008
- [30]
-
[31]
Regularization in kernel learning.The Annals of Statistics, 38(1):526 – 565, 2010
Shahar Mendelson and Joseph Neeman. Regularization in kernel learning.The Annals of Statistics, 38(1):526 – 565, 2010
work page 2010
-
[32]
On learning vector-valued functions.Neural Computation, 17(1):177–204, 2005
Charles A Micchelli and Massimiliano Pontil. On learning vector-valued functions.Neural Computation, 17(1):177–204, 2005
work page 2005
-
[33]
Philip Miller and Thorsten Hohage. Maximal spaces for approximation rates inℓ 1- regularization.Numerische Mathematik, 149(2):341–374, 2021
work page 2021
-
[34]
Jennifer L Mueller and Samuli Siltanen.Linear and nonlinear inverse problems with practical applications. SIAM, Philadelphia, 2012
work page 2012
- [35]
-
[36]
Garvesh Raskutti, Martin J. Wainwright, and Bin Yu. Minimax rates of estimation for high-dimensional linear regression overℓ q-balls.IEEE Transactions on Information Theory, 57(10):6976–6994, 2011
work page 2011
-
[37]
Abhishake Rastogi, Gilles Blanchard, and Peter Math´ e. Convergence analysis of Tikhonov regularization for non-linear statistical inverse problems.Electronic Journal of Statistics, 14(2):2798–2841, 2020
work page 2020
-
[38]
Inverse learning in Hilbert scales.Machine Learning, 112:2469–2499, 2023
Abhishake Rastogi and Peter Math´ e. Inverse learning in Hilbert scales.Machine Learning, 112:2469–2499, 2023
work page 2023
-
[39]
Abhishake Rastogi and Sivananthan Sampath. Optimal rates for the regularized learning algorithms under general source condition.Frontiers in Applied Mathematics and Statistics, 3:3, 2017
work page 2017
-
[40]
Thomas Schuster, Barbara Kaltenbacher, Bernd Hofmann, and Kamil S Kazimierski.Reg- ularization methods in Banach spaces, volume 10 ofRadon Series on Computational and Applied Mathematics. Walter de Gruyter GmbH & Co. KG, Berlin, 2012
work page 2012
-
[41]
Khemraj Shukla, Patricio Clark Di Leoni, James Blackshire, Daniel Sparkman, and George Em Karniadakis. Physics-informed neural network for ultrasound nondestructive quantification of surface breaking cracks.Journal of Nondestructive Evaluation, 39(3):61, 2020
work page 2020
-
[42]
Optimal rates for regularized least squares regression
Ingo Steinwart, Don Hush, and Clint Scovel. Optimal rates for regularized least squares regression. InS. Dasgupta and A. Klivans, editors, Proceedings of the 22nd Annual Conference on Learning Theory, pages 79–93, 2009
work page 2009
-
[43]
Robert Tibshirani. Regression shrinkage and selection via the lasso.Journal of the Royal Statistical Society Series B: Statistical Methodology, 58(1):267–288, 1996. 47
work page 1996
-
[44]
Birkh¨ auser Verlag, Basel, 1983
Hans Triebel.Theory of function spaces. Birkh¨ auser Verlag, Basel, 1983
work page 1983
-
[45]
Introduction to nonparametric estimation
Alexandre Tsybakov. Introduction to nonparametric estimation. InSpringer Series in Statis- tics, 2008
work page 2008
-
[46]
W. Van Aarle, W. J. Palenstijn, J. Cant, E. Janssens, F. Bleichrodt, A. Dabravolski, J. De Beenhouwer, K. J. Batenburg, and J. Sijbers. Fast and flexible X-ray tomography using the ASTRA toolbox.Optics express, 24(22):25129–25147, 2016
work page 2016
-
[47]
Aad W. van der Vaart and Jon A. Wellner.Weak convergence and empirical processes. Springer Series in Statistics, Springer-Verlag, New York, 1996. With Applications to Statistics. 48
work page 1996
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.