Pith. sign in

REVIEW 3 major objections 7 minor 47 references

Sparse recovery rates for nonlinear inverse problems proven minimax-optimal

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 09:55 UTC pith:ANTEEP4Y

load-bearing objection Minimax-optimal rates for ℓ¹-regularized nonlinear statistical inverse learning; proofs are sound and the k_t-to-VSC bridge holds for the minimax claim as stated. the 3 major comments →

arxiv 2607.07468 v1 pith:ANTEEP4Y submitted 2026-07-08 stat.ML cs.LGmath.STstat.TH

Statistical inverse learning and ell¹-regularization

classification stat.ML cs.LGmath.STstat.TH
keywords ratesassumptionsconvergenceinverseoperatorsourcestatisticalapproximation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper studies the problem of recovering a sparse, infinite-dimensional signal from finitely many noisy and indirect observations, as arise in computed tomography or PDE coefficient identification. The authors propose an ℓ¹-regularized empirical risk minimizer in a vector-valued reproducing kernel Hilbert space framework, where the forward operator mapping the unknown to the data can be nonlinear. The central result is a complete convergence-rate theory: under a variational source condition (characterizing solution smoothness via a parameter r) and polynomial spectral decay of the covariance operator (characterizing statistical complexity via an exponent b), the estimator achieves the reconstruction rate n^{-r/(1+b-br)} in ℓ¹ norm. The authors prove matching minimax lower bounds, establishing that no estimator can converge faster over the same class of problems. They further show that membership in an approximation space k_t — equivalently, polynomial decay of best n-term approximation errors — implies the required variational source condition, bridging sparse approximation theory with statistical inverse learning. The assumptions are verified concretely for reaction coefficient identification in elliptic PDEs and for filtered Radon transforms in computed tomography, yielding explicit rates for standard image models and sparsifying systems such as wavelets and shearlets.

Core claim

The minimax-optimal convergence rate for ℓ¹-regularized sparse statistical inverse learning with nonlinear forward operators is n^{-r/(1+b-br)} in ℓ¹ reconstruction norm, where r encodes the solution's sparsity-driven smoothness (via a variational source condition) and b encodes the effective dimension of the learning problem (via spectral decay of the covariance operator). This rate is achieved by the regularized estimator with parameter λ* = n^{-(1-r)/(1+b-br)} and cannot be improved by any estimator over the prior class P_{r,b}. A key structural finding is that the variational source condition — the analytical engine for the rates — is equivalent to membership in an approximation space kₜ

What carries the argument

The variational source condition (Assumption 5) with index function φ(t) = t^r, which quantifies how well the true solution can be approximated relative to the forward operator's sensitivity; the effective dimension N(λ) with polynomial bound N(λ) ≤ Cλ^{-b}, capturing statistical complexity via covariance spectral decay; the weighted bi-Lipschitz property (Assumption 6), which provides two-sided stability of the forward operator between weighted sequence space and data space; the approximation space kₘ

Load-bearing premise

The weighted bi-Lipschitz property (Assumption 6) requires the forward operator to satisfy a lower Lipschitz bound — meaning small changes in the data must reflect proportionally small changes in the solution. This fails for severely ill-posed problems like the backward heat equation or electrical impedance tomography, restricting the theory to finitely smoothing operators.

What would settle it

Construct a probability distribution in P_{r,b} for which the ℓ¹-regularized estimator with λ* = n^{-(1-r)/(1+b-br)} converges slower than n^{-r/(1+b-br)}, or exhibit any learning algorithm achieving a faster rate over the full class P_{r,b}.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For filtered Radon transforms with a boxcar filter, the effective dimension exponent is b = 2/3; with a Gaussian filter, eigenvalues decay super-polynomially so any b > 0 is admissible, yielding near-parametric rates n^{-r} for sufficiently sparse signals.
  • For cartoon-like images represented in shearlet frames (best n-term approximation rate n^{-1}), the theory predicts r = 1/4, giving concrete convergence rates of n^{-1/6} (boxcar) or n^{-1/4} (Gaussian) in ℓ¹ reconstruction norm.
  • The framework recovers classical kernel ridge regression results as a degenerate case when the forward operator is the synthesis identity, with the weighted bi-Lipschitz property holding as an equality and the same rate structure governing both sparse and non-sparse regimes.
  • The equivalence chain σ_n(f) = O(n^{1/2-1/t}) ⟺ f ∈ k_t ⟹ VSC with φ(s) = s^{(1-t)/(2-t)} provides a practical diagnostic: measuring best n-term approximation decay for a given signal class directly determines the achievable statistical convergence rate.
  • The parameter choice λ* = n^{-(1-r)/(1+b-r)} depends on unknown r and b, but the dual-function characterization in Corollary 4.4 suggests a data-driven balancing principle analogous to the discrepancy principle, potentially enabling adaptive selection without prior knowledge of smoothness.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper studies the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning. The unknown is modeled as an element of ℓ¹, and observations are generated through a possibly nonlinear forward operator A: ℓ¹ → H, where H is a vector-valued reproducing kernel Hilbert space (vv-RKHS). The authors propose an ℓ¹-regularized empirical risk minimizer and establish almost-sure consistency, non-asymptotic high-probability convergence rates in both prediction and ℓ¹ reconstruction norms, and matching minimax lower bounds. The rates depend on a source smoothness parameter r (characterized by a variational source condition) and an effective dimension exponent b (polynomial spectral decay of the covariance operator). The theory is connected to practical sparsity models via approximation spaces k_t, and the assumptions are verified for two representative inverse problems: reaction coefficient identification in elliptic PDEs and sparse computed tomography.

Significance. The paper makes a substantial contribution by extending the statistical inverse learning framework to ℓ¹-regularization in a Banach-space setting with possibly nonlinear forward operators. The minimax optimality of the derived rates n^{-r/(1+b-br)} (in ℓ¹ reconstruction norm) is a central and significant claim, established through matching upper and lower bounds over the prior class P_{r,b}. The connection between approximation spaces k_t, variational source conditions, and best n-term approximation errors (Theorem 5.7, Lemma 5.9) provides a concrete and verifiable bridge between sparse approximation theory and statistical convergence rates. The application to filtered Radon transforms, including explicit effective-dimension asymptotics for boxcar and Gaussian filters (Proposition 6.5), yields concrete convergence rates for standard image models and sparsifying systems, adding practical value to the theoretical contributions.

major comments (3)
  1. §5.2, Theorem 5.7 and the text following it: The converse direction of the k_t-to-VSC bridge is asserted but not proved. The text states: 'Conversely, again under Assumption 6(i), if f_ρ satisfies the variational source condition of Assumption 5 with φ(s)=s^r, then f_ρ belongs to the smoothness space k_t with t=(2r-1)/(r-1).' However, Theorem 5.7 only establishes the forward direction (k_t → VSC). While the skeptic's analysis confirms that the minimax claim in Corollary 5.8 (restricted to r ∈ (0,1/2)) does not logically depend on the unproved converse—since the lower bound constructs specific ρ* ∈ P_{r,b} via the proved forward direction—the asserted converse should either be proved, stated as a conjecture, or removed. As written, it could mislead readers into thinking the equivalence is fully established within the paper.
  2. §5.2, Corollary 5.8 vs. §4, Corollary 4.7: There is a mismatch in the range of r for which minimax optimality is claimed. The upper bound (Corollary 4.7) is stated for r ∈ (0,1), while the lower bound (Theorem 5.4) and the minimax optimality claim (Corollary 5.8) are restricted to r ∈ (0,1/2). Remark 5.10 acknowledges this restriction, but the abstract and introduction state that 'matching minimax lower bounds' are established without clearly communicating this restriction on r. The abstract should be amended to accurately reflect that minimax optimality is established for r ∈ (0,1/2), or the scope of the lower bound should be extended.
  3. §2.5, Assumption 6: The weighted bi-Lipschitz property (both parts (i) and (ii) with the same weight w) is load-bearing for the minimax optimality claim (Corollary 5.8), as the authors note. However, the verification in §6.1 for finitely smoothing operators A = G ∘ S shows that parts (i) and (ii) hold with different weights in the shearlet case (w_λ = 2^{-|λ|(a+1/2)} for (i) and w_λ = 2^{-|λ|(a-1/2)} for (ii)). The minimax optimality result requires the same weight w for both parts. The paper should clarify whether the minimax claim applies to the shearlet case, or whether it is limited to cases (like the wavelet case or the direct synthesis operator in §6.4) where the weights coincide.
minor comments (7)
  1. §1.1, Table 1: The 'Rate' column for the present work lists 'n^{-r/(1+b-br)}', which corresponds to the ℓ¹ reconstruction norm rate (p=1). The table would benefit from also listing the general interpolation norm rate n^{-(2r-pr+p-1)/(p(1+b-br))} for completeness, or noting that the displayed rate is the special case p=1.
  2. §4, Corollary 4.6, Eq. (28): The exponent in the log factor is written as log^{2/p}(4/η), but the derivation from Theorem 4.3 (which has log²(4/η)) should make this log^{2/p}(4/η). This appears correct but could be stated more explicitly for the reader.
  3. §6.3, Proposition 6.5: The boxcar filter case yields b = 2/3, and the Gaussian filter case yields any b ∈ (0,1). It would be helpful to explicitly state the resulting convergence rates (as done at the end of §6.3 for specific examples) in terms of the general formula n^{-r/(1+b-br)} for these two filter choices, to make the practical implications more immediately visible.
  4. §3, proof of Theorem 3.1: The application of Scheffé's lemma to conclude strong convergence from weak-* convergence and norm convergence in ℓ¹ is correct but could benefit from a brief justification or citation, as this is a less commonly used tool in this context.
  5. §2.7, Definition 2.4: The class P_{r,b} is defined with 0 ≤ r ≤ 1, but the main results (e.g., Corollary 4.5) require 0 < r < 1. The boundary cases r = 0 and r = 1 should be discussed or excluded from the definition for consistency.
  6. §5.1, Theorem 5.3: The condition on the weight sequence (49) involves constants c_0, ε_0, q. Remark 5.5 verifies this for polynomially decaying weights w_m = m^{-a}, but the relationship between the constant a and the parameters r, b is not fully explicit. Stating the constraint on a in terms of r and b would improve clarity.
  7. Typographical: §6.1, the sentence 'In the shearlet case, using H^{-a} ↪ S^{-a-1/2}_{2,2}, thus and Assumption 6(ii) is verified with w_λ = 2^{-|λ|(a+1/2)}' contains a grammatical error ('thus and').

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for a careful reading and three substantive comments, all of which are correct. We address each below and describe the revisions we will make.

read point-by-point responses
  1. Referee: §5.2, Theorem 5.7 and the text following it: The converse direction of the k_t-to-VSC bridge is asserted but not proved. The text states the converse but Theorem 5.7 only establishes the forward direction. The asserted converse should either be proved, stated as a conjecture, or removed.

    Authors: The referee is correct. Theorem 5.7 proves only the forward implication: membership in k_t (together with Assumption 6(ii)) implies the variational source condition with r = (1-t)/(2-t). The converse statement — that a variational source condition with φ(s) = s^r implies membership in k_t with t = (2r-1)/(r-1) under Assumption 6(i) — appears in the text following Theorem 5.7 without a proof. We do not currently have a complete proof of this converse direction. As the referee notes, the minimax claim in Corollary 5.8 does not logically depend on the converse: the lower bound construction in Theorem 5.4 uses the proved forward direction to exhibit specific ρ* ∈ P_{r,b}. Nevertheless, the unqualified assertion of the converse is misleading as written. We will revise the text to state the converse explicitly as a conjecture (Conjecture 5.8 or similar), clearly marking it as unproved, and will adjust the surrounding discussion to avoid any implication that the equivalence is fully established within the paper. revision: yes

  2. Referee: §5.2, Corollary 5.8 vs. §4, Corollary 4.7: There is a mismatch in the range of r for which minimax optimality is claimed. The upper bound is stated for r ∈ (0,1), while the lower bound and minimax optimality claim are restricted to r ∈ (0,1/2). The abstract and introduction state that 'matching minimax lower bounds' are established without clearly communicating this restriction on r.

    Authors: The referee is correct. The upper convergence rate (Corollary 4.7) is established for r ∈ (0,1), while the minimax lower bound (Theorem 5.4) and the minimax optimality statement (Corollary 5.8) are restricted to r ∈ (0,1/2). This restriction arises because the mapping t ↦ r = (1-t)/(2-t) from the approximation space k_t to the source condition index r has range (0,1/2), and the lower bound construction relies on this connection. Remark 5.10 acknowledges this, but the abstract and introduction do not. We will amend the abstract to read: 'We further prove matching minimax lower bounds for r ∈ (0,1/2), showing that the obtained convergence rates are optimal in this regime.' We will make a corresponding adjustment in the introduction (contribution (iii)) and in the statement of Corollary 5.8 to make the restriction on r explicit and prominent. revision: yes

  3. Referee: §2.5, Assumption 6: The weighted bi-Lipschitz property requires the same weight w for both parts (i) and (ii), but the verification in §6.1 for finitely smoothing operators A = G ∘ S shows that parts (i) and (ii) hold with different weights in the shearlet case. The minimax optimality result requires the same weight w for both parts. The paper should clarify whether the minimax claim applies to the shearlet case, or whether it is limited to cases where the weights coincide.

    Authors: The referee is correct. The minimax optimality result in Corollary 5.8 requires both parts of Assumption 6 to hold with the same weight w. In the wavelet case (Section 6.1), both parts (i) and (ii) are verified with w_λ = 2^{-|λ|a}, so the minimax claim applies. In the shearlet case, part (i) holds with w_λ = 2^{-|λ|(a+1/2)} and part (ii) holds with w_λ = 2^{-|λ|(a-1/2)}, which are different. Consequently, the minimax optimality result of Corollary 5.8 does not directly apply to the shearlet case as currently stated. The upper convergence rates (Corollary 4.7) still hold for the shearlet case, since they require only Assumption 6(ii), but the matching lower bound is not established in this setting. We will add a clarifying remark in Section 6.1 explicitly stating that: (a) the minimax optimality claim applies to the wavelet case and to the direct synthesis operator case (Section 6.4), where the weights coincide; (b) for the shearlet case, the upper rates remain valid but the minimax lower bound is not established, because the two parts of Assumption 6 hold with different weights. We view extending the lower bound to the shearlet case — either by refining the construction to accommodate distinct weights or by identifying a common weight under which both bounds hold — as an interesting open problem, which we will mention. revision: yes

Circularity Check

0 steps flagged

No significant circularity found; the minimax argument is self-contained with a minor self-citation for a standard concentration inequality.

full rationale

The paper's central minimax claim (Corollary 5.8) is derived through genuinely independent upper and lower bound arguments. The upper bound (Corollary 4.7) applies the variational source condition (Assumption 5, φ(t)=t^r) and polynomial spectral decay (Assumption 7) directly to the ℓ¹-regularized estimator, yielding rate n^{-(2r-pr+p-1)/(p(1+b-br))}. The lower bound (Theorem 5.4) constructs hard instances f_i ∈ k_t (approximation space), then uses Theorem 5.7 — which is PROVED in the forward direction (k_t → VSC under Assumption 6(ii)) — to certify that these instances belong to P_{r,b}. Fano's inequality then yields the matching lower rate. The forward direction of Theorem 5.7, which is the only direction needed for the lower bound (constructing specific ρ* ∈ P_{r,b}), is rigorously proved via Lemma 5.6 (itself a variant of Lemma 13 in [33], an external reference with no author overlap) combined with Assumption 6(ii). The converse direction (VSC → k_t) is asserted but not proved; however, it is NOT load-bearing for the minimax claim, since minimax optimality only requires that the hard instances lie in P_{r,b} (forward direction), not that all of P_{r,b} lies in k_t. The self-citation to [37] (Rastogi, Blanchard, Mathé) for Proposition A.1 is a standard Bernstein-type concentration inequality, not a premise that defines the conclusion. The regularization parameter λ* = n^{-(1-r)/(1+b-br)} is an a priori choice balancing bias and variance, not a data-driven fit. No step in the derivation chain reduces to its inputs by construction.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 0 invented entities

The paper introduces no new physical entities, particles, forces, or dimensions. All mathematical objects (vv-RKHS, covariance operator, approximation spaces k_t, weighted sequence spaces ℓ^p_w) are standard in the literature. The framework is constructive in the sense that it connects existing objects (variational source conditions, effective dimension, best n-term approximation) rather than postulating new ones.

free parameters (7)
  • r (source smoothness)
    Exponent in the variational source condition ϕ(t) = t^r. Not fitted to data but assumed as a property of the true solution. Enters the optimal λ* and the convergence rate.
  • b (effective dimension exponent)
    Polynomial decay exponent: N(λ) ≤ C_β λ^{-b}. Not fitted but assumed as a property of the covariance operator. Enters the optimal λ* and the convergence rate.
  • λ* (regularization parameter) = n^{-(1-r)/(1+b-br)}
    Optimal regularization parameter depending on r, b, and n. In practice r and b are unknown; the paper notes (Section 7) that b can be estimated spectrally and r via cross-validation, but no data-driven procedure with guarantees is provided.
  • C_β (spectral decay constant)
    Constant in N(λ) ≤ C_β λ^{-b}. Not numerically specified.
  • M, Σ (noise parameters)
    Sub-exponential noise parameters in Assumption 2. Not fitted; assumed fixed and known up to constants.
  • L_A (Lipschitz constant of A)
    Lipschitz constant in Assumption 4. Enters the high-probability bounds through C''. Not numerically specified.
  • L (weighted bi-Lipschitz constant)
    Constants in Assumption 6(i) and 6(ii). Required for upper and lower bounds. Verified as L=1 for the synthesis operator (Section 6.4) but not for general operators.
axioms (6)
  • domain assumption Sub-exponential noise (Assumption 2): the noise satisfies a Bernstein-type condition with constants M, Σ.
    Standard in statistical learning theory; excludes Gaussian white noise in infinite-dimensional output spaces, as noted in Section 2.3.
  • domain assumption Variational source condition (Assumption 5): ∥f−f_ρ∥_{ℓ¹} ≤ ∥f∥_{ℓ¹} − ∥f_ρ∥_{ℓ¹} + ϕ(∥A(f)−A(f_ρ)∥²_{H_μ}).
    Central smoothness assumption on the true solution. Shown to be implied by membership in approximation space k_t (Theorem 5.7), which is itself equivalent to polynomial decay of best n-term approximation (Lemma 5.9).
  • domain assumption Polynomial spectral decay (Assumption 7): N(λ) ≤ C_β λ^{-b} for 0 < b < 1.
    Controls the effective dimension of the learning problem. Verified for the PDE example with b = d/4 (Section 6.2) and for the filtered Radon transform with b = 2/3 (boxcar) or any b > 0 (Gaussian) (Section 6.3).
  • domain assumption Weighted bi-Lipschitz property (Assumption 6): A is Lipschitz and lower-Lipschitz with respect to ∥·∥_{w,2}.
    Required for upper rates (part ii), lower rates (part i), and the VSC verification (Theorem 5.7). Verified for finitely smoothing operators G satisfying condition (54). Excludes severely ill-posed problems.
  • domain assumption Weak-to-weak sequential continuity of S_μ ∘ A (Theorem 3.1).
    Required for consistency. Verified when A is linear and bounded, or nonlinear, continuous, and compact (Remark 3.2).
  • domain assumption Eigenvalue decay s_j ≤ β j^{-1/b} of the covariance operator (Section 5.1, before Theorem 5.3).
    Used in the lower bound construction to ensure Assumption 7 holds for the packing measures.

pith-pipeline@v1.1.0-glm · 46145 in / 4288 out tokens · 530085 ms · 2026-07-09T09:55:48.034715+00:00 · methodology

0 comments
read the original abstract

We study the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning. The unknown is modeled as an element of $\ell^1$, and observations are generated through a possibly nonlinear forward operator $A:\ell^1\to H$, where $H$ is a vector-valued reproducing kernel Hilbert space. We propose an $\ell^1$-regularized empirical risk minimizer and develop a theoretical analysis of its statistical properties. Under mild assumptions, we establish almost-sure consistency and derive non-asymptotic high-probability convergence rates in both the prediction and $\ell^1$ reconstruction norms. The rates depend on the source smoothness parameter $r$, characterized by a variational source condition, and the effective dimension exponent $b$, describing the polynomial spectral decay of the covariance operator. We further prove matching minimax lower bounds, showing that the obtained convergence rates are optimal. To relate the theory to practical sparsity models, we consider finitely smoothing operators of the form $A=G\circ S$, where $S$ is a synthesis operator, and show that approximation-space assumptions imply the required variational source conditions. In particular, we prove that membership in the approximation space $k_t$ is equivalent to polynomial decay of the best $n$-term approximation error. Finally, we verify the assumptions for two representative inverse problems: reaction coefficient identification in elliptic PDEs and sparse computed tomography. For filtered Radon transforms, we derive explicit effective-dimension asymptotics, yielding concrete convergence rates for standard image models and sparsifying systems.

Figures

Figures reproduced from arXiv: 2607.07468 by Abhishake Rastogi, Luca Ratti, Tapio Helin, Tatiana A. Bubba.

Figure 1
Figure 1. Figure 1: SVD decay for a 128×128 target with angular views in [0, π) using the ASTRA toolbox using parallel beam geometry and the strip model. We can finally detail the decay obtained in Corollary 4.5 for some specific examples, in which a specific decay of σn(f † ) is expected. Recall that if σn(f † ) = O(n −β ) with β > 1 2 , than, as described in 6.1, Assumption 5 is satisfied with ϕ(s) = s r , where r = 1 2 − 1… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

47 extracted references · 47 canonical work pages · 1 internal anchor

  1. [1]

    Laplace priors and spatial inhomogeneity in Bayesian inverse problems.Bernoulli, 30(2):878–910, 2024

    Sergios Agapiou and Sven Wang. Laplace priors and spatial inhomogeneity in Bayesian inverse problems.Bernoulli, 30(2):878–910, 2024

  2. [2]

    Compressed sensing for inverse problems II: applications to deconvolution, source recovery, and MRI

    Giovanni S Alberti, Alessandro Felisi, Matteo Santacesaria, and S Ivan Trapasso. Compressed sensing for inverse problems ii: applications to deconvolution, source recovery, and mri.arXiv preprint arXiv:2501.01929, 2025

  3. [3]

    Compressed sensing for inverse problems and the sample complexity of the sparse radon transform.Journal of the European Mathematical Society, pages 1–56, 2025

    Giovanni S Alberti, Alessandro Felisi, Matteo Santacesaria, and Salvatore Ivan Trapasso. Compressed sensing for inverse problems and the sample complexity of the sparse radon transform.Journal of the European Mathematical Society, pages 1–56, 2025

  4. [4]

    ´Alvarez, Lorenzo Rosasco, and Neil D

    Mauricio A. ´Alvarez, Lorenzo Rosasco, and Neil D. Lawrence.Kernels for vector-valued functions: A review, volume 4. Now Foundations and Trends, 2012

  5. [5]

    Theory of reproducing kernels.Transactions of the American Mathe- matical Society, 68:337–404, 1950

    Nachman Aronszajn. Theory of reproducing kernels.Transactions of the American Mathe- matical Society, 68:337–404, 1950

  6. [6]

    On regularization algorithms in learn- ing theory.Journal of Complexity, 23(1):52 – 72, 2007

    Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco. On regularization algorithms in learn- ing theory.Journal of Complexity, 23(1):52 – 72, 2007

  7. [7]

    Bickel, Ya’acov Ritov, and Alexandre B

    Peter J. Bickel, Ya’acov Ritov, and Alexandre B. Tsybakov. Simultaneous analysis of Lasso and Dantzig selector.Annals of Statistics, 37(4):1705–1732, 2009. 45

  8. [8]

    Optimal rates for regularization of statistical inverse learning problems.Foundations of Computational Mathematics, 18(4):971–1013, 2018

    Gilles Blanchard and Nicole M¨ ucke. Optimal rates for regularization of statistical inverse learning problems.Foundations of Computational Mathematics, 18(4):971–1013, 2018

  9. [9]

    Kernel regression, minimax rates and effective dimen- sionality: Beyond the regular case.Analysis and Applications, 18(04):683–696, 2020

    Gilles Blanchard and Nicole M¨ ucke. Kernel regression, minimax rates and effective dimen- sionality: Beyond the regular case.Analysis and Applications, 18(04):683–696, 2020

  10. [10]

    Bubba, Martin Burger, Tapio Helin, and Luca Ratti

    Tatiana A. Bubba, Martin Burger, Tapio Helin, and Luca Ratti. Convex regularization in statistical inverse learning problems.Inverse Problems and Imaging, 17(6):1193–1225, 2023

  11. [11]

    Shearlet-based regularization in statistical inverse learning with an application to x-ray tomography.Inverse Problems, 38(5):054001, 2022

    Tatiana A Bubba and Luca Ratti. Shearlet-based regularization in statistical inverse learning with an application to x-ray tomography.Inverse Problems, 38(5):054001, 2022

  12. [12]

    Emmanuel J Cand` es and David L Donoho. New tight frames of curvelets and optimal rep- resentations of objects with piecewise c2 singularities.Communications on Pure and Applied Mathematics: A Journal Issued by the Courant Institute of Mathematical Sciences, 57(2):219– 266, 2004

  13. [13]

    Cand` es and Terence Tao

    Emmanuel J. Cand` es and Terence Tao. The dantzig selector: Statistical estimation whenp is much larger thann.The Annals of Statistics, 35(6):2313–2351, 2007

  14. [14]

    Optimal rates for the regularized least-squares algorithm.Foundations of Computational Mathematics, 7(3):331–368, 2007

    Andrea Caponnetto and Ernesto De Vito. Optimal rates for the regularized least-squares algorithm.Foundations of Computational Mathematics, 7(3):331–368, 2007

  15. [15]

    Vector valued reproducing ker- nel Hilbert spaces of integrable functions and Mercer theorem.Analysis and Applications, 4(04):377–408, 2006

    Claudio Carmeli, Ernesto De Vito, and Alessandro Toigo. Vector valued reproducing ker- nel Hilbert spaces of integrable functions and Mercer theorem.Analysis and Applications, 4(04):377–408, 2006

  16. [16]

    Vector valued reproducing kernel Hilbert spaces and universality.Analysis and Applications, 8(01):19–61, 2010

    Claudio Carmeli, Ernesto De Vito, Alessandro Toigo, and Veronica Umanit´ a. Vector valued reproducing kernel Hilbert spaces and universality.Analysis and Applications, 8(01):19–61, 2010

  17. [17]

    Nonlinear approximation and the space bv(r2).American Journal of Mathematics, 121(3):587–628, 1999

    Albert Cohen, Ronald DeVore, Pencho Petrushev, and Hong Xu. Nonlinear approximation and the space bv(r2).American Journal of Mathematics, 121(3):587–628, 1999

  18. [18]

    Ap- proximation methods for supervised learning.Foundations of Computational Mathematics, 6(1):3–58, 2006

    Ronald DeVore, Gerard Kerkyacharian, Dominique Picard, and Vladimir Temlyakov. Ap- proximation methods for supervised learning.Foundations of Computational Mathematics, 6(1):3–58, 2006

  19. [19]

    Applications of Mathematics, Springer, New York, 1996

    Luc Devroye, L´ aszl´ o Gy¨ orfi, and G´ abor Lugosi.A probabilistic theory of pattern recognition, volume 31. Applications of Mathematics, Springer, New York, 1996

  20. [20]

    Injectivity and weak*-to-weak continuity suffice for con- vergence rates inℓ 1-regularization.Journal of Inverse and Ill-posed Problems, 26(1):85–94, 2018

    Jens Flemming and Daniel Gerth. Injectivity and weak*-to-weak continuity suffice for con- vergence rates inℓ 1-regularization.Journal of Inverse and Ill-posed Problems, 26(1):85–94, 2018

  21. [21]

    Springer, 1998

    David Gilbarg, Neil S Trudinger, David Gilbarg, and NS Trudinger.Elliptic partial differential equations of second order, volume 2. Springer, 1998

  22. [22]

    Optimally sparse multidimensional representation using shearlets.SIAM journal on mathematical analysis, 39(1):298–318, 2007

    Kanghui Guo and Demetrio Labate. Optimally sparse multidimensional representation using shearlets.SIAM journal on mathematical analysis, 39(1):298–318, 2007

  23. [23]

    Nonparametric goodness-of-fit testing for parametric covariate models in pharmacometric analyses.CPT: Pharmacometrics & Systems Pharmacology, 10(6):564–576, 2021

    Niklas Hartung, Martin Wahl, Abhishake Rastogi, and Wilhelm Huisinga. Nonparametric goodness-of-fit testing for parametric covariate models in pharmacometric analyses.CPT: Pharmacometrics & Systems Pharmacology, 10(6):564–576, 2021

  24. [24]

    Optimal convergence rates for sparsity promoting wavelet- regularization in besov spaces.Inverse Problems, 35(6):065005, 2019

    Thorsten Hohage and Philip Miller. Optimal convergence rates for sparsity promoting wavelet- regularization in besov spaces.Inverse Problems, 35(6):065005, 2019

  25. [25]

    Springer- Verlag, Berlin Heidelberg, 1985

    Lars H¨ ormander.The Analysis of Linear Partial Differential Operators, volume IV. Springer- Verlag, Berlin Heidelberg, 1985. 46

  26. [26]

    Inverse problems involving pdes with applications to imaging

    Taufiquar Khan. Inverse problems involving pdes with applications to imaging. In Pammy Manchanda, Ren´ e Pierre Lozi, and Abul Hasan Siddiqi, editors,Mathematical Modelling, Optimization, Analytic and Numerical Solutions, pages 181–195. Springer, Singapore, 2020

  27. [27]

    Shearlet smoothness spaces.Journal of Fourier Analysis and Applications, 19(3):577–611, 2013

    Demetrio Labate, Lucia Mantovani, and Pooran Negi. Shearlet smoothness spaces.Journal of Fourier Analysis and Applications, 19(3):577–611, 2013

  28. [28]

    Optimal rates for spec- tral algorithms with least-squares regression over Hilbert spaces.Applied and Computational Harmonic Analysis, 48(3):868–890, 2020

    Junhong Lin, Alessandro Rudi, Lorenzo Rosasco, and Volkan Cevher. Optimal rates for spec- tral algorithms with least-squares regression over Hilbert spaces.Applied and Computational Harmonic Analysis, 48(3):868–890, 2020

  29. [29]

    Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri

    L. Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri. Spectral algorithms for supervised learning.Neural Computation, 20(7):1873–1897, 2008

  30. [30]

    Elsevier, 1999

    St´ ephane Mallat.A wavelet tour of signal processing. Elsevier, 1999

  31. [31]

    Regularization in kernel learning.The Annals of Statistics, 38(1):526 – 565, 2010

    Shahar Mendelson and Joseph Neeman. Regularization in kernel learning.The Annals of Statistics, 38(1):526 – 565, 2010

  32. [32]

    On learning vector-valued functions.Neural Computation, 17(1):177–204, 2005

    Charles A Micchelli and Massimiliano Pontil. On learning vector-valued functions.Neural Computation, 17(1):177–204, 2005

  33. [33]

    Maximal spaces for approximation rates inℓ 1- regularization.Numerische Mathematik, 149(2):341–374, 2021

    Philip Miller and Thorsten Hohage. Maximal spaces for approximation rates inℓ 1- regularization.Numerische Mathematik, 149(2):341–374, 2021

  34. [34]

    SIAM, Philadelphia, 2012

    Jennifer L Mueller and Samuli Siltanen.Linear and nonlinear inverse problems with practical applications. SIAM, Philadelphia, 2012

  35. [35]

    SIAM, 2001

    Frank Natterer.The mathematics of computerized tomography. SIAM, 2001

  36. [36]

    Wainwright, and Bin Yu

    Garvesh Raskutti, Martin J. Wainwright, and Bin Yu. Minimax rates of estimation for high-dimensional linear regression overℓ q-balls.IEEE Transactions on Information Theory, 57(10):6976–6994, 2011

  37. [37]

    Convergence analysis of Tikhonov regularization for non-linear statistical inverse problems.Electronic Journal of Statistics, 14(2):2798–2841, 2020

    Abhishake Rastogi, Gilles Blanchard, and Peter Math´ e. Convergence analysis of Tikhonov regularization for non-linear statistical inverse problems.Electronic Journal of Statistics, 14(2):2798–2841, 2020

  38. [38]

    Inverse learning in Hilbert scales.Machine Learning, 112:2469–2499, 2023

    Abhishake Rastogi and Peter Math´ e. Inverse learning in Hilbert scales.Machine Learning, 112:2469–2499, 2023

  39. [39]

    Optimal rates for the regularized learning algorithms under general source condition.Frontiers in Applied Mathematics and Statistics, 3:3, 2017

    Abhishake Rastogi and Sivananthan Sampath. Optimal rates for the regularized learning algorithms under general source condition.Frontiers in Applied Mathematics and Statistics, 3:3, 2017

  40. [40]

    Walter de Gruyter GmbH & Co

    Thomas Schuster, Barbara Kaltenbacher, Bernd Hofmann, and Kamil S Kazimierski.Reg- ularization methods in Banach spaces, volume 10 ofRadon Series on Computational and Applied Mathematics. Walter de Gruyter GmbH & Co. KG, Berlin, 2012

  41. [41]

    Physics-informed neural network for ultrasound nondestructive quantification of surface breaking cracks.Journal of Nondestructive Evaluation, 39(3):61, 2020

    Khemraj Shukla, Patricio Clark Di Leoni, James Blackshire, Daniel Sparkman, and George Em Karniadakis. Physics-informed neural network for ultrasound nondestructive quantification of surface breaking cracks.Journal of Nondestructive Evaluation, 39(3):61, 2020

  42. [42]

    Optimal rates for regularized least squares regression

    Ingo Steinwart, Don Hush, and Clint Scovel. Optimal rates for regularized least squares regression. InS. Dasgupta and A. Klivans, editors, Proceedings of the 22nd Annual Conference on Learning Theory, pages 79–93, 2009

  43. [43]

    Regression shrinkage and selection via the lasso.Journal of the Royal Statistical Society Series B: Statistical Methodology, 58(1):267–288, 1996

    Robert Tibshirani. Regression shrinkage and selection via the lasso.Journal of the Royal Statistical Society Series B: Statistical Methodology, 58(1):267–288, 1996. 47

  44. [44]

    Birkh¨ auser Verlag, Basel, 1983

    Hans Triebel.Theory of function spaces. Birkh¨ auser Verlag, Basel, 1983

  45. [45]

    Introduction to nonparametric estimation

    Alexandre Tsybakov. Introduction to nonparametric estimation. InSpringer Series in Statis- tics, 2008

  46. [46]

    Van Aarle, W

    W. Van Aarle, W. J. Palenstijn, J. Cant, E. Janssens, F. Bleichrodt, A. Dabravolski, J. De Beenhouwer, K. J. Batenburg, and J. Sijbers. Fast and flexible X-ray tomography using the ASTRA toolbox.Optics express, 24(22):25129–25147, 2016

  47. [47]

    van der Vaart and Jon A

    Aad W. van der Vaart and Jon A. Wellner.Weak convergence and empirical processes. Springer Series in Statistics, Springer-Verlag, New York, 1996. With Applications to Statistics. 48