{"id":"abc628f9-da4c-434d-8167-5617f803ad79","arxiv_id":"2607.07468","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":7,"one_line_summary":"The ℓ¹-regularized empirical risk minimizer achieves minimax-optimal convergence rate n^{-r/(1+b-br)} for nonlinear statistical inverse learning under variational source conditions and polynomial effective-dimension decay.","lead":"This paper proves optimal convergence rates for recovering sparse signals from noisy, indirect observations using ℓ¹-regularization in a statistical inverse learning framework. It matters because it provides theoretical guarantees for sparse recovery in nonlinear inverse problems like computed tomography, bridging compressed sensing and statistical learning theory.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Minimax optimality hinges on whether the k_t-to-VSC bridge closes the upper/lower class gap under the same weight w; the construction is sound but the r-range mismatch deserves explicit verification.","rationale":"The reader correctly identified Assumption 6 as the weakest link, and my analysis confirms this is where the argument is most delicate: the weighted bi-Lipschitz property is essential for the k_t → VSC bridge (Theorem 5.7), which in turn is essential for the lower bound to apply to the same class P_{r,b} as the upper bound. However, upon careful examination, the argument is internally sound for r ∈ (0,1/2): the forward implication (k_t → VSC) is proved, and this is the only direction needed for the minimax claim. The converse (VSC → k_t) is mentioned but not proved, and is not needed for the minimax result as stated in Corollary 5.8 (which restricts to r ∈ (0,1/2)). The restriction to finitely smoothing operators (verified in Section 6.1) is a genuine scope limitation but is openly acknowledged and does not affect the correctness of the theory within its stated domain. The proofs follow established techniques (Caponnetto–De Vito, Blanchard–Mücke, Miller–Hohage) adapted to the Banach-space ℓ¹ setting. The upper and lower rates match: n^{-r/(1+b-br)} in ℓ¹ norm (p=1). No data-driven parameter selection or numerical validation is provided, but these are acknowledged as future work and do not undermine the theoretical contribution. The verdict of ACCEPT with HIGH confidence is appropriate.","tokens_in":43368,"tokens_out":1089,"duration_ms":707658,"concrete_test":"Independently verify the chain: starting from f_ρ ∈ k_t with ||f_ρ||_{k_t} ≤ ϱ, apply Lemma 5.6 to obtain inequality (50) with constant γ=γ(ϱ,t), then apply Assumption 6(ii) to substitute ||f_ρ - f||_{w,2} ≤ L||A(f)-A(f_ρ)||_{H_μ}. Check that the resulting constant C = γL^{(2-2t)/(2-t)} is finite and that the exponent r=(1-t)/(2-t) correctly maps (0,1) to (0,1/2). If any step fails, the VSC verification—and hence the lower bound's applicability to P_{r,b}—is compromised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The minimax claim (Corollary 5.8) requires that the upper-bound class (distributions satisfying VSC with φ(s)=s^r, Assumption 5) and the lower-bound class (functions in the approximation space k_t, used in Theorem 5.4) coincide. The bridge is Theorem 5.7: under Assumption 6(ii), f_ρ ∈ k_t implies VSC with r=(1-t)/(2-t) ∈ (0,1/2). The paper states (end of Section 5.2) that the converse also holds under Assumption 6(i), citing the same theorem. However, Theorem 5.7 as proved only establishes the forward direction (k_t → VSC). The converse direction (VSC with φ(s)=s^r → f_ρ ∈ k_t with t=2r-1/(r-1)) is asserted but not proved in the text. This matters because the upper bounds in Corollary 4.7 are stated for r ∈ (0,1) (via Assumption 5 directly), while the lower bound construction in Theorem 5.4 only covers r ∈ (0,1/2) (since t ∈ (0,1) maps to r ∈ (0,1/2)). For r ∈ [1/2,1), the upper bound applies but no lower bound is constructed, so minimax optimality is unverified in that range. The paper does note this restriction in Remark 5.10, but Corollary 5.8 states r ∈ (0,1/2), which is consistent. The real question is whether the unproved converse direction is actually needed for the minimax claim. Since the upper bound is proved for all ρ ∈ P_{r,b} (which requires VSC, not k_t membership), and the lower bound constructs specific ρ* ∈ P_{r,b} (which requires k_t membership → VSC, the proved direction), the classes do match: the lower bound's ρ* lies in P_{r,b} via the proved forward implication, and the upper bound covers all of P_{r,b}. So the converse is not actually needed for the minimax claim as stated. The concern reduces to whether the forward implication (Theorem 5.7) is correctly proved, which hinges on Lemma 5.6 (the k_t ↔ weighted-ℓ² inequality) and Assumption 6(ii). Both are standard adaptations from Miller–Hohage [33]. The argument appears sound.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper studies the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning. The unknown is modeled as an element of ℓ¹, and observations are generated through a possibly nonlinear forward operator A: ℓ¹ → H, where H is a vector-valued reproducing kernel Hilbert space (vv-RKHS). The authors propose an ℓ¹-regularized empirical risk minimizer and establish almost-sure consistency, non-asymptotic high-probability convergence rates in both prediction and ℓ¹ reconstruction norms, and matching minimax lower bounds. The rates depend on a source smoothness parameter r (characterized by a variational source condition) and an effective dimension exponent b (polynomial spectral decay of the covariance operator). The theory is connected to practical sparsity models via approximation spaces k_t, and the assumptions are verified for two representative inverse problems: reaction coefficient identification in elliptic PDEs and sparse computed tomography.","tokens_in":43608,"tokens_out":1794,"duration_ms":349999,"significance":"The paper makes a substantial contribution by extending the statistical inverse learning framework to ℓ¹-regularization in a Banach-space setting with possibly nonlinear forward operators. The minimax optimality of the derived rates n^{-r/(1+b-br)} (in ℓ¹ reconstruction norm) is a central and significant claim, established through matching upper and lower bounds over the prior class P_{r,b}. The connection between approximation spaces k_t, variational source conditions, and best n-term approximation errors (Theorem 5.7, Lemma 5.9) provides a concrete and verifiable bridge between sparse approximation theory and statistical convergence rates. The application to filtered Radon transforms, including explicit effective-dimension asymptotics for boxcar and Gaussian filters (Proposition 6.5), yields concrete convergence rates for standard image models and sparsifying systems, adding practical value to the theoretical contributions.","major_comments":[{"comment":"§5.2, Theorem 5.7 and the text following it: The converse direction of the k_t-to-VSC bridge is asserted but not proved. The text states: 'Conversely, again under Assumption 6(i), if f_ρ satisfies the variational source condition of Assumption 5 with φ(s)=s^r, then f_ρ belongs to the smoothness space k_t with t=(2r-1)/(r-1).' However, Theorem 5.7 only establishes the forward direction (k_t → VSC). While the skeptic's analysis confirms that the minimax claim in Corollary 5.8 (restricted to r ∈ (0,1/2)) does not logically depend on the unproved converse—since the lower bound constructs specific ρ* ∈ P_{r,b} via the proved forward direction—the asserted converse should either be proved, stated as a conjecture, or removed. As written, it could mislead readers into thinking the equivalence is fully established within the paper.","section":null},{"comment":"§5.2, Corollary 5.8 vs. §4, Corollary 4.7: There is a mismatch in the range of r for which minimax optimality is claimed. The upper bound (Corollary 4.7) is stated for r ∈ (0,1), while the lower bound (Theorem 5.4) and the minimax optimality claim (Corollary 5.8) are restricted to r ∈ (0,1/2). Remark 5.10 acknowledges this restriction, but the abstract and introduction state that 'matching minimax lower bounds' are established without clearly communicating this restriction on r. The abstract should be amended to accurately reflect that minimax optimality is established for r ∈ (0,1/2), or the scope of the lower bound should be extended.","section":null},{"comment":"§2.5, Assumption 6: The weighted bi-Lipschitz property (both parts (i) and (ii) with the same weight w) is load-bearing for the minimax optimality claim (Corollary 5.8), as the authors note. However, the verification in §6.1 for finitely smoothing operators A = G ∘ S shows that parts (i) and (ii) hold with different weights in the shearlet case (w_λ = 2^{-|λ|(a+1/2)} for (i) and w_λ = 2^{-|λ|(a-1/2)} for (ii)). The minimax optimality result requires the same weight w for both parts. The paper should clarify whether the minimax claim applies to the shearlet case, or whether it is limited to cases (like the wavelet case or the direct synthesis operator in §6.4) where the weights coincide.","section":null}],"minor_comments":[{"comment":"§1.1, Table 1: The 'Rate' column for the present work lists 'n^{-r/(1+b-br)}', which corresponds to the ℓ¹ reconstruction norm rate (p=1). The table would benefit from also listing the general interpolation norm rate n^{-(2r-pr+p-1)/(p(1+b-br))} for completeness, or noting that the displayed rate is the special case p=1.","section":null},{"comment":"§4, Corollary 4.6, Eq. (28): The exponent in the log factor is written as log^{2/p}(4/η), but the derivation from Theorem 4.3 (which has log²(4/η)) should make this log^{2/p}(4/η). This appears correct but could be stated more explicitly for the reader.","section":null},{"comment":"§6.3, Proposition 6.5: The boxcar filter case yields b = 2/3, and the Gaussian filter case yields any b ∈ (0,1). It would be helpful to explicitly state the resulting convergence rates (as done at the end of §6.3 for specific examples) in terms of the general formula n^{-r/(1+b-br)} for these two filter choices, to make the practical implications more immediately visible.","section":null},{"comment":"§3, proof of Theorem 3.1: The application of Scheffé's lemma to conclude strong convergence from weak-* convergence and norm convergence in ℓ¹ is correct but could benefit from a brief justification or citation, as this is a less commonly used tool in this context.","section":null},{"comment":"§2.7, Definition 2.4: The class P_{r,b} is defined with 0 ≤ r ≤ 1, but the main results (e.g., Corollary 4.5) require 0 < r < 1. The boundary cases r = 0 and r = 1 should be discussed or excluded from the definition for consistency.","section":null},{"comment":"§5.1, Theorem 5.3: The condition on the weight sequence (49) involves constants c_0, ε_0, q. Remark 5.5 verifies this for polynomially decaying weights w_m = m^{-a}, but the relationship between the constant a and the parameters r, b is not fully explicit. Stating the constraint on a in terms of r and b would improve clarity.","section":null},{"comment":"Typographical: §6.1, the sentence 'In the shearlet case, using H^{-a} ↪ S^{-a-1/2}_{2,2}, thus and Assumption 6(ii) is verified with w_λ = 2^{-|λ|(a+1/2)}' contains a grammatical error ('thus and').","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a strong contribution to the statistical inverse learning literature. The core mathematical content is sound, and the minimax optimality claim is well-supported for r ∈ (0,1/2). The main issues are presentational: the unproved converse direction of the k_t-to-VSC bridge is stated in a way that could mislead, and the r-range mismatch between upper and lower bounds is not clearly communicated in the abstract. The shearlet weight mismatch for Assumption 6 is a more substantive concern, but it appears to limit the scope of the minimax claim rather than invalidate it. I recommend minor revision to address these points."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful reading and three substantive comments, all of which are correct. We address each below and describe the revisions we will make.","responses":[{"response":"The referee is correct. Theorem 5.7 proves only the forward implication: membership in k_t (together with Assumption 6(ii)) implies the variational source condition with r = (1-t)/(2-t). The converse statement — that a variational source condition with φ(s) = s^r implies membership in k_t with t = (2r-1)/(r-1) under Assumption 6(i) — appears in the text following Theorem 5.7 without a proof. We do not currently have a complete proof of this converse direction. As the referee notes, the minimax claim in Corollary 5.8 does not logically depend on the converse: the lower bound construction in Theorem 5.4 uses the proved forward direction to exhibit specific ρ* ∈ P_{r,b}. Nevertheless, the unqualified assertion of the converse is misleading as written. We will revise the text to state the converse explicitly as a conjecture (Conjecture 5.8 or similar), clearly marking it as unproved, and will adjust the surrounding discussion to avoid any implication that the equivalence is fully established within the paper.","revision_made":"yes","referee_comment":"§5.2, Theorem 5.7 and the text following it: The converse direction of the k_t-to-VSC bridge is asserted but not proved. The text states the converse but Theorem 5.7 only establishes the forward direction. The asserted converse should either be proved, stated as a conjecture, or removed."},{"response":"The referee is correct. The upper convergence rate (Corollary 4.7) is established for r ∈ (0,1), while the minimax lower bound (Theorem 5.4) and the minimax optimality statement (Corollary 5.8) are restricted to r ∈ (0,1/2). This restriction arises because the mapping t ↦ r = (1-t)/(2-t) from the approximation space k_t to the source condition index r has range (0,1/2), and the lower bound construction relies on this connection. Remark 5.10 acknowledges this, but the abstract and introduction do not. We will amend the abstract to read: 'We further prove matching minimax lower bounds for r ∈ (0,1/2), showing that the obtained convergence rates are optimal in this regime.' We will make a corresponding adjustment in the introduction (contribution (iii)) and in the statement of Corollary 5.8 to make the restriction on r explicit and prominent.","revision_made":"yes","referee_comment":"§5.2, Corollary 5.8 vs. §4, Corollary 4.7: There is a mismatch in the range of r for which minimax optimality is claimed. The upper bound is stated for r ∈ (0,1), while the lower bound and minimax optimality claim are restricted to r ∈ (0,1/2). The abstract and introduction state that 'matching minimax lower bounds' are established without clearly communicating this restriction on r."},{"response":"The referee is correct. The minimax optimality result in Corollary 5.8 requires both parts of Assumption 6 to hold with the same weight w. In the wavelet case (Section 6.1), both parts (i) and (ii) are verified with w_λ = 2^{-|λ|a}, so the minimax claim applies. In the shearlet case, part (i) holds with w_λ = 2^{-|λ|(a+1/2)} and part (ii) holds with w_λ = 2^{-|λ|(a-1/2)}, which are different. Consequently, the minimax optimality result of Corollary 5.8 does not directly apply to the shearlet case as currently stated. The upper convergence rates (Corollary 4.7) still hold for the shearlet case, since they require only Assumption 6(ii), but the matching lower bound is not established in this setting. We will add a clarifying remark in Section 6.1 explicitly stating that: (a) the minimax optimality claim applies to the wavelet case and to the direct synthesis operator case (Section 6.4), where the weights coincide; (b) for the shearlet case, the upper rates remain valid but the minimax lower bound is not established, because the two parts of Assumption 6 hold with different weights. We view extending the lower bound to the shearlet case — either by refining the construction to accommodate distinct weights or by identifying a common weight under which both bounds hold — as an interesting open problem, which we will mention.","revision_made":"yes","referee_comment":"§2.5, Assumption 6: The weighted bi-Lipschitz property requires the same weight w for both parts (i) and (ii), but the verification in §6.1 for finitely smoothing operators A = G ∘ S shows that parts (i) and (ii) hold with different weights in the shearlet case. The minimax optimality result requires the same weight w for both parts. The paper should clarify whether the minimax claim applies to the shearlet case, or whether it is limited to cases where the weights coincide."}],"tokens_in":43146,"tokens_out":1894,"duration_ms":80910,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main result: this paper establishes minimax-optimal convergence rates n^{-r/(1+b-br)} for ℓ¹-regularized statistical inverse learning with nonlinear forward operators, with matching lower bounds. The combination of sparsity-promoting regularization, nonlinear forward maps, Banach-space reconstruction, and minimax lower bounds is genuinely new. Table 1 positions the contribution clearly against prior work (Caponnetto-De Vito, Blanchard-Mücke, Rastogi et al., Miller-Hohage). The rate depends on the source smoothness r (via variational source conditions) and the effective dimension exponent b (polynomial spectral decay), which is the right parameterization for this setting. The proofs follow established patterns — variational inequalities, concentration bounds, Fano-type lower bounds — adapted carefully to the ℓ¹/Banach-space setting. The Fenchel conjugate step in Theorem 4.3 is clean. The two concrete applications (elliptic PDE coefficient identification, filtered Radon transform) are worked out with explicit effective-dimension asymptotics, which gives the theory real grounding rather than leaving it purely abstract. The connection between approximation-space sparsity (k_t membership ↔ best n-term approximation decay) and variational source conditions via Theorem 5.7 and Lemma 5.9 is a useful bridge between sparse approximation theory and statistical inverse learning. On the stress-test concern about whether the k_t-to-VSC bridge closes the upper/lower class gap: the concern does not land as a problem for the minimax claim. The forward direction (k_t → VSC) is what Theorem 5.7 proves, and that is the direction needed — the lower bound constructs ρ* ∈ P_{r,b} via this forward implication, and the upper bound covers all of P_{r,b}. The converse (VSC → k_t) is asserted in the text but not proved; however, it is not actually needed for Corollary 5.8 as stated. The minimax claim is restricted to r ∈ (0, 1/2), which is consistent with the range of the k_t → VSC mapping. This is acknowledged in Remark 5.10. So the circularity burden is low. The soft spots are real but proportionate. Assumption 6 (weighted bi-Lipschitz) is the main structural limitation — it excludes severely ill-posed problems and is verified only for finitely smoothing operators. The optimal λ* depends on unknown r and b with no data-driven procedure. The numerical content is minimal (one SVD decay plot). The unproved converse direction should either be proved or flagged as conjectural rather than stated as established. None of these undermine the central result. This paper is for researchers in statistical learning theory and inverse problems who work on convergence rates and sparsity. It deserves a serious referee. The proofs are detailed, the result is new, and the applications are concrete enough to verify the abstract assumptions.","headline":"Minimax-optimal rates for ℓ¹-regularized nonlinear statistical inverse learning; proofs are sound and the k_t-to-VSC bridge holds for the minimax claim as stated.","tokens_in":44365,"tokens_out":1188,"would_cite":true,"duration_ms":103102,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Sparse recovery rates for nonlinear inverse problems proven minimax-optimal","keywords":[],"falsifier":"Construct a probability distribution in P_{r,b} for which the ℓ¹-regularized estimator with λ* = n^{-(1-r)/(1+b-br)} converges slower than n^{-r/(1+b-br)}, or exhibit any learning algorithm achieving a faster rate over the full class P_{r,b}.","tokens_in":43435,"feed_emoji":"📐","tokens_out":1068,"duration_ms":148911,"temperature":0.7,"pith_summary":"This paper studies the problem of recovering a sparse, infinite-dimensional signal from finitely many noisy and indirect observations, as arise in computed tomography or PDE coefficient identification. The authors propose an ℓ¹-regularized empirical risk minimizer in a vector-valued reproducing kernel Hilbert space framework, where the forward operator mapping the unknown to the data can be nonlinear. The central result is a complete convergence-rate theory: under a variational source condition (characterizing solution smoothness via a parameter r) and polynomial spectral decay of the covariance operator (characterizing statistical complexity via an exponent b), the estimator achieves the reconstruction rate n^{-r/(1+b-br)} in ℓ¹ norm. The authors prove matching minimax lower bounds, establishing that no estimator can converge faster over the same class of problems. They further show that membership in an approximation space k_t — equivalently, polynomial decay of best n-term approximation errors — implies the required variational source condition, bridging sparse approximation theory with statistical inverse learning. The assumptions are verified concretely for reaction coefficient identification in elliptic PDEs and for filtered Radon transforms in computed tomography, yielding explicit rates for standard image models and sparsifying systems such as wavelets and shearlets.","feed_headline":"Sparse recovery rates for nonlinear inverse problems proven optimal","feed_subtitle":"ℓ¹-regularized estimators achieve minimax-optimal convergence for CT and PDE inverse problems, with rates set by sparsity and spectral decay","key_machinery":"The variational source condition (Assumption 5) with index function φ(t) = t^r, which quantifies how well the true solution can be approximated relative to the forward operator's sensitivity; the effective dimension N(λ) with polynomial bound N(λ) ≤ Cλ^{-b}, capturing statistical complexity via covariance spectral decay; the weighted bi-Lipschitz property (Assumption 6), which provides two-sided stability of the forward operator between weighted sequence space and data space; the approximation space kₘ","core_discovery":"The minimax-optimal convergence rate for ℓ¹-regularized sparse statistical inverse learning with nonlinear forward operators is n^{-r/(1+b-br)} in ℓ¹ reconstruction norm, where r encodes the solution's sparsity-driven smoothness (via a variational source condition) and b encodes the effective dimension of the learning problem (via spectral decay of the covariance operator). This rate is achieved by the regularized estimator with parameter λ* = n^{-(1-r)/(1+b-br)} and cannot be improved by any estimator over the prior class P_{r,b}. A key structural finding is that the variational source condition — the analytical engine for the rates — is equivalent to membership in an approximation space kₜ","pith_inferences":[],"forward_implications":["For filtered Radon transforms with a boxcar filter, the effective dimension exponent is b = 2/3; with a Gaussian filter, eigenvalues decay super-polynomially so any b > 0 is admissible, yielding near-parametric rates n^{-r} for sufficiently sparse signals.","For cartoon-like images represented in shearlet frames (best n-term approximation rate n^{-1}), the theory predicts r = 1/4, giving concrete convergence rates of n^{-1/6} (boxcar) or n^{-1/4} (Gaussian) in ℓ¹ reconstruction norm.","The framework recovers classical kernel ridge regression results as a degenerate case when the forward operator is the synthesis identity, with the weighted bi-Lipschitz property holding as an equality and the same rate structure governing both sparse and non-sparse regimes.","The equivalence chain σ_n(f) = O(n^{1/2-1/t}) ⟺ f ∈ k_t ⟹ VSC with φ(s) = s^{(1-t)/(2-t)} provides a practical diagnostic: measuring best n-term approximation decay for a given signal class directly determines the achievable statistical convergence rate.","The parameter choice λ* = n^{-(1-r)/(1+b-r)} depends on unknown r and b, but the dual-function characterization in Corollary 4.4 suggests a data-driven balancing principle analogous to the discrepancy principle, potentially enabling adaptive selection without prior knowledge of smoothness."],"fun_headline_variants":["ℓ¹-regularized sparse recovery achieves minimax-optimal rates in inverse learning","Optimal convergence rates proven for ℓ¹-regularized statistical inverse learning","Sparse inverse learning: ℓ¹ regularization is minimax-optimal under nonlinear operators","Minimax-optimal rates for ℓ¹-regularized sparse inverse problems","ℓ¹ regularization attains optimal rates for nonlinear statistical inverse learning"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The weighted bi-Lipschitz property (Assumption 6) requires the forward operator to satisfy a lower Lipschitz bound — meaning small changes in the data must reflect proportionally small changes in the solution. This fails for severely ill-posed problems like the backward heat equation or electrical impedance tomography, restricting the theory to finitely smoothing operators.","fun_headline_variants_meta":{"raw":{"variants":["ℓ¹-regularized sparse recovery achieves minimax-optimal rates in inverse learning","Optimal convergence rates proven for ℓ¹-regularized statistical inverse learning","Sparse inverse learning: ℓ¹ regularization is minimax-optimal under nonlinear operators","Minimax-optimal rates for ℓ¹-regularized sparse inverse problems","ℓ¹ regularization attains optimal rates for nonlinear statistical inverse learning","ℓ¹-regularized estimators match minimax lower bounds for sparse inverse learning","Optimal sparse recovery rates for nonlinear inverse problems via ℓ¹ regularization","Variational source conditions shown equivalent to approximation space membership","ℓ¹-regularized estimators provably minimax-optimal for sparse CT and PDE recovery"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":1490,"prompt_tokens":640,"completion_tokens":850,"prompt_tokens_details":null},"tokens_in":640,"tokens_out":850,"duration_ms":50855,"temperature":1.0,"reasoning_tokens":708,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T09:55:48.034715+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Construct a probability distribution in P_{r,b} for which the ℓ¹-regularized estimator with λ* = n^{-(1-r)/(1+b-br)} converges slower than n^{-r/(1+b-br)}, or exhibit any learning algorithm achieving a faster rate over the full class P_{r,b}.","supporting_citations":[],"review_version":1}