Pith. sign in

REVIEW 4 major objections 5 minor 47 references

The paper proves that the nonconvex variable-projection objective in multi-snapshot spike deconvolution has an explicit, computable basin of convexity—a ball around the true spike locations whose radius is determined by PSF spectral flatnes

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 07:30 UTC pith:DYKS4A54

load-bearing objection A serious, novel basin-of-convexity analysis for VarProSD with a real new result, but the central conditioning bounds ride on an unpublished preprint and a couple of lemmas have missing hypotheses — worth a careful referee, not a desk reject. the 4 major comments →

arxiv 2607.09593 v2 pith:DYKS4A54 submitted 2026-07-10 stat.ML cs.NAeess.SPmath.NAmath.OC

Characterization of the Basin of Convexity for Multi-Snapshot Spike Deconvolution via Variable Projection

classification stat.ML cs.NAeess.SPmath.NAmath.OC MSC 65K1065T9994A1294A2094A08
keywords spike deconvolutionvariable projectionbasin of convexityBeurling-Selberg approximationgradient descentmulti-snapshotpoint spread functionlocal convergence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that the nonconvex least-squares problem over spike locations—obtained by eliminating amplitudes via variable projection—is strongly convex on a ball around the true locations whose radius is given in closed form. The radius is set by the flatness of the point spread function's power spectrum, the minimum spike separation, the amplitude dynamic range, and the sampling bandwidth. Within this ball, a unique local minimizer exists, gradient descent converges to it linearly, and the minimizer's error decays as one over the square root of the number of snapshots under stochastic noise; an adversarial-noise bound is also derived. If true, this is the first interpretable basin-of-convexity guarantee for arbitrary point spread functions under multiple snapshots, and it yields a practical way to select the sampling bandwidth.

Core claim

The central claim is Theorem 1: when the spikes are separated by more than 2/(3 ρ κ²) and the residual covariance satisfies the smallness condition (16), the VarProSD objective ℓ(γ)= (1/2L)∥P⊥_γ Y∥_F² is strongly convex with Lipschitz gradient on the ball N(τ,ϱ) defined in (14)–(15), so a unique local minimizer γ⋆ lies in that ball. Theorem 6 adds that gradient descent from any initialization in the ball converges linearly to γ⋆, and Theorem 2 bounds the recovery error by c₂√K √(Eg/Eg′) ∥R∥/(T Eg r_min(X)²), which under i.i.d. Gaussian noise becomes O(1/√L). The paper also constructs a worst-case rank-one adversarial noise and shows the resulting deterministic bound is √K sharper in the high

What carries the argument

The argument rests on the variable-projection objective and the structured Vandermonde matrix A_γ = G Φ_γ, where G is the diagonal PSF sampling matrix and Φ_γ is the Fourier-Vandermonde matrix of spike locations. Its curvature is controlled through Beurling–Selberg extremal approximations: bandlimited majorants and minorants of the truncated power spectral density of the PSF yield sharp bounds on the singular values of A_γ, Λ A_γ, and the derived matrix S_γ = A_γ^H Λ^H P⊥_γ Λ A_γ (Lemmas 9–10). These bounds, independent of the number of spikes K, translate directly into the strong-convexity constant and the gradient Lipschitz constant of the objective on the ball, and hence into the contract

Load-bearing premise

The proof's load-bearing premise is that the truncated power spectral density of the point spread function and its first two derivatives are of bounded variation and integrable with an f² weight, so that the Beurling–Selberg majorant/minorant bounds on the conditioning of the structured matrices carry through; if that spectral regularity fails, the basin radius and curvature bounds collapse.

What would settle it

Evaluate the smallest eigenvalue of the Hessian of ℓ(γ) at a point γ on the sphere d₂(γ,τ)=ϱ for the Gaussian-PSF setup of Figure 2(b) (σ=0.3, K=2, Δ=0.4); if it is below the strong-convexity constant (1/3)Eg′Tr_min(X)², the claimed basin radius is too large.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any initialization inside the explicit ball N(τ,ϱ) is guaranteed to yield linear convergence to the unique local minimizer, so the radius quantifies exactly how accurate an initialization must be.
  • Under i.i.d. Gaussian noise, the local minimizer achieves consistency with O(1/√L) error decay in the number of snapshots, matching the rate of ESPRIT for trivial PSFs but now for an arbitrary PSF.
  • The spectral parameter ρ gives a principled, PSF-driven rule for selecting sampling bandwidth: the bandwidth minimizing ρ predicts the empirically optimal B for Gaussian and Morlet PSFs, and larger B is not always better because it shrinks the basin.
  • The adversarial-noise analysis shows that a rank-one perturbation aligned with the top singular vector of the inverse-map Jacobian is asymptotically worst case, and the resulting deterministic error bound is sharper by a factor of √K than the stochastic bound in the high-SNR large-L regime.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Our inference: the same Beurling–Selberg conditioning machinery could extend to single-snapshot or unknown-PSF settings, where no explicit basin characterization currently exists.
  • Our inference: the residual-covariance condition (16), which the paper assumes but does not verify in experiments, is the most likely place for the theory to be conservative; a testable extension is to check empirically how tightly (16) holds at the bandwidths predicted optimal.
  • Our inference: the inverse-map Lipschitz analysis is a general template—the explicit Jacobian expression could be adapted to other separable nonlinear least-squares problems to derive adversarial stability bounds.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies multi-snapshot spike deconvolution with a known point-spread function (PSF). It adopts a variable-projection formulation (VarProSD) in which the amplitudes are eliminated, leaving a nonconvex least-squares problem over the spike locations. The main contribution is an explicit characterization of a basin of convexity around the ground truth: Theorem 1 states that, under a minimum-separation condition on the spikes, a spectral-regularity condition on the PSF, and a residual-covariance condition, the objective is strongly convex with Lipschitz gradient on a ball and admits a unique local minimizer. Theorem 2 gives an estimation-error bound for that minimizer under stochastic noise, Theorem 3 gives a deterministic adversarial-noise bound via the local Lipschitz property of the inverse map, and Theorem 6 establishes local linear convergence of gradient descent. The proofs rely on Beurling–Selberg extremal approximations to bound the conditioning of generalized Vandermonde matrices. Numerical experiments illustrate the theory and propose a PSF-driven sampling-bandwidth selection criterion.

Significance. If correct, this is the first explicit, interpretable basin-of-convexity guarantee for multi-snapshot spike deconvolution under a nontrivial PSF, with the radius expressed in terms of PSF spectral parameters, spike separation, amplitude dynamic range, and bandwidth. The paper also contributes a deterministic adversarial-noise analysis and a practical bandwidth-selection heuristic. The proof structure is coherent and detailed, with closed-form gradient/Hessian expressions and a reproducible code repository. However, the central theorem rests on conditioning lemmas imported from an unpublished preprint, and several stated hypotheses are weaker than those actually used in the proofs, so the correctness risk is concentrated in a few load-bearing points.

major comments (4)
  1. [Section 4, Lemma 9; Theorem 1] The proof of Lemma 9 asserts that U is full column rank using 'since G has at least 2K non-zero diagonal entries'. This condition on the PSF is not stated in Lemma 9 or in Theorem 1's hypotheses. If the PSF has spectral nulls on the sampled frequency grid, the bound on Sγ—and hence the curvature lower bound (48) and the basin radius (15)—is unsupported. Add the assumption to the theorems or modify the proof, and qualify the 'arbitrary PSF' claim in Sections 1.3 and 6.
  2. [Section 4, Lemma 7] The stated hypotheses of Lemma 7 require only BV and f²-integrability for Pg, but the proof bounds −Ĉ″₊(0)−Eg′ by applying Lemma 13 to Pg′ (see inequality (38)). This requires the same regularity for Pg′; since the conclusion already contains Eg′ and ρg′, the hypothesis is incomplete. Update Lemma 7 and the dependent results (Lemma 8 and Lemma 11) to state the needed regularity for Pg′ and Pg″.
  3. [Appendix A, Lemmas 10, 14, 15; Lemma 11] The key conditioning bounds are paraphrased from the unpublished preprint [15] and are not proved in this manuscript. Lemma 11's lower bound σmin(∇²ℓ) ≥ (1/3) Eg′ T rmin(X)², and therefore Theorem 1's basin radius (15) and the existence/uniqueness of the local minimizer, depend directly on these lemmas. Appendix A claims self-containedness, but the central proof is conditional on an unreviewed source. Include complete proofs of Lemmas 14 and 15 (or of Lemma 10) or clearly state that verification requires [15] and ensure it is publicly available by the time this paper appears.
  4. [Section 2.1, condition (16); Section 3] The residual-covariance condition (16) is a load-bearing hypothesis for Theorems 1, 2, and 6, but the numerical experiments in Section 3 do not verify whether this condition holds in the simulated parameter regimes. Since the experiments are presented as corroborating the theory, please report the relevant ratio (or an upper bound) for the settings in Figures 4, 6, 8, and 11, or state explicitly that the experiments only validate the qualitative behavior rather than the sufficient condition.
minor comments (5)
  1. [Section 1.4] The quantity rmax(X) is used in Theorems 1, 2, and 6 and in condition (16), but only rmin(X) is defined in the notation section. Define rmax(X) = max_j ||e_j^H X||_2 alongside rmin(X).
  2. [Section 2.2.3] The text 'Using ∥Aγ∥≲√TEg∥X∥ (cf. [15, Theorem 1])' is dimensionally/mathematically imprecise: Aγ is N×K, so its norm is bounded by √(TEg) times a separation-dependent factor, not by a quantity proportional to ∥X∥. The intended statement is likely ∥Y0∥ = ∥Aτ X∥ ≲ √(TEg)∥X∥. Please correct.
  3. [Section 2.2.3, Eq. (25)] The bound (25) is referred to as the 'gradient descent expression', but it is derived from Theorem 2, not from the gradient-descent convergence theorem (Theorem 6). Clarify the attribution to avoid confusion.
  4. [Section 2.2.2, Remark 2] The phrase 'XX^T in (23a) converges to the complex conjugate of the autocorrelation matrix RX = 1/L XX^H' is confusing because XX^T and XX^H differ, and R_X is defined with a different normalization. Please state the convergence in terms of (1/L)XX^T → I or similar, with the precise assumptions on X.
  5. [General editorial] Typos and minor issues: 'Frobenious' should be 'Frobenius' (Section 2.2.1); MSC classification '9408' should be '94A08'; the arXiv header dates conflict (v2, 24 Jul 2026 vs. July 27, 2026). Also, Figure 11(b) includes a '√CRB' curve whose definition for the L-vs-error experiment is not specified.

Circularity Check

0 steps flagged

No significant circularity; the basin-of-convexity proof is conditional on independent conditioning bounds, not on its own conclusion.

full rationale

The paper's central claim (Theorem 1) is a conditional statement: under a minimum-separation bound, the neighborhood radius (15), and the residual-covariance condition (16), the VarProSD objective is strongly convex with Lipschitz gradient and has a unique local minimizer. No displayed equation identifies the conclusion with an input by construction; the radius and residual condition are sufficient assumptions, not restatements of strong convexity or of the error bound. The proof of Lemma 11 (Appendix D) does import the dominant conditioning bounds from prior work: Lemma 14 and Lemma 15 are paraphrases of [15, Theorem 1], and Lemma 13 is a special case of [14, Theorem 3]. These are load-bearing, and [15] is a co-author's unpublished preprint, so the paper is not fully self-contained. However, the cited results are parameter-free theorems about singular values of weighted non-harmonic Fourier matrices and Beurling–Selberg majorants; their stated assumptions (bounded variation, integrability, minimum separation) do not include strong convexity of the VarProSD objective or any basin-size conclusion. Under the review rules, such parameter-free external results count as independent support and do not by themselves make the derivation circular. The paper also uses [15] in Lemma 9 via Lemma 10 for eigenvalue bounds on a Schur complement; this is again a prior conditioning result, not a restatement of the target. The bandwidth-selection criterion in Section 3.2 is an empirical heuristic: Bopt is chosen by minimizing ρ, and the numerical section compares it with the empirically optimal bandwidth. The paper explicitly disclaims providing a mathematically optimal B, so this is not a fitted input disguised as a prediction. One genuine internal gap exists but is not circular: Lemma 7's hypotheses only assume regularity of Pg, while its proof applies Lemma 13 to Pg' to control -C''_+(0)-Eg' (see the passage following Eq. (38)), requiring additional unstated BV/integrability assumptions on Pg'. This is a correctness/assumption issue, not a reduction of the conclusion to the input. Similarly, Lemma 9 requires G to have at least 2K nonzero diagonal entries, a hypothesis not listed in Theorem 1; again a proof gap, not circularity. No instance of self-definition, fitted-input-as-prediction, ansatz-smuggling, or renaming a known result was found. The negative result is therefore: no significant circularity, with score 0.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

No data-fitted free parameters appear in the theory; the absolute constants c1, c2, c5, c6 are existential and are not tuned to experiments. The central assumptions are spectral regularity of the PSF, the imported Beurling-Selberg approximation bounds, and the small-residual regime. No new physical or mathematical entities are postulated; the adversarial noise Z_adv is a constructed worst-case matrix, not an unobserved latent variable.

axioms (5)
  • domain assumption The truncated PSDs P_g, P_g′ and P_g′′ are in L1 and of bounded variation, with f ↦ 4π²f² P_g(f) also in L1.
    Invoked in Lemma 7 and Lemma 13 to apply the Beurling-Selberg majorant/minorant construction; for every PSF used in the experiments this holds, but it is a real restriction on the class of 'arbitrary' PSFs.
  • standard math The Beurling-Selberg extremal approximation lemma (Lemma 13) with the stated β-bandlimited majorant/minorant and residual bounds is valid.
    This is the core external Fourier-analysis result, cited to [14]; the paper states it as a theorem but does not prove it, and all later conditioning bounds inherit its exact constants and the even-symmetry properties.
  • domain assumption The generalized Vandermonde matrices Aγ, ΛAγ and U are full column rank, which requires G to have at least 2K nonzero diagonal entries.
    Used in Lemma 9 and throughout the Hessian analysis; if the PSF spectrum vanishes on too many sample frequencies, the projection and pseudoinverse identities break.
  • domain assumption The problem regime satisfies Δ > (2/3) ρ κ² and the residual-covariance condition (16) with a sufficiently small absolute constant.
    These are the explicit sufficient conditions in Theorem 1 for strong convexity; if the noise is large or the empirical covariance is far from the noiseless covariance, the basin need not exist.
  • domain assumption Noise is either i.i.d. Gaussian for Theorem 2 or bounded in spectral norm by (19) for Theorem 3.
    The Gaussian concentration argument and the adversarial Lipschitz argument depend on these two separate noise models, which cover very different perturbation regimes.

pith-pipeline@v1.3.0-alltime-deepseek · 41192 in / 14990 out tokens · 170910 ms · 2026-08-02T07:30:07.103573+00:00 · methodology

0 comments
read the original abstract

The problem of multi-snapshot spike deconvolution is studied, where the goal is to recover the locations of sparse impulses from their noisy convolution with a known point spread function (PSF) across multiple snapshots. A variable-projection formulation is adopted, in which the amplitudes are eliminated in closed form, thereby reducing the task to a nonconvex least-squares problem over the spike locations alone. This formulation is referred to as the variable-projection formulation of spike deconvolution (VarProSD). An explicit characterization of the basin of convexity of the VarProSD objective is provided in terms of key PSF properties, including its power spectral density and smoothness, revealing how sampling bandwidth and spike separation affect the local geometry. Within this basin, consistency of the estimator in the number of snapshots is established under stochastic noise, and a complementary, sharper error bound is derived under adversarial noise through the local Lipschitz property of the inverse map. Local convergence guarantees for gradient descent are further established when initialization is performed within the basin. A central role throughout the analysis is played by Beurling--Selberg extremal approximations, which enable sharp, PSF-agnostic bounds on the conditioning of the structured matrices arising in the optimization landscape. Numerical experiments are presented to corroborate the theoretical findings and demonstrate the effectiveness of modified ESPRIT initialization followed by gradient-based refinement.

Figures

Figures reproduced from arXiv: 2607.09593 by Kiryung Lee, Maxime Ferreira Da Costa, Meghna Kalra.

Figure 1
Figure 1. Figure 1: Estimation error vs. SNR under a Gaussian PSF for different initialization strategies ( [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Visualization of the region of convergence of the VarProSD objective [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: Performance of ESPRIT initialization methods and CRB across varying sampling bandwidth [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: PSF characteristic ρ plotted as a function of sampling bandwidth B = N/T and inverse pulse width 1/σ on a log-log scale. N is fixed; B is varied by adjusting T. The black curve indicates the optimal bandwidth Bopt that minimizes ρ for each 1/σ. time-domain shape and a PSD consisting of two Lorentzian-like lobes centered at ±5π ≈ ±15.7, which also decay sharply but away from the origin. These spectral chara… view at source ↗
Figure 6
Figure 6. Figure 6: Median log10(d2) error of gradient descent as a function of B and 1/σ, under oracle initialization (a) and modified ESPRIT initialization (b). The black curve indicates Bopt from ρ (as in [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Illustration of the Gaussian and Morlet PSFs in the time domain (left) and frequency domain [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Behavior of ρ and estimation error as functions of sampling bandwidth B for Gaussian and Morlet PSFs. (a) ρ vs. B, with vertical dashed lines indicating Bopt minimizing ρ for each PSF. (b) Median d2(τ , τb) error versus B for modified ESPRIT + GD, with vertical dashed lines indicating the empirically optimal B minimizing the error for each PSF. C+(f) ≥ Pg(f) for all f ∈ R. Hence, the right-hand side of (29… view at source ↗
Figure 9
Figure 9. Figure 9: Parameter ρ as a function of B and N for three Gaussian PSF widths σ. Colormaps show log10 ρ; the black curve indicates Bopt minimizing ρ for each N. (a) GD, σ = 0.01 (b) GD, σ = 0.05 (c) GD, σ = 0.10 [PITH_FULL_IMAGE:figures/full_fig_p020_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Performance of Mod-ESPRIT + GD versus B and N for three PSF widths. Colormaps show log10 of the median d2 error. The black curve is Bopt predicted from ρ, and the red dashed curve is the empirically optimal B that minimizes the median error. where Cb+(ξ) = R ∞ −∞ C+(f)e −i2πξf df, by the Poisson summation formula, we obtain X∞ l=−∞ C+  l T  e i2π(γj−γj ′ )l/T = T X∞ l=−∞ Cb+(lT − γj + γj ′ ). Plugging i… view at source ↗
Figure 11
Figure 11. Figure 11: Performance of various methods with varying (a) [PITH_FULL_IMAGE:figures/full_fig_p021_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Mean d2(τ , τb) error versus SNR under random Gaussian and adversarial noise for GD (left) and GN (right). bound |Cb′′ +(α)| in (34). For all α ∈ R, it comes that |Cb′′ +(α) − Pb′′ g (α)| = [PITH_FULL_IMAGE:figures/full_fig_p022_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

47 extracted references · 13 canonical work pages

  1. [1]

    Identification of parametric underspread linear systems and super-resolution radar

    Waheed U Bajwa, Kfir Gedalyahu, and Yonina C Eldar. “Identification of parametric underspread linear systems and super-resolution radar”. In:IEEE Transactions on Signal Processing59.6 (2011), pp. 2548–2561. doi:10.1109/TSP.2011.2112655

  2. [2]

    The stability of nonlinear least squares problems and the Cramér-Rao bound

    Samit Basu and Yoram Bresler. “The stability of nonlinear least squares problems and the Cramér-Rao bound”. In:IEEE Transactions on Signal Processing48.12 (2000), pp. 3426–3436.doi:10.1109/78.887032

  3. [3]

    Super-resolution of near-colliding point sources

    Dmitry Batenkov, Gil Goldman, and Yosef Yomdin. “Super-resolution of near-colliding point sources”. In: Information and Inference: A Journal of the IMA10.2 (2021), pp. 515–572.doi:10.1093/imaiai/iaaa005

  4. [4]

    Princeton, NJ: Princeton university press, 2009

    Dennis S Bernstein.Matrix mathematics: theory, facts, and formulas. Princeton, NJ: Princeton university press, 2009

  5. [5]

    Resolution of overlapping echoes of unknown shape

    Yoram Bresler and Alexander H Delaney. “Resolution of overlapping echoes of unknown shape”. In:ICASSP. 1989, pp. 2657–2660.doi:10.1109/ICASSP.1989.267014

  6. [6]

    Multi-way analysis in the food industry

    Rasmus Bro. “Multi-way analysis in the food industry”. In:Models, Algorithms, and Applications. Academish proefschrift. Dinamarca(1998)

  7. [7]

    Super-resolution from noisy data

    Emmanuel J Candès and Carlos Fernandez-Granda. “Super-resolution from noisy data”. In:J. Fourier Anal. Appl.19 (2013), pp. 1229–1254.doi:10.1007/s00041-013-9292-3

  8. [8]

    Towards a mathematical theory of super-resolution

    Emmanuel J Candès and Carlos Fernandez-Granda. “Towards a mathematical theory of super-resolution”. In:Commun. Pure. Appl. Math.67.6 (2014), pp. 906–956.doi:10.1002/cpa.21455

  9. [9]

    Harnessing sparsity over the continuum: Atomic norm minimization for superresolution

    Yuejie Chi and Maxime Ferreira Da Costa. “Harnessing sparsity over the continuum: Atomic norm minimization for superresolution”. In:IEEE Signal Processing Magazine37.2 (2020), pp. 39–57.doi: 10.1109/MSP.2019.2962209

  10. [10]

    Exact reconstruction using Beurling minimal extrapolation

    Yohann De Castro and Fabrice Gamboa. “Exact reconstruction using Beurling minimal extrapolation”. In: Journal of Mathematical Analysis and applications395.1 (2012), pp. 336–354

  11. [11]

    Inter-electrode spacing of surface EMG sensors: reduction of crosstalk contamination during voluntary contractions

    Carlo J De Luca et al. “Inter-electrode spacing of surface EMG sensors: reduction of crosstalk contamination during voluntary contractions”. In:Journal of biomechanics45.3 (2012), pp. 555–561.doi: 10.1016/j. jbiomech.2011.11.010

  12. [12]

    Philadelphia: SIAM, 1996

    John E Dennis Jr and Robert B Schnabel.Numerical methods for unconstrained optimization and nonlinear equations. Philadelphia: SIAM, 1996

  13. [13]

    Local Convergence Analysis of a Variable Projection Method for Regularized Separable Nonlinear Inverse Problems

    Malena I Español and Gabriela Jeronimo. “Local Convergence Analysis of a Variable Projection Method for Regularized Separable Nonlinear Inverse Problems”. In:SIAM Journal on Matrix Analysis and Applications 46.2 (2025), pp. 858–878.doi:10.1137/24M1639087

  14. [14]

    Second-order Beurling approximations and super-resolution from bandlimited functions

    Maxime Ferreira Da Costa. “Second-order Beurling approximations and super-resolution from bandlimited functions”. In:2023 International Conference on Sampling Theory and Applications (SampTA). IEEE. 2023, pp. 1–5.doi:10.1109/SampTA59647.2023.10301405

  15. [15]

    Maxime Ferreira Da Costa.The condition number of weighted non-harmonic Fourier matrices with applica- tions to super-resolution. Oct. 2025. hal:hal-04261330. preprint

  16. [16]

    Local geometry of nonconvex spike deconvolution from low- pass measurements

    Maxime Ferreira Da Costa and Yuejie Chi. “Local geometry of nonconvex spike deconvolution from low- pass measurements”. In:IEEE Journal on Selected Areas in Information Theory4 (2023), pp. 1–15.doi: 10.1109/JSAIT.2023.3262689

  17. [17]

    On the Stability of Super-Resolution and a Beurling–Selberg Type Extremal Problem

    Maxime Ferreira Da Costa and Urbashi Mitra. “On the Stability of Super-Resolution and a Beurling–Selberg Type Extremal Problem”. In:2022 IEEE International Symposium on Information Theory (ISIT). 2022, pp. 1737–1742.doi:10.1109/ISIT50566.2022.9834831

  18. [18]

    Preconditioned gradient descent for sketched mixture learning

    Joseph Gabet and Maxime Ferreira Da Costa. “Preconditioned gradient descent for sketched mixture learning”. In:2024 IEEE International Symposium on Information Theory (ISIT). Los Alamitos, CA: IEEE, 2024, pp. 3504–3509.doi:10.1109/ISIT57864.2024.10619105

  19. [19]

    The differentiation of pseudo-inverses and nonlinear least squares problems whose variables separate

    Gene H Golub and Victor Pereyra. “The differentiation of pseudo-inverses and nonlinear least squares problems whose variables separate”. In:SIAM Journal on numerical analysis10.2 (1973), pp. 413–432.doi: 10.1137/0710036

  20. [20]

    Stable estimation of pulses of unknown shape from multiple snapshots via ESPRIT

    Meghna Kalra and Kiryung Lee. “Stable estimation of pulses of unknown shape from multiple snapshots via ESPRIT”. In:IEEE Transactions on Signal Processing(2024).doi:10.1109/TSP.2024.3403494

  21. [21]

    On the condition number of Vandermonde matrices with pairs of nearly- colliding nodes

    Stefan Kunis and Dominik Nagel. “On the condition number of Vandermonde matrices with pairs of nearly- colliding nodes”. In:Numerical Algorithms87 (2021), pp. 473–496.doi:10.1007/s11075-020-00974-x

  22. [22]

    Channel State Information-Free Location-Privacy Enhancement: Fake Path Injection

    Jianxiu Li and Urbashi Mitra. “Channel State Information-Free Location-Privacy Enhancement: Fake Path Injection”. In:IEEE Transactions on Signal Processing72 (2024), pp. 3745–3760.doi:10.1109/TSP.2024. 3439315

  23. [23]

    Super-resolution limit of the ESPRIT algorithm

    Weilin Li, Wenjing Liao, and Albert Fannjiang. “Super-resolution limit of the ESPRIT algorithm”. In:IEEE Trans. Inf. Theory66.7 (2020), pp. 4593–4608.doi:10.1109/TIT.2020.2974174. 43

  24. [24]

    Stabilityandsuper-resolutionofMUSICandESPRITformulti-snapshotspectralestimation

    WeilinLietal.“Stabilityandsuper-resolutionofMUSICandESPRITformulti-snapshotspectralestimation”. In:IEEE Trans. Signal Process.70 (2022), pp. 4555–4570.doi:10.1109/TSP.2022.3204454

  25. [25]

    Yurii Nesterov.Introductory lectures on convex optimization: A basic course. Vol. 87. New York: Springer Science & Business Media, 2013

  26. [26]

    Performance analysis of the total least squares ESPRIT algorithm

    Bjorn Ottersten, Mats Viberg, and Thomas Kailath. “Performance analysis of the total least squares ESPRIT algorithm”. In:IEEE transactions on signal processing39.5 (2002), pp. 1122–1135.doi:10.1109/78.80967

  27. [27]

    Essai experimental et analytique sur les lois de la dilatabilite de fluides elastiques et sur celles da la force expansion de la vapeur de l’alcool, a differentes temperatures

    GRB Prony. “Essai experimental et analytique sur les lois de la dilatabilite de fluides elastiques et sur celles da la force expansion de la vapeur de l’alcool, a differentes temperatures”. In:J. Ec. Polytech.1.2 (1795)

  28. [28]

    Singapore: World Scientific, 1998

    Calyampudi Radhakrishna Rao and M Bhaskara Rao.Matrix algebra and its applications to statistics and econometrics. Singapore: World Scientific, 1998

  29. [29]

    ESPRIT-estimation of signal parameters via rotational invariance techniques

    Richard Roy and Thomas Kailath. “ESPRIT-estimation of signal parameters via rotational invariance techniques”. In:IEEE Trans. Acoust., Speech, Signal Process.37.7 (1989), pp. 984–995.doi:10.1109/29. 32276

  30. [30]

    New York: McGraw-Hill, 1976

    Walter Rudin.Principles of Mathematical Analysis. New York: McGraw-Hill, 1976

  31. [31]

    Algorithms for separable nonlinear least squares problems

    Axel Ruhe and Per Åke Wedin. “Algorithms for separable nonlinear least squares problems”. In:SIAM review22.3 (1980), pp. 318–337.doi:10.1137/1022057

  32. [32]

    Geometry of the Cramer-Rao bound

    Louis L Scharf and L Todd McWhorter. “Geometry of the Cramer-Rao bound”. In:Signal Processing31.3 (1993), pp. 301–311.doi:10.1016/0165-1684(93)90088-R

  33. [33]

    Superresolution without separation

    Geoffrey Schiebinger, Elina Robeva, and Benjamin Recht. “Superresolution without separation”. In:Infor- mation and Inference: A Journal of the IMA7.1 (2018), pp. 1–30.doi:10.1093/imaiai/iax006

  34. [34]

    Multiple emitter location and signal parameter estimation

    Ralph Schmidt. “Multiple emitter location and signal parameter estimation”. In:IEEE Trans. Antennas Propag.34.3 (1986), pp. 276–280.doi:10.1109/TAP.1986.1143830

  35. [35]

    Methods for blind equalization and resolution of overlapping echoes of unknown shape

    A Lee Swindlehurst and Jacob H Gunther. “Methods for blind equalization and resolution of overlapping echoes of unknown shape”. In:IEEE Trans. Signal Process.47.5 (1999), pp. 1245–1254.doi:10.1109/78. 757212

  36. [36]

    Compressed sensing off the grid

    Gongguo Tang et al. “Compressed sensing off the grid”. In:IEEE Trans. Inf. Theory59.11 (2013), pp. 7465– 7490.doi:10.1109/TIT.2013.2277451

  37. [37]

    The basins of attraction of the global minimizers of the non- convex sparse spike estimation problem

    Yann Traonmilin and Jean-François Aujol. “The basins of attraction of the global minimizers of the non- convex sparse spike estimation problem”. In:Inverse Problems36.4 (2020), p. 045003.doi:10.1088/1361- 6420/ab5aa3

  38. [38]

    On strong basins of attractions for non-convex sparse spike estimation: upper and lower bounds

    Yann Traonmilin et al. “On strong basins of attractions for non-convex sparse spike estimation: upper and lower bounds”. In:Journal of Mathematical Imaging and Vision66.1 (2024), pp. 57–74.doi:10.1007/s10851- 023-01163-w

  39. [39]

    On the Number of Lattice Points in a Ball

    Jeffrey D. Vaaler. “On the Number of Lattice Points in a Ball”. In:Communications in Mathematics31.2 (2023).doi:10.46298/cm.11119

  40. [40]

    Some Extremal Functions in Fourier Analysis

    Jeffrey D. Vaaler. “Some Extremal Functions in Fourier Analysis”. In:Bulletin of the American Mathematical Society12.2 (2 1985), pp. 183–216.issn: 0273-0979, 1088-9485.doi:10.1090/S0273-0979-1985-15349-2

  41. [41]

    Variable projection for nonsmooth problems

    Tristan Van Leeuwen and Aleksandr Y Aravkin. “Variable projection for nonsmooth problems”. In:SIAM journal on scientific computing43.5 (2021), S249–S268.doi:10.1137/20M1348650

  42. [42]

    Stochastic sparse-spike deconvolution

    Danilo R Velis. “Stochastic sparse-spike deconvolution”. In:Geophysics73.1 (2008), R1–R9.doi:10.1190/1. 2790584

  43. [43]

    Roman Vershynin.High-dimensional probability: An introduction with applications in data science. Vol. 47. Cambridge: Cambridge university press, 2018

  44. [44]

    Martin J Wainwright.High-dimensional statistics: A non-asymptotic viewpoint. Vol. 48. Cambridge: Cam- bridge university press, 2019

  45. [45]

    Channel estimation for OFDM transmission in multipath fading channels based on parametric channel modeling

    Baoguo Yang et al. “Channel estimation for OFDM transmission in multipath fading channels based on parametric channel modeling”. In:IEEE transactions on communications49.3 (2001), pp. 467–479.doi: 10.1109/26.911454

  46. [46]

    Nonasymptotic performance analysis of ESPRIT and spatial-smoothing ESPRIT

    Zai Yang. “Nonasymptotic performance analysis of ESPRIT and spatial-smoothing ESPRIT”. In:IEEE Trans. Inf. Theory69.1 (2022), pp. 666–681.doi:10.1109/TIT.2022.3199405

  47. [47]

    Inequalities for the singular values of Hadamard products

    Xingzhi Zhan. “Inequalities for the singular values of Hadamard products”. In:SIAM Journal on Matrix Analysis and Applications18.4 (1997), pp. 1093–1095.doi:10.1137/S0895479896309645. 44