{"id":"ad7c546d-315d-43af-984d-18db93d64503","arxiv_id":"2608.13201","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A single spectral sandwich on the Sinkhorn linearization yields identifiability, sparsistency, well-posedness, and convergence bounds for feature-parameterized inverse optimal transport.","lead":"This paper derives a single spectral bound for inverse optimal transport with feature-parameterized costs, and shows that identifiability, sparse recovery, well-posedness, and convergence all follow from it. It also introduces a cheap \"spectral proxy\" formula that shares the bound and is faster to evaluate.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"T2's irrepresentability bridge from the feature Gram Σ to the OT Hessian H is unproven and empirically unsatisfied (0/40 in E8b), so the sparsistency theorem lacks a demonstrated non-vacuous instantiation; the unified theory remains conditional.","rationale":"The reader's weakest-assumption identification is correct: the unproven bridge from irrepresentability of the feature Gram Σ to irrepresentability of the OT information matrix H is the most load-bearing unresolved point. I checked the central spectral argument independently and found Proposition 3.3, Lemma 3.1, and the derived T1/T3(a)/T4 bounds to be internally consistent: the exact linearization δx = −B H_T⁻¹ Bᵀ δc, the spectral sandwich, the Jacobian lower bound via σ_min(R_a) = 1/a_max, and the zero second-derivative residual in the Hessian of the cross-entropy all check out. The difficulty is specifically that T2's Assumption A2.2 is stated on H, while the paper's own Lemma 5.2 and the strength-of-claims section explicitly say the transfer from Σ to H is not assumed; E8b then finds the H condition violated in all 40 random-feature trials. This does not make T2 false—it is a valid conditional statement—but it means the paper has not established a feature-level condition under which sparsistency is guaranteed, so the advertised unification is incomplete. The experimental finding that recovery succeeds despite IR violation suggests IR may not be necessary, but that only widens the gap between the theorem's sufficient conditions and the actual IOT setting. No further concern warrants moving to REJECT: the paper is transparent about the missing bridge, the core spectral framework is sound, and the other theorems are conditional on clearly stated assumptions. The appropriate verdict remains CONDITIONAL, which is unchanged from the reader's assessment.","tokens_in":25094,"tokens_out":15375,"duration_ms":138082,"concrete_test":"Reproduce E8b with features orthogonalized in the P_T inner product so that Σ = I and the Lemma 5.2 mutual-incoherence condition holds trivially; for the same marginals and a sparse true θ*, compute H = Σ_l J_l(θ*)ᵀ diag(ρ_i/Q*(i,j)) J_l(θ*) and check whether ∥H_{SᶜS} H_{SS}⁻¹∥∞ < 1. If it fails, the Σ-to-H bridge is false rather than merely unproven; if it holds, that construction provides a concrete, non-vacuous regime where T2 applies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The core spectral bound (Proposition 3.3) and the derived T1, T3(a), and T4 arguments appear internally sound: the spectral sandwich (10) is correctly propagated through the exact Sinkhorn linearization, and the Hessian computation in §7.2 correctly uses the vanishing second-derivative residual. The load-bearing weakness is in T2. Assumption A2.2 is an irrepresentability condition on H = ∇²ℓ(θ*), the OT information matrix, but the only a priori sufficient condition supplied, Lemma 5.2, bounds mutual incoherence of the feature Gram Σ. Lemma 5.2's own note states that transferring Σ-irrepresentability to H 'requires an additional bridge' that is 'not separately assumed in this paper.' The experiments (E8b, 40 seeds) report the H condition satisfied 0/40 times, with the quantity in [1.36, 7.43]. Consequently, for generic feature parameterizations there is no known feature-level condition under which A2.2 holds, so Theorem 5.1 is a conditional generic-Lasso statement rather than an established consequence of the spectral framework. The paper openly documents this limitation, but the abstract's claim of four established theorems leans on T2 as part of the unified package; without the bridge, the package is incomplete rather than false. The failure does not undermine the core spectral bound or T1/T3/T4, but it keeps the overall contribution conditional and scope-limited.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a frequentist statistical theory for inverse entropic optimal transport (IOT) under feature-parameterized costs C_theta(i,j)=-theta^T phi(i,j), with observations given by conditional transition operators. The main technical object is the Sinkhorn linearization, the exact sensitivity of the entropic OT plan to the cost, whose tangent-space restricted Hessian satisfies the spectral sandwich (10). From this sandwich the authors derive a core singular-value lower bound (Proposition 3.3) and use it to drive four theorems and one observation: T1 (global identifiability on a gauge quotient with dimension bound F <= (K-1)^2), T2 (l1-support recovery under a Hessian irrepresentability condition), T3 (strong monotonicity and Lipschitz stability of the feature-moment map), T4 (local strong convexity and monotone gradient-descent convergence), and O5 (convergence to the pseudo-true projection under misspecification). The numerical section reports experiments E1-E10, including support-recovery phase transitions, irrepresentability diagnostics, perturbation transfer, initialization comparisons, and empirical scaling of pi_min with epsilon.","tokens_in":25395,"tokens_out":12123,"duration_ms":109144,"significance":"If the results are taken as stated, the paper makes a useful contribution: it reduces identifiability, well-posedness, and convergence constants to a single spectral bound sigma_min(J_theta) >= (pi_min/(a_max epsilon)) sqrt(lambda_min(Sigma)), with the Sinkhorn linearization as the underlying calculus. The derivation of Proposition 3.3 and the arguments for T1, T3(a), and T4 appear internally sound, and the paper is commendably explicit about which claims are conditional, conjectural, or empirical (e.g., the Holder exponent and the pi_min scaling law). The main caveat is that T2's irrepresentability hypothesis is stated for the OT information matrix H = grad^2 ell(theta*), whereas the only a priori sufficient condition supplied (Lemma 5.2) concerns the feature Gram matrix Sigma, and the transfer bridge is explicitly not assumed. The experiments find the H-irrepresentability condition satisfied 0/40 times, so the sparse-recovery theorem currently has no demonstrated non-vacuous instantiation in the feature-parameterized setting. This limits, but does not destroy, the claimed unification.","major_comments":[{"comment":"The central issue is the gap between Assumption A2.2 and its only a priori sufficient condition. A2.2 requires irrepresentability of H = grad^2 ell(theta*), while Lemma 5.2 verifies irrepresentability of the feature Gram matrix Sigma; the lemma's own note states that transferring Sigma-irrepresentability to H 'requires an additional bridge' that is 'not separately assumed in this paper.' Since no such bridge is proved and the numerical check E8b reports the H-condition satisfied 0/40 times (with values in [1.36, 7.43]), Theorem 5.1 is currently a conditional Lasso-type statement rather than an established consequence of the spectral framework. I recommend either proving a transfer result (e.g., a quantitative comparison between H and a constant multiple of Sigma under explicit conditions on pi_min, epsilon, and feature normalization) or revising the abstract, introduction, and conclusion so that T2 is presented as conditional on an unverified structural condition instead of as one of the theorems established on the core bound.","section":"Section 5.1, Lemma 5.2 and Assumption A2.2; Section 9.2, E8b"},{"comment":"The global selection condition is not formalized as a precise hypothesis. The proof's primal-dual witness construction establishes that the constructed restricted candidate satisfies the KKT conditions; equality with the global l1-penalized estimator requires the additional condition, stated inside A2.1, that the neighborhood U contains both the PDW candidate and the global solution and that no lower minimum exists outside U. This condition is only illustrated by examples ('for example, the penalized objective is convex...'), and no sufficient condition is established in the IOT setting. As written, Theorem 5.1 therefore proves support recovery for the PDW local solution, not necessarily for the global estimator btheta_lasso_n. The hypothesis must be stated formally if the theorem is to support the claimed conclusion about the global l1 estimator.","section":"Section 5.1, Theorem 5.1 and Assumption A2.1"}],"minor_comments":[{"comment":"The sentence '(PAP)|_T^{-1} != P A^{-1}P' uses the unqualified inverse of a singular operator; please write the inverse on the range explicitly to avoid ambiguity.","section":"Section 3.2"},{"comment":"The 'core spectral bound' is referred to as (3.3) in many places, but no displayed equation with that number appears; renumber the displays or fix the cross-references.","section":"Section 3.5 and throughout"},{"comment":"The abstract says the spectral proxy is 'spectrally exact (it preserves all singular-value bounds)', but Definition 3.2 only establishes equality of the two-sided operator-norm envelope; suggest 'preserves the spectral bounds' for accuracy.","section":"Abstract and Definition 3.2"},{"comment":"The subsection title 'The epsilon-scaling law of pi_min' overstates the status of a finite-window log-log fit; the surrounding text correctly labels alpha_eff as an empirical exponent, so consider retitling the subsection.","section":"Section 9.6"},{"comment":"The text in Section 5.1 mentions a (6x-156x) comparison between the range-based Hoeffding constant and the Bernstein constant, while E9 reports ratios in [13.7, 33.9]; reconcile the two ranges or clarify that they refer to different thresholds and settings.","section":"Section 5.1 and Section 9.2, E9"},{"comment":"The implementation-specific path '09 iot theory/output/experiments e1 e10.json' should be replaced by a stable artifact reference or removed from the final manuscript.","section":"Section 9, Experimental protocol"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest about its limitations, and the core spectral-sandwich results appear sound. The main risk is that the abstract and title promise more than T2 delivers. If the authors can either supply the Sigma-to-H irrepresentability bridge or reframe the claims so that T2 is explicitly conditional, I would be comfortable with publication after a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First the punchline: the paper's core spectral bound is sound and the identifiability/well-posedness/convergence theorems that follow from it are derived cleanly; the sparsistency theorem is the weak link, and the authors know it. The spectral sandwich on the restricted Hessian—pi_min/epsilon I ≤ H_T^{-1} ≤ pi_max/epsilon I—is a direct consequence of the Hessian being epsilon times diag(1/pi), and it propagates correctly through the Sinkhorn linearization to the core bound sigma_min ≥ (pi_min/(a_max epsilon)) sqrt(lambda_min(Sigma)). I checked the steps in Proposition 3.3 and the proof of T1; they're correct. The dimension bound F ≤ (K-1)^2 and the quotient-space injectivity statement are a real contribution, and the Lipschitz/strong-convexity constants are explicit. The Hessian computation in §7.2 has a nice vanishing-residual argument. The paper is also admirably open about its own status: the 'Strength of the claims' paragraph and the open problems list say which parts are complete and which are not.\n\nThe soft spot is the irrepresentability bridge, exactly as your note says. T2 requires irrepresentability of H = grad^2 ell(theta*), but Lemma 5.2 only supplies a sufficient condition on the feature Gram Sigma, and the paper states that transferring it to H 'requires an additional bridge' that is 'not separately assumed.' That bridge never appears. The experiments confirm the worry: H-IR is satisfied 0/40 times. Recovery still works empirically, so the condition is sufficient but not necessary, but the theorem has no demonstrated non-vacuous instantiation. O5's Hölder continuity is honestly labeled a conjecture, and the one favorable numeric setting (epsilon=0.1) collapses once the Adam residual is accounted for. These are disclosed limitations, not hidden ones.\n\nThe citation pattern looks fair and relevant. The absence of code/data is a moderate issue but not decisive.\n\nOverall the central framework survives. I would send this to peer review, with a referee instructed to push for a proof of the Sigma-to-H transfer or a downgrade of T2's status in the abstract. If I were working in IOT I'd cite the spectral sandwich and the T1/T3 results. It belongs in a reading group precisely because the gap is instructive.\n\nRecommendation: engage; serious referee.","headline":"A mostly sound spectral framework for inverse OT, with a clean core bound driving T1/T3/T4, but the sparsistency theorem rests on an openly admitted and empirically unsatisfied irrepresentability bridge.","tokens_in":25942,"tokens_out":3975,"would_cite":true,"duration_ms":34008,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49Q22","62F12","62J07","90C25"],"pacs":[],"model":"deepseek-v4-flash","headline":"One spectral bound on the Sinkhorn linearization drives identifiability, stability, convergence, and sparse recovery in inverse optimal transport.","keywords":["inverse optimal transport","Sinkhorn linearization","spectral proxy","entropic regularization","identifiability","sparsistency","strong monotonicity","misspecified estimation"],"falsifier":"At a dense grid of parameter values with fixed positive marginals and $\\lambda_{\\min}(\\Sigma)>0$, compute the exact Jacobian $J_\\theta$ by automatic differentiation and test whether $\\sigma_{\\min}(J_\\theta) \\ge (\\pi_{\\min}(\\theta)/(a_{\\max}\\varepsilon))\\sqrt{\\lambda_{\\min}(\\Sigma)}$; any single violation would falsify Proposition 3.3 and, with it, the common core of T1, T3, and T4.","tokens_in":24890,"feed_emoji":"🧮","tokens_out":9339,"duration_ms":75884,"temperature":0.7,"pith_summary":"The paper sets out to show that feature-parameterized inverse optimal transport—recovering the cost parameters of an entropic transport problem from observed transition probabilities—is a well-posed statistical problem rather than a heuristic inversion. Its key move is to differentiate the Sinkhorn map with respect to the cost and prove a two-sided spectral bound on the restricted Hessian: the inverse of the Hessian is sandwiched between $(\\pi_{\\min}/\\varepsilon)I$ and $(\\pi_{\\max}/\\varepsilon)I$. From this one sandwich the paper derives four theorems and an observation: global identifiability up to a gauge kernel, support recovery for $\\ell_1$-penalized estimation, strong monotonicity of the feature-moment map with a Lipschitz inverse, monotone gradient-descent convergence, and convergence of misspecified estimators to the projection of the truth onto the OT model set. If the claims hold, practitioners get explicit dimension limits, Lipschitz constants, strong-convexity moduli, and exponential rate constants for a class of problems that previously lacked them.","feed_headline":"One spectral inequality governs inverse optimal transport","feed_subtitle":"From one Hessian sandwich come identifiability, Lipschitz stability, monotone convergence, and support recovery.","key_machinery":"The Sinkhorn linearization is the implicit-function derivative of the entropic plan with respect to the cost, $\\delta x = -B H_T^{-1} B^\\top \\delta c$, obtained by differentiating the KKT conditions of the entropy-regularized problem. Its spectral proxy, $\\delta x_{\\mathrm{SSP}} = -(1/\\varepsilon) P_T D_\\pi P_T \\delta c$, replaces the dense inverse by two tangent-space projections surrounding elementwise multiplication by the plan entries, preserving all spectral bounds. The load-bearing identity is the spectral sandwich $(\\pi_{\\min}/\\varepsilon)I \\preceq H_T^{-1} \\preceq (\\pi_{\\max}/\\varepsilon)I$: because every plan entry lies between $\\pi_{\\min}$ and $\\pi_{\\max}$, the inverse restricted Hessian inherits uniform spectral control, and that control propagates directly to the singular-value lower bound that all downstream theorems use.","core_discovery":"The central discovery is that the statistical and algorithmic theory of feature-parameterized inverse optimal transport radiates from a single spectral bound. The paper proves that the restricted Hessian of the entropic OT objective, pulled back to the tangent space of feasible plans, has inverse sandwiched between $(\\pi_{\\min}/\\varepsilon)I$ and $(\\pi_{\\max}/\\varepsilon)I$; consequently the Jacobian of the conditional transition operator with respect to the cost parameters obeys $\\sigma_{\\min}(J_\\theta) \\ge (\\pi_{\\min}(\\theta)/(a_{\\max}\\varepsilon))\\sqrt{\\lambda_{\\min}(\\Sigma)}$ at every $\\theta$ with positive plan entries and positive $\\lambda_{\\min}(\\Sigma)$. This bound yields global injectivity on the quotient by the gauge kernel (T1), strong monotonicity and a Lipschitz inverse for the feature-moment map (T3), local strong convexity of the cross-entropy objective with monotone gradient descent (T4), and support recovery for the $\\ell_1$-penalized estimator under an additional irrepresentability condition (T2); the misspecification analysis (O5) identifies the estimator's limit as the projection of the truth onto the OT model set.","pith_inferences":["Because the core bound degrades as $\\pi_{\\min}\\to 0$, a practical design rule the paper leaves implicit is that $\\varepsilon$ should be chosen to keep $\\pi_{\\min}(\\theta,\\varepsilon)\\sqrt{\\lambda_{\\min}(\\Sigma)}/\\varepsilon$ bounded away from zero, rather than merely set small.","The paper's own 40-seed check finds the irrepresentability condition violated in every trial, which suggests that plain $\\ell_1$ penalization in raw feature coordinates may not be the right sparse-recovery device for generic feature parameterizations; a preconditioned or debiased estimator whose effective information matrix inherits Gram-matrix irrepresentability is a natural testable alternative.","The spectral proxy's geometric form—project, weight by plan entries, project back—predicts that inverse-problem conditioning is controlled by where the plan mass sits; one could test this by placing features on low-mass regions of the plan and measuring the condition number of the sensitivity Gram matrix."],"forward_implications":["On the quotient space $\\mathbb{R}^F/N_\\Phi$, the map from cost parameters to conditional transition operators is globally injective; on the original space it is injective exactly when $\\mathrm{rank}(\\Sigma)=F$, which forces the feature-dimension cap $F\\le(K-1)^2$.","On any compact convex parameter domain, the inverse map from observed transition operators to cost parameters is Lipschitz with constant at most $\\varepsilon\\|\\Phi^\\top S_a\\|_{\\mathrm{op}}/(\\pi_{\\min}\\lambda_{\\min}(\\Sigma))$, so small observation noise cannot produce arbitrarily large parameter error.","Near the true parameter, the population cross-entropy is strongly convex with modulus at least $\\pi_{\\min}^2\\lambda_{\\min}(\\Sigma)/\\varepsilon^2$, so fixed-step-size gradient descent converges monotonically from any initialization inside its basin.","Under irrepresentability of the information matrix and all-coordinate score concentration, the $\\ell_1$-penalized estimator recovers the true support with failure probability at most $C_1\\exp(-2n t_n^2/\\Delta_{\\max}^2)$, and the debiased refit on the recovered support is asymptotically normal.","When the data are not generated by the OT model, the estimator converges to the pre-image (under the T3 inverse map) of the cross-entropy projection of the truth onto the OT model set, rather than to any true parameter."],"supporting_citations":[{"why":"Establishes the multiplicative form of the entropic plan, the object whose cost-sensitivity the paper differentiates.","marker":"[23]"},{"why":"Defines the entropic OT forward map and its computational Sinkhorn scheme, the data-generating process for the observed transition operators.","marker":"[7]"},{"why":"Introduces the inverse OT problem as a recovery-from-transition-operator setting that this paper recasts in frequentist terms.","marker":"[24]"},{"why":"Supplies the sparsistency and irrepresentability framework for inverse OT that T2 extends with an explicit exponential rate constant.","marker":"[1]"},{"why":"Provides the primal-dual witness support-recovery conditions used in the proof of Theorem 5.1.","marker":"[30]"},{"why":"Provides the Lasso sharp-threshold and irrepresentability theory that underlies the $\\ell_1$-penalized estimator analysis.","marker":"[26]"},{"why":"Gives the misspecified M-estimation consistency theory on which Observation O5 relies.","marker":"[28]"},{"why":"Underpins the Lipschitz-inverse conclusion for strongly monotone operators used in Theorem T3.","marker":"[17]"}],"fun_headline_variants":["One spectral sandwich ties together inverse OT theory","Spectral sandwich: the one inequality behind inverse OT","From one Hessian bound: all of inverse OT's key results","Inverse OT's master inequality: the spectral sandwich","A spectral sandwich unifies inverse OT stats and algorithms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the information matrix at the true parameter satisfies an irrepresentability condition, a requirement the paper cannot check a priori and that its own 40-seed experiment finds violated in every trial.","fun_headline_variants_meta":{"raw":{"variants":["One spectral sandwich ties together inverse OT theory","Spectral sandwich: the one inequality behind inverse OT","From one Hessian bound: all of inverse OT's key results","Inverse OT's master inequality: the spectral sandwich","A spectral sandwich unifies inverse OT stats and algorithms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000939,"raw_usage":{"total_tokens":4108,"prompt_tokens":1133,"completion_tokens":2975,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":749,"completion_tokens_details":{"reasoning_tokens":2898}},"tokens_in":749,"tokens_out":2975,"duration_ms":20623,"temperature":1.0,"reasoning_tokens":2898,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:32:21.714984+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"At a dense grid of parameter values with fixed positive marginals and $\\lambda_{\\min}(\\Sigma)>0$, compute the exact Jacobian $J_\\theta$ by automatic differentiation and test whether $\\sigma_{\\min}(J_\\theta) \\ge (\\pi_{\\min}(\\theta)/(a_{\\max}\\varepsilon))\\sqrt{\\lambda_{\\min}(\\Sigma)}$; any single violation would falsify Proposition 3.3 and, with it, the common core of T1, T3, and T4.","supporting_citations":[{"cited_title":"Zhao and B","cited_arxiv_id":null,"evidence_quote":"Provides the primal-dual witness support-recovery conditions used in the proof of Theorem 5.1."},{"cited_title":"Sinkhorn , A relationship between arbitrary positive matrices and doubly stochastic matrices , The Annals of Mathematical Statistics, 35 (1964), pp","cited_arxiv_id":null,"evidence_quote":"Establishes the multiplicative form of the entropic plan, the object whose cost-sensitivity the paper differentiates."},{"cited_title":"Cuturi , Sinkhorn distances: Lightspeed computation of optimal transport , Advances in Neural Information Processing Systems, 26 (2013)","cited_arxiv_id":null,"evidence_quote":"Defines the entropic OT forward map and its computational Sinkhorn scheme, the data-generating process for the observed transition operators."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Lasso sharp-threshold and irrepresentability theory that underlies the $\\ell_1$-penalized estimator analysis."},{"cited_title":"White , Maximum likelihood estimation of misspecified models , Econometrica, 50 (1982), pp","cited_arxiv_id":null,"evidence_quote":"Gives the misspecified M-estimation consistency theory on which Observation O5 relies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Underpins the Lipschitz-inverse conclusion for strongly monotone operators used in Theorem T3."}],"review_version":1}