{"id":"c469ba37-695e-4f77-9963-17b6cec62499","arxiv_id":"2502.01164","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A penalized optimal transport relaxation interpolates between covariate-free and covariate-conditioned causal bounds, converging to the sharp conditional bound as the penalty grows.","lead":"This paper shows how to compute tighter bounds on cause-and-effect questions by matching treated and control groups on background variables, using ordinary optimal transport software. The method interpolates between no covariate adjustment and full covariate conditioning, giving a practical tool for partial identification in randomized experiments.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Finite-sample rate is not uniform in the penalty: Theorem 4.3's C λ η γ_{N,d} requires controlling the Brenier curvature λ(η), which the paper concedes it cannot bound; if λ grows superlinearly in η the advertised convergence guarantee degrades exactly in the regime that approaches the sharp COT…","rationale":"The reader's weakest assumption is the same as mine. The population-level interpolation (Proposition 3.6) is proved carefully: monotonicity, continuity, and the η→∞ limit all check out, and the Gaussian example provides a closed-form consistency check. The finite-sample rate is the main claimed contribution beyond prior OT relaxations, and its advertised form Cληγ_{N,d} is only meaningful if the curvature λ can be controlled as a function of η. The paper openly disclaims such a bound in Section 4.2, leaving Lemma 4.6 for Gaussians and a heuristic in Appendix C.7. Without a general bound, the theorem is conditional on an unverified hypothesis in exactly the regime where the method is supposed to converge to the sharp COT bound. This does not invalidate the method's population validity or its numerical usefulness, but it makes the statistical guarantee incomplete, supporting the CONDITIONAL verdict rather than ACCEPT. I would not move the verdict; the paper already flags the limitation, and my proposed 1D check would settle whether the concern actually lands.","tokens_in":34974,"tokens_out":12262,"duration_ms":130229,"concrete_test":"Compute the 1D Brenier map T_η between P_{Y(1),Z} and P_{Y(0),√η Z} for a non-Gaussian location model, e.g., Y(0)=0.5Z+ε0, Y(1)=1.5Z+ε1 with ε's drawn from a centered t-distribution or a two-component Gaussian mixture. In one dimension T_η = F_target^{-1}∘F_source is available by quantile functions, so evaluate ∇T_η (the Hessian of the Brenier potential) on a fine grid and record λ(η)=max(max_x ∇T_η(x), max_x 1/∇T_η(x)) for η∈{1,10,100,1000}. If λ(η) grows faster than linearly in η, Theorem 4.3's rate Cληγ_{N,d} is not usable in the sharp-bound limit, and the paper would need a dimension-independent curvature bound or a revised statistical guarantee.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing concern is the unquantified dependence of the Brenier curvature λ on the penalty η in Theorem 4.3. The theorem states E|Vip,n,m(η)-Vip(η)| ≤ C λ η γ_{N,d}, where λ enters through the assumption 1/λ I ⪯ ∇² φ̃ ⪯ λ I on the potential between P_{A12^T Y(0), η Z} and P_{Y(1),Z}. Section 4.2 explicitly admits 'we currently lack a general method to bound the curvature λ with respect to η', and Appendix C.7 offers only an informal scaling argument. This matters because the population-level tightening Vip(η)↑Vc is only realized as η→∞; to exploit it one must send η with N, but the error bound contains η and any η-dependent λ. If λ grows faster than O(η), the bound can diverge as η grows, so the finite-sample guarantee does not cover the asymptotically exact regime. Even if λ is only linear in η, the rate becomes η²γ_{N,d}, which after balancing a 1/η bias gap yields a slow rate for d≥5. The paper does not provide a bias-variance tradeoff or data-driven η selection, so the central claim of 'narrower intervals, asymptotically exact' is supported at the population level but not statistically at any stated rate. The concern is about incompleteness of the guarantee, not an inconsistency in Proposition 3.6, which appears correct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a covariate-aware relaxation of conditional optimal transport (COT) for partial identification of causal estimands. For a penalty parameter η, the method introduces mirror covariates Z(0), Z(1) and solves a standard (unconditional) optimal transport problem with augmented cost h(Y(0),Y(1)) + η∥Z(0)-Z(1)∥². The population value Vip(η) is shown to be a valid lower bound that is no smaller than the covariate-free OT bound Vu and no larger than the sharp COT bound Vc, and to interpolate between Vu and Vc as η ranges from 0 to infinity. A plug-in estimator based on empirical marginals is proposed, implemented through a standard LP/Sinkhorn solver, and analyzed for consistency and convergence rate. Synthetic and real-data experiments (including the STAR dataset) illustrate that the method can substantially tighten the partial identification interval relative to covariate-free OT and to a competing COT-based method.","tokens_in":35340,"tokens_out":7123,"duration_ms":84289,"significance":"If fully established, the population-level interpolation result would be a clean and practically useful bridge: it turns the harder COT problem into a sequence of standard OT problems while preserving covariate adjustment, and it gives a valid, monotone family of lower bounds. The proof of Proposition 3.6 uses compactness and uniqueness arguments that appear correct, and the estimator is simple to implement with off-the-shelf OT packages. The paper also provides reproducible code and a useful set of quadratic-function examples (Neymanian variance, correlation of potential outcomes, null-effect testing). The main weakness is that the finite-sample rate in Theorem 4.3 is not uniform in the penalty η, because the curvature constant λ is allowed to depend on η without a general bound; this blocks the advertised 'asymptotically exact as η→∞' statistical claim. This is a completeness issue in the finite-sample theory rather than an error in the population-level interpolation.","major_comments":[{"comment":"The finite-sample guarantee is not uniform in the penalty η, and this is load-bearing for the paper's central statistical claim. The bound is E[|Vip,n,m(η)-Vip(η)|] ≤ C λ η γ_{N,d}, where λ appears in the curvature assumption 1/λ I ⪯ ∇²φ̃ ⪯ λ I on the Brenier potential between P_{A12^T Y(0), ηZ} and P_{Y(1),Z}. Since the source measure depends on η, λ is generally a function of η, but the paper states in Section 4.2 that it 'currently lack[s] a general method to bound the curvature λ with respect to η'; Appendix C.7 offers only an informal scaling heuristic and Lemma 4.6 covers only jointly Gaussian marginals. The population gain Vip(η)↑Vc is realized only as η→∞, yet the error bound contains η and any η-dependent λ; if λ grows faster than O(η), the bound diverges exactly in the regime that approaches the sharp COT bound. Moreover, no bias-variance tradeoff or data-driven rule for selecting η is provided. The theorem as stated is a fixed-η statement, but the abstract and Remark 3.7 use the limit η→∞ as a central motivation, so the finite-sample theory is incomplete for that regime.","section":"Section 4.2, Theorem 4.3"},{"comment":"The appeal to Caffarelli regularity does not close the gap left by Theorem 4.3. Lemma 4.5 asserts the existence of some λ > 0 under C² uniformly convex supports and Hölder densities, but it gives no quantitative control of λ in terms of η, the covariance structure, or the distributions. Consequently, combining Lemma 4.5 with Theorem 4.3 still yields an unspecified constant and does not establish a rate that remains meaningful when η is sent to infinity with the sample size. The paper should either prove a general bound, e.g., λ(η)=O(η) under explicit conditions, or restrict the asymptotic-exactness claim to the population level and present the finite-sample theory for fixed η with a separate treatment of the bias Vc - Vip(η).","section":"Section 4.2, Lemma 4.5"},{"comment":"The experimental section evaluates the estimator by L1 error against the oracle Vc and reports that the error 'decreases quickly for small η and then stabilizes' as η grows. This empirical behavior is consistent with a favorable λ(η), but it does not substitute for the missing general bound. In addition, the experiments always choose η from a fixed grid (0 to 100) rather than by a principled rule, so the finite-sample message in Section 5.2 is stronger than what Theorem 4.3 currently guarantees. I would ask the authors to either provide a practical selection procedure that balances the bias term Vc - Vip(η) against the variance term Cληγ_{N,d}, or to state explicitly that the rate theory covers only fixed η and that the near-exact regime is justified only at the population level.","section":"Section 4.2, Theorem 4.3 and Section 5.2"}],"minor_comments":[{"comment":"The words 'narrower PI intervals for any value of the penalty parameter' should be 'no wider', because Proposition 3.6 guarantees only Vip(η) ≥ Vu and equality can occur, for example when the covariates are independent of the potential outcomes.","section":"Abstract and Remark 3.7"},{"comment":"There are several typos that should be corrected in a revision: 'Founier' should be 'Fournier' in the C.2 heading, 'uniqueneness' should be 'uniqueness' in the proof of Proposition 3.5, 'certein' should be 'certain' in Remark 4.4, and 'satistified' should be 'satisfied' in the proof of Proposition 3.5.","section":"Appendix C.2, C.3, C.4, Remark 4.4"},{"comment":"The constraint notation '1^T π = (1/m)1^T, π1 = (1/n)1' is easy to misread; I suggest writing π ∈ R^{n×m}_+ with row sums 1/m and column sums 1/n, and stating explicitly that the displayed constraint is over matrix couplings.","section":"Algorithm 1, Eq. (4)"},{"comment":"The 'relative sample size' row in Table 2 is not defined in the text. The authors should state that the ratio is Vu/Vip,n,m(η) (or the equivalent normalized variance ratio) and explain why it corresponds to the relative sample size required for a given level of statistical power.","section":"Section 5.3.1, Table 2"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the population-level interpolation result (Proposition 3.6) appears correct and is a solid contribution, and the method is easy to implement and well motivated. The main issue is that the finite-sample convergence guarantee is not uniform in η, and the paper explicitly concedes that no general bound on λ(η) is available. I would not reject on this basis, because the central population claim is sound, but the advertised 'asymptotically exact' statistical statement needs either a substantial new theoretical argument or a careful restatement of the scope of the rate result. The authors should also be encouraged to separate their comparisons from DualBounds instances that produce invalid estimates (e.g., correlation upper bound above 1 in Table 3) and to report a principled choice of η."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this paper gives a practical way to tighten causal partial-identification bounds using covariates without solving conditional OT. The idea is simple: couple (Y(0), Z) and (Y(1), Z) with standard OT on an augmented cost h + eta||Z0 - Z1||^2. Proposition 3.6 shows Vip(eta) starts at the covariate-free OT bound, increases monotonically, and converges to the sharp COT bound as eta -> infinity. That is clean and proved properly under compactness. The mirror relaxation itself is not new — Carlier et al. 2010 and the conditional OT literature are cited — but the causal PI interpretation and the interpolation result in this context are new. The plug-in estimator is easy to implement with off-the-shelf OT solvers, code is provided, and simulations plus the STAR analysis show it tightens intervals and is robust where the DualBounds nuisance-based method struggles. Credit where due: this is a useful, model-lean method.\n\nThe soft spot is exactly the one the stress-test flags. Theorem 4.3's rate is C lambda eta gamma_{N,d}, where lambda is the curvature of a Brenier potential that depends on eta. The paper admits in Section 4.2 that it has no general bound on lambda(eta). Lemma 4.6 gives linear growth only for Gaussian marginals; Appendix C.7 is informal. To actually use the asymptotic exactness, you need eta -> infinity with N, and the error bound has eta and potentially lambda(eta) multiplying the empirical Wasserstein rate. If lambda grows superlinearly, the advertised rate is not a rate. Even with linear growth, the bias-variance tradeoff is not worked out and there is no data-driven eta selection. So the population-level story is solid, but the statistical guarantee, as stated, is incomplete. I would not call this fatal: Proposition 3.6 is correct, the method is useful with finite eta, and the experiments are honest. But the Theorem 4.3 claim should be softened or accompanied by a curvature bound in relevant cases.\n\nWho is this for? Applied people doing PI with covariates and methodologists working on OT relaxations. It deserves a serious referee — an editor should send it out, not desk-reject. I would cite it for the interpolation result and the practical recipe, and I'd want the authors to address the curvature dependence in revision.","headline":"Clean interpolation result with a real but incomplete finite-sample guarantee; worth refereeing and probably citing.","tokens_in":35857,"tokens_out":1590,"would_cite":true,"duration_ms":17052,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49Q22","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A penalty on duplicate covariates converts conditional optimal transport into ordinary optimal transport, producing valid partial-identification bounds that tighten toward the sharp covariate-adjusted limit.","keywords":["causal inference","partial identification","conditional optimal transport","mirror covariates","optimal transport","covariate adjustment","finite-sample rates","quadratic causal estimands"],"falsifier":"Compute or estimate the curvature λ of the Brenier potential between $P_{A_{12}^\\top Y(0),\\,\\eta Z}$ and $P_{Y(1),Z}$ for a non-Gaussian family as $\\eta$ grows and check whether $\\lambda=O(\\eta)$; a superlinear growth would break the uniformity of the stated bound. Alternatively, simulate the plug-in estimator at fixed $\\eta$ across increasing sample sizes $n,m$ and compare the empirical $L^1$ error to $C\\lambda\\eta\\,\\gamma_{N,d}$.","tokens_in":34792,"feed_emoji":"🎯","tokens_out":6771,"duration_ms":68983,"temperature":0.7,"pith_summary":"When a causal effect is only partially identified, the range of plausible values shrinks if covariates are used, but the sharp covariate-adjusted bound requires solving conditional optimal transport, which is hard to estimate and compute. This paper proposes a mirror relaxation: duplicate the covariates as Z(0) and Z(1), couple the two covariate-outcome distributions, and add a penalty η‖Z(0)−Z(1)‖² to the transport cost. The resulting problem is standard optimal transport, so existing solvers apply, and its lower bound Vip(η) lies between the covariate-free bound Vu and the sharp conditional bound Vc: it equals Vu at η=0 and converges to Vc as η→∞. A plug-in estimator is consistent and, for quadratic cost functions, has an explicit finite-sample convergence rate. If these claims hold, practitioners can obtain tighter covariate-adjusted bounds without estimating conditional distributions or solving a separate transport problem at each covariate value.","feed_headline":"Mirror-covariate transport interpolates plain to covariate-sharp causal bounds","feed_subtitle":"A penalty that drives duplicate covariates together turns hard conditional optimal transport into standard OT and narrows…","key_machinery":"The central object is the mirror relaxation Vip(η): it introduces mirror covariate copies Z(0) and Z(1), requires that the marginals of (Y(0),Z(0)) and (Y(1),Z(1)) match the observed covariate-outcome laws, and minimizes h(Y(0),Y(1)) plus η‖Z(0)−Z(1)‖²₂ over all such couplings. This is an unconditional optimal transport problem on the augmented space (Y×Z)², so off-the-shelf OT solvers apply; the penalty term carries the covariate information, forcing the two covariate copies to align as η grows and thereby recovering the conditional transport constraint Z(0)=Z(1) almost surely in the limit. For quadratic cost functions, the finite-sample analysis is carried by the Brenier potential between the transformed marginal distributions, whose strong convexity and smoothness constant λ enters the rate.","core_discovery":"The paper's central result, Proposition 3.6, is that the mirror relaxation Vip(η) interpolates between the two partial-identification boundaries: Vip(0)=Vu and lim_{η→∞} Vip(η)=Vc, with Vu≤Vip(η)≤Vc and monotonic continuity in η. In words, a single family of ordinary optimal transport problems, each minimizing h(Y(0),Y(1)) plus η‖Z(0)−Z(1)‖²₂ over couplings of the two observed covariate-outcome laws, produces a valid lower bound that is never worse than the covariate-free bound and can be driven arbitrarily close to the sharp conditional bound by increasing the penalty. The paper also constructs a plug-in estimator from the empirical covariate-outcome measures and proves consistency; for quadratic cost functions, the expected absolute error is bounded by Cλη·γ_{N,d}, where γ_{N,d} is the empirical Wasserstein convergence rate and λ is the curvature of the associated Brenier potential.","pith_inferences":["The mirror penalty can be viewed as a Lagrangian penalty on the constraint Z(0)=Z(1); a natural extension the paper does not develop is to select η by monitoring the expected mismatch E‖Z(0)−Z(1)‖² under the optimal coupling, treating it as a duality gap.","The same penalty-on-copies device should apply to other hard OT variants, including the causal or adapted OT relaxation sketched in the appendix, suggesting a general recipe for converting conditional constraints into unconditional OT penalties.","Because the rate γ_{N,d} worsens with dimension, the practical benefit of covariate adjustment may be offset by estimation error unless covariate selection, which the paper lists as future work, is applied.","The population interpolation result is proved under compactness and smoothness assumptions; the Gaussian example shows the conclusion can survive without compactness, but a general extension would require explicit tail or moment conditions."],"forward_implications":["For any finite η the mirror-relaxation bound is a valid partial-identification bound and is never wider than the covariate-free OT bound, so covariates can be exploited without estimating conditional marginals.","Increasing η moves the bound monotonically toward the sharp covariate-adjusted COT bound, so one algorithmic template covers the full range from Vu to Vc.","The plug-in estimator is consistent and can be computed with standard OT solvers, avoiding the per-covariate computation and nuisance-function estimation required by direct COT approaches.","For quadratic causal estimands such as Neymanian variance, null-effect tests, and the correlation of potential outcomes, the finite-sample error decays at the empirical Wasserstein rate γ_{N,d} up to the factor λη.","In the STAR data experiment the method narrows the Neymanian confidence interval and produces a correlation upper bound below one, which the paper reads as evidence of treatment-effect heterogeneity."],"supporting_citations":[{"why":"Supplies the existence, uniqueness, and gluing results for optimal transport couplings used in the proofs of Proposition 3.5 and Proposition 3.6.","marker":"Villani et al., 2009"},{"why":"Provides the convex-potential representation of optimal maps for quadratic costs that underlies Theorem 4.3.","marker":"Brenier, 1991"},{"why":"Introduces the block-triangular-map and continuation-method idea that the mirror relaxation extends.","marker":"Carlier et al., 2010"},{"why":"Supplies the plugin-estimation rate strategy for smooth optimal transport maps that Theorem 4.3 adapts to general quadratic costs.","marker":"Manole et al., 2024"},{"why":"Gives the empirical Wasserstein convergence rates that define the rate γ_{N,d} in Theorem 4.3.","marker":"Fournier & Guillin, 2015"},{"why":"Provides the existing covariate-assisted conditional-optimal-transport method used as a comparison in the experiments.","marker":"Ji et al., 2023"},{"why":"Gives the multi-marginal optimal transport formulation for causal partial identification that serves as the unconditional OT baseline.","marker":"Gao et al., 2025"},{"why":"Documents the failure of naive plug-in conditional OT estimators and motivates the relaxation to ordinary OT.","marker":"Chemseddine et al., 2024"}],"fun_headline_variants":["Penalty-driven OT tightens causal bounds to covariate sharpness","Covariate-aware OT relaxation tightens causal bounds with a penalty","Single OT family interpolates causal bounds via covariate penalty","From plain to sharp causal bounds: one OT penalty adjusts","Penalty tunes OT to move causal bounds from loose to sharp"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The finite-sample rate theorem assumes the Brenier potential between the transformed covariate-outcome measures is globally strongly convex and smooth, and the paper has no general bound on how its curvature constant λ scales with the penalty η, so if λ grows faster than linearly the advertised rate degrades.","fun_headline_variants_meta":{"raw":{"variants":["Penalty-driven OT tightens causal bounds to covariate sharpness","Covariate-aware OT relaxation tightens causal bounds with a penalty","Single OT family interpolates causal bounds via covariate penalty","From plain to sharp causal bounds: one OT penalty adjusts","Penalty tunes OT to move causal bounds from loose to sharp"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000839,"raw_usage":{"total_tokens":3664,"prompt_tokens":956,"completion_tokens":2708,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":2622}},"tokens_in":572,"tokens_out":2708,"duration_ms":19741,"temperature":1.0,"reasoning_tokens":2622,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T16:21:03.757782+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute or estimate the curvature λ of the Brenier potential between $P_{A_{12}^\\top Y(0),\\,\\eta Z}$ and $P_{Y(1),Z}$ for a non-Gaussian family as $\\eta$ grows and check whether $\\lambda=O(\\eta)$; a superlinear growth would break the uniformity of the stated bound. Alternatively, simulate the plug-in estimator at fixed $\\eta$ across increasing sample sizes $n,m$ and compare the empirical $L^1$ error to $C\\lambda\\eta\\,\\gamma_{N,d}$.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the existence, uniqueness, and gluing results for optimal transport couplings used in the proofs of Proposition 3.5 and Proposition 3.6."},{"cited_title":"Polar factorization and monotone rearrangement of vector-valued functions","cited_arxiv_id":null,"evidence_quote":"Provides the convex-potential representation of optimal maps for quadratic costs that underlies Theorem 4.3."},{"cited_title":"From K nothe's transport to B renier's map and a continuation method for optimal transport","cited_arxiv_id":null,"evidence_quote":"Introduces the block-triangular-map and continuation-method idea that the mirror relaxation extends."},{"cited_title":"Plugin estimation of smooth optimal transport maps","cited_arxiv_id":null,"evidence_quote":"Supplies the plugin-estimation rate strategy for smooth optimal transport maps that Theorem 4.3 adapts to general quadratic costs."},{"cited_title":"and Guillin, A","cited_arxiv_id":null,"evidence_quote":"Gives the empirical Wasserstein convergence rates that define the rate γ_{N,d} in Theorem 4.3."},{"cited_title":"Bridging multiple worlds: Multi-marginal optimal transport for causal partial-identification problem","cited_arxiv_id":null,"evidence_quote":"Gives the multi-marginal optimal transport formulation for causal partial identification that serves as the unconditional OT baseline."}],"review_version":1}