{"id":"3d2e4e54-d1fd-44fa-9f17-9cf91ef4f768","arxiv_id":"2411.08998","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Optimal-transport alignment of pre- and post-deployment distributions recovers the Bregman cost behind strategic agent responses.","lead":"To predict how people will behave after a model judges them, the authors estimate the hidden cost of changing one's profile by matching pre-deployment and post-deployment data with optimal transport. If the method holds up, modelers can skip the guesswork about strategic responses and still use fast performative optimization.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central estimator only recovers the Bregman component of the true cost; if actual agent costs are not Bregman divergences, the paper provides no bound on the error in the response map, so the claimed plug-in guarantee does not follow.","rationale":"The reader's weakest_assumption is exactly the Bregman restriction on the true agent cost, and I agree that this is the single most load-bearing premise of the paper. The estimator is internally coherent when the true cost is a Bregman divergence: the first-order condition (3.2) is correct, the population identity is valid, and Theorem 4.3 provides a plausible parametric-rate statement for the correctly specified case. The code is released, which is independent support for reproducibility. The central weakness is external validity: the method estimates only the Bregman projection of the true cost, and no misspecification bound is supplied for the response map, which is the object actually used in plug-in performative risk minimization. The paper explicitly acknowledges that a general bivariate cost is unidentifiable from finitely many deployments, so the Bregman restriction is a deliberate modeling choice; but the abstract and summary describe the method as estimating the cost without emphasizing that this guarantee is conditional on the cost lying in the Bregman class. The proposed simulation with an L1 cost is a direct check of whether the recovered Bregman potential yields accurate response maps when the assumption is violated. If the test shows vanishing error, the concern is mitigated; if not, the gap is real. Since the reader's verdict is already CONDITIONAL and this concern is the basis for that condition, no change to the verdict is needed.","tokens_in":19950,"tokens_out":20050,"duration_ms":193157,"concrete_test":"Run Algorithm 1 on data generated from a non-Bregman cost, e.g., c(z,z') = ||z - z'||_1 with z in [0,1]^2 and B_theta(z') = theta'z' - ||z'||^2/2, using the convex-network Bregman class of Appendix B. Compute the exact response map T_theta by grid search and compare T_theta_hat for held-out theta at sample sizes n = 10^3, 10^4, and 10^5. If the mean transport error does not decay to zero as n grows, the Bregman assumption is load-bearing; if it does, the concern is empirically mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is the Bregman assumption in Section 3.1. The derivation of the population identity behind (3.1) starts from the first-order condition (3.2), which holds only when the true cost is c_phi*(z,z') = phi*(z') - phi*(z) - grad phi*(z)'(z'-z) for some strictly convex phi*. If the real cost is not in this class, no phi satisfies (3.2), and the equality (grad phi*)#P = (grad phi* - grad B_theta)#Q_theta that justifies the estimator is unavailable. In that case the minimizer of (3.1) is an uncontrolled projection of the true response map onto the Bregman family: Theorem 4.3 only guarantees convergence to gamma*, the population minimizer inside the assumed class, and the plug-in bound (4.7) contains a MisspErr term for which the paper provides no bound. The experiments do not test this: Section 5.2 and Appendix B simulate quadratic or other Bregman costs, and the B_theta-misspecification robustness in Figures 1 and 2 is empirical only, with no supporting theory. The abstract's claim that the methodology estimates the cost and the distribution map is therefore conditional on a structural assumption that is neither validated nor quantified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an optimal-transport-based method for estimating, up to an affine adjustment of its Bregman potential, the cost function in a utility-maximizing microfoundation model of performative prediction. The method minimizes a Wasserstein-variance objective (3.1) that aligns push-forwards of the ex-ante distribution and one or more ex-post distributions, and it is extended to an ex-post-only variant (3.4). The authors establish an identifiability condition (Theorem 4.1) and a corollary for the single-deployment case (Corollary 4.2), and they prove a parametric convergence rate for the cost parameters (Theorem 4.3). Experiments on a credit-scoring dataset show accurate recovery of the Bregman potential when the benefit function is well specified, empirical robustness of the response map to benefit-function misspecification, and competitive plug-in performance for performative risk minimization.","tokens_in":20186,"tokens_out":10012,"duration_ms":91787,"significance":"If the structural assumptions hold, this is a valuable contribution: it converts a misspecified-microfoundation problem into a semi-parametric estimation problem, offers a verifiable identifiability criterion, and provides a plug-in route to performative optimization with fast rates. The population equality underlying (3.1) is correct at the true potential, and the paper is transparent about the non-identifiability of general bivariate costs. The ex-post-only variant and the empirical finding that response-map estimates are robust to benefit misspecification are useful additions. The public code and the use of a standard dataset support reproducibility. However, the strength of the claims is currently limited by the unquantified cost-misspecification gap and by an over-stated convergence rate in low dimensions.","major_comments":[{"comment":"The estimator is only guaranteed to recover the Bregman component of the cost. If the true cost is not a Bregman divergence, the first-order condition (3.2) has no solution for any φ, and the distributional equality behind (3.1) fails; the minimizer of (3.1) is then an uncontrolled projection of the response map onto the Bregman family. Theorem 4.3 bounds only the deviation from γ*, the population minimizer inside the assumed class, and the plug-in guarantee in Eq. (4.7) contains a MisspErr term for which no bound is supplied. Since the abstract claims that the methodology estimates the cost and the distribution map, this gap needs either a quantitative misspecification analysis or a more carefully scoped statement of the claims.","section":"Section 3.1 and Eq. (4.7)"},{"comment":"The stated rate E||bγ - γ*||^2 ≤ K n^{-2/d} is not valid for d = 1 and d = 2. For empirical W2^2, the sharp rates are O(n^{-1}) in d = 1 and O(n^{-1} log n) in d = 2, whereas n^{-2/d} would give n^{-2} and n^{-1} respectively; the d = 1 rate is impossible. The cited result of Manole and Niles-Weed (2024) covers the regime d ≥ 3 (with corrections for d = 2). The theorem should either restrict to d ≥ 3 or state the correct rates for d ≤ 2, and this matters because Section 5.1 reports one-dimensional experiments.","section":"Theorem 4.3, Eq. (4.4)"},{"comment":"The proof of the one-ex-ante/one-ex-post identifiability claim relies on the assertion that T^n(z) → z* for any strictly concave Bθ with a finite maximizer, where T is the proximal-type map in (A.9). This convergence is not an immediate consequence of strict concavity; it requires additional hypotheses such as strong concavity or coercivity/essential smoothness of φ - Bθ, or an explicit appeal to a Bregman proximal point convergence theorem with stated conditions. Without such support, the corollary is not fully established.","section":"Corollary 4.2 and its proof in Appendix A"},{"comment":"The strong convexity assumptions on γ ↦ min_{Π∈Δ(P,Qθ)} L(Π,γ) and on its empirical counterpart are stated without any primitive conditions under which they hold. They are not verified for the quadratic class φ(z) = (1/2)z^T M z used in Section 5.2, nor for the nonparametric isotonic-regression setting in Section 5.1. Since these assumptions are exactly what turns the Wasserstein stability bound into a parameter-rate, the theorem's hypotheses are not connected to the practical instantiations of the method.","section":"Theorem 4.3 assumptions"}],"minor_comments":[{"comment":"The text says 'induce a Bergman divergence' but the intended term is 'Bregman divergence'.","section":"Section 4, after Eq. (4.1)"},{"comment":"The phrase 'from a finite number of ex-ante distributions' appears to be a slip; it should presumably read 'ex-post distributions'.","section":"Section 3.1"},{"comment":"The caption says 'well estimated when he benefit function'; it should read 'when the benefit function'.","section":"Figure 1 caption"},{"comment":"In the final sentence, 'Since θ is continuous and strictly increasing' should read 'Since f is continuous and strictly increasing'.","section":"Proof of Lemma A.1"},{"comment":"The term 'conservative solutions' is nonstandard for 'gradients of convex functions'; consider using 'functions that are derivatives of convex functions' or 'subgradients of convex functions' instead.","section":"Theorem 4.1"},{"comment":"The alternating minimization is stated to converge by Tseng (2001), but no verification is given that the required regularity conditions hold for the nonparametric class of potentials; a brief comment would help.","section":"Algorithm 1 and Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The core idea is sound and the paper is likely acceptable after revision. The most urgent fixes are the low-dimensional rate statement in Theorem 4.3, which should be reconciled with the sharp empirical W2 literature, and the unstated regularity needed for the convergence of T^n in Corollary 4.2. The Bregman assumption is the central limitation; it is acknowledged in the paper but the plug-in claim in Eq. (4.7) is incomplete without a bound on the misspecification term. I would encourage the authors to add a scope caveat in the abstract or a short misspecification analysis, even if only for a tractable subclass."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper has a real new idea, the identifiability analysis is worth stealing, and the Bregman restriction is more than a technical convenience—it is the load-bearing assumption, and the paper doesn't tell you what happens when it fails.\n\nThe new thing is the formulation: instead of assuming a manipulation graph or a linear response, they cast the unknown cost as a Bregman divergence and estimate its potential by aligning the pushforwards of ex-ante and ex-post distributions with a Wasserstein barycenter. The population equality behind (3.1) is correct, and the idea to use first-order conditions to get the alignments is elegant. The identifiability condition in Theorem 4.1 is genuinely useful—it gives practitioners a check—and the corollary that one ex-ante and one ex-post distribution can identify the cost, given strict concavity of B, is a nice result. The paper also ships code, and the experiments show the estimator works when the model is correctly specified.\n\nThe soft spots, in order. First, the Bregman assumption. The paper is upfront that general costs are unidentifiable from finitely many deployments, and restricting to Bregman is a natural response. But the method inherits all its guarantees from that assumption. If the true cost is not Bregman, the population equality fails and the minimizer of (3.1) is an uncontrolled projection; the plug-in bound (4.7) has a MisspErr term that is never bounded. The experiments do not test this—they simulate Bregman costs, and the benefit-misspecification robustness in Figures 1–2 is empirical only, with no theory. Second, the rate theorem (Theorem 4.3) is narrower than the abstract suggests: it covers a parametric class, assumes strong convexity of the objective, and the n^{-2/d} rate inherits the dimension dependence of empirical Wasserstein. The extension to multiple ex-post distributions is left open. These are addressable rather than fatal, and I would not block publication on them, but they should be stated clearly.\n\nThird, a minor thing: the text says 'Bergman' in a few places. Not important.\n\nWho is this for? Anyone working in performative prediction or strategic classification who wants to go beyond assuming a known distribution map. The method gives a principled way to estimate the map from data, and the identifiability check is a practical contribution. This deserves a serious referee: the idea is novel, the analysis is honest, and the limitations are mostly explicit. I would ask the referee to push on the misspecification question—either a bound on the response-map error under non-Bregman costs or a sensitivity analysis—before accepting.","headline":"A genuinely new estimator for Bregman costs in strategic prediction, with the usual caveat that the structural assumption is doing more work than the paper fully acknowledges.","tokens_in":20742,"tokens_out":2445,"would_cite":true,"duration_ms":22248,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62R07","62F12","49Q22"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that the hidden cost driving agents' strategic responses to a predictive model can be estimated, up to an affine shift, by aligning pre- and post-deployment distributions with optimal transport.","keywords":["performative prediction","strategic classification","microfoundation inference","Bregman divergence","optimal transport","Wasserstein barycenter","distribution shift","cost estimation"],"falsifier":"Simulate an environment whose true cost is a non-Bregman bivariate function, for example $c(z,z')=\\|z-z'\\|_2^3$ or $c(z,z')=\\|z-z'\\|_2^2+g(z)$ with a state-dependent term, while agents solve (2.1) with a known strictly concave benefit. If, with large samples, the recovered Bregman potential fails to reproduce the true response map $T_\\theta$, the central claim fails. A cheaper check is to compare the predicted push-forward $(\\hat T_\\theta)_\\#P$ against held-out ex-post samples under such a deliberately non-Bregman cost.","tokens_in":19709,"feed_emoji":"🎯","tokens_out":11873,"duration_ms":87867,"temperature":0.7,"pith_summary":"The paper asks whether a practitioner can learn how a population of strategic agents reacts to a deployed predictive model without being told the agents' cost of changing their attributes. It answers yes, provided that cost is a Bregman divergence: from samples of the population before deployment and after one or more deployments, one can recover the gradient of the Bregman potential up to a constant shift, and hence recover the response map. The identification rests on the fact that in the potential's gradient coordinates, the pre-deployment and post-deployment push-forward distributions coincide. If the paper is right, performative prediction can move from derivative-free optimization or grossly misspecified microfoundations to plug-in estimation of the performative distribution map, with a parametric convergence rate.","feed_headline":"Unknown agent costs recoverable from before-and-after samples","feed_subtitle":"Estimating the cost turns performative prediction into a plug-in problem with fast convergence, instead of guessing a misspecified model.","key_machinery":"The load-bearing object is the Bregman divergence $c_\\phi(z,z') = \\phi(z')-\\phi(z)-\\nabla\\phi(z)^\\top(z'-z)$ induced by a strictly convex potential $\\phi$; restricting costs to this class reduces the infinite-dimensional problem of estimating a bivariate cost to estimating the gradient $\\nabla\\phi$. The first-order condition of the agent's utility maximization, $\\nabla\\phi(T_\\theta(z))-\\nabla B_\\theta(T_\\theta(z))=\\nabla\\phi(z)$, implies that $(\\nabla\\phi)_\\#P$ and $(\\nabla\\phi-\\nabla B_\\theta)_\\#Q_\\theta$ coincide, so the estimator aligns these measures by minimizing the sum of squared 2-Wasserstein distances to a common barycenter. Identifiability is characterized by whether the only conservative (gradient-of-convex) functions $h$ satisfying $h\\circ T_0 = h\\circ T_1 = \\cdots = h\\circ T_m$ are constants, and the convergence rate follows from sharp empirical optimal-transport bounds for smooth costs.","core_discovery":"The central claim is that the unknown cost $c(z,z')$ in the utility-maximization model $T_\\theta(z) \\in \\arg\\max_{z'} B_\\theta(z') - c(z,z')$ is identifiable from the ex-ante distribution $P$ and one or more ex-post distributions $Q_\\theta=(T_\\theta)_\\#P$, up to an affine adjustment of the Bregman potential, whenever $c=c_\\phi$ is a Bregman divergence $\\phi(z')-\\phi(z)-\\nabla\\phi(z)^\\top(z'-z)$ with strictly convex potential $\\phi$. The proposed estimator minimizes the Wasserstein variance over potentials $\\phi$ and a barycenter $\\mu$, namely $\\min_{\\phi,\\mu} \\sum_{k=0}^m W_2^2(\\mu,\\, (\\nabla\\phi-\\nabla B_{\\theta_k})_\\# Q_{\\theta_k})$ with $(B_{\\theta_0},Q_{\\theta_0})=(0,P)$, forcing the aligned push-forward distributions to coincide. Theorem 4.1 reduces identifiability to a condition on conservative solutions of a system of equations induced by the optimal transport maps, and Corollary 4.2 shows that one ex-post distribution suffices when the benefit is strictly concave with a finite maximizer. Theorem 4.3 establishes the parametric rate $\\mathbb{E}\\|\\hat\\gamma-\\gamma^\\star\\|_2^2 \\le K n^{-2/d}$ for estimating the parameters of the potential from $n$ i.i.d. samples of $P$ and of $Q_\\theta$. The authors also show experimentally that the estimated response map remains accurate even when the benefit function is misspecified, although the estimated potential is then biased.","pith_inferences":["If the true cost is not a Bregman divergence, the estimator returns the Bregman projection of the true cost onto the assumed class, but the paper gives no bound on the error this induces in the response map; bounding or testing this projection error with held-out ex-post samples is a natural next step.","Theorem 4.1's identifiability condition can be checked empirically from the estimated transport maps, so a practitioner with finitely many deployments can verify whether the only conservative solutions of the system are constants before trusting the cost estimate.","The $n^{-2/d}$ rate suggests the method will struggle in high-dimensional attribute spaces; incorporating structural assumptions such as separable or low-rank costs is an open direction the paper's parametric experiments only begin to explore.","The same distribution-alignment principle could apply to non-strategic performative shifts driven by a deterministic map, which the paper names as future work, though the cost-recovery interpretation would then need a different microfoundation."],"forward_implications":["With access to the ex-ante distribution and a single ex-post distribution, the cost is identifiable whenever the known benefit function is strictly concave with a finite maximizer, so one model deployment can suffice to learn the microfoundation.","The estimated potential parameters converge at the rate $K n^{-2/d}$ in squared error, so the statistical error in plug-in performative risk minimization inherits this rate through the decomposition into misspecification and statistical error.","The estimated response map $\\hat T_\\theta$ solves $\\nabla\\hat\\phi(\\hat T_\\theta(z))-\\nabla B_\\theta(\\hat T_\\theta(z))=\\nabla\\hat\\phi(z)$ and can be plugged directly into performative risk minimization, enabling fast white-box optimization algorithms.","The method extends to settings with only ex-post distributions by aligning the push-forwards of the agent responses alone when no pre-deployment sample is available.","Empirically, misspecification of the benefit function biases the estimated potential but not the estimated response map, so downstream performative risk minimization remains accurate under this misspecification."],"supporting_citations":[{"why":"Defines the performative prediction setting and the microfoundation model in Section 2 that the paper's estimator is built on.","marker":"Perdomo et al., 2020"},{"why":"Introduces strategic classification with a squared-distance cost, the canonical example of a Bregman cost that the method generalizes.","marker":"Hardt et al., 2016"},{"why":"Supplies the causal strategic linear regression example whose misspecified-cost failure motivates cost estimation.","marker":"Shavit et al., 2020"},{"why":"Gives the decomposition of plug-in performative risk error into misspecification and statistical terms used in eq. (4.7).","marker":"Lin and Zrnic, 2023"},{"why":"Provides the sharp empirical optimal-transport convergence bound that yields the $n^{-2/d}$ rate in Theorem 4.3.","marker":"Manole and Niles-Weed, 2024"},{"why":"Supplies the optimal-transport results used in the identifiability proof of Theorem 4.1.","marker":"Villani, 2009"},{"why":"Introduces Bregman divergences, the class of costs within which the unknown cost is estimated.","marker":"Bregman, 1967"}],"fun_headline_variants":["Recover agent costs from before-and-after samples","Infer hidden utility costs from pre-post distributions","Optimal transport unveils strategic costs in prediction","From ex-ante to ex-post: identify agent costs","Learn performative costs from distribution shift data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the agents' true cost is a Bregman divergence with a strictly convex potential, and that the benefit function is known; if the real cost is not of this form, the distributional equality the estimator relies on fails and the recovered potential is only an unquantified projection of the true cost.","fun_headline_variants_meta":{"raw":{"variants":["Recover agent costs from before-and-after samples","Infer hidden utility costs from pre-post distributions","Optimal transport unveils strategic costs in prediction","From ex-ante to ex-post: identify agent costs","Learn performative costs from distribution shift data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1494,"prompt_tokens":1036,"completion_tokens":458,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":386}},"tokens_in":652,"tokens_out":458,"duration_ms":5197,"temperature":1.0,"reasoning_tokens":386,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:11:01.415880+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an environment whose true cost is a non-Bregman bivariate function, for example $c(z,z')=\\|z-z'\\|_2^3$ or $c(z,z')=\\|z-z'\\|_2^2+g(z)$ with a state-dependent term, while agents solve (2.1) with a known strictly concave benefit. If, with large samples, the recovered Bregman potential fails to reproduce the true response map $T_\\theta$, the central claim fails. A cheaper check is to compare the predicted push-forward $(\\hat T_\\theta)_\\#P$ against held-out ex-post samples under such a deliberately non-Bregman cost.","supporting_citations":[{"cited_title":"Strategic Classification","cited_arxiv_id":null,"evidence_quote":"Introduces strategic classification with a squared-distance cost, the canonical example of a Bregman cost that the method generalizes."},{"cited_title":"Edelman, and Brian Axelrod","cited_arxiv_id":null,"evidence_quote":"Supplies the causal strategic linear regression example whose misspecified-cost failure motivates cost estimation."},{"cited_title":"Sharp convergence rates for empirical optimal transport with smooth costs","cited_arxiv_id":null,"evidence_quote":"Provides the sharp empirical optimal-transport convergence bound that yields the $n^{-2/d}$ rate in Theorem 4.3."},{"cited_title":"Optimal Transport: Old and New","cited_arxiv_id":null,"evidence_quote":"Supplies the optimal-transport results used in the identifiability proof of Theorem 4.1."}],"review_version":1}