Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

Optimal treatment assignment rules under capacity constraints

T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper proves that capacity-constrained treatment assignment is optimally solved by a Bayesian rule over couplings of covariates and treatment quotas, while the plug-in rule fails under nonsmooth welfare.

desk verdict The optimal transport reformulation is the genuinely useful part; the paper's main theorem rests on Assumption 3.6, which fails for the very reasons its own convex setup implies. read the letter →

arxiv 2506.12225 v2 pith:NJ3BM6FE submitted 2025-06-13 econ.EM

classification econ.EM MSC 62C1062F1249Q2262P20
keywords treatmentassignmentcapacityconstraintsoptimaltransportBayesiandecisionruleslocalasymptoticoptimalitydirectionaldifferentiabilitysocialwelfarestatisticaltheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

When a treatment is scarce and a fixed share of the population can receive it, the planner must coordinate assignment probabilities across the entire covariate distribution, which breaks the pointwise logic of unconstrained rules. The paper's central proposal is to recast this constrained welfare maximization as an unconstrained optimal transport problem: choose a coupling between the covariate distribution and the treatment-quota distribution, so any feasible coupling automatically respects the capacity constraint. Within a local asymptotic framework, the paper proves that a penalized Bayesian rule is average optimal among rules that are asymptotically optimal at the true parameter, while the standard plug-in rule—rank by estimated benefit and fill the quota—is average optimal only when the welfare function is fully differentiable. When welfare is only directionally differentiable, as under distributionally robust (maxmin) preferences, the Bayesian rule remains optimal and the plug-in rule generally does not, a distinction the simulations and a Colombian voucher application illustrate.

What carries the argument

The central object is the action space M of all couplings of the covariate distribution FX and the treatment-quota distribution FT, equipped with the Wasserstein distance d_W; solving max over M of the welfare functional W(θ,µ) automatically enforces the capacity constraint, so the constrained problem becomes an unconstrained one over a compact metric space. To select a well-defined Bayesian rule from the possibly multiple maximizers of the posterior welfare integral, the paper adds the penalty H(µ) = (d_W(µ,ν))^2 for a fixed reference coupling ν, whose strict convexity yields a unique penalized rule that converges to the minimum-penalty Bayesian rule. The asymptotic arguments then ride on an asymptotic representation theorem for rules in D, a Bernstein–von Mises posterior concentration bound, and the directional differentiability of the welfare map, which together identify the limit-experiment loss that average optimality requires.

What would settle it

Run the paper's simulation with the robustness kink parameter ε set so the welfare contrast is exactly zero on a region of positive covariate mass, so that the set of optimal couplings has a flat boundary and Assumption 3.6's uniform gap fails, and check whether √n P(µ_B ∉ A0) still converges to zero and the Bayesian rule's average risk stays below the plug-in rule's; if either fails, the separation assumption is doing the work.

Watch

Extended reading notes

Core claim

The paper's central claim is that under Assumptions 3.1–3.6 the penalized Bayesian rule {µ_B_n} is average optimal in the class D of decision rules that are asymptotically optimal at the true parameter, and that the plug-in rule attains average optimality under full differentiability but generally not when the welfare function is only directionally differentiable. The key to the result is the observation that the planner's constrained problem is equivalent to maximizing W(θ,µ) over the Wasserstein space M of couplings of FX and FT, which makes the action space metric and compact and the capacity constraint automatic. In the limit experiment delivered by the asymptotic representation theorem, any optimal rule must minimize a loss built from the directional derivative of the welfare function; the Bayesian rule converges to that minimizing coupling, whereas the plug-in rule converges to the maximizer of a different objective that coincides only under linearity of the derivative.

Load-bearing premise

The proof depends on Assumption 3.6, which says the welfare value of every suboptimal treatment coupling is uniformly below the optimum value, with a gap that never shrinks to zero; if a sequence of near-optimal couplings approaches the optimal set, the argument that the Bayesian rule lands in the optimal set with probability tending to one collapses.

Editorial extensions

If this is right

  • When the welfare function is smooth, both the plug-in and Bayesian rules are average optimal and coincide in large samples, so practitioners can safely use the simpler ranking-based rule.
  • Under a directionally differentiable welfare function, the plug-in rule should be avoided because its limit objective differs from the one optimal rules must solve.
  • The optimal transport formulation reduces computation to a linear program when covariate and treatment distributions are discrete or discretized, so optimal rules are computable with existing software.
  • The framework extends beyond binary treatments to discrete and continuous treatment quotas, and to semiparametric models via least favorable submodels with quasi-Bayesian implementation.
  • In finite samples the Bayesian rule shows lower risk than the plug-in rule even under smooth welfare, suggesting a practical small-sample advantage.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The uniform separation assumption in Assumption 3.6 is likely to be the binding constraint in real applications where the welfare function is flat near the optimum; practitioners should test for near-ties in the optimal coupling before relying on the Bayesian rule's asymptotic guarantee.
  • Entropic regularization of the coupling problem would give a smoothed, strictly convex objective whose Sinkhorn algorithm approximates the penalized Bayesian rule; the paper's penalty construction suggests a principled way to tune the regularization parameter.
  • The plug-in failure under directional differentiability probably extends to minimax optimality, as the paper conjectures; a minimax version of the Bayesian rule in this capacity-constrained setting would be a natural next step.
  • The same optimal-transport-of-couplings device could be applied to other quota-constrained decisions, such as matching policies in two-sided markets or budget-constrained algorithmic fairness, wherever the constraint is a fixed marginal distribution.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper reformulates the capacity-constrained treatment assignment problem as an optimal transport problem: the planner chooses a coupling of the covariate distribution and the treatment distribution, so the capacity constraint is automatically satisfied. Within a local asymptotic limits-of-experiments framework, the paper claims that a penalized Bayesian rule is average optimal when the planner's welfare function is directionally differentiable, while the plug-in rule is average optimal only under full differentiability. The paper also provides a simulation study and an empirical illustration based on Angrist et al. (2006).

Significance. If the main theorem were correct, the paper would be a useful extension of Hirano and Porter (2009) and Christensen et al. (2025) to capacity constraints and fractional assignments, and the optimal transport formulation is elegant and computationally attractive. However, the central optimality result rests on Assumption 3.6, which is impossible in every nontrivial instance of the model. As a result, the main claim is not established in the setting the paper actually targets, including the paper's own simulation design. The optimal transport reformulation itself is a genuine contribution, and the exposition is generally clear, but the core decision-theoretic result is vacuous as stated.

major comments (3)
  1. [Section 3.2, Assumption 3.6 and Lemma A.4] Assumption 3.6 is internally inconsistent with the maintained structure whenever A0 is not all of M. In Section 2.2, M is the convex set of couplings of FX and FT, and W(θ0,µ)=∫w(θ0,x,t)dµ is linear in µ. Fix µ*∈A0 and ν∈M\A0. For α∈(0,1), µ_α=(1-α)µ*+αν belongs to M, and W(θ0,µ_α)=(1-α)c+αW(θ0,ν)<c, where c=max_{µ∈M}W(θ0,µ). Thus W(θ0,µ_α)→c as α↓0, so sup_{ν∉A0}W(θ0,ν)=c. No K<c can satisfy the displayed inequality in Assumption 3.6. The assumption therefore forces A0=M, in which case welfare is constant across all couplings and the capacity constraint is irrelevant. This is not a missing derivation from primitives: equation (A.7) in Lemma A.4 requires a strictly positive uniform gap η, and no such gap can exist in any nontrivial setting with heterogeneous treatment effects, including the Section 4 simulation. Consequently Theorem 3.3 does not establish average optimality in the claimed setting.
  2. [Section 3.3.1, equation (3.4) and surrounding text] The paper's claim that the plug-in rule is generally not average optimal under directional differentiability is not formalized. The argument invokes a 'high-level condition that the process {B_n(µ):µ∈A0} is asymptotically tight', but this condition is neither stated among Assumptions 3.1-3.6 nor proved. Showing that the two limit maximization problems (3.3) and (3.4) can differ is not the same as proving that the plug-in rule fails average optimality; no theorem with a precise non-optimality statement or a counterexample is provided. Since the abstract presents this as a main finding, a formal statement with explicit assumptions is needed.
  3. [Section 2.3, definition of D in (2.6)] The admissible class D includes the condition √nP^n_{θ_nh}(µ_n∉A0)→0, and Theorem 3.3 must show that the Bayesian rule belongs to D. Lemma A.3(i) is the only place where this is established, and its proof depends on Lemma A.4, which in turn relies on the impossible uniform separation in Assumption 3.6. Therefore the proof provides no route to membership of the Bayesian rule in D for exactly the heterogeneous-treatment-effect cases the paper intends to cover.
minor comments (3)
  1. [Example 2.3] The text refers to the 'PACES program in Columbia'; the country should be Colombia, which is also the setting of Angrist et al. (2006).
  2. [Section 4] The text refers to 'the upper panels of Figure 4' before Figure 4 is introduced; the figure numbering and placement should be adjusted so that the reference is immediate.
  3. [Appendix A, Lemma A.7] There are typographical artifacts in the appendix, such as 'Q n =Q n +o_{P...}(1)' where the tilde over the second Q_n is missing; these should be cleaned up.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: Theorem 3.3 is derived from stated assumptions and external limits-of-experiments results, with no fitted input or load-bearing self-citation; the main caveat is an assumption-consistency concern, not circularity.

full rationale

The paper's central result is a derivation, not a repackaging of its inputs. Theorem 3.3 is proved from Assumptions 3.1–3.6 using external mathematical machinery: van der Vaart's asymptotic representation theorem (Proposition 2.4), Clarke-Barron/Bernstein-von Mises posterior behavior, and Christensen, Moon, and Schorfheide's lemmas for posterior concentration. No parameter or constant is fitted to the theorem's conclusion, and the Bayesian rule's optimality is established through the limit-experiment loss L∞, not assumed. The plug-in comparison also rests on a genuine distinction: the plug-in limit solves max_{µ∈A0} ∫ ˙w_{θ0}(x,t;Δ)dµ, whereas the optimal limit rule solves (3.3), max_{µ∈A0} ∫∫ ˙w_{θ0}(x,t;s)dN(Δ,I0^{-1})(s)dµ; these differ exactly when ˙w is nonlinear, which is the paper's claimed finding. There is no load-bearing self-citation: the authors cite Hirano and Porter, Christensen et al., Nutz, van der Vaart, and others, none of which are their own prior works. A substantive concern exists but it is not circular: Assumption 3.6 requires a uniform welfare gap between A0 and M\A0, and because W(θ0,µ) is linear in µ and M is convex, such a gap can only hold when A0=M, making the assumption incompatible with the intended heterogeneous-effect applications and with the Section 4 simulation. That is an internal consistency or assumption-coverage problem, not a case where a prediction is equivalent to its input by construction, and it does not make the derivation self-referential.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central theoretical claim is analytical; no numerical parameters are fitted to data in the proof. The optimal transport reformulation is an exact equivalence under the fixed-marginal capacity model. The strongest added assumptions are the uniform separation condition and the high-level tightness condition for plug-in rules, both of which are not derived from more primitive conditions.

assumptions (7)
  • standard math The model satisfies differentiability in quadratic mean with nonsingular information matrix I0 at each θ0.
    Section 2.3 and Assumption 3.1; standard local asymptotic normality regularity.
  • domain assumption The planner knows the target covariate distribution FX and the capacity constraint binds exactly with known treatment marginal FT.
    Section 2 and Section 2.2; the optimal transport reformulation is exact only when capacity is a fixed known marginal constraint.
  • domain assumption X×T is compact in the product metric space.
    Assumption 3.2; needed for compactness of the coupling space, Wasserstein metric, and uniform convergence arguments.
  • ad hoc to paper Uniform separation between A0 and M\A0 holds (Assumption 3.6).
    Ensures argmax invariance near θ0 and is used to prove √n P(µ_B_n ∉ A0)→0; it is not derived from primitive conditions and may fail under ties.
  • domain assumption Welfare w is bounded, continuous, and directionally differentiable with polynomial growth of the derivative (Assumptions 3.3 and 3.4).
    These regularity conditions are needed for the delta method and for process convergence on ℓ∞(M).
  • ad hoc to paper For the plug-in analysis, a best regular estimator exists and the process B_n(µ) is asymptotically tight.
    Section 3.3.1 imposes these high-level conditions; they are not part of Assumptions 3.1-3.6 and are not proved.
  • standard math Parametric posterior consistency holds via Schwartz's theorem under local quadratic and soundness conditions.
    Assumption 3.1(iv)-(v) and Lemma A.3; used to show the Bayesian rule belongs to the class D.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimal treatment assignment rules under capacity constraints." pith.science (2026). https://pith.science/paper/NJ3BM6FE

@misc{pith2026250612225,
  author       = {Pith},
  title        = {Pith review of: Optimal treatment assignment rules under capacity constraints},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NJ3BM6FE}},
  note         = {Machine review of arXiv:2506.12225}
}
read the original abstract

We study treatment assignment problems under capacity constraints, where a planner aims to maximize social welfare by assigning treatments based on observable covariates. Such constraints, common when treatments are costly or limited in supply, introduce nontrivial challenges for deriving optimal statistical assignment rules because the planner needs to coordinate treatment assignment probabilities across the entire covariate distribution. To address these challenges, we reformulate the planner's constrained maximization problem as an optimal transport problem, which makes the problem effectively unconstrained. We then establish local asymptotic optimality results of assignment rules using a limits of experiments framework. Finally, we illustrate our method with a voucher assignment problem for private secondary school attendance using data from Angrist et al. (2006)

Figures

Figures reproduced from arXiv: 2506.12225 by the authors.

Figure 1
Figure 1. Welfare contrasts under smooth and directionally differentiable welfare at θ0 We set β0 = (−2, −3), α0 = 4, and ui ∼ N(0, σ2 0 ) with σ0 = 10. The observed data is an i.i.d. sample Z n = {(Yi , Xi , Ti)} n i=1. The parameters θ0 = (β0, α0, σ0) can be estimated by the maximum likelihood using Z n . In this Tobit model, the conditional mean of the potential outcomes in the training population is given by w(θ, x, t) = … view at source ↗
Figure 2
Figure 2. Comparisons of estimated risks: n = 200 (left: smooth welfare, right: directionally differentiable welfare) [PITH_FULL_IMAGE:figures/full_fig_p017_2.png] view at source ↗
Figure 3
Figure 3. Comparisons of estimated risks: n = 500 (left: smooth welfare, right: directionally differentiable welfare) [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Comparison of (average) treatment assignment under directionally dif￾ferentiable welfare Note: The upper panels show the oracle rule, the middle show the Bayesian rule, and the lower show the plug-in rule under θ0 and n = 200 [PITH_FULL_IMAGE:figures/full_fig_p019_4.png]
Figure 5
Figure 5. Figure 5: Voucher allocations under smooth welfare Note: The upper panels show the plug-in rule, and the lower panels show the Bayesian rule. The color intensity represents the density of each cell in FX, with darker shades indicating higher density [PITH_FULL_IMAGE:figures/ful…
Figure 6
Figure 6. Figure 6: Voucher allocations under directionally differentiable welfare Note: The upper panels show the plug-in rule, and the lower panels show the Bayesian rule. The color intensity represents the density of each cell in FX, with darker shades indicating higher density. τ for …

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Who With Whom? Learning Optimal Matching Policies

    econ.EM 2025-07 conditional novelty 6.0 of 10

    An entropy-regularized optimal transport method learns welfare-optimal two-sided matching policies with estimated costs, supported by a non-asymptotic regret bound and calibrated simulations suggesting about one perce...

Reference graph

Works this paper leans on

4 extracted references · 2 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Externally valid policy choice

    Adjaho, C. and T. Christensen (2023). “Externally valid policy choice.” arXiv:2205.05561 [econ.EM]. Angrist, J., E. Bettinger, and M. Kremer (2006). “Long-term educational consequences of secondary school vouchers: Evidence from administrative records in Colombia.”American Economic Review96.3, 847–862.doi:10.1257/aer. 96.3.847. Angrist, J., E. Bettinger, ...

  2. [3]

    van der Vaart, A

    Cambridge University Press. van der Vaart, A. W. and J. A. Wellner (1996).Weak Convergence and Empirical Processes with Applications to Statistics. Springer-Verlag.doi:10.1007/978-1-4757-2545-2. Villani, C. (2009).Optimal Transport: Old and New. Vol

  3. [44]

    Robust Bayesian inference for set-identified models

    Cambridge University Press. Ghosh, J. K. and R. V. Ramamoorthi (2003).Bayesian Nonparametrics. Springer New York. Giacomini, R. and T. Kitagawa (2021). “Robust Bayesian inference for set-identified models.”Econometrica89.4, 1519–1556.doi:10.3982/ECTA16773. Gilboa, I. and D. Schmeidler (1989). “Maxmin expected utility with non-unique prior.”Journal of Math...

  4. [338]

    Asymptotic analysis of point decisions with general loss functions

    Springer. Xu, H. (2024). “Asymptotic analysis of point decisions with general loss functions.”Working Paper. Yata, K. (2023). “Optimal decision rules under partial identification.” arXiv:2111.04926 [econ.EM]. Department of Economics, University of Rochester, Rochester, NY 14627, USA. Email address:ksunada@ur.rochester.edu Department of Economics, Universi...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.