{"id":"ca82c8dd-515e-4ed9-b464-8cce09301809","arxiv_id":"2608.04669","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Training outcome models by differentiating a value estimate through dual prices yields allocation policies that respect long-run capacity limits and beat decision-blind predict-then-optimize baselines on deployment-adjusted value across six datasets.","lead":"Allocation policies for scarce resources are usually trained in two disconnected steps: predict outcomes, then price the resource. This paper trains prediction and pricing together by passing the training signal through the dual prices, and reports better deployment-adjusted value and fewer capacity violations across six datasets, including a 70,000-patient hospital cohort.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline DAPV ranking at κ=0 depends on the fixed simulation horizon and the V0 valuation of unserved arrivals; the robustness re-simulation asserted in §4.4 is not shown, and Fig. 3b shows two-stage leaders at κ=0 on three datasets.","rationale":"The paper's theoretical core is sound: the entropy-based F–G bound (Prop. 1), the in-expectation feasibility of the convex surrogate (Props. 2–3), and the excess-value decomposition (Thm. 6) are carefully stated and proved, with the global-maximizer assumption (A4) explicitly left undischarged (Remark 8). The evaluation is unusually transparent, including a refuted pre-specified hypothesis and negative ablations. The central claim is nonetheless the empirical one, and it rests on the DAPV index. The κ=0 term penalizes unserved arrivals at V0, measured at a fixed horizon; this is a legitimate design choice, but the paper's assertion that the ranking is unchanged across horizon multipliers is not shown, and Fig. 3b reveals that on three of five datasets the best two-stage method already leads at κ=0. The aggregate 'top slots at every delay cost' therefore depends on how the normalization and pooled mean weight Diabetes 130, where end-to-end wins. A longer horizon is the natural stress test: it reduces u for methods that leave queues, potentially flipping the κ=0 ranking. Since the robustness check is cheap to run with the existing harness and the claim appears in the abstract, this is the load-bearing concern. It does not invalidate the paper; it makes the headline conditional on an unverified evaluation choice. The reader's CONDITIONAL verdict is appropriate, and this concern reinforces it rather than moving it.","tokens_in":23407,"tokens_out":13906,"duration_ms":157952,"concrete_test":"Re-run the full queueing evaluation at T_max ∈ {1.25, 1.5, 2, 3, 5, 10} × last arrival, using the same paired arrival and resource streams and the same seeds. Recompute DAPV(κ) per dataset and the combined index at κ=0 and κ=0.01 for all methods, and record the rank of F and G against the best two-stage method. If the top two slots change for any horizon multiplier, the claimed 'every delay cost, including zero' result is not robust to the horizon choice and the central empirical claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that the two end-to-end variants hold the top two slots on the combined deployment-adjusted value index at every delay cost, including κ=0. At κ=0, DAPV(0) = V0 + (1−u)(V_served − V0), so an arrival still queued at T_max = 1.5× the last arrival is valued at the no-treatment outcome, regardless of how much value it would have realized had the simulation run longer. Section 4.4 asserts that re-simulating at multipliers 1.25–3 'changes no method comparison' but includes no results to support this. The Adult semi-synthetic raw served-value leader is PtO-mlp (1.49 vs. 1.30 for F), so a longer horizon that clears its queues should raise its DAPV(0); if it overtakes F at κ=0, the headline fails. Figure 3b's per-dataset crossover values (Adult 0.0212, ACTG 0.0422, non-nested 0.0013) already imply that the best two-stage method leads at κ=0 on those datasets; the combined-index lead at κ=0 is carried by the per-dataset normalization and the unweighted mean, not by per-dataset dominance. The robustness of the ranking to the horizon and to the unserved-arrival valuation is therefore the load-bearing condition for the paper's headline, and it is currently unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies learning assignment policies for sequentially arriving individuals under long-run capacity constraints on scarce treatments. The authors formulate a bilevel program in which outcome models are trained end-to-end by differentiating an inverse-propensity-weighted estimate of deployed policy value through the dual prices of the inner allocation problem. Two inner objectives are considered: the exact nonconvex softmax-weighted dual and a convex log-sum-exp surrogate. The paper proves that the convex surrogate is uniformly close to the exact objective (Proposition 1), that its optimum induces a softmax allocation satisfying capacity constraints in expectation with complementary slackness (Propositions 2 and 3), and derives an excess-value decomposition for the end-to-end estimator and a boundary-layer bound for the two-stage plug-in (Theorem 6). Empirically, the authors deploy all methods in a queueing simulation with Poisson resource replenishment and summarize performance with a deployment-adjusted value index DAPV(kappa). They claim that the two end-to-end variants hold the top two slots on the combined index at every delay cost kappa, including kappa=0, and that decision-blind baselines overshoot capacities and incur longer queues. On the largest dataset, end-to-end training achieves higher raw policy value than decision-blind baselines.","tokens_in":23522,"tokens_out":5581,"duration_ms":63744,"significance":"If the empirical claims are supported, the paper makes a meaningful contribution to decision-focused learning under resource constraints: it treats the dual prices as part of the learned policy class rather than as a post-hoc correction, and it gives a convex relaxation with an exact feasibility-in-expectation guarantee and a logarithmic-in-number-of-arms optimality gap. The theoretical core is genuinely derived: the entropy proof of the F-G bound is clean, the KKT analysis in Proposition 3 is sound, and Theorem 6(iii) explicitly identifies the boundary-layer mass as the channel through which prediction error affects deployed value. The paper is also commendably honest: it reports a pre-specified hypothesis that was refuted, acknowledges the un-discharged optimization-oracle assumption (A4), and provides code. The main weakness is that the headline empirical claim, especially the kappa=0 ranking, rests on evaluation choices whose robustness is asserted but not demonstrated.","major_comments":[{"comment":"The claim that end-to-end methods lead the combined index at kappa=0 is not robustly supported because DAPV(0) values arrivals still queued at the fixed horizon T_max = 1.5 times the last arrival at the no-treatment value V0, and the paper asserts, without presenting results, that re-simulating at multipliers 1.25-3 'changes no method comparison'. On the Adult semi-synthetic set, the raw served-value leader is PtO-mlp (1.49 vs. 1.30 for F), so a longer horizon that lets PtO-mlp's queues clear would raise its DAPV(0) and could overturn the kappa=0 ranking. Please add the re-simulation results or restrict the headline claim to the specific horizon used.","section":"§4.4 and §5.4, Eq. (9)"},{"comment":"The per-dataset crossover values (Adult 0.0212, ACTG 0.0422, non-nested 0.0013) show that the best two-stage method leads at kappa=0 on those datasets, so the combined-index lead at kappa=0 is carried by the per-dataset normalization (random=0, best at kappa=0 is 1) and by the unweighted mean rather than by per-dataset dominance. Since the abstract's 'including zero' is a central empirical claim, the paper should report per-dataset DAPV(0) values and justify why the unweighted normalized mean is the appropriate summary rather than a measure that obscures the underlying per-dataset comparisons.","section":"§5.4 and Figure 3b"},{"comment":"The temperature study reports that F's capacity excess grows from 1.1% at tau=0.01 to 28.2% at tau=1, which is why deployment routes through a buffered LP (0.92 capacity). This means the feasibility-in-expectation guarantee of Proposition 2 applies to the convex surrogate G, while the feasibility of F in deployment depends on an additional, a-priori-fixed buffer that is not part of the theoretical guarantee. The paper should state this distinction explicitly in the main text, since the abstract and introduction sometimes present feasibility as a property of the trained end-to-end policy without this caveat.","section":"§5.6 and Appendix G"}],"minor_comments":[{"comment":"The abstract says 'Across six datasets, the two end-to-end variants take the top slots on a deployment-adjusted value index at every delay cost,' but Figure 3 aggregates only five datasets; the mechanism dataset is excluded from the combined index. Please clarify which datasets the combined-index claim covers.","section":"Abstract and §5.4"},{"comment":"The definition of T_max as '1.5 times the last arrival time' should state whether the last arrival time is measured from the start of the simulation and whether the multiplier applies to the inter-arrival scale; a precise timing convention would make the re-simulation claim easier to verify.","section":"§4.4"},{"comment":"In the ACTG 175 discussion, the end-to-end methods deploy the capped arms at 0.24-0.28 against a cap of 0.30, i.e., below capacity. The text interprets this as respecting the constraint, but it should also comment on whether the under-utilization indicates a value loss relative to the oracle, especially since the methods tie on value within noise.","section":"§5.2"},{"comment":"The implicit-differentiation derivation assumes a locally constant active set and strict complementarity. The paper notes empirically that the Hessian was positive-definite at every iterate, but it would help to state how often the active set changed between outer iterations and whether any iterate came close to violating strict complementarity.","section":"§4.2 and Appendix F"},{"comment":"The combined-index curves are shown without error bars or confidence bands, even though each sweep uses 10 seeds. Adding per-seed variability or at least reporting whether the end-to-end lead is statistically significant at kappa=0 would strengthen the headline claim.","section":"Figure 3a"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for cs.LG and the theoretical contributions are solid. The main gap is empirical: the kappa=0 headline depends on the horizon choice and the unserved-arrival valuation, and the asserted robustness re-simulation is not shown. This is fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth taking seriously. The paper trains outcome models by differentiating an IPW estimate of deployed policy value through the dual prices of the capacity-constrained allocation problem, and compares against decision-blind predict-then-optimize. The math is solid: the entropy-based F–G bound (Prop 1), exact feasibility in expectation for the convex surrogate (Props 2–3), and the excess-value decomposition (Thm 6) all check out. The evaluation is more honest than most: six datasets, paired tests, a pre-specified hypothesis that was refuted and reported, negative ablations (SNIPS, buffer), and a capacity-matched neural baseline. The code link exists, though no commit hash is given, so I could not verify the runs.\n\nWhat is actually new is the bilevel formulation itself, plus the convex relaxation whose optimum is feasible in expectation by construction. The alternating baseline (Alt) is a good control for isolating the value of the implicit gradient, and the queueing simulation is the right way to measure deployment behavior. The paper also does not oversell where it loses: flexible regression keeps the raw-value lead where ground truth is measurable, and they say so.\n\nThe soft spots are in proportion. The biggest is the headline claim that the end-to-end variants top the combined index at every delay cost, including zero. At kappa=0, DAPV prices still-queued arrivals at the no-treatment value and the simulation stops at 1.5x the last arrival. The paper asserts that re-simulating at multipliers 1.25–3 changes no method comparison, but shows no results. Figure 3b even shows the best two-stage method leading at kappa=0 on three of five datasets, so the combined-index lead at kappa=0 is carried by the per-dataset normalization and the unweighted mean, not by per-dataset dominance. That is not a fatal flaw, but it means the headline, as stated, is over-broad. Second, Theorem 6 bounds the regret of a global IPW maximizer, not of any training iterate; the paper admits this in Remark 8, but it is worth remembering when citing the theory. Third, the code is unreviewed, which matters for a paper this empirical.\n\nWho is this for? Anyone working on decision-focused learning, off-policy policy learning, or allocation of scarce resources will get something from the formulation, the feasibility guarantee, and the evaluation design. It deserves a serious referee. My recommendation: send it to review, and ask the authors to either show the horizon robustness or soften the \"every delay cost including zero\" claim.","headline":"A genuinely new training recipe for capacity-constrained allocation, with clean math and an unusually honest evaluation; the headline empirical claim is real but rests on a horizon choice whose robustness is asserted, not shown.","tokens_in":24275,"tokens_out":1792,"would_cite":true,"duration_ms":20628,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training allocation policies end-to-end — pushing the gradient of an off-policy value estimate through the capacity-enforcing dual prices — beats decision-blind predict-then-optimize on deployed value across six datasets.","keywords":["end-to-end policy learning","dual prices","capacity constraints","inverse propensity weighting","bilevel optimization","predict-then-optimize","queueing simulation","scarce resource allocation"],"falsifier":"Extend the queueing horizon on the Adult semi-synthetic suite from the fixed multiplier of 1.5 times the last arrival to 3 and 6 times, recompute the deployment-adjusted value at $\\kappa=0$ for the raw served-value leader (PtO-mlp), and check whether it overtakes the end-to-end policies once its longer queue clears; the paper asserts robustness to multipliers 1.25–3 without displaying the sweep, so this is the direct test of the claim that end-to-end ranks first at every delay cost, including zero.","tokens_in":22984,"feed_emoji":"🏥","tokens_out":27445,"duration_ms":258360,"temperature":0.7,"pith_summary":"The paper tries to establish that an assignment policy for scarce resources — each arrival decided immediately, long-run usage of every resource capped — is best trained end-to-end, by differentiating an inverse-propensity-weighted (IPW) estimate of the deployed policy's value through the dual prices that enforce the caps. The standard predict-then-optimize pipeline fits one outcome model per arm by regression, blind to the decision boundary, and only then prices the resources; the cost of that blindness appears at deployment, where small prediction errors near a price boundary push usage over capacity and create long queues. Its bilevel method makes the prices part of the learner, so the same prices that keep the deployed policy feasible also carry the training signal that says which prediction errors matter. The theoretical core is a convex surrogate whose first-order optimality conditions are exactly the capacity constraints of the smoothed policy, guaranteeing feasibility in expectation (Proposition 2) at a cost of at most $\\tau\\log|T|$ relative to the exact objective (Proposition 1). Evaluated in a queueing simulation with resources replenished at their capacity rates, the end-to-end variants top the deployment-adjusted value index at every delay cost across six datasets, win raw value at the largest scale, and the paper reports where it does not win: where ground truth is measurable, flexible decision-blind regression remains the stronger pure predictor.","feed_headline":"Train through the prices, beat predict-then-optimize at deployment","feed_subtitle":"Gradients flow through the capacity prices, so policies hold caps, queue less, and win raw value at scale.","key_machinery":"The load-bearing object is the bilevel program of Equation 6, in which the inner problem computes the dual price vector $\\mu_\\theta$ that the capacity LP would pick if the current outcome scores were the truth, and the outer problem maximizes the inverse-propensity-weighted estimate of the deployed softmax policy's value, $\\hat{V}_{\\mathrm{IPW}}(\\theta,\\mu_\\theta)$. The gradient of that objective passes through the prices by implicit differentiation of the inner optimality conditions, which reduces the backward pass to one small dense linear solve on a KKT saddle-point system of dimension at most $2|T|$. The convex surrogate $G$ — the log-sum-exp, i.e. Nesterov's entropic smoothing, of the sample dual — is what makes the mechanism transparent: its derivative with respect to $\\mu_t$ is $b_t$ minus the average allocation to treatment $t$ under the smoothed policy, so the inner optimum is exactly the point where the policy is feasible in expectation, and the bound between $G$ and the exact objective $F$ is $\\tau\\log|T|$, linear in the smoothing temperature and logarithmic in the number of arms. The same entropy identity that proves the bound also prices the smoothing bias of the whole policy class.","core_discovery":"The paper's central claim is that making the dual prices part of the learner changes what gets deployed: training the outcome models by differentiating an inverse-propensity-weighted estimate of the deployed policy's value through the price map — the bilevel program in which the inner problem solves for prices exactly as deployment would — yields policies whose feasibility is a property of the trained weights, not a post-hoc correction. The mechanism is visible in the convex inner objective $G$: its derivative with respect to the price $\\mu_t$ is the residual $b_t - \\frac{1}{N}\\sum_i \\sigma_{t,i}$, the capacity minus the smoothed policy's average allocation, so at the inner optimum every priced treatment saturates its capacity and every unpriced one stays under it, in expectation, with prices positive exactly where the cap binds (Proposition 2). The paper proves $G$ is close to the exact nonconvex inner objective $F$, with the gap at most $\\tau \\log |T|$ (Proposition 1), where $\\tau$ is the softmax temperature and $|T|$ the number of arms. Deployed in a queueing system with resources replenished at their capacity rates, the two end-to-end variants take the top two slots on the deployment-adjusted value index at every delay cost, including zero; decision-blind baselines frequently violate the caps and wait several times longer. On the largest dataset, a seventy-thousand-patient hospital cohort, end-to-end training also wins held-out policy value outright, a margin that survives a capacity-matched neural baseline — and where the method does not win, the paper reports it: flexible decision-blind regression remains the stronger pure predictor where ground truth is measurable.","pith_inferences":["At $\\kappa=0$ the index prices still-queued arrivals at the no-treatment value, which systematically penalizes methods with longer queues; on the Adult semi-synthetic set the two-stage value leader is exactly such a method, so the \"first at every delay cost, including zero\" ranking is contingent on the fixed 1.5-times-last-arrival horizon rather than a steady-state ordering.","The $\\tau\\log|T|$ gap suggests a concrete recipe the paper does not run: train with the convex surrogate at moderate temperature, where feasibility in expectation and conditioning hold, then anneal $\\tau$ toward zero at deployment to approach the hard price rule while staying within the proved gap.","The excess-value bound in the appendix assumes the training loop reaches the global IPW maximiser (assumption A4, flagged by the authors as open), so the theoretical guarantee covers the statistically optimal policy in the priced-softmax class rather than any particular gradient-ascent run on the nonconvex outer problem.","The mechanism dataset isolates the regime where the pipelines separate — dense population near the decision boundary, a strictly positive price, and a harmful low-capacity arm — which yields a rough a-priori diagnostic: use the bilevel training when boundary-near prediction errors dominate and capacity binds, and flexible two-stage regression when the outcome surface is smooth and caps are generou"],"forward_implications":["Feasibility in expectation becomes a property of the trained policy rather than a post-hoc repair: with the convex inner objective, the trained allocation respects every capacity at the inner optimum, whatever the accuracy of the fitted outcome models.","Deployment behaviour changes: in the queueing simulations the end-to-end policies hold their caps and clear their queues in 12–16 periods, while two-stage baselines sit at or above the caps with waits of 31–56 periods on the same arrival streams.","The implicit gradient earns its cost: full end-to-end differentiation beats the alternating dual-refresh shortcut in 24 of 25 cells on ground-truth regime grids, and the dedicated KKT backward is 70 to 282 times faster than routing the inner problem through a generic differentiable-optimization layer.","Raw policy-value wins are possible at scale: on the 70,000-patient hospital cohort the end-to-end method wins held-out value outright (0.989 vs 0.944 for the best decision-blind method) and beats the capacity-matched neural baseline by +0.051.","Flexible decision-blind regression keeps the raw prediction lead wherever ground truth is measurable, at every training size; the paper reports this as a refuted pre-specified hypothesis, and on its mechanism dataset no baseline wins both axes at once — value and feasibility."],"supporting_citations":[{"why":"Supplies the predict-then-optimize baseline, the dual-price deployment rule the paper trains through (Equation 2), and the queueing-simulation evaluation framework it inherits.","marker":"(Tang et al. 2024)"},{"why":"Supplies the inverse-propensity-weighted estimator used as the end-to-end outer objective (Equation 4) together with its unbiasedness guarantee.","marker":"(Swaminathan and Joachims 2015a)"},{"why":"The entropic smoothing that defines the convex surrogate G in Equation 7, whose optimality conditions yield the feasibility-in-expectation guarantee.","marker":"(Nesterov 2005)"},{"why":"The differentiable-optimization-layer machinery used to differentiate through the inner price solve by implicit differentiation of its KKT conditions.","marker":"(Amos and Kolter 2017)"},{"why":"The generic differentiable convex layer that serves as the scaling baseline in the backward-pass speed comparison (70 to 282 times slower than the paper's dedicated KKT solve).","marker":"(Agrawal et al. 2019)"},{"why":"The empirical-welfare-maximisation uniform-deviation argument that the paper's excess-value decomposition (Theorem 6) extends to priced softmax policy classes.","marker":"(Kitagawa and Tetenov 2018)"},{"why":"Supplies the policy-learning-with-observational-data framework and the doubly-robust extension used when propensities must be estimated.","marker":"(Athey and Wager 2021)"},{"why":"The dual-guided shortcut whose alternating analogue (Alt) is the paper's baseline for measuring when the implicit gradient matters.","marker":"(Rodriguez-Diaz, Bansak, and Paulson 2025)"}],"fun_headline_variants":["End-to-end learning that sees the price of capacity","Differentiate through dual prices for feasible policies","When capacity binds, train through the prices","Gradients through prices keep policies within caps","Train outcome models through the price map"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline ranking's reach — including its claim of leading at zero delay cost — rests on the fixed simulation horizon of 1.5 times the last arrival, since on the Adult semi-synthetic set the raw served-value leader is a two-stage method (1.49 vs 1.30) whose zero-delay score would rise if a longer horizon let its queue clear.","fun_headline_variants_meta":{"raw":{"variants":["End-to-end learning that sees the price of capacity","Differentiate through dual prices for feasible policies","When capacity binds, train through the prices","Gradients through prices keep policies within caps","Train outcome models through the price map"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000301,"raw_usage":{"total_tokens":1849,"prompt_tokens":1172,"completion_tokens":677,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":788,"completion_tokens_details":{"reasoning_tokens":609}},"tokens_in":788,"tokens_out":677,"duration_ms":8448,"temperature":1.0,"reasoning_tokens":609,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:21:11.286053+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Extend the queueing horizon on the Adult semi-synthetic suite from the fixed multiplier of 1.5 times the last arrival to 3 and 6 times, recompute the deployment-adjusted value at $\\kappa=0$ for the raw served-value leader (PtO-mlp), and check whether it overtakes the end-to-end policies once its longer queue clears; the paper asserts robustness to multipliers 1.25–3 without displaying the sweep, so this is the direct test of the claim that end-to-end ranks first at every delay cost, including zero.","supporting_citations":[],"review_version":1}