{"id":"8596060f-311f-476b-b1d3-25c0d60cb391","arxiv_id":"2607.09820","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Learned predictive ambiguity sets output a finite scenario distribution and context-dependent Wasserstein radius for decision-focused DRO, matching most fixed-radius portfolio performance with a smaller average radius.","lead":"A neural model learns both a forecast distribution of future returns and a state-dependent Wasserstein radius that sets how much a robust optimizer should distrust that forecast. This can make predict-then-optimize pipelines less brittle by adapting robustness to market regimes instead of using one fixed safety margin.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Single chronological split on one 20-name universe is the load-bearing empirical weakness for the superiority and adaptivity claims.","rationale":"The dual reformulation (Eqs. 14–19) and the portfolio specialization are standard and internally consistent; the staged training objective is coherent. The only place the strongest claim is insecure is the empirical support for “substantially improves \to recovers most of Fixed-DRO with smaller radius and better regime adaptivity.” That support is a single split whose fragility the authors already acknowledge. The Reader correctly identified this as the weakest assumption and issued a CONDITIONAL verdict with moderate confidence. No deeper mathematical inconsistency or hidden assumption in the dual/layer construction appears; therefore the stress-test does not move the verdict. The concrete multi-fold check is exactly the experiment the paper itself says is needed.","tokens_in":11300,"tokens_out":509,"duration_ms":5260,"concrete_test":"Re-run the exact LPAS-W vs Fixed-DRO pipeline on at least three additional non-overlapping chronological folds (or three random 20-name S&P 500 universes) with the same hyper-parameter protocol; if the mean radius reduction and the ranking of Sharpe / CVaR95 reverse or become statistically insignificant on a majority of folds, the out-of-sample superiority claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (LPAS recovers most of Fixed-DRO performance with smaller radius and better regime adaptivity) rests almost entirely on Table 2 and the regime slices in Table 4. Both are produced from one chronological 1129/410/515 split of a single 20-name S&P 500 basket (2018–2026). No multi-seed, multi-fold, or multi-universe results are reported, and the paper itself flags this in §7. Because the Transformer scenario generator, radius network, and decision-aware validation score are all fit on the same path, the reported 26.28 % return / Sharpe 1.30 / radius 24.3 vs 35.4 edge could be an artifact of that particular market path and hyper-parameter selection rather than a stable property of learned radii. Ablations (Table 3) and volatility-bin diagnostics (Fig. 3–4) remain informative but cannot substitute for independent replications.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes learned predictive ambiguity sets (LPAS): a contextual deep model that outputs a finite nominal scenario distribution and a state-dependent Wasserstein radius (optionally an anisotropic ground metric), which define a contextual ambiguity set for a DRO decision layer. The radius is trained by combining pinball quantile calibration, size regularization, and downstream decision loss (Eqs. 24–28; Algorithm 1). The finite dual of the Wasserstein DRO layer is derived (Eqs. 14–15) and specialized to a long-only portfolio problem with Euclidean cost, yielding a closed-form robust objective (Eqs. 19–20). On a single chronological split of 20 S&P 500 names (2018–2026), LPAS-W reports 26.28% annualized return, Sharpe 1.30, final wealth 1.61, and a smaller average radius than a deep fixed-radius DRO baseline while remaining competitive on tail metrics and improving some regime slices (Tables 2–4).","tokens_in":11582,"tokens_out":1096,"duration_ms":8684,"significance":"If the empirical claims hold under stronger validation, the work would be a useful bridge between decision-focused learning and Wasserstein DRO: instead of a hand-tuned fixed radius centered on historical samples, both the nominal distribution and the radius become contextual and trainable. The dual specialization for the portfolio layer is standard and correctly applied, and the staged training objective (prediction + calibration + size + decision loss) is a concrete, implementable recipe. The main contribution is therefore architectural and empirical rather than theoretical; its value hinges on whether adaptive radii reliably reduce unnecessary conservatism while preserving robustness. The paper is transparent about the single-split limitation (§7), which is appropriate.","major_comments":[{"comment":"§6.1–6.3 and Tables 2–4: The central superiority and regime-adaptivity claims rest on a single chronological 1129/410/515 split of one 20-name S&P 500 universe, with no multi-seed, multi-fold, or multi-universe results. Because the Transformer scenario generator, radius network, and decision-aware validation score are all selected on this path, the reported edge (26.28% return / Sharpe 1.30 / radius 24.3 vs Fixed-DRO’s 35.4) and the high-ρ / drawdown regime gains in Table 4 could be path-specific. §7 already flags this; for the claims as stated, at least one additional rolling fold or multi-seed summary is load-bearing.","section":null},{"comment":"§4.5 and §6.5 / Fig. 4–5: After decision-aware tuning the empirical coverage of LPAS-W is 0.755 versus the nominal τ=0.9 target. The paper notes that size regularization and decision loss trade exact coverage for performance, and suggests conformal post-calibration if strict coverage is required. That is fine as a design choice, but the abstract and introduction still present the radius as “calibrated”; the manuscript should either report a post-calibrated variant or qualify the calibration claim so that readers do not over-read statistical coverage guarantees.","section":null},{"comment":"§3–4 and experiments: The optional anisotropic ground metric c_ψ (Eq. 7) is part of the stated framework and contributions but is never evaluated; all results use fixed Euclidean cost. Either evaluate a learned-metric ablation or narrow the contribution statement so that the unevaluated component is not presented as part of the empirical method.","section":null}],"minor_comments":[{"comment":"Table 1: Candidate numbering and the “Selected” column are hard to parse; a clearer description of the decision-aware validation score (weights on loss, tail, calibration, size) would help reproducibility.","section":null},{"comment":"Figures 1–2 legends truncate method names (“Deep Predictive Fixed-radius D”); expand for readability.","section":null},{"comment":"Eq. (26): The scenario prediction loss is written as a min-i reconstruction; clarify whether this is the sole L_pred used or whether NLL/energy-score variants were also tried.","section":null},{"comment":"§6.1: List of 20 tickers and feature construction details are useful; a short note on data source and any survivorship handling would strengthen the experimental section.","section":null},{"comment":"Notation: ρ_t, ρ_ϕ(z_t), and ρ_min appear in several places; a single consistent symbol table would reduce minor ambiguity.","section":null}],"recommendation":"major_revision","confidential_remarks":"The technical core (Wasserstein dual + staged radius learning) is sound and the portfolio specialization is standard. The main risk is over-claiming from a single-split finance experiment; if the authors add even modest multi-fold or multi-seed evidence and tighten the calibration/metric language, the paper becomes a reasonable fit for a methods-oriented ML or decision-focused learning venue. Without that, the empirical claims remain fragile."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: treat the ambiguity radius as a contextual output of the same model that produces the nominal scenarios, train it with pinball calibration plus size and decision loss, and feed the resulting Wasserstein ball into a dual DRO layer. That package is what they call LPAS.\n\nWhat is actually new is the joint object—finite predictive scenarios plus a state-dependent radius—and the staged training that makes it trainable. The dual (Eq. 14–15) and the portfolio specialization (Eq. 19–20) are textbook under affine loss and Euclidean cost; they are written cleanly and used correctly. The ablations are honest: kill calibration and the radius collapses and performance dies; size and decision fine-tuning matter less at the chosen scale. Against equal-weight, P2O, and historical WDRO the gains are large. Against a deep fixed-radius DRO that shares the same Transformer scenarios, LPAS keeps most of the return and Sharpe with a clearly smaller average radius and slightly better worst-month/CVaR numbers, and the radius tracks volatility the way you would want.\n\nThe soft spot is exactly the one the stress note flags, and the paper already owns it in §7: one chronological 1129/410/515 split on one 20-name S&P universe, no seeds, no rolling folds, no second universe. Decision-aware validation on that same path can favor the adaptive story. Coverage also undershoots the nominal τ after decision tuning. Those are real limits on how hard you can lean on the 26% / 1.30 / radius-24-vs-35 numbers; they are not reasons to dismiss the architecture.\n\nCitations are appropriate (Wang, Sun, Chenreddy–Delage, classic Wasserstein DRO). No circularity in the objective. Free parameters are the usual portfolio/ML set and are selected on validation.\n\nThis is for people who already run predict-then-optimize or Wasserstein DRO pipelines and want a practical adaptive-radius layer, especially in finance-style sequential decisions. It is not a theory paper. I would send it to referees: the method is coherent, the dual is checkable, the empirics are directionally informative, and the limitations are stated. Multi-fold/seed results and code would turn a conditional accept into a much stronger one. Worth engaging if you work in this stack; not urgent if you do not.","headline":"Clean, usable recipe for adaptive Wasserstein radii in decision-focused DRO; the math is standard and the portfolio gains look real, but the superiority claim still sits on one chronological path of one 20-name basket.","tokens_in":12175,"tokens_out":607,"would_cite":false,"duration_ms":12583,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A deep model can learn how large a robustness ball should be for each market state, recovering most of the gains of strong fixed-radius DRO while using a smaller average radius.","keywords":["distributionally robust optimization","Wasserstein ambiguity sets","decision-focused learning","portfolio optimization","predictive uncertainty","adaptive robustness","scenario generation"],"falsifier":"Re-run the same portfolio experiment over multiple rolling windows, random seeds, and a larger asset universe and find that the learned-radius model no longer matches fixed-radius DRO on return and Sharpe while keeping a smaller radius, or that the learned radius fails to rise with market volatility.","tokens_in":12155,"feed_emoji":"📈","tokens_out":914,"duration_ms":15303,"temperature":0.7,"pith_summary":"Predict-then-optimize systems treat a point forecast as reliable and can amplify small prediction errors into large decision mistakes. Classical distributionally robust optimization protects against that by optimizing against a whole ball of distributions, but the ball is usually centered on historical samples with one fixed radius, so it is often too conservative in calm regimes and still brittle under regime shift. This paper proposes learned predictive ambiguity sets: from context, a neural model outputs a finite nominal scenario distribution and a state-dependent Wasserstein radius that together define the ambiguity set for a robust decision layer. The radius is trained by conditional quantile calibration, size regularization, and realized decision loss so that robustness expands when forecasts are unreliable and contracts when they are trustworthy. On long-only portfolio optimization with 20 S&P 500 stocks from 2018–2026, the method substantially beats equal-weight, pure predict-then-optimize, and historical Wasserstein DRO, and nearly matches a strong deep fixed-radius baseline (26.28% annualized return, Sharpe 1.30, final wealth 1.61) while using a smaller average radius and adapting better across volatility and drawdown regimes.","feed_headline":"Learned DRO radii match fixed-radius portfolios with less conservatism","feed_subtitle":"On 20 S&P stocks, state-dependent Wasserstein balls nearly equal fixed-radius DRO while shrinking the average radius.","key_machinery":"Learned predictive ambiguity sets (LPAS): a contextual Wasserstein ball whose center is a neural finite nominal scenario distribution and whose radius is a state-dependent network; the dual of that ball supplies a tractable robust decision layer that is trained jointly with the radius.","core_discovery":"Learned predictive ambiguity sets—a contextual finite scenario distribution plus a state-dependent Wasserstein radius trained by quantile calibration, size regularization, and downstream decision loss—can make distributionally robust optimization adaptive rather than globally fixed. On the reported 20-asset portfolio task they recover most of the out-of-sample performance of a deep fixed-radius DRO baseline while using a smaller average radius, slightly better tail metrics, and stronger regime adaptivity.","pith_inferences":["The same radius-learning pattern should transfer to inventory, routing, and energy dispatch, where context likewise modulates forecast reliability.","When decision-aware tuning leaves empirical coverage below the nominal quantile, conformal post-calibration of the learned radius can restore strict risk control without fully undoing performance.","The optional anisotropic ground metric (proposed but not tested) could further cut conservatism by stretching the ball only in decision-sensitive directions."],"forward_implications":["Robust optimizers can shrink the ambiguity radius in calm regimes without giving up protection when forecasts are unreliable.","How large the robustness ball should be can be driven by decision quality, not only by predictive coverage.","Historical fixed-radius Wasserstein DRO underuses context and can be outperformed by predictive centers plus adaptive radii.","Most of the reported gains can be obtained by staged training: pretrain scenarios, calibrate the radius, then decision-focused fine-tuning."],"fun_headline_variants":["Learned DRO radii nearly match fixed-radius portfolios with less radius","Contextual Wasserstein balls cut conservatism while holding portfolio gains","Adaptive ambiguity sets recover fixed-radius DRO with smaller average radius","State-dependent DRO radii match deep baselines and improve regime adaptivity","Decision-trained Wasserstein radii equal fixed DRO at lower average size"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The out-of-sample superiority and regime-adaptivity claims rest on a single chronological train/validation/test split of one 20-stock S&P 500 universe, without multi-seed or multi-fold checks.","fun_headline_variants_meta":{"raw":{"variants":["Learned DRO radii nearly match fixed-radius portfolios with less radius","Contextual Wasserstein balls cut conservatism while holding portfolio gains","Adaptive ambiguity sets recover fixed-radius DRO with smaller average radius","State-dependent DRO radii match deep baselines and improve regime adaptivity","Decision-trained Wasserstein radii equal fixed DRO at lower average size"]},"model":"grok-4.5","effort":"low","cost_usd":0.003138,"raw_usage":{"total_tokens":1137,"prompt_tokens":825,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":31380000,"prompt_tokens_details":{"text_tokens":825,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":242,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":825,"tokens_out":70,"duration_ms":2776,"temperature":1.0,"reasoning_tokens":242,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T15:18:17.549470+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same portfolio experiment over multiple rolling windows, random seeds, and a larger asset universe and find that the learned-radius model no longer matches fixed-radius DRO on return and Sharpe while keeping a smaller radius, or that the learned radius fails to rise with market volatility.","supporting_citations":[],"review_version":1}