{"id":"e2f5c0e1-b619-4fd3-b1ad-be6233579329","arxiv_id":"1908.03907","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":15,"one_line_summary":"A dynamic programming framework for VaR-constrained portfolio choice under heavy-tailed returns, but the numerical implementation fits prices instead of returns and omits the actual optimization.","lead":"This paper uses Markov Decision Process methods to decide when to hold a risky stock versus cash while keeping Value at Risk within a limit, assuming heavy-tailed returns. The numerical analysis fits the distributions to price levels rather than returns, and the claimed optimal strategies are not actually derived from a maximization.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed optimality is not established: the Bellman maximization is never performed; the recursive policy (6) is derived only from solving a VaR inequality, with no objective function or action-space optimization, so it does not maximize expected utility.","rationale":"The reader's verdict of REJECT is appropriate. The reader's weakest_assumption centered on the price/return distribution mismatch, but also mentioned 'absence of any actual maximization.' I identify the missing maximization as the single most load-bearing concern because it directly defeats the central claim (optimal expected utility), independent of the empirical fitting errors. Even if the Pareto fit were corrected to returns, the optimization step would still be absent: equations (6), (8), (10), (12), (13) are all derived from the VaR inequality alone, not from maximizing the Bellman objective. The paper's action-space description is also inconsistent (continuous π∈[0,M] vs. two alternatives buy/sell), and no selection rule is given for choosing among the many policies satisfying the inequalities. The proposed test (one-step Bellman maximization vs. inequality boundary) would settle whether (6) actually characterizes an optimum; given the derivation, it cannot, but the test provides a concrete falsification path. I do not object to the topic (VaR-constrained portfolio choice is legitimate), and the authors reference relevant literature; the concern is internal to the argument. No machine-checked proofs, code, or reproducible data are provided, so there is no independent support to offset the gap. Verdict remains REJECT as the reader concluded; no adjustment needed.","tokens_in":16399,"tokens_out":8226,"duration_ms":77838,"concrete_test":"Settle the concern by implementing the one-step Bellman maximization with the paper's own definitions: for a fixed terminal policy π_t, compute the expected reward E[W_{t-1}] = E[π_{t-1}S_{t-1} + (1-π_{t-1})r - (π_t S_t + (1-π_t)r)] under the fitted return distribution, restricted to policies satisfying the VaR constraint (4). Maximize this expectation over π_{t-1} ∈ [0,1] numerically (e.g., grid or gradient) and compare the argmax to the boundary of inequality (6) for several π_t values. If the maximizer is not the boundary value of (6), or if the feasible set of (6) contains points with higher expected reward, then (6) is not the dynamic-programming optimum. Also verify whether (6) is even binding for typical π_t values (e.g., for π_t=0.5 the upper bound exceeds 1, so the inequality is vacuous).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that dynamic programming yields strategies maximizing expected utility under a VaR constraint. Section 2 states the problem as 'maximize {W_{t-1}(s_{t-1},a) + sum_j p_{t-1}(j|s_{t-1},a) u^π_t}' subject to P(L_{t-1}>Q_{0.05}|L_t>Q_{0.05}) ≥ 0.95. However, in Section 3.1 the recursive relation (6) is obtained solely by solving inequality (5) for π_{t-1}; no derivative is taken and no supremum over actions in the Bellman equation (2) is evaluated. The result π_{t-1} ≤ π_t(19.1977π_t+1) defines a feasible set, not an optimal policy. The statement that 'for each π_{t-1} obtained for a period the maximum value of W_{t-1} is calculated' only evaluates W at a feasible point; it does not maximize the expected reward. Consequently, the 'optimal strategies' and 'optimal value function' asserted in the abstract and conclusions are not derived from any objective-function maximization. This is an internal gap, not a disagreement with consensus: even if every distributional fit were correct, the Bellman optimality step is absent. A secondary empirical error (fitting Pareto to prices, λ=85.34, so Q_{0.05}=5e-6λ lies below the support x≥λ and the integrals in (5) are applied outside the distribution's valid range) compounds the problem, but the missing maximization alone invalidates the optimality claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a discrete-time portfolio problem with one risky asset and a risk-free asset, where the investor maximizes expected utility (elsewhere, the 'median return') subject to a Value-at-Risk constraint, assuming heavy-tailed return distributions. The authors formulate the problem as a finite-horizon Markov decision process, derive backward recursions for the portfolio weight under Pareto, Weibull, and Inverse Gaussian return distributions, and also use a kernel density estimator for the unknown-distribution case. They calibrate the distributions to daily stock price data of Entergy Corporation, perform out-of-sample Kolmogorov-Smirnov tests, and numerically plot the implied optimal strategy, value function, and portfolio wealth with and without transaction costs. The central claims are that the recursions provide optimal strategies and that the value function is maximized at each time step.","tokens_in":16920,"tokens_out":5155,"duration_ms":53479,"significance":"If the claims were correct, the paper would offer a simple discrete-time rule for when to buy or liquidate a single risky asset under heavy-tailed returns while controlling VaR, with an explicit treatment of transaction costs. The empirical comparison of Pareto, Weibull, Inverse Gaussian, mixture-normal, and variance-gamma fits is a useful exercise, and the paper is transparent about its data source and KS-test methodology. However, the central optimality result is not established: the Bellman maximization in Eq. (2) is never carried out, the Pareto calibration is performed on stock prices rather than returns, the VaR quantile lies outside the fitted distribution's support, and the backward recursions are anchored on arbitrary terminal values. These are internal logical and empirical errors, not a matter of disagreement with current consensus. The paper's framework as presented does not support its stated conclusions, and the numerical outputs are not reliable.","major_comments":[{"comment":"The Bellman optimality step is never performed. Equation (2) defines the optimal value function as a supremum over actions in the action space, but the recursive relation (6) is obtained solely by solving the VaR inequality (5) for π_{t-1} given π_t. No derivative is taken, no supremum or maximum over alternative actions is evaluated, and the reward W_{t-1}=L_{t-1}-L_t is only evaluated at a feasible point, not maximized. Consequently, the text's claim that these are 'optimal strategies' and that the value function is maximized is unsupported by the derivations. This is an internal gap: even if every distributional fit were correct, the paper would not have solved the optimization problem it states.","section":"Section 2 and Section 3.1, Eqs. (2), (5), (6)"},{"comment":"The Pareto distribution is fitted to stock prices, not returns. The paper defines returns as (Price_t - Price_{t-1})/Price_{t-1}, but the estimated scale parameter λ = 85.34364 equals the minimum daily closing price of Entergy stock, and the estimated tail index α ≈ 10346 is implausible for return data. The quantile Q0.05 = 5×10^-6 λ ≈ 4.3×10^-4 lies far below the support of the fitted Pareto distribution (x ≥ λ = 85.34). As a result, the integral limits K_t in Eq. (5) are below the distribution's support, making the numerator and denominator of the constraint meaningless and the derived recursion (6) invalid. The same prices-versus-returns confusion affects the out-of-sample KS test, which validates the fit on prices rather than on the returns the model claims to address.","section":"Section 3.1, Table 1 and Eq. (5)"},{"comment":"The density used in Eq. (5) is not the density of the Pareto distribution defined earlier. The text defines the survival function as F(x) = (λ/x)^α for x ≥ λ, whose density is α λ^α / x^{α+1} on x ≥ λ. Equation (5), however, integrates α λ^α / (x+λ)^{α+1}, a shifted form with support x ≥ 0. This is an internal inconsistency, and the numerical value of the integral in Eq. (5) is therefore not a consequence of the stated Pareto model. The recursion (6) cannot be justified from the model as written.","section":"Section 3.1, Eq. (5)"},{"comment":"The backward recursions are anchored on arbitrary terminal strategy values that are assumed without justification: 0.00001 for Weibull, 0.1 for Inverse Gaussian, and 0.001 for the kernel-density case. These terminal values are not derived from the optimization problem or from any economic condition, yet they determine the entire path of the 'optimal' strategy through Eqs. (10), (13), and the KDE recursion. The dependence of the reported results on these externally imposed endpoints is not discussed, and it undermines the claim that the computed policies solve the stated maximization problem.","section":"Sections 3.2, 3.3, and 4.1"}],"minor_comments":[{"comment":"The abstract says the investor maximizes 'expected utility', while Section 1.1 says the investor seeks to 'maximize the median return of the portfolio'. These are different objectives, and the paper does not clarify which one is actually used in the MDP formulation.","section":"Abstract and Section 1.1"},{"comment":"The sentence 'Solving (8) and putting the values of λ and r = 0.00014...' should refer to Eq. (9), not Eq. (8), which is the transaction-cost recursion of the previous subsection.","section":"Section 3.2, after Eq. (9)"},{"comment":"Figures 7-9 are described in the text as showing the Inverse Gaussian results, but the captions in the list of figures state 'considering Weibull distribution'. The captions should be corrected.","section":"Figure captions 7-9"},{"comment":"The KS critical values reported in Table 6 are inconsistent: for a sample of size 50 the standard 5% critical value is about 0.192, but the table reports 0.12323 for the mixture normal and 0.156 for the variance gamma. The origin of these different critical values is not explained.","section":"Table 6"},{"comment":"In Eq. (15), the term F(0) appears without a clear definition of F; if F is the kernel CDF, then F(0) may not be meaningful for returns data that can take negative values. The expression should be stated more carefully.","section":"Section 4.1, Eq. (15)"}],"recommendation":"reject","confidential_remarks":"The manuscript's central claims are not supported by the derivations or the empirical analysis. The absence of any actual Bellman maximization and the fitting of the Pareto distribution to prices rather than returns are fundamental errors that cannot be repaired by a local revision; the proposed recursions would need to be re-derived and the empirical work redone from scratch. I do not see a path to acceptance within the scope of the current manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This one can go back. The main claim—that the recursions give optimal VaR-constrained strategies under heavy tails—does not hold up, and the numerical illustration is built on a fit to the wrong variable.\n\nWhat the paper does well is narrow but real: it sets up a finite-horizon MDP with a single risky asset, and it compares five distributions (Pareto, Weibull, Inverse Gaussian, mixture normal, variance gamma) using out-of-sample KS tests. That comparison is routine but honestly executed.\n\nThe soft spots are large. First, the 'optimal' policy in Sections 3.1–3.3 is obtained by solving the VaR inequality for π_{t-1} (e.g., equation (6)). No supremum over actions is ever evaluated in the Bellman equation (2). The authors say they calculate the maximum W_{t-1} for each π_{t-1}, but that is a feasible value, not an optimum. There is no objective-function maximization anywhere. That missing step invalidates the optimality claim on its own.\n\nSecond, the Pareto fit is confounded. The estimated scale λ = 85.34364 is a stock price, not a return. The 5% quantile Q = 5×10⁻⁶λ ≈ 0.00043 lies far below the Pareto support x ≥ 85.34, so the integrals in (5) are evaluated outside the distribution's valid range. The density in (5) also uses (x+λ) in the denominator instead of x for the Type I Pareto defined earlier, so the algebra is internally inconsistent. The Weibull and inverse Gaussian fits are less obviously wrong, but the same concern about applying the quantile threshold to returns while using price-based notation runs through the paper.\n\nThird, the paper is incomplete: most figures and tables are placeholders, and no code or data are provided. The terminal strategy values are assumed arbitrarily (0.00001, 0.1, 0.001) without sensitivity analysis.\n\nThe literature review is adequate but not deep; the novelty over Lwin et al. and Chow–Ghavamzadeh is incremental at best. The paper does not deserve a serious referee: the load-bearing pieces are absent or incorrect. It could serve as a rough draft for a real study, but as submitted, it is not a contribution.\n\nRecommendation: desk reject.","headline":"The paper's central optimality claim collapses because no maximization is ever performed, and the Pareto fit is to prices, not returns.","tokens_in":17354,"tokens_out":3235,"would_cite":false,"duration_ms":41138,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G10","91G80"],"pacs":[],"model":"deepseek-v4-flash","headline":"For one risky asset plus cash, the paper derives a discrete-time recursion for the optimal portfolio weight from a 5% Value-at-Risk constraint, covering Pareto, Weibull, Inverse Gaussian, and kernel-estimated returns.","keywords":["Portfolio Optimization","Markov Decision Process","Value at Risk","Heavy-tailed returns","Dynamic programming","Kernel density estimation","Parametric distribution","Non-parametric distribution"],"falsifier":"Evaluate the integral in equation (5) with the reported Pareto estimates $\\lambda=85.34364$ and $\\alpha=10346.37374$ and the quantile $Q=5\\times 10^{-6}\\lambda\\approx 0.00043$: the fitted Pareto density is zero below $\\lambda$, so the integration limits in (5) fall outside the distribution's support and the survival-function calculation cannot produce the printed recursion. A direct empirical check would simulate returns from the fitted distribution, apply the recursion, and count how often the conditional probability in (4) stays above 0.95.","tokens_in":16214,"feed_emoji":"📉","tokens_out":16208,"duration_ms":149260,"temperature":0.7,"pith_summary":"This paper tries to establish that a risk-managed portfolio choice remains tractable when stock returns have heavy tails, where moment-based methods and Brownian-motion tools are unavailable. The authors model the investor as facing a finite-horizon Markov decision problem with one risky asset and a risk-free bank account, and they impose a Value-at-Risk constraint: the probability that wealth stays above the 5% loss quantile, given that it was above it one period earlier, must be at least 95%. From this constraint they derive a recursive inequality for the optimal fraction of wealth held in the risky asset, worked out explicitly for Pareto, Weibull, and Inverse Gaussian returns and, when the distribution is unknown, for a kernel density estimate of returns. If the recursions are correct, they give a direct rule for when to build up or liquidate the risky asset while keeping tail risk under a fixed quantile, with and without transaction costs.","feed_headline":"One VaR inequality tells when to build or liquidate a risky asset","feed_subtitle":"The recursive rule comes from maximizing expected utility under a 5% loss quantile for heavy-tailed returns.","key_machinery":"The engine of the argument is the one-period conditional VaR survival constraint, $$P(L_{t-1}>Q_{0.05}\\mid L_t>Q_{0.05})\\ge 0.95,$$ combined with the finite-horizon Markov decision process policy-evaluation identity $u^\\pi_{t-1}=(L_{t-1}-L_t)+E(u^\\pi_t)$. Substituting the wealth equation and the survival function of the fitted return distribution turns the probability constraint into an integral inequality in the portfolio weights $\\pi_{t-1}$ and $\\pi_t$; solving that inequality at each time step yields the recursive optimal holding. In the unknown-distribution case the same inequality is evaluated with a kernel density estimate and integrated numerically. This constraint-to-recursion conversion is what carries the whole paper: it turns a tail-risk target into an explicit rebalancing rule.","core_discovery":"The central claim is that the Value-at-Risk constraint itself determines the optimal policy, and that dynamic programming then evaluates it. Starting from the wealth equation $L_t=\\pi_t S_t+(1-\\pi_t)r$, the authors require $P(L_{t-1}>Q_{0.05}\\mid L_t>Q_{0.05})\\ge 0.95$. For Pareto-distributed returns with scale $\\lambda$ and tail index $\\alpha$, substituting the survival function turns this into an integral inequality that solves to the recursion $\\pi_{t-1}\\le \\pi_t(19.1977\\pi_t+1)$ at a daily rate $r=0.00014$; the Weibull case gives $\\pi_{t-1}\\ge \\pi_t(1-13.08\\pi_t)$ and the Inverse Gaussian case an implicit relation solved numerically. With a 10% transaction cost the recursions become second-order. The value function is then obtained from the finite-horizon dynamic programming identity $u^\\pi_{t-1}=(L_{t-1}-L_t)+E(u^\\pi_t)$. In the non-parametric case the same conditional-probability constraint is evaluated by numerical integration against a kernel density estimate of the return distribution. The paper concludes that the resulting strategy tells the investor when to accumulate and when to liquidate the risky asset, and that transaction costs make the strategy nearly constant until the terminal period.","pith_inferences":["A natural extension would replace the single 95% quantile condition with a family of quantiles or with expected shortfall, turning the one-path recursion into a full risk-return frontier.","For multiple risky assets, the second moment may fail to exist, so a multi-asset version would need a rank-based or copula dependence measure in place of covariance; the paper flags this as future work, and the constraint-template approach suggests how to build it.","In the kernel-density case, the value function and wealth fluctuate widely; smoothing or bandwidth regularization would likely be needed before using the recursion as a live trading rule.","A holdout test could compare the realised frequency of wealth falling below the 5% quantile against the promised 5%: if the recursion is right, the violation rate should match the target over a long out-of-sample period."],"forward_implications":["Under a fitted Pareto model, the next optimal holding is obtained directly from the current one by $\\pi_{t-1}\\le \\pi_t(19.1977\\pi_t+1)$, so the strategy is computable by hand, without Monte Carlo simulation.","When the return distribution is unknown, the same VaR constraint solved by numerical integration against a kernel density estimate gives an implementable rule from historical data alone.","With transaction costs, the recursion becomes second-order and the optimal policy is nearly flat until the horizon, predicting infrequent rebalancing followed by a sharp build-up or liquidation near the terminal date.","The direction of the interest-rate effect matches standard intuition: lower bank rates push the optimal holding in the risky asset upward.","Out-of-sample KS comparisons suggest the three two-parameter heavy-tailed candidates fit the data almost as well as the six-parameter mixture normal, so the approach controls tail risk at low estimation cost."],"supporting_citations":[{"why":"Supplies the finite-horizon Markov decision process policy-evaluation recursion used for the value function.","marker":"Puterman,2005"},{"why":"Establishes that heavy-tailed returns can fail to have finite moments, motivating a quantile-based risk constraint.","marker":"Campbell et al., 1997"},{"why":"Provides the kernel density estimator used in the unknown-distribution numerical case.","marker":"Markovich (2007)"},{"why":"Supplies the extreme-value distribution framework behind the Pareto and Weibull candidates.","marker":"Embrechts et. al., 1997"},{"why":"Provides the daily closing-price dataset for Entergy Corporation used in all numerical calibrations and plots.","marker":"Quantopian, 2018"}],"fun_headline_variants":["VaR rule alone decides when to buy or sell risky assets","Heavy-tailed returns: a single VaR condition sets policy","Optimal portfolio strategy from one VaR inequality","VaR constraint dictates accumulation and liquidation timing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fitted heavy-tailed distribution is the true law of the daily returns, with the 5% VaR quantile used in the constraint lying inside the support of that distribution; if the quantile falls outside the support, the integral inequalities and recursions derived from them do not follow.","fun_headline_variants_meta":{"raw":{"variants":["VaR rule alone decides when to buy or sell risky assets","Heavy-tailed returns: a single VaR condition sets policy","Optimal portfolio strategy from one VaR inequality","VaR constraint dictates accumulation and liquidation timing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000291,"raw_usage":{"total_tokens":1681,"prompt_tokens":908,"completion_tokens":773,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":709}},"tokens_in":524,"tokens_out":773,"duration_ms":7784,"temperature":1.0,"reasoning_tokens":709,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:59:10.069038+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the integral in equation (5) with the reported Pareto estimates $\\lambda=85.34364$ and $\\alpha=10346.37374$ and the quantile $Q=5\\times 10^{-6}\\lambda\\approx 0.00043$: the fitted Pareto density is zero below $\\lambda$, so the integration limits in (5) fall outside the distribution's support and the survival-function calculation cannot produce the printed recursion. A direct empirical check would simulate returns from the fitted distribution, apply the recursion, and count how often the conditional probability in (4) stays above 0.95.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the kernel density estimator used in the unknown-distribution numerical case."},{"cited_title":"and Mikosch T","cited_arxiv_id":null,"evidence_quote":"Supplies the extreme-value distribution framework behind the Pareto and Weibull candidates."}],"review_version":1}