{"id":"df246975-218b-4443-8a73-bac4ad7c0998","arxiv_id":"1908.08684","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A nonlinear optimization model that minimizes portfolio drawdown and, in backtests without transaction costs, outperforms three major equity indices on average.","lead":"The paper builds portfolios that are chosen to have the smallest dips in value over recent weeks, then tests them on three major stock indices over 2010 to 2016. On average these portfolios beat the plain index in return and drawdown, but the test ignores trading costs, which would be large for this strategy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline dominance claim rests on the §4.1 assumption of zero transaction costs; at 10-day rebalancing across up to 500 assets, realistic costs can plausibly erase the modest gross excess returns in Table 2.","rationale":"I read the paper as a methodological contribution plus an empirical demonstration. The optimisation model and the partial linearisation argument in Section 3 are mathematically sound; replacing Mt = max Pτ by Mt ≥ Pτ is valid for the stated minimisation objectives, and the authors correctly note the objective direction is what makes the replacement work. The computational study is transparent about solvers, time limits, data curation against survivor bias, and the fact that not all rebalances are proven optimal. Those features support the paper's internal validity. However, the headline conclusion is an out-of-sample performance claim, and its weakest link is not the model but the evaluation protocol: the reported dominance is gross of transaction costs. With rebalancing every 10 days, even moderately realistic costs can consume the 1–2.5% annualised gross excess return shown in the average row of Table 2. The paper's own Section 4.1 flags the zero-cost assumption, but Section 5 repeats the dominance claim without it. The reader's conditional verdict is appropriate: the mathematical contribution can stand, but the empirical claim should be re-reported net of costs, or at least accompanied by turnover statistics and a cost break-even analysis. I do not see a more basic flaw that would require changing the verdict to accept or reject; the concern is exactly the one identified by the reader.","tokens_in":18537,"tokens_out":10379,"duration_ms":117463,"concrete_test":"Re-run the MINAVG 0.1 and MINMAX 0.1 cases (EURO STOXX 50, FTSE 100, S&P 500) using the same 30-day in-sample rule and 10-day rebalancing, but with proportional transaction costs of 10 bps per unit traded on both buy and sell, enforced through equations (6)–(8) with γ set to the corresponding limit. Compare the net-of-cost out-of-sample average daily return, maximum drawdown and average drawdown against the index row of Table 2; if the average net excess return becomes negative or the drawdown metrics cease to dominate, the headline claim is overturned. If the code is unavailable, the minimal alternative is to report average one-way turnover per rebalance for each Table 2 row and compute net returns under 5, 10 and 20 bps costs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim in Section 5 — that minimal drawdown portfolios dominate the indices out-of-sample on return, Sharpe ratio, maximum drawdown and average drawdown — is generated under the explicit assumption in Section 4.1 that 'transaction costs are zero.' This is not a minor bookkeeping point. The strategy rebalances every 10 trading days (Section 4.1) over up to 500 assets, so roughly 25 round trips per year. The average row of Table 2 shows gross excess daily returns over the index of 0.000041–0.000100, i.e. about 1.0–2.5% annualised. A one-way cost of 10 bps on a portfolio that turns over only 50% per rebalance costs about 2.5% per year; full turnover at 10 bps costs 5% per year. Since the reported optimisations did not penalise turnover (costs were set to zero), nothing in the model prevents such turnover. The conclusion states the dominance result without the gross-of-costs qualification. Therefore the most load-bearing assumption is the one the authors themselves flag, and whether it survives depends on turnover and cost levels not reported in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a nonlinear programming formulation for constructing portfolios that minimise average drawdown, maximum drawdown, or a weighted combination of the two, using a rolling in-sample window and periodic rebalancing. The optimisation model includes cash inflows/outflows, transaction-cost variables, short-selling constraints, and a partial linearisation that replaces the maximum-value equality with equivalent linear inequalities. Computational results are reported for EURO STOXX 50, FTSE 100, and S&P 500 daily data over 2010–2016, with a 30-day in-sample window, 20-day drawdown lookback, and rebalancing every 10 days. The paper claims that, on average, the resulting minimal drawdown portfolios dominate the respective market indices out-of-sample on return, Sharpe ratio, maximum drawdown, and average drawdown over roughly 1800 trading days.","tokens_in":18814,"tokens_out":4259,"duration_ms":45126,"significance":"The mathematical contribution is sound and the paper is clearly written: the partial linearisation argument in Section 3.3 is correct, and the use of a global nonlinear solver (SCIP) with reported optimality gaps is a strength. The out-of-sample evaluation is a genuine predictive test rather than a fitted in-sample exercise, because the parameters T=30, D=20, and the 10-day rebalancing interval are fixed by hand and not tuned to the out-of-sample period. If the empirical dominance claim were robust, the paper would be a useful OR contribution to drawdown-based portfolio construction. However, the evidence for the central claim is thin: it rests on three single-path histories over one historical period, is computed under an explicit zero-transaction-cost assumption, and is reported without any statistical error measure. These gaps currently limit the strength of the conclusions that can be drawn.","major_comments":[{"comment":"The headline out-of-sample dominance is computed under the explicit assumption that transaction costs are zero (stated in Section 4.1). With rebalancing every 10 trading days and up to 500 assets, turnover can be substantial, and the model does not penalise turnover because costs are set to zero. The average row of Table 2 shows gross excess daily returns over the index of only about 0.00004–0.00010 (roughly 1–2.5% annualised). A one-way cost of 10 basis points on a portfolio that turns over only 50% per rebalance would cost about 2.5% per year, which is of the same order as the reported excess return. The conclusion in Section 5 states the dominance result without repeating this gross-of-costs qualification. The authors should report turnover per rebalance, perform a sensitivity analysis with realistic cost levels (e.g., 5–20 basis points one-way), or include a turnover penalty in the optimisation; without this, the central empirical claim is not robust.","section":"Section 5; Table 2"},{"comment":"The empirical claim of out-of-sample dominance is based on three indices and one historical period (2010–2016). Each instance provides a single out-of-sample path, and the average row in Table 2 is the mean over only these three paths. No standard errors, confidence intervals, bootstrap resamples, or tests across subperiods are reported. Given the small gross excess returns, sampling variation alone could overturn the dominance conclusion. The authors should at least report the distribution of out-of-sample returns across rebalances, or validate the strategy over additional periods or asset universes, before claiming that the proposed portfolios 'dominated the market indices' on average.","section":"Section 5; Table 2"},{"comment":"The central computational results depend on several hand-chosen parameters: T=30, D=20, the 10-day rebalancing interval, the proportion limits δi, and the shorting limits. No sensitivity analysis is provided, so the reader cannot judge whether the reported dominance is robust to reasonable changes in these settings. For example, the 10-day rebalancing interval interacts directly with the transaction-cost issue raised above, and the choice of T=30 means the in-sample drawdown is computed from only 30 observations. A sensitivity study over T, D, and rebalancing frequency would substantially strengthen the empirical claims.","section":"Section 4.1; Section 4.2"}],"minor_comments":[{"comment":"The formulation assumes Mt > 0 whenever Pt > 0. If Pt = 0 and all Pτ in the lookback window are also zero, drawdown in equation (2) is undefined (0/0). With real price data this does not arise, but a short comment on the assumption would avoid ambiguity.","section":"Section 3.2, equations (1)–(2)"},{"comment":"The validity of replacing the equality definition of dt by the inequality dt ≥ 100(Mt−Pt)/Mt is stated to follow by an argument 'very similar' to that for equation (12). For the MINMAX objective, the reasoning is slightly different because dt does not appear directly in the objective; the authors could spell out why alternative optimal solutions make the inequality tight.","section":"Section 3.5, equation (23)"},{"comment":"Table 2 is dense and somewhat hard to read; the column grouping into in-sample and out-of-sample blocks would benefit from a separator or subheadings, and the 'Proportion limit' column could be labelled more explicitly (e.g., 'δ limit' for long-only and 'δL/δS limit' for shorting cases).","section":"Table 2"},{"comment":"The data are described as 'manually curated' to avoid survivor bias, but no details are given on how the index compositions were obtained or verified. A brief description of the curation procedure would improve reproducibility.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's mathematical core is sound and the topic fits the journal's scope, but the empirical headline is currently overstated relative to the evidence. The zero-transaction-cost assumption combined with the modest gross excess returns is the most serious issue, and the lack of any statistical error measure on the three-path average further weakens the claim. I believe these points are addressable within a revision, so major revision rather than rejection seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. First, the modeling contribution is solid: a rolling-window drawdown-minimization NLP with a clear partial linearization, solved to proven global optimality with SCIP on up to 500 assets. Second, the out-of-sample \"dominance\" claim is real but gross of costs, and realistic transaction costs can plausibly erase the reported edge. Read Table 2 as a stress test of the method, not as a deployable alpha result.\n\nThe formulation is the usual epigraph trick applied carefully, and the authors are appropriately careful about when the linearization is valid. The proof by contradiction is correct, and they explicitly note that the replacement of the max constraint by inequalities would fail if the objective were to maximize drawdown. The computational setup is more careful than much of this literature: manually curated index membership avoids survivor bias, T=30 and D=20 are fixed rather than tuned to the out-of-sample period, and they report the percentage of rebalances solved to proven global optimality. The shorting extension is sensible, and the honest reporting that shorting helps in-sample but not out-of-sample is a point in the paper's favor.\n\nThe main limitation is the headline claim. Section 4.1 states that transaction costs are assumed to be zero, but Section 5 repeats the dominance conclusion without that qualification. With rebalancing every 10 days across hundreds of assets, turnover is unconstrained and likely substantial. Gross excess returns average roughly 1 to 2.5 percent annualized; a one-way cost of 10 basis points on 50 percent turnover per rebalance costs about 2.5 percent per year, so the edge can disappear. Also, there is only one historical period (2010-2016) and three indices, with no standard errors, bootstrap confidence intervals, or subperiod checks. The word \"dominate\" is loose: on the S&P 500, the MINAVG 0.1 portfolio has a lower average daily return than the index (0.000357 vs. 0.000387), and the averaged row hides per-instance losses. The absence of code and curated data makes independent verification harder, though the equations are complete enough to reimplement.\n\nWho is this for? OR researchers working on portfolio optimization and practitioners who want a transparent drawdown-control baseline. It deserves peer review: the formulation is competent, the empirical work is honest about its main friction, and the central limitation is stated rather than hidden. My recommendation: send it out, but require a transaction-cost sensitivity analysis or a reframing of the conclusion to gross-of-cost, per-instance results.","headline":"A clean, honest formulation for rolling drawdown minimization whose headline out-of-sample edge is gross of costs and probably fragile to realistic frictions.","tokens_in":19295,"tokens_out":2530,"would_cite":true,"duration_ms":27811,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G10","90C30"],"pacs":[],"model":"deepseek-v4-flash","headline":"A portfolio built to minimise drawdown can beat its market index on return, Sharpe ratio, maximum drawdown, and average drawdown out of sample.","keywords":["portfolio drawdown","nonlinear programming","portfolio optimisation","index out-performance","rebalancing","Sharpe ratio","partial linearisation","long-short portfolios"],"falsifier":"Recompute the same rebalancing strategy while charging a conservative transaction cost, for example 10 to 50 basis points of traded value, and compare net out-of-sample average daily return and Sharpe ratio with the index; if net performance no longer beats the index, the headline dominance result fails.","tokens_in":18353,"feed_emoji":"📉","tokens_out":7817,"duration_ms":74611,"temperature":0.7,"pith_summary":"The paper claims that a portfolio constructed to minimise drawdown—the percentage of value lost from the best recent peak—can beat the market index on four measures at once: return, Sharpe ratio, maximum drawdown, and average drawdown, out of sample. The strategy uses only the previous 30 days of prices, rebalances every 10 trading days, and was tested on three major equity indices over roughly seven years of daily data. The claim matters because drawdown is a path-dependent risk measure that standard mean-variance optimisation does not capture; two portfolios with identical mean and variance can have very different drawdowns. If the reported dominance holds, drawdown minimisation is a tractable, data-driven route to index out-performance.","feed_headline":"Drawdown-minimising portfolios beat market indices","feed_subtitle":"Rebalancing every 10 days on 30 days of prices, the model beat the index on return, Sharpe, and drawdown.","key_machinery":"The load-bearing object is the path-dependent drawdown sequence $\\{d_t\\}$, defined against a rolling maximum $M_t$ over the last $D$ periods. The nonlinear parts of the model are the maximum in $M_t$ and the fraction in $d_t$; the paper partially linearises them by writing $M_t \\ge P_\\tau$ for every $\\tau$ in the lookback window and $d_t \\ge 100(M_t-P_t)/M_t$. Since the objective is monotone decreasing in drawdown, an optimal solution forces these inequalities to hold with equality, so no maximum operator or division needs to be solved directly. That replacement is what makes the formulation tractable enough for a nonlinear programming solver to find proven global optima in a large fraction of the rebalances, and it is the technical step the computational results rely on.","core_discovery":"The central discovery is that minimal-drawdown portfolios, obtained by solving a nonlinear program that minimises either average drawdown, maximum drawdown, or a weighted combination of the two, dominate their market index out of sample. Drawdown at time $t$ is defined as $d_t = 100(M_t-P_t)/M_t$, where $P_t$ is portfolio value and $M_t$ is the maximum value over the current and preceding $D$ periods, so it is the opportunity cost of the single best missed sale-and-repurchase trade. The paper shows that the nonlinear maximum defining $M_t$ can be replaced by linear inequalities $M_t \\ge P_\\tau$, and the equality defining $d_t$ can be relaxed to $d_t \\ge 100(M_t-P_t)/M_t$; because the objective pushes drawdown down, these inequalities are tight at an optimum. With $T=30$, $D=20$, rebalancing every 10 days, and zero transaction costs, the long-only portfolios beat the index on average daily return, annualised Sharpe ratio, maximum drawdown, and average drawdown over roughly 1,800 out-of-sample days. Allowing short positions improves in-sample performance but does not consistently improve out-of-sample returns.","pith_inferences":["The paper does not test realistic transaction costs; because the strategy rebalances every 10 days across up to 500 assets, charging even a few basis points per trade could erode or reverse the reported return edge.","The in-sample windows overlap (30 days of data reused every 10 days), so the in-sample drawdown improvements may partly reflect fitting overlapping windows; a non-overlapping robustness check would isolate the effect.","The benchmark is the market index only; comparing the same rebalancing rule with an equal-weight or momentum portfolio would show whether the benefit comes from drawdown minimisation itself or from active rebalancing in rising markets.","A natural extension is to vary $T$ and $D$ jointly to map how much price history and drawdown memory are needed for the out-of-sample dominance to persist."],"forward_implications":["If the out-of-sample dominance holds, drawdown minimisation is a viable stand-alone index out-performance rule that needs only recent prices and an optimiser, not forecasts or scenario paths.","The objective can be tuned between average drawdown, maximum drawdown, or a weighted mix, so an investor can choose whether to penalise frequent shallow dips or rare deep troughs.","Long-only minimal-drawdown portfolios appear preferable to allowing short positions under the tested parameter settings, since shorting helped in-sample metrics but did not consistently help out-of-sample returns.","The formulation already includes cash inflows and outflows as well as transaction-cost constraints, so it extends to realistic rebalancing settings with modest changes."],"supporting_citations":[{"why":"Supplies the mean-variance baseline that the paper contrasts with path-dependent drawdown.","marker":"[42]"},{"why":"Provides the drawdown-measure framework that the rolling-window formulation adapts.","marker":"[18]"},{"why":"Documents market-dependent transaction-cost estimates and motivates the zero-cost assumption used in the results.","marker":"[43]"},{"why":"Gives the annualisation formula used to compute the out-of-sample Sharpe ratios.","marker":"[46]"},{"why":"Supplies the nonlinear solver used to optimise the formulations and produce the reported results.","marker":"[53]"},{"why":"Documents the solver's ability to reach proven global optimality, supporting the reported optimality percentages.","marker":"[63]"},{"why":"Shows earlier use of drawdown lookback windows and drawdown-based CAPM, providing context for the rolling $D$-period definition.","marker":"[68]"}],"fun_headline_variants":["Minimal drawdown portfolios beat market indices on return and risk","Drawdown-minimising model dominates indices over 1,800 trading days","Nonlinear optimisation for low drawdown beats the market","Out-of-sample: minimal drawdown portfolios beat the index"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that transaction costs are zero, while the strategy rebalances every 10 days across up to 500 assets; realistic trading costs would reduce the reported out-of-sample excess return and could overturn the dominance claim.","fun_headline_variants_meta":{"raw":{"variants":["Minimal drawdown portfolios beat market indices on return and risk","Drawdown-minimising model dominates indices over 1,800 trading days","Nonlinear optimisation for low drawdown beats the market","Out-of-sample: minimal drawdown portfolios beat the index"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00135,"raw_usage":{"total_tokens":5503,"prompt_tokens":985,"completion_tokens":4518,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":4447}},"tokens_in":601,"tokens_out":4518,"duration_ms":31495,"temperature":1.0,"reasoning_tokens":4447,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:31:25.415593+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the same rebalancing strategy while charging a conservative transaction cost, for example 10 to 50 basis points of traded value, and compare net out-of-sample average daily return and Sharpe ratio with the index; if net performance no longer beats the index, the headline dominance result fails.","supporting_citations":[{"cited_title":"Portfolio selection","cited_arxiv_id":null,"evidence_quote":"Supplies the mean-variance baseline that the paper contrasts with path-dependent drawdown."},{"cited_title":"Drawdown measure in portfolio optimization","cited_arxiv_id":null,"evidence_quote":"Provides the drawdown-measure framework that the rolling-window formulation adapts."},{"cited_title":"Detection of momentum eﬀects using an index out-performance strategy","cited_arxiv_id":null,"evidence_quote":"Documents market-dependent transaction-cost estimates and motivates the zero-cost assumption used in the results."},{"cited_title":"Discovering errors in tracking error","cited_arxiv_id":null,"evidence_quote":"Gives the annualisation formula used to compute the out-of-sample Sharpe ratios."},{"cited_title":"Available from http://scip.zib.de/ Last accessed August 13 2019","cited_arxiv_id":null,"evidence_quote":"Supplies the nonlinear solver used to optimise the formulations and produce the reported results."},{"cited_title":"SCIP: Global optimization of mixed-integer nonlinear programs in a branch-and-cut framework","cited_arxiv_id":null,"evidence_quote":"Documents the solver's ability to reach proven global optimality, supporting the reported optimality percentages."},{"cited_title":"Capital Asset Pricing Model (CAPM) with draw- down measure","cited_arxiv_id":null,"evidence_quote":"Shows earlier use of drawdown lookback windows and drawdown-based CAPM, providing context for the rolling $D$-period definition."}],"review_version":1}