{"id":"386ecb9c-8298-45f8-ba35-bd12e18af4f9","arxiv_id":"1908.04962","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"Ellipsoidal and separable robust portfolio models show higher in-sample average Sharpe ratios than Markowitz on BSE 30 and BSE 100 data, but the evaluation is entirely in-sample.","lead":"The paper compares three robust portfolio optimization models against the Markowitz model on Indian stock index data and simulated data, finding that ellipsoidal and separable uncertainty models often show higher average Sharpe ratios in a narrow risk-aversion window. A generalist should read this to see whether robust optimization, designed to handle estimation error, actually helps practitioners in an emerging market.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-sample Sharpe ratios do not support the paper's claim of real-market viability; no out-of-sample validation is provided.","rationale":"The paper's formal development is standard and clearly presented, but the empirical evidence for the central claim is entirely in-sample. The claim that robust optimization is a viable alternative in a real market is forward-looking, so the relevant performance measure must be based on data not used for estimation. The paper does not provide this: Section 3 computes Sharpe ratios on the same sample used to estimate all inputs, and the simulated experiments, where the true generating parameters are known, are also evaluated in-sample. In-sample Sharpe comparisons are not trustworthy for this purpose because the Markowitz optimizer exploits in-sample moments directly, while robust models apply worst-case penalties; their relative in-sample performance may reflect different estimation biases rather than superior out-of-sample behavior. The paper itself notes unstable and partly unexplained trends in Section 4, especially around the Box model and the number of samples, and provides no confidence intervals or statistical tests. The reader's weakest assumption is exactly the load-bearing weakness: without out-of-sample validation, the conclusion that robust approaches are viable alternatives to Markowitz is unsupported. The proposed rolling-window test would settle whether the observed superiority persists in a realistic setting. Thus the REJECT verdict remains appropriate and no change is needed.","tokens_in":10570,"tokens_out":5117,"duration_ms":55732,"concrete_test":"Implement a rolling-window out-of-sample test on the same S&P BSE 30 and S&P BSE 100 daily log-returns. For each rebalance date, estimate the mean and covariance on the trailing 120 trading days; solve Mark, Box, Ellip, and Sep at each lambda in {2, 2.5, 3, 3.5, 4}; hold the resulting portfolios for the next 21 trading days; compute the realized Sharpe ratio over that holding period. Compare the average realized Sharpe of Ellip and Sep against Mark across all windows and lambda values, with bootstrap confidence intervals. If the robust models do not beat Markowitz out-of-sample, the in-sample advantage reported in Tables 3 and 6 is an artifact and the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that robust optimization is a viable alternative to Markowitz in real markets, but every supporting table (Tables 1-6) reports Sharpe ratios computed on the same sample used to estimate means, covariances, uncertainty-set parameters, and portfolio weights. Section 3 describes no train/test split, no rolling window, and no out-of-sample evaluation. In-sample Sharpe is not a valid comparison for this claim: the Markowitz weights are chosen to optimize the in-sample mean-variance tradeoff, so its in-sample Sharpe is optimistically biased; the robust models solve a different in-sample objective, so their higher tabulated Sharpe may reflect a different in-sample bias rather than better realized performance. The simulated-data experiments could have avoided this entirely by evaluating portfolios against the known true mean and covariance, but they instead report in-sample Sharpe on the generated sample. The paper itself flags unstable and unexplained trends in Section 4, further weakening the inference. Absent out-of-sample evidence, the observed Ellip/Sep advantage is not evidence of practical viability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper empirically compares the Markowitz mean-variance model with three robust portfolio optimization models (box, ellipsoidal, and separable uncertainty sets) on simulated data and on daily returns of stocks in the S&P BSE 30 and S&P BSE 100 indices. The authors report average Sharpe ratios over a risk-aversion range λ∈[2,4] and conclude that the ellipsoidal and separable robust models are viable alternatives to the Markowitz model in practice. Sections 2 and 3 present the model formulations and the computational results; Section 4 discusses the influence of the number of stocks, sample size, and data type; Section 5 concludes.","tokens_in":10753,"tokens_out":2964,"duration_ms":31510,"significance":"If the central claim were supported, the paper would provide practical guidance for Indian equity portfolio construction and contribute to the ongoing debate on whether robust optimization improves on Markowitz in real markets. The manuscript has several strengths: the robust formulations (box, ellipsoidal, separable) are standard and correctly derived; the computational experiments cover two index sizes and simulated data with two sample sizes; and the tabulated results allow easy comparison across models and settings. However, the significance of the conclusion is undermined by the evaluation methodology: all Sharpe ratios are computed in-sample, without any out-of-sample validation, statistical significance testing, or adjustment for the look-ahead bias inherent in comparing models on the data used to estimate their inputs.","major_comments":[{"comment":"The central claim that robust models are 'a viable alternative to the Markowitz model' in a real market setup is not supported by the evidence because all reported Sharpe ratios are computed on the same sample used to estimate means, covariances, uncertainty-set parameters, and portfolio weights. Section 3 describes no train/test split, rolling window, or any other out-of-sample procedure. The Markowitz weights are chosen to optimize the in-sample mean-variance objective, so its in-sample Sharpe ratio is optimistically biased; the robust models solve a different in-sample objective, so their higher tabulated Sharpe ratios may reflect a different in-sample bias rather than better realized performance. To substantiate the practical-viability claim, the authors should evaluate portfolios on a holdout sample or with rolling-window re-estimation, and for the simulated data they should compare portfolios against the known population mean and covariance or on an independently generated test sample.","section":"Section 3, Tables 1-6"},{"comment":"The general statement that 'larger the number of stocks, better is the performance of the portfolios constructed using robust optimization' is contradicted by the authors' own market-data results: Table 7 reports a maximum average Sharpe ratio of 0.2 for N=31 versus 0.194 for N=98. The subsequent paragraph acknowledges this opposite behavior for market data and attributes it to limited data and estimation error, but the paragraph's opening claim is still stated without qualification. This internal inconsistency weakens the discussion section and should be corrected by either revising the general claim or explicitly conditioning it on data type.","section":"Section 4.1, Table 7"},{"comment":"The paper states that simulated samples are generated using 'the true mean and covariance matrix' of the historical log-returns, but the text in Section 3 says these are the mean and covariance 'obtained from' or 'set to those' of the historical data. These are sample estimates, not population truths. This imprecision matters because the subsequent in-sample evaluation on simulated data is then not a test against known parameters; it is a test on a finite sample drawn from an estimated distribution. The authors should use the phrase 'estimated mean and covariance' or, better, generate a test sample from the same estimated parameters and evaluate the portfolios on that independent sample, which would provide a genuine out-of-sample check.","section":"Section 3, simulated data description"},{"comment":"No statistical significance is reported for the differences in average Sharpe ratios. For example, in Table 1 the Markowitz average is 0.181 and the separable model average is 0.182; in Table 6 the ellipsoidal model is 0.194 versus 0.182 for Markowitz. Given that these are in-sample averages over only five λ values, the differences may be within sampling noise. The authors should provide standard errors, confidence intervals, bootstrap tests, or paired tests across the λ grid (and ideally across replications for the simulated data) before concluding that one model outperforms another.","section":"Tables 1-6"}],"minor_comments":[{"comment":"The title contains a typo: 'PERFORMAN CE' should be 'PERFORMANCE'.","section":"Title"},{"comment":"There is a typo: 'condidence' should be 'confidence'.","section":"Section 2.2"},{"comment":"The paper assumes a 6% annualized risk-free rate but does not state how daily log-returns are annualized for the Sharpe ratio. Clarify the annualization convention (e.g., multiply daily Sharpe by sqrt(252) and adjust for log vs. simple returns).","section":"Section 3, Sharpe ratio calculation"},{"comment":"The description of the entries as 'maximum possible Sharpe Ratio' is ambiguous. It appears the entries are the maximum, across the four models, of the average Sharpe ratio over λ∈[2,4]; please state this explicitly in the table captions or the text.","section":"Tables 7-9"},{"comment":"Reference [15] contains a long tracking parameter (fbclid) and a broken URL format; the reference should be cleaned up or replaced with a stable citation.","section":"References"},{"comment":"The observation that the cross-over in performance with sample size is 'not obvious' for the smaller-stock case is honest, but the discussion would benefit from a hypothesis (e.g., estimation error in the covariance matrix) rather than leaving the trend unexplained.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The reader's verdict of reject is understandable given the lack of out-of-sample validation, but I see the manuscript's core flaw as a fixable methodological gap rather than an irreparable one. The model formulations are standard and the computational experiments are reproducible in structure; a major revision that adds holdout or rolling-window evaluation, statistical significance tests, and corrects the internal inconsistency in Section 4.1 could bring the evidence in line with the claim. I would not reject outright, but I would require such experiments before publication. One further concern for the editor: the paper's framing of simulated data as using 'true' parameters when they are estimated from the historical sample should be corrected; this is a presentation issue but it affects how readers interpret the simulation results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a clearly written application of three standard robust portfolio formulations to Indian index data, but the evidence for the headline claim is an in-sample comparison and that is not enough.\n\nWhat's actually new: applying box, ellipsoidal and separable uncertainty-set models to S&P BSE 30 and S&P BSE 100, and comparing simulated data against market data. The formulations are textbook but they are stated correctly, and the authors make a genuine effort to discuss practical dimensions like number of stocks and sample size. The non-parametric bootstrap construction for the separable set is a reasonable choice.\n\nWhere it falls short: the central comparison is average Sharpe ratios computed on the same sample used to estimate means, covariances, uncertainty-set radii, and the risk-free rate. There is no train/test split, no rolling window, no out-of-sample evaluation, and no measure of sampling uncertainty. That is not a minor gap; it is the whole argument. The Markowitz portfolio is optimised for that same data, so its in-sample Sharpe is biased upward, and the robust portfolios are solving a different in-sample problem — comparing the two on the training data tells you about estimation bias, not about which would do better going forward. The simulated-data experiments could have avoided this by testing against the known true mean and covariance, but they also report in-sample Sharpe on the generated samples. The paper itself flags unstable or unexplained trends in Section 4 (the sample-size effect for 31 stocks, and the Box model's inconsistency), which underscores the fragility. I also note the risk-aversion window λ∈[2,4] is fixed after seeing the results, with no sensitivity analysis on α or β.\n\nIt is not a bad paper in conception. The Indian-market angle is a legitimate extension of work by Tütüncü-Koenig and Santos, and it is well written enough that the method is reproducible. But the abstract's claim that robust approaches are a 'viable alternative' in real markets is not supported by the evidence as presented.\n\nBottom line: I would not desk-reject this outright — with an out-of-sample evaluation protocol it could be a useful empirical note — but it needs major revision before it makes any claim about practical viability. If I were handling it, I'd send it to a referee with instructions to focus on the evaluation protocol, and I'd expect the paper to come back either with rolling-window results or with a much more modest conclusion.","headline":"The math is fine but the performance claim is built on in-sample Sharpe ratios, so the Indian-market evidence doesn't support the conclusion.","tokens_in":11291,"tokens_out":2714,"would_cite":false,"duration_ms":27923,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Robust portfolio optimization matches or beats Markowitz on Indian index data.","keywords":["robust portfolio optimization","worst-case scenario","uncertainty sets","S&P BSE 30","S&P BSE 100","Markowitz model","Sharpe ratio","Indian stock market"],"falsifier":"A reader could refute the central claim by recomputing the six comparisons with a train/test split or a rolling window: estimate means, covariances, and uncertainty sets on data through September 2018, then evaluate realized Sharpe ratios on 2019–2020 returns for S&P BSE 30 and S&P BSE 100. If the ellipsoidal and separable models no longer meet or beat Markowitz on average, the central claim is falsified.","tokens_in":10361,"feed_emoji":"📈","tokens_out":14705,"duration_ms":122513,"temperature":0.7,"pith_summary":"This paper asks whether robust portfolio optimization, which builds uncertainty about estimated returns and covariances directly into the portfolio choice, can stand in for the classical Markowitz mean-variance model in practice. Using daily log-returns from the Indian S&P BSE 30 (31 stocks) and S&P BSE 100 (98 stocks) indices, plus simulated returns matched to those stocks, it compares Markowitz with three robust formulations based on box, ellipsoidal, and separable uncertainty sets. The paper argues that the answer is yes: in the risk-aversion range $\\lambda \\in [2,4]$, the ellipsoidal and separable models deliver average Sharpe ratios equal to or higher than Markowitz's, on market data as well as simulated data. The box model, by contrast, mostly tracks Markowitz. If the comparison holds, Indian practitioners have a straightforward, risk-aware alternative to mean-variance optimization that does not sacrifice risk-adjusted return.","feed_headline":"Robust portfolios match or beat Markowitz on Indian index data","feed_subtitle":"Testing on S&P BSE 30 and 100 index stocks, ellipsoidal and separable models match or beat Markowitz's Sharpe ratios.","key_machinery":"The argument turns on replacing the Markowitz objective with a worst-case (max-min) counterpart in which the expected-return vector $\\mu$ and possibly the covariance matrix $\\Sigma$ are allowed to vary inside uncertainty sets. For a general set $\\mathcal{U}$, the robust problem maximizes $\\min_{(\\mu,\\Sigma)\\in\\mathcal{U}} \\mu^\\top x - \\lambda x^\\top \\Sigma x$ subject to $x^\\top 1 = 1$ and $x \\ge 0$. A box uncertainty set around the expected return produces the penalty $-\\delta^\\top |x|$ (the Box model); an ellipsoidal set produces $-\\delta \\sqrt{x^\\top \\Sigma_\\mu x}$ (the Ellip model); and a separable set treats lower and upper bounds on every entry of $\\mu$ and $\\Sigma$ independently, yielding a tractable maximization (the Sep model). These three formulations are the machinery: they convert parameter estimation error into an explicit penalty or constraint, and the paper's empirical comparison of their Sharpe ratios is what carries the conclusion.","core_discovery":"The paper's central claim, stated in the abstract and supported by Tables 1–6, is that robust approaches are a viable alternative to the Markowitz model in a real market setup, not only in simulated data. Concretely, for S&P BSE 30 data (193 daily log-returns, December 2017 to September 2018) and S&P BSE 100 data (442 daily log-returns, December 2016 to September 2018), the ellipsoidal and separable uncertainty-set formulations produce average Sharpe ratios that are greater than or equal to the Markowitz benchmark over the risk-aversion interval $\\lambda \\in [2,4]$. On the 31-stock market data the separable model has the highest average Sharpe ratio (0.200 versus 0.189 for Markowitz), while on the 98-stock market data the ellipsoidal model is marginally ahead (0.194 versus 0.182). The paper also reports that these two models outperform Markowitz on simulated data in the same risk-aversion range, and it treats the lower-lying efficient frontiers of the robust models as evidence against the over-optimism of the Markowitz frontier.","pith_inferences":["Editorial extension: because all Sharpe ratios are computed in-sample, the paper does not itself establish that the robust advantage persists out-of-sample; a rolling-window replication would turn its central claim into a testable trading rule.","Editorial extension: the persistent edge of the ellipsoidal and separable models over the box model suggests that the geometry of the uncertainty set, not the mere act of adding robustness, is the active ingredient in the improvement.","Editorial extension: the paper leaves unexplained why, for 31 stocks, fewer simulated samples produce better performance than 1000 samples; identifying that mechanism would sharpen the practical guidance.","Editorial extension: re-running the comparison with a different assumed risk-free rate or with data after September 2018 would reveal how sensitive the model ranking is to these calibration choices."],"forward_implications":["The paper implies that an Indian large-cap investor can adopt the ellipsoidal or separable robust formulation instead of Markowitz and expect at least the same average Sharpe ratio in the $\\lambda \\in [2,4]$ risk-aversion range.","The box uncertainty model is not a useful upgrade on this evidence: its average Sharpe ratios hover at or just above Markowitz, and its Sharpe-ratio behavior is inconsistent across the data sets.","Increasing the number of stocks helps the robust models in simulated data: the maximum average Sharpe ratio rises from 0.200 at 31 stocks to 0.244 at 98 stocks when 1000 samples are used, a gain the paper attributes to diversification.","Because the robust models' efficient frontiers lie below Markowitz's, the paper reads their performance as consistent with the known over-estimation of the Markowitz frontier, meaning the robust portfolios achieve comparable risk-adjusted return with less optimistic inputs.","The paper's conclusion that robust optimization is practically useful across numbers of stocks, sample sizes, and data types flows from these Sharpe comparisons rather than from a separate out-of-sample test."],"supporting_citations":[{"why":"It defines the classical mean-variance portfolio selection problem that all robust models are measured against.","marker":"[9, 10]"},{"why":"It supplies the box and ellipsoidal uncertainty-set formulations and the risk-aversion range $\\lambda \\in [2,4]$ used in the experiments.","marker":"[5]"},{"why":"It provides the worst-case robust Markowitz formulation on which the Box, Ellip, and Sep models are built.","marker":"[7]"},{"why":"It introduces the separable uncertainty sets for expected returns and covariance that form the basis of the Sep model.","marker":"[8]"},{"why":"It develops the separable-uncertainty robust asset allocation approach whose empirical success the paper extends.","marker":"[14]"},{"why":"It gives the chi-square confidence calibration for the ellipsoidal uncertainty set and the robust variants that the Ellip model uses.","marker":"[3]"},{"why":"It reports earlier comparisons of robust versus mean-variance portfolios on simulated and real data, the debate the paper revisits.","marker":"[12]"},{"why":"It provides the skeptical result that robust portfolios can underperform Markowitz, against which the paper's positive evidence is set.","marker":"[13]"},{"why":"It documents overestimation of the estimated efficient frontier, cited to explain why robust frontiers lie below the Markowitz frontier.","marker":"[2]"},{"why":"It supplies the adjusted closing prices from which the S&P BSE 30 and S&P BSE 100 daily log-returns are computed.","marker":"[17]"}],"fun_headline_variants":["Robust models match or beat Markowitz on Indian indices","On BSE 30 and 100, robust Sharpe beats Markowitz","Ellipsoidal and separable robust models outdo Markowitz","Indian data: robust optimization rivals Markowitz performance","Robust portfolio methods offer Markowitz alternative in India"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's conclusion assumes that a portfolio's Sharpe ratio calculated on the same data used to build the portfolio tells you how it will perform in the future.","fun_headline_variants_meta":{"raw":{"variants":["Robust models match or beat Markowitz on Indian indices","On BSE 30 and 100, robust Sharpe beats Markowitz","Ellipsoidal and separable robust models outdo Markowitz","Indian data: robust optimization rivals Markowitz performance","Robust portfolio methods offer Markowitz alternative in India"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000754,"raw_usage":{"total_tokens":3346,"prompt_tokens":932,"completion_tokens":2414,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":2333}},"tokens_in":548,"tokens_out":2414,"duration_ms":17359,"temperature":1.0,"reasoning_tokens":2333,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:27:04.734089+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could refute the central claim by recomputing the six comparisons with a train/test split or a rolling window: estimate means, covariances, and uncertainty sets on data through September 2018, then evaluate realized Sharpe ratios on 2019–2020 returns for S&P BSE 30 and S&P BSE 100. If the ellipsoidal and separable models no longer meet or beat Markowitz on average, the central claim is falsified.","supporting_citations":[{"cited_title":"and Focardi, S.M","cited_arxiv_id":null,"evidence_quote":"It supplies the box and ellipsoidal uncertainty-set formulations and the risk-aversion range $\\lambda \\in [2,4]$ used in the experiments."},{"cited_title":"and Fabozzi, F.J","cited_arxiv_id":null,"evidence_quote":"It provides the worst-case robust Markowitz formulation on which the Box, Ellip, and Sep models are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It introduces the separable uncertainty sets for expected returns and covariance that form the basis of the Sep model."},{"cited_title":"and Koenig, M","cited_arxiv_id":null,"evidence_quote":"It develops the separable-uncertainty robust asset allocation approach whose empirical success the paper extends."},{"cited_title":"and Stubbs, R.A","cited_arxiv_id":null,"evidence_quote":"It gives the chi-square confidence calibration for the ellipsoidal uncertainty set and the robust variants that the Ellip model uses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It reports earlier comparisons of robust versus mean-variance portfolios on simulated and real data, the debate the paper revisits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the skeptical result that robust portfolios can underperform Markowitz, against which the paper's positive evidence is set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It documents overestimation of the estimated efficient frontier, cited to explain why robust frontiers lie below the Markowitz frontier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the adjusted closing prices from which the S&P BSE 30 and S&P BSE 100 daily log-returns are computed."}],"review_version":1}