{"id":"b520b0d6-933d-4c95-951b-486587210ef3","arxiv_id":"2607.27461","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A three-matrix, rank-based Markov-chain portfolio reports out-of-sample Sharpes of 1.06 and 1.32 vs the market's 0.78 and 1.14, net of costs, even though its return rank is itself 'close to unforecastable.'","lead":"This paper tests a simple stock-picking rule for S&P 500 names built from three matrices of price, volume, and market-cap data: a correlation-distance matrix plus two monthly ranking-transition matrices. On two out-of-sample periods it reports higher Sharpe ratios than the market (1.06 and 1.32 vs 0.78 and 1.14) after trading costs, though the cleanest window's edge is only marginally statistically significant.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Return-chain unforecastability contradicts the reported long-short edge; until the actual long-short's positive spread is traced to the π_R forecasts or another explicit channel, the three-matrix attribution is unsupported.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing tension: the paper's own calibration says the return rank is unforecastable and the momentum spread has no positive forward edge, while Section 4 reports a profitable long-short selected by that forecast. My reading agrees that this must be resolved before the central attribution claim can be accepted. I give credit where it is due: the paper has a genuinely non-overlapping second out-of-sample window, honest reporting of failed ablations, and a block bootstrap on the second test. But those do not resolve the attribution problem. They show that some strategy on this data beat the market; they do not show that the three-matrix forecast representation is what generated the edge. The concrete test I propose is a direct ablation: randomize the π_R scores while keeping everything else identical. If the Sharpe survives, the forecasts are not the driver; if it collapses, the paper's Section 3.4 claims need to be reconciled with the predictive content of π_R. Because the manuscript provides no public code or data (the repository is private), this test cannot currently be run by a reviewer, which is an additional reason to keep the verdict conditional rather than accepting the market-beating attribution on the current evidence.","tokens_in":39275,"tokens_out":8234,"duration_ms":90882,"concrete_test":"Re-run the Section 4 backtest on both out-of-sample windows with a single change: each month, before selecting the 2K long-only names and the K-long/K-short legs, randomly permute the calibrated π_R scores across the cross section, preserving the set of scores. Keep K=15, tol=0.08, the λ/θ regime timing, five-basis-point costs, and walk-forward refitting identical. If the randomized-score book retains a Sharpe near the reported 1.06/1.32 (or the long-short near 0.6), the return-chain forecast is not the source of the edge, confirming the contradiction. If the Sharpe collapses, the residual signal in π_R is doing the work despite Section 3.4, and the realized monthly spread of the actual long and short legs should be reported to locate the source.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the portfolio's outperformance is generated by the three-matrix representation. The only return-based selection signal is π_R (Eq. 20), the covariate-conditioned probability of landing in the top return decile. Section 3.4 reports, on the same data and horizon, that the one-month return rank is near-random, the six-month persistence is mechanical overlap, covariates add nothing to the point forecast, and 'the next-month return of the top decile minus the bottom ... never a positive forward edge' at any ranking window. If that is correct, π_R is almost uninformative about next-month returns, so the Section 4 long-short selected by π_R should have approximately zero expected gross spread before costs, not an out-of-sample Sharpe near 0.6 (and not the combined book's 1.06/1.32). The paper never isolates the actual channel that makes the traded long-short profitable: it could be a volatility tilt, regime luck, no-trade-band path dependence, or a backtest artifact. This is an internal inconsistency, not a disagreement with consensus: either Section 3.4's spread is computed on a different object than the traded portfolio, or the traded portfolio's edge is not attributable to the forecasts. Table 5's finding that the long-only size gain is 'market breadth, not selection alpha' reinforces the concern that the market-beating result may be driven by a different mechanism than the three-matrix forecasts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a three-matrix representation of an equity market—an arccos correlation-distance matrix and two Markov-chain transition matrices on monthly return and volatility deciles—and conditions the chains on market-derived covariates. It estimates covariate-conditioned forward-rank probabilities π_R and π_V, builds a monthly-rebalanced portfolio that blends a market-neutral momentum long-short with a regime-timed long-only sleeve, and reports out-of-sample Sharpe ratios of 1.06 and 1.32 (1.08 and 1.44 with residual-distance diversification) against market Sharpes of 0.78 and 1.14, net of five-basis-point costs. The paper also reports diagnostics (entropy production, transfer entropy), a second clean test window on an independent data source, and several ablations of fundamentals, spectral early-warning signals, and volume-based signals.","tokens_in":39719,"tokens_out":8722,"duration_ms":85543,"significance":"The framework is refreshingly simple, mostly walk-forward, and the paper is unusually honest about the ways its first out-of-sample window was inspected and about failures of candidate improvements on the clean holdout. If the central attribution is correct, the result would be practically important and would sit naturally alongside documented momentum and low-volatility anomalies. The strongest features are the second independent test window with frozen hyperparameters, the explicit cost modeling, and the many negative results reported in the appendices. However, the central attribution of the outperformance to the three-matrix forecasts is undermined by an internal inconsistency between the forecast calibration results and the traded long-short's reported edge, and the statistical significance of the headline Sharpe advantage is not established.","major_comments":[{"comment":"Section 3.4 states that the one-month return rank is near-random, that covariates add nothing to the return-rank point forecast at any window, and that 'the next-month return of the top decile minus the bottom ... never a positive forward edge' at any ranking window. Section 4 then reports that the market-neutral long-short selected by π_R of Eq. (20) — with λ=0, so the volatility score drops out — has a 'real gross edge' (Sharpe 0.17) and an out-of-sample Sharpe of 0.63. Since §3.4 says the covariates add nothing and the book ranks on a six-month return window, π_R should be nearly equivalent to the trailing-return decile, making the traded long-short exactly the top-minus-bottom spread that §3.4 reports as never positive. This is a load-bearing internal inconsistency: either the §3.4 statistic is computed on a different object than the traded book, or the long-short's edge comes from a","section":"§3.4, §4"},{"comment":"The statistical support for the headline claim is weak and not adjusted for multiple testing. The clean-window base-book excess-return p-value is 0.09, and the full-record base-book p-value is 0.08; only the diversified clean-window estimate reaches p=0.04. No correction is made for the many variants, periods, and refinements examined in Sections 4–5 and Appendices A–C. Moreover, the abstract's Sharpe comparisons (1.32 versus 1.14, 1.44 versus 1.14) are never tested; the bootstrap addresses only mean excess return. The paper should either provide multiple-testing-adjusted or pre-registered significance statements, report confidence intervals for the Sharpe differences, or temper the abstract's claims of beating the market.","section":"§5, Table 11"},{"comment":"The attribution of the long-only sleeve's contribution is unresolved. Table 5 shows that the long-only book's Sharpe is essentially flat in book size on the validation period and first test, and that the gain on the 2025–2026 window is 'market breadth, not selection alpha.' Because the combined book is fully long-only in rising markets (θ=1), the clean-window Sharpe of 1.44 could owe substantially to holding a broad equal-weight set of names in a strong bull market rather than to the return-chain forecasts. The paper should decompose the combined book's excess return into cap-weight market, equal-weight breadth, and selection/risk-tilt components to show what the three-matrix forecasts actually add.","section":"§4.1, Table 5"}],"minor_comments":[{"comment":"The text says the no-band long-short has a Sharpe of 0.17, while the Figure 10 caption reports 0.13. Please reconcile.","section":"Fig. 10"},{"comment":"The abstract calls both test sets 'non-overlapping out-of-sample,' but §5 concedes that the 2022–2024 period was used repeatedly to refine the book and is 'no longer perfectly clean.' Please qualify the first test in the abstract.","section":"§5"},{"comment":"The distance matrix is N×N with N varying as the S&P 500 constituent list changes; calling the three matrices 'fixed-size' is misleading. Clarify that only the transition-matrix dimensions are fixed.","section":"§2.1"},{"comment":"The Yahoo panel of 'the 475 names with a full history since 2018' still induces survivorship for the 2022–2024 cross-source slice and for the beginning of the clean window. The paper acknowledges a mild survivorship tilt for the earlier slice; please quantify the effect on the clean window as well.","section":"§5"},{"comment":"The code repository is currently private. For a paper whose central claim is a strong empirical outperformance, public code and data or a detailed pseudocode appendix are important for verification.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is thoughtful and unusually candid about its limitations, but the contradiction between §3.4's 'never a positive forward edge' and §4's profitable long-short is at the core of the market-beating claim. If the author can trace the long-short's gross edge to an explicit mechanism or show that §3.4's spread is a different object, the paper could become publishable after also addressing the statistical-significance concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a serious attempt, but the central attribution does not hold as written. The three-matrix representation is elegant, and the second out-of-sample test is genuinely clean—frozen hyperparameters, independent data source, daily marking. What the paper does well: the covariate-conditioned ranking chains are a natural extension of OMD-Stocks; the entropy-production convexity bound (conditioning cannot decrease it) is correct and gives a useful attribution; the asymmetry between the forecastable volatility rank and the near-unforecastable return rank is documented carefully; and the negative results on fundamentals, spectral timing, and sizing are reported honestly rather than buried.\n\nThe soft spot is load-bearing. Section 3.4 states that the one-month return rank is near-random, that at any ranking window the next-month return of the top decile minus the bottom is never positive, and that covariates add nothing to the return-chain point forecast. Section 4 then builds the tradeable long-short entirely on π_R, the covariate-conditioned probability of landing in the top return decile, with λ=0 so the volatility score drops out. If the return rank is unforecastable, π_R is essentially current rank, and the long-short is a six-month momentum book whose expected next-month spread the paper itself says is negative. Yet the long-short earns a positive out-of-sample Sharpe. The paper never traces this edge to a specific channel—volatility tilt, lead-lag hedge, no-trade path dependence, or regime luck. Until that is done, the claim that outperformance comes from the three-matrix forecasts is unsupported.\n\nOther issues are milder. The clean-window excess-return p-values are 0.09 (base) and 0.04 (diversified)—borderline, and the Sharpe differences are not significant. The code and data are not public, so the backtests cannot be independently checked. The long-only size effect is, as the paper itself says, market breadth rather than selection alpha.\n\nNet: this deserves a serious referee, but the referee should demand a direct reconciliation of the Section 3.4 and Section 4 results, and an isolation of the actual source of the long-short's edge. Without that, the market-beating headline is not established.","headline":"Honest, transparent framework with a genuinely clean second test set, but the market-beating long-short's edge is never traced to the forecasts, and the paper's own unforecastability results contradict it.","tokens_in":40115,"tokens_out":6293,"would_cite":false,"duration_ms":65484,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G10","62M05","94A17"],"pacs":[],"model":"deepseek-v4-flash","headline":"A portfolio built from three price-only matrices outperforms the S&P 500 out of sample, the paper reports.","keywords":["portfolio optimization","Markov chains","momentum long-short","cross-sectional rankings","correlation distance matrix","entropy production","transfer entropy","out-of-sample Sharpe"],"falsifier":"Rebuild the combined book with the return-chain score replaced by a placebo score (e.g., permuted or reverse-ranked π^R, or random selection within the same deciles) while keeping the distance-matrix covariates, the regime timing, the no-trade band, and the long-only sleeve unchanged; if the out-of-sample Sharpe stays near 1.06 and 1.32, the claim that the three-matrix return forecast carries the edge is falsified. A complementary check is to regress the book's daily excess returns on standard factor returns (market, SMB, HML, momentum, low-vol) and require that the intercept remain economical","tokens_in":39216,"feed_emoji":"📈","tokens_out":4491,"duration_ms":44119,"temperature":0.7,"pith_summary":"The paper argues that a dynamic portfolio strategy for the S&P 500 can be driven entirely by three fixed-size matrices derived from daily prices, volumes, and market caps: the arccos (geodesic) distance matrix of return correlations, and the transition matrices of two Markov chains that rank names monthly by trailing return and trailing volatility. On two non-overlapping out-of-sample windows (2022–2024 and 2025–2026), the resulting market-neutral momentum long-short blended with an opportunistic long-only sleeve beats the cap-weighted index, with Sharpe ratios of 1.06 and 1.32 versus the market's 0.78 and 1.14, net of 5 basis point trading costs. The paper also reports that the volatility rank is forecastable while the return rank is near-unforecastable, and that the edge survives a cross-source rebuild with frozen hyperparameters. A sympathetic reader would care because the claim is that a transparent, invertible, parameter-light representation of the market—not a black-box predictor—can systematically beat the index from public data alone; if true, it sits alongside documented momentum and low-volatility anomalies.","feed_headline":"Three price-only matrices beat the S&P 500","feed_subtitle":"Out-of-sample Sharpes of 1.06 and 1.32 versus the index's 0.78 and 1.14, net of 5 bp costs.","key_machinery":"The object that carries the argument is a triple of fixed-size matrices: M(t) = arccos C(t), the geodesic distance matrix of trailing one-year return correlations; and P^R, P^V, the 10×10 transition matrices of two decile-ranking Markov chains (return rank, volatility rank). The chains are extended to depend on covariates by a log-linear conditional-intensity model whose coefficients are estimated by convex maximum likelihood; conditioning can only increase entropy production, and the gap Δσ = σ_cond − σ is used to attribute directional structure to each covariate. The selection score is the calibrated probability of landing in the top return-decile next month (π^R), sometimes blended with t","core_discovery":"The central claim is that the three-matrix representation—the correlation-distance matrix together with return-rank and volatility-rank transition matrices, conditioned on cross-sectionally bucketed price and volume covariates—is sufficient to build a portfolio that beats the cap-weighted S&P 500 out of sample. The author reports daily-marked, cost-net Sharpe ratios of 1.06 (2022–2024) and 1.32 (2025–2026) for a market-neutral momentum long-short blended with a regime-timed long-only sleeve, against the market's 0.78 and 1.14, and further improvement (to 1.08 and 1.44) when the long sleeve is diversified by residual distance. The paper is explicit that the volatility chain is forecastable wh","pith_inferences":["The paper's own calibration—one-month return rank near-random, six-month persistence mechanical, and no positive forward edge for the top-minus-bottom decile spread—implies that the reported long-short Sharpe cannot be coming from the return-chain forecast alone; an editor's inference is that the edge must be investigated for hidden channels such as a volatility tilt, the regime-timed long sleeve,","If the volatility chain carries genuine forecastability, a natural extension the author leaves implicit is to build the long-short from volatility-rank scores (or a blended score with lambda > 0) and test whether the edge persists; the paper's own lambda = 0 calibration suggests the author found this unhelpful, but the tension is worth resolving.","The information-dissipation length and spectral early-warning signals are shown to be contemporaneous rather than predictive outside 2008; a testable extension is to combine the regime-aware entropy production with the volume-turnover signal (real out of sample, too fast to trade) as a conditioning covariate at a slower horizon.","The strongest falsifiable prediction is the magnitude of the edge on a third, longer clean window; the paper's block bootstrap p-values (e.g., 0.04 for the diversified book on 2025–2026) leave non-trivial sampling uncertainty, so an independent replication over 2–3 years would be the decisive check."],"forward_implications":["If the reported out-of-sample results are reproducible, then a market-neutral momentum book built from rank-chain probabilities can carry a tradable edge even when univariate return forecasts look random, provided turnover is controlled by a no-trade threshold.","The framework's differentiation of the volatility chain (forecastable, long memory) from the return chain (near-random at monthly horizon) suggests that risk-ranking dynamics are the more reliable substrate for dynamic portfolio construction.","Because the method uses only prices, volumes, and market caps and never inverts a covariance matrix, it offers an interpretable, low-complexity alternative to mean-variance optimization that practitioners could adopt without proprietary data.","The clean 2025–2026 test, rebuilt from an independent data source with frozen hyperparameters, is the paper's strongest evidence that the edge is not an artifact of repeated inspection of the first test period.","The convex information-leader overlay shows an explicit mechanism to convert a network-derived signal into crash insurance without changing the Sharpe ratio, which is a concrete design for drawdown control."],"fun_headline_variants":["Three price-only matrices beat S&P 500 out-of-sample","Three matrices, no inversion, beat S&P 500 net of costs","Three-matrix momentum beats S&P 500, Sharpe 1.32","Price-only matrices outperform S&P 500 in tests","Three matrices outdo S&P 500 without covariance"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The reported out-of-sample long-short edge is attributed to the return-chain selection score, but the paper's own calibration states that the one-month return rank is near-random and that the top-minus-bottom forward edge is never positive, so the load-bearing premise is that this long-short outperformance truly comes from the three-matrix representation rather than from volatility or regime tilts, look-ahead in the walk-forward refit, or trading-cost assumptions.","fun_headline_variants_meta":{"raw":{"variants":["Three price-only matrices beat S&P 500 out-of-sample","Three matrices, no inversion, beat S&P 500 net of costs","Three-matrix momentum beats S&P 500, Sharpe 1.32","Price-only matrices outperform S&P 500 in tests","Three matrices outdo S&P 500 without covariance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000972,"raw_usage":{"total_tokens":4046,"prompt_tokens":901,"completion_tokens":3145,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":3058}},"tokens_in":645,"tokens_out":3145,"duration_ms":23090,"temperature":1.0,"reasoning_tokens":3058,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:35:00.983253+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rebuild the combined book with the return-chain score replaced by a placebo score (e.g., permuted or reverse-ranked π^R, or random selection within the same deciles) while keeping the distance-matrix covariates, the regime timing, the no-trade band, and the long-only sleeve unchanged; if the out-of-sample Sharpe stays near 1.06 and 1.32, the claim that the three-matrix return forecast carries the edge is falsified. A complementary check is to regress the book's daily excess returns on standard factor returns (market, SMB, HML, momentum, low-vol) and require that the intercept remain economical","supporting_citations":[],"review_version":1}