REVIEW 4 major objections 5 minor 3 references
Conformal Kelly: Conformal Prediction Intervals as the Scale in Fractional Kelly Position Sizing
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that a 75% conformal interval's width can serve as the volatility scale in fractional Kelly position sizing, and that when it does, a slow, unweighted rolling quantile beats every locally adaptive alternative.
desk verdict Unusually honest pre-registered study of conformal intervals as Kelly scale, but the central design rule—slow beats fast—is a development-window finding with no out-of-sample confirmation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the per-asset conformal half-width $q^{eff} = (q^{roll})^{0.7}(q^{anchor})^{0.3}$, divided by a fixed constant to form $\hat{\sigma}$ and squared into the fractional Kelly denominator. It serves three roles: inverse-variance model averaging, position sizing via $\hat{\mu}/\hat{\sigma}^2$, and (through its downside miscoverage) a book-level leverage dial. The contrast between the two uses is the argument: the scale must be slow and per asset, while the regime signal must be pooled. The hard gross cap (binding 98% of days) keeps the book below the Kelly optimum so cross-sectional ratios, not the overall scale, drive performance.
What would settle it
Run the identical frozen harness and ablation grid on a second asset universe — the paper names this its highest-value follow-up: if any locally adaptive conformal method (adaptive conformal inference, recency weighting, volatility-scaled scores) beats the slow rolling quantile there, the design rule is a development-window artifact. A cheaper test: record the daily log-change variance of each scale estimator, as the paper did for holding-period residuals (0.00343 vs 0.00391); if estimator smoothness does not track the growth ordering, the 'noisy scale is charged to wealth' mechanism fails.
Extended reading notes
Core claim
Using a 75% conformal interval's width as the Kelly volatility scale, per asset from a 500-day rolling window with anchor shrinkage, feeds $f = \kappa\hat{\mu}/\hat{\sigma}^2$ and compounds at 28.5% annualized net log growth (Sharpe 1.34) on 2016–2021 ETF data under costs and a 2.0 gross cap. Ablations show the growth ordering matches tail-sensitivity: bounded quantile beats linear MAD beats quadratic SD by 2.1 pp at matched leverage; every locally adaptive variant loses. A pooled downside-miscoverage dial cuts drawdown from 27.7% to 20.3% at better Sharpe. The sealed 2022–2024 test kept coverage (0.745 vs 0.750) but not growth (~30% of DEV).
Load-bearing premise
The load-bearing premise is that six years (2016–2021) on eight US-listed ETFs is representative enough that the observed ordering — slow, unweighted, per-asset rolling conformal quantiles beating every faster alternative — is a general property of the Kelly sizing map and not regime luck or search selection; the paper itself says external validity is untested and that the binding 2.0 gross leverage cap and 2020 do enormous structural work.
Editorial extensions
If this is right
- If the design rule holds, any volatility estimate consumed by a nonlinear sizing map should be chosen for stability, not local accuracy: the paper's growth ordering matches the influence-function ordering of the estimators (bounded quantile beats linear MAD beats quadratic SD).
- Conformal coverage can transfer to regimes the estimator never saw even when the economic value built on it does not: the sealed test showed 0.745 realized coverage against 0.750 nominal through the 2022 drawdown.
- The two conformal signals have opposite aggregation requirements: width should be per asset and slow, while the downside miscoverage dial should be pooled across assets and used only to cut leverage, never to raise it.
- A disclosed large agentic search can carry confirmatory weight only when paired with a sealed window and a pre-registered interpretation rule written before unsealing; the paper's registered prediction (growth shrinks to roughly 30% of DEV) held quantitatively.
- Practically, the paper concludes the specific book is not investable as configured: at 50 bps costs and 4% financing, lockbox growth falls to roughly +1–5% per year, below every unlevered passive bar.
Reading between the lines
- My inference: the slow-stability principle should generalize to any strategy that consumes a scale estimate nonlinearly — vol-targeting, risk parity, option pricing — and is directly testable by rerunning the same slow-versus-adaptive ablation on a second universe, which the paper names as the highest-value follow-up.
- My inference: the width-as-scale reading predicts a measurable diagnostic the authors did not run — that squared conformal half-width tracks realized variance and that the growth gap between estimators is predictable from the variance of daily changes in the scale; the paper measured this smoothness only for the holding-period case (daily log-change std 0.00343 at $H_s=21$ vs 0.00391 at $H_s=5$).
- My inference: the DEV/lockbox reversal between the post-hoc equity bar (B7) and its commodity mirror (B8) quantifies how much apparent alpha in an agentic search can be hindsight asset selection — roughly 20 pp of annual growth on DEV — and implies frozen-harness comparisons should include the mirror control by default.
- My inference: since the drawdown dial's Sharpe benefit vanished out of sample while its drawdown truncation survived, trailing miscoverage is best read as a slow regime indicator whose value is tail insurance, not return enhancement; a longer lockbox window or a stress-universe test would settle whether the DEV placebo result was luck.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces "Conformal Kelly": a fractional-Kelly position-sizing rule whose volatility input is read from a 75% conformal prediction interval for 21-day returns, applied to eight US-listed ETFs. On the 2016-2021 development window the best configuration earns 28.5% annualised net log growth (Sharpe 1.34, max drawdown 27.7%), and a drawdown-dial variant earns 25.8% (Sharpe 1.39, max drawdown 20.3%). The paper's principal development-window finding is a design rule: slow, unweighted, per-asset rolling conformal quantiles outperform locally adaptive conformal methods and a rolling standard deviation as the Kelly denominator. The authors disclose an autonomous LLM-agent search over roughly 200 configurations, seal all data from 2022 onward, and pre-register configurations, benchmarks, and interpretation rules before one lockbox evaluation. The lockbox shows coverage transfers (0.745 vs 0.75) but growth does not (8.5% and 7.0% per year vs passive bars), and the registered risk-adjusted defence fails; the authors conclude that the calibration mechanism transfers but the economic value of the sizing rule does not.
Significance. If the "slow beats fast" design rule were externally valid, the paper would make a substantial contribution to both conformal prediction and Kelly sizing: a scale estimator with a measured interior optimum in adaptation speed, opposite to the time-series conformal-prediction literature's emphasis on local adaptivity. The execution has real strengths: a sealed, pre-registered lockbox with verbatim committed output; explicit disclosure of the agentic search and of multiple failed hypotheses; honest reporting of a known z inconsistency, a post-hoc 4-equity benchmark, and the failure of the registered risk-adjusted defence. The central limitation is equally clear: the design rule is established only on the same DEV window used for the ~200-configuration search, and the lockbox did not re-run the ablation suite, so the estimator ordering has zero out-of-sample confirmation. The only out-of-sample positive is marginal calibration of the intervals, which tests coverage, not whether interval width is a useful Kelly scale. The significance is therefore conditional on an independent validation of the ordering, or on a re-framing of the paper as a primarily negative result.
major comments (4)
- [§6.1, §9, §11] The load-bearing claim that 'slow beats fast' is established only on the searched DEV window. §9 reports roughly 200 configurations scored on DEV; §6.1 reports all adaptive-device contrasts as DEV values; Limitation #1 states external validity is untested. The lockbox (§10) tested only the two frozen configurations against passive bars and did not re-run the ablation suite, so the estimator ordering has no out-of-sample support. Because the claim contradicts the standard time-series CP recommendation, it cannot be promoted as a general design principle from a single searched window. A revision should either add an independent validation of the ablation ordering (e.g., a re-registered second universe or a within-sample pseudo-OOS scheme) or explicitly demote the claim to a DEV observation in the title and abstract.
- [§5.1, Table 4] The conformal-vs-standard-deviation ablation has no paired confidence intervals, and the headline gap is not exactly matched. The text reports a 3.5 pp leg-3 gap at post-cap gross 1.960 vs 1.922, and the gross-matched replay narrows it to roughly 2.2 pp; the exactly matched leg-2 gap is 2.1 pp, yet no paired interval is given. One standard error on annualised growth is quoted as about 0.095 (§8.1), so the reader cannot judge whether 2.1 pp is distinguishable from zero. In addition, the decomposition of the 3.5 pp gap into updating (+1.5 pp) and functional change (+1.9 pp) is not identified because the frozen-SD cell is missing. Add block-bootstrap paired intervals as listed in Limitation #1, or weaken the corresponding claims.
- [§4, §11 Limitation 2] The gross cap is load-bearing for the design rule. §4 reports the book is pinned at the cap on 97.7% of DEV days and states that under the cap anything tilting toward low-volatility assets loses; §11 Limitation 2 concedes that where the cap does not bind the conclusions would probably be different. The slow conformal estimator wins partly by understating sigma for fat-tailed assets, keeping more weight on them under the cap—a regime-specific mechanism. Without checking the estimator ordering at other cap levels, or in an unconstrained variant, the proposed 'stability of the width' principle is confounded with cap saturation. A concrete cap-interaction test, or an explicit restriction of the principle to cap-constrained books, is needed.
- [§10, §8.1, Abstract] The lockbox results, honestly reported, refute the economic claim: growth fell to roughly 30% of DEV, both configurations rank last on Sharpe and Calmar, and the drawdown dial fails its pre-registered 'Sharpe no worse' condition. The paper does disclose this, but the abstract's first two paragraphs still lead with DEV Sharpe/Calmar and the placebo p-value. Given the pre-registered interpretation is a partial refutation, the framing should lead with the failed confirmation and identify the paper as a negative-result study for the economic value of the sizing map, with the DEV design rule as hypothesis-generating only. This is not a request to hide the DEV results; it is a request to align emphasis with the confirmatory evidence.
minor comments (5)
- [Table L2] Bar B8 is labeled 'control' but on the lockbox it is the best performer on growth. A more neutral label such as 'commodity-block post-hoc book' would avoid implying it is a null control.
- [Figure 1] The legend lists 'SPY (true, unlevered 1.0×)' while the caption mentions the harness-capped SPY bar at 0.75×; clarify that the true SPY series is outside the harness and is not scored by prepare.py.
- [§5.3] The first row of the contribution table gives the horizon change as 'daily to 21-day forecast horizon'; for completeness, state the horizon value h=21 in the table itself.
- [§7] The placebo p-value appears as '0.024' in the abstract and '1/41' in the text; unify the notation (for example, 'p = 1/41 ≈ 0.024' in both places).
- [§3] The disclosed z-mismatch is important: z=1.2816 rather than 1.1503 changes the effective Kelly fraction by 24%. Since the paper is precluded from a post-unsealing sensitivity rerun, consider adding a footnote quantifying the implied scale change and reminding readers that the gross cap makes the effect bounded rather than exactly measured.
Circularity Check
No significant circularity: the conformal-width-as-Kelly-scale construction is an empirically tested rule, not a definitional identity, and the pre-registered lockbox gives independent evaluation.
full rationale
Walking the derivation chain: nonconformity scores are causal absolute residuals, the conformal quantile q_eff is formed from those scores, sigma = q_eff/z, and the sizing map is f = kappa * mu / sigma^2. Each stage consumes the previous stage's output without re-supplying the claim being tested; growth, Sharpe, and drawdown are measured by simulation, not derived from the interval width by construction. The main design rule ('slow, unweighted, per-asset rolling conformal quantiles beat locally adaptive alternatives') comes from paired DEV ablations; the paper itself labels these exploratory: 'These DEV comparisons are exploratory: the grids were searched, and the winning entries are selected values rather than pre-specified hypotheses; only the lockbox carries confirmatory weight.' That is in-sample selection, which the paper discloses and mitigates with a sealed lockbox, not circular reasoning. The lockbox did not re-run the ablation ordering, so the positive design rule lacks out-of-sample confirmation; this is an external-validity limitation, which the paper states: 'External validity is untested; a second universe was forbidden by the frozen harness and is the highest-value follow-up.' Similarly, the gross cap doing 'enormous structural work' and 2020 being a large fraction of DEV are acknowledged regime/selection concerns, not equation-level circularity. No load-bearing self-citation was found: the cited prior work (Vovk, Kelly, Gibbs & Candes, etc.) supplies standard tools, not the paper's conclusion. The closest thing to a circular flavor — using the conformal width both as an ensemble weight and as the sizing denominator — is repeated use of the same computed quantity, not a definitional equivalence between input and output. Under the stated review rules, none of these rises to a circular step; score 0.
Assumptions & free parameters
free parameters (11)
- Kelly fraction kappa =
0.15 (reduced from 0.25)
- nominal miscoverage alpha =
0.25 (swept 0.05-0.45)
- calibration window W =
500 days (from 250)
- anchor shrinkage lambda =
0.3
- forecast horizon H and ensemble =
21 days; ensemble {12,16,21,27,34}
- per-asset position clip =
0.75 (harness MAX_PER_ASSET)
- gross leverage cap =
2.0
- drawdown dial weight beta =
1.0
- drawdown dial window M =
21 days
- Gaussian constant z =
1.2816 (inconsistent with alpha=0.25)
- ridge regularization lambda =
10
assumptions (7)
- domain assumption Fractional Kelly under expected log utility is the appropriate objective for a wealth-compounding strategy.
- domain assumption Conformal interval width divided by a Gaussian quantile is an adequate estimate of the scale sigma that Kelly sizing needs.
- domain assumption Realized 75% coverage is the relevant calibration property of the interval for the sizing map.
- domain assumption The frozen Kaggle price snapshot for the eight ETFs is accurate and free of survivorship or look-ahead issues.
- ad hoc to paper The 2016-2021 DEV window is representative enough to infer a general design principle.
- domain assumption A flat 5 bps per unit of turnover and a one-day execution lag are adequate proxies for real trading frictions.
- domain assumption A zero risk-free rate is a harmless simplification on the development window.
Cite this review
Pith. "Pith review of Conformal Kelly: Conformal Prediction Intervals as the Scale in Fractional Kelly Position Sizing." pith.science (2026). https://pith.science/paper/JZ4QRTFA
@misc{pith2026260801494,
author = {Pith},
title = {Pith review of: Conformal Kelly: Conformal Prediction Intervals as the Scale in Fractional Kelly Position Sizing},
year = {2026},
howpublished = {\url{https://pith.science/paper/JZ4QRTFA}},
note = {Machine review of arXiv:2608.01494}
}
read the original abstract
Conformal prediction has traditionally been used to quantify prediction uncertainty. We put that uncertainty to a second use, combining a 75% conformal interval with fractional Kelly to size portfolio positions: as the range widens we shrink the position, and as it narrows we grow it. On a six-year development window (2016-2021), with trading costs and strict leverage caps, this compounds at 28.5% annualised net log growth with a Sharpe ratio of 1.34 and a 27.7% maximum drawdown, versus 15.9% for holding the S&P 500 and 21-22% for passive portfolios at the same leverage. Our main development-window finding runs against the literature's advice for conformal prediction on time series. Every tweak that adapts the interval faster to market conditions costs 0.7 to 5.3 points of annual growth; the winner is the simplest method: slow, unweighted, per-asset rolling quantiles. When an interval sizes a position rather than describing one forecast, width stability beats local sharpness. It also beats the textbook standard deviation by 2.1 points at matched leverage. We also implement a risk control: when the intervals miss on the downside far more than their historical rate, we cut leverage. On the development window this cut maximum drawdown from 27.7% to 20.3% while raising the Sharpe ratio, beating all 40 placebo timings (rank-based p = 1/41). These numbers came from an autonomous LLM-agent search over 200 configurations, so we sealed all data from 2022 onward and pre-registered configurations, benchmarks, and interpretation rules before one evaluation. Calibration held (0.745 coverage against 0.750, weakest through 2022); growth did not: the two configurations earned 8.5% and 7.0% per year, below the passive benchmarks, and a pre-registered hindsight benchmark beat them on raw growth while taking a 46% drawdown. All outcomes are reported as pre-registered.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Barber, R.F., Candès, E.J., Ramdas, A. and Tibshirani, R.J. (2023). Conformal prediction beyond ex- changeability.TheAnnalsofStatistics 51(2),816–845. Browne, S. (1997). Survival and growth with a liability: optimal portfolio strategies in continuous time. MathematicsofOperationsResearch 22(2),468–493. Chan, J.S. et al. (2024). MLE-bench: Evaluating Machi...
arXiv 2023
-
[32]
arXiv:1905.03222. Vovk, V. and Bendtsen, C. (2018). Conformal predictive decision making.Proceedings of the Seventh WorkshoponConformalandProbabilisticPredictionandApplications ,PMLR91,52–62. Vovk, V., Gammerman, A. and Shafer, G. (2005; 2nd ed. 2022).Algorithmic Learning in a Random World. Springer. Vovk, V., Lindsay, D., Nouretdinov, I. and Gammerman, A...
arXiv 1905
-
[34]
arXiv:2106.00170. Grossman, S.J. and Zhou, Z. (1993). Optimal investment strategies for controlling drawdowns.Mathe- maticalFinance 3(3),241–276. Jegadeesh, N. and Titman, S. (1993). Returns to Buying Winners and Selling Losers: Implications for StockMarketEfficiency. TheJournalofFinance 48(1),65–91. Jia,Y.andHan,B.(2026). PortfolioSelectionwithAdaptiveCon...
arXiv 1993
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.