Pith. sign in

REVIEW 4 major objections 5 minor 3 references

Conformal Kelly: Conformal Prediction Intervals as the Scale in Fractional Kelly Position Sizing

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that a 75% conformal interval's width can serve as the volatility scale in fractional Kelly position sizing, and that when it does, a slow, unweighted rolling quantile beats every locally adaptive alternative.

desk verdict Unusually honest pre-registered study of conformal intervals as Kelly scale, but the central design rule—slow beats fast—is a development-window finding with no out-of-sample confirmation. read the letter →

arxiv 2608.01494 v1 pith:JZ4QRTFA submitted 2026-08-02 q-fin.PM q-fin.RM

classification q-fin.PMq-fin.RM MSC 62G1591G10
keywords conformalpredictionKellycriterionfractionalpositionsizingvolatilityestimationdrawdowncontrolpre-registeredevaluationtimeseries
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish a new use for conformal prediction intervals — a calibration method that wraps a forecast in an interval with a stated coverage rate — by reading their width as the volatility scale in fractional Kelly position sizing: widen the interval, shrink the bet; narrow it, bet more. Its central design claim, tested on 2016–2021 data for eight US-listed ETFs under costs, a one-day lag, and a 2.0 leverage cap, is that the simplest scale estimator wins: a slow, unweighted, per-asset rolling conformal quantile beat every locally adaptive variant by 0.7 to 5.3 percentage points of annual growth, because a noisy scale is charged directly to compounded wealth. It also claims that the pooled one-sided downside miscoverage rate is risk information that pays in drawdown reduction (27.7% to 20.3%) rather than growth, beating 40 placebo timings. The sealed 2022–2024 test splits the pair of claims: coverage transferred nearly exactly (0.745 against 0.750), while growth did not — the two configurations earned 8.5% and 7.0% per year, below the passive benchmarks, as the pre-registration predicted.

What carries the argument

The central object is the per-asset conformal half-width $q^{eff} = (q^{roll})^{0.7}(q^{anchor})^{0.3}$, divided by a fixed constant to form $\hat{\sigma}$ and squared into the fractional Kelly denominator. It serves three roles: inverse-variance model averaging, position sizing via $\hat{\mu}/\hat{\sigma}^2$, and (through its downside miscoverage) a book-level leverage dial. The contrast between the two uses is the argument: the scale must be slow and per asset, while the regime signal must be pooled. The hard gross cap (binding 98% of days) keeps the book below the Kelly optimum so cross-sectional ratios, not the overall scale, drive performance.

What would settle it

Run the identical frozen harness and ablation grid on a second asset universe — the paper names this its highest-value follow-up: if any locally adaptive conformal method (adaptive conformal inference, recency weighting, volatility-scaled scores) beats the slow rolling quantile there, the design rule is a development-window artifact. A cheaper test: record the daily log-change variance of each scale estimator, as the paper did for holding-period residuals (0.00343 vs 0.00391); if estimator smoothness does not track the growth ordering, the 'noisy scale is charged to wealth' mechanism fails.

Watch

Extended reading notes

Core claim

Using a 75% conformal interval's width as the Kelly volatility scale, per asset from a 500-day rolling window with anchor shrinkage, feeds $f = \kappa\hat{\mu}/\hat{\sigma}^2$ and compounds at 28.5% annualized net log growth (Sharpe 1.34) on 2016–2021 ETF data under costs and a 2.0 gross cap. Ablations show the growth ordering matches tail-sensitivity: bounded quantile beats linear MAD beats quadratic SD by 2.1 pp at matched leverage; every locally adaptive variant loses. A pooled downside-miscoverage dial cuts drawdown from 27.7% to 20.3% at better Sharpe. The sealed 2022–2024 test kept coverage (0.745 vs 0.750) but not growth (~30% of DEV).

Load-bearing premise

The load-bearing premise is that six years (2016–2021) on eight US-listed ETFs is representative enough that the observed ordering — slow, unweighted, per-asset rolling conformal quantiles beating every faster alternative — is a general property of the Kelly sizing map and not regime luck or search selection; the paper itself says external validity is untested and that the binding 2.0 gross leverage cap and 2020 do enormous structural work.

Editorial extensions

If this is right

  • If the design rule holds, any volatility estimate consumed by a nonlinear sizing map should be chosen for stability, not local accuracy: the paper's growth ordering matches the influence-function ordering of the estimators (bounded quantile beats linear MAD beats quadratic SD).
  • Conformal coverage can transfer to regimes the estimator never saw even when the economic value built on it does not: the sealed test showed 0.745 realized coverage against 0.750 nominal through the 2022 drawdown.
  • The two conformal signals have opposite aggregation requirements: width should be per asset and slow, while the downside miscoverage dial should be pooled across assets and used only to cut leverage, never to raise it.
  • A disclosed large agentic search can carry confirmatory weight only when paired with a sealed window and a pre-registered interpretation rule written before unsealing; the paper's registered prediction (growth shrinks to roughly 30% of DEV) held quantitatively.
  • Practically, the paper concludes the specific book is not investable as configured: at 50 bps costs and 4% financing, lockbox growth falls to roughly +1–5% per year, below every unlevered passive bar.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the slow-stability principle should generalize to any strategy that consumes a scale estimate nonlinearly — vol-targeting, risk parity, option pricing — and is directly testable by rerunning the same slow-versus-adaptive ablation on a second universe, which the paper names as the highest-value follow-up.
  • My inference: the width-as-scale reading predicts a measurable diagnostic the authors did not run — that squared conformal half-width tracks realized variance and that the growth gap between estimators is predictable from the variance of daily changes in the scale; the paper measured this smoothness only for the holding-period case (daily log-change std 0.00343 at $H_s=21$ vs 0.00391 at $H_s=5$).
  • My inference: the DEV/lockbox reversal between the post-hoc equity bar (B7) and its commodity mirror (B8) quantifies how much apparent alpha in an agentic search can be hindsight asset selection — roughly 20 pp of annual growth on DEV — and implies frozen-harness comparisons should include the mirror control by default.
  • My inference: since the drawdown dial's Sharpe benefit vanished out of sample while its drawdown truncation survived, trailing miscoverage is best read as a slow regime indicator whose value is tail insurance, not return enhancement; a longer lockbox window or a stress-universe test would settle whether the DEV placebo result was luck.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces "Conformal Kelly": a fractional-Kelly position-sizing rule whose volatility input is read from a 75% conformal prediction interval for 21-day returns, applied to eight US-listed ETFs. On the 2016-2021 development window the best configuration earns 28.5% annualised net log growth (Sharpe 1.34, max drawdown 27.7%), and a drawdown-dial variant earns 25.8% (Sharpe 1.39, max drawdown 20.3%). The paper's principal development-window finding is a design rule: slow, unweighted, per-asset rolling conformal quantiles outperform locally adaptive conformal methods and a rolling standard deviation as the Kelly denominator. The authors disclose an autonomous LLM-agent search over roughly 200 configurations, seal all data from 2022 onward, and pre-register configurations, benchmarks, and interpretation rules before one lockbox evaluation. The lockbox shows coverage transfers (0.745 vs 0.75) but growth does not (8.5% and 7.0% per year vs passive bars), and the registered risk-adjusted defence fails; the authors conclude that the calibration mechanism transfers but the economic value of the sizing rule does not.

Significance. If the "slow beats fast" design rule were externally valid, the paper would make a substantial contribution to both conformal prediction and Kelly sizing: a scale estimator with a measured interior optimum in adaptation speed, opposite to the time-series conformal-prediction literature's emphasis on local adaptivity. The execution has real strengths: a sealed, pre-registered lockbox with verbatim committed output; explicit disclosure of the agentic search and of multiple failed hypotheses; honest reporting of a known z inconsistency, a post-hoc 4-equity benchmark, and the failure of the registered risk-adjusted defence. The central limitation is equally clear: the design rule is established only on the same DEV window used for the ~200-configuration search, and the lockbox did not re-run the ablation suite, so the estimator ordering has zero out-of-sample confirmation. The only out-of-sample positive is marginal calibration of the intervals, which tests coverage, not whether interval width is a useful Kelly scale. The significance is therefore conditional on an independent validation of the ordering, or on a re-framing of the paper as a primarily negative result.

major comments (4)
  1. [§6.1, §9, §11] The load-bearing claim that 'slow beats fast' is established only on the searched DEV window. §9 reports roughly 200 configurations scored on DEV; §6.1 reports all adaptive-device contrasts as DEV values; Limitation #1 states external validity is untested. The lockbox (§10) tested only the two frozen configurations against passive bars and did not re-run the ablation suite, so the estimator ordering has no out-of-sample support. Because the claim contradicts the standard time-series CP recommendation, it cannot be promoted as a general design principle from a single searched window. A revision should either add an independent validation of the ablation ordering (e.g., a re-registered second universe or a within-sample pseudo-OOS scheme) or explicitly demote the claim to a DEV observation in the title and abstract.
  2. [§5.1, Table 4] The conformal-vs-standard-deviation ablation has no paired confidence intervals, and the headline gap is not exactly matched. The text reports a 3.5 pp leg-3 gap at post-cap gross 1.960 vs 1.922, and the gross-matched replay narrows it to roughly 2.2 pp; the exactly matched leg-2 gap is 2.1 pp, yet no paired interval is given. One standard error on annualised growth is quoted as about 0.095 (§8.1), so the reader cannot judge whether 2.1 pp is distinguishable from zero. In addition, the decomposition of the 3.5 pp gap into updating (+1.5 pp) and functional change (+1.9 pp) is not identified because the frozen-SD cell is missing. Add block-bootstrap paired intervals as listed in Limitation #1, or weaken the corresponding claims.
  3. [§4, §11 Limitation 2] The gross cap is load-bearing for the design rule. §4 reports the book is pinned at the cap on 97.7% of DEV days and states that under the cap anything tilting toward low-volatility assets loses; §11 Limitation 2 concedes that where the cap does not bind the conclusions would probably be different. The slow conformal estimator wins partly by understating sigma for fat-tailed assets, keeping more weight on them under the cap—a regime-specific mechanism. Without checking the estimator ordering at other cap levels, or in an unconstrained variant, the proposed 'stability of the width' principle is confounded with cap saturation. A concrete cap-interaction test, or an explicit restriction of the principle to cap-constrained books, is needed.
  4. [§10, §8.1, Abstract] The lockbox results, honestly reported, refute the economic claim: growth fell to roughly 30% of DEV, both configurations rank last on Sharpe and Calmar, and the drawdown dial fails its pre-registered 'Sharpe no worse' condition. The paper does disclose this, but the abstract's first two paragraphs still lead with DEV Sharpe/Calmar and the placebo p-value. Given the pre-registered interpretation is a partial refutation, the framing should lead with the failed confirmation and identify the paper as a negative-result study for the economic value of the sizing map, with the DEV design rule as hypothesis-generating only. This is not a request to hide the DEV results; it is a request to align emphasis with the confirmatory evidence.
minor comments (5)
  1. [Table L2] Bar B8 is labeled 'control' but on the lockbox it is the best performer on growth. A more neutral label such as 'commodity-block post-hoc book' would avoid implying it is a null control.
  2. [Figure 1] The legend lists 'SPY (true, unlevered 1.0×)' while the caption mentions the harness-capped SPY bar at 0.75×; clarify that the true SPY series is outside the harness and is not scored by prepare.py.
  3. [§5.3] The first row of the contribution table gives the horizon change as 'daily to 21-day forecast horizon'; for completeness, state the horizon value h=21 in the table itself.
  4. [§7] The placebo p-value appears as '0.024' in the abstract and '1/41' in the text; unify the notation (for example, 'p = 1/41 ≈ 0.024' in both places).
  5. [§3] The disclosed z-mismatch is important: z=1.2816 rather than 1.1503 changes the effective Kelly fraction by 24%. Since the paper is precluded from a post-unsealing sensitivity rerun, consider adding a footnote quantifying the implied scale change and reminding readers that the gross cap makes the effect bounded rather than exactly measured.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the conformal-width-as-Kelly-scale construction is an empirically tested rule, not a definitional identity, and the pre-registered lockbox gives independent evaluation.

full rationale

Walking the derivation chain: nonconformity scores are causal absolute residuals, the conformal quantile q_eff is formed from those scores, sigma = q_eff/z, and the sizing map is f = kappa * mu / sigma^2. Each stage consumes the previous stage's output without re-supplying the claim being tested; growth, Sharpe, and drawdown are measured by simulation, not derived from the interval width by construction. The main design rule ('slow, unweighted, per-asset rolling conformal quantiles beat locally adaptive alternatives') comes from paired DEV ablations; the paper itself labels these exploratory: 'These DEV comparisons are exploratory: the grids were searched, and the winning entries are selected values rather than pre-specified hypotheses; only the lockbox carries confirmatory weight.' That is in-sample selection, which the paper discloses and mitigates with a sealed lockbox, not circular reasoning. The lockbox did not re-run the ablation ordering, so the positive design rule lacks out-of-sample confirmation; this is an external-validity limitation, which the paper states: 'External validity is untested; a second universe was forbidden by the frozen harness and is the highest-value follow-up.' Similarly, the gross cap doing 'enormous structural work' and 2020 being a large fraction of DEV are acknowledged regime/selection concerns, not equation-level circularity. No load-bearing self-citation was found: the cited prior work (Vovk, Kelly, Gibbs & Candes, etc.) supplies standard tools, not the paper's conclusion. The closest thing to a circular flavor — using the conformal width both as an ensemble weight and as the sizing denominator — is repeated use of the same computed quantity, not a definitional equivalence between input and output. Under the stated review rules, none of these rises to a circular step; score 0.

Assumptions & free parameters 11 free parameters · 7 assumptions · 0 invented entities

The central result depends on a large set of hyperparameters selected on the same window used to establish the design principle, and on the assumption that one 6-year window on eight ETFs is sufficient to infer a general property of conformal-based sizing. No new physical entity is introduced; the drawdown dial is a strategy component, not an entity.

free parameters (11)
  • Kelly fraction kappa = 0.15 (reduced from 0.25)
    Selected on DEV; §5.3 attributes +0.7pp to pure cap-saturation geometry.
  • nominal miscoverage alpha = 0.25 (swept 0.05-0.45)
    Chosen as the design point; alpha is empirically flat once leverage is controlled, with a shallow unadopted max at 0.35.
  • calibration window W = 500 days (from 250)
    Tuned on DEV; +0.4pp in §5.3, balanced against delaying the first usable date.
  • anchor shrinkage lambda = 0.3
    Tuned on DEV; +0.24pp and reduces max drawdown, §5.3.
  • forecast horizon H and ensemble = 21 days; ensemble {12,16,21,27,34}
    The 21-day horizon is the dominant DEV choice (+7.5pp); the ensemble adds +0.7pp.
  • per-asset position clip = 0.75 (harness MAX_PER_ASSET)
    Fixed by the harness but sits near the interior optimum of the DEV clip sweep by luck, not design.
  • gross leverage cap = 2.0
    Chosen in the harness; the book is pinned at the cap on 98% of DEV days, making many conclusions dependent on this constraint.
  • drawdown dial weight beta = 1.0
    Selected on DEV; the frontier is a step in growth and then flat, with beta=1 reported as the parameter-free point.
  • drawdown dial window M = 21 days
    Selected on DEV; results are noisy across M and would need re-selection on other data.
  • Gaussian constant z = 1.2816 (inconsistent with alpha=0.25)
    Leftover from an earlier alpha=0.20; disclosed as equivalent to a 24% change in kappa.
  • ridge regularization lambda = 10
    Fixed in the harness; the forecaster is deliberately weak, with a drift-only forecast reaching 82% of the final result.
assumptions (7)
  • domain assumption Fractional Kelly under expected log utility is the appropriate objective for a wealth-compounding strategy.
    Used throughout as the sizing map and the evaluation metric (annualized net log growth).
  • domain assumption Conformal interval width divided by a Gaussian quantile is an adequate estimate of the scale sigma that Kelly sizing needs.
    The paper states this directly in §3: it is a heuristic scale estimator, not a split-conformal predictor.
  • domain assumption Realized 75% coverage is the relevant calibration property of the interval for the sizing map.
    The whole evaluation treats coverage against nominal 0.75 as the calibration criterion.
  • domain assumption The frozen Kaggle price snapshot for the eight ETFs is accurate and free of survivorship or look-ahead issues.
    Data comes from a single Kaggle snapshot; the paper validates causality but not the price source against an independent vendor.
  • ad hoc to paper The 2016-2021 DEV window is representative enough to infer a general design principle.
    The paper's key generalizability assumption, acknowledged in §11.1 as untested.
  • domain assumption A flat 5 bps per unit of turnover and a one-day execution lag are adequate proxies for real trading frictions.
    The harness applies only these costs; the paper explicitly disclaims market impact and execution risk.
  • domain assumption A zero risk-free rate is a harmless simplification on the development window.
    The paper states this is benign on 2016-2021 but flatters lockbox risk-adjusted numbers, since cash paid 4-5% in 2022-2024.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Conformal Kelly: Conformal Prediction Intervals as the Scale in Fractional Kelly Position Sizing." pith.science (2026). https://pith.science/paper/JZ4QRTFA

@misc{pith2026260801494,
  author       = {Pith},
  title        = {Pith review of: Conformal Kelly: Conformal Prediction Intervals as the Scale in Fractional Kelly Position Sizing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JZ4QRTFA}},
  note         = {Machine review of arXiv:2608.01494}
}
read the original abstract

Conformal prediction has traditionally been used to quantify prediction uncertainty. We put that uncertainty to a second use, combining a 75% conformal interval with fractional Kelly to size portfolio positions: as the range widens we shrink the position, and as it narrows we grow it. On a six-year development window (2016-2021), with trading costs and strict leverage caps, this compounds at 28.5% annualised net log growth with a Sharpe ratio of 1.34 and a 27.7% maximum drawdown, versus 15.9% for holding the S&P 500 and 21-22% for passive portfolios at the same leverage. Our main development-window finding runs against the literature's advice for conformal prediction on time series. Every tweak that adapts the interval faster to market conditions costs 0.7 to 5.3 points of annual growth; the winner is the simplest method: slow, unweighted, per-asset rolling quantiles. When an interval sizes a position rather than describing one forecast, width stability beats local sharpness. It also beats the textbook standard deviation by 2.1 points at matched leverage. We also implement a risk control: when the intervals miss on the downside far more than their historical rate, we cut leverage. On the development window this cut maximum drawdown from 27.7% to 20.3% while raising the Sharpe ratio, beating all 40 placebo timings (rank-based p = 1/41). These numbers came from an autonomous LLM-agent search over 200 configurations, so we sealed all data from 2022 onward and pre-registered configurations, benchmarks, and interpretation rules before one evaluation. Calibration held (0.745 coverage against 0.750, weakest through 2022); growth did not: the two configurations earned 8.5% and 7.0% per year, below the passive benchmarks, and a pre-registered hindsight benchmark beat them on raw growth while taking a 46% drawdown. All outcomes are reported as pre-registered.

Figures

Figures reproduced from arXiv: 2608.01494 by the authors.

Figure 1
Figure 1. DEV wealth (log scale) and drawdown for Config A and Config B, against the harness-capped [PITH_FULL_IMAGE:figures/full_fig_p023_1.png] view at source ↗
Figure 2
Figure 2. DEV monthly net return distributions, Config A and Config B (n = 72). [PITH_FULL_IMAGE:figures/full_fig_p024_2.png] view at source ↗
Figure 3
Figure 3. Growth-drawdown frontier over the downside-miscoverage weight, with the constant-leverage [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Sigma-source ablation as growth gains over the rolling residual sd; both ablation configurations [PITH_FULL_IMAGE:figures/full_fig_p025_4.png]
Figure 5
Figure 5. Figure 5: Per-year DEV net log growth: Config A, Config B, harness-capped SPY, and the post-hoc 4- [PITH_FULL_IMAGE:figures/full_fig_p025_5.png]
Figure 6
Figure 6. Figure 6: Realized post-cap gross leverage on DEV; Config A dashed and drawn on top, pinned at the cap [PITH_FULL_IMAGE:figures/full_fig_p026_6.png]
Figure 7
Figure 7. Figure 7: Per-asset annualised net contribution to DEV growth, Config A and Config B. [PITH_FULL_IMAGE:figures/full_fig_p026_7.png]
Figure 8
Figure 8. Figure 8: Trailing 63-day realized conformal coverage on the LOCKBOX window, pooled and per asset, [PITH_FULL_IMAGE:figures/full_fig_p027_8.png]
Figure 9
Figure 9. Figure 9: Lockbox wealth (log scale) and drawdown, Config A and Config B (primary variant), with the [PITH_FULL_IMAGE:figures/full_fig_p028_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

3 extracted references · 3 linked inside Pith

  1. [1]

    and Tibshirani, R.J

    Barber, R.F., Candès, E.J., Ramdas, A. and Tibshirani, R.J. (2023). Conformal prediction beyond ex- changeability.TheAnnalsofStatistics 51(2),816–845. Browne, S. (1997). Survival and growth with a liability: optimal portfolio strategies in continuous time. MathematicsofOperationsResearch 22(2),468–493. Chan, J.S. et al. (2024). MLE-bench: Evaluating Machi...

  2. [32]

    arXiv:1905.03222. Vovk, V. and Bendtsen, C. (2018). Conformal predictive decision making.Proceedings of the Seventh WorkshoponConformalandProbabilisticPredictionandApplications ,PMLR91,52–62. Vovk, V., Gammerman, A. and Shafer, G. (2005; 2nd ed. 2022).Algorithmic Learning in a Random World. Springer. Vovk, V., Lindsay, D., Nouretdinov, I. and Gammerman, A...

  3. [34]

    Grossman, S.J

    arXiv:2106.00170. Grossman, S.J. and Zhou, Z. (1993). Optimal investment strategies for controlling drawdowns.Mathe- maticalFinance 3(3),241–276. Jegadeesh, N. and Titman, S. (1993). Returns to Buying Winners and Selling Losers: Implications for StockMarketEfficiency. TheJournalofFinance 48(1),65–91. Jia,Y.andHan,B.(2026). PortfolioSelectionwithAdaptiveCon...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.