REVIEW 3 major objections 5 minor 16 references
Manifold Constrained Conformal Prediction for Spatial Events
T0 review · 3 major / 5 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Conformal prediction for spatial event clouds stays calibrated by scoring with Wasserstein distance and restricting sets to the training data manifold.
desk verdict Solid spherical OT-conformal method with a clean coverage bound and useful constrained sampler; real-data coverage is only extrapolated, which is the main soft spot. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The empirical manifold-restricted conformal set {Y : SSW(μ_Y, f(X)) ≤ τ_α and d_M(Y) ≤ ε}, sampled by a constrained flow whose velocity is the minimum-norm solution of the two linear constraints that drive the score to the conformal quantile and the manifold distance into the ε-tube.
What would settle it
On held-out cyclone seasons or earthquake years, empirical coverage of the constrained sets falls well below 1−α even when ε is set large enough that every training cloud lies inside the tube defined by the other training clouds.
Extended reading notes
Core claim
Spatial point clouds can be embedded as empirical measures so that conformity is measured by spherical sliced Wasserstein distance, turning conformal sets into metric balls in Wasserstein space. Intersecting those balls with an empirical tube around the training clouds yields manifold-restricted sets whose coverage is at least 1−α minus a term that vanishes exponentially in the training size when the tube covers the support. A minimum-norm joint velocity field drives candidates onto the conformal shell and into the tube, producing calibrated, geophysically coherent ensembles.
Load-bearing premise
True event configurations live on a compact set that a finite training sample can cover with a fixed-radius tube; large holes in that cover let coverage fall below the nominal level.
Editorial extensions
If this is right
- Any forecasting system that outputs a density or ensemble over event locations can be post-hoc calibrated without a parametric point-process model.
- Prediction ensembles can be forced onto physically admissible supports such as fault networks or cyclone basins while retaining a finite-sample coverage lower bound.
- Energy distance and manifold distance both improve relative to highest-density-region and generative baselines when the manifold tube is active.
- A data-adaptive rule that sets ε large enough to cover all training clouds makes the coverage gap small in practice.
- Integrating the same constrained flow over α produces a full calibrated predictive distribution represented as an ensemble.
Reading between the lines
- The same tube-plus-flow construction should transfer to other variable-cardinality spatial catalogs—wildfire ignitions, disease case clusters—whenever a training atlas of plausible patterns exists.
- Hard geometric constraints (known fault geometry, ocean-only genesis) could replace or tighten the empirical tube and further shrink both the coverage gap and off-manifold mass.
- When the base forecast is badly misspecified, the manifold constraint can act as a practical regularizer that keeps risk maps usable even if the raw predictive density is diffuse.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a conformal prediction procedure for variable-cardinality spatial event clouds on S² (tropical cyclone genesis, earthquakes). Point clouds are embedded as empirical measures and scored by spherical sliced Wasserstein (SSW) distance, yielding ambient SSW balls with the usual split-conformal guarantee. These balls are intersected with an empirical manifold tube {d_M ≤ ε} around the training targets; Proposition 3.1 supplies a coverage lower bound 1−α−N_ε(1−β_ε)^{n_0}. Because the intersected set is intractable, a minimum-norm constrained velocity field is derived that drives candidates to the boundary {S=τ_α, d_M≤ε}, producing ensembles and a flow-based predictive distribution. Synthetic experiments recover near-nominal coverage under a data-adaptive ε rule and improve energy and manifold distance relative to HDR and generative baselines; the same skill metrics are reported for cyclone and California earthquake data.
Significance. Uncertainty quantification for discrete spatial event configurations is practically important and poorly served by residual-based conformal methods. The combination of an SSW conformity score on P_2(S²), an explicit empirical manifold restriction with a proved coverage gap, and a constrained flow sampler is a coherent and novel contribution. Strengths include a clean proof of Proposition 3.1, a transparent leave-one-out-style ε rule that is shown to restore coverage on four synthetic processes, and consistent gains in energy distance and manifold distance against both HDR-style and generative baselines. If the real-data coverage claims can be substantiated (or appropriately qualified), the method would be a useful post-hoc calibration tool for hazard ensembles.
major comments (3)
- [Abstract, §§4.3–4.4] Abstract and §1 claim “near-nominal coverage … on … tropical cyclone genesis and earthquake occurrences,” yet §§4.3–4.4 report only energy distance and manifold distance (Tables 2–3). Coverage of C_α is the indicator 1{S(Y_test)≤τ_α and d_M(Y_test)≤ε} and can be evaluated without the flow; it is never shown for the 2015–2025 cyclone or 2016–2025 earthquake years. With n_0 only ~9–15 yearly aggregates, the exponential gap term in Prop. 3.1 need not be negligible. Either report empirical coverage on the held-out seasons or temper the abstract/intro claim to match what is demonstrated (synthetic coverage + real-data skill).
- [§3.2 Prop. 3.1, §4.1] The data-adaptive ε rule of §4.1 (smallest ε covering all training clouds) is validated only under n_0=100 synthetic clouds (Figs. 2–3). For the geophysical tasks the same rule is applied with far smaller n_0. Prop. 3.1 makes the dependence on n_0 and β_ε explicit; the manuscript should quantify or bound the residual gap for the cyclone/earthquake training sizes (or show leave-one-year-out coverage on the calibration window) so that readers can judge whether “near-nominal” remains plausible.
- [§3.1, §§4.3–4.4] Exchangeability of yearly event clouds is assumed throughout. Tropical-cyclone and earthquake series exhibit temporal dependence, non-stationarity, and possible climate-driven drift. The paper never discusses whether split conformal remains approximately valid under these violations, nor does it apply any of the existing non-exchangeable corrections (e.g., Barber et al. 2023, already cited). A short sensitivity check or explicit caveat is needed before the real-data coverage language can stand.
minor comments (5)
- [References / body] Encoding artifacts appear throughout (“V ovk”, “Peyr´e”, “Cand `es”). Clean for camera-ready.
- [Table 1] Table 1 reports a negative energy distance for CVF on IHPP-3 (−0.18). While possible for U-statistic estimators, a one-sentence note that the estimator is unbiased for a non-negative quantity would avoid confusion.
- [Figure 4] Figure 4 caption says “5 methods” but the text discusses VF, CVF and several baselines; label panels consistently with the method codes M1–M5, VF, CVF used in the tables.
- [§3.3 Eq. (20)] Eq. (20) defines a predictive distribution by integrating ν_{x,α} over α; no diagnostic is given that the resulting ensemble is calibrated as a full distribution (only set coverage of individual levels is discussed). A brief remark on what is and is not claimed would help.
- [§4 / Appendix E.7] Appendix E.7 timing table is useful; consider moving a one-line summary of wall-clock cost into the main numerical section so practitioners can gauge feasibility.
Circularity Check
No significant circularity: coverage bound and manifold restriction are independently derived from exchangeability plus a covering argument; only a minor non-load-bearing self-citation of the coauthor's concurrent flow sampler appears.
-
self citation load bearing
[Section 3.3 (and Introduction)]
"Building on the flow-based conformal predictive distribution framework of (Harris, 2026), we introduce a constrained gradient flow that simultaneously targets the SSW prediction set boundary and the data manifold…"
The sampling procedure that represents the non-tractable set C_α as an ensemble is obtained by modifying a concurrent method of co-author Harris. The modification itself is derived in Appendix A, and the citation is not used to establish coverage, the lower bound of Prop. 3.1, or the energy/manifold-distance comparisons; it is therefore only a minor, non-load-bearing self-citation.
full rationale
The ambient conformal guarantee (Eq. 10) is the standard split-conformal argument under exchangeability of the SSW scores. The manifold-restricted set is defined by intersection with an empirical tube (Eq. 13); Proposition 3.1 then supplies an explicit, non-tautological lower bound 1-α-N_ε(1-β_ε)^n0 obtained by a finite measurable cover of M and a union bound, with the proof given in full in Appendix B. The data-adaptive rule for ε (smallest value that covers the training clouds) is a transparent leave-one-out construction on D0 and is not optimized against the reported test energy or manifold distances. The constrained velocity field (Eq. 19) is derived from first principles via the Moore-Penrose solution of the two linear constraints in Appendix A; the only external reference is the unconstrained flow of Harris (2026), which is used merely as a starting point for the sampling procedure and is not invoked to justify coverage, the lower bound, or the skill comparisons. Energy distance is an external proper scoring rule evaluated on held-out clouds. Consequently no claimed guarantee or numerical result reduces by construction to a fitted constant or to an unverified self-citation chain.
Assumptions & free parameters
free parameters (4)
- ε (manifold tolerance)
- λ1, λ2 (flow decay rates)
- L (number of SSW slices)
- α (miscoverage level)
assumptions (4)
- standard math Calibration and test pairs are exchangeable (or i.i.d.).
- domain assumption The data-generating distribution is supported on a compact manifold M ⊂ Y.
- standard math SSW metrizes the same weak topology as W2 on P2(S2).
- ad hoc to paper The minimum-norm velocity field of the under-determined linear system drives trajectories to the boundary of the intersected set.
invented entities (2)
-
Empirical manifold tube {Y : d_M(Y) ≤ ε}
-
Constrained velocity field v(y) = −J^+ c
Cite this review
Pith. "Pith review of Manifold Constrained Conformal Prediction for Spatial Events." pith.science (2026). https://pith.science/paper/YNEOV7KY
@misc{pith2026260710008,
author = {Pith},
title = {Pith review of: Manifold Constrained Conformal Prediction for Spatial Events},
year = {2026},
howpublished = {\url{https://pith.science/paper/YNEOV7KY}},
note = {Machine review of arXiv:2607.10008}
}
read the original abstract
We introduce a new conformal prediction method that constructs calibrated prediction sets over collections of spatial events, such as tropical cyclone genesis and earthquake locations. Forecasting natural hazards has become increasingly important, due to their significant economic impact, and quantifying the uncertainty of predictions is critical for accurate risk assessment. Our approach works by representing spatial point clouds as empirical measures so that we can score them using (sliced) Wasserstein distance, then constraining the resulting distribution-valued prediction set to be supported only near the training data manifold. We derive a coverage lower bound for the intersected sets and show that, in practice, this gap can be made small through a simple data-adaptive selection criterion. Because the resulting set is not analytically tractable, we introduce a modified flow-based sampling procedure, which allows us to represent and apply these prediction sets in practice as ensembles. Numerical experiments on synthetic data, tropical cyclone genesis, and earthquake occurrences show that our method achieves near-nominal coverage, with significantly lower energy distance and manifold distance than highest predictive density region (HDR) baselines along with generative model baselines.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Knapp, Carl J
Preprint July 14, 2026 Jennifer Gahtan, Kenneth R. Knapp, Carl J. III Schreck, Howard J. Diamond, James P. Kossin, and Michael C. Kruk. International best track archive for climate stewardship (IBTrACS) project, version 4.01,
2026
-
[2]
Dataset. DOI: 10.25921/82ty-9e16. Available at https://www.ncei.noaa.gov/access/metadata/landing-page/bin/iso? id=gov.noaa.ncdc%3AC01552. Accessed 2026-02-23. Tilmann Gneiting and Adrian E. Raftery. Strictly proper scoring rules, prediction, and es- timation.Journal of the American Statistical Association, 102(477):359–378,
-
[3]
Flow-Based Conformal Predictive Distributions
doi: 10.1198/016214506000001437. Trevor Harris. Flow-based conformal predictive distributions.arXiv preprint arXiv:2602.07633,
-
[4]
doi: 10.1080/00031305.1996.10474359. Thomas H Jordan, Yun-Tai Chen, Paolo Gasparini, Raul Madariaga, Ian Main, Warner Marzocchi, Gerassimos Papadopoulos, Gennady Sobolev, Koshun Yamaoka, and Jochen Zschau. Operational earthquake forecasting: State of knowledge and guidelines for utilization.Annals of Geophysics, 54(4):315–391,
-
[5]
Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114,
Diederik P Kingma and Max Welling. Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114,
-
[6]
Multivariate conformal prediction using optimal transport.arXiv preprint arXiv:2502.03609,
Michal Klein, Louis Bethune, Eugene Ndiaye, and Marco Cuturi. Multivariate conformal prediction using optimal transport.arXiv preprint arXiv:2502.03609,
-
[7]
Conformal graph prediction with Z-Gromov–Wasserstein distances.arXiv preprint arXiv:2603.02460,
Gabriel Melo, Thibaut de Saivre, Anna Calissano, and Florence d’Alch ´e-Buc. Conformal graph prediction with Z-Gromov–Wasserstein distances.arXiv preprint arXiv:2603.02460,
-
[8]
Neural spatiotemporal point processes: Trends and challenges.arXiv preprint arXiv:2502.09341,
Sumantrak Mukherjee, Mouad Elhamdi, George Mohler, David A Selby, Yao Xie, Sebastian V ollmer, and Gerrit Grossmann. Neural spatiotemporal point processes: Trends and challenges.arXiv preprint arXiv:2502.09341,
Show all 16 references
-
[9]
QiFeng Qian, XiaoJing Jia, and Yanluan Lin
doi: 10.1016/S0304-4149(97)00028-8. QiFeng Qian, XiaoJing Jia, and Yanluan Lin. Reduced tropical cyclone genesis in the future as predicted by a machine learning model.Earth’s Future, 10(2):e2021EF002455,
-
[10]
URL https://agupubs.onlinelibrary
doi: https://doi.org/10.1029/2021EF002455. URL https://agupubs.onlinelibrary. wiley.com/doi/abs/10.1029/2021EF002455. e2021EF002455 2021EF002455. Preprint July 14, 2026 Danilo Rezende and Shakir Mohamed. Variational inference with normalizing flows. InInternational conference ...
2026 doi
-
[11]
Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, and Kenji Fukumizu
doi: 10.1002/wics.1375. Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, and Kenji Fukumizu. Equivalence of distance-based and rkhs-based statistics in hypothesis testing.The annals of statistics, pp. 2263– 2291,
-
[12]
Optimal transport-based conformal prediction
Gauthier Thurin, Kimia Nadjahi, and Claire Boyer. Optimal transport-based conformal prediction. arXiv preprint arXiv:2501.18991,
-
[13]
C´edric Villani.Optimal Transport: Old and New, volume
Accessed 2026-04-24. C´edric Villani.Optimal Transport: Old and New, volume
2026
-
[14]
doi: 10.4401/ag-4839
ISSN 1593-5213. doi: 10.4401/ag-4839. URL http://dx.doi. org/10.4401/ag-4839. Zihao Zhou, Xingyi Yang, Ryan Rossi, Handong Zhao, and Rose Yu. Neural point process for learning spatiotemporal event dynamics. InLearning for Dynamics and Control Conference, pp. 777–789,
-
[15]
FixX n+1 and write S(y) =S(Y, f(X n+1)), d(y) =d M(Y), wherey∈R 3|Y| is the vectorized point cloud
A DERIVATION OF THEVELOCITYFIELD We derive the minimum-norm velocity field used in Section 3.3. FixX n+1 and write S(y) =S(Y, f(X n+1)), d(y) =d M(Y), wherey∈R 3|Y| is the vectorized point cloud. The desired scalar dynamics are d dt S(y(t)) =−λ 1{S(y(t))−τ α}, d dt d(y(t)) =−λ...
2026
-
[16]
, πA(m) be an ordering such that DπA(1) ≥D πA(2) ≥ · · · ≥D πA(m)
Dij = Pij Aij , then as in HDR, keep the smallest number of cells with cumulative probability of at least1−α Let πA(1), . . . , πA(m) be an ordering such that DπA(1) ≥D πA(2) ≥ · · · ≥D πA(m). Then define KMA(α) = min ( k: kX ℓ=1 PπA(ℓ) ≥1−α ) , and the minimum-area region is ...
2026
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.