Pith. sign in

REVIEW 2 major objections 4 minor 62 references

From Jumps to Signatures: a Generative Method for Temporal Point Processes

T0 review · 2 major / 4 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read An interarrival embedding turns event sequences into continuous paths so signature methods can generate and score whole TPP trajectories.

desk verdict Solid pathwise fix for signatures on TPPs: new embedding with real proofs, first signature generative model, and three justified metrics; determinacy gap is real but already flagged and not load-bearing for the empirical claim. read the letter →

arxiv 2607.06652 v1 pith:XLZSXWDE submitted 2026-07-07 cs.LG stat.ML

classification cs.LGstat.ML
keywords temporalpointprocessesroughpathsignaturesinterarrivalembeddinggenerativemodelssignatureWassersteindistancecountingpathsdistributionaldiscrepancy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Temporal point processes produce sequences of event times, but their natural sample paths are discontinuous step functions. Signature methods from rough path theory give universal features and distributional characterisations for continuous paths of bounded variation, so those tools have been largely out of reach for event data. This paper builds a stable, injective lift—the interarrival embedding—that maps each counting path to a continuous piecewise-linear path that records successive waiting times, then applies the signature on the time-augmented curve. With that lift in hand, the authors train SIGTPP, a generative model that minimises a global signature-Wasserstein loss between complete generated and observed trajectories rather than a sum of per-event losses. The same path metric also yields three mathematically justified distributional discrepancies (energy, Wasserstein-1, and signature-Wasserstein-1) for evaluating generative TPPs. On synthetic and real-world benchmarks, SIGTPP obtains the best average rank across eight complementary metrics and improves relative scores against every baseline by at least 19 percent on average.

What carries the argument

The interarrival embedding Φ: it interpolates each counting path so that the value at successive event times equals the successive interarrival durations, producing a continuous piecewise-linear path of bounded variation on which the (time-augmented) signature is well-defined and injective.

What would settle it

Find a pair of distinct TPP laws whose interarrival embeddings have identical expected truncated signatures at every finite degree (or whose empirical Sig-W1 distance collapses to zero while energy or W1 distances remain large), which would show that the signature loss fails to separate counting-path distributions.

Watch

Extended reading notes

Core claim

The interarrival embedding is a Lipschitz, injective map from the space of unit-jump counting paths into continuous paths of bounded variation, with a constructive inverse that is Hölder continuous under a mild separation of interarrival times. This lift makes the expected-signature characterisation available for TPP laws and supports SIGTPP, the first signature-based generative model for temporal point processes trained with a single path-level Sig-W1 loss on complete trajectories.

Load-bearing premise

That the laws of the embedded continuous paths satisfy the infinite-radius moment condition needed for the expected signature to uniquely determine the distribution, so that matching truncated signatures separates distinct event-sequence laws.

Editorial extensions

If this is right

  • Signature methods can be applied to discrete event sequences without treating them as continuous time series or forcing a parametric intensity.
  • Generative TPP training can target a single global discrepancy between complete trajectories instead of a sum of local conditional losses.
  • Energy, W1, and Sig-W1 distances on counting paths become rigorously justified evaluation metrics for generative TPPs.
  • Pointwise errors such as MAE and MSE are shown to favour deterministic regressors and should not be primary metrics for generative quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same embedding could serve as a drop-in feature map for supervised or forecasting models on event sequences, not only for generative training.
  • Because the lift is constructive and invertible, one could decode signature-space interventions back into event times, enabling controllable generation of sequences with prescribed higher-order statistics.
  • Marked or high-dimensional TPPs may require a carefully chosen multi-dimensional analogue of the interarrival lift if signature dimension is not to explode.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper develops a pathwise framework for generative modelling and evaluation of temporal point processes. It introduces the interarrival embedding Φ, a Lipschitz injective lift from càdlàg counting paths to continuous bounded-variation paths (Theorem 6), so that the signature transform and expected-signature tools apply. Building on Φ, SIGTPP is trained by matching truncated expected signatures of complete embedded trajectories (Sig-W1 loss, Eq. 4) rather than per-event conditional losses. The same counting-path metric d_N is used to justify three distributional discrepancies (energy, W1, Sig-W1) for evaluation. Empirically, SIGTPP is compared to VAE, DDPM, WGAN, GAMMA and DETER on four synthetic and five real datasets under eight metrics, reporting best average ranks, competitive pairwise wins, and relative-score gains of at least 19% against every baseline.

Significance. If the claims hold, the work supplies a usable bridge from rough-path signatures to discrete event sequences and a principled alternative to local likelihood or adversarial objectives for generative TPPs. The embedding proofs (Appendix B), the metric d_N with its OT interpretation, and the explicit energy/W1/Sig-W1 evaluation suite are concrete contributions that the community can reuse even if SIGTPP is not adopted as a default generator. Strengths include detailed stability proofs, bootstrap standard errors, multi-metric evaluation, truncation ablations, and released code. The main theoretical soft spot is the unverified infinite-radius moment condition needed for full expected-signature determinacy of the pushforwards; the paper already flags this in Section 3.4 and Appendix E, so the empirical ranking claims do not rest on it.

major comments (2)
  1. Section 3.4 and Theorem 3: the argument that Sig-W1 separates distinct laws on N relies on the pushforward laws Φ♯P satisfying the infinite-radius moment condition of the expected-signature determinacy theorem. The paper only shows injectivity of Φ plus continuous-path determinacy would imply separation for large enough M, and Appendix E correctly notes that such conditions are hard to verify even for continuous processes and are not checked here. This does not invalidate Lipschitz/injectivity of Φ or the well-definedness of E and W1 under d_N, but it does leave the theoretical status of Sig-W1 as a separating metric incomplete. A short discussion of what can be said without the radius condition (e.g., that Sig-W1 is always a pseudometric, and when it is positive in practice) would make the claim precise.
  2. Section 4.1 and Tables 1–2: model selection uses a rank aggregate over validation diagnostics while checkpoints are chosen by validation L_log(τ). Because several reported metrics (including L_log(τ) and L_λ) enter both selection and evaluation, and because CRPS consistently favours the conditional baselines, it would strengthen the central empirical claim to report a sensitivity check under an alternative selection criterion (e.g., validation W1 or Sig-W1 only) or to hold out one metric family from selection. The current protocol is transparent but leaves open whether the average-rank advantage is partly selection-driven.
minor comments (4)
  1. Definition 4 / Figure 1: the figure caption and surrounding text are clear, but a one-line statement that the signature is applied to the time-augmented path t ↦ (t, Φ(η)_t) would help readers who skip the paragraph after Eq. (2).
  2. Table 3: the ablation is only on TX and SO; a sentence on whether M=3 was also preferred on the synthetic suite would make the truncation choice more uniform.
  3. Appendix E: the linear-interpolation lift is listed as a limitation; a brief pointer to why step or other schemes were not used (beyond the standard signature literature) would be useful for follow-up work.
  4. Notation: N is used both for the space of counting paths and for a random counting path; a consistent distinction (e.g., script N vs. N) would reduce occasional ambiguity in Section 3.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: embedding stability, d_N metric, and Sig-W1 loss are derived from first-principles definitions and external rough-path theorems without reducing to fitted inputs or self-citation chains.

full rationale

The core derivation chain is self-contained. Definition 4 constructs the interarrival embedding Φ by piecewise-linear interpolation of interarrival times on the augmented grid (Convention 1); Theorem 6 then proves Lipschitz continuity (via Lemmas S6–S9 and the L1 triangle inequality), injectivity/bijectivity (via the constructive backshift inverse Ψ of Lemmas S14–S16), and Hölder stability of the inverse on the sieve N_δ (via recursive error propagation and compactness separation in Lemmas S18–S29). These are direct analytic arguments on (N,d_N) and (C,d_1); none defines the target property in terms of itself. The counting-path metric d_N (Definition 5) is introduced as the L1 integral of path differences, shown to be a genuine metric (Proposition S3), of negative type by isometric embedding into L1 (Theorem S4, citing external Bretagnolle et al.), and equal to the Wasserstein-1 distance on T_max-padded Dirac measures (Appendix A.3). Energy, W1 and Sig-W1 are then the standard lifts of this ground metric (Section 3.4); their well-definedness follows immediately and does not presuppose the empirical ranking claims. SIGTPP’s training objective (Eq. 4) is the Euclidean distance between expected truncated signatures of Φ-embedded paths; this is an optimisation target, not a prediction forced by a prior fit. Theorem 3 (determinacy of the expected signature) is cited from Chevyrev & Lyons (external); the paper only claims that injectivity of Φ plus the continuous-path moment condition would separate laws, and explicitly flags in Appendix E that the radius-of-convergence condition is unverified. No load-bearing uniqueness theorem is imported from the present authors, no ansatz is smuggled via self-citation, and no fitted parameter is renamed a prediction. Empirical ranks and relative scores are independent experimental outcomes. The derivation therefore contains no circular step of the enumerated kinds.

Assumptions & free parameters 3 free parameters · 4 assumptions · 2 invented entities

The central claims rest on standard rough-path and optimal-transport facts plus one new embedding whose properties are proved under mild separation. Free parameters are ordinary neural-network and truncation choices; no new physical entities are postulated. The main external reliance is the expected-signature determinacy theorem for continuous paths of bounded variation.

free parameters (3)
  • signature truncation degree M = 3 (eval), 8 (train)
    Chosen by ablation (Table 3) and fixed at 3 for evaluation / 8 for training; controls feature dimension and loss fidelity.
  • LSTM hidden size H and decoder architecture = 16 or 32
    Tuned on validation rank aggregate; standard capacity hyperparameters of the autoregressive generator.
  • learning rates and teacher-forcing / detach flags
    Grid-searched per dataset (Table S2); affect optimization trajectory but not the mathematical claims.
assumptions (4)
  • standard math Expected signature determines the law of continuous bounded-variation paths when the radius of convergence is infinite (Chevyrev-Lyons determinacy).
    Invoked in Section 3.4 to argue that Sig-W1 separates distinct pushforward laws after embedding.
  • standard math L1 is of negative type, hence so is any isometric subspace (including counting paths under d_N).
    Used in Appendix A.1 to justify the energy distance as a pseudometric.
  • domain assumption Event sequences are simple (strictly increasing times, unit jumps, no jump at T_max) and recorded at finite precision so a positive interarrival lower bound δ exists.
    Required for Hölder continuity of the inverse embedding (Theorem 6 part 3 and sieve N_δ).
  • ad hoc to paper Linear interpolation of interarrival heights is an adequate continuous lift for signature methods.
    Chosen as the standard signature practice; other schemes (step, spline) are possible but not analyzed.
invented entities (2)
  • interarrival embedding Φ
    purpose: Injective Lipschitz lift from càdlàg counting paths to continuous BV paths so signatures apply.
    Defined in Definition 4; stability proved in Theorem 6; no external independent evidence beyond the paper's own analysis and experiments.
  • counting-path metric d_N
    purpose: Genuine metric on N that coincides with W1 of padded event-time measures and induces energy and W1 discrepancies.
    Definition 5; shown equivalent to optimal-transport distance in Appendix A.3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Jumps to Signatures: a Generative Method for Temporal Point Processes." pith.science (2026). https://pith.science/paper/XLZSXWDE

@misc{pith2026260706652,
  author       = {Pith},
  title        = {Pith review of: From Jumps to Signatures: a Generative Method for Temporal Point Processes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XLZSXWDE}},
  note         = {Machine review of arXiv:2607.06652}
}
read the original abstract

Rough path signatures are a universal feature map for continuous paths and, via the expected signature, characterise path distributions. These guarantees do not directly extend to cadlag paths of Temporal Point Processes (TPPs), limiting the use of signature methods for event sequences. Furthermore, neural TPP models, including recent generative approaches, optimise per-event objectives with no global sequence-level loss, while evaluation of variable-length event sequences lacks distributional discrepancy measures. This paper proposes a common pathwise framework for addressing these limitations. We introduce the interarrival embedding, a stable, injective lift from jump paths to continuous paths of bounded variation, extending signature methods to discrete event sequences. Our theoretical contributions give rise to sigTPP, the first signature-based generative model for TPPs, trained using a path-level loss on complete trajectories. We further analyse the space of counting paths and derive three distributional discrepancies, providing mathematically justified tools for evaluating generative TPP models. Across synthetic and real-world datasets, sigTPP achieves the best average rank based on eight complementary metrics, outperforms or is within a standard error of the strongest baseline in 64% of the dataset-metric pairs, and according to a relative score, improves against every baseline by at least 19% on average.

Figures

Figures reproduced from arXiv: 2607.06652 by the authors.

Figure 1
Figure 1. Illustration of the interarrival em￾bedding Φ (Definition 4) for a sampled Pois￾son process with λ= 1 and Tmax = 7. Equivalently, Φ(η) is the unique continuous function that is affine on every interarrival time and satisfies Φ(η)tk = τk, k≤m+1 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. SIGTPP’s optimisation. A counting path sampled from µθ is embedded with Φ to a continuous path, mapped to its truncated signature in T M(R 2 ), and summarised by its expectation under the model. The loss (Sig-W1 ) compares this against the expected signature of the embedded targets’ paths, opposed to the usual per-event loss. Loss gradient flows back through µθ. signature property, any 1-Lipschitz test function can … view at source ↗
Figure 3
Figure 3. (a) Quantile-quantile (QQ) plots (generated vs. reference sample quantiles on x-axis and [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Relative scores with respect to DETER, on all datasets (see Table S6 for the actual values). The leftmost group shows the average relative score across all metrics, followed by the metric-wise scores for each metric (see Appendix C for metric definitions). A lower scor…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

62 extracted references · 62 canonical work pages

  1. [1]

    Maddix, Hao Wang, Michael W

    Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Olek- sandr Shchur, Syama Sundar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C. Maddix, Hao Wang, Michael W. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke-Schneider, and Yuyang Wang. Chronos: Learning the Languag...

  2. [2]

    Towards Principled Methods for Training Generative Adversarial Networks

    Martin Arjovsky and Léon Bottou. Towards Principled Methods for Training Generative Adversarial Networks. InProceedings of the 5th International Conference on Learning Representations, 2017

  3. [3]

    Hawkes Processes in Finance.Market Microstructure and Liquidity, 1(1), 2015

    Emmanuel Bacry, Iacopo Mastromatteo, and Jean-François Muzy. Hawkes Processes in Finance.Market Microstructure and Liquidity, 1(1), 2015

  4. [4]

    Bretagnolle, D

    J. Bretagnolle, D. Dacunha-Castelle, and J. L. Krivine. Lois stables et espaces Lp. InSymposium on Probability Methods in Analysis, pages 48–54, 1967. doi: 10.1007/BFb0061106

  5. [5]

    Deep Continuous-Time State-Space Models for Marked Event Sequences

    Yuxin Chang, Alex James Boyd, Cao Xiao, Taha Kass-Hout, Parminder Bhatia, Padhraic Smyth, and Andrew Warrington. Deep Continuous-Time State-Space Models for Marked Event Sequences. In Advances in Neural Information Processing Systems, 2025

  6. [6]

    Springer, 2026

    Ilya Chevyrev and Andrey Kormilitzin.A Primer on the Signature Method in Machine Learning, pages 3–64. Springer, 2026

  7. [7]

    Characteristic Functions of Measures on Geometric Rough Paths.The Annals of Probability, 44(6):4049–4082, 2016

    Ilya Chevyrev and Terry Lyons. Characteristic Functions of Measures on Geometric Rough Paths.The Annals of Probability, 44(6):4049–4082, 2016

  8. [8]

    Signature Moments to Characterize Laws of Stochastic Processes

    Ilya Chevyrev and Harald Oberhauser. Signature Moments to Characterize Laws of Stochastic Processes. Journal of Machine Learning Research, 23(176):7928–7969, 2022

Show all 62 references
  1. [9]

    Learning to simulate realistic limit order book markets from data as a World Agent

    Andrea Coletta, Aymeric Moulin, Svitlana Vyetrenko, and Tucker Balch. Learning to simulate realistic limit order book markets from data as a World Agent. InProceedings of the Third ACM International Conference on AI in Finance, pages 428–436, 2022

  2. [10]

    Limit Order Book Simulation with Generative Adversarial Networks.SSRN working paper 4512356, 2023

    Rama Cont, Mihai Cucuringu, Jonathan Kochems, and Felix Prenzel. Limit Order Book Simulation with Generative Adversarial Networks.SSRN working paper 4512356, 2023

  3. [11]

    Universal approximation theorems for continuous functions of càdlàg paths and Lévy-type signature models.Finance and Stochastics, 29(2): 289–342, 2025

    Christa Cuchiero, Francesca Primavera, and Sara Svaluto-Ferro. Universal approximation theorems for continuous functions of càdlàg paths and Lévy-type signature models.Finance and Stochastics, 29(2): 289–342, 2025

  4. [12]

    Daley and David Vere-Jones.An Introduction to the Theory of Point Processes: Volume I: Elementary Theory and Methods

    Daryl J. Daley and David Vere-Jones.An Introduction to the Theory of Point Processes: Volume I: Elementary Theory and Methods. Springer, 2003

  5. [13]

    A Decoder-Only Foundation Model for Time-Series Forecasting

    Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. A Decoder-Only Foundation Model for Time-Series Forecasting. InProceedings of the 41st International Conference on Machine Learning, pages 10148–10167, 2024

  6. [14]

    Brodsky, and Stephan Günnemann

    Kelian Dascher-Cousineau, Oleksandr Shchur, Emily E. Brodsky, and Stephan Günnemann. Using Deep Learning for Flexible and Scalable Earthquake Forecasting.Geophysical Research Letters, 50(17): e2023GL103909, 2023

  7. [15]

    Springer, 1997

    Michel Marie Deza and Monique Laurent.Geometry of Cuts and Metrics. Springer, 1997

  8. [16]

    Recurrent Marked Temporal Point Processes: Embedding Event History to Vector

    Nan Du, Hanjun Dai, Rakshit Trivedi, Utkarsh Upadhyay, Manuel Gomez-Rodriguez, and Le Song. Recurrent Marked Temporal Point Processes: Embedding Event History to Vector. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages...

  9. [17]

    Fleming and John J

    Philip J. Fleming and John J. Wallace. How Not to Lie with Statistics: The Correct Way to Summarize Benchmark Results.Communications of the ACM, 29(3):218–221, 1986

  10. [18]

    Folland.Real Analysis: Modern Techniques and Their Applications

    Gerald B. Folland.Real Analysis: Modern Techniques and Their Applications. Wiley, 1999. 10

  11. [19]

    Friz and Nicolas B

    Peter K. Friz and Nicolas B. Victoir.Multidimensional Stochastic Processes as Rough Paths: Theory and Applications. Cambridge University Press, 2010

  12. [20]

    Point process models for COVID-19 cases and deaths.Journal of Applied Statistics, 50(11-12):2294–2309, 2023

    Álvaro Gajardo and Hans-Georg Müller. Point process models for COVID-19 cases and deaths.Journal of Applied Statistics, 50(11-12):2294–2309, 2023

  13. [21]

    Strictly Proper Scoring Rules, Prediction, and Estimation.Journal of the American Statistical Association, 102(477):359–378, 2007

    Tilmann Gneiting and Adrian E Raftery. Strictly Proper Scoring Rules, Prediction, and Estimation.Journal of the American Statistical Association, 102(477):359–378, 2007

  14. [22]

    Large Language Models Are Zero-Shot Time Series Forecasters

    Nate Gruver, Marc Finzi, Shikai Qiu, and Andrew Gordon Wilson. Large Language Models Are Zero-Shot Time Series Forecasters. InAdvances in Neural Information Processing Systems, 2023

  15. [23]

    Uniqueness for the Signature of a Path of Bounded Variation and the Reduced Path Group.Annals of Mathematics, 171(1):109–167, 2010

    Ben Hambly and Terry Lyons. Uniqueness for the Signature of a Path of Bounded Variation and the Reduced Path Group.Annals of Mathematics, 171(1):109–167, 2010

  16. [24]

    Alan G. Hawkes. Spectra of Some Self-Exciting and Mutually Exciting Point Processes.Biometrika, 58 (1):83–90, 1971

  17. [25]

    Personalized Dynamic Treatment Regimes in Continuous Time: A Bayesian Approach for Optimizing Clinical Decisions with Timing.Bayesian Analysis, 17(3):849–878, 2022

    William Hua, Hongyuan Mei, Sarah Zohar, Magali Giral, and Yanxun Xu. Personalized Dynamic Treatment Regimes in Continuous Time: A Bayesian Approach for Optimizing Clinical Decisions with Timing.Bayesian Analysis, 17(3):849–878, 2022

  18. [26]

    EventFlow: Forecasting Temporal Point Processes with Flow Matching

    Gavin Kerrigan, Kai Nelson, and Padhraic Smyth. EventFlow: Forecasting Temporal Point Processes with Flow Matching. InProceedings of the 29th International Conference on Artificial Intelligence and Statistics, 2026

  19. [27]

    Deep Signature Transforms

    Patrick Kidger, Patric Bonnier, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons. Deep Signature Transforms. InAdvances in Neural Information Processing Systems, pages 3082–3092, 2019

  20. [28]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. InProceedings of the 3rd International Conference on Learning Representations, 2015

  21. [29]

    Learning Temporal Point Processes via Reinforcement Learning

    Shuang Li, Shuai Xiao, Shixiang Zhu, Nan Du, Yao Xie, and Le Song. Learning Temporal Point Processes via Reinforcement Learning. InAdvances in Neural Information Processing Systems, pages 10804–10814, 2018

  22. [30]

    Sig- Wasserstein GANs for Conditional Time Series Generation.Mathematical Finance, 34(2):622–670, 2024

    Shujian Liao, Hao Ni, Marc Sabate-Vidales, Lukasz Szpruch, Magnus Wiese, and Baoren Xiao. Sig- Wasserstein GANs for Conditional Time Series Generation.Mathematical Finance, 34(2):622–670, 2024

  23. [31]

    Haitao Lin, Lirong Wu, Guojiang Zhao, Pai Liu, and Stan Z. Li. Exploring Generative Neural Temporal Point Process.Transactions on Machine Learning Research, 2022

  24. [32]

    Haitao Lin, Cheng Tan, Lirong Wu, Zhangyang Gao, Zicheng Liu, and Stan Z. Li. An Extensive Survey With Empirical Studies on Deep Temporal Point Processes.IEEE Transactions on Knowledge and Data Engineering, 37(4):1599–1619, 2025

  25. [33]

    Add and Thin: Diffusion for Temporal Point Processes

    David Lüdke, Marin Biloš, Oleksandr Shchur, Marten Lienen, and Stephan Günnemann. Add and Thin: Diffusion for Temporal Point Processes. InAdvances in Neural Information Processing Systems, pages 56784–56801, 2023

  26. [34]

    Unlocking Point Processes through Point Set Diffusion

    David Lüdke, Enric Rabasseda Raventós, Marcel Kollovieh, and Stephan Günnemann. Unlocking Point Processes through Point Set Diffusion. InProceedings of the 13th International Conference on Learning Representations, 2025

  27. [35]

    Edit-Based Flow Match- ing for Temporal Point Processes

    David Lüdke, Marten Lienen, Marcel Kollovieh, and Stephan Günnemann. Edit-Based Flow Match- ing for Temporal Point Processes. InProceedings of the 14th International Conference on Learning Representations, 2026

  28. [36]

    Distance covariance in metric spaces.The Annals of Probability, 41(5):3284–3305, 2013

    Russell Lyons. Distance covariance in metric spaces.The Annals of Probability, 41(5):3284–3305, 2013

  29. [37]

    Lyons, Michael Caruana, and Thierry Lévy.Differential Equations Driven by Rough Paths: École d’Été de Probabilités de Saint-Flour XXXIV - 2004

    Terry J. Lyons, Michael Caruana, and Thierry Lévy.Differential Equations Driven by Rough Paths: École d’Été de Probabilités de Saint-Flour XXXIV - 2004. Springer, 2007

  30. [38]

    Hongyuan Mei and Jason M. Eisner. The Neural Hawkes Process: A Neurally Self-Modulating Multivariate Point Process. InAdvances in Neural Information Processing Systems, pages 6754–6764, 2017

  31. [39]

    Selby, Yao Xie, Sebastian V ollmer, and Gerrit Grossmann

    Sumantrak Mukherjee, Mouad Elhamdi, George Mohler, David A. Selby, Yao Xie, Sebastian V ollmer, and Gerrit Grossmann. Neural Spatiotemporal Point Processes: Trends and Challenges.Transactions on Machine Learning Research, 2025

  32. [40]

    Sig- Wasserstein GANs for Time Series Generation

    Hao Ni, Lukasz Szpruch, Marc Sabate-Vidales, Baoren Xiao, Magnus Wiese, and Shujian Liao. Sig- Wasserstein GANs for Time Series Generation. InProceedings of the Second ACM International Confer- ence on AI in Finance, 2021

  33. [41]

    Deep Mixture Point Processes: Spatio-temporal Event Prediction with Rich Contextual Information

    Maya Okawa, Tomoharu Iwata, Takeshi Kurashima, Yusuke Tanaka, Hiroyuki Toda, and Naonori Ueda. Deep Mixture Point Processes: Spatio-temporal Event Prediction with Rich Contextual Information. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery...

  34. [42]

    Context-aware spatio-temporal event prediction via convolutional Hawkes processes.Machine Learning, 111(8):2929–2950, 2022

    Maya Okawa, Tomoharu Iwata, Yusuke Tanaka, Takeshi Kurashima, Hiroyuki Toda, and Hisashi Kashima. Context-aware spatio-temporal event prediction via convolutional Hawkes processes.Machine Learning, 111(8):2929–2950, 2022

  35. [43]

    Fully Neural Network Based Model for General Temporal Point Processes

    Takahiro Omi, Naonori Ueda, and Kazuyuki Aihara. Fully Neural Network Based Model for General Temporal Point Processes. InAdvances in Neural Information Processing Systems, pages 2120–2129, 2019. 11

  36. [44]

    Goodwin, John R

    Imanol Perez Arribas, Guy M. Goodwin, John R. Geddes, Terry Lyons, and Kate E. A. Saunders. A signature-based machine learning model for distinguishing bipolar disorder and borderline personality disorder.Translational Psychiatry, 8(274), 2018

  37. [45]

    Romain Pic, Clément Dombry, Philippe Naveau, and Maxime Taillardat. Proper scoring rules for multivari- ate probabilistic forecasts based on aggregation and transformation.Advances in Statistical Climatology, Meteorology and Oceanography, 11(1):23–58, 2025

  38. [46]

    On Wasserstein Two-Sample Testing and Related Families of Nonparametric Tests.Entropy, 19(2):47, 2017

    Aaditya Ramdas, Nicolás García Trillos, and Marco Cuturi. On Wasserstein Two-Sample Testing and Related Families of Nonparametric Tests.Entropy, 19(2):47, 2017

  39. [47]

    I. J. Schoenberg. Metric Spaces and Positive Definite Functions.Transactions of the American Mathemati- cal Society, 44(3):522–536, 1938

  40. [48]

    Fast and Flexible Temporal Point Processes with Triangular Maps

    Oleksandr Shchur, Nicholas Gao, Marin Biloš, and Stephan Günnemann. Fast and Flexible Temporal Point Processes with Triangular Maps. InAdvances in Neural Information Processing Systems, pages 73–84, 2020

  41. [49]

    Neural Temporal Point Processes: A Review

    Oleksandr Shchur, Ali Caner Türkmen, Tim Januschowski, and Stephan Günnemann. Neural Temporal Point Processes: A Review. InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 4585–4593, 2021

  42. [50]

    Székely and Maria L

    Gábor J. Székely and Maria L. Rizzo. Energy statistics: A class of statistics based on distances.Journal of Statistical Planning and Inference, 143(8):1249–1272, 2013

  43. [51]

    Time is of the Essence: A Joint Hierarchical RNN and Point Process Model for Time and Item Predictions

    Bjørnar Vassøy, Massimiliano Ruocco, Eliezer de Souza da Silva, and Erlend Aune. Time is of the Essence: A Joint Hierarchical RNN and Point Process Model for Time and Item Predictions. InProceedings of the Twelfth ACM International Conference on Web Search and Data Mining, pag...

  44. [52]

    Springer, 2009

    Cédric Villani.Optimal Transport: Old and New. Springer, 2009

  45. [53]

    Wasserstein Learning of Deep Generative Point Process Models

    Shuai Xiao, Mehrdad Farajtabar, Xiaojing Ye, Junchi Yan, Xiaokang Yang, Le Song, and Hongyuan Zha. Wasserstein Learning of Deep Generative Point Process Models. InAdvances in Neural Information Processing Systems, pages 3247–3257, 2017

  46. [54]

    Zhang, Qingsong Wen, Jun Zhou, and Hongyuan Mei

    Siqiao Xue, Xiaoming Shi, Zhixuan Chu, Yan Wang, Hongyan Hao, Fan Zhou, Caigao Jiang, Chen Pan, James Y . Zhang, Qingsong Wen, Jun Zhou, and Hongyuan Mei. EasyTPP: Towards Open Bench- marking Temporal Point Processes. InProceedings of the 12th International Conference on Learn...

  47. [55]

    Spatio-Temporal Diffusion Point Processes

    Yuan Yuan, Jingtao Ding, Chenyang Shao, Depeng Jin, and Yong Li. Spatio-Temporal Diffusion Point Processes. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 3173–3184, 2023

  48. [56]

    Self-Attentive Hawkes Process

    Qiang Zhang, Aldo Lipani, Omer Kirnap, and Emine Yilmaz. Self-Attentive Hawkes Process. In Proceedings of the 37th International Conference on Machine Learning, pages 11183–11193, 2020

  49. [57]

    Automatic Integration for Spatiotemporal Neural Point Processes

    Zihao Zhou and Rose Yu. Automatic Integration for Spatiotemporal Neural Point Processes. InAdvances in Neural Information Processing Systems, pages 50237–50253, 2023

  50. [58]

    Neural Point Process for Learning Spatiotemporal Event Dynamics

    Zihao Zhou, Xingyi Yang, Ryan Rossi, Handong Zhao, and Rose Yu. Neural Point Process for Learning Spatiotemporal Event Dynamics. InProceedings of the 4th Annual Learning for Dynamics and Control Conference, pages 777–789, 2022

  51. [59]

    Imitation Learning of Neural Spatio-Temporal Point Processes.IEEE Transactions on Knowledge and Data Engineering, 34(11):5391–5402, 2022

    Shixiang Zhu, Shuang Li, Zhigang Peng, and Yao Xie. Imitation Learning of Neural Spatio-Temporal Point Processes.IEEE Transactions on Knowledge and Data Engineering, 34(11):5391–5402, 2022. 12 Appendix A A Pathwise Metric on Counting Paths A.1 Distances induced by the Pathwise...

  52. [60]

    The mapΦis Lipschitz (Theorem S11)

  53. [61]

    The map Φ is injective on N , hence a bijection between N and its image Φ(N) (Theo- rem S16)

  54. [62]

    Continuity fails on all ofΦ(N); see Remark S17

    The inverse mapΨ=Φ| −1 N is continuous on Φ(Nδ) for each δ >0; more precisely, it satisfies a 1 2-Hölder bound on the separated subset Φ(Nδ) defined in Appendix B.3 (Theorem S29). Continuity fails on all ofΦ(N); see Remark S17. B.1Φis Lipschitz Lemma S6(AffineL 1 Bound).Leta <...

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.