Pith. sign in

REVIEW 2 major objections 4 minor 157 references

Advances in Temporal Point Processes: Bayesian, Neural, and LLM Approaches

T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey tries to establish a complete map of temporal point process research as of 2020-2025, arguing that the field splits into Bayesian, neural, and LLM-based approaches that all revolve around the same conditional-intensity object.

desk verdict A timely and useful survey of TPPs, with an excellent LLM section and a solid Bayesian nonparametric roadmap, but the branching likelihood in Eq. (11) is wrong and must be fixed. read the letter →

arxiv 2501.14291 v3 pith:67QQH7UB submitted 2025-01-24 cs.LG stat.ML

classification cs.LGstat.ML MSC 60G5562M0962F1568T07
keywords temporalpointprocessesconditionalintensityfunctionHawkesprocessBayesiannonparametricinferenceneuralTransformerattentionlargelanguagemodelsnoise-contrastiveestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Temporal point processes (TPPs) model event sequences whose timestamps are irregular, such as tweets, trades, neural spikes, and earthquakes, and this survey tries to establish where the field stands in 2020-2025. The authors argue that current research falls into three coherent camps: Bayesian TPPs, mostly nonparametric and placing priors on the intensity function; neural TPPs, using recurrent, autoregressive, and differential-equation architectures; and LLM-based TPPs, using language models for prediction, reasoning, retrieval, and multimodal events. Across all three camps, they claim the same mathematical object, the conditional intensity function, and the same estimation question recur, with training objectives ranging from maximum likelihood to Wasserstein, noise-contrastive, and score-matching criteria. The survey positions itself as an update to earlier reviews that stopped around 2020 and largely ignored Bayesian nonparametric TPPs, so it aims to be the place to see the current shape of the field and its open problems.

What carries the argument

The load-bearing object is the conditional intensity function $\lambda^*(t)$, the expected number of events per unit time given the event history, together with its one-to-one correspondence to the conditional density $f(t \mid H_{t_n}) = \lambda^*(t) \exp(-\int_{t_n}^t \lambda^*(\tau)\,d\tau)$. Because every TPP can be specified either through this function or through the next-event distribution, the whole survey rests on that equivalence. Three secondary mechanisms carry the three branches: Gaussian process priors on the intensity, split into independent Poisson components via a branching latent variable for Bayesian nonparametrics; learned history embeddings $h_n$, whether updated recurrently, computed by self-attention, or evolved continuously by an ODE or SDE with jumps, for neural TPPs; and the divergence-minimization objective $D(f(T) \| f_\theta(T))$ that defines training across all frameworks.

What would settle it

Check whether a published temporal point process model with an ODE- or SDE-driven intensity antedates the work named as first in Section IV-C, since that priority claim is falsifiable; likewise, a complete enumeration of TPP papers from 2020 to 2025 would either confirm or break the Figure 1 taxonomy, because the survey reports no systematic corpus-selection protocol.

Watch

Extended reading notes

Core claim

The central claim this survey asserts is a taxonomy: recent TPP research is organized by modeling philosophy into Bayesian, neural, and LLM-based approaches, and each branch has a small number of canonical mechanisms. Bayesian nonparametric TPPs treat the intensity function as an infinite-dimensional parameter under a Gaussian process or Dirichlet process prior, with the doubly intractable posterior handled by MCMC, Laplace, or variational methods; for Hawkes processes, a branching latent variable factorizes the likelihood so the baseline and triggering functions can be estimated separately. Neural TPPs divide architecturally into recurrent models with event-by-event history updates, autoregressive models based on Transformer self-attention, and differential-equation models with hidden states that flow between events and jump at events, alongside a thread on equivalent parameterizations where modeling the cumulative intensity turns the log-likelihood integral into a derivative. LLM-based TPPs are treated as a distinct emerging branch that includes prompt-inspired continual learning, direct fusion of LLMs with temporal encoding, retrieval, and multimodal benchmarks. The survey further claims that architectures are only half the story: the other half is the training objective, and it reviews KL divergence and maximum likelihood, Wasserstein distance, noise-contrastive estimation, and Fisher divergence as alternative answers to the same distribution-matching problem.

Load-bearing premise

The survey's claim to be comprehensive rests on the assumption that its three-way taxonomy, and the placement of individual methods within it, accurately reflects the entire field; if a significant line of research is omitted or a key attribution is wrong, the map loses its value as a guide.

Editorial extensions

If this is right

  • A new TPP modeler can be located in the field by two choices: which family supplies the history representation, and which parameterization (intensity, cumulative intensity, density, or quantile) defines the next-event distribution.
  • Bayesian nonparametric Hawkes inference reduces to a template: augment a branching variable, then alternate between updating the branching posterior and updating the two Gaussian-process-modulated Poisson components, a recipe the survey identifies across many papers.
  • Modeling the cumulative conditional intensity $\Lambda^*(t)$ with monotonic networks removes numerical integration from maximum likelihood estimation and turns the log-likelihood into a derivative, a trick that applies beyond any single architecture.
  • The taxonomy implies that the LLM branch is not a separate mathematical framework but a new interface for the same conditional-intensity object, adding text, retrieval, and multimodal context to event sequences.
  • The stated lack of a standard benchmark protocol means cross-paper performance claims in neural TPPs should be read with caution until a consistent evaluation setup becomes standard practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable consequence the authors leave implicit: if the taxonomy is right, nearly every neural TPP paper from 2020 to 2025 should fall into exactly one of the three architectural families, so a systematic coding of the literature would measure how much of the field the survey actually covers.
  • The paper's emphasis on the equivalence among intensity, density, and cumulative-intensity parameterizations suggests a natural comparison it does not run: hold the architecture fixed and train the same TPP with MLE, Wasserstein, noise-contrastive, and score-matching objectives on identical datasets to see which estimator pays for its variance.
  • Since the survey notes that earlier score-matching estimators for point processes were shown to be incomplete, an immediate open question is whether the corrected weighted estimator also fixes the earlier spatio-temporal and multivariate variants.
  • The LLM branch may progressively dissolve into the neural branch, because LLM-based TPPs are themselves neural TPPs; the three-way split is best read as a snapshot of research momentum rather than a permanent division.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This survey reviews temporal point processes from three perspectives: Bayesian methods (parametric and nonparametric), neural TPPs (recurrent, autoregressive, and differential-equation based), and LLM-based TPPs (inspired, integrated, and multimodal extensions). It also covers model training objectives and applications in event prediction and causal discovery. The paper aims to provide a comprehensive and current update to earlier surveys, with particular emphasis on Bayesian nonparametric Hawkes processes and on work from 2020 through 2024.

Significance. If its formulas and characterizations are correct, the survey would be a valuable and current map of the TPP field: the organization is clear, the reference list is broad and up to date, and the treatment of Bayesian nonparametric Hawkes inference, SDE-based neural TPPs, and LLM-TPP integration fills genuine gaps in earlier surveys. The paper is expository and contains no fitting loop or circular derivation; its contribution is organizational and pedagogical. However, the correctness of the central augmented Hawkes likelihood in Section III-B is essential for the survey to serve as a reliable reference, and that equation is currently wrong.

major comments (2)
  1. [III-B, Eq. (11)] Equation (11) states the triggering component of the augmented Hawkes likelihood as the product over n of exp(-∫_0^{Tφ} φ(τ)dτ), which does not depend on event times, and then claims that marginalizing X recovers the original Hawkes likelihood. This is incorrect. In the branching representation, each event t_j contributes a survival factor exp(-∫_{t_j}^T φ(t-t_j)dt) = exp(-∫_0^{min(Tφ,T-t_j)} φ(τ)dτ), because the triggering kernel from event j can only excite later times and the observation window ends at T. For a single event at t_1, Eq. (11) gives exp(-∫_0^{Tφ}φ(τ)dτ), whereas the correct marginalized likelihood has exp(-∫_0^{T-t_1}φ(τ)dτ), and the two differ whenever T-t_1<Tφ. The equation also omits the single-parent constraint ∑_{m≤n} x_{nm}=1. Since Section III-B builds its entire account of GP-Hawkes inference on this factorization, the formula should be corrected before the survey is used as a reference for Bayesian nonparametric Hawkes processes.
  2. [II-C] The blanket statement that "the TPP likelihood is not conjugate to any prior" is too strong. Conjugate analyses exist for basic Poisson process models, such as Gamma priors for a homogeneous Poisson intensity, and later sections of this paper rely on conditionally conjugate augmented likelihoods, e.g., the Pólya-Gamma augmentation discussed in Section III-A. The sentence should be qualified to refer to the general lack of convenient conjugacy for Hawkes or Cox-process likelihoods rather than to all TPPs.
minor comments (4)
  1. [Fig. 1 and Section VI-B] In Figure 1, the Wasserstein Distance entry cites [86,87], but Section VI-B cites [151] for the Wasserstein GAN framework and [87] for the later likelihood-free method. Reference [86] is a duplicate of [51], the RNN intensity model, and is not the Wasserstein result. The figure entry should likely read [151,87].
  2. [Introduction] The Introduction calls the survey comprehensive but does not describe the literature search or inclusion criteria. Given the claim of comprehensiveness, a short statement of the period, venues, and selection rules would help readers calibrate coverage and would make the taxonomy in Figure 1 easier to evaluate.
  3. [Various] There are several typographical issues: "K-varaite" in Section VII-B should be "K-variate"; the footnote in Section II-A contains a stray "1" after "one-to-one"; reference [139] misspells "Good" as "Goodd"; reference [100] misspells "Conference" as "Confernce". These should be corrected in a final pass.
  4. [IV-C] The claim that "the earliest work, to the best of our knowledge, that integrates differential equations with point processes is Chen et al. [68]" is worth softening or documenting. The statement is hedged, but the characterization of a Neural ODE paper as the first integration of differential equations with point processes is the kind of historical claim that readers may rely on; a citation to the relevant section of [68] or a more explicit caveat would strengthen it.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey restates external results and cites prior work descriptively; the flagged Eq. (11) issue is a correctness concern, not a circular reduction.

full rationale

This is a survey, not a derivation or prediction pipeline. Its central claim—a three-way taxonomy of Bayesian, neural, and LLM-based TPPs with literature characterizations—is organizational and descriptive; it is not obtained by fitting parameters, by a theorem whose assumptions include the conclusion, or by renaming a measured quantity as a forecast. The mathematical statements it does derive, such as Eqs. (1)–(3) and the likelihood in Eq. (7), follow directly from standard definitions and are not equivalent to the survey's inputs by construction. Self-citations such as [39], [40], [42], [81], [83], [91], [93], [128], and [134] appear throughout the taxonomy and in descriptive passages, but they are cited as prior published contributions and are not load-bearing premises that force an otherwise unsupported conclusion. The branching-variable likelihood in Eq. (11), as the skeptic notes, appears to replace the per-event exposure interval [t_i, T] with a fixed [0, T_phi], making the formula incorrect for the Bayesian nonparametric Hawkes inference account in Section III-B. However, a mistaken formula is a correctness or accuracy problem, not a circularity problem: the text does not derive Eq. (11) from the posterior it claims to compute, and no fitted parameter or self-citation chain is being presented as an independent prediction. Accordingly, no circular step can be exhibited under the required reductions, and the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No parameters fitted, no new entities postulated. All assumptions are standard mathematical facts or the meta-assumption that the survey's categorization is faithful. The main caveat is Eq. (11)'s branching likelihood, which is presented imprecisely.

assumptions (4)
  • standard math Conditional intensity and conditional density are related by the hazard formula in Eq. (3), used as the bridge between TPP parameterizations.
    Invoked in Section II-A and used throughout the survey to translate between density, intensity, and cumulative intensity parameterizations; standard measure-theoretic background.
  • standard math The likelihood of a marked TPP is the product of conditional intensities times the exponential survival term, Eq. (7).
    Used in Section II-C and later for MLE; a standard result for simple point processes.
  • domain assumption The branching representation of Hawkes processes with latent parent assignments, Eq. (11), is a valid decomposition of the Hawkes likelihood.
    Used in Section III-B to explain Bayesian nonparametric Hawkes inference; the formula as written is imprecise (survival exponents independent of event times, no single-parent constraint), so the assumption should be read as the standard branching-process construction rather than the displayed equation.
  • domain assumption The three-way classification of TPP research into Bayesian, neural, and LLM approaches, and the subdivision of neural TPPs into recurrent/autoregressive/differential-equation families, partitions the recent literature.
    The entire survey structure depends on this partition; the paper asserts it in the Introduction and Figure 1 but provides no systematic corpus-selection protocol, so coverage depends on author selection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Advances in Temporal Point Processes: Bayesian, Neural, and LLM Approaches." pith.science (2026). https://pith.science/paper/67QQH7UB

@misc{pith2026250114291,
  author       = {Pith},
  title        = {Pith review of: Advances in Temporal Point Processes: Bayesian, Neural, and LLM Approaches},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/67QQH7UB}},
  note         = {Machine review of arXiv:2501.14291}
}
read the original abstract

Temporal point processes (TPPs) are stochastic process models used to characterize event sequences occurring in continuous time. Traditional statistical TPPs have a long-standing history, with numerous models proposed and successfully applied across diverse domains. In recent years, advances in deep learning have spurred the development of neural TPPs, enabling greater flexibility and expressiveness in capturing complex temporal dynamics. The emergence of large language models (LLMs) has further sparked excitement, offering new possibilities for modeling and analyzing event sequences by leveraging their rich contextual understanding. This survey presents a comprehensive review of recent research on TPPs from three perspectives: Bayesian, deep learning, and LLM approaches. We begin with a review of the fundamental concepts of TPPs, followed by an in-depth discussion of model design and parameter estimation techniques in these three frameworks. We also revisit classic application areas of TPPs to highlight their practical relevance. Finally, we outline challenges and promising directions for future research.

Figures

Figures reproduced from arXiv: 2501.14291 by the authors.

Figure 1
Figure 1. The taxonomy of Bayesian, neural, and LLM-based TPPs. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the conditional density function and conditional intensity function in TPPs. (a) The conditional density function of the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Illustration of neural TPPs. (a) Recurrent neural TPPs, where the hidden state is updated recurrently using the current event and then used to model [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Four common parameterizations of TPPs: (a) the probability density function (PDF) of the next event, (b) the cumulative distribution function (CDF) [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: An overview of LLM-based TPPs. The event time [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: An illustration of DanmakuTPPBench [83], a multi-modal benchmark [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: An overview of applications of TPPs. TPPs are widely used for event prediction in domains such as social networks, epidemiology, fraud detection, [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

157 extracted references · 71 canonical work pages

  1. [68]

    Neural ordinary differential equations,

    R. T. Chen, Y . Rubanova, J. Bettencourt, and D. K. Duvenaud, “Neural ordinary differential equations,” Advances in neural information pro- cessing systems, vol. 31, 2018

  2. [1]

    D. J. Daley and D. Vere-Jones, An introduction to the theory of point processes: volume II: general theory and structure . Springer Science & Business Media, 2007

  3. [2]

    Scalable Bayesian inference for excitatory point process networks,

    S. W. Linderman and R. P. Adams, “Scalable Bayesian inference for excitatory point process networks,” arXiv preprint, 2015

  4. [3]

    Hawkes model for price and trades high- frequency dynamics,

    E. Bacry and J.-F. Muzy, “Hawkes model for price and trades high- frequency dynamics,” Quantitative Finance, vol. 14, no. 7, pp. 1147– 1166, 2014

  5. [4]

    Interval-censored Transformer Hawkes: Detecting information oper- ations using the reaction of social systems,

    Q. Kong, P. Calderon, R. Ram, O. Boichak, and M.-A. Rizoiu, “Interval-censored Transformer Hawkes: Detecting information oper- ations using the reaction of social systems,” in The Web Conference , 2023

  6. [5]

    J. F. C. Kingman, Poisson processes. Clarendon Press, 1992, vol. 3

  7. [6]

    Spectra of some self-exciting and mutually exciting point processes,

    A. G. Hawkes, “Spectra of some self-exciting and mutually exciting point processes,” Biometrika, vol. 58, no. 1, pp. 83–90, 1971

  8. [7]

    A self-correcting point process,

    V . Isham and M. Westcott, “A self-correcting point process,” Stochastic Processes and Their Applications , vol. 8, no. 3, pp. 335–347, 1979

Show all 157 references
  1. [8]

    Recent advance in temporal point process: from machine learning perspective,

    J. Yan, “Recent advance in temporal point process: from machine learning perspective,” SJTU Technical Report, 2019

  2. [9]

    Neural temporal point processes: A review,

    O. Shchur, A. C. T ¨urkmen, T. Januschowski, and S. G ¨unnemann, “Neural temporal point processes: A review,” arXiv preprint, 2021

  3. [10]

    Hawkes processes in finance,

    E. Bacry, I. Mastromatteo, and J.-F. Muzy, “Hawkes processes in finance,” Market Microstructure and Liquidity , vol. 1, no. 01, p. 1550005, 2015

  4. [11]

    Hawkes processes and their applications to finance: a review,

    A. G. Hawkes, “Hawkes processes and their applications to finance: a review,” Quantitative Finance, vol. 18, no. 2, pp. 193–198, 2018

  5. [12]

    Bayesian analysis of a Poisson process with a change-point,

    A. E. Raftery and V . E. Akman, “Bayesian analysis of a Poisson process with a change-point,” Biometrika, pp. 85–89, 1986

  6. [13]

    Bayesian estimation of multivari- ate hawkes processes with inhibition and sparsity,

    I. Deutsch and G. J. Ross, “Bayesian estimation of multivari- ate hawkes processes with inhibition and sparsity,” arXiv preprint arXiv:2201.05009, 2022

  7. [14]

    Bayesian inference for hawkes processes,

    J. G. Rasmussen, “Bayesian inference for hawkes processes,” Method- ology and Computing in Applied Probability , vol. 15, pp. 623–642, 2013

  8. [15]

    Approximate methods in bayesian point process spatial models,

    M. M. Hossain and A. B. Lawson, “Approximate methods in bayesian point process spatial models,” Computational statistics & data analysis, vol. 53, no. 8, pp. 2831–2842, 2009

  9. [16]

    Fitting complex ecological point process models with integrated nested laplace approximation,

    J. B. Illian, S. Martino, S. H. Sørbye, J. B. Gallego-Fern ´andez, M. Zun- zunegui, M. P. Esquivias, and J. M. Travis, “Fitting complex ecological point process models with integrated nested laplace approximation,” Methods in Ecology and Evolution , vol. 4, no. 4, pp. 305–315, 2013

  10. [17]

    Vari- ational estimation in spatiotemporal systems from continuous and point-process observations,

    A. Zammit-Mangion, G. Sanguinetti, and V . Kadirkamanathan, “Vari- ational estimation in spatiotemporal systems from continuous and point-process observations,” IEEE Transactions on Signal Processing , vol. 60, no. 7, pp. 3449–3459, 2012

  11. [18]

    Log Gaussian Cox processes,

    J. Møller, A. R. Syversveen, and R. P. Waagepetersen, “Log Gaussian Cox processes,” Scandinavian journal of statistics , vol. 25, no. 3, pp. 451–482, 1998

  12. [19]

    Tractable nonparametric Bayesian inference in Poisson processes with Gaussian process inten- sities,

    R. Adams, I. Murray, and D. MacKay, “Tractable nonparametric Bayesian inference in Poisson processes with Gaussian process inten- sities,” in International Conference on Machine Learning , 2009

  13. [20]

    The permanental process,

    P. McCullagh and J. Møller, “The permanental process,” Advances in applied probability, vol. 38, no. 4, pp. 873–888, 2006

  14. [21]

    MCMC for doubly- intractable distributions,

    I. Murray, Z. Ghahramani, and D. J. MacKay, “MCMC for doubly- intractable distributions,” in Conference on Uncertainty in Artificial Intelligence, 2006

  15. [22]

    Efficient Bayesian nonparametric modelling of structured point processes,

    T. Gunter, C. Lloyd, M. A. Osborne, and S. J. Roberts, “Efficient Bayesian nonparametric modelling of structured point processes,” in Conference on Uncertainty in Artificial Intelligence , 2014

  16. [23]

    Scalable nonparametric Bayesian inference on point processes with Gaussian processes,

    Y .-L. K. Samo and S. Roberts, “Scalable nonparametric Bayesian inference on point processes with Gaussian processes,” in International Conference on Machine Learning , 2015

  17. [24]

    Fast Gaussian pro- cess methods for point process intensity estimation,

    J. P. Cunningham, K. V . Shenoy, and M. Sahani, “Fast Gaussian pro- cess methods for point process intensity estimation,” in International Conference on Machine Learning , 2008

  18. [25]

    Fast Bayesian intensity estimation for the permanental process,

    C. J. Walder and A. N. Bishop, “Fast Bayesian intensity estimation for the permanental process,” in International Conference on Machine Learning, 2017

  19. [26]

    Fast Kronecker inference in Gaussian processes with non-Gaussian likeli- hoods,

    S. Flaxman, A. Wilson, D. Neill, H. Nickisch, and A. Smola, “Fast Kronecker inference in Gaussian processes with non-Gaussian likeli- hoods,” in International Conference on Machine Learning . PMLR, 2015, pp. 607–616

  20. [27]

    Sparse spectral Bayesian permanental process with generalized kernel,

    J. Sellier and P. Dellaportas, “Sparse spectral Bayesian permanental process with generalized kernel,” in International Conference on Arti- ficial Intelligence and Statistics , 2023

  21. [28]

    Nonstationary sparse spectral permanental process,

    Z. Sun, Y . Zhang, Z. Ling, X. Fan, and F. Zhou, “Nonstationary sparse spectral permanental process,” in Advances in Neural Information Processing Systems, 2024

  22. [29]

    Variational inference for Gaussian process modulated Poisson processes,

    C. Lloyd, T. Gunter, M. Osborne, and S. Roberts, “Variational inference for Gaussian process modulated Poisson processes,” in International Conference on Machine Learning , 2015

  23. [30]

    A multitask point process predictive model,

    W. Lian, R. Henao, V . Rao, J. Lucas, and L. Carin, “A multitask point process predictive model,” in International Conference on Machine Learning, 2015

  24. [31]

    Large-scale Cox process inference using variational Fourier features,

    S. John and J. Hensman, “Large-scale Cox process inference using variational Fourier features,” in International Conference on Machine Learning, 2018

  25. [32]

    Structured variational inference in continuous cox process models,

    V . Aglietti, E. V . Bonilla, T. Damoulas, and S. Cripps, “Structured variational inference in continuous cox process models,” Advances in Neural Information Processing Systems , vol. 32, 2019

  26. [33]

    Efficient Bayesian inference of sigmoidal Gaussian Cox processes,

    C. Donner and M. Opper, “Efficient Bayesian inference of sigmoidal Gaussian Cox processes,” Journal of Machine Learning Research , vol. 19, no. 1, pp. 2710–2743, 2018

  27. [34]

    Heterogeneous multi-task Gaussian Cox processes,

    F. Zhou, Q. Kong, Z. Deng, F. He, P. Cui, and J. Zhu, “Heterogeneous multi-task Gaussian Cox processes,” Machine Learning , vol. 112, no. 12, pp. 5105–5134, 2023

  28. [35]

    Dirichlet process mixtures of Beta distributions, with applications to density and intensity estimation,

    A. Kottas, “Dirichlet process mixtures of Beta distributions, with applications to density and intensity estimation,” in Workshop on Learning with Nonparametric Bayesian Methods, ICML , 2006

  29. [36]

    Bayesian mixture modeling for spatial Pois- son process intensities, with applications to extreme value analysis,

    A. Kottas and B. Sans ´o, “Bayesian mixture modeling for spatial Pois- son process intensities, with applications to extreme value analysis,” Journal of Statistical Planning and Inference , vol. 137, no. 10, pp. 3151–3163, 2007

  30. [37]

    Efficient non- parametric Bayesian Hawkes processes,

    R. Zhang, C. J. Walder, M. Rizoiu, and L. Xie, “Efficient non- parametric Bayesian Hawkes processes,” in International Joint Con- ference on Artificial Intelligence , 2019

  31. [38]

    Variational inference for sparse Gaussian process modulated Hawkes process,

    R. Zhang, C. Walder, and M.-A. Rizoiu, “Variational inference for sparse Gaussian process modulated Hawkes process,” in AAAI Confer- ence on Artificial Intelligence , 2020

  32. [39]

    Ef- ficient EM-variational inference for nonparametric Hawkes process,

    F. Zhou, S. Luo, Z. Li, X. Fan, Y . Wang, A. Sowmya, and F. Chen, “Ef- ficient EM-variational inference for nonparametric Hawkes process,” Statistics and Computing , vol. 31, no. 4, p. 46, 2021

  33. [40]

    Efficient inference for nonparametric Hawkes processes using auxiliary latent variables,

    F. Zhou, Z. Li, X. Fan, Y . Wang, A. Sowmya, and F. Chen, “Efficient inference for nonparametric Hawkes processes using auxiliary latent variables,” Journal of Machine Learning Research , vol. 21, no. 241, pp. 1–31, 2020

  34. [41]

    Variational bayesian inference for nonlinear hawkes process with gaussian process self- effects,

    N. Malem-Shinitski, C. Ojeda, and M. Opper, “Variational bayesian inference for nonlinear hawkes process with gaussian process self- effects,” Entropy, vol. 24, no. 3, p. 356, 2022

  35. [42]

    Efficient inference for dynamic flexible interactions of neural popu- lations,

    F. Zhou, Q. Kong, Z. Deng, J. Kan, Y . Zhang, C. Feng, and J. Zhu, “Efficient inference for dynamic flexible interactions of neural popu- lations,” Journal of Machine Learning Research , vol. 23, no. 211, pp. 1–49, 2022

  36. [43]

    Bayesian estimation of nonlinear hawkes processes,

    D. Sulem, V . Rivoirard, and J. Rousseau, “Bayesian estimation of nonlinear hawkes processes,” Bernoulli, vol. 30, no. 2, pp. 1257–1286, JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 2024

  37. [44]

    Bayesian nonparametric hawkes processes with appli- cations,

    D. Markwick, “Bayesian nonparametric hawkes processes with appli- cations,” Ph.D. dissertation, UCL (University College London), 2020

  38. [45]

    Bayesian nonparametric learning for point processes with spatial homogeneity: A spatial analysis of nba shot locations,

    F. Yin, J. Jiao, J. Yan, and G. Hu, “Bayesian nonparametric learning for point processes with spatial homogeneity: A spatial analysis of nba shot locations,” in International Conference on Machine Learning. PMLR, 2022, pp. 25 523–25 551

  39. [46]

    Semiparametric estimation for mul- tivariate hawkes processes using dependent dirichlet processes: An application to order flow data in financial markets,

    A. Z. Jiang and A. Rodriguez, “Semiparametric estimation for mul- tivariate hawkes processes using dependent dirichlet processes: An application to order flow data in financial markets,” arXiv preprint arXiv:2502.17723, 2025

  40. [47]

    Online nonparametric bayesian hawkes processes,

    J. Worrall, “Online nonparametric bayesian hawkes processes,” Ph.D. dissertation, Queensland University of Technology, 2024

  41. [48]

    Nonparametric bayesian estimation for multivariate hawkes processes,

    S. Donnet, V . Rivoirard, and J. Rousseau, “Nonparametric bayesian estimation for multivariate hawkes processes,” The Annals of statistics , vol. 48, no. 5, pp. 2698–2727, 2020

  42. [49]

    Recurrent marked temporal point processes: embedding event history to vector,

    N. Du, H. Dai, R. Trivedi, U. Upadhyay, M. Gomez-Rodriguez, and L. Song, “Recurrent marked temporal point processes: embedding event history to vector,” inInternational Conference on Knowledge Discovery and Data Mining , 2016

  43. [50]

    The neural Hawkes process: A neurally self-modulating multivariate point process,

    H. Mei and J. Eisner, “The neural Hawkes process: A neurally self-modulating multivariate point process,” in Advances in Neural Information Processing Systems , 2017

  44. [51]

    Modeling the intensity function of point process via recurrent neural networks,

    S. Xiao, J. Yan, X. Yang, H. Zha, and S. Chu, “Modeling the intensity function of point process via recurrent neural networks,” in AAAI Conference on Artificial Intelligence , 2017

  45. [52]

    Recurrent spatio-temporal point process for check-in time prediction,

    G. Yang, Y . Cai, and C. K. Reddy, “Recurrent spatio-temporal point process for check-in time prediction,” in International Conference on Information and Knowledge Management , 2018

  46. [53]

    Fully neural network based model for general temporal point processes,

    T. Omi, N. Ueda, and K. Aihara, “Fully neural network based model for general temporal point processes,” in Advances in Neural Information Processing Systems, 2019

  47. [54]

    Unipoint: Uni- versally approximating point processes intensities,

    A. Soen, A. Mathews, D. Grixti-Cheng, and L. Xie, “Unipoint: Uni- versally approximating point processes intensities,” in AAAI conference on artificial intelligence , 2021

  48. [55]

    Learning tem- poral point processes with intermittent observations,

    V . Gupta, S. Bedathur, S. Bhattacharya, and A. De, “Learning tem- poral point processes with intermittent observations,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2021, pp. 3790–3798

  49. [56]

    On non-asymptotic theory of re- current neural networks in temporal point processes,

    Z. Chen, G. Fang, and W. Yu, “On non-asymptotic theory of re- current neural networks in temporal point processes,” arXiv preprint arXiv:2406.00630, 2024

  50. [57]

    Mamba hawkes process,

    A. Gao, S. Dai, and Y . Hu, “Mamba hawkes process,” arXiv preprint, 2024

  51. [58]

    Deep linear Hawkes processes,

    Y . Chang, A. Boyd, C. Xiao, T. Kass-Hout et al., “Deep linear Hawkes processes,” arXiv preprint, 2024

  52. [59]

    Transformer Hawkes process,

    S. Zuo, H. Jiang, Z. Li, T. Zhao, and H. Zha, “Transformer Hawkes process,” in International Conference on Machine Learning , 2020

  53. [60]

    Self-attentive Hawkes process,

    Q. Zhang, A. Lipani, ¨O. Kirnap, and E. Yilmaz, “Self-attentive Hawkes process,” in International Conference on Machine Learning , 2020

  54. [61]

    Decomposable Transformer point processes,

    A. Panos, “Decomposable Transformer point processes,” in Advances in Neural Information Processing Systems , 2024

  55. [62]

    Deep fourier kernel for self-attentive point processes,

    S. Zhu, M. Zhang, R. Ding, and Y . Xie, “Deep fourier kernel for self-attentive point processes,” in International Conference on Artificial Intelligence and Statistics , 2021

  56. [63]

    Neural point process for learning spatiotemporal event dynamics,

    Z. Zhou, X. Yang, R. Rossi, H. Zhao, and R. Yu, “Neural point process for learning spatiotemporal event dynamics,” in Learning for Dynamics and Control Conference, 2022

  57. [64]

    Transformer embeddings of irregularly spaced events and their participants,

    C. Yang, H. Mei, and J. Eisner, “Transformer embeddings of irregularly spaced events and their participants,” in Proceedings of the Tenth International Conference on Learning Representations (ICLR) , 2022

  58. [65]

    Sparse Transformer Hawkes process for long event sequences,

    Z. Li and M. Sun, “Sparse Transformer Hawkes process for long event sequences,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases , 2023

  59. [66]

    Interpretable Transformer Hawkes processes: Unveiling complex interactions in social networks,

    Z. Meng, K. Wan, Y . Huang, Z. Li, Y . Wang, and F. Zhou, “Interpretable Transformer Hawkes processes: Unveiling complex interactions in social networks,” in International Conference on Knowledge Discovery and Data Mining , 2024

  60. [67]

    Federated Transformer Hawkes processes for distributed event se- quence prediction,

    X. Wang, F. Qiang, L. Ma, P. Zhang, H. Yang, Z. Li, and J. Zhang, “Federated Transformer Hawkes processes for distributed event se- quence prediction,” in International Joint Conference on Neural Net- works, 2024

  61. [69]

    A stochastic differential equation framework for guiding online user activities in closed loop,

    Y . Wang, E. Theodorou, A. Verma, and L. Song, “A stochastic differential equation framework for guiding online user activities in closed loop,” in International Conference on Artificial Intelligence and Statistics. PMLR, 2018, pp. 1077–1086

  62. [70]

    Neural jump stochastic differential equations,

    J. Jia and A. R. Benson, “Neural jump stochastic differential equations,” in Advances in Neural Information Processing Systems , 2019

  63. [71]

    Neural spatio-temporal point processes,

    R. T. Chen, B. Amos, and M. Nickel, “Neural spatio-temporal point processes,” in International Conference on Learning Representations , 2021

  64. [72]

    Latent ordinary differential equations for irregularly-sampled time series,

    Y . Rubanova, R. T. Chen, and D. K. Duvenaud, “Latent ordinary differential equations for irregularly-sampled time series,” Advances in Neural Information Processing Systems , 2019

  65. [73]

    Neural jump-diffusion temporal point processes,

    S. Zhang, C. Zhou, Y . A. Liu, P. Zhang, X. Lin, and Z.-M. Ma, “Neural jump-diffusion temporal point processes,” in International Conference on Machine Learning , 2024

  66. [74]

    Intensity-free learning of temporal point processes,

    O. Shchur, M. Bilos, and S. G ¨unnemann, “Intensity-free learning of temporal point processes,” in International Conference on Learning Representations, 2020

  67. [75]

    Learning quantile functions for temporal point processes with recurrent neural splines,

    S. B. Taieb, “Learning quantile functions for temporal point processes with recurrent neural splines,” in International Conference on Artificial Intelligence and Statistics , 2022

  68. [76]

    Fast and flexible temporal point processes with triangular maps,

    O. Shchur, N. Gao, M. Bilos, and S. G ¨unnemann, “Fast and flexible temporal point processes with triangular maps,” in Advances in neural information processing systems , 2020

  69. [77]

    Cumulative hazard function based efficient multivariate tem- poral point process learning,

    B. Liu, “Cumulative hazard function based efficient multivariate tem- poral point process learning,” arXiv preprint, 2024

  70. [78]

    Prompt-augmented temporal point process for streaming event sequence,

    S. Xue, Y . Wang, Z. Chu, X. Shi et al., “Prompt-augmented temporal point process for streaming event sequence,” in Advances in Neural Information Processing Systems , 2023

  71. [79]

    Language models can improve event prediction by few-shot abductive reasoning,

    X. Shi, S. Xue, K. Wang, F. Zhou, J. Zhang, J. Zhou, C. Tan, and H. Mei, “Language models can improve event prediction by few-shot abductive reasoning,” in Advances in Neural Information Processing Systems, 2024

  72. [80]

    TPP-LLM: Modeling temporal point processes by efficiently fine-tuning large language models,

    Z. Liu and Y . Quan, “TPP-LLM: Modeling temporal point processes by efficiently fine-tuning large language models,” arXiv preprint, 2024

  73. [81]

    Language- tpp: Integrating temporal point processes with language models for event analysis,

    Q. Kong, Y . Zhang, Y . Liu, P. Tong, E. Liu, and F. Zhou, “Language- tpp: Integrating temporal point processes with language models for event analysis,” arXiv preprint arXiv:2502.07139 , 2025

  74. [82]

    Retrieval of temporal event sequences from textual descriptions,

    Z. Liu and Y . Quan, “Retrieval of temporal event sequences from textual descriptions,” in Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing , 2025, pp. 37–49

  75. [83]

    Danmakutpp- bench: A multi-modal benchmark for temporal point process modeling and understanding,

    Y . Jiang, J. Li, Y . Liu, D. Yang, F. Zhou, and Q. Kong, “Danmakutpp- bench: A multi-modal benchmark for temporal point process modeling and understanding,” arXiv preprint arXiv:2505.18411 , 2025

  76. [84]

    Maximum likelihood estimation of Hawkes’ self-exciting point processes,

    T. Ozaki, “Maximum likelihood estimation of Hawkes’ self-exciting point processes,” Annals of the Institute of Statistical Mathematics , vol. 31, no. 1, pp. 145–155, 1979

  77. [85]

    Maximum likelihood estimation of cascade point-process neural encoding models,

    L. Paninski, “Maximum likelihood estimation of cascade point-process neural encoding models,” Network: Computation in Neural Systems , vol. 15, no. 4, pp. 243–262, 2004

  78. [86]

    Modeling the intensity function of point process via recurrent neural networks,

    S. Xiao, J. Yan, X. Yang, H. Zha, and S. M. Chu, “Modeling the intensity function of point process via recurrent neural networks,” in Proceedings of the Thirty-First AAAI Conference on Artificial Intelli- gence, February 4-9, 2017, San Francisco, California, USA , S. Singh and...

  79. [87]

    Learning conditional generative models for temporal point processes,

    S. Xiao, H. Xu, J. Yan, M. Farajtabar, X. Yang, L. Song, and H. Zha, “Learning conditional generative models for temporal point processes,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018

  80. [88]

    Initiator: noise-contrastive estimation for marked temporal point process,

    R. Guo, J. Li, and H. Liu, “Initiator: noise-contrastive estimation for marked temporal point process,” in Proceedings of the 27th Interna- tional Joint Conference on Artificial Intelligence, 2018, pp. 2191–2197

  81. [89]

    Noise-contrastive estimation for multi- variate point processes,

    H. Mei, T. Wan, and J. Eisner, “Noise-contrastive estimation for multi- variate point processes,” in Advances in Neural Information Processing Systems, 2020

  82. [90]

    Score-matching estimators for continuous-time point-process regression models,

    M. Sahani, G. Bohner, and A. Meyer, “Score-matching estimators for continuous-time point-process regression models,” in 26th IEEE International Workshop on Machine Learning for Signal Processing, MLSP 2016, Vietri sul Mare, Salerno, Italy, September 13-16, 2016 , F. A. N. Palm...

  83. [91]

    Integration-free training for spatio- temporal multimodal covariate deep kernel point processes,

    Y . Zhang, Q. Kong, and F. Zhou, “Integration-free training for spatio- temporal multimodal covariate deep kernel point processes,” in Ad- vances in Neural Information Processing Systems , 2023

  84. [92]

    Smurf- thp: score matching-based uncertainty quantification for transformer hawkes process,

    Z. Li, Y . Xu, S. Zuo, H. Jiang, C. Zhang, T. Zhao, and H. Zha, “Smurf- thp: score matching-based uncertainty quantification for transformer hawkes process,” in International Conference on Machine Learning . PMLR, 2023, pp. 20 210–20 220. JOURNAL OF LATEX CLASS FILES, VOL. 14,...

  85. [93]

    Is score matching suitable for estimating point processes?

    H. Cao, Z. Meng, T. Ke, and F. Zhou, “Is score matching suitable for estimating point processes?” in Advances in Neural Information Processing Systems, 2024

  86. [94]

    Vigdet: Knowledge informed neural temporal point process for coordination detection on social media,

    Y . Zhang, K. Sharma, and Y . Liu, “Vigdet: Knowledge informed neural temporal point process for coordination detection on social media,” Advances in Neural Information Processing Systems, vol. 34, pp. 3218– 3231, 2021

  87. [95]

    Mush: Multi-stimuli hawkes process based sybil attacker detector for user-review social networks,

    Z. Qu, C. Lyu, and C.-H. Chi, “Mush: Multi-stimuli hawkes process based sybil attacker detector for user-review social networks,” IEEE Transactions on Network and Service Management , vol. 19, no. 4, pp. 4600–4614, 2022

  88. [96]

    Framework for detecting fake retweets using deep neural network,

    S. D. Hegde, A. Shetty, N. Manoj, A. Kalasad, and R. Bharathi, “Framework for detecting fake retweets using deep neural network,” in 2022 IEEE 7th International conference for Convergence in Technology (I2CT). IEEE, 2022, pp. 1–6

  89. [97]

    Temporal properties of higher-order interactions in social networks,

    G. Cencetti, F. Battiston, B. Lepri, and M. Karsai, “Temporal properties of higher-order interactions in social networks,” Scientific reports , vol. 11, no. 1, p. 7028, 2021

  90. [98]

    Public opinion field effect and hawkes process join hands for information popularity prediction,

    J. Li, Y . Yang, Y . Zhang, Q. Hu, A. Zhao, and H. Gao, “Public opinion field effect and hawkes process join hands for information popularity prediction,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 11, 2025, pp. 12 076–12 083

  91. [99]

    Identifying hidden patterns of fake covid-19 news: An in-depth sentiment analysis and topic modeling approach,

    T. Ahammad, “Identifying hidden patterns of fake covid-19 news: An in-depth sentiment analysis and topic modeling approach,” Natural Language Processing Journal, vol. 6, p. 100053, 2024

  92. [100]

    Sir- hawkes: on the relationship between epidemic models and Hawkes point processes,

    M. A. Rizoiu, S. Mishra, Q. Kong, M. Carman, and L. Xie, “Sir- hawkes: on the relationship between epidemic models and Hawkes point processes,” in The Web Confernce, 2018

  93. [101]

    Hawkes process modeling of covid-19 with mobility leading indicators and spatial covariates,

    W.-H. Chiang, X. Liu, and G. Mohler, “Hawkes process modeling of covid-19 with mobility leading indicators and spatial covariates,” International journal of forecasting, vol. 38, no. 2, pp. 505–520, 2022

  94. [102]

    Predicting covid-19 spread from large-scale mobility data,

    A. Schwabe, J. Persson, and S. Feuerriegel, “Predicting covid-19 spread from large-scale mobility data,” in Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining , 2021, pp. 3531–3539

  95. [103]

    Estimating covid-19 transmission time using hawkes point processes,

    F. Schoenberg, “Estimating covid-19 transmission time using hawkes point processes,” The Annals of Applied Statistics , vol. 17, no. 4, pp. 3349–3362, 2023

  96. [104]

    Network self-exciting point processes to measure health impacts of covid-19,

    P. Giudici, P. Pagnottoni, and A. Spelta, “Network self-exciting point processes to measure health impacts of covid-19,” Journal of the Royal Statistical Society Series A: Statistics in Society , vol. 186, no. 3, pp. 401–421, 2023

  97. [105]

    Space-time point-process models for earthquake occur- rences,

    Y . Ogata, “Space-time point-process models for earthquake occur- rences,” Annals of the Institute of Statistical Mathematics , vol. 50, no. 2, pp. 379–402, 1998

  98. [106]

    Flexible spatio-temporal hawkes process models for earthquake occurrences,

    J. Kwon, Y . Zheng, and M. Jun, “Flexible spatio-temporal hawkes process models for earthquake occurrences,” Spatial Statistics, vol. 54, p. 100728, 2023

  99. [107]

    Hawkes-based models for high frequency financial data,

    K. Nystr ¨om and C. Zhang, “Hawkes-based models for high frequency financial data,” Journal of the Operational Research Society , vol. 73, no. 10, pp. 2168–2185, 2022

  100. [108]

    Hawkes processes in finance: market structure and impact,

    J. Chen, N. Taylor, S. Yang, and Q. Han, “Hawkes processes in finance: market structure and impact,” The European Journal of Finance , vol. 28, no. 7, pp. 621–626, 2022

  101. [109]

    State dependent parallel neural hawkes process for limit order book event stream prediction and simulation,

    Z. Shi and J. Cartlidge, “State dependent parallel neural hawkes process for limit order book event stream prediction and simulation,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2022, pp. 1607–1615

  102. [110]

    A fractional hawkes process model for earthquake aftershock sequences,

    L. Davis, B. Baeumer, and T. Wang, “A fractional hawkes process model for earthquake aftershock sequences,” Journal of the Royal Statistical Society Series C: Applied Statistics, vol. 73, no. 5, pp. 1185– 1202, 2024

  103. [111]

    Sequential recommendation based on multivariate hawkes process embedding with attention,

    D. Wang, X. Zhang, Z. Xiang, D. Yu, G. Xu, and S. Deng, “Sequential recommendation based on multivariate hawkes process embedding with attention,” IEEE transactions on cybernetics , vol. 52, no. 11, pp. 11 893–11 905, 2021

  104. [112]

    Tideh: Time-dependent hawkes process for predicting retweet dynamics,

    R. Kobayashi and R. Lambiotte, “Tideh: Time-dependent hawkes process for predicting retweet dynamics,” in Proceedings of the in- ternational AAAI conference on web and social media , vol. 10, no. 1, 2016, pp. 191–200

  105. [113]

    Multivariate hawkes process for cyber insurance,

    Y . Bessy-Roland, A. Boumezoued, and C. Hillairet, “Multivariate hawkes process for cyber insurance,” Annals of Actuarial Science , vol. 15, no. 1, p. 14–39, 2021

  106. [114]

    Neural network based temporal point processes for attack detection in industrial control systems,

    G. Fortino, C. Greco, A. Guzzo, and M. Ianni, “Neural network based temporal point processes for attack detection in industrial control systems,” in 2022 IEEE international conference on cyber security and resilience (CSR). IEEE, 2022, pp. 221–226

  107. [115]

    Hawkes process modeling of adverse drug reactions with longitudinal observational data,

    Y . Bao, Z. Kuang, P. Peissig, D. Page, and R. Willett, “Hawkes process modeling of adverse drug reactions with longitudinal observational data,” in Machine learning for healthcare conference . PMLR, 2017, pp. 177–190

  108. [116]

    Detecting and quantifying causal associations in large nonlinear time series datasets,

    J. Runge, P. Nowack, M. Kretschmer, S. Flaxman, and D. Sejdinovic, “Detecting and quantifying causal associations in large nonlinear time series datasets,” Science advances, vol. 5, no. 11, p. eaau4996, 2019

  109. [117]

    Causal screening in dynamical systems,

    S. W. Mogensen, “Causal screening in dynamical systems,” in Con- ference on Uncertainty in Artificial Intelligence . PMLR, 2020, pp. 310–319

  110. [118]

    Learning granger causality for hawkes processes,

    H. Xu, M. Farajtabar, and H. Zha, “Learning granger causality for hawkes processes,” in International conference on machine learning . PMLR, 2016, pp. 1717–1726

  111. [119]

    Learning social infectivity in sparse low-rank networks using multi-dimensional hawkes processes,

    K. Zhou, H. Zha, and L. Song, “Learning social infectivity in sparse low-rank networks using multi-dimensional hawkes processes,” in Artificial Intelligence and Statistics , 2013, pp. 641–649

  112. [120]

    ℓ0-regularized sparsity for probabilistic mixture models,

    D. T. Phan and T. Id ´e, “ℓ0-regularized sparsity for probabilistic mixture models,” in Proceedings of the 2019 SIAM International Conference on Data Mining . SIAM, 2019, pp. 172–180

  113. [121]

    Causal discov- ery in hawkes processes by minimum description length,

    A. Jalaldoust, K. Hlav ´aˇckov´a-Schindler, and C. Plant, “Causal discov- ery in hawkes processes by minimum description length,” in Proceed- ings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 6, 2022, pp. 6978–6987

  114. [122]

    Granger causal inference in multivariate hawkes processes by minimum message length,

    K. Hlav ´aˇckov´a-Schindler, A. Melnykova, and I. Tubikanec, “Granger causal inference in multivariate hawkes processes by minimum message length,” Journal of Machine Learning Research , vol. 25, no. 133, pp. 1–26, 2024

  115. [123]

    Uncovering causality from multivariate hawkes integrated cumulants,

    M. Achab, E. Bacry, S. Ga ¨ıffas, I. Mastromatteo, and J.-F. Muzy, “Uncovering causality from multivariate hawkes integrated cumulants,” Journal of Machine Learning Research , vol. 18, no. 192, pp. 1–28, 2018

  116. [124]

    CAUSE: Learning granger causality from event sequences using attribution methods,

    W. Zhang, T. K. Panum, S. Jha, P. Chalasani, and D. Page, “CAUSE: Learning granger causality from event sequences using attribution methods,” in Proceedings of the 37th International Conference on Machine Learning, 2020

  117. [125]

    Learning granger causality from instance-wise self-attentive hawkes processes,

    D. Wu, T. Id ´e, G. Kollias, J. Navratil, A. Lozano, N. Abe, Y . Ma, and R. Yu, “Learning granger causality from instance-wise self-attentive hawkes processes,” in International Conference on Artificial Intelli- gence and Statistics . PMLR, 2024, pp. 415–423

  118. [126]

    Detection of short-term temporal dependencies in hawkes processes with heterogeneous background dynamics,

    Y . Chen, F. Li, A. Schneider, Y . Nevmyvaka, A. Amarasingham, and H. Lam, “Detection of short-term temporal dependencies in hawkes processes with heterogeneous background dynamics,” in Uncertainty in Artificial Intelligence . PMLR, 2023, pp. 369–380

  119. [127]

    Learning granger causality for non-stationary hawkes processes,

    W. Chen, J. Chen, R. Cai, Y . Liu, and Z. Hao, “Learning granger causality for non-stationary hawkes processes,” Neurocomputing, vol. 468, pp. 22–32, 2022

  120. [128]

    Thps: Topological hawkes processes for learning causal structure on event sequences,

    R. Cai, S. Wu, J. Qiao, Z. Hao, K. Zhang, and X. Zhang, “Thps: Topological hawkes processes for learning causal structure on event sequences,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 1, pp. 479–493, 2022

  121. [129]

    Causalnet: Unveiling causal structures on event sequences by topology-informed causal attention,

    H. Zhu, H. Huang, K. Yin, Z. Fan, H. Jin, and B. Liu, “Causalnet: Unveiling causal structures on event sequences by topology-informed causal attention,” in Proceedings of the IJCAI , 2024, pp. 7144–7152

  122. [130]

    A simple yet scalable granger causal structural learning approach for topological event sequences,

    M. Li, S. Liu, H. Qian, and A. Zhou, “A simple yet scalable granger causal structural learning approach for topological event sequences,” Advances in Neural Information Processing Systems , vol. 37, pp. 97 124–97 140, 2024

  123. [131]

    Cumulants of hawkes point processes,

    S. Jovanovi ´c, J. Hertz, and S. Rotter, “Cumulants of hawkes point processes,” Physical Review E , vol. 91, no. 4, p. 042802, 2015

  124. [132]

    Cumulants of hawkes processes are robust to observation noise,

    W. Trouleau, J. Etesami, M. Grossglauser, N. Kiyavash, and P. Thiran, “Cumulants of hawkes processes are robust to observation noise,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 444–10 454

  125. [133]

    Causal discovery from event sequences by local cause-effect attribution,

    J. C ¨uppers, S. Xu, A. Musa, and J. Vreeken, “Causal discovery from event sequences by local cause-effect attribution,” Advances in Neural Information Processing Systems , vol. 37, pp. 24 216–24 241, 2024

  126. [134]

    Structural hawkes processes for learning causal structure from discrete-time event sequences,

    J. Qiao, R. Cai, S. Wu, Y . Xiang, K. Zhang, and Z. Hao, “Structural hawkes processes for learning causal structure from discrete-time event sequences,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence , 2023, pp. 5702–5710

  127. [135]

    Lecture notes: Temporal point processes and the conditional intensity function,

    J. G. Rasmussen, “Lecture notes: Temporal point processes and the conditional intensity function,” arXiv preprint, 2018

  128. [136]

    Variational learning of inducing variables in sparse Gaus- sian processes,

    M. Titsias, “Variational learning of inducing variables in sparse Gaus- sian processes,” in Artificial Intelligence and Statistics , 2009, pp. 567– 574

  129. [137]

    Extending earthquakes’ reach through cascading,

    D. Marsan and O. Lengline, “Extending earthquakes’ reach through cascading,” Science, vol. 319, no. 5866, pp. 1076–1079, 2008. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 19

  130. [138]

    A nonparametric EM algorithm for multiscale Hawkes processes,

    E. Lewis and G. Mohler, “A nonparametric EM algorithm for multiscale Hawkes processes,” Journal of Nonparametric Statistics , vol. 1, no. 1, pp. 1–20, 2011

  131. [139]

    Nonparametric roughness penalties for probability densities,

    I. Goodd and R. A. Gaskins, “Nonparametric roughness penalties for probability densities,” Biometrika, vol. 58, no. 2, pp. 255–277, 1971

  132. [140]

    Learning triggering kernels for multi-dimensional Hawkes processes,

    K. Zhou, H. Zha, and L. Song, “Learning triggering kernels for multi-dimensional Hawkes processes,” in International Conference on Machine Learning, 2013

  133. [141]

    First-and second-order statistics character- ization of Hawkes processes and non-parametric estimation,

    E. Bacry and J.-F. Muzy, “First-and second-order statistics character- ization of Hawkes processes and non-parametric estimation,” IEEE Transactions on Information Theory , vol. 62, no. 4, pp. 2184–2202, 2016

  134. [142]

    Graphical modeling for multivariate Hawkes processes with nonparametric link functions,

    M. Eichler, R. Dahlhaus, and J. Dueck, “Graphical modeling for multivariate Hawkes processes with nonparametric link functions,” Journal of Time Series Analysis , vol. 38, no. 2, pp. 225–242, 2017

  135. [143]

    Adaptive estimation for Hawkes processes; application to genome analysis,

    P. Reynaud-Bouret, S. Schbath et al., “Adaptive estimation for Hawkes processes; application to genome analysis,” The Annals of Statistics , vol. 38, no. 5, pp. 2781–2822, 2010

  136. [144]

    Nonparametric estimation of hawkes processes with rkhss,

    A. Bonnet and M. Sangnier, “Nonparametric estimation of hawkes processes with rkhss,” arXiv preprint arXiv:2411.00621 , 2024

  137. [145]

    Poisson intensity estimation with reproducing kernels,

    S. Flaxman, Y . W. Teh, D. Sejdinovic et al. , “Poisson intensity estimation with reproducing kernels,” Electronic Journal of Statistics , vol. 11, no. 2, pp. 5081–5104, 2017

  138. [146]

    Long short-term memory,

    S. Hochreiter, “Long short-term memory,” Neural Computation MIT- Press, 1997

  139. [147]

    Empirical evaluation of gated recurrent neural networks on sequence modeling,

    J. Chung, C. Gulcehre, K. Cho, and Y . Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” arXiv preprint arXiv:1412.3555, 2014

  140. [148]

    RWKV: Rein- venting RNNs for the Transformer era,

    B. Peng, E. Alcaide, Q. G. Anthony, A. Albalak et al., “RWKV: Rein- venting RNNs for the Transformer era,” in Conference on Empirical Methods in Natural Language Processing , 2023

  141. [149]

    Efficiently modeling long sequences with structured state spaces,

    A. Gu, K. Goel, and C. Re, “Efficiently modeling long sequences with structured state spaces,” in International Conference on Learning Representations, 2022

  142. [150]

    Mamba: Linear-time sequence modeling with selective state spaces,

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint, 2023

  143. [151]

    Wasserstein learning of deep generative point process models,

    S. Xiao, M. Farajtabar, X. Ye, J. Yan et al. , “Wasserstein learning of deep generative point process models,” in Advances in Neural Information Processing Systems , 2017

  144. [152]

    Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics,

    M. U. Gutmann and A. Hyv ¨arinen, “Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics,” The journal of machine learning research , vol. 13, no. 1, pp. 307–361, 2012

  145. [153]

    Exploring the limits of language modeling,

    R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y . Wu, “Exploring the limits of language modeling,” arXiv preprint arXiv:1602.02410, 2016

  146. [154]

    Estimation of non-normalized statistical models by score matching,

    A. Hyv ¨arinen, “Estimation of non-normalized statistical models by score matching,” J. Mach. Learn. Res. , vol. 6, pp. 695–709, 2005

  147. [155]

    Investigating causal relations by econometric models and cross-spectral methods,

    C. W. Granger, “Investigating causal relations by econometric models and cross-spectral methods,” Econometrica: journal of the Econometric Society, pp. 424–438, 1969

  148. [156]

    Cardinality-regularized hawkes-granger model,

    T. Id ´e, G. Kollias, D. Phan, and N. Abe, “Cardinality-regularized hawkes-granger model,” Advances in Neural Information Processing Systems, vol. 34, pp. 2682–2694, 2021

  149. [157]

    On the predictive accuracy of neural temporal point process models for continuous-time event data,

    T. Bosser and S. B. Taieb, “On the predictive accuracy of neural temporal point process models for continuous-time event data,” arXiv preprint arXiv:2306.17066, 2023

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.