{"id":"779ffba2-0afe-4bb6-b25e-9b0cce71e872","arxiv_id":"2501.14291","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review that groups recent temporal point process research into Bayesian, neural, and LLM-based approaches and catalogs training, applications, and open challenges.","lead":"This paper surveys temporal point processes, the mathematics of event sequences in continuous time, from Bayesian, neural network, and large language model perspectives. It organizes recent model designs, training methods, and applications into a single reference guide for the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (11)'s branching likelihood is wrong: it replaces each event's exposure interval [t_i,T] with a fixed [0,Tφ], so the central inference recipe for GP-Hawkes models is misstated; a one-event counterexample confirms.","rationale":"Good-faith reading: this is a survey, not a new estimator, so the central claim is about providing a comprehensive and accurate map of TPP research. That claim is weakened if a core technical formula underlying one of the three pillars is wrong. Eq. (11) is load-bearing because Section III-B uses it to justify the entire branching-variable GP-Hawkes inference program, including the claim that the likelihood factorizes into two independent components that can each be handled with the GP-Poisson methods of Section III-A. The error changes the likelihood's dependence on event times and would mislead readers implementing or extending those methods. The reader's verdict was already CONDITIONAL, and this concern supports that verdict rather than moving it to ACCEPT or REJECT; a corrigendum to Eq. (11) and a qualification to the §II-C non-conjugacy claim would be sufficient to make the survey's technical core credible. I do not see a reason to reject the survey outright, since its taxonomy and literature coverage remain useful despite this mathematical imprecision.","tokens_in":29139,"tokens_out":7254,"duration_ms":70426,"concrete_test":"Set T=2, φ(τ)=1 on [0,Tφ] with Tφ=2, μ constant, and compare the augmented likelihood for one event at t1=0.5 vs t1=1.5. Eq. (11) yields the same value in both cases, μ e^{−2μ−2}. The correct branching likelihood gives μ e^{−2μ−1.5} for t1=0.5 and μ e^{−2μ−0.5} for t1=1.5, because the exposure integral is ∫_0^{T−t1} 1 dτ. If the two correct evaluations differ while Eq. (11) does not, the equation is missing the event-time-dependent exposure and the GP-Hawkes inference description in §III-B needs correction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In §III-B the survey presents the augmented branching likelihood. In the triggering component it writes ∏_{n=1}^N exp(−∫_0^{Tφ} φ(τ)dτ), independent of event times. The correct factor for event i is exp(−∫_0^{min(Tφ,T−t_i)} φ(τ)dτ), equivalently exp(−∫_{t_i}^T φ(t−t_i)dt) capped by the support. For a single event at time t1, the augmented likelihood is μ(t1) exp(−∫_0^T μ − ∫_0^{T−t1} φ); Eq. (11) gives μ(t1) exp(−∫_0^T μ − ∫_0^{Tφ} φ). These differ whenever T−t1<Tφ. This is not a cosmetic typo: §III-B's entire GP-Hawkes inference program depends on this factorization, using the branching variable X to decouple the posterior updates of μ(t) and φ(τ). Misstating this likelihood means the survey's account of Bayesian nonparametric Hawkes inference is mathematically incorrect, not merely incomplete. The rest of the survey's organization may still be useful, but the comprehensiveness claim requires an accurate core formula. A secondary overstatement, the blanket 'likelihood is not conjugate to any prior' claim in §II-C, should also be qualified, but the branching-likelihood error is the decisive issue.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews temporal point processes from three perspectives: Bayesian methods (parametric and nonparametric), neural TPPs (recurrent, autoregressive, and differential-equation based), and LLM-based TPPs (inspired, integrated, and multimodal extensions). It also covers model training objectives and applications in event prediction and causal discovery. The paper aims to provide a comprehensive and current update to earlier surveys, with particular emphasis on Bayesian nonparametric Hawkes processes and on work from 2020 through 2024.","tokens_in":29391,"tokens_out":7846,"duration_ms":76860,"significance":"If its formulas and characterizations are correct, the survey would be a valuable and current map of the TPP field: the organization is clear, the reference list is broad and up to date, and the treatment of Bayesian nonparametric Hawkes inference, SDE-based neural TPPs, and LLM-TPP integration fills genuine gaps in earlier surveys. The paper is expository and contains no fitting loop or circular derivation; its contribution is organizational and pedagogical. However, the correctness of the central augmented Hawkes likelihood in Section III-B is essential for the survey to serve as a reliable reference, and that equation is currently wrong.","major_comments":[{"comment":"Equation (11) states the triggering component of the augmented Hawkes likelihood as the product over n of exp(-∫_0^{Tφ} φ(τ)dτ), which does not depend on event times, and then claims that marginalizing X recovers the original Hawkes likelihood. This is incorrect. In the branching representation, each event t_j contributes a survival factor exp(-∫_{t_j}^T φ(t-t_j)dt) = exp(-∫_0^{min(Tφ,T-t_j)} φ(τ)dτ), because the triggering kernel from event j can only excite later times and the observation window ends at T. For a single event at t_1, Eq. (11) gives exp(-∫_0^{Tφ}φ(τ)dτ), whereas the correct marginalized likelihood has exp(-∫_0^{T-t_1}φ(τ)dτ), and the two differ whenever T-t_1<Tφ. The equation also omits the single-parent constraint ∑_{m≤n} x_{nm}=1. Since Section III-B builds its entire account of GP-Hawkes inference on this factorization, the formula should be corrected before the survey is used as a reference for Bayesian nonparametric Hawkes processes.","section":"III-B, Eq. (11)"},{"comment":"The blanket statement that \"the TPP likelihood is not conjugate to any prior\" is too strong. Conjugate analyses exist for basic Poisson process models, such as Gamma priors for a homogeneous Poisson intensity, and later sections of this paper rely on conditionally conjugate augmented likelihoods, e.g., the Pólya-Gamma augmentation discussed in Section III-A. The sentence should be qualified to refer to the general lack of convenient conjugacy for Hawkes or Cox-process likelihoods rather than to all TPPs.","section":"II-C"}],"minor_comments":[{"comment":"In Figure 1, the Wasserstein Distance entry cites [86,87], but Section VI-B cites [151] for the Wasserstein GAN framework and [87] for the later likelihood-free method. Reference [86] is a duplicate of [51], the RNN intensity model, and is not the Wasserstein result. The figure entry should likely read [151,87].","section":"Fig. 1 and Section VI-B"},{"comment":"The Introduction calls the survey comprehensive but does not describe the literature search or inclusion criteria. Given the claim of comprehensiveness, a short statement of the period, venues, and selection rules would help readers calibrate coverage and would make the taxonomy in Figure 1 easier to evaluate.","section":"Introduction"},{"comment":"There are several typographical issues: \"K-varaite\" in Section VII-B should be \"K-variate\"; the footnote in Section II-A contains a stray \"1\" after \"one-to-one\"; reference [139] misspells \"Good\" as \"Goodd\"; reference [100] misspells \"Conference\" as \"Confernce\". These should be corrected in a final pass.","section":"Various"},{"comment":"The claim that \"the earliest work, to the best of our knowledge, that integrates differential equations with point processes is Chen et al. [68]\" is worth softening or documenting. The statement is hedged, but the characterization of a Neural ODE paper as the first integration of differential equations with point processes is the kind of historical claim that readers may rely on; a citation to the relevant section of [68] or a more explicit caveat would strengthen it.","section":"IV-C"}],"recommendation":"major_revision","confidential_remarks":"The main problem is confined to one load-bearing equation, but it is a genuine mathematical error in the central formula of Section III-B, so the manuscript needs a revision rather than a minor edit. I do not see evidence of circularity or citation manipulation, and the survey's overall organization is sound once Eq. (11) is corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: if you want one place to get oriented on temporal point processes as of 2025, including the new LLM-based work, this survey is worth reading. The Bayesian nonparametric and LLM sections are genuinely new relative to Yan (2019) and Shchur et al. (2021), and the taxonomy into recurrent, autoregressive, and differential-equation neural TPPs is serviceable. The authors know the literature and give credit where it is due.\n\nThe soft spot is real and it sits in a load-bearing place. In Section III-B, Eq. (11) writes the augmented branching likelihood with the triggering-component survival factor as exp(-∫_0^{T_φ} φ(τ)dτ) for every event, independent of the event time. That is not the correct factor. For an event at t_i, the exposure interval for triggering ends at T - t_i (capped by the support T_φ), not at T_φ. The stress-test is right: a single event at t_1 separates the two expressions. This matters because the whole subsequent discussion of GP-Hawkes inference in that section rests on this factorization. It is a mathematical error in the central formula, not a typo.\n\nThere is also a secondary overstatement in Section II-C: the claim that the TPP likelihood is not conjugate to any prior. Taken literally, that is false; later in Section III-A the authors themselves discuss conditionally conjugate augmentation schemes (e.g., Pólya-Gamma). The claim should be qualified.\n\nTwo more minor points. The survey asserts comprehensiveness but includes no documented corpus-selection protocol, so coverage is whatever the authors chose to include. That is common in surveys and not disqualifying, but it means the 'comprehensive' claim is weaker than it sounds. And the attribution in Section IV-C that Chen et al. [68] is the earliest work integrating ODEs with point processes is a stretch; Neural ODEs is not really a point-process paper, and there are earlier precedents.\n\nNet: the paper's organization and coverage are useful, especially for the LLM-TPP subarea, and the authors clearly know the material. The branching-likelihood error is fixable but it is not cosmetic. I would send this to peer review, ask the authors to correct Eq. (11) and soften the non-conjugacy claim, and after that it is a reasonable survey to accept.","headline":"A timely and useful survey of TPPs, with an excellent LLM section and a solid Bayesian nonparametric roadmap, but the branching likelihood in Eq. (11) is wrong and must be fixed.","tokens_in":29936,"tokens_out":2497,"would_cite":false,"duration_ms":22145,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60G55","62M09","62F15","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey tries to establish a complete map of temporal point process research as of 2020-2025, arguing that the field splits into Bayesian, neural, and LLM-based approaches that all revolve around the same conditional-intensity object.","keywords":["temporal point processes","conditional intensity function","Hawkes process","Bayesian nonparametric inference","neural temporal point processes","Transformer attention","large language models","noise-contrastive estimation"],"falsifier":"Check whether a published temporal point process model with an ODE- or SDE-driven intensity antedates the work named as first in Section IV-C, since that priority claim is falsifiable; likewise, a complete enumeration of TPP papers from 2020 to 2025 would either confirm or break the Figure 1 taxonomy, because the survey reports no systematic corpus-selection protocol.","tokens_in":28941,"feed_emoji":"🗺️","tokens_out":7871,"duration_ms":71207,"temperature":0.7,"pith_summary":"Temporal point processes (TPPs) model event sequences whose timestamps are irregular, such as tweets, trades, neural spikes, and earthquakes, and this survey tries to establish where the field stands in 2020-2025. The authors argue that current research falls into three coherent camps: Bayesian TPPs, mostly nonparametric and placing priors on the intensity function; neural TPPs, using recurrent, autoregressive, and differential-equation architectures; and LLM-based TPPs, using language models for prediction, reasoning, retrieval, and multimodal events. Across all three camps, they claim the same mathematical object, the conditional intensity function, and the same estimation question recur, with training objectives ranging from maximum likelihood to Wasserstein, noise-contrastive, and score-matching criteria. The survey positions itself as an update to earlier reviews that stopped around 2020 and largely ignored Bayesian nonparametric TPPs, so it aims to be the place to see the current shape of the field and its open problems.","feed_headline":"One survey maps temporal point processes into three research camps","feed_subtitle":"From Gaussian-process priors to LLM encoders, the field's models, training criteria, and open problems in one picture.","key_machinery":"The load-bearing object is the conditional intensity function $\\lambda^*(t)$, the expected number of events per unit time given the event history, together with its one-to-one correspondence to the conditional density $f(t \\mid H_{t_n}) = \\lambda^*(t) \\exp(-\\int_{t_n}^t \\lambda^*(\\tau)\\,d\\tau)$. Because every TPP can be specified either through this function or through the next-event distribution, the whole survey rests on that equivalence. Three secondary mechanisms carry the three branches: Gaussian process priors on the intensity, split into independent Poisson components via a branching latent variable for Bayesian nonparametrics; learned history embeddings $h_n$, whether updated recurrently, computed by self-attention, or evolved continuously by an ODE or SDE with jumps, for neural TPPs; and the divergence-minimization objective $D(f(T) \\| f_\\theta(T))$ that defines training across all frameworks.","core_discovery":"The central claim this survey asserts is a taxonomy: recent TPP research is organized by modeling philosophy into Bayesian, neural, and LLM-based approaches, and each branch has a small number of canonical mechanisms. Bayesian nonparametric TPPs treat the intensity function as an infinite-dimensional parameter under a Gaussian process or Dirichlet process prior, with the doubly intractable posterior handled by MCMC, Laplace, or variational methods; for Hawkes processes, a branching latent variable factorizes the likelihood so the baseline and triggering functions can be estimated separately. Neural TPPs divide architecturally into recurrent models with event-by-event history updates, autoregressive models based on Transformer self-attention, and differential-equation models with hidden states that flow between events and jump at events, alongside a thread on equivalent parameterizations where modeling the cumulative intensity turns the log-likelihood integral into a derivative. LLM-based TPPs are treated as a distinct emerging branch that includes prompt-inspired continual learning, direct fusion of LLMs with temporal encoding, retrieval, and multimodal benchmarks. The survey further claims that architectures are only half the story: the other half is the training objective, and it reviews KL divergence and maximum likelihood, Wasserstein distance, noise-contrastive estimation, and Fisher divergence as alternative answers to the same distribution-matching problem.","pith_inferences":["A testable consequence the authors leave implicit: if the taxonomy is right, nearly every neural TPP paper from 2020 to 2025 should fall into exactly one of the three architectural families, so a systematic coding of the literature would measure how much of the field the survey actually covers.","The paper's emphasis on the equivalence among intensity, density, and cumulative-intensity parameterizations suggests a natural comparison it does not run: hold the architecture fixed and train the same TPP with MLE, Wasserstein, noise-contrastive, and score-matching objectives on identical datasets to see which estimator pays for its variance.","Since the survey notes that earlier score-matching estimators for point processes were shown to be incomplete, an immediate open question is whether the corrected weighted estimator also fixes the earlier spatio-temporal and multivariate variants.","The LLM branch may progressively dissolve into the neural branch, because LLM-based TPPs are themselves neural TPPs; the three-way split is best read as a snapshot of research momentum rather than a permanent division."],"forward_implications":["A new TPP modeler can be located in the field by two choices: which family supplies the history representation, and which parameterization (intensity, cumulative intensity, density, or quantile) defines the next-event distribution.","Bayesian nonparametric Hawkes inference reduces to a template: augment a branching variable, then alternate between updating the branching posterior and updating the two Gaussian-process-modulated Poisson components, a recipe the survey identifies across many papers.","Modeling the cumulative conditional intensity $\\Lambda^*(t)$ with monotonic networks removes numerical integration from maximum likelihood estimation and turns the log-likelihood into a derivative, a trick that applies beyond any single architecture.","The taxonomy implies that the LLM branch is not a separate mathematical framework but a new interface for the same conditional-intensity object, adding text, retrieval, and multimodal context to event sequences.","The stated lack of a standard benchmark protocol means cross-paper performance claims in neural TPPs should be read with caution until a consistent evaluation setup becomes standard practice."],"supporting_citations":[{"why":"Foundational theory that supplies the definition of point processes and the conditional intensity function on which the entire survey rests.","marker":"[1]"},{"why":"Introduces the self-exciting Hawkes process, the classical history-dependent model that anchors the statistical, neural, and causal-discovery discussions.","marker":"[6]"},{"why":"A prior neural TPP review covering work before 2020, which the survey positions its neural section as updating.","marker":"[9]"},{"why":"Introduces the sigmoidal Gaussian Cox process with MCMC inference, the starting point of the Bayesian nonparametric Poisson branch.","marker":"[19]"},{"why":"Introduces Polya-Gamma augmentation to Bayesian nonparametric TPPs, enabling conditionally conjugate Gibbs, EM, and variational inference.","marker":"[33]"},{"why":"The first recurrent neural TPP, anchoring the recurrent family and the history-embedding idea.","marker":"[49]"},{"why":"The Transformer Hawkes process, anchoring the autoregressive and self-attention family.","marker":"[59]"},{"why":"Neural jump stochastic differential equations, anchoring the continuous-time differential-equation family.","marker":"[70]"},{"why":"A direct LLM-TPP integration approach that defines the LLM branch's strategy for encoding temporal information.","marker":"[80]"},{"why":"A large-scale empirical comparison under a consistent setup, cited as evidence for the lack of a standard benchmark protocol.","marker":"[157]"}],"fun_headline_variants":["Three frameworks, one survey: Bayesian, neural, and LLM TPPs","Event sequences reimagined: TPP survey covers three paradigms","From GP priors to LLM encoders: TPP survey unites all","Temporal point processes: Bayesian, neural, and now LLM mindsets","One taxonomy to rule them all: Bayesian, neural, and LLM TPPs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim to be comprehensive rests on the assumption that its three-way taxonomy, and the placement of individual methods within it, accurately reflects the entire field; if a significant line of research is omitted or a key attribution is wrong, the map loses its value as a guide.","fun_headline_variants_meta":{"raw":{"variants":["Three frameworks, one survey: Bayesian, neural, and LLM TPPs","Event sequences reimagined: TPP survey covers three paradigms","From GP priors to LLM encoders: TPP survey unites all","Temporal point processes: Bayesian, neural, and now LLM mindsets","One taxonomy to rule them all: Bayesian, neural, and LLM TPPs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1640,"prompt_tokens":969,"completion_tokens":671,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":569}},"tokens_in":585,"tokens_out":671,"duration_ms":6199,"temperature":1.0,"reasoning_tokens":569,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:14:26.874977+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether a published temporal point process model with an ODE- or SDE-driven intensity antedates the work named as first in Section IV-C, since that priority claim is falsifiable; likewise, a complete enumeration of TPP papers from 2020 to 2025 would either confirm or break the Figure 1 taxonomy, because the survey reports no systematic corpus-selection protocol.","supporting_citations":[{"cited_title":"On the Predictive Accuracy of Neural Temporal Point Process Models for Continuous-time Event Data","cited_arxiv_id":"2306.17066","evidence_quote":"A large-scale empirical comparison under a consistent setup, cited as evidence for the lack of a standard benchmark protocol."}],"review_version":1}