Pith. sign in

REVIEW 3 major objections 5 minor 167 references

Model-free Methods for Event History Analysis and Efficient Adjustment (PhD Thesis)

T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This thesis establishes that a model-free parameter called the Local Covariance Measure can be estimated by double machine learning so that, under conditional local independence, the estimator converges uniformly at a $\sqrt{n}$ rate to a…

desk verdict The X-LCT in Chapter 2 is a genuine advance in nonparametric event-history testing; read it for that result, but note that the rate conditions for the recommended nuisance estimators remain unverified. read the letter →

arxiv 2502.07906 v1 pith:XKH6OEMV submitted 2025-02-11 stat.ME math.STstat.MLstat.TH

classification stat.MEmath.STstat.MLstat.TH MSC 62G1062G0562G2062N0162N02
keywords conditionallocalindependenceCovarianceMeasuredoublemachinelearningcountingprocessnonparametrictestingcovariateadjustmentefficiencybounddoublyrobustestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This thesis argues that many statistical targets can be defined model-free, meaning they have a clear meaning for any data-generating distribution instead of only inside a parametric model, and can still be estimated at parametric speed by combining machine-learning nuisance estimates with double machine learning. Its main contribution is the Local Covariance Measure, a time-indexed covariance between a residual process and the compensated counting-process martingale, which is identically zero under the hypothesis of conditional local independence. The thesis proves that a cross-fitted estimator of this measure converges uniformly at a $\sqrt{n}$ rate to a mean-zero Gaussian martingale under the null, provided the nuisance functions are consistent with modest rates, and uses this to build the Local Covariance Test with uniform asymptotic level and power against local alternatives. The same model-free logic yields the Debiased Outcome-adapted Propensity Estimator for efficient covariate adjustment and the Aalen Covariance Measure for assumption-lean regression in event-history settings.

What carries the argument

The load-bearing object is the residual process $G_t$, for example the additive residual $X_t - E[X_t\mid\mathcal{F}_{t-}]$, together with the stochastic integral $I_t=\int_0^t G_s\,dM_s$. The Local Covariance Measure is $\gamma_t=E[I_t]$, and its defining property is that residualization makes the integrand orthogonal to the past, giving the estimator a Neyman orthogonal structure: substituting slow, nonparametric nuisance estimates does not create a first-order bias. Under the alternative, $\gamma_t=\int_0^t \mathrm{Cov}(G_s,\lambda_s-\underline{\lambda}_s)\,ds$, so the LCM is nonzero exactly when the two intensities differ in a way the residual can detect. The proof machinery combines a uniform extension of the functional martingale central limit theorem, chaining arguments for stochastic equicontinuity, sample splitting and cross-fitting to remove dependence, and the empirical variance estimator $\hat V_n(t)=|J_n|^{-1}\sum_j\int_0^t(\hat G_{j,s})^2\,dN_{j,s}$.

What would settle it

Simulate data from the Cox-type model of Section 2.6 under the null with the historical functional linear model and compute the $L^2$ error product $\sqrt{|J_n|}\,g(n)h(n)$ for the kernel-based estimates of $\lambda$ and $\Pi$; if this product fails to converge to zero, the uniform level and power of the X-LCT are not guaranteed.

Watch

Extended reading notes

Core claim

Under the hypothesis that a counting process $N$ is conditionally locally independent of a c\`adl\`ag process $X$ given a filtration $(\mathcal{F}_t)$, the Local Covariance Measure $\gamma_t = E[\int_0^t G_s\,dM_s]$ is the zero function, where $G_t$ is a residual process satisfying $E[G_t\mid\mathcal{F}_{t-}]=0$ and $M_t=N_t-\int_0^t\lambda_s\,ds$ is the compensated martingale. The thesis shows that estimating $\gamma$ by plugging machine-learning estimates of the intensity $\lambda$ and of the residual map into the stochastic integral, with sample splitting or cross-fitting, yields a process that is asymptotically indistinguishable from a Gaussian martingale: under Assumptions 2.4.1 and 2.4.2, $\sqrt{|J_n|}(\hat\gamma^{(\cdot)}-\gamma)$ converges uniformly to a mean-zero Gaussian martingale, and the resulting test statistic converges to the supremum of a standard Brownian motion. For covariate adjustment, the thesis identifies the optimal subset of covariates and shows that DOPE attains the semiparametric efficiency bound; for event-history association, it shows that the Aalen Covariance Measure is estimable in a doubly robust way at a $\sqrt{n}$ rate under modest nuisance rates.

Load-bearing premise

The load-bearing premise is that the machine-learning estimates of the conditional intensity and the residual process are accurate enough that the product of their $L^2$ errors shrinks faster than $n^{-1/2}$; for the historical functional linear model used in the simulations, the thesis states this rate is only conjectured, not rigorously established.

Editorial extensions

If this is right

  • Researchers can test whether a time-varying exposure directly drives an event process without committing to a parametric hazard model, as long as the nuisance functions can be estimated at the required rates.
  • The X-LCT critical values are distribution-free: under the null the statistic converges to $\sup_{0\le t\le 1}|B_t|$ for a standard Brownian motion, so no simulation of the null distribution is needed.
  • The uniform asymptotic level means the test can serve as a subroutine in constraint-based procedures for learning local independence graphs, replacing an oracle test with a practical one.
  • For covariate adjustment, DOPE provides a single estimator that attains the efficiency bound no matter which subset of covariates is used, and it identifies the optimal subset.
  • The Aalen Covariance Measure gives an assumption-lean replacement for the exposure coefficient in an Aalen additive hazards model, so effect estimates retain a clear interpretation when the model is misspecified.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the uniform asymptotic theory holds for counting processes, the same proof pattern should extend to tests for local independence in general semimartingale systems, provided nuisance estimation rates can be established; the discrete-time observation issue would be the main obstacle.
  • The residual process is a user choice: transformations, time shifts, and linear or nonlinear filters of $X$ can be inserted, which the paper notes should shift power toward the corresponding departure from the null; a systematic power comparison across such filters is a natural next step.
  • The rate condition $\sqrt{|J_n|}g(n)h(n)\to 0$ means a fast rate for one nuisance can compensate for a slow rate for the other; a practical benchmark would be to verify the conjectured kernel-estimation rates for the historical functional linear model used in the simulations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This PhD thesis (arXiv:2502.07906) contains a methodological introduction and three stand-alone papers. Chapter 2 defines the Local Covariance Measure (LCM), a time-indexed functional parameter that quantifies deviations from conditional local independence of a counting process, and proposes estimation by double machine learning with sample splitting or cross-fitting, leading to the (X-)LCT test. The paper proves, under Assumptions 2.4.1, 2.4.2, and 2.5.1, uniform weak convergence of the estimator to a mean zero Gaussian martingale under the null, uniform asymptotic level, and power against sqrt(n) local alternatives for the additive residual process. Chapter 3 presents the Debiased Outcome-adapted Propensity Estimator (DOPE) for efficient covariate adjustment, and Chapter 4 introduces the Aalen Covariance Measure (ACM) with double robustness. The full text supplied for review covers Chapter 1 and Chapter 2 in detail; my detailed assessment focuses on those chapters.

Significance. If the results hold as stated, Chapter 2 is a substantial contribution: it is the first nonparametric test of conditional local independence with explicit uniform asymptotic guarantees, and it extends double machine learning from scalar parameters to a time-indexed functional parameter. The proof machinery, including a uniform version of Rebolledo's martingale CLT and uniform chaining arguments in Section 2.B, is of independent interest and is presented carefully. Chapter 1 gives a clear, self-contained introduction to Neyman orthogonality and rate double robustness. The theorems are carefully stated as conditional results, and the thesis openly discusses the limitations of the available rate theory; however, the central gap between the rate assumptions in the theorems and the practical estimators used in the simulation study needs to be addressed in the thesis version.

major comments (3)
  1. [§2.D and §2.6.1, Assumption 2.4.2] Assumption 2.4.2 is load-bearing for Theorem 2.4.6, Theorem 2.5.1, and Theorem 2.5.4, because the remainder R_3^{(n)} in (2.4.21) is controlled only by the product rate sqrt(|J_n|) g(n) h(n) -> 0. Yet Section 2.D states that for the full historical functional linear model no published rate results are available, and Section 2.6.1 states that sufficient rate results 'should be possible but have not yet been established rigorously.' The simulation study in Section 2.6.2 uses exactly this model to estimate both rho_X and rho_Y (see (2.6.36) and the four kernels), so the implemented X-LCT in Algorithm 2 is not verified to satisfy the assumptions of the theorems. If the true rates are slower than n^{-1/4+epsilon}, the uniform level and power conclusions do not apply to the recommended implementation. The thesis should either establish the required rates under explicit regularity conditions for the historical functional linear model, or present the uniform guarantees as conditional on Assumption 2.4.2 and describe the simulations as illustrative rather than confirmatory.
  2. [§2.5.1, §2.5.2, Theorems 2.5.2 and 2.5.4] The uniform power result Theorem 2.5.2 is proved only for the sample-split LCT with the additive residual process, while the cross-fitted X-LCT of Definition 2.5.3, which the paper recommends for practical use, has only the uniform level statement in Theorem 2.5.4. No theorem gives uniform power for the cross-fitted test Psi^K_n; its power is supported only by the simulation study in Section 2.6.4. If the proof extends to cross-fitting, the extension should be stated; otherwise the abstract and Section 2.7 should specify that the uniform power guarantee applies to the sample-split version and that the X-LCT's power is an empirical finding.
  3. [§2.D, §2.B, Assumption 2.4.2] Assumption 2.4.2 requires convergence uniformly over the parameter set Theta. The rate survey in Section 2.D gives only pointwise rates: for example, the n^{-(1+m)/(2m+3)} heuristic for the historical functional linear model is derived from a fixed-t prediction error bound in Cai and Yuan (2012), and the convolution-model rate from Manrique (2016) is likewise stated for a fixed model. No argument in Section 2.D establishes uniformity over theta in Theta or over t in [0,1]. The simulation settings vary kernels and beta_2 over finite grids in Section 2.6.2, which does not fill this gap. The manuscript should state which additional regularity conditions on Theta, such as smoothness classes with uniform constants, make the uniform convergence in Assumption 2.4.2 hold.
minor comments (5)
  1. [Abstract and Sammenfatning] The text contains a rendering error: '?n-consistency' appears where '\sqrt{n}-consistency' is intended, both in the English abstract and the Danish summary.
  2. [§2.6.2, Assumption 2.4.1] The simulation uses unbounded Gaussian innovations, while Assumption 2.4.1 requires uniformly bounded processes; the text says the reported results were generated without caps. Please either report a capped version or state explicitly that the simulations are intended as an approximation to, rather than a verification of, Assumption 2.4.1.
  3. [§2.5.2 and §2.6.4] The main text does not state the number of folds K used for the X-LCT in the simulation study; please give K in the main text or confirm the value given in Section 2.G.1.
  4. [§2.3.2] The statement that the statistic in (2.3.16) is 'closely related' to the partial copula could be made precise; it would be useful to specify the exact transformation and whether the asymptotic results of Petersen and Hansen (2021) apply directly to this version.
  5. [Chapter 1, Example 1.0.1] The notation uses P both for a single distribution and for a collection of distributions in the same section (e.g., 'P in \mathcal{P}'); a brief notation table or a change of symbol for the collection would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the X-LCT null distribution follows from a martingale CLT, not from the fitted nuisance estimators; the unverified rate condition is an honest limitation, not a circular reduction.

full rationale

The central derivation chain is self-contained. The LCM is defined as a functional of the true distribution, and the estimator is analyzed through the explicit decomposition (2.4.17), where the leading term U^(n) uses the true residual process and the true martingale, while the nuisance-dependent remainders R_1, R_2, R_3 and D_2 are shown to vanish under Assumptions 2.4.1 and 2.4.2. The null distribution in Theorems 2.5.1 and 2.5.4 follows from Rebolledo's martingale CLT together with the Brownian time-change and scale invariance property, so the test is not constructed to have its null distribution by definition. The paper explicitly flags, in Section 2.D and Section 2.7, that rigorous rate results for the historical functional linear model used in the simulations have not yet been established: 'For the full historical functional linear model we are not aware of any published rate results' and 'we regard it is as an independent research project to establish rates for general historical regression methods.' This is a stated limitation and a correctness/robustness gap for the recommended practical implementation, but it is not a circular reduction of the theoretical claim: the theorems are conditional on Assumption 2.4.2, not derived from the conjecture that the assumption holds. Self-citations to the published version [Christgau et al., 2023b] and to the supplement [Christgau et al., 2023c] are bibliographic references to the same results included in the thesis, and no load-bearing argument reduces to an unverified self-citation. The proofs and technical lemmas are contained in the included supplement, and the paper builds on external, established results for the martingale CLT and double machine learning. Therefore no circularity is present; the main risk is the unverified rate condition, which belongs to correctness assessment rather than to circularity analysis.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The main assumptions are the i.i.d. data-generating process, the martingale definition of local independence, boundedness and rate conditions on nuisance estimators, and usual causal assumptions for DOPE and ACM. No free parameters fitted to data appear in the derivations; hyperparameters like the number of cross-fitting folds K are user choices and do not affect the asymptotic claims. No new physical or ontological entities are proposed; the new statistical functionals (LCM, ACM, DOPE) are defined directly in terms of observable distributions.

assumptions (6)
  • domain assumption Observations are i.i.d. according to an unknown distribution P.
    Chapter 1 states 'it will be assumed throughout this thesis' (Section 1). This is standard for the statistical theory developed.
  • domain assumption Local independence is defined via the martingale property: N_t - Lambda_t remains a martingale after enlarging the filtration.
    Definition 2.2.1 in Chapter 2. This is the formal mathematical meaning of the target concept.
  • ad hoc to paper Assumption 2.4.1: The intensity process and the residual process are uniformly bounded almost surely on [0,1].
    This strong boundedness assumption is introduced to make the martingale CLT and remainder bounds tractable. The authors note it can be relaxed at the cost of more complex proofs (Section 2.7).
  • ad hoc to paper Assumption 2.4.2: Nuisance estimators satisfy g(n), h(n) -> 0 and sqrt(n) g(n) h(n) -> 0 uniformly.
    This is the key rate double robustness assumption. It is not verified for the historical functional linear model estimators used in the simulation; Section 2.D says such rate results are scarce.
  • ad hoc to paper Assumption 2.5.1: The asymptotic variance function V(t) is bounded away from zero at t=1 uniformly over the parameter set.
    Needed to avoid a degenerate null distribution for the supremum-based test statistic.
  • domain assumption For DOPE and ACM: Standard causal identification conditions such as positivity, consistency, and no unmeasured confounding.
    These are not stated in the provided text but are typical for adjustment estimators. We cannot verify their exact form from the abstract only.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Model-free Methods for Event History Analysis and Efficient Adjustment (PhD Thesis)." pith.science (2026). https://pith.science/paper/XKH6OEMV

@misc{pith2026250207906,
  author       = {Pith},
  title        = {Pith review of: Model-free Methods for Event History Analysis and Efficient Adjustment (PhD Thesis)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XKH6OEMV}},
  note         = {Machine review of arXiv:2502.07906}
}
abstract

This thesis contains a series of independent contributions to statistics, unified by a model-free perspective. The first chapter elaborates on how a model-free perspective can be used to formulate flexible methods that leverage prediction techniques from machine learning. Mathematical insights are obtained from concrete examples, and these insights are generalized to principles that permeate the rest of the thesis. The second chapter studies the concept of local independence, which describes whether the evolution of one stochastic process is directly influenced by another. To test local independence, we define a model-free parameter called the Local Covariance Measure (LCM). We formulate an estimator for the LCM, from which a test of local independence is proposed. We discuss how the size and power of the proposed test can be controlled uniformly and investigate the test in a simulation study. The third chapter focuses on covariate adjustment, a method used to estimate the effect of a treatment by accounting for observed confounding. We formulate a general framework that facilitates adjustment for any subset of covariate information. We identify the optimal covariate information for adjustment and, based on this, introduce the Debiased Outcome-adapted Propensity Estimator (DOPE) for efficient estimation of treatment effects. An instance of DOPE is implemented using neural networks, and we demonstrate its performance on simulated and real data. The fourth and final chapter introduces a model-free measure of the conditional association between an exposure and a time-to-event, which we call the Aalen Covariance Measure (ACM). We develop a model-free estimation method and show that it is doubly robust, ensuring $\sqrt{n}$-consistency provided that the nuisance functions can be estimated with modest rates. A simulation study demonstrates the use of our estimator in several settings.

Figures

Figures reproduced from arXiv: 2502.07906 by the authors.

Figure 2.1
Figure 2.1. 1: Local independence graph illustrating a dependence structure among the three processes X, Z and N. Here N is the indicator of death for an individual, X is their cumulative pension savings and Z is a covariate process. All nodes in this graph have implicit self-loops. There is no edge from X to N, which indicates that death is not directly influenced by pension savings. This can be formalized as N being condition… view at source ↗
Figure 2.2
Figure 2.2. 2: Local independence graphs illustrating how the three processes X, Y , and Z could affect each other and time of death in the Cox example. There is no direct influence of X (pension savings) on time of death in either of the two graphs, but in the left graph the death indicator is furthermore conditionally locally independent of X given the history of Z and N. In the right graph, Z and N does not block all paths f… view at source ↗
Figure 2.2
Figure 2.2. 3: Histograms of the distributions of three different estimators of γ1. Each histogram contains 1000 estimates fitted to samples of size n “ 500. The samples were sam￾pled from a model that satisfies the hypothesis of conditional local independence and hence the ground truth is γ1 “ 0. See Section 2.6.2 for further details of the data generating process. However, we cannot expect the plug-in estimator to have a ? n-… view at source ↗
Figures from the paper (16 more)
Figure 2.2
Figure 2.2. Figure 2.2: 4: A time dependent extension of [PITH_FULL_IMAGE:figures/full_fig_p041_2_2.png]
Figure 2.6
Figure 2.6. Figure 2.6: 5: Empirical cumulative distribution functions of simulated p-values for the cross-fitted local covariance test and the hazard ratio test. The simulated data satisfies the hypothesis of conditional local independence, so the p￾values are supposed to be uniformly dist…
Figure 2.6
Figure 2.6. Figure 2.6: 6: For each ρ0 P t0, 5, 10u, the lines show the average rejection rates of our proposed test X-LCT (blue) and the hazard ratio test (orange) as functions of sample size, with each average taken over 8 different settings. For each setting, the rejection rate is comput…
Figure 2
Figure 2. Figure 2: G.1: Empirical distribution functions of p-values for the three different condi￾tional local independence tests considered, simulated under the sampling scheme described in Section 2.6. The dotted line shows y “ x correspond￾ing to a uniform distribution. 91 [PITH_FUL…
Figure 2
Figure 2. Figure 2: G.2: Sample paths of ˇγ K,p500q fitted on data sampled from three different alter￾natives as described in Section 2.6.4. Here pX, Y, Zq are sampled from the scheme described in Section 2.6, with both ρX and ρY being the constant kernel and with β “ ´1. For each alterna…
Figure 2
Figure 2. Figure 2: G.4: The plots show the average rejection rate of the double machine learning tests based on the supremum statistic (blue) and the endpoint statistic (red). n “ 2000. This is different from the previous settings, and can be explained by a slower convergence of the inte…
Figure 3.1
Figure 3.1. Figure 3.1: 1: The covariate W can have a complex data structure, even if the information it represents is structured and can be categorized into components that influence treatment and outcome separately. confounding. If the underlying confounding mechanisms are captured by a s…
Figure 3.3
Figure 3.3. Figure 3.3: 2: The σ-algebra Q given in Definition 3.3.2 as a description of W. Definition 3.3.2. Define the σ-algebras Q – ł PPP QP , QP – σpFpy |t,W; Pq; y P R, t P Tq, R – ł PPP RP , RP – σpbtpW; Pq; t P Tq. ♣ Note that QP , Q, RP and R are all descriptions of W, see [PITH_F…
Figure 3.5
Figure 3.5. Figure 3.5: 3: Root mean square errors for various estimators of χ1 plotted against sample size. Each data point is an average over 900 datasets, 300 for each choice of d P t4, 12, 36u. The bands around each line correspond to ˘1 standard error. For this plot, the outcome regres…
Figure 3.5
Figure 3.5. Figure 3.5: 4: Root mean square errors for various estimators of χ1 plotted against sample size. Each data point is an average over 900 datasets, 300 for each choice of d P t4, 12, 36u. The bands around each line correspond to ˘1 standard error. For this plot, the outcome regres…
Figure 3.5
Figure 3.5. Figure 3.5: 5: Coverage of asymptotic confidence intervals of the adjusted mean χ1, ag￾gregated over d P t4, 12, 36u and with the true β given in (3.5.19). 120 [PITH_FULL_IMAGE:figures/full_fig_p132_3_5.png]
Figure 3.5
Figure 3.5. Figure 3.5: 6: Distribution of estimated propensity scores based on the full covariate set W, the single-index WJ ˆθ, and the outcome predictions gpp1,Wq, respec￾tively. we see that the full covariate set yields propensity scores that essentially violate positiv￾ity, with this e…
Figure 3
Figure 3. Figure 3: B.1: Neural network architecture for single-index model. 3.B. Details of simulation study Our experiments were conducted in Python [Van Rossum et al., 1995]. The linear and logistic regression was imported from the scikit-learn package [Pedregosa et al., 2011], and the…
Figure 3
Figure 3. Figure 3: B.2: Comparison of DOPE estimators with and without sample splitting. Estimator Estimate BS se BS CI Regr. (NN) 0.020 0.008 (0.002, 0.035) Regr. (Logistic) 0.027 0.009 (0.012, 0.048) DOPE-BCL (Logistic) 0.024 0.010 (0.004, 0.040) Naive contrast 0.388 0.010 (0.369, 0.40…
Figure 3
Figure 3. Figure 3: B.3: Coverage of asymptotic confidence intervals of the adjusted mean χ1, ag￾gregated over d P t4, 12, 36u and with the true β given in (3.5.19). The numbers in the parentheses denote the corresponding average width of the confidence interval. 138 [PITH_FULL_IMAGE:fig…
Figure 4.5
Figure 4.5. Figure 4.5: 1: Scaled RMSE for various estimators with respect to the cumulative direct effect. 161 [PITH_FULL_IMAGE:figures/full_fig_p173_4_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

167 extracted references · 58 canonical work pages

  1. [1]

    O. Aalen. A model for nonparametric regression analysis of counting processes. In Mathematical Statistics and Probability Theory: Proceedings, Sixth International Conference, Wis a (Poland), 1978 , pages 1--25. Springer, 1980

  2. [2]

    O. O. Aalen. Dynamic modelling and causality. Scandinavian Actuarial Journal, pages 177--190, 1987

  3. [3]

    O. O. Aalen. A linear regression model for the analysis of life times. Statistics in Medicine, 8 0 (8): 0 907--925, 1989

  4. [4]

    O. O. Aalen, K. R ysland, J. M. Gran, and B. Ledergerber. Causality, mediation and time: a dynamic viewpoint. Journal of the Royal Statistical Society. Series A (Statistics in Society), 175 0 (4): 0 831--861, 2012

  5. [5]

    Achab, E

    M. Achab, E. Bacry, S. Ga\" ffas, I. Mastromatteo, and J.-F. Muzy. Uncovering causality from multivariate H awkes integrated cumulants. In Proceedings of the 34th International Conference on Machine Learning, volume 70, pages 1--10. PMLR, 06--11 Aug 2017

  6. [6]

    R. J. Adler and J. E. Taylor. Random Fields and Geometry, volume 80 of Springer Monographs in Mathematics. Springer, 2007

  7. [7]

    P. K. Andersen, . Borgan, R. D. Gill, and N. Keiding. Statistical models based on counting processes. Springer Series in Statistics. Springer-Verlag, New York, 1993

  8. [8]

    Bacry, M

    E. Bacry, M. Bompaire, P. Deegan, S. Ga \" ffas, and S. V. Poulsen. tick: a P ython library for statistical learning, with an emphasis on H awkes processes and time-dependent models. Journal of Machine Learning Research, 18 0 (214): 0 1--5, 2018

Show all 167 references
  1. [9]

    O utcome-adaptive lasso: Variable selection for causal inference

    I. Bald \'e , Y. A. Yang, and G. Lefebvre. Reader reaction to “ O utcome-adaptive lasso: Variable selection for causal inference” by S hortreed and E rtefaie (2017). Biometrics, 79 0 (1): 0 514--520, 2023

  2. [10]

    Bender, D

    A. Bender, D. R \"u gamer, F. Scheipl, and B. Bischl. A general machine learning framework for survival analysis. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 158--173. Springer, 2020

  3. [11]

    Bengs and H

    V. Bengs and H. Holzmann. Uniform approximation in classical weak convergence theory. arXiv preprint arXiv:1903.09864, 2019

  4. [12]

    Benkeser and M

    D. Benkeser and M. van der Laan . The highly adaptive lasso estimator. In 2016 IEEE international conference on data science and advanced analytics (DSAA), pages 689--696. IEEE, 2016

  5. [13]

    Benkeser, W

    D. Benkeser, W. Cai, and M. van der Laan . A nonparametric super-efficient estimator of the average treatment effect. Statistical Science, 35: 0 484--495, 08 2020

  6. [14]

    R. Berk, A. Buja, L. Brown, E. George, A. K. Kuchibhotla, W. Su, and L. Zhao. Assumption lean regression. The American Statistician, 2019

  7. [15]

    Bianchi, A

    A. Bianchi, A. M. Christgau, and J. S. Pedersen. On generators and relations of the rational cohomology of H ilbert schemes. Journal of Algebraic Combinatorics, 57 0 (3): 0 829--857, 2023

  8. [16]

    P. J. Bickel, C. A. Klaassen, Y. Ritov, and J. A. Wellner. Efficient and Adaptive Estimation for Semiparametric Models. Springer Book Archive. Springer-Verlag New York, 1 edition, 1998. ISBN 978-0-387-98473-5

  9. [17]

    Billingsley

    P. Billingsley. Convergence of probability measures. John Wiley & Sons, 2013

  10. [18]

    Billingsley

    P. Billingsley. Probability and measure. John Wiley & Sons, 2017

  11. [19]

    S. M. Bischofberger, M. Hiabu, E. Mammen, and J. P. Nielsen. Smooth backfitting for additive hazard rates. arXiv preprint arXiv:2302.09510, 2023

  12. [20]

    C. S. Bojer and J. P. Meldgaard. Kaggle forecasting competitions: An overlooked learning opportunity. International Journal of Forecasting, 37 0 (2): 0 587--603, 2021

  13. [21]

    Boucheron, G

    S. Boucheron, G. Lugosi, and P. Massart. Concentration inequalities: A nonasymptotic theory of independence. Oxford University Press, 2013

  14. [22]

    Bousquet and A

    O. Bousquet and A. Elisseeff. Stability and generalization. Journal of Machine Learning Research, 2: 0 499--526, 2002

  15. [23]

    L. Breiman. Statistical modeling: The two cultures (with comments and a rejoinder by the author). Statistical science, 16 0 (3): 0 199--231, 2001

  16. [24]

    Br \'e maud

    P. Br \'e maud. Point processes and queues. Springer-Verlag, New York, 1981

  17. [25]

    N. E. Breslow and N. E. Day. Statistical methods in cancer research. V olume II -- T he design and analysis of cohort studies. IARC Scientific Publications, 82: 0 1--406, 1987

  18. [26]

    R. Cai, S. Wu, J. Qiao, Z. Hao, K. Zhang, and X. Zhang. THP s: Topological H awkes processes for learning causal structure on event sequences. IEEE Transactions on Neural Networks and Learning Systems, pages 1--15, 2022

  19. [27]

    T. T. Cai and M. Yuan. Minimax and adaptive prediction for functional linear regression. Journal of the American Statistical Association, 107 0 (499): 0 1201--1216, 2012

  20. [28]

    Q. Chen, V. Syrgkanis, and M. Austern. Debiased machine learning without sample-splitting for stable estimators. Advances in Neural Information Processing Systems, 35: 0 3096--3109, 2022

  21. [29]

    Chernozhukov, D

    V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins. Double/debiased machine learning for treatment and structural parameters . The Econometrics Journal, 21 0 (1): 0 C1--C68, 01 2018. ISSN 1368-4221

  22. [30]

    Chickering, D

    M. Chickering, D. Heckerman, and C. Meek. Large-sample learning of bayesian networks is NP -hard. Journal of Machine Learning Research, 5: 0 1287--1330, 2004

  23. [31]

    A. M. Christgau and N. R. Hansen. Assumption-lean A alen regression. In preparation, 2024+

  24. [32]

    A. M. Christgau and N. R. Hansen. Efficient adjustment for complex covariates: G aining efficiency with DOPE . arXiv preprint arXiv:2402.12980, 2024

  25. [33]

    A. M. Christgau, A. Arnaudon, and S. Sommer. Moment evolution equations and moment matching for stochastic image EPD iff. Journal of Mathematical Imaging and Vision, 65 0 (4): 0 563--576, 2023 a

  26. [34]

    A. M. Christgau, L. Petersen, and N. R. Hansen. Nonparametric conditional local independence testing. Annals of Statistics, 51 0 (5): 0 2116--2144, 2023 b

  27. [35]

    A. M. Christgau, L. Petersen, and N. R. Hansen. Supplement to: Nonparametric conditional local independence testing . Annals of Statistics, 51 0 (5), 2023 c . doi:10.1214/23-AOS2323SUPPA

  28. [36]

    S. N. Cohen and R. J. Elliott. Stochastic calculus and applications, volume 2. Springer, 2015

  29. [37]

    Commenges and A

    D. Commenges and A. G \'e gout-Petit. A general dynamical statistical model with causal interpretation. Journal of the Royal Statistical Society. Series B (Statistical Methodology), 71 0 (3): 0 719--736, 2009

  30. [38]

    C. S. Cox, J. J. Feldman, C. D. Golden, M. A. Lane, J. H. Madans, M. E. Mussolino, and S. T. Rothwell. Plan and operation of the NHANES I epidemiologic follow-up study, 1992. Vital and health statistics., 12 1997

  31. [39]

    D. R. Cox. Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological), 34 0 (2): 0 187--202, 1972

  32. [40]

    Daniel, J

    R. Daniel, J. Zhang, and D. Farewell. Making apples from oranges: comparing noncollapsible effect estimators and their standard errors after adjustment for different covariate sets. Biometrical Journal, 63 0 (3): 0 528--557, 2021

  33. [41]

    Davidson-Pilon

    C. Davidson-Pilon. lifelines: survival analysis in P ython. Journal of Open Source Software, 4 0 (40): 0 1317, 2019. doi:10.21105/joss.01317

  34. [42]

    Delecroix, W

    M. Delecroix, W. H \"a rdle, and M. Hristache. Efficient estimation in conditional single-index regression. Journal of Multivariate Analysis, 86 0 (2): 0 213--226, 2003

  35. [43]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805, 2018

  36. [44]

    V. Didelez. Graphical models for composable finite M arkov processes. Scandinavian Journal of Statistics, 34 0 (1): 0 169--185, 2006

  37. [45]

    V. Didelez. Graphical models for marked point processes based on local independence. Journal of the Royal Statistical Society. Series B (Statistical Methodology), 70 0 (1): 0 245--264, 2008

  38. [46]

    V. Didelez. Causal reasoning for events in continuous time: A decision-theoretic approach. In Proceedings of the UAI 2015 Workshop on Advances in Causal Inference, 2015

  39. [47]

    Didelez and M

    V. Didelez and M. J. Stensrud. On the logic of collapsibility for causal effect measures. Biometrical Journal, 64 0 (2): 0 235--242, 2022

  40. [48]

    Dukes, T

    O. Dukes, T. Martinussen, E. J. Tchetgen Tchetgen, and S. Vansteelandt. On doubly robust estimation of the hazard difference. Biometrics, 75 0 (1): 0 100--109, 2019

  41. [49]

    \'E mery and W

    M. \'E mery and W. Schachermayer. On V ershik’s standardness criterion and T sirelson’s notion of cosiness. In S \'e minaire de Probabilit \'e s XXXV , pages 265--305. Springer, 2001

  42. [50]

    E. X. Fang, Y. Ning, and H. Liu. Testing and confidence intervals for high dimensional proportional hazards models. Journal of the Royal Statistical Society. Series B (Statistical Methodology), pages 1415--1437, 2017

  43. [51]

    T. R. Fleming and D. P. Harrington. Counting processes and survival analysis, volume 169. John Wiley & Sons, 2011

  44. [52]

    Forr \'e and J

    P. Forr \'e and J. M. Mooij. A mathematical introduction to causality. Lecture notes, 2023

  45. [53]

    S. S. Franklin, S. A. Khan, N. D. Wong, M. G. Larson, and D. Levy. Is pulse pressure useful in predicting risk for coronary heart disease? T he F ramingham heart study. Circulation, 100 0 (4): 0 354--360, 1999

  46. [54]

    J. H. Friedman. Greedy function approximation: a gradient boosting machine. Annals of Statistics, pages 1189--1232, 2001

  47. [55]

    Gnecco, J

    N. Gnecco, J. Peters, S. Engelke, and N. Pfister. Boosted control functions. arXiv:2310.05805, 2023

  48. [56]

    C. W. J. Granger. Investigating causal relations by econometric models and cross-spectral methods. Econometrica, 37 0 (3): 0 424--438, 1969

  49. [57]

    Greenewald, K

    K. Greenewald, K. Shanmugam, and D. Katz. High-dimensional feature selection for sample efficient treatment effect estimation. In International Conference on Artificial Intelligence and Statistics, pages 2224--2232. PMLR, 2021

  50. [58]

    F. R. Guo, A. R. Lundborg, and Q. Zhao. Confounder selection: Objectives and approaches. arXiv:2208.13871, 2022

  51. [59]

    Györfi, M

    L. Györfi, M. Kohler, A. Krzyżak, and H. Walk. A Distribution-Free Theory of Nonparametric Regression. Springer Series in Statistics. Springer New York, NY, 1 edition, 2002. ISBN 978-0-387-22442-8. doi:10.1007/b97848

  52. [60]

    J. Hahn. On the role of the propensity score in efficient semiparametric estimation of average treatment effects. Econometrica, pages 315--331, 1998

  53. [61]

    B. B. Hansen. The prognostic analogue of the propensity score. Biometrika, 95 0 (2): 0 481--488, 2008

  54. [62]

    Hardt, B

    M. Hardt, B. Recht, and Y. Singer. Train faster, generalize better: Stability of stochastic gradient descent. In International conference on Machine Learning, pages 1225--1234. PMLR, 2016

  55. [63]

    Harezlak, B

    J. Harezlak, B. A. Coull, N. M. Laird, S. R. Magari, and D. C. Christiani. Penalized solutions to functional regression problems. Computational statistics & data analysis, 51 0 (10): 0 4911--4925, 2007

  56. [64]

    Hastie, R

    T. Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman. The elements of statistical learning: data mining, inference, and prediction, volume 2. Springer, 2009

  57. [65]

    Henckel, E

    L. Henckel, E. Perkovi \'c , and M. H. Maathuis. Graphical criteria for efficient total effect estimation via adjustment in causal linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 84 0 (2): 0 579--599, 2022

  58. [66]

    M. A. Hern \'a n. The hazards of hazard ratios. Epidemiology, 21 0 (1): 0 13--15, 2010

  59. [67]

    Hiabu, E

    M. Hiabu, E. Mammen, M. D. Mart \' nez-Miranda, and J. P. Nielsen. Smooth backfitting of proportional hazards with multiplicative components. Journal of the American Statistical Association, 116 0 (536): 0 1983--1993, 2021

  60. [68]

    Hines, O

    O. Hines, O. Dukes, K. Diaz-Ordaz, and S. Vansteelandt. Demystifying statistical learning based on efficient influence functions. The American Statistician, 76 0 (3): 0 292--304, 2022

  61. [69]

    Hines, K

    O. Hines, K. Diaz-Ordaz, and S. Vansteelandt. Optimally weighted average derivative effects. arXiv preprint arXiv:2308.05456, 2023

  62. [70]

    Homan, S

    T. Homan, S. Bordes, and E. Cichowski. Physiology, pulse pressure. In: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing, 2024. Updated 2023 Jul 10. Retrived from: https://www.ncbi.nlm.nih.gov/books/NBK482408/

  63. [71]

    J. Hou, J. Bradic, and R. Xu. Treatment effect estimation under additive hazards models with high-dimensional confounding. Journal of the American Statistical Association, 118 0 (541): 0 327--342, 2023

  64. [72]

    J. Huang. Efficient estimation of the partly linear additive C ox model. Annals of Statistics, 27 0 (5): 0 1536--1563, 1999

  65. [73]

    Ichimura

    H. Ichimura. Semiparametric least squares ( SLS ) and weighted SLS estimation of single-index models. Journal of Econometrics, 58 0 (1-2): 0 71--120, 1993

  66. [74]

    Jiang and J.-L

    C.-R. Jiang and J.-L. Wang. Functional single index models for longitudinal data. Annals of Statistics, 39 0 (1): 0 362--388, 2011

  67. [75]

    C. Ju, S. Gruber, S. D. Lendle, A. Chambaz, J. M. Franklin, R. Wyss, S. Schneeweiss, and M. J. van der Laan . Scalable collaborative targeted learning for high-dimensional data. Statistical methods in medical research, 28 0 (2): 0 532--554, 2019

  68. [76]

    C. Ju, D. Benkeser, and M. J. van der Laan . Robust inference on the average treatment effect using the outcome highly adaptive lasso. Biometrics, 76 0 (1): 0 109--118, 2020

  69. [77]

    Kallenberg

    O. Kallenberg. Foundations of modern probability, volume 3. Springer, 2021

  70. [78]

    M. Kasy. Uniformity and the delta method. Journal of Econometric Methods, 8 0 (1), 2019

  71. [79]

    E. H. Kennedy. Semiparametric doubly robust targeted double machine learning: a review. arXiv preprint arXiv:2203.06469, 2022

  72. [80]

    Klyne and R

    H. Klyne and R. D. Shah. Average partial effect estimation using double machine learning. arXiv preprint arXiv:2308.09207, 2023

  73. [81]

    Kolmogoroff

    A. Kolmogoroff. Grundbegriffe der Wahrscheinlichkeitsrechnung. Ergebnisse der Mathematik und Ihrer Grenzgebiete. 1. Folge. Springer Berlin, Heidelberg, 1 edition, 1933. ISBN 978-3-642-49596-0

  74. [82]

    The attractiveness of an additive hazard model: An example from medical demography

    Kravdal. The attractiveness of an additive hazard model: An example from medical demography. European Journal of Population / Revue européenne de Démographie, 13 0 (1): 0 33--47, 1997. ISSN 1572-9885. doi:10.1023/A:1005768425808

  75. [83]

    Ledoux and M

    M. Ledoux and M. Talagrand. Probability in Banach Spaces: Isoperimetry and Processes, volume 23. Springer Science & Business Media, 1991

  76. [84]

    D. K. Lee, N. Chen, and H. Ishwaran. Boosted nonparametric hazards with time-dependent covariates. Annals of Statistics, 49 0 (4): 0 2101, 2021

  77. [85]

    M. Y. Lee S. McDaniel and R. Chappell. Analysis and design of clinical trials using additive hazards survival endpoints. Statistics in Biopharmaceutical Research, 11 0 (3): 0 274--282, 2019. doi:10.1080/19466315.2019.1575278

  78. [86]

    D. Y. Lin and Z. Ying. Semiparametric analysis of the additive risk model . Biometrika, 81 0 (1): 0 61--71, 03 1994. ISSN 0006-3444. doi:10.1093/biomet/81.1.61

  79. [87]

    J. J. Lok. Statistical modeling of causal effects in continuous time. Annals of Statistics, 36 0 (3): 0 1464--1507, 06 2008

  80. [88]

    J. J. Lok. Mimicking counterfactual outcomes to estimate causal effects. Annals of Statistics, 45 0 (2): 0 461, 2017

  81. [89]

    S. M. Lundberg, G. Erion, H. Chen, A. DeGrave, J. M. Prutkin, B. Nair, R. Katz, J. Himmelfarb, N. Bansal, and S.-I. Lee. From local explanations to global understanding with explainable AI for trees. Nature machine intelligence, 2 0 (1): 0 56--67, 2020

  82. [90]

    A. R. Lundborg. Modern methods for variable significance testing. Apollo - University of Cambridge Repository, 2023. doi:10.17863/CAM.93556

  83. [91]

    A. R. Lundborg and N. Pfister. Perturbation-based analysis of compositional data. arXiv:2311.18501, 2023

  84. [92]

    A. R. Lundborg, I. Kim, R. D. Shah, and R. J. Samworth. The projected covariance measure for assumption-lean variable significance testing. arXiv preprint arXiv:2211.02039, 2022 a

  85. [93]

    A. R. Lundborg, R. D. Shah, and J. Peters. Conditional independence testing in H ilbert spaces with applications to functional data analysis. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 84 0 (5): 0 1821--1850, 2022 b . doi:10.1111/rssb.12544

  86. [94]

    A. Maity. Nonparametric functional concurrent regression models. Wiley Interdisciplinary Reviews: Computational Statistics, 9 0 (2): 0 e1394, 2017

  87. [95]

    Malfait and J

    N. Malfait and J. O. Ramsay. The historical functional linear model. The Canadian Journal of Statistics, 31 0 (2): 0 115--128, 2003

  88. [96]

    Manrique

    T. Manrique. Functional linear regression models: application to high-throughput plant phenotyping functional data. PhD thesis, Université de Montpellier, 2016

  89. [97]

    Manrique, C

    T. Manrique, C. Crambes, and N. Hilgert. Ridge regression for the functional concurrent model. Electronic Journal of Statistics, 12 0 (1): 0 985--1018, 2018

  90. [98]

    Martinussen and T

    T. Martinussen and T. H. Scheike. Dynamic regression models for survival data, volume 1. Springer, 2006

  91. [99]

    Martinussen and S

    T. Martinussen and S. Vansteelandt. On collapsibility and confounding bias in C ox and A alen regression models. Lifetime Data Analysis, 19 0 (3): 0 279--296, 2013. ISSN 1380-7870. doi:10.1007/s10985-013-9242-z

  92. [100]

    Martinussen, S

    T. Martinussen, S. Vansteelandt, and P. K. Andersen. Subtleties in the interpretation of hazard contrasts. Lifetime Data Analysis, 26: 0 833--855, 2020

  93. [101]

    I. W. McKeague and P. D. Sasieni. A partly parametric additive risk model. Biometrika, 81 0 (3): 0 501--514, 1994

  94. [102]

    S. W. Mogensen and N. R. Hansen. Markov equivalence of marginalized local independence graphs. Annals of Statistics, 48 0 (1): 0 539--559, 2020

  95. [103]

    S. W. Mogensen and N. R. Hansen. Graphical modeling of stochastic processes driven by correlated noise. Bernoulli, 28 0 (4): 0 3023--3050, 2022

  96. [104]

    S. W. Mogensen, D. Malinsky, and N. R. Hansen. Causal learning for partially observed stochastic dynamical systems. In Proceedings of the 34th conference on Uncertainty in Artificial Intelligence, pages 350--360, 2018

  97. [105]

    W. K. Newey. Uniform convergence in probability and stochastic equicontinuity. Econometrica, 59 0 (4): 0 1161--1167, 1991

  98. [106]

    W. K. Newey. The asymptotic variance of semiparametric estimators. Econometrica: Journal of the Econometric Society, pages 1349--1382, 1994

  99. [107]

    Neykov, S

    M. Neykov, S. Balakrishnan, and L. Wasserman. Minimax optimal conditional independence testing. Annals of Statistics, 49 0 (4): 0 2151--2177, 2021

  100. [108]

    Z. Niu, A. Chakraborty, O. Dukes, and E. Katsevich. Reconciling model-x and doubly robust approaches to conditional independence testing. arXiv preprint arXiv:2211.14698, 2022

  101. [109]

    Pakbin, X

    A. Pakbin, X. Wang, B. J. Mortazavi, and D. K. Lee. Bo XHED 2.0: Scalable boosting of dynamic survival analysis. arXiv preprint arXiv:2103.12591, 2021

  102. [110]

    Parkinson, G

    S. Parkinson, G. Ongie, and R. Willett. Linear neural network layers promote learning single-and multiple-index models. arXiv:2305.15598, 2023

  103. [111]

    Paszke, S

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. Pytorch: An imperative style, high-per...

  104. [112]

    J. Pearl. Causality. Cambridge University Press, 2009

  105. [113]

    Pedregosa, G

    F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in P ython. Journal of Machine Learnin...

  106. [114]

    Perkovi\'c, J

    E. Perkovi\'c, J. Textor, M. Kalisch, and M. H. Maathuis. Complete graphical characterization and construction of adjustment sets in markov equivalence classes of ancestral graphs. Journal of Machine Learning Research, 18 0 (220): 0 1--62, 2018

  107. [115]

    Peters, D

    J. Peters, D. Janzing, and B. Sch \"o lkopf. Elements of causal inference: foundations and learning algorithms. The MIT Press, 2017

  108. [116]

    Petersen and N

    L. Petersen and N. R. Hansen. Testing conditional independence via quantile regression based partial copulas. Journal of Machine Learning Research, 22 0 (70): 0 1--47, 2021

  109. [117]

    Pfanzagl and W

    J. Pfanzagl and W. Wefelmeyer. Contributions to a general asymptotic statistical theory. Statistics & Risk Modeling, 3 0 (3-4): 0 379--388, 1985

  110. [118]

    D. Pollard. Convergence of stochastic processes. Springer Series in Statistics. Springer-Verlag, New York, 1984

  111. [119]

    P \"o lsterl

    S. P \"o lsterl. scikit-survival: A library for time-to-event analysis built on top of scikit-learn. Journal of Machine Learning Research, 21 0 (212): 0 1--6, 2020

  112. [120]

    Popoviciu

    T. Popoviciu. Sur les \'e quations alg \'e briques ayant toutes leurs racines r \'e elles. Mathematica, 9 0 (129-145): 0 20, 1935

  113. [121]

    J. L. Powell, J. H. Stock, and T. M. Stoker. Semiparametric estimation of index coefficients. Econometrica: Journal of the Econometric Society, pages 1403--1430, 1989

  114. [122]

    Rebolledo

    R. Rebolledo. Central limit theorems for local martingales. Zeitschrift f \"u r Wahrscheinlichkeitstheorie und verwandte Gebiete , 51 0 (3): 0 269--286, 1980

  115. [123]

    Revuz and M

    D. Revuz and M. Yor. Continuous martingales and Brownian motion, volume 293. Springer Science & Business Media, 2013

  116. [124]

    J. M. Robins and A. Rotnitzky. Semiparametric efficiency in multivariate regression models with missing data. Journal of the American Statistical Association, 90 0 (429): 0 122--129, 1995

  117. [125]

    J. M. Robins, A. Rotnitzky, and L. P. Zhao. Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association, 89 0 (427): 0 846--866, 1994

  118. [126]

    P. M. Robinson. Root-n-consistent semiparametric regression. Econometrica: Journal of the Econometric Society, pages 931--954, 1988

  119. [127]

    L. C. G. Rogers and D. Williams. Diffusions, M arkov processes, and martingales , volume 2. Cambridge University Press, Cambridge, 2000

  120. [128]

    P. R. Rosenbaum and D. B. Rubin. The central role of the propensity score in observational studies for causal effects. Biometrika, 70 0 (1): 0 41--55, 1983

  121. [129]

    Rotnitzky and E

    A. Rotnitzky and E. Smucler. Efficient adjustment sets for population average causal treatment effect estimation in graphical models. Journal of Machine Learning Research, 21 0 (188): 0 1--86, 2020

  122. [130]

    R ysland, P

    K. R ysland, P. Ryalen, M. Nyg rd, and V. Didelez. Graphical criteria for the identification of marginal causal effects in continuous-time survival and event-history analyses. arXiv preprint arXiv:2202.02311, 2022

  123. [131]

    H. C. Rytgaard, T. A. Gerds, and M. J. van der Laan. Continuous-time targeted minimum loss-based estimation of intervention-specific mean outcomes. Annals of Statistics, 50 0 (5): 0 2469--2491, 2022

  124. [132]

    H. C. W. Rytgaard, F. Eriksson, and M. van der Laan. Estimation of time-specific intervention effects on continuously distributed time-to-event outcomes by targeted maximum likelihood estimation. arXiv:2106.11009, 2021

  125. [133]

    P. Sasieni. Information bounds for the conditional hazard ratio in a nested family of regression models. Journal of the Royal Statistical Society: Series B (Methodological), 54 0 (2): 0 617--635, 1992

  126. [134]

    Calculating the expecation of the supremum of absolute value of a B rownian motion

    saz. Calculating the expecation of the supremum of absolute value of a B rownian motion. Mathematics Stack Exchange, 2019. Retrieved (2019-06-06) from: https://math.stackexchange.com/q/3252132

  127. [135]

    o rrmann, and P. B \

    C. Scheidegger, J. H \"o rrmann, and P. B \"u hlmann. The weighted generalised covariance measure. Journal of Machine Learning Research, 23 0 (273): 0 1--68, 2022

  128. [136]

    R. L. Schilling. Measures, integrals and martingales. Cambridge University Press, 2017

  129. [137]

    R. L. Schilling and L. Partzsch. Brownian Motion: An Introduction to Stochastic Processes. De Gruyter, 2012. ISBN 9783110278989. doi:doi:10.1515/9783110278989

  130. [138]

    Sch \"o lkopf, F

    B. Sch \"o lkopf, F. Locatello, S. Bauer, N. R. Ke, N. Kalchbrenner, A. Goyal, and Y. Bengio. Toward causal representation learning. Proceedings of the IEEE, 109 0 (5): 0 612--634, 2021

  131. [139]

    Schweder

    T. Schweder. Composable M arkov processes. Journal of Applied Probability, 7 0 (2): 0 400--410, 1970

  132. [140]

    u rk and H.-G. M \

    D. S ent \"u rk and H.-G. M \"u ller. Functional varying coefficient models for longitudinal data. Journal of the American Statistical Association, 105 0 (491): 0 1256--1264, 2010

  133. [141]

    R. D. Shah and J. Peters. The hardness of conditional independence testing and the generalised covariance measure. Annals of Statistics, 48 0 (3): 0 1514--1538, 2020

  134. [142]

    C. Shi, D. Blei, and V. Veitch. Adapting neural networks for the estimation of treatment effects. Advances in neural information processing systems, 32, 2019

  135. [143]

    S. M. Shortreed and A. Ertefaie. Outcome-adaptive lasso: variable selection for causal inference. Biometrics, 73 0 (4): 0 1111--1122, 2017

  136. [144]

    Smucler, A

    E. Smucler, A. Rotnitzky, and J. M. Robins. A unifying approach for doubly-robust l1-regularized estimation of causal contrasts. arXiv:1904.03737, 2019

  137. [145]

    J. A. Soloff, R. F. Barber, and R. Willet. Stability via resampling: statistical problems beyond the real line. arXiv preprint arXiv:2405.09511, 2024 a

  138. [146]

    J. A. Soloff, R. F. Barber, and R. Willett. Bagging provides assumption-free stability. Journal of Machine Learning Research, 25 0 (131): 0 1--35, 2024 b

  139. [147]

    M. J. Stensrud and M. A. Hern \'a n. Why test for proportional hazards? Jama, 323 0 (14): 0 1401--1402, 2020

  140. [148]

    A. B. Tsybakov. Introduction to Nonparametric Estimation. Springer Series in Statistics. Springer New York, NY, 1 edition, 2009. ISBN 978-0-387-79052-7. doi:10.1007/b13794

  141. [149]

    Uhler, G

    C. Uhler, G. Raskutti, P. B \"u hlmann, and B. Yu. Geometry of the faithfulness assumption in causal inference. Annals of Statistics, pages 436--463, 2013

  142. [150]

    M. J. van der Laan and S. Gruber. Collaborative double robust targeted maximum likelihood estimation. The international journal of biostatistics, 6 0 (1), 2010

  143. [151]

    M. J. van der Laan and S. Rose. Targeted learning: causal inference for observational and experimental data, volume 4. Springer, 2011

  144. [152]

    A. W. van der Vaart. Asymptotic statistics, volume 3. Cambridge university press, 2000

  145. [153]

    A. W. van der Vaart and J. A. Wellner. Weak convergence and empirical processes. Springer Series in Statistics. Springer-Verlag, New York, 1996

  146. [154]

    A. W. van der Vaart and J. A. Wellner. Weak Convergence and Empirical Processes: With Applications to Statistics. Springer Series in Statistics. Springer Cham, 2 edition, 2023. ISBN 978-3-031-29038-1

  147. [155]

    Van Rossum, F

    G. Van Rossum, F. L. Drake, et al. Python reference manual, volume 111. Centrum voor Wiskunde en Informatica Amsterdam, 1995

  148. [156]

    Vansteelandt and O

    S. Vansteelandt and O. Dukes. Assumption-lean inference for generalised linear model parameters. Journal of the Royal Statistical Society Series B: Statistical Methodology (with discussion), 84 0 (3): 0 657--685, 2022

  149. [157]

    Vansteelandt, O

    S. Vansteelandt, O. Dukes, K. Van Lancker, and T. Martinussen. Assumption-lean C ox regression. Journal of the American Statistical Association, pages 1--10, 2022

  150. [158]

    Veitch, D

    V. Veitch, D. Sridhar, and D. Blei. Adapting text embeddings for causal inference. In Conference on Uncertainty in Artificial Intelligence, pages 919--928. PMLR, 2020

  151. [159]

    X. Wang, A. Pakbin, B. Mortazavi, H. Zhao, and D. Lee. Bo XHED : Boosted e X act H azard E stimator with D ynamic covariates. In International Conference on Machine Learning, pages 9973--9982. PMLR, 2020

  152. [160]

    M. T. Wells. Nonparametric kernel estimation in counting processes with explanatory variables. Biometrika, 81 0 (4): 0 795--801, 1994

  153. [161]

    Xiao , J

    S. Xiao , J. Yan , M. Farajtabar , L. Song , X. Yang , and H. Zha . Learning time series associated event sequences with recurrent point process networks. IEEE Transactions on Neural Networks and Learning Systems, 30 0 (10): 0 3124--3136, 2019

  154. [162]

    H. Xu, M. Farajtabar, and H. Zha. Learning G ranger causality for H awkes processes. In Proceedings of The 33rd International Conference on Machine Learning, volume 48, pages 1717--1726, 2016

  155. [163]

    Yao, H.-G

    F. Yao, H.-G. M \"u ller, and J.-L. Wang. Functional linear regression analysis for longitudinal data . Annals of Statistics, 33 0 (6): 0 2873 -- 2903, 2005

  156. [164]

    Yuan and T

    M. Yuan and T. T. Cai. A reproducing kernel H ilbert space approach to functional linear regression. Annals of Statistics, 38 0 (6): 0 3412--3444, 2010

  157. [165]

    Zhong, J

    Q. Zhong, J. Mueller, and J.-L. Wang. Deep learning for the partially linear C ox model. Annals of Statistics, 50 0 (3): 0 1348--1375, 2022

  158. [166]

    K. Zhou, H. Zha, and L. Song. Learning social infectivity in sparse low-rank networks using multi-dimensional H awkes processes. In Proceedings of the 16th International Conference on Artificial Intelligence and Statistics, 2013

  159. [167]

    P. N. Zivich and A. Breskin. Machine learning for causal inference: on the use of cross-fit estimators. Epidemiology (Cambridge, Mass.), 32 0 (3): 0 393, 2021

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.