{"id":"373d2b08-81af-4432-90bf-930849c70fc0","arxiv_id":"2507.12424","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Bayesian hierarchical Hawkes model with edge correction estimates a population branching factor of 0.742 for aggressive behavior onsets in autistic inpatient youth, roughly 17% below the pooled model estimate.","lead":"The authors fit a hierarchical self-exciting model to the timing of aggressive behavior onsets in 70 autistic inpatient youths and estimate that each onset triggers about 0.74 later onsets on average, lower than the 0.90 obtained from a model that pools all patients together. The result matters because it changes estimates of how aggressive episodes cascade and how much clinical weight to give to internal versus external triggers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hierarchical branching-factor advantage is not validated: no simulation or benchmark shows partial pooling recovers a known branching factor, and the reported ELPD/GOF differences only weakly support the model.","rationale":"The reader's weakest assumption focuses on the model family (stationary exponential Hawkes with constant baseline), which is a legitimate source of misspecification. I agree that the kernel and constant-baseline assumptions are not validated on this population, and that the threefold cascade-size claim depends entirely on the parametric model. My more load-bearing concern is narrower and more specific: the paper claims the hierarchical model is 'less biased' and 'consistently favored' by GOF, but it provides no ground-truth benchmark for bias, and its own Table 4 contradicts the prose—the unpooled model has higher non-rejection rates at several significance levels, and the PSIS-LOO difference between unpooled and hierarchical is not statistically significant. The paper's reported MCMC diagnostics and power-scaling sensitivity analyses are credible evidence that the hierarchical posterior is well-behaved, but they do not substantiate the bias-reduction claim, which is the central advance over prior pooled work. A simulation study with known branching factors is the direct, feasible test that would settle whether partial pooling actually reduces bias in this sparse, heterogeneous regime; absent that, the verdict should remain CONDITIONAL rather than REJECT, because the core estimate is plausible and not internally contradicted, but the headline claim of reduced bias is unsupported as stated.","tokens_in":55475,"tokens_out":2780,"duration_ms":24410,"concrete_test":"Run a simulation study that generates synthetic inpatient-style datasets under the paper's exponential Hawkes model (Eq. 5) with known person-level branching factors drawn from a realistic distribution, including heavy-tailed event counts and sessions with zero events, then fit pooled, unpooled, and hierarchical models and report bias, RMSE, and coverage of the population branching factor. If the hierarchical estimator does not recover the true branching factor with lower bias than the pooled estimator across the simulated regime, the central bias-reduction claim fails. Separately, recompute Table 4's Lewis test rates and correct the prose to match the numbers, or supply the reported session-level proportions so the discrepancy can be resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the hierarchical model gives a less biased population branching factor (0.742 ± 0.026) than the pooled model (0.899 ± 0.015), with the difference framed as reducing bias from high-frequency individuals. But 'bias' is never defined against a ground truth. The paper has no simulation study, no synthetic-data recovery check, and no external benchmark; the only evidence is that the hierarchical estimate is lower and more precise than the pooled estimate and more stable than the unpooled estimate. A lower estimate is not automatically less biased: the pooled and hierarchical models are different estimators, and their difference could reflect prior shrinkage, kernel misspecification, or the nonstationary exogenous rate (e.g., staff changes, medication adjustments, session-specific context) rather than true bias correction. The sensitivity analysis (Tables 10-11) shows the hierarchical model is robust to prior/likelihood power-scaling, but power-scaling perturbs the assumed model; it does not test the model itself. In addition, Table 4 contradicts the prose: the paper states the hierarchical model has the highest proportion of non-rejections, but at the session level the unpooled model has higher non-rejection rates at 0.05 (0.88 vs 0.86) and 0.10 (0.83 vs 0.81), and at the person level the unpooled model is also higher at 0.05 (0.90 vs 0.86) and 0.10 (0.84 vs 0.83). Thus the claim that GOF 'consistently favored' the hierarchical model is not supported by the paper's own Table 4. The PSIS-LOO difference is also small (ELPD diff 4.81, se 17.22; p=0.88), so the hierarchical model is not statistically better than the unpooled model in predictive performance. These issues do not refute the central estimate, but they substantially weaken the claim that the hierarchical model is less biased and better fitting.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian hierarchical (partially pooled) Hawkes process with an exponential kernel and edge-effect correction to estimate the population branching factor for aggressive behavior onsets in 70 psychiatric inpatient youth with autism. It compares this model with pooled and unpooled alternatives, reporting a population branching factor of 0.742 ± 0.026 for the hierarchical model, versus 0.899 ± 0.015 for the pooled model and 0.717 ± 0.139 for the unpooled model. From these estimates it derives expected cascade sizes of about 2.92 versus 9.09 descendants per parent onset. The paper also reports MCMC diagnostics, power-scaling sensitivity analyses, PSIS-LOO cross-validation, Lewis tests with Durbin's modification, and RTCT residual analyses, and uses the inferred branching structure to illustrate potential links with physiological signals.","tokens_in":55799,"tokens_out":3344,"duration_ms":41611,"significance":"If the hierarchical estimate is valid, the result is clinically consequential: it would place the population dynamics in a more strongly subcritical regime than previously reported, changing estimated cascade sizes by roughly a factor of three. The paper has genuine strengths: the inference pipeline is careful, with convergence diagnostics, posterior predictive checks, power-scaling sensitivity analysis, and multiple goodness-of-fit criteria; the appendices provide detailed per-parameter diagnostics and individual-level data; and the software choices (NumPyro, ArviZ) support reproducibility. However, the central claim that the hierarchical model is 'less biased' is not validated against any ground truth, and some goodness-of-fit statements are not supported by the reported tables. The contribution is therefore promising but not yet established at the level claimed.","major_comments":[{"comment":"The central claim that partial pooling 'reduces bias from high-frequency individuals' is not supported by any ground-truth comparison. No simulation study, synthetic-data recovery check, or external benchmark is reported; the evidence for bias reduction is that the hierarchical estimate is lower and more precise than the pooled estimate and more stable than the unpooled estimate. A lower estimate is not automatically a less biased estimate: the difference between the pooled and hierarchical posteriors could reflect prior shrinkage under the Gamma(2.5, 0.4) hyperprior on mu_alpha, kernel misspecification, or nonstationarity of the exogenous rate, rather than correction of a known bias. The paper should add a simulation study in which data are generated from a known hierarchical Hawkes process with a known population branching factor and compare the recovery properties of the pooled, unpooled, and partially pooled estimators. This is load-bearing because the entire clinical interpretation rests on the claim that 0.742 is closer to the truth than 0.899.","section":"Section 3.1, Figure 2"},{"comment":"The prose states that 'the partially pooled model yields the highest proportion of sessions that do not reject the null hypothesis,' but Table 4 contradicts this. At the session level, the unpooled model has higher non-rejection rates at significance levels 0.05 (0.88 vs. 0.86) and 0.10 (0.83 vs. 0.81); at the person level, the unpooled model is also higher at 0.05 (0.90 vs. 0.86) and 0.10 (0.84 vs. 0.83). The partially pooled model is higher only at the 0.15 level. Similarly, Table 3 shows an ELPD difference of only 4.81 with a standard error of 17.22 between the partially pooled and unpooled models, which is not statistically significant and is much smaller than the reported uncertainty. The claim that goodness-of-fit measures 'consistently favored' the hierarchical model is therefore not supported by the reported evidence. The GOF section should be rewritten to state accurately which comparisons favor which model, and the strength of the evidence should be calibrated accordingly.","section":"Section 3.4, Table 4"},{"comment":"The model assumes a stationary exponential Hawkes process with a constant baseline intensity and a single exponential excitation kernel. This assumption is load-bearing because the branching factor is a fitted parameter of this assumed model. The paper justifies the exponential kernel by citing Filimonov and Sornette's financial-data robustness results, but no validation is provided for this population, where unobserved context such as staff changes, medication adjustments, or session-specific conditions could make the baseline rate nonstationary, or where memory of past events may not be exponential. If the exogenous rate varies with context, a fitted exponential kernel can absorb that variation into the self-excitation term and inflate alpha; the lower hierarchical estimate could then reflect shrinkage rather than a true property of the behavior. A concrete and feasible check would be a simulation study using a nonstationary baseline or a non-exponential kernel, comparing the estimated branching factor under the three models; alternatively, the authors could fit a model with time-varying or session-level covariates and report whether the branching factor estimate changes materially.","section":"Section 2.2.2, Eqs. (4)-(5)"},{"comment":"The 'threefold smaller cascade' result is a re-expression of the fitted branching factor through the formula mu_alpha/(1 - mu_alpha), not an independent empirical finding. The expected number of descendants inherits any bias or misspecification in the branching factor estimate, so the clinical contrast between 2.92 and 9.09 descendants cannot be presented as additional evidence for the hierarchical model. The paper should either describe this quantity explicitly as a derived summary of the posterior estimate, or provide an out-of-sample or simulation-based validation that the implied cascade sizes are accurate. As written, the cascade comparison is a transformation of the same parameters being compared, and the uncertainty reported for the ratio does not account for model misspecification.","section":"Section 3.1, Figure 3"}],"minor_comments":[{"comment":"The text contains a typo: 'Postetior Predictive Checks' should be 'Posterior Predictive Checks'.","section":"Section 2.5.1"},{"comment":"There is a misspelling: 'estaimates' should be 'estimates'.","section":"Section 3.1"},{"comment":"The caption lists panels (a) through (f) but the text refers to '(g) Session 41'; the panel labeling should be made consistent.","section":"Figure 5 caption"},{"comment":"The caption says that at the person level the table reports 'the average proportion where the null hypothesis was rejected,' but the values (0.86, 0.90, etc.) are consistent with non-rejection rates, not rejection rates; the caption should be corrected.","section":"Table 4 caption"},{"comment":"The text states an ESS target of at least 100, while the reported diagnostics show ESS values exceeding 3,000. The target should either be stated as a diagnostic threshold rather than a target for the actual run, or the text should explain why the run exceeded it by such a large margin.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the paper has one solid empirical result and one overclaimed interpretation. The solid result is that a hierarchical Hawkes model with edge-effect correction gives a lower and much tighter branching factor estimate (0.742 ± 0.026 vs 0.899 ± 0.015) on 70 inpatient youth. The overclaim is that this is 'less biased.' There is no simulation, no synthetic-data recovery, and no external benchmark. Lower is not automatically closer to the truth.\n\nWhat is genuinely new is the package: partial pooling plus edge-effect correction plus power-scaling sensitivity on a small, messy clinical dataset. The MCMC diagnostics are carefully reported and the sensitivity analysis showing unpooled instability for sparse individuals is useful. That is real applied work and it earns credit.\n\nThe soft spots are in the interpretation and the GOF reporting. The bias claim is load-bearing and unsupported. The pooled and hierarchical models are different estimators with different priors; their difference could reflect shrinkage, kernel misspecification, or a nonstationary exogenous rate (staff changes, meds) rather than bias correction. Power-scaling perturbs the assumed model, it does not validate the model. The exponential kernel is justified by a finance result, not by checking this population. Second, Table 4 contradicts the prose. The text says the hierarchical model is 'consistently favored' by the Lewis test, but at 0.05 and 0.10 the unpooled model has higher non-rejection rates at both session and person level. The PSIS-LOO difference between hierarchical and unpooled is also tiny (ELPD diff 4.81, se 17.22, p=0.88). So 'superior GOF' is not supported by their own tables. Third, the threefold cascade claim is a re-expression of the fitted branching factor via mu/(1-mu); it is not an independent validation.\n\nNone of this refutes the central estimate. The hierarchical estimate may well be closer to the truth; the paper just has not shown it. What it needs is a simulation study where a known branching factor is recovered under their kernel, a reconciled GOF table, and toned-down wording.\n\nWho this is for: applied statisticians and clinicians working on behavioral event time series, and Hawkes modelers who want a case study in partial pooling. It deserves a serious referee; the direction is credible and the execution is competent. I would send it out, but I would tell the authors the bias claim and the GOF summaries need to be fixed before acceptance.","headline":"Solid empirical comparison, but 'less biased' is asserted, not demonstrated; the GOF claim is contradicted by their own Table 4.","tokens_in":56441,"tokens_out":2594,"would_cite":false,"duration_ms":28767,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Aggression in autistic inpatient youth is less contagious than pooled models claim","keywords":["Hawkes process","branching factor","hierarchical model","aggressive behavior","autism inpatient","temporal point process","partial pooling","Bayesian inference"],"falsifier":"Fit the same hierarchical model with a power-law or nonparametric self-excitation kernel and compare the posterior of the population branching factor; if the subcritical estimate near 0.742 moves substantially or crosses into the supercritical regime, the exponential-kernel assumption is the deciding factor and the branching factor estimate is not robust.","tokens_in":55234,"feed_emoji":"📉","tokens_out":3321,"duration_ms":38777,"temperature":0.7,"pith_summary":"This paper argues that previous estimates of how strongly one aggressive behavior onset triggers the next were inflated, because they pooled all patients into one homogeneous model. The authors fit a hierarchical Hawkes process with an exponential kernel and an edge-effect correction, allowing each patient's own triggering rate to borrow strength from the group. They report a sample-population branching factor of 0.742 ± 0.026, statistically significantly lower than the pooled model's 0.899 ± 0.015, and far more precise than the unpooled model's 0.717 ± 0.139. If correct, this changes the predicted cascade size from about nine descendant onsets per parent onset to about three, shifting clinical emphasis from internal escalation toward external triggers and individualized monitoring.","feed_headline":"Autism aggression cascades are threefold smaller than pooled models suggest","feed_subtitle":"A hierarchical Hawkes model lowers the branching factor to 0.742, changing predicted cascade size from 9 to about 3 descendants.","key_machinery":"The central object is the branching factor α in a Hawkes process with exponential kernel φ(t) = αβ exp(−βt), where α is the expected number of direct offspring per event and the process is subcritical when α < 1. The model adds an edge-effect initial intensity term (μ0 − μ)β exp(−βt) to account for unobserved history before each observation session. The key mechanism is partial pooling: patient-level parameters μn, αn, βn are drawn from LogNormal distributions whose population-level means and scales are learned from the data, so that sparse patients borrow strength from the group while high-frequency patients do not dominate.","core_discovery":"The central claim is that partially pooling patient-specific Hawkes process parameters corrects a systematic upward bias in the sample-population branching factor for aggressive behavior onset. Using a conditional intensity with baseline plus exponential self-excitation, the hierarchical model yields a mean branching factor of 0.742 with a standard deviation of 0.026, versus 0.899 ± 0.015 for the pooled model and 0.717 ± 0.139 for the unpooled model. The paper further claims this reduces the expected total number of descendants per parent onset from 9.09 ± 1.58 (pooled) to 2.92 ± 0.40 (hierarchical), and that the hierarchical model is robust to prior and likelihood power-scaling perturbations while the unpooled model is not, especially for individuals with sparse data.","pith_inferences":["The threefold cascade reduction is a direct mathematical consequence of the formula α/(1−α): moving α from about 0.90 to about 0.74 changes expected descendants from 9 to under 3, so the clinical difference is driven entirely by the point estimate of α.","A natural extension would be to test the model on data that include known external triggers such as staff changes or medication timing, to validate whether events classified as exogenous actually align with documented environmental events.","The choice of exponential kernel was justified by robustness results from financial data, not validated on this population; a power-law or nonparametric kernel comparison would test whether the subcritical estimate is kernel-independent.","Individual branching factors from the hierarchical model could be used directly as patient-level risk scores, stratifying youth for different intensities of monitoring or intervention."],"forward_implications":["The sample-population branching factor for aggressive behavior onset in this cohort is below the critical value of 1, meaning cascades are not self-sustaining.","Each aggressive onset is expected to produce about 2.9 subsequent onsets, roughly three times fewer than the pooled model's 9.1, directly affecting predicted escalation risk.","The hierarchical model provides more precise population-level estimates and more stable individual-level estimates than the unpooled model for patients with sparse data.","Separating exogenous from endogenous onsets can support linkage of aggressive behavior to physiological signals and individualized early-warning systems.","Pooled Hawkes models applied to heterogeneous clinical populations are likely to overstate self-excitation and misclassify exogenous events as endogenous."],"supporting_citations":[{"why":"Prior pooled-model study on the same data that this work extends and corrects.","marker":"[16]"},{"why":"Filimonov and Sornette supplies the justification for the exponential kernel's robustness to outlier inter-arrival times.","marker":"[21]"},{"why":"Laub, Lee, and Taimre supplies the edge-effect correction formulation and the expected-descendants formula α/(1−α).","marker":"[22]"},{"why":"Gelman 2006 grounds the argument that pooled models underfit and unpooled models overfit, motivating partial pooling.","marker":"[24]"},{"why":"Gelman's prior recommendations for hierarchical variance parameters support the heavy-tailed hyperprior choices.","marker":"[26]"},{"why":"Kallioinen et al. supplies the power-scaling sensitivity analysis methodology used to assess prior and likelihood robustness.","marker":"[36]"},{"why":"Vehtari et al. supplies the PSIS-LOO criterion used for comparing predictive performance across models.","marker":"[32]"}],"fun_headline_variants":["Hierarchical Hawkes model cuts autism aggression cascades to 3 events","Autism aggression branching factor drops to 0.742 with hierarchical model","New model: autism aggression cascades 3-fold smaller than pooled estimate","Hierarchical Hawkes correction shrinks autism aggression cascade estimate","Partially pooled Hawkes model lowers autism aggression branching factor"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that aggressive behavior onsets are generated by a stationary exponential Hawkes process with constant baseline intensity and a single exponential self-excitation kernel; if the true exogenous rate varies with unobserved context or if memory of past events is not exponential, the estimated branching factor and downstream cascade sizes are model artifacts rather than properties of the behavior.","fun_headline_variants_meta":{"raw":{"variants":["Hierarchical Hawkes model cuts autism aggression cascades to 3 events","Autism aggression branching factor drops to 0.742 with hierarchical model","New model: autism aggression cascades 3-fold smaller than pooled estimate","Hierarchical Hawkes correction shrinks autism aggression cascade estimate","Partially pooled Hawkes model lowers autism aggression branching factor"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000666,"raw_usage":{"total_tokens":3085,"prompt_tokens":1038,"completion_tokens":2047,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":654,"completion_tokens_details":{"reasoning_tokens":1957}},"tokens_in":654,"tokens_out":2047,"duration_ms":13723,"temperature":1.0,"reasoning_tokens":1957,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:46:20.846968+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the same hierarchical model with a power-law or nonparametric self-excitation kernel and compare the posterior of the population branching factor; if the subcritical estimate near 0.742 moves substantially or crosses into the supercritical regime, the exponential-kernel assumption is the deciding factor and the branching factor estimate is not robust.","supporting_citations":[{"cited_title":"Temporal Point Process Modeling of Aggressive Behavior Onset in Psychiatric Inpatient Youths with Autism","cited_arxiv_id":null,"evidence_quote":"Prior pooled-model study on the same data that this work extends and corrects."},{"cited_title":"Apparent criticality and calibration issues in the Hawkes self-excited point process model: application to high-frequency financial data","cited_arxiv_id":null,"evidence_quote":"Filimonov and Sornette supplies the justification for the exponential kernel's robustness to outlier inter-arrival times."},{"cited_title":"The elements of Hawkes processes","cited_arxiv_id":null,"evidence_quote":"Laub, Lee, and Taimre supplies the edge-effect correction formulation and the expected-descendants formula α/(1−α)."},{"cited_title":"Multilevel (hierarchical) modeling: what it can and cannot do","cited_arxiv_id":null,"evidence_quote":"Gelman 2006 grounds the argument that pooled models underfit and unpooled models overfit, motivating partial pooling."},{"cited_title":"Prior distributions for variance parameters in hierarchical models (comment on article by Browne and Draper)","cited_arxiv_id":null,"evidence_quote":"Gelman's prior recommendations for hierarchical variance parameters support the heavy-tailed hyperprior choices."},{"cited_title":"Detecting and diagnosing prior and likelihood sensitivity with power-scaling","cited_arxiv_id":null,"evidence_quote":"Kallioinen et al. supplies the power-scaling sensitivity analysis methodology used to assess prior and likelihood robustness."},{"cited_title":"Practical Bayesian model evaluation using leave-one-out cross-validation and W AIC","cited_arxiv_id":null,"evidence_quote":"Vehtari et al. supplies the PSIS-LOO criterion used for comparing predictive performance across models."}],"review_version":1}