{"id":"ea15383b-ddf7-4a26-8368-daaf54ea1560","arxiv_id":"2507.23034","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A temporal-network community test averages e-values obtained from per-snapshot p-values, retaining valid type I error control under arbitrary dependence.","lead":"This paper proposes a statistical test for community structure in networks that change over time, combining per-snapshot p-values into an averaged e-value. It offers one of the first formal ways to test temporal networks, though the method detects average structure rather than evolving communities.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The invalid 'max' calibrator is used as a primary comparison in simulations; since it is not a calibrator, the paper's experimental support for the e-value test is partly based on a procedure that can inflate type I error.","rationale":"The reader's verdict is CONDITIONAL and their rationale already mentions the invalid max calibrator, but their weakest_assumption field is about the meaningfulness of averaging per-snapshot evidence. I focus on the max calibrator because it is the most concrete, correctable, load-bearing issue: unlike the 'average community structure' limitation, which the authors acknowledge and position as a scope condition, the max calibrator is presented throughout the simulation study as one of the main options, so the paper's experimental evidence for type I error and power is compromised for that option. The theorem is mathematically sound for valid calibrators, so the issue is not fatal; it requires a revision of the simulations and recommendations. I therefore do not think the verdict should change from CONDITIONAL, hence UNCHANGED.","tokens_in":17219,"tokens_out":11423,"duration_ms":136082,"concrete_test":"Re-run the Section 4.2 correlated-SBM null setting (T=10, n=1000, rho=0.25, delta=0) with 10,000 Monte Carlo replicates using the released code, and compute the rejection rate of the max-calibrated average e-value at threshold 20. If the rejection rate exceeds 0.05 by more than sampling error (e.g., >0.06), the max results in Figures 1-3 and the Supplement should be removed or explicitly labeled as invalid; if it is approximately 0.05, the concern is mitigated for this regime and the max calibrator's inclusion is less harmful, though still unsupported in general.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The formal claim (Theorem 1) is correct for any valid p-to-e calibrator: with a proper calibrator, each E_t is an e-value and the average is an e-value. The load-bearing problem is that Section 3.1 explicitly states that g_max(p) = -exp(-1)/(p log p) for p <= e^{-1} and 1 otherwise is not a calibrator, yet Section 4.1 lists 'max' among the five compared calibrators and Figures 1-3 and the Supplement present it as the best performer. Because the integral of g_max over [0,1] is infinite, E[g_max(P)] is not uniformly bounded by 1 under the null even when P is exactly uniform; the large e-values reported for 'max' can be an artifact of invalid type I error inflation, not power. This matters for the central claim because the paper's practical message is that e-values retain validity under arbitrary temporal dependence; a user who follows the simulation comparison and selects 'max' loses the advertised type I error guarantee. The issue is acknowledged in passing ('max is not a proper calibrator'), but the authors continue to use it as a primary recommendation, so the experimental evidence for the method's performance is partly based on a non-valid procedure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hypothesis test for community structure in temporal networks. The test computes a static p-value for each snapshot, converts each p-value to an e-value using a p-to-e calibrator, and averages the e-values across snapshots. The central theoretical claim (Theorem 1) is that this average is a valid e-value under arbitrary temporal dependence. The method is illustrated on three simulated network processes (correlated SBM, dynamic SBM, dynamic DCBM) and five real-world networks, using the Bickel-Sarkar and Yanchenko-Sengupta static tests. The authors also candidly discuss limitations, including the implicit 'average community strength' interpretation, the sensitivity to zero p-values, and the difficulty of defining temporal community structure.","tokens_in":17465,"tokens_out":6099,"duration_ms":69189,"significance":"The elementary averaging argument is correct and gives a genuinely simple way to combine dependent snapshot-level evidence while preserving type I error guarantees. The paper also ships reproducible code and is unusually transparent about the meaning and limitations of the proposed null. If the novelty claim can be properly qualified relative to Wilson et al. (2017) and the invalid 'max' calibrator is removed from the main performance claims, the result would be a useful practical addition to the network testing toolkit. However, the current simulation narrative is partly based on a non-valid procedure, so the empirical support for the method's power comparisons must be reworked.","major_comments":[{"comment":"The 'max' calibrator is explicitly defined in Section 3.1 as not a calibrator, yet it is included as one of the five compared calibrators in all simulation studies and is repeatedly described as giving the best performance (Sections 4.2.3, 4.3.3, 4.4.3). Because the integral of g_max over [0,1] diverges, a uniform p-value under the null can yield E[g_max(P)] = infinity, so the large e-values and high rejection rates reported for 'max' do not demonstrate valid power. The authors should remove 'max' from the main comparisons or clearly label it as an invalid oracle, and re-run the headline simulation summaries using only valid calibrators (kappa = 0.25, 0.5, 0.75, avg) to confirm that the qualitative conclusions change.","section":"Section 3.1, Section 4.1, Figures 1-3"},{"comment":"The claim that there are 'no statistical methods available' to test for community structure in temporal networks and that the proposed test is 'the first hypothesis test for community structure in temporal networks' is difficult to reconcile with the authors' own description of Wilson et al. (2017) as deriving a hypothesis test for community structure in multilayer networks, of which temporal networks are a special case. The authors should either specify the precise difference in null hypotheses or objectives that makes their test novel, or soften the novelty claim.","section":"Section 1 and Section 3.2"},{"comment":"The paper acknowledges the infinite e-value issue in the Reality network, but the same phenomenon occurs in the dynamic SBM and dynamic DCBM simulations, where 'all p-values were 0... not plotted' (Sections 4.3.3 and 4.4.3). This means that a single snapshot with p=0 dominates the average, so the reported e-value is not a finite measure of 'average' community strength in exactly the settings used to claim power. The authors should report the frequency of zero p-values and either use a censored or regularized calibrator, or discuss this limitation more prominently in the simulation sections.","section":"Section 6 and Sections 4.3.3, 4.4.3"}],"minor_comments":[{"comment":"The statement 'Then \\bar E_T is an e-value in the sense of (1)' refers to the p-value definition in equation (1); it should refer to the e-value definition in equation (2).","section":"Section 3.2, Theorem 1"},{"comment":"There is a typo in 'We stress that he goal of these simulations' which should read 'the goal'.","section":"Section 4.1"},{"comment":"The transition probability notation '\\pi_{k,k+1}' appears to be a typo for '\\pi_{k,k'}' given the surrounding text; the current subscript is confusing.","section":"Section 4.3.1"},{"comment":"The phrase 'underling strength' should be 'underlying strength'.","section":"Section 6"},{"comment":"The figure legends label 'max' as one of the 'Calibrator' series even though the text states it is not a proper calibrator; renaming the series to 'max (invalid)' or a similar label would avoid misleading readers.","section":"Figures 1-3 and 8-10"}],"recommendation":"major_revision","confidential_remarks":"The core theorem is correct and the paper is honest about limitations. My main concern is that the invalid 'max' calibrator is used as a primary point of comparison and is presented as the best performer, which can mislead readers about the validity guarantee; this should be addressed before publication. The novelty claim also needs qualification relative to Wilson et al. (2017)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core claim is correct but modest: average e-values constructed from valid per-snapshot p-values form an e-value, so you get valid type-I error control under arbitrary temporal dependence. The new application—testing for community structure in temporal networks by averaging calibrated per-snapshot evidence—is not in the literature, and the paper does a decent job laying out the idea and its limitations. The writing is clear, there's code on GitHub, and the real-data section using the Chung-Lu null is a nice touch. It is a useful baseline and a good tutorial on e-values.\n\nThe main soft spot is the 'max' calibrator. The authors correctly state in Section 3.1 that g_max is not a calibrator, but then they use it in every simulation figure, report it as the best performer, and base the real-data calibrator choice on the simulation results. Because E[g_max(P)] is not bounded by 1, the large e-values it produces are not valid evidence. This doesn't affect Theorem 1, which only covers proper calibrators, but it undermines the experimental section and could mislead a practitioner who follows the comparison. It should either be removed or clearly flagged as invalid in the plots and tables, not presented as a competitive option.\n\nThe 'first hypothesis test for temporal community structure' claim is also too strong. The paper itself cites Wilson et al. (2017), which gives a test for multilayer networks and notes temporal networks are a special case. The authors should either soften the claim or argue more carefully why that test doesn't apply. The averaging also makes the test invariant to snapshot order, and the authors admit it only captures roughly constant-strength community structure. That's a real scope limit, but it's acknowledged in the Conclusion, so it's a caveat rather than a hidden flaw.\n\nI think the paper deserves a serious referee. It's a legitimate, clearly written contribution, but the calibrator issue needs to be addressed before publication. The fix is straightforward: drop 'max' from the comparisons, or mark it as invalid and use it only to illustrate the risk. Also, the firstness claim should be qualified.","headline":"Correct but modest: valid e-value averaging applied to temporal community testing, but an invalid calibrator is showcased as best and the firstness claim is overstated.","tokens_in":18008,"tokens_out":2999,"would_cite":false,"duration_ms":34418,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Averaging per-snapshot e-values produces the first statistically valid test for community structure in temporal networks.","keywords":["e-values","temporal networks","community detection","hypothesis testing","stochastic block model","dynamic networks","p-value calibration","type I error control"],"falsifier":"Generate a temporal network with $T=10$ snapshots, where snapshots 1-5 are drawn from a stochastic block model with strong community structure and snapshots 6-10 from an Erdős–Renyi model. If the average of the calibrated e-values remains below the rejection threshold (say, 20) even though the per-snapshot p-values in the first half are tiny, then the claimed test does not detect the temporal community structure it is meant to test.","tokens_in":16995,"feed_emoji":"🕸️","tokens_out":9590,"duration_ms":96272,"temperature":0.7,"pith_summary":"This paper proposes the first hypothesis test for the presence of community structure in a temporal network, where data arrive as a sequence of network snapshots over time. The procedure computes a valid p-value for static community structure on each snapshot, converts each p-value into an e-value using a calibrator, and averages the e-values. Because the arithmetic mean of e-values is itself an e-value, the resulting average preserves type I error control under arbitrary dependence among snapshots, which is exactly the complication that makes p-value combination intractable. The method is demonstrated on correlated and dynamic stochastic block models and on five real-world networks, with the test proposed as a modular extension any static community-structure test can be plugged into. A stated caveat is that averaging implicitly targets community structure whose strength is roughly constant over time, so the test does not capture intermittent, merging, splitting, or accumulating structures.","feed_headline":"E-values yield first temporal community-structure test","feed_subtitle":"Averaging per-snapshot e-values preserves validity under unknown dependence.","key_machinery":"The central object is the e-value, a nonnegative random variable whose expectation under the null hypothesis is at most one; large e-values indicate evidence against the null, and $1/E$ is a valid p-value. The argument turns on the combination property that the arithmetic mean of arbitrarily dependent e-values is again an e-value. Each snapshot p-value is transformed by a p-to-e calibrator, including $g_\\kappa(p)=\\kappa p^{\\kappa-1}$, the averaged calibrator $g_{\\mathrm{avg}}(p)=(1-p+p\\log p)/(p(-\\log p)^2)$, and the comparison 'max' function $g_{\\max}(p)=-e^{-1}/(p\\log p)$ for $p\\le e^{-1}$, which the paper notes is not itself a calibrator. Theorem 1 then follows directly from the definition of an e-value, and the average is the test statistic.","core_discovery":"Let $G^{(1)},\\dots,G^{(T)}$ be snapshots of a temporal network, let $P_t$ be a valid p-value for the null hypothesis of no community structure on the $t$-th snapshot, and let $E_t = g(P_t)$ be the e-value obtained by applying a p-to-e calibrator such as $g_\\kappa(p)=\\kappa p^{\\kappa-1}$. Theorem 1 asserts that $\\bar E_T = \\frac{1}{T}\\sum_{t=1}^T E_t$ is an e-value, meaning $\\sup_{P\\in\\mathcal{P}_0}\\mathbb{E}_P(\\bar E_T)\\le 1$; hence a large average is evidence against the null, and the test's validity does not depend on how the snapshots are correlated. To the authors' knowledge, this is the first hypothesis test for community structure in temporal networks. The construction is modular: any static community-structure hypothesis test that yields valid p-values can be used, and the paper implements it both with an Erdős–Renyi null using Tracy-Widom asymptotics and with a Chung-Lu null using bootstrap.","pith_inferences":["Because the average is invariant to snapshot ordering, the test cannot distinguish communities that appear, merge, split, or disappear; detecting such phenomena would require a change-point or state-switching formulation rather than an averaging one.","A weighted average updated sequentially as new snapshots arrive would remain an e-value at every time, suggesting an anytime-valid monitoring procedure for emerging community structure without multiple-testing corrections.","The calibrator step inflates small p-values to maintain validity, which the paper identifies as a source of inefficiency; e-values constructed directly from a network model, rather than via p-value calibration, are the natural route to a more powerful temporal test."],"forward_implications":["Any valid static community-structure test can be substituted into the framework, so the temporal test automatically inherits future improvements in static testing.","Weighted averaging, with weights summing to one, remains a valid e-value, allowing users to place more emphasis on recent snapshots without losing error control.","Using the rejection rule $\\bar E_T > 20$ bounds the type I error at $0.05$, and the simulations report low rejection rates under the null and increasing power as community structure strengthens.","The same averaging scheme applies to multilayer networks, where each layer plays the role of a snapshot, extending the test beyond the temporal setting."],"supporting_citations":[{"why":"Supplies the static community-structure p-value test used on each snapshot, based on the largest eigenvalue under an Erdős–Renyi null.","marker":"Bickel and Sarkar (2016)"},{"why":"Source of the p-to-e calibrators and the result that averaging e-values preserves the e-value property.","marker":"Vovk and Wang (2021)"},{"why":"Provides the convention that an e-value above 20 corresponds to a type I error of at most 0.05.","marker":"Wang (2023)"},{"why":"Provides the bootstrap static test with a Chung-Lu null used in the real-data analysis.","marker":"Yanchenko and Sengupta (2024)"},{"why":"The dynamic stochastic block model used to generate temporal networks with time-varying community labels.","marker":"Matias and Miele (2017)"},{"why":"Frames the taxonomy of temporal community structure that motivates the paper's discussion of what the test can and cannot detect.","marker":"Cazabet and Rossetti (2023)"}],"fun_headline_variants":["E-values unlock first temporal network community test","First community-structure test for evolving networks","E-values enable temporal community detection test","Temporal networks: e-values test community structure","New e-value test for community structure over time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The test is meaningful only if temporal community structure can be represented by the average of per-snapshot static community evidence, so the strength of the structure must be roughly constant over time and the ordering of snapshots is treated as irrelevant.","fun_headline_variants_meta":{"raw":{"variants":["E-values unlock first temporal network community test","First community-structure test for evolving networks","E-values enable temporal community detection test","Temporal networks: e-values test community structure","New e-value test for community structure over time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2687,"prompt_tokens":892,"completion_tokens":1795,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":1728}},"tokens_in":508,"tokens_out":1795,"duration_ms":16393,"temperature":1.0,"reasoning_tokens":1728,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:07:27.665117+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a temporal network with $T=10$ snapshots, where snapshots 1-5 are drawn from a stochastic block model with strong community structure and snapshots 6-10 from an Erdős–Renyi model. If the average of the calibrated e-values remains below the rejection threshold (say, 20) even though the per-snapshot p-values in the first half are tiny, then the claimed test does not detect the temporal community structure it is meant to test.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the static community-structure p-value test used on each snapshot, based on the largest eigenvalue under an Erdős–Renyi null."},{"cited_title":"and Wang, R","cited_arxiv_id":null,"evidence_quote":"Source of the p-to-e calibrators and the result that averaging e-values preserves the e-value property."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the convention that an e-value above 20 corresponds to a type I error of at most 0.05."},{"cited_title":"and Miele, V","cited_arxiv_id":null,"evidence_quote":"The dynamic stochastic block model used to generate temporal networks with time-varying community labels."},{"cited_title":"and Rossetti, G","cited_arxiv_id":null,"evidence_quote":"Frames the taxonomy of temporal community structure that motivates the paper's discussion of what the test can and cannot detect."}],"review_version":1}