{"id":"40aa7556-9447-4ca0-b9f4-9933b9036154","arxiv_id":"2504.17701","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A comparative study shows that network sampling method performance depends on network type and target metric, with simple methods often beating advanced ones on temporal networks.","lead":"This paper compares eight network sampling methods on a static collaboration network and a temporal messaging network. It finds that no single method dominates, and simpler methods can outperform advanced ones on temporal data, so method choice should depend on network type and target metrics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Temporal-network conclusion rests on a two-method, single-dataset comparison with no confidence intervals, so the claimed 'simpler methods work better on temporal networks' is not established.","rationale":"The reader's weakest_assumption correctly identifies the temporal conclusion as resting on a two-method, one-dataset comparison. My stress test reaches the same point: the abstract and conclusion make a general claim about temporal networks, but §2.3 provides only UNS versus PRS on CollegeMsg, with no statistical uncertainty quantified. This is the least secure link in the central argument because the static-network part of the paper is more substantial: it uses eight methods (despite the text saying six) on CA-HepTh, with 100 replicates and multiple metrics, and includes a CLT consistency check in Fig. 6. The static result is a reasonable empirical contribution. However, the temporal result is what distinguishes the paper's takeaway from the existing 'no universal winner' consensus, and that part is under-supported. I do not think this requires rejection, because the paper is transparent about its small scope and frames itself as a modest comparative study; the appropriate remedy is to add temporal experiments or soften the generalization. Since the reader already recommended CONDITIONAL, my read does not change that verdict, so I set verdict_should_be to UNCHANGED.","tokens_in":7637,"tokens_out":3890,"duration_ms":41044,"concrete_test":"Run the temporal experiment on at least three additional temporal networks (e.g., email-Eu-core-temporal, Bitcoin OTC alpha, and sx-Mathoverflow) using the full method suite from the static experiment—UNS, WNS, UES, IES, RWS, MHRWS, SS, BFS, and PRS—at comparable sample fractions, with at least 100 replicates per method. Report mean relative error and 95% bootstrap confidence intervals for average degree, clustering coefficient, edge percentage, and largest component size at each snapshot. If UNS does not outperform the exploration/edge-based methods on most temporal datasets and metrics, the abstract's 'simpler techniques can be more effective' generalization should be removed or explicitly limited to the two methods and one dataset studied.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central and most distinctive claim is the static/temporal contrast: advanced methods do well on static networks, but simpler methods are more effective on temporal networks (abstract, §5). The temporal evidence for this claim is confined to §2.3 and Fig. 7, where only two methods are compared—Uniform Node Sampling (UNS) and PageRank Sampling (PRS)—on one dataset (CollegeMsg). Both methods are node-based, so the entire 'advanced methods' category on temporal networks is represented by a single PageRank variant; no edge-based or exploration-based method is tested temporally. The sample is fixed at 30 nodes drawn from the t=0 snapshot, and the later snapshots contain as few as 69 edges (Table 2), so 30-node induced subnetworks can be extremely sparse and high-variance. The paper reports no confidence intervals, bootstrap errors, effect sizes, or statistical tests for any temporal metric; Fig. 7 is qualitative. Consequently, the conclusion that 'simpler techniques often yield better estimates in temporal settings' could be an artifact of choosing PageRank as the only 'advanced' comparator, of the single temporal dataset, or of the small fixed sample size, rather than a general property of temporal networks. The Discussion (§4) acknowledges the two-dataset and fixed-metric limitations but does not acknowledge that the temporal result relies on only two methods, which is the more severe gap for the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents an empirical comparison of network sampling methods on two real-world datasets: a static scientific collaboration network (CA-HepTh/HEP-TH) and a temporal message-sending network (CollegeMsg). The authors organize sampling methods into node-based, edge-based, and exploration-based categories and compare their ability to preserve metrics such as average degree, clustering coefficient, largest component size, average shortest path, edge percentage, and s-metric. Static experiments use 100 independent samples fixed at 1,000 nodes; temporal experiments fix the sampled node set at 30 nodes drawn from the initial snapshot and track the induced subnetworks over time. The paper reports that no single method consistently outperforms others, that exploration-based methods perform well on static networks, and that simpler methods such as uniform node sampling can be more effective on temporal networks. It also reports a Central Limit Theorem check for uniform node sampling and concludes that sampling strategy should be tailored to network type and target metric.","tokens_in":7909,"tokens_out":5850,"duration_ms":58700,"significance":"If the static/temporal contrast were firmly established, the paper would offer practically useful guidance for practitioners choosing sampling methods, and the explicit taxonomy of node-, edge-, and exploration-based methods is a clear organizing framework. The static experiment design, with 100 replications at a sample size of 1,000 nodes, is a reasonable empirical setup, and the authors are transparent about using public datasets and about some limitations of the study. However, the paper's most distinctive claim, that advanced methods underperform simpler ones on temporal networks, rests on very thin evidence: two methods, one temporal dataset, and no inferential statistics. The static 'most robust method' conclusion is also presented without quantitative uncertainty measures. These gaps make the current claims broader than the evidence supports.","major_comments":[{"comment":"The central claim that 'simpler techniques can be more effective' on temporal networks is not established by the evidence presented. The temporal experiment compares only Uniform Node Sampling and PageRank Sampling on a single dataset (CollegeMsg), with the sample fixed at 30 nodes taken from G(t=0). Table 2 shows that later snapshots are very small: at t=4 there are only 69 edges and an average degree of 1.19, so induced 30-node subnetworks can be extremely sparse and high-variance. The paper reports no confidence intervals, bootstrap errors, effect sizes, or statistical tests for any temporal metric, and Fig. 7 is purely qualitative. The observed reversal could be an artifact of the choice of PageRank as the only 'advanced' method, of the un-described selection of 116 persistent users, or of the fixed small sample size, rather than a general property of temporal networks. Section 4's limitation list acknowledges only the two-dataset and fixed-metric limitations and does not mention that the temporal conclusion rests on a two-method comparison. The authors should either substantially expand the temporal evaluation (more methods, more datasets, inferential statistics) or materially restrict the claims in the abstract and conclusion.","section":"§2.3, Fig. 7, §4"},{"comment":"The static conclusion that RWS and SS are 'the most robust' methods is based on visual inspection of boxplots. The manuscript does not report error measures, confidence intervals, or pairwise comparisons across the 100 samples; statements such as 'consistently approximated' and 'significant deviations' are not quantified. Given the visible separation in Fig. 5, this is likely fixable by reporting summary statistics such as mean absolute error or root mean square error with associated uncertainties, but as written the ranking of methods is not quantitatively established.","section":"§3, Fig. 5"},{"comment":"The manuscript is internally inconsistent about the number of methods used in the static experiment. Section 2.2 says 'we compare six methods' but then lists eight methods (UNS, WNS, UES, IES, RWS, MHRWS, SS, BFS). Figure 4's caption says six methods, Fig. 5 says eight methods, while Sections 4 and 5 say six. The reader cannot determine which methods were actually included in each figure. Please reconcile the method lists, figure captions, and text, and state explicitly which methods are included in each experimental setting.","section":"§2.2, Fig. 4, Fig. 5, §4, §5"}],"minor_comments":[{"comment":"The dataset name is inconsistent across the manuscript: 'CA-HepTH' in the abstract/introduction, 'Arxiv HEP-TH' in Section 2.2, and 'CA-HepTh' in Table 1. Please use one consistent name.","section":"§2.2, Table 1"},{"comment":"The formula for average degree is missing the normalization by the number of nodes: it should be ⟨k⟩ = (1/n) Σ_i k_i. The numerical values in Table 2 are consistent with this normalized definition, so this appears to be a typographical error, but it should be corrected.","section":"Eq. (1)"},{"comment":"PageRank Sampling is not fully specified: the manuscript does not state on which graph PageRank is computed (the t=0 snapshot or the full temporal aggregate), what damping factor is used, or how many nodes are selected. This is needed for reproducibility.","section":"§2.3"},{"comment":"The text says the distributions approach normality 'as the number of samples increases,' but the experimental design uses a fixed number of 100 samples and varies the sample size (number of nodes). Please correct the wording to refer to increasing sample size.","section":"§3, Fig. 6"},{"comment":"There are several typographical errors: 'Methematics' and 'Unites States' in the author affiliation, 'egdes' in Table 2, and an incomplete word in Eq. (5) ('connecte'). These should be corrected.","section":"Throughout"},{"comment":"The data availability statement names the Stanford Network Analysis Project but does not give dataset versions, access dates, or any code. Releasing the analysis code would substantially improve reproducibility.","section":"§7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a workmanlike empirical comparison, but its most distinctive claim about temporal networks requires substantially more evidence than is currently provided. If the authors only fix the method-count inconsistency and typos without strengthening the temporal analysis or appropriately limiting the claims, I would not recommend acceptance. The static comparison would also benefit from quantitative uncertainty measures, but that issue is more readily addressable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the static half of this paper is a decent, if unoriginal, empirical replication; the temporal half is where it overreaches. The headline claim that 'simpler techniques can be more effective' on temporal networks is supported by exactly one comparison—UNS vs. PageRank on CollegeMsg—with no confidence intervals, no statistical tests, and no other temporal dataset. That is not enough to carry a general conclusion.\n\nWhat the paper does well: the static-network comparison is reasonably careful. It uses 100 samples of 1,000 nodes from CA-HepTh, evaluates eight methods across six metrics, and checks normality of the UNS sampling distribution under the CLT. The result that exploration-based methods (RWS, SS, MHRWS) generally do better on static networks while uniform methods lag is consistent with prior work, including Blagus et al. [24], which the paper cites. As a teaching or reference piece, the static section has some value.\n\nThe soft spots are real. First, the temporal conclusion is not just weak; it is basically anecdotal. The experiment samples 30 nodes once at t=0 from a 116-node subgraph and then evaluates induced subnetworks over later snapshots. Two methods, one dataset, no error bars. PageRank is the only 'advanced' method tested temporally, so the entire claim about advanced methods failing in temporal settings rests on a single comparator. Second, the paper says 'six methods' in several places but lists eight in Section 2.2 and Figure 5; that inconsistency should have been caught. Third, the Discussion lists limitations about dataset diversity and metrics but misses the more serious problem: the temporal evidence is under-powered by design.\n\nIf the paper's purpose is to survey sampling methods, the static results are fine as a replication. If the point is the static/temporal contrast, that needs substantially more work: more methods on the temporal side, more datasets, and proper uncertainty quantification. As it stands, I would not cite the temporal claim. For peer review, I would send this back for major revision at best; honestly, in its current form I would be inclined to desk reject and invite a resubmission after the temporal experiment is expanded.","headline":"A decent but unoriginal static-network comparison is undermined by an over-generalized temporal claim resting on two methods, one dataset, and no statistical tests.","tokens_in":8361,"tokens_out":2858,"would_cite":false,"duration_ms":29187,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that no network sampling method consistently beats the others, with advanced methods doing better on static networks and simple methods on temporal ones, so method choice must adapt to network type and target metric.","keywords":["network sampling","temporal networks","static networks","random walk sampling","PageRank sampling","uniform node sampling","clustering coefficient","degree distribution"],"falsifier":"Repeat the temporal experiment on a second temporal network, such as an email or phone-call dataset with snapshots, and compare uniform node sampling with PageRank sampling on the same metrics. If PageRank sampling matches or beats uniform node sampling on structural metrics in that network, the claimed temporal inversion fails. A second check is to compute the degree heterogeneity of CollegeMsg over time: the paper's suggestion that the temporal network has uniformly random features predicts low heterogeneity, so a strongly scale-free temporal degree distribution would undermine the explanation.","tokens_in":7461,"feed_emoji":"📊","tokens_out":5819,"duration_ms":56135,"temperature":0.7,"pith_summary":"Network sampling selects a small subgraph to stand in for a full network, so the choice of method determines whether sampled metrics reflect reality. This paper compares representative methods from three families — node-based, edge-based, and exploration-based — on a static scientific collaboration network and a temporal message-sending network. Its central finding is that no single method wins everywhere: random-walk and snowball sampling track static structure best, while on the temporal network the simple uniform node sampling beats the more sophisticated PageRank sampling on structural metrics. The practical message is that sampling strategy should be tuned to network type and target metric rather than chosen once for all graphs.","feed_headline":"No sampling method wins on every network","feed_subtitle":"A two-dataset comparison finds static and temporal networks favor different sampling strategies.","key_machinery":"The central object is the sampled subgraph $G_s=(V_s,E_s)$ generated from $G=(V,E)$ by each method, and the evaluation protocol that surrounds it. For the static network the protocol draws 100 independent samples at fixed node counts and compares six metrics — average degree, clustering coefficient, largest component ratio, average shortest path, density, and the s-metric — against the full network's values. For the temporal network the protocol samples at time $t=0$ and keeps the same node set across later 40-day snapshots, which separates the method's sampling bias from the network's temporal decay.","core_discovery":"The central claim is that sampling-method performance is context-dependent and the direction of the effect can invert with network type. On the static CA-HepTh collaboration network, exploration-based methods are superior: random walk sampling and snowball sampling preserve clustering, degree, and the largest component, while uniform node and edge sampling fragment the graph. On the temporal CollegeMsg network, the ordering flips: uniform node sampling estimates node and edge structure well but approximates connectivity poorly, while PageRank node sampling does the reverse. The paper also reports that uniform node sampling produces metric estimates whose distributions converge to normal as the number of samples grows, consistent with the Central Limit Theorem despite the network's power-law degree distribution.","pith_inferences":["If the temporal inversion holds beyond CollegeMsg, a practical decision rule emerges: measure activity concentration or degree heterogeneity first, then choose a sampling bias direction — toward hubs for static structure, away from hubs for temporal structure.","The Central Limit Theorem observation could be developed into a formal error-bar method for arbitrary network metrics, a step the paper does not take.","The paper's suggestion that the temporal network has uniformly random features is testable against a temporal null model that reshuffles messages in time while preserving the aggregate degree sequence.","The fixed-node-set temporal design points toward an adaptive streaming sampler: sample once, track the same nodes, and re-sample only when metric drift exceeds a threshold."],"forward_implications":["On static networks, exploration-based methods such as random walk and snowball sampling are the safer default when connectivity and clustering matter.","On temporal networks, centrality-biased sampling can distort structural metrics, so uniform node sampling deserves a place as a baseline rather than being dismissed as naive.","Benchmarking sampling methods should be decomposed by network type and by metric, since overall rankings hide opposite orderings.","Uniform node sampling can yield statistically reliable global estimates even when the sample destroys local structure, which supports using sample-mean confidence intervals for network metrics.","A single early sampling time can capture later temporal dynamics when the sampled node set is held fixed, at least on the CollegeMsg network."],"supporting_citations":[{"why":"Defines uniform node sampling, the baseline that wins on temporal structural metrics.","marker":"[9]"},{"why":"Supplies weighted node sampling, the degree-based method tested on the static network.","marker":"[10]"},{"why":"Defines PageRank node sampling, the advanced method that underperforms on the temporal network.","marker":"[11]"},{"why":"Defines uniform edge sampling, one of the edge-based baselines.","marker":"[16]"},{"why":"Defines induced edge sampling, another edge-based baseline.","marker":"[17]"},{"why":"Defines random walk sampling, a strong performer on the static network.","marker":"[18]"},{"why":"Defines Metropolis-Hastings random walk sampling, another exploration-based method.","marker":"[19]"},{"why":"Defines snowball sampling, which along with random walk tracks static structure best.","marker":"[21]"},{"why":"Supplies the static CA-HepTh collaboration dataset used for the static comparison.","marker":"[35]"},{"why":"Supplies the temporal CollegeMsg dataset used for the temporal comparison.","marker":"[36]"}],"fun_headline_variants":["Sampling method success flips on temporal networks","Static vs temporal networks: different sampling winners","No one-size-fits-all strategy for network sampling","Temporal networks flip the sampling rulebook","Best network sampling depends on network type"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The temporal half of the conclusion — that simpler methods can outperform advanced ones on temporal networks — rests on comparing only two methods, uniform node sampling and PageRank sampling, on a single temporal dataset, CollegeMsg; if that two-method, one-dataset comparison is atypical, the temporal claim does not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Sampling method success flips on temporal networks","Static vs temporal networks: different sampling winners","No one-size-fits-all strategy for network sampling","Temporal networks flip the sampling rulebook","Best network sampling depends on network type"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1600,"prompt_tokens":812,"completion_tokens":788,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":720}},"tokens_in":428,"tokens_out":788,"duration_ms":7903,"temperature":1.0,"reasoning_tokens":720,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:32:30.797303+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the temporal experiment on a second temporal network, such as an email or phone-call dataset with snapshots, and compare uniform node sampling with PageRank sampling on the same metrics. If PageRank sampling matches or beats uniform node sampling on structural metrics in that network, the claimed temporal inversion fails. A second check is to compute the degree heterogeneity of CollegeMsg over time: the paper's suggestion that the temporal network has uniformly random features predicts low heterogeneity, so a strongly scale-free temporal degree distribution would undermine the explanation.","supporting_citations":[{"cited_title":"Metropolis algorithms for representative subgraph sampling","cited_arxiv_id":null,"evidence_quote":"Defines Metropolis-Hastings random walk sampling, another exploration-based method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines uniform node sampling, the baseline that wins on temporal structural metrics."},{"cited_title":"Adamic, Rajan M","cited_arxiv_id":null,"evidence_quote":"Supplies weighted node sampling, the degree-based method tested on the static network."},{"cited_title":"Sampling from large graphs","cited_arxiv_id":null,"evidence_quote":"Defines PageRank node sampling, the advanced method that underperforms on the temporal network."},{"cited_title":"Krishnamurthy, M","cited_arxiv_id":null,"evidence_quote":"Defines uniform edge sampling, one of the edge-based baselines."},{"cited_title":"Ahmed, Jennifer Neville, and Ramana Kompella","cited_arxiv_id":null,"evidence_quote":"Defines induced edge sampling, another edge-based baseline."},{"cited_title":"Butts, and Athina Markopoulou","cited_arxiv_id":null,"evidence_quote":"Defines random walk sampling, a strong performer on the static network."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines snowball sampling, which along with random walk tracks static structure best."},{"cited_title":"Graph evolution: Densification and shrinking diameters","cited_arxiv_id":null,"evidence_quote":"Supplies the static CA-HepTh collaboration dataset used for the static comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the temporal CollegeMsg dataset used for the temporal comparison."}],"review_version":1}