{"id":"86a985b1-c4ed-46a6-83b9-1f6a134e6c4f","arxiv_id":"2608.09788","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The power-law exponent of a growing hypergraph's degree distribution depends only on the average ratio of new nodes to hyperedge size, and preferential attachment monotonically raises the simplicial fraction up to the gelation transition.","lead":"A new model of how hypergraphs grow shows that a single ratio, the average number of new nodes per new hyperedge, sets the tail of the degree distribution, no matter how hyperedge sizes vary. The authors use this to argue that preferential attachment, the 'rich get richer' rule, is what makes many real hypergraphs look simplicial.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Small-p regime invalidates the core prediction: Appendix C shows the simulated tail deviates from Theorem 1 exactly where all real datasets sit (p ≈ 0.001), and the paper offers no convergence proof.","rationale":"I read the paper's central claim as the ratio-only exponent. The mean-field algebra in Appendix A is internally consistent, and the closed form in Theorem 1 appears correct once the fraction is parsed appropriately. The most vulnerable point is not the algebra but the relevance of the limit: the paper's own validation shows a systematic discrepancy in the small-p regime, and all datasets are there. The reader's chosen weakest assumption—the backward-stepping estimator's meaning in closed populations—is worth noting, but it is partly mitigated by the fact that the estimated p̂ reduces to |V|/Σ|e|, a static incidence ratio that is well-defined even for fixed populations; the real unresolved issue is whether the asymptotic degree-distribution prediction is accurate enough to support the 'fit without parameters' claim. The monotone simpliciality finding is plausible and supported by null-model comparisons, though the α-fitting is not predictive. Overall, the reader's conditional verdict is appropriate: the paper should either provide finite-size scaling or restrict the universality claim to the regime where it is numerically supported.","tokens_in":10486,"tokens_out":42340,"duration_ms":381300,"concrete_test":"Run the Appendix C simulation for p = 0.001, 0.005, and 0.01 with increasing final times t = 10^5, 10^6, 10^7 (or extrapolate from smaller t), and measure the empirical CCDF tail, e.g., the 99.9th percentile of hyperdegree. If the deviation from Corollary 1 does not shrink monotonically with t, or if the fitted tail exponent stays below 1/(1−p)+1 without tracking the theoretical curve, the finite-size explanation is unsupported and the theorem's applicability to the small-p regime fails. As a second check, simulate two growth rules with identical p but different Y_t distributions (e.g., Y_t ≡ 6 vs. Y_t uniform on {2,10}) at p = 0.001 and compare their CCDFs; material divergence would directly contradict the 'independent of shapes' universality claim at realistic sizes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the ratio-only power law of Theorem 1/Corollary 1. Its derivation rests on the mean-field approximation in Eq. (8), which replaces Δ_t = Y_t − X_t by its mean and treats selections as independent. Appendix C explicitly reports that for small p the simulated CCDF is systematically heavier-tailed than the theoretical prediction, calling it a 'finite-size effect that becomes increasingly pronounced as p→0'. All eight empirical datasets in Table 1 have fitted p between 5e−4 and 1.5e−2, precisely the small-p regime. The paper never demonstrates convergence to the thermodynamic limit in this regime: for p ≈ 0.001, the mean-field condition Δ_t d/D_t ≪ 1 holds only when t ≫ p^{−1/p} ≈ 10^{3000}, so the promised exponent γ = 1/(1−p)+1 ≈ 2 is not testable at any accessible size, and the 'parameter-free fit' does not yield a usable degree-distribution prediction for the data it is fitted to. The authors decline to validate Theorem 1 against empirical degree sequences (Section 6.1), so the paper's applied claims rest on an unverified asymptotic in the regime of interest.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes a generalized preferential-attachment hypergraph model in which each new hyperedge has size Y_t and contains X_t new nodes and Y_t − X_t existing nodes selected with probability proportional to hyperdegree. The authors derive, by a mean-field approximate master equation in the thermodynamic limit, that the stationary hyperdegree distribution is P(d_H) ∝ d_H^{-(1/(1-p)+1)}, with p = E[X_t]/E[Y_t], and that the exponent depends only on p, not on the shapes of the input distributions (Theorem 1, Corollary 1). They propose a backward-stepping estimator of {X_t, Y_t} from timestamped hypergraphs, fit it to eight XGI-DATA datasets (p̂ from 5e-4 to 1.5e-2), and, using a nonlinear preferential-attachment variant with exponent α, report that the simplicial fraction σ_SF increases monotonically with α up to the gelation transition at α=1 and then decreases. The paper concludes that preferential attachment is a simpliciality-enforcing mechanism. Appendices contain the AME derivation, the CCDF proof, and simulation validation (Y_t ~ Poisson(5), X_t ~ Binomial(Y_t,p)).","tokens_in":10735,"tokens_out":36916,"duration_ms":299682,"significance":"If Theorem 1 holds, the ratio-only universality is a clean, publishable contribution: it generalizes the Avin et al. and Wang et al. models, the recursion in Appendix A solves correctly, the gamma-function form is properly normalized (verified for test values of p and d_H), and Corollary 1's CCDF is consistent. Credit is due for shipping open code, for stating the mean-field nature of the derivation up front, and for honestly reporting in Appendix C that the simulated tail is heavier than the theoretical prediction for small p and declining in Section 6.1 to claim empirical validation against degree sequences. However, the manuscript's empirical significance is weaker than its framing: all eight datasets lie in the small-p regime where the paper's own simulations deviate from the theory, the backward-stepping estimator embeds an unflagged node-birth assumption that is violated by four of the eight datasets, and the simpliciality 'confirmation' is a fit of α to σ_SF rather than a sharp test. These are fixable framing and validation gaps rather than defects in the analytic core.","major_comments":[{"comment":"The central universality claim is left unvalidated precisely in the regime where all the empirical data sit. Table 1 reports p̂ between 5e-4 and 1.5e-2 for all eight datasets, yet Appendix C reports that in this small-p range the simulated CCDF is systematically heavier-tailed than the prediction of Corollary 1, attributing the discrepancy to a finite-size effect that 'becomes increasingly pronounced as p→0'. No convergence proof, error bound, or scaling analysis is provided to show that the deviation vanishes in the thermodynamic limit; the mean-field gain probability differs from the exact without-replacement selection probability by a relative amount of order (Y_t−X_t)·d_H/D_t ∼ t^{-p} for the dominant nodes, so for p≈1e-3 the approach to the limit occurs only on astronomically large timescales. Section 6.1 then explicitly declines to test Theorem 1 against the empirical degree sequences. The abstract's claims that the model can be fitted to any timestamped dataset 'without parametric assumptions' and that the exponent depends only on p therefore overstate what is established; the authors should either supply a quantitative convergence analysis for the small-p regime or explicitly restrict the universality and fitting claims and present γ≈2 as an untested limiting prediction.","section":"§3.2 (Theorem 1), §6.1, Appendix C"},{"comment":"The backward-stepping estimator asserts that a node left isolated by removing e_T 'therefore first appeared when e_T was added.' This inference is valid only for a growing process in which nodes are born at their first hyperedge and never pre-exist the observation window. Four of the eight datasets (Contact-high-school, Contact-primary-school, Hospital-Lyon, Malawi-Village) are closed-population proximity contact networks whose individuals pre-exist; a node that appears in exactly one hyperedge is then counted as a new node, which inflates X_t and biases p̂. The paper neither flags this assumption in Section 5.2 nor assesses the bias (for example, by recomputing p̂ under the opposite convention that all nodes are present from the start). Since the fitted p̂ underpins the small-p and γ≈2 statements and the 'parameter-free fitting' contribution, this assumption and a sensitivity analysis must be added.","section":"§5.2"},{"comment":"The evidence for 'preferential attachment as a simpliciality-enforcing mechanism' is presented as confirmation, but the empirical side of the argument is a fit: α is chosen so that the NLPA model reproduces the observed σ_SF, and the monotone model trend is then read as confirming the mechanism. Since σ_SF is matched by construction at the fitted α, the match itself carries limited evidential weight; what the data actually support is that the empirical σ_SF lies on the model curve in the plausible range α∈[0,1.1]. The monotone increase of σ_SF with α is a genuine, falsifiable prediction of the model, but it is not tested out-of-sample here. A sharper test would compare additional statistics at the single fitted α without further tuning (for example, σ_SF restricted to hyperedges of a given size, or degree-conditional downward closure), and the text should be rephrased so that the confirmation claim matches the strength of that test.","section":"§6.2, Figures 2–9"}],"minor_comments":[{"comment":"The text states that the first mean-field approximation replaces Y_t−X_t by its conditional expectation, but Eq. (8) retains the random quantity Y_t−X_t; the two-stage conditioning-and-averaging procedure should be described consistently.","section":"§3.2 and Appendix A, Eq. (8)"},{"comment":"The sentence 'Entries marked in red are to be completed' and the phrase 'for the eight datasets for which they are currently available' indicate that the manuscript is an unfinished draft; the captions should be brought in line with the presented content or the missing entries supplied.","section":"Table 1 caption"},{"comment":"The definition of a simplicial complex is incomplete: 'A simplicial complex is a hypergraph with an additional structural property: for every hyperedge present' is not a complete sentence and does not state the downward-closure condition.","section":"§2"},{"comment":"The backward-stepping procedure requires a total order on hyperedges, but several datasets (notably 20-second-resolution contact networks) contain simultaneous hyperedges; the tie-breaking rule should be stated.","section":"§5.2"},{"comment":"The simulation study does not report the number of time steps or the final number of nodes used; this is essential for evaluating the claimed finite-size nature of the small-p deviation.","section":"Appendix C"},{"comment":"The abstract locates the gelation transition 'at α>1', while Section 4 defines it at α=1; the wording should be made consistent.","section":"Abstract and §4"},{"comment":"The model description does not specify the initial condition or what happens when Y_t−X_t exceeds the current node count |V_t|; a minimal initial configuration and a rule for the early-time regime are needed.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a preliminary draft: the arXiv date is future-dated and the Table 1 caption states that entries are still to be completed. The paper's content is network science (physics.soc-ph/cs.SI); the math.AT classification fits only through the simplicial-complex motivation and seems mismatched if the journal is an algebraic-topology venue. The authors' own Appendix C limitation is the right place to focus a revision: the stated small-p discrepancy is the point on which the empirical claims turn, and the revision should either supply a scaling or convergence argument or explicitly narrow the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Things you should know up front: this paper has one clean theoretical result and one suggestive empirical study, but the central claim is not tested in the regime where it matters. The generalization of the Avin et al. ratio-only degree distribution to arbitrary X_t and Y_t is a genuine contribution; the derivation in Appendix A is internally coherent, the recursion solves, and the gamma-function form normalizes. The simpliciality study is the first I've seen connecting preferential attachment strength to simpliciality across real datasets, and the null-model comparisons are carefully done. Credit also for shipping code and data.\n\nThe soft spots are real. Appendix C shows that for small p—the regime where all eight datasets sit, p on the order of 0.0005 to 0.015—the simulated CCDF is systematically heavier than Theorem 1 predicts. The paper calls it a finite-size effect but supplies no proof of convergence in that regime; a back-of-the-envelope using their own mean-field condition suggests t would need to be absurdly large before the approximation kicks in. They openly say they don't validate Theorem 1 against empirical degree sequences (Section 6.1), which is honest but means the headline applied claim is unsupported at the sizes anyone can simulate.\n\nSecond, the backward-stepping estimator assumes a growing population: a node left isolated after removing a hyperedge is assumed to have first appeared when that hyperedge was added. That's fine for co-authorship or email, but for closed-population proximity networks (Hospital-Lyon, Malawi-Village) the individuals pre-exist and may appear in one hyperedge only; the estimator then counts them as new, biasing p downward. This assumption is load-bearing for the empirical small-p conclusion and is not flagged.\n\nThird, the simpliciality empirical argument: they fit alpha to reproduce the observed simplicial fraction and then present the monotone trend as confirmation. That's a fitting-to-target component; the qualitative monotonic increase up to gelation is still worth reporting, and the attachment-driven vs distribution-driven distinction is interesting.\n\nWho should read this: network scientists working on higher-order models, and anyone who fits preferential attachment to hypergraph data. It deserves a serious referee: the derivation is worth engaging and the flaws are addressable. The authors should either prove convergence in the small-p regime or explicitly reframe the claim for moderate/large p, and they should fix the estimator or state the growing-process assumption. Send it to review with major-revision expectations.","headline":"The ratio-only mean-field result is clean and worth engaging, but the paper's own simulations undercut it in the small-p regime where all its datasets sit, and the backward-stepping estimator has a load-bearing assumption.","tokens_in":11247,"tokens_out":3202,"would_cite":true,"duration_ms":26514,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["05C65","05C82","60C05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that in a generalized preferential attachment hypergraph model, the stationary hyperdegree distribution follows a power law whose exponent depends only on the ratio p of the expected number of new nodes per step to the…","keywords":["hypergraph","preferential attachment","simpliciality","hyperdegree distribution","power law","mean-field approximation","gelation transition","backward-stepping estimation"],"falsifier":"Construct a synthetic timestamped hypergraph from a closed population with known node birth times (all nodes present before the first hyperedge) and run the backward-stepping procedure: if it reports p substantially above the true value of 0, or if the degree tail predicted by Theorem 1 with the fitted p fails to match the simulated tail, the empirical small-p claims and the parameter-free fitting claim are called into question.","tokens_in":10296,"feed_emoji":"","tokens_out":2852,"duration_ms":25565,"temperature":0.7,"pith_summary":"This paper asks what generative rules produce the high simpliciality observed in real hypergraph data—where subsets of hyperedges tend to also be hyperedges. It introduces a preferential attachment model in which each new hyperedge has a random size and recruits a random number of new nodes, and proves that the tail of the resulting hyperdegree distribution is a power law controlled by a single number: the ratio p of the average number of new nodes per step to the average hyperedge size. The shape of the two distributions otherwise does not matter, so any two growth rules sharing the same ratio produce the same degree tail. The paper also shows that this model can be fit to timestamped data without parametric assumptions, and that strengthening preferential attachment raises the simplicial fraction until the gelation transition at α > 1, where a single hub dominates and simpliciality drops. If true, the ratio p becomes a compact, data-estimable summary of how a higher-order network grows, and preferential attachment is a concrete mechanism behind observed simplicial structure.","feed_headline":"One ratio sets the power-law tail of growing hypergraphs","feed_subtitle":"New nodes vs. hyperedge size alone fixes the degree exponent, and stronger attachment builds simplicial structure up to gelation.","key_machinery":"The key machinery is a mean-field approximate master equation (AME) for the probability that a node has hyperdegree d_H at time t, in which random quantities are replaced by their conditional expectations and the strong law of large numbers provides deterministic limits for total node count and total hyperdegree. Solving the stationary recursion yields the closed-form gamma-function distribution and the power-law tail. Two further components carry the empirical argument: a backward-stepping procedure that reconstructs the sequence (X_t, Y_t) from any timestamped hypergraph by removing hyperedges in reverse order and counting newly isolated nodes as new arrivals, and the nonlinear preferential attachment kernel p(u) ∝ d(u)^α, whose gelation transition at α = 1 is shown to divide a simpliciality-enforcing regime from a simpliciality-destroying one.","core_discovery":"The central discovery is a universality result for growing hypergraphs: the stationary hyperdegree distribution of the generalized preferential attachment model is P(d_H) = Γ(d_H) Γ(1/(1−p)+1) / ((1−p) Γ(1/(1−p)+d_H+1)), which scales as $d_H^{{-γ}}$ with γ = 1/(1−p) + 1, where p = E[X_t]/E[Y_t] is the limiting ratio of expected new nodes to expected hyperedge size. The exponent depends on the two random processes only through this ratio, not on the shapes of their distributions. The paper further finds, via a nonlinear generalization, that the simplicial fraction σ_SF increases monotonically with the preferential attachment exponent α for α ≤ 1 and decreases for α > 1, with the gelation transition at α = 1 marking the boundary. Across eight empirical datasets, the fitted ratios are all very small (p ≈ 0.0005 to 0.015), situating real systems in the small-p regime where preferential attachment has the strongest potential to enforce simpliciality; the mechanism is that moderate attachment concentrates hyperedges around high-degree nodes, making their neighborhoods dense enough to satisfy downward closure, while superlinear attachment creates a dominant hub whose hyperedges lack their sub-hyperedges.","pith_inferences":["The backward-stepping estimator treats any node that becomes isolated after removing a hyperedge as having first appeared at that step; in closed-population datasets such as proximity contact networks, where individuals pre-exist, this biases the fitted p upward, so the empirical 'small-p' conclusion may be partly an artifact of the estimation procedure.","If the universality holds, p could serve as a one-number fingerprint for comparing higher-order growth processes across domains, and the simpliciality-enforcement curve as a function of α offers a template for classifying systems by how strongly attachment drives their simplicial structure.","A direct testable extension would be to run the backward-stepping procedure on synthetic hypergraphs with known node birth times and closed populations; if the fitted p deviates systematically from the true value, the empirical analysis of real datasets should be re-examined with a birth-time-aware estimator."],"forward_implications":["Any two growth rules for hypergraphs with the same ratio p produce identical stationary degree tails, so the ratio is a universal parameter for degree heterogeneity in growing higher-order networks.","Hyperedge sizes and new-node counts can be read off directly from timestamped data, allowing the model to be fit without assuming particular distributions for X_t or Y_t.","In the small-p regime, which includes all eight empirical datasets, preferential attachment is the dominant driver of simpliciality in sparse systems like email and legislative networks, while proximity networks owe most of their simpliciality to size and degree distributions.","Superlinear preferential attachment actively undermines simpliciality past the gelation threshold, predicting that systems with very strong rich-get-richer dynamics should have lower simplicial fractions."],"supporting_citations":[{"why":"Provides the earlier preferential attachment hypergraph model with one pre-existing node per hyperedge, whose degree distribution form this paper shows is universal.","marker":"[16]"},{"why":"The original Barabási–Albert preferential attachment mechanism that the hypergraph model generalizes.","marker":"[14]"},{"why":"Supplies the nonlinear preferential attachment kernel and the gelation transition for α > 1 that organizes the simpliciality results.","marker":"[19]"},{"why":"Defines the simplicial fraction σ_SF used throughout to quantify simpliciality.","marker":"[12]"},{"why":"The XGI-DATA repository from which all eight empirical timestamped hypergraph datasets are drawn.","marker":"[5]"},{"why":"Establishes simplicial closure and higher-order link prediction as motivation for why simplicial structure matters in real data.","marker":"[4]"}],"fun_headline_variants":["Hypergraph power-law exponent pinned to single ratio","Preferential attachment builds simpliciality in hypergraphs","One ratio controls hypergraph degree distribution","Attachment strength enforces simpliciality up to gelation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The backward-stepping estimator assumes that a node left isolated after removing a hyperedge first appeared exactly when that hyperedge was added, which holds only for growing populations where nodes are born at their first hyperedge and not for closed populations where individuals pre-exist.","fun_headline_variants_meta":{"raw":{"variants":["Hypergraph power-law exponent pinned to single ratio","Preferential attachment builds simpliciality in hypergraphs","One ratio controls hypergraph degree distribution","Attachment strength enforces simpliciality up to gelation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000652,"raw_usage":{"total_tokens":3045,"prompt_tokens":1053,"completion_tokens":1992,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":1930}},"tokens_in":669,"tokens_out":1992,"duration_ms":12984,"temperature":1.0,"reasoning_tokens":1930,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:46:11.469096+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a synthetic timestamped hypergraph from a closed population with known node birth times (all nodes present before the first hyperedge) and run the backward-stepping procedure: if it reports p substantially above the true value of 0, or if the degree tail predicted by Theorem 1 with the fitted p fails to match the simulated tail, the empirical small-p claims and the parameter-free fitting claim are called into question.","supporting_citations":[{"cited_title":"Connectivity of Growing Random Networks","cited_arxiv_id":null,"evidence_quote":"Supplies the nonlinear preferential attachment kernel and the gelation transition for α > 1 that organizes the simpliciality results."},{"cited_title":"The simpliciality of higher- order networks","cited_arxiv_id":null,"evidence_quote":"Defines the simplicial fraction σ_SF used throughout to quantify simpliciality."},{"cited_title":"https: //github.com/xgi-org/xgi-data","cited_arxiv_id":null,"evidence_quote":"The XGI-DATA repository from which all eight empirical timestamped hypergraph datasets are drawn."},{"cited_title":"Simplicial closure and higher-order link prediction","cited_arxiv_id":null,"evidence_quote":"Establishes simplicial closure and higher-order link prediction as motivation for why simplicial structure matters in real data."}],"review_version":1}