{"id":"1cf0840a-e7e5-4fd4-b823-e9647cc1d922","arxiv_id":"2412.16647","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Degree correlations in growing citation networks follow from the marginal dependencies of two mechanisms: individual growth dynamics and causal inheritance between citing and cited papers.","lead":"This paper derives a growth model for causal networks that explains why highly cited papers are cited by other highly cited papers. The model uses a compact parameter set and matches citation data across four scientific fields.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal kernel in Eq. (8) is fitted on the same four networks used for validation, so the reported degree correlations may reflect in-sample fitting rather than emergent prediction; an out-of-sample test is needed.","rationale":"The reader's weakest_assumption correctly identifies the causal kernel factorization in Eq. (8) and the in-sample fitting issue as the key risk. My independent reading confirms that the analytical framework is coherent: Eqs. (5)-(6) are plausible consequences of the stationary master equation, and the derivation of knn and assortativity follows. The load-bearing problem is empirical validation, not mathematics. The paper's own text admits that the kernel is 'best fitted by the mixed Weibull distribution function' from the same data, and the prior work [42] that supplies λ̄ was also calibrated on WOS. Because no parameter values are given in the main text, a reader cannot even reproduce Table I without the supplement, and no out-of-sample network is tested. A leave-one-out or holdout test would cleanly resolve whether the kernel generalizes. This concern does not invalidate the framework; it means the predictive claim is stronger than the current evidence justifies. The reader's CONDITIONAL verdict is appropriate, so I recommend no change.","tokens_in":8187,"tokens_out":2484,"duration_ms":24605,"concrete_test":"Perform a leave-one-out validation: fit the mixed-Weibull K̃ and the λ̄ relation in Eq. (8) using only three of the four disciplines (e.g., biology, chemistry, mathematics), then use those fixed functional forms to predict the fourth discipline's P(k,k'), knn(k'), and assortativity from Eq. (4). If the predicted quantities deviate from empirical measurements by more than the reported theory-experiment differences, the kernel is not universal and the in-sample agreement does not support the emergence claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that degree correlations in causal networks emerge from marginal dependencies on dynamic correlations (Green's function G) and causal correlations (kernel K). The mathematical derivation of Eqs. (5)-(6) is internally consistent, but the empirical validation rests on a load-bearing assumption: the causal kernel K(λ|λ',t) = (1/λ̄(λ',t)) K̃(λ/λ̄(λ',t)) in Eq. (8), where λ̄(λ',t) is taken from prior work [42] and K̃ is fitted to a mixed Weibull distribution using the same WOS data. If K̃'s shape parameters are chosen from the four networks whose P(k), P(k,k'), knn(k'), and assortativity are then reported as theoretical predictions, the agreement in Figs. 1-3 and Table I is partially circular: the kernel can absorb the degree correlations the paper claims to derive. The paper does not list the fitted values (k0, Weibull coefficients, etc.) in the main text, and it does not test the kernel on any network outside the four used for calibration. The Discussion's claim of reducing parameters from O(N) to O(1) is only meaningful if the kernel is universal and specified once, not re-fit per network. This is a validation gap, not an internal inconsistency, but it directly undermines the strength of the 'emergence' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes a mean-field growth framework for causal (directed acyclic) networks in which each vertex is endowed with a latent state (fitness) and an individual-level growth process described by a Green's function. The authors derive a stationary master equation, Eq. (4), that self-consistently determines the state distribution and growth rate from the causal kernel and the Green's function. From that solution they compute the degree distribution, joint degree distribution, k-nearest-neighbor function, and assortativity, via Eqs. (5), (6), (9), and (10). The theory is applied to four Web of Science citation networks (biology, chemistry, mathematics, physics), and the authors report quantitative agreement for network growth rates, fitness distributions, degree distributions, joint degree distributions, knn(k'), and assortativity. The central mathematical claim is that degree correlations are not directly imposed but emerge from marginal dependencies in the joint state-age distribution. The paper argues that its framework reduces the number of free parameters from O(N) in fitness models to O(1), because causal dependencies among fitness values are encoded in a universal causal kernel.","tokens_in":8553,"tokens_out":2587,"duration_ms":25152,"significance":"If the claims hold, the paper would provide a valuable analytic unification: a single self-consistent equation linking individual growth (dynamic correlations) and edge-formation rules (causal correlations) to macroscopic degree correlations in DAGs. The derivation of Eqs. (5)-(6) is nontrivial and internally consistent, and the conditional-independence observation is a clean conceptual contribution. The empirical validation on four large citation networks is a strength, and the paper is generally clear about the distinction between dynamic and causal correlation. However, the significance of the empirical results is currently tempered by the fact that the central input to the theory, the causal kernel K in Eq. (8), is fitted on the same four networks that are later used for validation, and the fitted parameter values are not reported in the main text. The 'O(1) parameters' claim is therefore not yet fully supported. Still, the theoretical framework is sufficiently well posed that an out-of-sample test or a fully specified kernel could substantially raise its value.","major_comments":[{"comment":"The empirical validation is partially in-sample: the causal kernel factorization K(λ|λ',t) = (1/λ̄(λ',t)) K̃(λ/λ̄(λ',t)) is measured and the mixed-Weibull form of K̃ is fitted on the same four WOS networks that are then used to report agreement in Figs. 1-3 and Table I. If the kernel absorbs the degree correlations the theory claims to derive, the reported agreement is not a strong test of the emergence mechanism. The manuscript should provide an out-of-sample validation (e.g., calibrate the kernel on three disciplines and predict the fourth) or otherwise demonstrate that the kernel is universal and not merely an efficient parametric fit to each target network.","section":"Empirical Validation, Eq. (8)"},{"comment":"The mean fitness λ̄(λ',t) is stated to be 'an exponential function of the ratio k(λ',t)/k∞(λ'), as shown in our previous work [42]', but the exact functional form is not given here. Because Eq. (8) is one of the two central inputs to the stationary master equation, the paper is not self-contained and the predictions cannot be reproduced without consulting a separate arXiv preprint. Please state the explicit form of λ̄(λ',t) or reproduce it in the Supplemental Material, and give the fitted values of k0, m, and the mixed-Weibull parameters used for each network or for the assumed universal kernel.","section":"Eq. (8) and preceding text"},{"comment":"The claim that the framework reduces parameters from O(N) to O(1) is not yet substantiated. The manuscript fixes σ=1 and μ=2.3 globally, but k0, m, and the mixed-Weibull parameters are not tabulated, and it is unclear whether they are re-fitted for each of the four disciplines. If these parameters are refitted per network, the effective number of free parameters is O(1) per network, not O(1) across all causal networks. The Discussion should either state one universal parameter vector valid for all four networks or clarify precisely which parameters are shared and which are discipline-specific.","section":"Discussion, parameter-count claim"},{"comment":"Equation (6) shows that k and k' are conditionally independent given state and age, with correlations mediated by Ψ(s,τ;s',τ'). This is a correct conditional-independence statement, but the phrase 'degree correlation is entirely encoded in the state-age correlation within Ψ' may overstate the content: the joint state-age distribution Ψ itself is determined by the same causal kernel K that is fitted to data. The manuscript should clarify that the emergence claim is about the conditional factorization, not about the origin of K. This would help the reader distinguish the analytic decomposition (which is new and useful) from the empirical claim that the kernel is fundamental rather than phenomenological.","section":"Eq. (4) and Section Network Characteristics"}],"minor_comments":[{"comment":"There is a typo in the Introduction: 'a general correlated growth framework for casual networks' should read 'causal networks'.","section":"Abstract and Introduction"},{"comment":"'Substituting Eq. (2) into Eq. (1) and taking t → ∞ yields leads the stationary master equation' should be 'yields' or 'leads to'.","section":"Theoretical Framework, paragraph after Eq. (4)"},{"comment":"The notation for the joint state-age distribution Ψ(s,τ;s',τ') is introduced in the text but the formula immediately after it is not assigned an equation number; giving it a number would make the cross-references in Eqs. (5) and (6) easier to follow.","section":"Network Characteristics, Eq. (6)"},{"comment":"The sentence 'the empirical ψ(λ) is obtained by fitting the RPP model to individual papers' raises a question: is that fit performed per paper (which would reintroduce O(N) fitting for the empirical baseline), and how is the empirical ψ(λ) then compared with the theory? A brief clarification of the fitting procedure would remove ambiguity.","section":"Empirical Validation, discussion of ψ(λ)"},{"comment":"In the sentence preceding Eq. (9), 'up to a normalization factor' is vague; the exact normalization of the conditional distribution P(λ,τ;λ',τ'|k') should be specified in the Supplemental Material to allow reproduction of the knn curves.","section":"Eq. (9)"},{"comment":"Table I would benefit from a column or footnote stating the number of vertices and the citation counts for each discipline, so the reader can judge the statistical weight of the reported errors.","section":"Table I caption"},{"comment":"In Figure 1, panel (d) shows the rescaled kernel for physics only; the caption notes similar plots are in the Supplementary Material. Since the kernel universality is a load-bearing assumption, it would be helpful to show the rescaled kernels for all four disciplines in the main text, even in a small inset.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid analytic core and a plausible empirical case, but the validation strategy currently mixes calibration and testing, and the parameter-count claim is stronger than the reported evidence. I would support publication after the authors provide a clearer separation between fitted and predicted quantities and, ideally, an out-of-sample test. The paper's scope fits physics.soc-ph, though the contribution is more methodological than discovery-driven. I did not find signs of problematic citation practices, but the heavy reliance on the authors' own prior work for the kernel form should be clearly stated in the revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"At bottom, this paper delivers a clean mean-field derivation of degree correlations in growing DAGs, and that part is genuinely new. The master equation (4) and the joint degree formula (6) are not in the cited literature, and they show something useful: conditional on a vertex's state and age, degrees are independent, so all observable degree correlations funnel through the state-age distribution. The math is internally consistent, and the authors are clear about the modeling assumptions.\n\nThe soft spot is empirical validation. The causal kernel K(λ|λ',t) in Eq. (8) is measured on the same four WOS citation networks that are then compared to the theory's predictions. The shape of K̃ is fit to a mixed Weibull, and key parameters (k0, m, the Weibull coefficients) are not given in the main text. So the agreement in Figs. 1–3 is partly in-sample: the fitted kernel can absorb whatever correlations the formula then 'predicts.' The O(1) parameter claim only holds if the kernel is universal and specified once, not re-fit for each network. The authors mention previous work [42] for the mean fitness scaling, but that doesn't resolve the circularity for K̃.\n\nTo be clear, this is a validation gap, not an internal inconsistency. The derivation stands on its own. What's missing is an out-of-sample test—e.g., using the kernel measured on one corpus to predict correlations on another—or at minimum a fully transparent statement of all fitted values and a sensitivity analysis. The discussion of super-exponential growth and connection to condensation is a nice pointer, and the paper is honest about the stationarity limitation.\n\nWho should read this? Network scientists working on growing DAGs or citation dynamics will find the framework useful. It deserves a serious referee: the theory is interesting, the derivation is sound, and the empirical overreach is fixable with more rigorous validation. I'd recommend conditional acceptance with a request for out-of-sample tests and full parameter reporting.","headline":"A clean mean-field theory for correlations in growing DAGs; the math is solid, but the empirical validation is in-sample and needs an out-of-sample test.","tokens_in":9036,"tokens_out":2347,"would_cite":true,"duration_ms":18941,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["05C82","05C20"],"pacs":["89.75.Fb","89.75.Hc"],"model":"deepseek-v4-flash","headline":"This paper claims that degree correlations in causal networks emerge from two microscopic correlations—memory in individual degree growth and state transmission from parent to child—and that a single stationary master equation predicts…","keywords":["causal networks","degree correlations","directed acyclic graphs","network growth","citation networks","assortativity","reinforced Poisson process","latent fitness"],"falsifier":"Fit G and K on one causal network, say physics citations, and use them with Eq. (4) to predict P(k,k') and assortativity on an out-of-sample causal network such as patent citations without refitting the kernel; if the predicted assortativity falls outside empirical error bars, the claim that degree correlations emerge from these two inputs fails. Alternatively, hold fitness values fixed but shuffle the parent-child pairing in the data while preserving the univariate marginals of K: if the predicted P(k,k') is unchanged, then the kernel is not carrying the causal correlation the paper attributes to it.","tokens_in":2006,"feed_emoji":"🔗","tokens_out":4399,"duration_ms":79806,"temperature":0.7,"pith_summary":"This paper tries to establish that the degree correlations observed in causal networks are not imposed by rewiring rules but emerge from marginalizing over two underlying correlations: dynamic correlation in how a single node's degree grows over time, and causal correlation in how new events inherit state from the nodes they link to. It packages these into a stationary master equation whose solution yields the growth rate, the fitness distribution, the degree distribution, the joint degree distribution, the nearest-neighbor degree, and the assortativity coefficient. Applied to four Web of Science citation networks spanning biology, chemistry, mathematics, and physics, the predictions match empirical measurements without per-node fitting. If the paper is right, topological correlations in a directed acyclic graph are a fingerprint of growth mechanics rather than incidental structural noise.","feed_headline":"Two hidden correlations drive citation networks' degree structure","feed_subtitle":"44M-paper test shows causal and dynamic correlations predict degree distribution and assortativity without per-paper fitting.","key_machinery":"The machinery is the stationary master equation $\\psi(s) = \\frac{1}{\\langle k\\rangle} \\int_0^\\infty d\\tau \\int ds'\\, K(s|s',\\tau)\\, \\partial_\\tau k(s',\\tau)\\, e^{-r\\tau}\\, \\psi(s')$, fed by two inputs: the Green's function $G_s(k,\\tau|0,0)$ for individual degree growth (dynamic correlation) and the causal kernel $K(s|s',\\tau)$ for state transmission along edges (causal correlation). For the reinforced Poisson process used here, the Green's function is a negative binomial distribution, and the kernel is factorized as $K(\\lambda|\\lambda',t) = \\frac{1}{\\bar{\\lambda}(\\lambda',t)} \\tilde{K}\\left(\\frac{\\lambda}{\\bar{\\lambda}(\\lambda',t)}\\right)$, with the mean fitness $\\bar{\\lambda}$ an exponential function of $k(\\lambda',t)/k_\\infty(\\lambda')$ and $\\tilde{K}$ a mixed Weibull fitted from the data. Solving the master equation self-consistently determines the growth rate $r$ and stationary state distribution $\\psi(s)$, from which all network observables follow by averaging Green's functions over the joint state-age distribution.","core_discovery":"The central claim is that in a growing directed acyclic graph, the joint state-age distribution $\\Psi(s,\\tau;s',\\tau') = \\frac{1}{\\langle k\\rangle} K(s|s',\\tau'-\\tau)\\, \\partial_{\\tau'} k(s',\\tau'-\\tau)\\, \\Psi(s',\\tau')$ fully encodes degree correlations. Because degrees are conditionally independent given state and age, the joint degree distribution factors as an average of Green's functions, $P(k,k') = \\left\\langle \\frac{k'}{k(s',\\tau')} G_{s'}(k',\\tau'|0,0)\\, G_s(k,\\tau|0,0)\\right\\rangle$, so all topological correlation is mediated by marginal dependencies in the state-age distribution. The paper derives closed forms for the nearest-neighbor degree $k_{\\mathrm{nn}}(k')$ and the assortativity coefficient $r_{\\mathrm{corr}}$ from the same average, and validates them against citation networks from four disciplines.","pith_inferences":["If the fitted mixed-Weibull kernel is universal, the same kernel shape should reproduce degree correlations in an out-of-sample causal system such as patent citations or legislative citations; relative success or failure there would test the universality claim.","Because the derivation only needs degrees to be conditionally independent given state and age, the mechanism likely extends to any growing DAG with a low-dimensional latent state, not just citation networks.","A practical use the authors do not spell out is forecasting: measuring G and K on a young network may predict its eventual assortativity, turning the framework into a prediction tool rather than only a fitting tool.","A subtle circularity check would be to ask whether the mixed-Weibull fit is already absorbing the degree correlations the paper claims to derive; this can be probed by holding univariate marginals fixed while destroying the parent-child pairing in the data."],"forward_implications":["The number of free parameters needed to model a causal network drops from O(N) to O(1): latent fitness values are not assigned per vertex but generated by the causal kernel.","The observed saturation of k_nn(k') at large k' is a genuine prediction of the causal kernel combined with the finite ultimate degree k_infinity, not a finite-size artifact.","Assortativity in citation networks is reproduced quantitatively, so the framework explains the sign and magnitude of degree correlation without invoking dynamic rewiring.","Because stationarity requires time-translation invariance, the same framework predicts a super-exponential, winner-takes-all regime when the causal kernel has sufficiently broad tails, analogous to Bose-Einstein condensation in fitness models.","The framework is written for general causal systems, so the same equations apply to social media event cascades, biological evolution, and economic growth wherever a directed acyclic causal structure grows in time."],"supporting_citations":[{"why":"Supplies the exponential form of the mean fitness $\\bar{\\lambda}(\\lambda',t)$ as a function of $k(\\lambda',t)/k_\\infty(\\lambda')$ used in the kernel factorization Eq. (8).","marker":"[42]"},{"why":"Establishes the reinforced Poisson process with preferential attachment and aging used as the individual-level Green's function G.","marker":"[12]"},{"why":"Provides the RPP popularity-dynamics model whose negative-binomial Green's function is used for individual citation growth.","marker":"[41]"},{"why":"Defines the assortativity coefficient r_corr whose covariance-and-variance form the paper evaluates through Eq. (10).","marker":"[35]"},{"why":"Supplies the randomized-network normalization R(k,k') = P(k,k')/P_rand(k,k') used to display joint degree correlations.","marker":"[28]"},{"why":"Provides the fitness-model baseline whose O(N) parameterization and Bose-Einstein condensation phase the paper contrasts with its O(1) causal-kernel approach.","marker":"[25]"},{"why":"Represents the intrinsic-fitness model class that assigns a fitness to each vertex, the limitation the causal kernel is designed to remove.","marker":"[38]"},{"why":"Provides the preferential-attachment growth scaffold for growing networks that the causal framework extends.","marker":"[24]"}],"fun_headline_variants":["Causal and dynamic correlations drive citation network structure","New growth model explains degree correlations in causal networks","Two correlations predict degree patterns in citation networks","Why citation networks correlate: a general growth framework","Unified theory reveals origin of degree correlations in causal nets"],"cache_read_input_tokens":11136,"weakest_assumption_plain":"The load-bearing premise is that the causal kernel factorizes as in Eq. (8) with a single universal mixed-Weibull shape, so the kernel fitted on the four observed networks is the actual mechanism generating degree correlations rather than an absorbing fit that merely encodes those correlations.","fun_headline_variants_meta":{"raw":{"variants":["Causal and dynamic correlations drive citation network structure","New growth model explains degree correlations in causal networks","Two correlations predict degree patterns in citation networks","Why citation networks correlate: a general growth framework","Unified theory reveals origin of degree correlations in causal nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000387,"raw_usage":{"total_tokens":1996,"prompt_tokens":854,"completion_tokens":1142,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":1070}},"tokens_in":470,"tokens_out":1142,"duration_ms":9167,"temperature":1.0,"reasoning_tokens":1070,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:22:44.667622+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit G and K on one causal network, say physics citations, and use them with Eq. (4) to predict P(k,k') and assortativity on an out-of-sample causal network such as patent citations without refitting the kernel; if the predicted assortativity falls outside empirical error bars, the claim that degree correlations emerge from these two inputs fails. Alternatively, hold fitness values fixed but shuffle the parent-child pairing in the data while preserving the univariate marginals of K: if the predicted P(k,k') is unchanged, then the kernel is not carrying the causal correlation the paper attributes to it.","supporting_citations":[{"cited_title":"Correlated Impact Dynamics in Science","cited_arxiv_id":"2303.03646","evidence_quote":"Supplies the exponential form of the mean fitness $\\bar{\\lambda}(\\lambda',t)$ as a function of $k(\\lambda',t)/k_\\infty(\\lambda')$ used in the kernel factorization Eq. (8)."},{"cited_title":"Quantifying long-term scientific impact","cited_arxiv_id":null,"evidence_quote":"Establishes the reinforced Poisson process with preferential attachment and aging used as the individual-level Green's function G."},{"cited_title":"Modeling and predicting pop- ularity dynamics via reinforced poisson processes","cited_arxiv_id":null,"evidence_quote":"Provides the RPP popularity-dynamics model whose negative-binomial Green's function is used for individual citation growth."},{"cited_title":"Assortative mixing in networks","cited_arxiv_id":null,"evidence_quote":"Defines the assortativity coefficient r_corr whose covariance-and-variance form the paper evaluates through Eq. (10)."},{"cited_title":"Specificity and stability in topology of protein networks","cited_arxiv_id":null,"evidence_quote":"Supplies the randomized-network normalization R(k,k') = P(k,k')/P_rand(k,k') used to display joint degree correlations."},{"cited_title":"Bose- einstein condensation in complex networks","cited_arxiv_id":null,"evidence_quote":"Provides the fitness-model baseline whose O(N) parameterization and Bose-Einstein condensation phase the paper contrasts with its O(1) causal-kernel approach."},{"cited_title":"Scale-free networks from vary- ing vertex intrinsic fitness","cited_arxiv_id":null,"evidence_quote":"Represents the intrinsic-fitness model class that assigns a fitness to each vertex, the limitation the causal kernel is designed to remove."},{"cited_title":"Emergence of scaling in random networks","cited_arxiv_id":null,"evidence_quote":"Provides the preferential-attachment growth scaffold for growing networks that the causal framework extends."}],"review_version":1}