{"id":"2fa628fe-d6c4-40dd-b4ee-5e68607dd1e9","arxiv_id":"2509.04946","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In the strong-timescale-separation limit, fast-to-slow couplings do not create statistical dependence between layers, while slow-to-fast feedback couplings do, and can propagate along minimal network paths.","lead":"For networks of coupled stochastic variables that evolve on widely separated timescales, the paper derives the leading-order joint distribution as a chain of conditional distributions and shows that mutual information across layers is created only by slow-to-fast feedback connections, not by fast-to-slow ones. The result could help scientists interpret information flow in neural, ecological, and biochemical multiscale models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Factorization and the MI principles presuppose unique, globally attracting stationary conditional densities for each fast Fokker-Planck operator; no such conditions are stated, so Eq. (33) can fail for metastable or absorbing fast dynamics.","rationale":"The reader's weakest-assumption analysis already identified the load-bearing premise: existence, uniqueness, and sufficiently fast relaxation of the stationary conditional densities for each Fokker-Planck operator. My pass sharpens that concern: it is not merely a formal technicality but a concrete failure mode for metastable or absorbing fast dynamics, directly relevant to the paper's claim of model-independence and to the numerical examples (which all use globally stable, ergodic fast layers). The mPP characterization and linear Gaussian checks are plausible and partly supported by previous work; the main gap is the missing ergodicity/mixing hypothesis. Since this is addressable by stating and verifying explicit conditions, and since the paper's core derivation is a standard adiabatic-elimination result under those conditions, I do not move the verdict: it remains conditionally acceptable provided the conditions are added and the metastable fast-layer check is performed. The typos and missing error bars noted by the reader are secondary and do not alter this assessment.","tokens_in":19342,"tokens_out":14333,"duration_ms":164408,"concrete_test":"Take the three-layer direct-interaction motif of Sec. IV.B.1, choose the fastest layer to be a one-dimensional bistable Ornstein-Uhlenbeck process U(x1)=a(x1^2−1)^2 with barrier height chosen so that the Kramers escape time is comparable to τ2 (e.g., set a so τ_K/τ2≈1 at τ1/τ2=τ2/τ3=10^{-4}). Solve the full Fokker-Planck equation or simulate the Langevin system from two distinct initial preparations (both wells), estimate the MIMMO entries I12, I13, I23, and compare with the factorization prediction I_μν=0. If the MIMMO entries differ from zero beyond sampling error or depend on the initial preparation, Eq. (52) and principle (ii) fail without an added ergodicity condition. Running the same test with the barrier lowered (τ_K≪τ2) should recover the predicted factorization and would confirm that the issue is the missing relaxation condition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central derivation solves, at each order, an equation of the form L_eff_μ p^{st,eff}_{μ|...}=0 (Eqs. (18), (22), (29), (45)) and uses that solution to build the factorized p_eff in Eq. (33). For a general nonlinear Fokker-Planck operator this stationary equation need not have a unique normalizable solution: non-confining forces give no stationary density, multiplicative diffusion can produce absorbing boundaries (as in the contact-process example with g=x(1−x)), and bistable or multistable potentials have stationary densities whose basin selection depends on initial conditions. The assumption in Sec. II.B that timescales are 'infinitely different' does not control relaxation across metastable barriers: a fast variable with Kramers time comparable to the next-slower timescale is not stationary with respect to the slower variables, no matter how small τ1/τ2 is. If the fast layer is non-ergodic, direct interactions can transmit the fast layer's retained initial-condition memory to slower layers, producing nonzero MIMMO entries in the direct-interaction motif of Sec. IV.B.1, in contradiction to principle (ii). Thus Eq. (33), and the universal information principles (i)–(iii), are only valid under an unstated ergodicity/mixing condition on every L_eff_μ and a bound on relaxation times relative to the next scale. Without stating these conditions, the claim that the result holds for 'any set' of SDEs is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a timescale-separation expansion for systems of coupled stochastic differential equations with N separated timescales. The main claim is that, in the limit of infinite timescale separation, the joint probability density factorizes as p_eff_N(x_N,t) times a product of stationary conditional densities for the faster clusters (Eq. (33)), with effective Fokker-Planck operators defined by averaging over faster variables. For pairwise multilayer interactions, the paper defines feedback and direct interactions, introduces minimal propagation paths (mPPs), and derives structural principles for the zero/nonzero pattern of the Mutual Information Matrix for Multiscale Observables (MIMMO). The formalism is tested on linear, neural, ecological, and contact-process dynamics, and a two-timescale feedback loop is shown to produce an effective self-interaction in the slow layer. A connection with previous discrete-state triadic-interaction results is also discussed.","tokens_in":19686,"tokens_out":14360,"duration_ms":148756,"significance":"The three-timescale derivation in Sec. II.B is explicit and formally clean, and the general factorization offers a compact way to understand multiscale stochastic systems beyond the Gaussian and discrete frameworks treated earlier by the authors. The MIMMO principles are falsifiable structural predictions and are tested on four different dynamical models, which is a genuine strength. If the result is properly qualified, this is a useful contribution to the statistical-mechanics understanding of information flow in multiscale networks. However, the paper overstates its generality: the derivation silently assumes unique, globally attracting stationary conditional densities, and the arbitrary-N expansion is not rigorously justified. These are load-bearing issues for Eq. (33) and the information principles.","major_comments":[{"comment":"The derivation assumes that each effective stationary equation L_eff_mu p=0 has a unique, globally attracting, normalizable solution. No such hypothesis is stated. Non-confining drifts admit no stationary density; bistable potentials can have multiple stationary distributions selected by initial conditions; and multiplicative noise with absorbing boundaries, as in the contact-process example with g=x(1-x), requires explicit boundary-condition analysis. The assumption that timescales are 'infinitely different' does not control relaxation across metastable barriers: a fast variable with a Kramers time comparable to the next-slower timescale is not stationary on the slow variables' time even if tau_1/tau_2 -> 0. The stated claims 'any set of SDEs' and Eq. (33) therefore need an explicit ergodicity/mixing condition on each L_eff_mu and a relaxation-time bound.","section":"Sec. II.B; Eqs. (18), (22), (29), (33)"},{"comment":"The general N-timescale expansion is not rigorously justified. Eq. (25) expands in all small parameters epsilon_mu, but Eq. (26) retains terms (epsilon_nu/epsilon_mu) L_mu p^(1_nu) with nu>mu, which are large rather than small. The three-timescale calculation uses a nested expansion (first in epsilon_1, then in epsilon_2), but the general argument does not show that these cross-ratio terms are subdominant or cancel, nor does it provide the solvability conditions that determine the leading-order factor. Without a reformulation in terms of the ratios epsilon_mu/epsilon_{mu+1}, or a matched-asymptotics proof of the order-by-order closure, the arbitrary-N factorization Eq. (33) is not established.","section":"Sec. II.C, Eqs. (25)-(31)"},{"comment":"The sentence that in the feedback motif 'the elements of MIMMO are always non-zero' is stronger than the derivation supports. The factorization p=p_{1|2}p_{2|3}p_3 shows that I_13 can be nonzero, but proving positivity requires non-degeneracy of the conditional dependencies; special coupling functions or parameter choices can make a particular MI vanish. The qualitative theorem is about which entries have nonzero leading-order structure, not literally 'always non-zero'. Please rephrase with the required generic conditions.","section":"Sec. IV.B.2, Eq. (57)"}],"minor_comments":[{"comment":"The denominator contains a typo: p_eff_mu(x_nu,t) p_eff_nu(x_nu,t) should read p_eff_mu(x_mu,t) p_eff_nu(x_nu,t).","section":"Eq. (48)"},{"comment":"The definition of i_mu uses M_mu inside the sum; this should be M_nu.","section":"Eq. (3)"},{"comment":"The text says 'see Fig. 2a' when referring to the propagation-path case; this should be Fig. 4a.","section":"Sec. IV.B.3"},{"comment":"The noise amplitude g_mu(x_mu) appears in the Langevin equation, but the Fokker-Planck operator in Eq. (36) is written with diffusion coefficient D_i only. Please clarify whether g_mu has been absorbed into D_i or explain the notation.","section":"Eqs. (35)-(36)"},{"comment":"There is an incomplete sentence: 'following. Crucially, each layer...' at the start of the discrete-state connection. A reference or explanatory phrase appears to be missing.","section":"Sec. V"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the central idea is likely salvageable, but the two technical gaps -- missing ergodicity assumptions and the unjustified N-timescale expansion -- are load-bearing. I recommend major revision rather than rejection, provided the authors add explicit hypotheses and repair the expansion argument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious, mostly clean paper that takes the authors' discrete-state and Gaussian timescale-separation framework and extends it to generic nonlinear continuous SDEs. The conditional factorization in Eq. (33), the mPP construction in Eq. (46), and the three information principles are genuinely useful. The derivation is self-contained, and the numerical tests across four different dynamics support the qualitative claim that the MIMMO structure is topology-driven rather than model-specific. The two-layer loop with the tanh nonlinearity is a nice worked example. The core factorization is standard adiabatic elimination, but the systematic conditional-dependence structure and the information asymmetry principles are the real contribution.\n\nThe soft spot is the one you flagged: the factorization assumes, at each stage, that the fast conditional density is the unique stationary state of the corresponding effective Fokker-Planck operator and that it is reached before the next timescale changes. That is not true for arbitrary SDEs. Non-confining forces, absorbing boundaries, or metastable wells all break the argument. The contact-process example with g=x(1-x) is precisely a system where boundary behavior matters, which makes the 'any set of SDEs' claim in the abstract too strong. Without an explicit ergodicity/mixing condition, principle (ii) is conditional, not universal. That is the load-bearing gap.\n\nThe general-N expansion in Section II.C is also sketched more than proved. The text drops cross-ratio terms epsilon_nu/epsilon_mu, which are large for nu>mu, with only a comment about retaining dominant terms. That needs either a proper ordering argument or a clearer statement of what is being neglected. The numerics are light—one order of magnitude separation, no error bars, no code/data—and there are a few typos (e.g., the marginal formula in Eq. (49) and the diffusion constant in Eq. (73)) that don't help precision. All of these are fixable.\n\nWho should read it: anyone working on multiscale stochastic dynamics, adiabatic elimination, or information propagation in networks. It deserves a real referee, but with a request for substantial revision: state the ergodicity conditions, fix the general-N argument, and add stronger numerical support. I would accept it for peer review, not desk reject.","headline":"A useful but overbroad generalization of the authors' earlier timescale-separation results; the factorization and MIMMO principles are right when the fast dynamics are ergodic, but the paper never says so.","tokens_in":20151,"tokens_out":3493,"would_cite":true,"duration_ms":36726,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Slow-to-fast feedback is the only source of mutual information in strongly separated stochastic systems; fast-to-slow couplings add none.","keywords":["timescale separation","stochastic differential equations","Fokker-Planck equation","mutual information","multilayer networks","feedback interactions","minimal propagation paths","effective dynamics"],"falsifier":"Observe a two-layer stochastic system with only a fast-to-slow coupling and set the timescale ratio small enough that the fast layer is effectively stationary; estimate the mutual information between layers from long trajectories. The paper predicts zero mutual information in the leading-order limit, so a systematic nonzero value that does not shrink as the ratio decreases would falsify the no-generation principle.","tokens_in":19234,"feed_emoji":"🔁","tokens_out":6352,"duration_ms":62738,"temperature":0.7,"pith_summary":"The paper asks how information flows between processes that evolve on very different timescales. Its answer is a general timescale-separation formula: for any set of coupled stochastic differential equations with widely separated rates, the joint probability of all variables splits into the slowest variable's distribution times a chain of stationary conditional distributions. Which conditional dependencies survive is fixed by the direction of the couplings, not by the equations' details. Slow-to-fast ('feedback') interactions generate mutual information; fast-to-slow ('direct') interactions cannot, and can only relay information that a feedback interaction has already created, along special minimal propagation paths. The authors test the claim on linear, neural, ecological, and contact-process dynamics and reproduce the predicted information structure in all cases.","feed_headline":"Feedback, not feed-forward, creates information across timescales","feed_subtitle":"Slow-to-fast links alone generate statistical dependence; fast-to-slow links only relay it.","key_machinery":"The machine is a hierarchy of effective Fokker-Planck operators: the operator for layer μ is obtained by integrating the original operator over all faster variables using their stationary conditional densities. Repeating this from fastest to slowest produces the factorization and also manufactures a self-interaction in the slow layer when two timescales are reciprocally coupled. The accompanying combinatorial device is the minimal propagation path (mPP), a directed path from a slower layer to a faster layer that remains a propagation path after all layers slower than the endpoint are removed; the set of conditional dependencies for layer μ is exactly the set of slower layers connected to μ b","core_discovery":"The central discovery is an explicit leading-order solution to the coupled Fokker-Planck equation when timescales are infinitely separated: the joint probability factorizes as the slowest layer's distribution times a product of stationary conditional distributions for every faster layer. The solution is obtained by solving each fast layer's operator at stationarity while slower layers are frozen, then averaging each next operator over all faster variables. For pairwise interactions between layers, the surviving conditional dependencies are exactly the layers reachable by a feedback edge or a minimal propagation path in the coarse-grained directed graph. The authors then define a matrix of pa","pith_inferences":["Beyond the paper, at finite (not infinite) timescale separation direct interactions should generate small but nonzero mutual information of order the timescale ratio; the paper's leading-order structure predicts where these corrections appear.","The minimal-propagation-path rule gives a design recipe: to create a functional coupling between a slow regulator and a target, place a feedback edge to any faster layer that path-reaches the target, not necessarily a direct edge to the target itself.","The effective self-interaction loop suggests a minimal stochastic model of slow-timescale memory: fast fluctuations leave a nonlinear trace in the slow equation, a feature that could be fitted in gene-regulatory or neural circuits.","A natural stress test is to break uniqueness of the fast stationary density, for example by bistability, and check whether residual information appears even as the timescale ratio goes to zero; the paper's assumptions exclude this case, so a positive result would define the boundary of the regime."],"forward_implications":["In strongly separated multiscale networks, the pattern of nonzero mutual information becomes a topological fingerprint: entries vanish wherever only fast-to-slow couplings exist.","Regulatory, slow-to-fast couplings are the only structural motifs that create functional information; purely feed-forward chains are information-neutral.","Information can be routed: adding one feedback link creates mutual information not only with its direct target but with every layer connected through minimal propagation paths.","A two-timescale reciprocal loop collapses to an effective self-interaction in the slow variable, explaining how fast fluctuations shape slow statistics without explicit fast variables.","Nonlinear dynamics do not break these principles; they only change the magnitudes of the mutual information, so the predicted structure is testable across model classes."],"supporting_citations":[{"why":"Establishes the discrete-state timescale-separation framework and the conditional-factorization structure that this paper extends to continuous Fokker-Planck dynamics.","marker":"[30]"},{"why":"Supplies the Gaussian multilayer process whose effective-operator solution is the linear baseline generalized here to nonlinear SDEs.","marker":"[40]"},{"why":"Provides the multiple-scale expansion machinery (decimation and averaging) used to derive the order-by-order factorization.","marker":"[47]"},{"why":"Earlier timescale coarse-graining with multiple reservoirs that motivates the timescale-rescaling ansatz in Eqs. (4)-(5).","marker":"[46]"},{"why":"Used to connect mutual-information divergence at linear instability and to fix the information-theoretic behavior of the two-layer loop.","marker":"[21]"},{"why":"One of the network-dynamics models used to test universality of the predicted mutual-information structure.","marker":"[48]"},{"why":"The mutual-information estimator used to compute the information-matrix entries from simulation trajectories.","marker":"[51]"}],"fun_headline_variants":["Slow-to-fast links create information; fast-to-slow only relay it","Feedback loops generate cross-timescale information","Timescale separation reveals feedback as information source","Information arises from slow-to-fast feedback, not fast-to-slow relay"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Every fast layer has a unique stationary conditional distribution that it reaches before slower layers move at all, so the timescale ratios can be treated as infinite; if a fast layer is bistable, absorbing, or only moderately faster, the factorization and its zero-information predictions can fail.","fun_headline_variants_meta":{"raw":{"variants":["Slow-to-fast links create information; fast-to-slow only relay it","Feedback loops generate cross-timescale information","Timescale separation reveals feedback as information source","Information arises from slow-to-fast feedback, not fast-to-slow relay"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000823,"raw_usage":{"total_tokens":3460,"prompt_tokens":793,"completion_tokens":2667,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":2610}},"tokens_in":537,"tokens_out":2667,"duration_ms":19552,"temperature":1.0,"reasoning_tokens":2610,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:46:13.609184+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Observe a two-layer stochastic system with only a fast-to-slow coupling and set the timescale ratio small enough that the fast layer is effectively stationary; estimate the mutual information between layers from long trajectories. The paper predicts zero mutual information in the leading-order limit, so a systematic nonzero value that does not shrink as the ratio decreases would falsify the no-generation principle.","supporting_citations":[{"cited_title":"Tostevin and P","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian multilayer process whose effective-operator solution is the linear baseline generalized here to nonlinear SDEs."},{"cited_title":"De Domenico, A","cited_arxiv_id":null,"evidence_quote":"Provides the multiple-scale expansion machinery (decimation and averaging) used to derive the order-by-order factorization."},{"cited_title":"Avanzini, N","cited_arxiv_id":null,"evidence_quote":"Used to connect mutual-information divergence at linear instability and to fix the information-theoretic behavior of the two-layer loop."},{"cited_title":"Bianconi,Multilayer networks: structure and function(Oxford university press, 2018)","cited_arxiv_id":null,"evidence_quote":"One of the network-dynamics models used to test universality of the predicted mutual-information structure."},{"cited_title":"St-Onge, H","cited_arxiv_id":null,"evidence_quote":"The mutual-information estimator used to compute the information-matrix entries from simulation trajectories."}],"review_version":1}