{"id":"19f57248-35a3-4af1-acd6-cde2bda38783","arxiv_id":"2411.11710","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In a multiscale stochastic network, combining signals before the nonlinearity (integration) encodes more input-output information than combining them after the nonlinearity (summation), with fast processing and certain layer sizes giving the largest gains.","lead":"This paper compares two pooling strategies in a stochastic three-layer network: combine incoming signals first and then apply a nonlinearity, or apply the nonlinearity to each signal first and then combine them. It finds that combining first, called nonlinear integration, usually lets the output encode more information about the input, and that the advantage is largest when the processing layer is fast and the layer sizes are matched to the input dimension.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's 'always' claim is internally contradicted by its own Fig. 2i-j: in the small-MP, large-sigma_OP, strong-coupling regime, nonlinear summation yields higher mutual information than integration.","rationale":"The reader correctly identifies the central claim as the 'always' assertion and notes both the internal qualification and the lack of error bars. My stress test agrees that this is the main load-bearing issue: the paper's own Fig. 2i-j and SM Fig. S3 display a regime where summation beats integration, directly contradicting the strongest claim as stated. I do not share the reader's weakest-assumption emphasis on the factorization imported from earlier work, because the timescale-separated Gaussian factorization is a controlled leading-order result in the small parameters tau_P/tau_O and tau_O/tau_I, and the paper validates it visually in Fig. 1. Strong coupling does not invalidate the singular perturbation; it only makes the conditional means more nonlinear. The more pressing problem is logical and statistical: the 'always' claim is too broad, and the numerical evidence lacks uncertainty quantification. The proposed concrete test would settle whether the reversal is genuine or an artifact of the approximation/noise. If genuine, the claim must be restricted and the verdict remains CONDITIONAL; if not, the claim survives but should be reworded to 'generally' with error bars. This is a partial agreement: we reach the same verdict, but through a different emphasis on the mechanism of the concern.","tokens_in":25775,"tokens_out":5026,"duration_ms":50216,"concrete_test":"Recompute Fig. 2i-j (or SM Fig. S3) for a representative reversal point (e.g., MP = 5, sigma_OP = 10, g_PI = g_OP = 10) using both the analytic factorized sampling and a direct Langevin simulation of the full three-unit system at timescale ratios tau_P/tau_O = 10^-2 and 10^-3, estimating I_IO with at least 10^3 random-matrix realizations and bootstrap standard errors. If I_IO^ns > I_IO^int persists in the full simulation, the 'always' claim is falsified; if it flips, the analytical result was an artifact and the headline claim gains support but still needs error bars.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim in Sec. II C is that 'independently of the internal timescale ordering, nonlinear integration always leads to higher mutual information with respect to summation.' This is directly undermined by the paper itself: in the same section and in SM Fig. S3, for small processing dimension MP, large interaction variance sigma_OP, and strong couplings (g_PI = g_OP around 10), the reported results show I_IO^ns > I_IO^int. This is not an external objection but the authors' own numerical finding. The abstract softens to 'over a wide range of parameters,' so the discrepancy could be patched by rewording. However, the figures provide no error bars on the mutual information estimates (only averages over 10^3 matrix realizations), so it is impossible to tell whether the claimed reversal is a robust effect or sampling noise, and likewise whether the claimed advantage of integration in other panels is significant. The load-bearing condition for the headline result is that integration outperforms summation in the highlighted regimes; the paper contains a regime where it does not, so the 'always' statement is false as written. The concern is not with the timescale-separation factorization itself (which is controlled by the small timescale ratios, not by coupling strength), but with the scope and verification of the ordering claim. A concrete test is to compute the mutual information in the reversal region with full Langevin simulation and with bootstrap error bars; if the reversal persists, the claim must be restricted to the demonstrated parameter ranges and the verdict should remain CONDITIONAL. If it disappears, the analytical approximation was misleading and the 'always' claim may be true after all, but only after adding uncertainty quantification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript studies a three-unit stochastic dynamical system (input, processing layer, output) with timescale separation and nonlinear activation functions. It derives a joint stationary distribution in the limits of fast and slow processing, expressed as products of Gaussian conditional distributions whose means depend nonlinearly on the conditioning variables. Using this factorization, the authors compute the input–output mutual information for two processing schemes—nonlinear summation (process-then-sum) and nonlinear integration (sum-then-process)—and report that integration yields higher mutual information over a wide parameter range, that fast processing enhances information, that an optimal processing-layer size exists, and that integration promotes tunable output bistability.","tokens_in":26090,"tokens_out":8298,"duration_ms":76726,"significance":"The analytical framework is a genuine strength: it avoids fitted parameters, provides an efficient sampling scheme, and is cross-checked against Langevin simulations in Fig. 1. If the ordering result is robust, the work offers a general principle for motif design in biological and artificial networks—integrate before nonlinearity to improve encoding. The practical relevance is high, as the model covers architectures such as random recurrent neural networks and signaling cascades. However, the headline claim is more general than the evidence presented, and the numerical support lacks uncertainty estimates; these issues must be resolved before the paper can serve as a foundation.","major_comments":[{"comment":"The assertion in Sec. II.C that, \"independently of the internal timescale ordering, nonlinear integration always leads to higher mutual information with respect to summation\" is false as written: the authors' own Figs. 2i-j and SM Fig. S3 show a regime with small processing dimension M_P, large interaction variance σ_OP, and strong couplings (g_PI = g_OP around 10) where I^ns_IO > I^int_IO. The main-text caption explicitly acknowledges this (\"at large variances ... such advantage is achieved only for large enough M_P\"). Please replace \"always\" with a qualified statement that explicitly carves out this regime, or restrict the claim to match the abstract's \"over a wide range of parameters.\"","section":"Sec. II.C; Figs. 2i-j; SM Fig. S3"},{"comment":"All mutual-information values are reported as averages over 10^3 random-matrix realizations, but no error bars, confidence intervals, or realization-to-realization statistics are provided. Because the central claim is an ordering between I^int_IO and I^ns_IO, the absence of uncertainty quantification makes it impossible to judge whether the reported reversal at small M_P is a statistically robust effect or sampling noise, and likewise whether the advantages in other panels are significant. Please provide bootstrap or standard-error estimates, and show the distribution of per-realization I_IO for at least one point in the reversal regime and one point in the advantage regime.","section":"Sec. II.C; Figs. 2a-k; SM Figs. S3-S5"},{"comment":"The factorization is imported from the authors' prior work and is controlled by timescale ratios rather than coupling strength, so I do not question its leading-order validity. However, the numerical comparisons are made at strong couplings (g = 5-10), while the only direct Langevin verification (Fig. 1) is at g = 5 for a single parameter set. A direct check of the mutual information at a strong-coupling point in the reversal region and at a point in the advantage region would confirm that the ordering is not an artifact of the conditional-Gaussian sampling at these parameters.","section":"Sec. II.B; Eqs. (3)-(4); SM Eqs. (S34), (S42)"}],"minor_comments":[{"comment":"There is a sign inconsistency in the definition of h_{O|I}: the main text defines h_{O|I} = ⟨p_{O|I} log2 p_{O|I}⟩_O without a minus sign, whereas SM Eq. (S7) correctly defines it as the conditional entropy with a minus sign. As written, Eq. (1) would give IIO = HO + H_{O|I}. Please fix the main-text definition so that h_{O|I} is the conditional entropy and Eq. (1) reads IIO = HO - ⟨h_{O|I}⟩_I.","section":"Eq. (1) and SM Eq. (S7)"},{"comment":"In the conditional Gaussian p_{μ|ν} = N(m_{μ|ν}(x_ν), Σ_ν), the covariance should be that of the conditioned variable μ (e.g., Σ_P for p_{P|I} and Σ_O for p_{O|P}), not Σ_ν. This notation is inconsistent with Eq. (7) and with SM Eqs. (S20) and (S36).","section":"Eq. (5)"},{"comment":"The abstract's \"systematically enhance\" and Sec. II.C's \"always\" should be harmonized; the latter is too strong given the exception stated later in the same section.","section":"Abstract and Sec. II.C"},{"comment":"The caption already contains a clear caveat about the large-σ_OP regime; the main text should refer to this caveat in the same paragraph as the \"always\" sentence, rather than only in the later discussion.","section":"Fig. 2 caption"},{"comment":"The caption uses \"Mp\" with a lowercase p in \"if Mp is small\"; this should be M_P for consistency with the rest of the paper.","section":"SM Fig. S3 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid analytical contribution but the overclaim in Sec. II.C and the missing error bars are the main obstacles. The sign error in Eq. (1) is easy to fix but should not be left in a published version. The reliance on the authors' prior work for the factorization is legitimate, but the verification at strong coupling should be extended to the parameter regimes used in the main figures."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper gives a clean analytical treatment of a three-unit stochastic processing chain and shows that nonlinear integration (combine first, then activate) generally beats nonlinear summation (activate first, then combine) for input-output mutual information. The genuinely new pieces are the explicit Gaussian-average expansion for tanh and the systematic comparison of the two schemes across fast and slow processing, including the dimensionality tradeoff between input and processing layers. The derivation is transparent, the conditional-Gaussian factorization is controlled by timescale ratios, and the authors cross-check the resulting pdf against Langevin simulations in Fig. 1. There are no fitted parameters; the mutual information is evaluated from derived distributions. The SM is thorough and the mathematical claim about the tanh average is machine-checkable in principle.\n\nThe main soft spot is the 'always' claim in Sec. II C. It is false as written: their own Fig. 2i-j and SM Fig. S3 show a strong-coupling, small-MP, large-sigma_OP regime where summation beats integration. They do qualify this a few sentences later, so it is a wording-and-scope problem rather than a deep mathematical flaw, but it is a load-bearing sentence and should be fixed. The second issue is that all numerical comparisons are averaged over 10^3 random-matrix realizations with no error bars or bootstrap. We cannot tell whether the reversal is a robust effect or sampling noise, and similarly whether the claimed integration advantage in other panels is significant. A concrete test would be Langevin simulation in the reversal region with bootstrap confidence intervals. The generalization beyond tanh is also asserted more broadly than demonstrated; the framework is tanh-specific, and the abstract's generality is overstated.\n\nThe timescale-separation factorization is imported from the authors' earlier work and is only spot-checked here at a few parameter points. That is acceptable for a letter, but it means the strong-coupling regime (g=5-10) deserves a few more simulation checks, since the Gaussian factorization is controlled by timescale ratios rather than coupling strength.\n\nWho is this for? People working on stochastic information processing, reservoir computing, and biochemical sensing. It deserves a serious referee: the analytical contribution is real, the question is meaningful, and the flaws are fixable with restricted claims and uncertainty quantification. I would send it out.","headline":"The paper's analytical machinery is solid and the integration-vs-summation comparison is a real contribution, but the 'always' claim in Sec. II C outruns the authors' own numerics and needs rewording plus error bars before the headline result can stand as stated.","tokens_in":26619,"tokens_out":1921,"would_cite":true,"duration_ms":20760,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["82C31","94A17","82C32"],"pacs":["05.40.-a","89.70.Cf"],"model":"deepseek-v4-flash","headline":"This paper claims that in multiscale stochastic processing networks, nonlinear integration (weighted sum first, then a tanh nonlinearity) yields higher input-output mutual information than nonlinear summation (tanh first, then sum), with…","keywords":["mutual information","nonlinear integration","nonlinear summation","timescale separation","stochastic multilayer networks","input-output encoding","optimal dimensionality","bistability"],"falsifier":"Run direct Langevin simulations of the same three-unit model—Gaussian input, $\\tanh$ couplings, summation versus integration—over the full grid of coupling strengths, layer sizes, and processing-to-output weight variances without imposing the factorized solution, and estimate input-output mutual information from long trajectories at each point; the central claim stands only if $I_{IO}^{\\mathrm{int}} > I_{IO}^{\\mathrm{ns}}$ wherever the paper's formulas predict it, especially at strong coupling with small $M_P$ and large $\\sigma_{OP}^2$, where only a few points are currently checked against simulation. A complementary check is to compute the first finite-timescale-ratio corrections to the factorized densities and see whether the sign of the difference survives.","tokens_in":25596,"feed_emoji":"🧠","tokens_out":9765,"duration_ms":86570,"temperature":0.7,"pith_summary":"The paper studies a minimal stochastic processing chain—an input unit, an intermediate processing unit, and an output unit, each containing many degrees of freedom on its own timescale—and asks how much information about the input survives to the output. It compares two wiring schemes: nonlinear summation, where each signal is passed through a $\\tanh$ nonlinearity and then averaged, and nonlinear integration, where the weighted average is formed first and then passed through $\\tanh$. Using a timescale-separated solution for the stationary distribution, the paper shows that integration almost always gives higher input-output mutual information than summation, for both fast and slow processing units, with fast processing making the advantage larger. The reason to care is architectural: if the claim holds, the order of pooling and nonlinear activation is a generic control knob for encoding accuracy, and the best processing-layer size is set by the input dimension rather than being a free parameter.","feed_headline":"Integration before nonlinearity carries more input information","feed_subtitle":"Ordering of pooling and activation sets how much signal survives in biological and artificial processing chains.","key_machinery":"The load-bearing object is the timescale-separated factorization of the joint stationary density into Gaussian conditionals: $p^{\\mathrm{fp}}_{IOP} = p_I p_{P|I} p^{\\mathrm{eff}}_{O|I}$ for fast processing and $p^{\\mathrm{sp}}_{IOP} = p_I p_{P|I} p_{O|P}$ for slow processing, with the input as the slowest unit and each conditional covariance independent of the conditioning variable. The nonlinearity is pushed entirely into the conditional means; for the hyperbolic tangent, the required Gaussian average is evaluated exactly through a convergent series expansion, yielding closed-form means in both summation and integration schemes. With these means, fast-processing mutual information reduces to the output entropy minus a constant conditional-entropy term, while slow-processing mutual information is obtained by sampling the conditional distribution; output bistability is then quantified with Sarle's bimodality coefficient.","core_discovery":"On the paper's own terms, the central result is an ordering: with a slow input and a processing unit that is either faster or slower than the output, the input-output mutual information of nonlinear integration exceeds that of nonlinear summation, $I_{IO}^{\\mathrm{int}} > I_{IO}^{\\mathrm{ns}}$, across a wide range of coupling strengths, layer sizes, and weight distributions. The documented exception is a strong-coupling regime at small processing dimension and large variance of the processing-to-output weights, where summation can win because the few sampled weights push the activation into saturation. The same calculation uncovers two design principles: for a fixed input dimension $M_I$ there is an optimal processing dimension $M_P^*$, with low-dimensional inputs served best by high-dimensional embeddings and high-dimensional inputs by low-dimensional projections; and integration spontaneously produces a bimodal output distribution, read by the authors as input discrimination whose mode balance can be tuned by adding biases to the activation.","pith_inferences":["Editorial extension: the reason integration wins is plausibly that averaging before a saturating nonlinearity keeps fluctuations in the linear regime of $\\tanh$; if so, the same ordering should hold for other sigmoidal activations, which the paper does not test.","Editorial extension: in practical reservoir or neural architectures, the result suggests placing the readout nonlinearity after the weighted sum of internal activities and choosing the hidden-layer width relative to the input dimension; this is a design heuristic the paper does not train or evaluate.","Editorial extension: the tunable output bistability could serve as a primitive classifier, with the active output mode labelling the input; turning the bimodality coefficient into a classification error rate would be a natural next step beyond the paper.","Editorial extension: the derivation assumes exact timescale separation, so whether integration remains superior at order-one timescale ratios is an open question that direct simulation at finite ratios could settle."],"forward_implications":["In multiscale processing chains with a slow input, combining signals before the nonlinear activation is a generic way to raise input-output mutual information relative to the activate-then-combine scheme.","A processing unit that is faster than the output increases information transfer on its own and amplifies the advantage of integration; slower processing reduces but does not erase it.","For a given input dimension there is an optimal processing-layer dimension, and the optimal strategy flips from high-dimensional embedding for small inputs to low-dimensional projection for large inputs.","Integration spontaneously yields a bistable output distribution, and introducing biases into the activation shifts the weight of the two modes, giving a tunable input-discrimination mechanism.","The same ordering and optimal-dimension behaviour carry over to chains with more than one processing unit, since the derivation is independent of the number of processing layers."],"supporting_citations":[{"why":"Supplies the timescale-separated factorization of multilayer stationary distributions from which the fast- and slow-processing joint densities are imported.","marker":"[36]"},{"why":"Provides the order-by-order solution method for Gaussian conditional distributions on multilayer networks used to derive the exact joint pdfs.","marker":"[55]"},{"why":"Supplies the tanh activation and random recurrent network setting that motivates the summation and integration nonlinearities.","marker":"[46]"},{"why":"Mentions reservoir-computing architectures whose summation and integration schemes the paper compares.","marker":"[49]"},{"why":"Entropy estimator used to evaluate output entropy and hence mutual information in the fast-processing case.","marker":"[56]"},{"why":"Nearest-neighbour entropy estimator used for conditional entropy estimates in the slow-processing case.","marker":"[57]"},{"why":"Defines Sarle's bimodality coefficient used to quantify the emergent output bistability.","marker":"[59]"}],"fun_headline_variants":["Nonlinear integration beats summation for info encoding","Order matters: integrate first, then activate","Fast processors boost information in nonlinear nets","Pooling before activation maximizes input information"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison stands or falls on the timescale-separated approximation imported from the authors' earlier work: at stationarity each unit's response to the unit before it is a Gaussian whose variance does not depend on the signal value, so all nonlinearity enters only through the conditional mean; if strong couplings produce corrections that change those variances, the information ordering between integration and summation could change.","fun_headline_variants_meta":{"raw":{"variants":["Nonlinear integration beats summation for info encoding","Order matters: integrate first, then activate","Fast processors boost information in nonlinear nets","Pooling before activation maximizes input information"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1201,"prompt_tokens":889,"completion_tokens":312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":259}},"tokens_in":505,"tokens_out":312,"duration_ms":4134,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:13:42.463152+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run direct Langevin simulations of the same three-unit model—Gaussian input, $\\tanh$ couplings, summation versus integration—over the full grid of coupling strengths, layer sizes, and processing-to-output weight variances without imposing the factorized solution, and estimate input-output mutual information from long trajectories at each point; the central claim stands only if $I_{IO}^{\\mathrm{int}} > I_{IO}^{\\mathrm{ns}}$ wherever the paper's formulas predict it, especially at strong coupling with small $M_P$ and large $\\sigma_{OP}^2$, where only a few points are currently checked against simulation. A complementary check is to compute the first finite-timescale-ratio corrections to the factorized densities and see whether the sign of the difference survives.","supporting_citations":[{"cited_title":"Nicoletti and D","cited_arxiv_id":null,"evidence_quote":"Supplies the timescale-separated factorization of multilayer stationary distributions from which the fast- and slow-processing joint densities are imported."},{"cited_title":"Nicoletti and D","cited_arxiv_id":null,"evidence_quote":"Provides the order-by-order solution method for Gaussian conditional distributions on multilayer networks used to derive the exact joint pdfs."},{"cited_title":"Sompolinsky, A","cited_arxiv_id":null,"evidence_quote":"Supplies the tanh activation and random recurrent network setting that motivates the summation and integration nonlinearities."},{"cited_title":"Vasicek, A test for normality based on sample en- tropy, Journal of the Royal Statistical Society Series B: Statistical Methodology 38, 54 (1976)","cited_arxiv_id":null,"evidence_quote":"Entropy estimator used to evaluate output entropy and hence mutual information in the fast-processing case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Nearest-neighbour entropy estimator used for conditional entropy estimates in the slow-processing case."},{"cited_title":"Multiscale nonlinear integration drives accurate encoding of input information","cited_arxiv_id":null,"evidence_quote":"Defines Sarle's bimodality coefficient used to quantify the emergent output bistability."}],"review_version":1}