Pith. sign in

REVIEW 3 major objections 7 minor 79 references

Branch-local shunting beats matched additive E/I readout only when a reliable divisor kills signal-aligned gain harder than it costs signal or adds denominator noise.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 04:10 UTC pith:OV4DHTDO

load-bearing objection Careful matched-resource theory that kills the “shunting is automatically better” story and replaces it with a falsifiable support–reliability rule; the surrogate is a real soft spot, the V1 transfer is thin, but the package deserves referees. the 3 major comments →

arxiv 2607.24990 v1 pith:OV4DHTDO submitted 2026-07-27 q-bio.NC

When Branch-Local Shunting Helps: A Gain-Load-Alignment Principle for Dendritic E/I Networks

classification q-bio.NC
keywords dendritic computationshunting inhibitiondivisive normalizationE/I balancepopulation codinggain controlmorphologyV1 decoding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Biological neurons can divide excitation by local inhibition on branched dendrites, but it is unclear when that helps population readout more than simply adding and subtracting the same nonnegative inputs. This paper shows that locally, any realizable shunting decision direction already lives inside the positive additive E/I cone, and every scalar shunting threshold has an exact affine additive twin—so shunting is not automatically more expressive. Beyond that local limit, help or harm follows a gain–load–alignment rule: branch-local division wins when a trustworthy divisor sits on the pathway carrying multiplicative nuisance gain and removes that gain more than it attenuates the task signal or imports variability through its denominator. Morphology matters as routing—where a reliable nuisance estimate can meet the affected signal—not as free capacity from depth. Designed hierarchies, frozen-feature normalization, and mouse V1 downstream decoding all show the same conditional pattern and its predicted reversals when support is shuffled, sensors are corrupted, or private noise and readout width change.

Core claim

Neither dendritic depth nor shunting is intrinsically advantageous over a resource-matched additive E/I readout of the same nonnegative inputs. Shunting helps only under a gain–load–alignment principle: a reliable, correctly placed divisor must suppress signal-aligned multiplicative gain more than it reduces baseline signal fidelity or adds residual denominator variability. Locally, shunting cannot open a new first-order decision direction outside the positive additive cone.

What carries the argument

The gain–load–alignment margin A = (B−1) + (B−Mg)Geff − Λ, built from baseline fidelity B, surviving gain factor Mg, aligned gain strength Geff, and residual denominator load Λ. Positive A predicts a fixed-template shunting advantage; morphology and support enter by changing where divisors meet nuisances and how B, Mg, and Λ trade off.

Load-bearing premise

The finite-amplitude story treats performance as approximately separable into signal fidelity, leftover shared gain, and independent denominator noise, so that a simple signed margin predicts which rule wins.

What would settle it

Hold contacts and parameters fixed, place a reliable nuisance sensor on the branch that carries the affected excitation, then shuffle that support or corrupt the sensor: the shunting increment should appear with matched support and reliable sensing, and reverse or vanish under shuffle or corruption, as the paper’s hierarchy and V1 width/noise contrasts predict.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Local normalization in machine learning should help under multiplicative group gain only when a reliable, support-matched divisor acts before aggregation; global or noisy divisors can hurt.
  • Dendritic morphology is best read as a routing graph that decides where nuisance estimates meet signals, not as a generic depth knob for capacity.
  • Downstream decoders with narrow access can show a shunting edge that shrinks or reverses as readout width and private noise grow.
  • Flexible nonlinear predictors can overtake a fixed shunting tree once labels are plentiful, so any shunting edge is an inductive-bias, small-sample effect when support is right.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Circuit mapping should prioritize branch-resolved inhibitory placement relative to gain-modulated excitatory pathways rather than bulk E/I ratios or depth counts.
  • The same support–reliability logic may organize when LayerNorm/GroupNorm-style blocks beat plain additive readouts in non-biological models under structured multiplicative shift.
  • If internal dendritic nonlinearities strongly rotate signal–nuisance geometry, the simple B/Mg/Λ bookkeeping may need Jacobian-level recalibration before it guides experiments.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The manuscript asks when branch-local shunting (divisive) inhibition outperforms a matched additive E/I readout of the same nonnegative inputs. It makes three contributions: (i) exact local boundaries — every scalar shunting threshold is an affine additive comparator (Prop. 1), the terminal shunting Jacobian is contained in the positive additive E/I cone with equality iff an additive-optimal ray admits a positive self-consistent shunting realization (Thm. 1, with a closed-form realizability test, Alg. 2), and passive additive trees flatten exactly to linear readouts (Prop. 3); (ii) a finite-amplitude "gain–load–alignment" criterion (Eqs. 11–16) in which a signed margin A = (B−1) + (B−Mg)Geff − Λ_T predicts the sign of the shunting advantage, validated by split-sample held-out sign prediction and by support-shuffle, load, and sensor-corruption reversals in designed hierarchies and trained DendriNets; (iii) transfer tests to frozen CIFAR features and three mouse V1 sessions, where the decoder gap is largest at narrow readout and reverses under private noise at the widest. The conclusion is explicitly conditional: neither depth nor shunting is intrinsically advantageous.

Significance. If the results hold, this is a useful and unusually disciplined contribution. The local theory is exact rather than asymptotic: the threshold identity is distribution-free, cone containment comes with a constructive self-consistency test verified to numerical precision (Suppl. Fig. S1, max residual < 8×10⁻¹⁶), and the additive flattening is proved. The beyond-local principle is falsifiable — Table 1 lists regime-by-regime predictions — and is stress-tested with sign-reversing interventions (support shuffle, load, sensor noise) rather than only confirmatory fits. The comparator ladder (Table 2) cleanly separates implemented-rule, tangent, cone-capacity, and fitted-linear questions, and resource matching is operationalized (Table 3). Reproducibility is exemplary: pinned commits, per-figure manifest, seed tables, and GPU-hour accounting; the failed Allen five-mouse replication is reported rather than suppressed. The work is primarily a computational/theoretical principle with bounded neural transfer, and is appropriately scoped as such.

major comments (3)
  1. [§2.2, Eqs. (11)–(16); Methods A.2] The central beyond-local claim is carried by Eq. (11)'s factorization d'²_sh ≈ B d'²_0/(1 + Mg Geff + Λ_T) and the additive signed margin in Eq. (13). As the authors themselves note (§A.2), for a general ratio N/Q first-order variability contains Var(δN), r0²Var(δQ), and −2r0Cov(δN,δQ), so gain suppression, fidelity, and load need not be decoupled penalties; the shunting arm of Eq. (16) additionally assumes an isotropic, direction-preserving surrogate with c_T = 1+Λ_T and tree-level alignment equal to the input-level α. Reserving Λ for independent δL is good practice, but the main-text criterion (Eqs. 12–13) and abstract are stated without this restriction, and a tree Jacobian that rotates signal/nuisance geometry voids the α used in Geff. The empirical calibration is supportive but leaves unexplained sign structure: 33/166 held-out sign errors (Fig. 2F, conditioned on the 166/320 realiz
  2. [§2.3, Fig. 4E; Methods A.3.5] The tree-level calibration in Fig. 4E reports G_sh/G_a only as normalized d'²-degradation ratios and states these 'are motivated by, but do not estimate, (B, Mg, Λ_T)' (§A.3.5). Relatedly, for activated trees Eqs. (11)–(14) retain their form 'only after B, Mg, and Λ_T are recomputed from the activated representations' (§2.2), which risks making the criterion descriptive bookkeeping on the realized representation rather than an independently predictive condition. At least one demonstration is needed in which B, Mg, and Λ_T are measured directly (this is straightforward in the passive hierarchy of Fig. 4A–C, where the surrogate's quantities are all estimable), the margin is computed from those measurements alone, and its sign predicts the held-out ordering. Without it, the claim that Eq. (13) is a predictive criterion — as opposed to a consistent decomposition — is not yet shown at tree le
  3. [§2.5, Fig. 7; Methods A.5.4] The qualitative V1 findings — gap largest at narrow readout (Fig. 7B–C) and reversal under private noise at P_E = 64 (Fig. 7E) — rest on three sessions, and the running-state interaction is heterogeneous across those sessions (Fig. 7D); the prespecified Allen five-mouse control did not meet its replication rule (§A.5.4). The text hedges appropriately in most places, but the reversal at the widest readout is a headline prediction of the principle and is presented with mean ± SEM over n = 3. Please report session-level intervals for the sign of the P_E = 64 reversal (i.e., in how many sessions the crossing actually occurs), and soften the abstract's 'reverses under strong private noise at the widest readout' to reflect the session count. More sessions would strengthen this, but clearer quantification is the minimum.
minor comments (7)
  1. [Fig. 2F; Table 7] Fig. 2F: the calibration is conditioned on positive realizations (154/320 attempts had none). Please state the selection rule and whether the excluded libraries differ systematically in alignment or operating point, since this bears on how general the 133/166 sign agreement is.
  2. [§2.1–2.2; Table 4] Notation: T is used both for the branch tree (topology) and for total conductance T_n; the collision is flagged in Table 4 but still forces careful reading in Eqs. (4)–(14). Consider T vs. T_n or a tree symbol such as 𝒯 throughout.
  3. [§A.4.1; Table 5] The implementation term 'reactivation' for post-voltage activation is confusing against its biological meaning; the reader-facing rename is acknowledged in §A.4.1 but the code-facing term still appears in main-text-adjacent passages. Unify terminology.
  4. [§2.2 vs. Methods A.2] The caveat that α is defined at input level while tree Jacobians can rotate signal–nuisance geometry appears only in Methods A.2 ('Generic tree-induced rotations... require numerical reoptimization'). Given its importance (major comment 1), surface it in §2.2 near the definition of Geff.
  5. [§2.3, Fig. 4D] Fig. 4D's depth sweep 'does not match parameters or contacts across depth' and is relegated to Suppl. Fig. S5 for diagnostics; state this limitation in the main-text sentence citing the panel so the depth-reversal claim is not over-read.
  6. [§2.5; Methods A.3.1] The d'²-versus-information disclaimers are well handled (ILB treated as a log-loss difference, BIAWGN mapping only under Gaussian assumptions). It would help to add one sentence in §2.5 noting explicitly that ΔILB = paired log-loss gap in bits, since readers may otherwise read Fig. 7B–C as information orderings.
  7. [§3 Discussion] Given the debate cited (Holt & Koch 1997; Prescott & De Koninck 2003) on whether shunting is truly divisive at the soma, a short Discussion paragraph on which physiological regimes (synaptic noise, saturation) map onto the model's operating-point assumptions would help neurophysiology readers place the scope.

Circularity Check

1 steps flagged

No load-bearing circularity: local theorems are algebraic identities, and the gain–load criterion is a stated surrogate tested by held-out composition and sign-reversing interventions.

specific steps
  1. fitted input called prediction [§2.2 Eqs. (11)–(16); Methods A.2 fixed-template / direction-preserving surrogate]
    "d′2_sh ≈ B(T, r)d′2_0 / (1 + Mg(T, r)Geff + ΛT(r)). ... The shunting identity additionally assumes an isotropic, direction-preserving surrogate with cT = 1 + ΛT. ... under the direction-preserving surrogate, define the surviving gain factor Mg(T, r) ... by Varg(w⊤_sh hsh)/Varind(w⊤_sh hsh) = α Mg(T, r)Γ"

    Mild only: when the surrogate assumptions hold exactly, Mg and Λ are read off the same shunting representation’s variance decomposition that enters the predicted d′2, so the signed margin A is partly rearranging that decomposition rather than an independent dynamical law. This is not a full circularity because the paper treats it as a leading-order surrogate, estimates components on calibration splits, and tests combined-nuisance signs and mechanism order on held-out data (not the same cells used to define the terms), with imperfect agreement (e.g., 133/166; 56/63) showing the claim is falsifiable rather than true by definition.

full rationale

The derivation chain does not reduce the central claims to their inputs by construction. Proposition 1 is an exact threshold equivalence (Vsh>τ ⟺ E−ρτI>bτ) from algebra on the scalar shunt. Theorem 1 is local Jacobian containment in the positive additive cone, with equality only under an independently checkable self-consistent realization—not an assumption that shunting helps. The beyond-local gain–load–alignment margin (Eqs. 11–14) is derived under an explicit direction-preserving / rank-one surrogate (Methods A.2; Eq. 16’s cT=1+ΛT); B, Mg, and Λ are intermediate operational quantities, and the paper tests their composition on held-out combined corruption (Fig. 2F split-sample surrogate; Fig. 4E clean-fit / calibrate / fresh combined-test; Fig. 5D delta-method vs Monte Carlo) and predicts reversals under load, support shuffle, and sensor noise rather than refitting the conclusion. The “designed hierarchy” is an existence construction with negative controls (shuffle and noise reverse the linear comparisons), not a tautology. The only related-work self-citation ([53], same architecture for credit assignment) is explicitly scoped out of the forward-readout claims and is not used as a uniqueness or forcing premise. Residual risk is assumption validity (surrogate separability), not circular reduction.

Axiom & Free-Parameter Ledger

6 free parameters · 8 axioms · 3 invented entities

The central conditional claim rests on passive conductance compartmental abstractions, nonnegative E/I pathway coding, matched-resource comparators, and a separable gain/load surrogate beyond the local linearization. Theory contributions are mostly standard math plus domain biophysics; the main modeling risk is treating the surrogate margin and designed support alignment as representative of trained/biological pathways. Empirical transfer adds small-n session assumptions without identifying circuits.

free parameters (6)
  • Hierarchical construction gains and sensors (μ, Δ, σ_pr, s, g, s_g, sensor log-SD) = e.g. μ=1, Δ=0.12, σ_pr=0.20, s=12, g=0.4; s_g grid 0–1
    Exact-inventory tree (Fig. 4) uses hand-chosen drive, private noise, axial g, and sensor conductances so aligned deep shunting can beat linear controls; pilots on seeds 0–1 precede held-out seeds.
  • Gain–load Monte Carlo operating point (Ē, Ī, L̄, ΔE, ΔI, σ_pr) = (Ē,Ī,L̄)=(3,1,0.5), (ΔE,ΔI)=(0.6,-0.2), σ_pr=0.25
    Single-branch regime map (Fig. 5D) fixes baseline drives and class shifts that set where the delta-method boundary lies.
  • Trained-network contact caps and widths (N_E, N_I, P_E, branch factors) = V1 example N_E/N_I=40/10; contextual N_E∈{16,32}, N_I∈{4,12}; trees [8],[2,3],[2,1,2], etc.
    Contextual, ordered-E/I, and V1 DendriNet runs depend on chosen Top-K budgets and tree shapes; equal-budget selection still leaves architecture as experimenter choice.
  • Shunting leak/input scale r in fixed-template V1 diagnostics = grid r∈{0.01,0.1,1,10,100}; often prefers larger r
    Fixed shunt V=E/(r+E+I) is only identifiable up to response units; leave-one-session selection among {0.01…100} can move the operating regime.
  • CIFAR group-divisive stabilizer and group partition = ε=0.05; 8 groups; 512×2×2 features
    Frozen ResNet bridge uses eight contiguous channel groups and hg/(0.05+mean), choices that define ‘aligned’ divisor support.
  • Validation-selected optimizers/initializers (LR, activation policy, epoch ceilings) = V1 selected analytical init, branch LR 3e-4, decoder 2e-3
    Complete-model comparisons freeze setups chosen on validation (analytical vs occupancy init, LRs, long V1 horizons up to 1600 epochs); not unique physical constants.
axioms (8)
  • domain assumption Steady-state passive conductance balance with pure shunting (E_I=E_L) yields the unit-leak node ratio V=N/(1+T); active NMDA/Ca spikes, strong backpropagation, and fast temporal dynamics are omitted.
    Methods A.1.1–A.1.2; scope section states high-conductance quasi-static regime.
  • domain assumption Upstream population activity follows X=g f(θ)+ε with shared gain g (E[g]=1) and independent residual covariance, so harmful variability is low-rank and potentially signal-aligned.
    Eqs. 1–2; standard information-limiting correlation setup.
  • domain assumption Strict conductance readouts use nonnegative pathway activities and nonnegative synaptic weights; signed LDA is diagnostic only.
    §2.1 and Methods A.1.3–A.1.4; cone statements live in positive E/I coordinates.
  • standard math First-order linearization at a positive self-consistent operating point is valid for local d′² comparisons; curvature remainders are higher-order (delta-method bound).
    Theorem 1, Methods A.1.6–A.1.8.
  • ad hoc to paper Beyond-local shunting advantage is summarized by direction-preserving surrogates where residual load rescales independent variance by c_T=1+Λ_T and surviving gain is Mg, with gain and load approximately separable in the denominator.
    Eqs. 11–16 and Methods A.2; authors note generic rotations need numerical reoptimization.
  • standard math Passive additive trees with nonnegative transfers flatten exactly to linear somatic scores; topology changes path gains but not function class once weights are free.
    Proposition 3 / Eq. 17.
  • ad hoc to paper Matched comparisons (shared samples, morphology, contact/parameter budgets, or equal validation search) isolate integration rule from resource confounds.
    Table 3 and throughout Results; defines the paper’s notion of ‘fair’ mechanism tests.
  • domain assumption V1 analyses test downstream decoder implications of shared variability/state, not biological identification of shunting circuits or cell types.
    §2.5 and Discussion; E/I paths are computational evidence assignments.
invented entities (3)
  • DendriNet independent evidence
    purpose: Trainable branched E/I population module that swaps additive vs shunting rules, morphology, Top-K allocation, divisor locality, and optional post-voltage activation under matched resources.
    Core methods contribution; implementation abstraction rather than a new physical particle/force. Independent use appears via public code and a related credit-assignment preprint.
  • Gain–load–alignment margin A(T,r) with factors B, M_g, Λ_T independent evidence
    purpose: Scalar bookkeeping criterion predicting when branch-local shunting beats additive readouts and when depth/support changes help or hurt.
    Paper-defined operational entities estimated from surrogates or simulations; falsifiable via predicted reversals (shuffle, load, sensor noise) rather than only in-sample fit.
  • Denominator load L / residual penalty Λ no independent evidence
    purpose: Name non-target pool activity that shifts operating point and injects divisor noise competing with gain suppression.
    Conceptual coordinate linking background conductance, masks, and distractors to the criterion; not a newly observed biophysical quantity uniquely measured in V1.

pith-pipeline@v1.2.0-grok45-kimik3 · 59283 in / 5152 out tokens · 87635 ms · 2026-07-31T04:10:52.741601+00:00 · methodology

0 comments
read the original abstract

Biological neurons combine excitatory and inhibitory (E/I) activity on branched dendrites through shunting, in which inhibition divisively attenuates excitation. Whether this improves population readout over additive E/I integration of the same nonnegative inputs remains unclear. We introduce DendriNet, a trainable framework that varies integration rule, morphology, synaptic allocation, divisor locality, and dendritic nonlinearities. For population codes with multiplicative gain, a local linearization of any realizable shunting readout yields a decision direction within the positive additive E/I cone; matching the additive optimum requires a positive self-consistent shunting realization. Every scalar shunting threshold also has an exact affine additive realization. Beyond this local limit, performance follows a gain-load-alignment principle: branch-local shunting helps when a reliable divisor suppresses signal-aligned gain more than it attenuates signal or adds denominator variability. Passive additive trees flatten to linear readouts, whereas shunting trees compose local divisors. In a designed hierarchy, deep shunting outperforms tangent and fitted-linear controls, but flexible nonlinear predictors overtake it with enough labels. Support shuffling reverses the linear comparisons, sensor corruption reverses the fitted-linear comparison, and resource-matched activated training shows no consistent depth benefit. The same support and reliability interaction appears in frozen-feature normalization. Across three mouse V1 sessions, the shunting-over-additive decoder gap is largest for narrow readouts, reverses under strong private noise at the widest readout, and varies across running states. Morphology can determine where reliable nuisance estimates meet task-relevant signals, but neither depth nor shunting is intrinsically advantageous.

Figures

Figures reproduced from arXiv: 2607.24990 by Bernardo L. Sabatini, Houman Safaai, Maceo Richards, Naeem Khoshnevis.

Figure 1
Figure 1. Figure 1: DendriNet and resource controls. (A) One unit combines nonnegative E/I drive additively or through a local shunting divisor. (B) Branched units form a population decoder. (C) Passive additive log-gain elasticity is 1; shunting elasticity is [1 + gT0] −1 . (D) A continuous mass constraint M = ∥w∥1 maps the local mechanism gap across total mass and inhibitory share (360 generated libraries). (E–F) A separate… view at source ↗
Figure 2
Figure 2. Figure 2: Gain suppression competes with load and signal loss. (A–B) Shared gain reaches numerator and divisor; independent pool noise reaches only the divisor. At B = α = 1 and Mg = 0.35, the boundary is (1 − Mg)Γ = Λ. (C) Across α ∈ [0, 1] at Γ = 1.1, reoptimization (solid) shifts the fixed-template crossing (dashed). (D) Conductance and gain change the mechanism ordering on identical E/I observations. (E) The nea… view at source ↗
Figure 3
Figure 3. Figure 3: Divisor support must match the nuisance. (A) Global, matched, and fragmented divisors repartition the same nuisance-bearing inputs. (B) Matched-minus-shuffled linear accessibility across 96 computed oracle conditions; the benefit peaks on the diagonal, where divisor granularity equals nuisance rank. (C–D) Matching inhibitory support preferentially improves shunting at zero I-stream background (I-bg), witho… view at source ↗
Figure 4
Figure 4. Figure 4: Aligned hierarchical routing can suppress shared gain. (A) The same eight observations are routed through flat [8], shallow [2, 3], and deep [2, 1, 2] trees; purple/blue/green mark fine/coarse/global inputs. (B) Only the deep shunting composition crosses above its tangent and the fitted-linear decoder. (C) Paired contrasts show the aligned benefit and its support/noise reversals. (D) A controlled sweep sho… view at source ↗
Figure 5
Figure 5. Figure 5: Training gives conditional effects and a load reversal. (A) Cue c selects target yc from two streams under gain g; below, a [2, 2] DendriNet predicts yˆ. (B) Selected shunting and additive accuracies at the prespecified alignment endpoints α = 0, 1 show the absolute contextual-task effect. (C) Paired complete-model and matched-operating-point gaps distinguish trained performance from rule-isolating control… view at source ↗
Figure 6
Figure 6. Figure 6: The gain–load interaction transfers to conventional feature normalization. A frozen ImageNet-pretrained ResNet-18 supplies CIFAR-10 features; all trained arms use the full 45,000/5,000 training/validation split and the same clean-validation search budget. (A) Aligned group gain separates the mean-only divisive readout from an unnormalized additive readout, while GroupNorm is comparably robust. (B) Support … view at source ↗
Figure 7
Figure 7. Figure 7: V1 decoder effects depend on readout width. (A) Training-selected V1 pools feed matched additive and shunting dendritic readouts (schematic). (B–C) Absolute bounds stay high, but their gap is largest at PE = 8. (D) The running-state difference-in-differences (DiD) varies by session. (E) Private noise favors shunting at PE = 8, 16 but reverses at 64. (F) A training-estimated multiplicative correction reduce… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

79 extracted references · 37 canonical work pages

  1. [1]

    D. J. Heeger. Normalization of cell responses in cat striate cortex.Visual Neuroscience, 9(2): 181–197, 1992. doi: 10.1017/S0952523800009640

  2. [2]

    J. H. Reynolds and D. J. Heeger. The normalization model of attention.Neuron, 61(2):168–185,

  3. [3]

    Carandini and D

    M. Carandini and D. J. Heeger. Normalization as a canonical neural computation.Nature Reviews Neuroscience, 13(1):51–62, 2012. doi: 10.1038/nrn3136

  4. [4]

    F. S. Chance, L. F. Abbott, and A. D. Reyes. Gain modulation from background synaptic input.Neuron, 35(4):773–782, 2002. doi: 10.1016/S0896-6273(02)00820-6

  5. [5]

    S. J. Mitchell and R. A. Silver. Shunting inhibition modulates neuronal gain during synaptic excitation.Neuron, 38(3):433–445, 2003. doi: 10.1016/S0896-6273(03)00200-9

  6. [6]

    London and M

    M. London and M. Häusser. Dendritic computation.Annual Review of Neuroscience, 28: 503–532, 2005. doi: 10.1146/annurev.neuro.28.061604.135703

  7. [7]

    R. A. Silver. Neuronal arithmetic.Nature Reviews Neuroscience, 11(7):474–489, 2010. doi: 10.1038/nrn2864

  8. [8]

    G. R. Holt and C. Koch. Shunting inhibition does not have a divisive effect on firing rates. Neural Computation, 9(5):1001–1013, 1997. doi: 10.1162/neco.1997.9.5.1001

  9. [9]

    Prescott and Yves De Koninck

    Steven A. Prescott and Yves De Koninck. Gain control of firing rate by shunting inhibition: roles of synaptic noise and dendritic saturation.Proceedings of the National Academy of Sciences, 100(4):2076–2081, 2003. doi: 10.1073/pnas.0337591100

  10. [10]

    Panayiota Poirazi, Terrence Brannon, and Bartlett W. Mel. Pyramidal neuron as two-layer neural network.Neuron, 37(6):989–999, 2003. doi: 10.1016/S0896-6273(03)00149-1

  11. [11]

    Polsky, B

    A. Polsky, B. W. Mel, and J. Schiller. Computational subunits in thin dendrites of pyramidal cells.Nature Neuroscience, 7(6):621–627, 2004. doi: 10.1038/nn1253

  12. [12]

    Beniaguev, I

    D. Beniaguev, I. Segev, and M. London. Single cortical neurons as deep artificial neural networks. Neuron, 109(17):2727–2739.e3, 2021. doi: 10.1016/j.neuron.2021.07.002

  13. [13]

    Chavlis and P

    S. Chavlis and P. Poirazi. Dendrites endow artificial neural networks with accurate, ro- bust and parameter-efficient learning.Nature Communications, 16:943, 2025. doi: 10.1038/ s41467-025-56297-9

  14. [14]

    Benjamin S. H. Lyo and Cristina Savin. Complex priors and flexible inference in recurrent circuits with dendritic nonlinearities. InInternational Conference on Learning Representations, 2024. 14

  15. [15]

    Anamika Agrawal and Michael A. Buice. Bounds on the computational complexity of neu- rons due to dendritic morphology. InAdvances in Neural Information Processing Systems, volume 38, 2025. URLhttps://proceedings.neurips.cc/paper_files/paper/2025/hash/ d9e56606b1bd3095fbb4241c4e15e420-Abstract-Conference.html

  16. [16]

    Principles governing the operation of synaptic inhibition in dendrites.Neuron, 75(2):330–341, 2012

    Albert Gidon and Idan Segev. Principles governing the operation of synaptic inhibition in dendrites.Neuron, 75(2):330–341, 2012. doi: 10.1016/j.neuron.2012.05.015

  17. [17]

    Sarah Jarvis, Konstantin Nikolic, and Simon R. Schultz. Neuronal gain modulability is deter- mined by dendritic morphology: A computational optogenetic study.PLOS Computational Biology, 14(3):e1006027, 2018. doi: 10.1371/journal.pcbi.1006027

  18. [18]

    Dynamic gain decomposition reveals functional effects of dendrites, ion channels, and input statistics in population coding.Journal of Neuroscience, 44(13):e0799232023, 2024

    Chenfei Zhang, Omer Revah, Fred Wolf, and Andreas Neef. Dynamic gain decomposition reveals functional effects of dendrites, ion channels, and input statistics in population coding.Journal of Neuroscience, 44(13):e0799232023, 2024. doi: 10.1523/JNEUROSCI.0799-23.2023

  19. [19]

    Caleb C. A. Stokes, Corinne M. Teeter, and Jeffry S. Isaacson. Single dendrite-targeting interneurons generate branch-specific inhibition.Frontiers in Neural Circuits, 8:139, 2014. doi: 10.3389/fncir.2014.00139

  20. [20]

    Bloss, Mark S

    Erik B. Bloss, Mark S. Cembrowski, Bill Karsh, Jennifer Colonell, Richard D. Fetter, and Nelson Spruston. Structured dendritic inhibition supports branch-selective integration in CA1 pyramidal cells.Neuron, 89(5):1016–1030, 2016. doi: 10.1016/j.neuron.2016.01.029

  21. [21]

    Murray, and Xiao-Jing Wang

    Guangyu Robert Yang, John D. Murray, and Xiao-Jing Wang. A dendritic disinhibitory circuit mechanism for pathway-specific gating.Nature Communications, 7:12815, 2016. doi: 10.1038/ncomms12815

  22. [22]

    Niell and Michael P

    Cristopher M. Niell and Michael P. Stryker. Modulation of visual responses by behavioral state in mouse visual cortex.Neuron, 65(4):472–479, 2010. doi: 10.1016/j.neuron.2010.01.033

  23. [23]

    McGinley, Stephen V

    Matthew J. McGinley, Stephen V. David, and David A. McCormick. Cortical membrane potential signature of optimal states for sensory signal detection.Neuron, 87(1):179–192, 2015. doi: 10.1016/j.neuron.2015.05.038

  24. [24]

    Robbe L. T. Goris, J. Anthony Movshon, and Eero P. Simoncelli. Partitioning neuronal variability.Nature Neuroscience, 17(6):858–865, 2014. doi: 10.1038/nn.3711

  25. [25]

    Rabinowitz, Robbe L

    Neil C. Rabinowitz, Robbe L. Goris, Marlene Cohen, and Eero P. Simoncelli. Attention stabilizes the shared gain of V4 populations.eLife, 4:e08998, 2015. doi: 10.7554/eLife.08998

  26. [26]

    B. B. Averbeck, P. E. Latham, and A. Pouget. Neural correlations, population coding and computation.Nature Reviews Neuroscience, 7(5):358–366, 2006. doi: 10.1038/nrn1888

  27. [27]

    Cohen and Adam Kohn

    Marlene R. Cohen and Adam Kohn. Measuring and interpreting neuronal correlations.Nature Neuroscience, 14(7):811–819, 2011. doi: 10.1038/nn.2842

  28. [28]

    Moreno-Bote, J

    R. Moreno-Bote, J. Beck, I. Kanitscheider, X. Pitkow, P. E. Latham, and A. Pouget. Information- limiting correlations.Nature Neuroscience, 17(10):1410–1417, 2014. doi: 10.1038/nn.3807

  29. [29]

    Jaffe, Selmaan N

    MohammadMehdi Kafashan, Anna W. Jaffe, Selmaan N. Chettih, Ramon Nogueira, Iñigo Arandia-Romero, Christopher D. Harvey, Rubén Moreno-Bote, and Jan Drugowitsch. Scaling of sensory information in large neural populations shows signatures of information-limiting correlations.Nature Communications, 12(1):473, 2021. doi: 10.1038/s41467-020-20722-y. 15

  30. [30]

    Sophie Denève and Christian K. Machens. Efficient codes and balanced networks.Nature Neuroscience, 19(3):375–382, 2016. doi: 10.1038/nn.4243

  31. [31]

    Fellous, M

    J.-M. Fellous, M. Rudolph, A. Destexhe, and T. J. Sejnowski. Synaptic background noise controls the input/output characteristics of single cells in an in vitro model of in vivo activity. Neuroscience, 122(3):811–829, 2003. doi: 10.1016/j.neuroscience.2003.08.027

  32. [32]

    Carandini, D

    M. Carandini, D. J. Heeger, and J. A. Movshon. Linearity and normalization in simple cells of the macaque primary visual cortex.Journal of Neuroscience, 17(21):8621–8644, 1997. doi: 10.1523/JNEUROSCI.17-21-08621.1997

  33. [33]

    Busse, A

    L. Busse, A. R. Wade, and M. Carandini. Representation of concurrent stimuli by population activity in visual cortex.Neuron, 64(6):931–942, 2009. doi: 10.1016/j.neuron.2009.11.004

  34. [34]

    Coen-Cagli and S

    R. Coen-Cagli and S. S. Solomon. Relating divisive normalization to neuronal response variability. Journal of Neuroscience, 39(37):7344–7356, 2019. doi: 10.1523/JNEUROSCI.0126-19.2019

  35. [35]

    Simoncelli

    Odelia Schwartz and Eero P. Simoncelli. Natural signal statistics and sensory gain control. Nature Neuroscience, 4(8):819–825, 2001. doi: 10.1038/90526

  36. [36]

    Simoncelli

    Johannes Ballé, Valero Laparra, and Eero P. Simoncelli. Density modeling of images using a gen- eralized normalization transformation. InInternational Conference on Learning Representations, 2016

  37. [37]

    Bryan P. Tripp. Decorrelation of spiking variability and improved information transfer through feedforward divisive normalization.Neural Computation, 24(4):867–894, 2012. doi: 10.1162/ NECO_a_00255

  38. [38]

    Bounds, Hillel Adesnik, and Ruben Coen-Cagli

    Oren Weiss, Hayley A. Bounds, Hillel Adesnik, and Ruben Coen-Cagli. Modeling the diverse effects of divisive normalization on noise correlations.PLOS Computational Biology, 19(11): e1011667, 2023. doi: 10.1371/journal.pcbi.1011667

  39. [39]

    Burg, Santiago A

    Max F. Burg, Santiago A. Cadena, George H. Denfield, Edgar Y. Walker, Andreas S. Tolias, Matthias Bethge, and Alexander S. Ecker. Learning divisive normalization in primary visual cortex.PLOS Computational Biology, 17(6):e1009028, 2021. doi: 10.1371/journal.pcbi.1009028

  40. [40]

    Feature-specific divisive normalization improves natural image encoding for depth perception.bioRxiv, 2024

    Long Ni and Johannes Burge. Feature-specific divisive normalization improves natural image encoding for depth perception.bioRxiv, 2024. doi: 10.1101/2024.09.05.611536

  41. [41]

    Jakob Jordan, João Sacramento, Willem A. M. Wybo, Mihai A. Petrovici, and Walter Senn. Conductance-based dendrites perform Bayes-optimal cue integration.PLOS Computational Biology, 20(6):e1012047, 2024. doi: 10.1371/journal.pcbi.1012047

  42. [42]

    Beyond divisive normalization: Scalable feedforward networks for multisensory integration across reference frames.Journal of Neuroscience, 45(41):e0104252025, 2025

    Arefeh Farahmandi, Parisa Abedi Khoozani, and Gunnar Blohm. Beyond divisive normalization: Scalable feedforward networks for multisensory integration across reference frames.Journal of Neuroscience, 45(41):e0104252025, 2025. doi: 10.1523/JNEUROSCI.0104-25.2025

  43. [43]

    Wakhloo, Will Slatton, and SueYeon Chung

    Albert J. Wakhloo, Will Slatton, and SueYeon Chung. Neural population geometry and optimal coding of tasks with shared latent structure.Nature Neuroscience, 29:682–692, 2026. doi: 10.1038/s41593-025-02183-y

  44. [44]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. doi: 10.1109/CVPR.2016.90. 16

  45. [45]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky. Learning multiple layers of features from tiny images. Techni- cal report, University of Toronto, 2009. URL https://www.cs.toronto.edu/~kriz/ learning-features-2009-TR.pdf

  46. [46]

    Group normalization

    Yuxin Wu and Kaiming He. Group normalization. InProceedings of the European Conference on Computer Vision, pages 3–19, 2018. doi: 10.1007/978-3-030-01261-8_1

  47. [47]

    Stringer, M

    C. Stringer, M. Pachitariu, N. Steinmetz, M. Carandini, and K. D. Harris. High-dimensional geometry of population responses in visual cortex.Nature, 571(7765):361–365, 2019. doi: 10.1038/s41586-019-1346-5

  48. [48]

    Recordings of ~20,000 neurons from V1 in response to oriented stimuli

    Marius Pachitariu, Michalis Michaelos, and Carsen Stringer. Recordings of ~20,000 neurons from V1 in response to oriented stimuli. Janelia Research Campus Figshare, 2019. URL https://doi.org/10.25378/janelia.8279387.v3. Version 3; CC BY-NC 4.0

  49. [49]

    On variational bounds of mutual information

    Ben Poole, Sherjil Ozair, Aaron van den Oord, Alex Alemi, and George Tucker. On variational bounds of mutual information. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 5171–5180. PMLR,

  50. [50]

    M. E. Larkum. A cellular mechanism for cortical associations: an organizing principle for the cerebral cortex.Trends in Neurosciences, 36(3):141–151, 2013. doi: 10.1016/j.tins.2012.11.006

  51. [51]

    Dendritic integration: 60 years of progress.Nature Neuro- science, 18(12):1713–1721, 2015

    Greg Stuart and Nelson Spruston. Dendritic integration: 60 years of progress.Nature Neuro- science, 18(12):1713–1721, 2015. doi: 10.1038/nn.4157

  52. [52]

    Zolnik, Paweł Fidzinski, Felix Bolduan, Athanasia Papoutsi, Panayiota Poirazi, Martin Holtkamp, Imre Vida, and Matthew E

    Albert Gidon, Timothy A. Zolnik, Paweł Fidzinski, Felix Bolduan, Athanasia Papoutsi, Panayiota Poirazi, Martin Holtkamp, Imre Vida, and Matthew E. Larkum. Dendritic ac- tion potentials and computation in human layer 2/3 cortical neurons.Science, 367(6473):83–87,

  53. [53]

    Sabatini

    Houman Safaai, Maceo Richards, and Bernardo L. Sabatini. Shunting inhibition and dendritic branching shape local credit assignment.arXiv preprint arXiv:2607.03556, 2026. URLhttps: //arxiv.org/abs/2607.03556

  54. [54]

    Destexhe and D

    A. Destexhe and D. Paré. Impact of network activity on the integrative properties of neocortical pyramidal neurons in vivo.Journal of Neurophysiology, 81(4):1531–1547, 1999. doi: 10.1152/jn. 1999.81.4.1531

  55. [55]

    Ozeki, I

    H. Ozeki, I. M. Finn, E. S. Schaffer, K. D. Miller, and D. Ferster. Inhibitory stabilization of the cortical network underlies visual surround suppression.Neuron, 62(4):578–592, 2009. doi: 10.1016/j.neuron.2009.03.028

  56. [56]

    W. Rall. Theory of physiological properties of dendrites.Annals of the New York Academy of Sciences, 96(4):1071–1092, 1962. doi: 10.1111/j.1749-6632.1962.tb54120.x

  57. [57]

    Koch.Biophysics of Computation: Information Processing in Single Neurons

    C. Koch.Biophysics of Computation: Information Processing in Single Neurons. Oxford University Press, 1999. doi: 10.1093/oso/9780195104912.001.0001

  58. [58]

    Destexhe, M

    A. Destexhe, M. Rudolph, and D. Paré. The high-conductance state of neocortical neurons in vivo.Nature Reviews Neuroscience, 4(9):739–751, 2003. doi: 10.1038/nrn1198. 17

  59. [59]

    J. C. Magee and E. P. Cook. Somatic epsp amplitude is independent of synapse location in hippocampal pyramidal neurons.Nature Neuroscience, 3(9):895–903, 2000. doi: 10.1038/78800

  60. [60]

    S. R. Williams and G. J. Stuart. Dependence of epsp efficacy on synapse location in neocortical pyramidal neurons.Science, 295(5561):1907–1910, 2002. doi: 10.1126/science.1067903

  61. [61]

    Tran-Van-Minh, R

    A. Tran-Van-Minh, R. D. Cazé, T. Abrahamsson, L. Cathala, B. S. Gutkin, and D. A. DiGregorio. Contribution of sublinear and supralinear dendritic integration to neuronal computations. Frontiers in Cellular Neuroscience, 9:67, 2015. doi: 10.3389/fncel.2015.00067

  62. [62]

    R. D. Cazé, M. Humphries, and B. Gutkin. Passive dendrites enable single neurons to compute linearly non-separable functions.PLOS Computational Biology, 9(2):e1002867, 2013. doi: 10.1371/journal.pcbi.1002867

  63. [63]

    Payeur, J.-C

    A. Payeur, J.-C. Béïque, and R. Naud. Classes of dendritic information processing.Current Opinion in Neurobiology, 58:78–85, 2019. doi: 10.1016/j.conb.2019.07.006

  64. [64]

    What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009

    Kevin Jarrett, Koray Kavukcuoglu, Marc’Aurelio Ranzato, and Yann LeCun. What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009. doi: 10.1109/ICCV.2009.5459469

  65. [65]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural networks. InAdvances in Neural Information Processing Systems, volume 25, pages 1097–1105, 2012

  66. [66]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InProceedings of the 32nd International Conference on Machine Learning, volume 37 ofProceedings of Machine Learning Research, pages 448–456,

  67. [67]

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization.arXiv preprint arXiv:1607.06450, 2016. doi: 10.48550/arXiv.1607.06450

  68. [68]

    Saskia E. J. de Vries, Jerome A. Lecoq, Michael A. Buice, et al. A large-scale standardized phys- iological survey reveals functional organization of the mouse visual cortex.Nature Neuroscience, 23:138–151, 2020. doi: 10.1038/s41593-019-0550-9

  69. [69]

    Guerguiev, T

    J. Guerguiev, T. P. Lillicrap, and B. A. Richards. Towards deep learning with segregated dendrites.eLife, 6:e22901, 2017. doi: 10.7554/eLife.22901

  70. [70]

    Sacramento, R

    J. Sacramento, R. Ponte Costa, Y. Bengio, and W. Senn. Dendritic cortical microcircuits approximate the backpropagation algorithm. InAdvances in Neural Information Processing Systems, volume 31, pages 8721–8732, 2018

  71. [71]

    tangent-matched

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. doi: 10.1109/CVPR.2009.5206848. 18 Appendix Overview This appendix provides the theoretical derivations, experimental methods, and additional ana...

  72. [76]

    [2,3] [2,1,2] same branch-unit, contact, and parameter budgets A Exact-resource redistribution

  73. [77]

    [2,3] [2,1,2] Exact-resource morphology −0.4 −0.2 0.0 Test accuracy change from [8] (pp)B Valid policies · I-bg 0

  74. [78]

    [2,3] [2,1,2] Exact-resource morphology −2 −1 0 Test accuracy change from [8] (pp)C Depth effect · I-bg 3

  75. [79]

    [2,3] [2,1,2] Exact-resource morphology −1 0 1 Shunting − additive accuracy (pp)D Valid same-policy gap Shunting · analytical Additive · analytical Additive · occupancy Shunting · occupancy Analytical · I-bg 0 Analytical · I-bg 3 Occupancy · I-bg 3 Supplementary Fig. S10:Redistributing a fixed branch budget does not produce a generic depth benefit after s...

  76. [2009]

    doi: 10.1016/j.neuron.2009.01.002

  77. [2015]

    URLhttps://proceedings.mlr.press/v37/ioffe15.html

  78. [2019]

    URLhttps://proceedings.mlr.press/v97/poole19a.html

  79. [2020]

    doi: 10.1126/science.aax6239