Pith. sign in

REVIEW 2 major objections 4 minor 90 references

A sparse-coding model of V1 trained by denoising score matching is equivalent to a minimal diffusion model, with a denoising Jacobian that decomposes into sparse-gated recurrent spreading along learned lateral connections.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 22:34 UTC pith:HYBT4HLT

load-bearing objection Empirically interesting, but the central Jacobian expansion in Eq. 8 is algebraically wrong; the mechanistic story does not hold as written. the 2 major comments →

arxiv 2607.15693 v1 pith:HYBT4HLT submitted 2026-07-17 q-bio.NC cs.AI

Toward a mechanistic understanding of inference in visual cortex and diffusion models

classification q-bio.NC cs.AI
keywords sparse codingdenoising score matchingnon-factorial priordiffusion modelsV1 horizontal connectionscontour completionJacobian decompositiongeneralization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to show that a recurrent sparse-coding model of primary visual cortex, trained with a denoising score-matching objective, is functionally a minimal diffusion model whose internal mechanics can be read directly from its parameters. The key move is adding a learned pairwise interaction matrix to the sparse-coding prior, so inference becomes a recurrent circuit. After training on natural images, the learned interactions resemble V1 horizontal connections, and the model denoises nearly as well as standard black-box diffusion models. By decomposing the network's Jacobian, the paper claims to explain mechanistically how contour completion arises and how diffusion models assign high probability to continuous natural deformations. If correct, this bridges neuroscience and machine learning: it turns a black-box generative model into an interpretable circuit and makes concrete, testable predictions about V1 connectivity.

Core claim

The paper's central claim is that a non-factorial sparse-coding model, with energy E = (1/2σ²)‖x − Φz‖² + γψ(σ)(‖λ∘z‖₁ + ½ zᵀMz), trained by denoising score matching, is equivalent to a minimal diffusion model. The learned interaction matrix M develops collinear facilitation that mirrors V1 horizontal connections. The denoising Jacobian J(x) = Φη[I − Σ(I − ηW)]⁻¹ΣΦᵀ, where W = ΦᵀΦ + σ²γ(σ)M and Σ gates only active latent variables, expands into an infinite Neumann series. This series shows that perturbations propagate as sparse-gated recurrent spreading along currently active, collinear neurons, generating the geometry-adaptive harmonic eigenvectors observed in both this model and standard d

What carries the argument

The central object is the denoising Jacobian decomposition J(x) = ΦJzΦᵀ with Jz = η[I − Σ(I − ηW)]⁻¹Σ, expanded as Σ + ηΣWΣ + η²ΣWΣWΣ + ⋯. Here Φ is a dictionary of oriented Gabor-like features, W combines the Gram matrix of the dictionary with the learned pairwise interaction matrix scaled by noise, and Σ is a binary diagonal matrix indicating which latent variables are active. This decomposition shows that pixel-space probability flow reduces to latent-space iterative spreading: perturbations are projected into latent space, gated by active neurons, propagated through lateral connections, and synthesized back to pixels. This machinery carries the argument because it turns the black-box dif

Load-bearing premise

The whole equivalence with diffusion models rests on assuming the latent posterior p(z|x) is sharply peaked at its mode, so that the score of the marginal density is approximated by the gradient at a single MAP point; if the posterior is broad or multimodal, especially at high noise, the trained model is not the diffusion model of the true marginal distribution.

What would settle it

Compute the exact or Monte Carlo posterior expectation of z given x for a trained model on ambiguous, high-noise images and compare Φ⟨z⟩ with the MAP reconstruction Φz*; a systematic discrepancy would show that the learned score is not the true marginal score, invalidating the Jacobian-decomposition interpretation. Alternatively, in V1, an optogenetic perturbation during low-contrast stimulation with context orthogonal to the target's preferred orientation should produce no collinear facilitation; if strong facilitation appears regardless of context orientation, the sparse-gating prediction fa

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A single-layer recurrent circuit trained only with denoising score matching can learn collinear facilitation resembling V1 horizontal connections, so such connectivity can emerge from natural-image statistics rather than being hand-specified.
  • The sparse-gated Neumann-series spreading explains why diffusion models develop geometry-adaptive harmonic eigenvectors, giving a neuronal-level account of why these models generalize from finite training sets.
  • The noise-dependent scaling of the interaction matrix predicts that V1 surround influence should shift from suppressive at high contrast to facilitatory at low contrast, with facilitation selectively routed along collinear, context-aligned paths.
  • Detached neurons with near-zero dictionary norms act as an emergent hierarchical layer; ablating them degrades denoising at high noise, suggesting that hierarchical computation can arise within a nominally single-layer circuit.
  • Sampling via the probability-flow ODE using the MAP denoiser reaches generative parity with a parameter-matched U-Net, indicating that the interpretable sparse-coding model captures the core generative behavior of black-box diffusion models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The diffusion-model equivalence rests on the MAP approximation, so the account is most secure in low-noise regimes; at high noise, where the posterior is likely broad or multimodal, the learned denoiser may deviate from the true marginal score, and the Jacobian decomposition would then describe a different denoiser than the one claimed.
  • The Neumann-series view suggests a direct analogy to graph diffusion on the support of active latent variables, which implies that higher-order terms correspond to multi-synaptic pathways; an ablation that truncates the series at first order could quantify how much of contour completion relies on long-range recurrent loops.
  • The detached-neuron phenomenon invites a concrete anatomical hypothesis: some V1 neurons could be largely unresponsive to local feedforward input yet participate in global grouping via lateral connections; this could be tested with recordings that dissociate local receptive-field drive from long-range contextual modulation.
  • The equivalence between this model and diffusion models suggests a transferable interpretability recipe: match the Jacobian eigenbases of a black-box diffusion model to the latent-space spreading of this sparse-coding model to identify which internal units implement contour completion and semantic deformation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes a sparse-coding model of V1 with a learned pairwise interaction matrix M over latent variables, trained on natural images by denoising score matching. The authors argue that, under a MAP approximation to the marginal score, the model is equivalent to a minimal diffusion model, that the learned M recapitulates V1 horizontal/association-field connectivity, and that the denoising Jacobian can be decomposed into a Neumann series in terms of M, mechanistically explaining contour completion and diffusion-model generalization. The paper also reports that a fraction of dictionary elements become 'detached' from the image and act as an emergent hierarchical layer, and it presents denoising and generation results comparing the model to a U-Net.

Significance. If the central derivation were correct, the paper would provide a valuable bridge: an interpretable recurrent circuit whose Jacobian is explicitly related to learned lateral interactions, connecting V1 horizontal connections to the behavior of diffusion models. The empirical observations — emergence of collinear facilitation in M, improved denoising over the factorial baseline, the detached-neuron phenomenon, and qualitative parity with a U-Net — are interesting and worth reporting. The training pipeline (ISTA inference, implicit differentiation, phantom gradients) is a useful engineering contribution. However, the central mechanistic claim rests on an algebraic identity that does not hold, and the MAP score approximation is not validated in the regime studied. The load-bearing interpretability results are therefore unsupported in their current form.

major comments (2)
  1. [Section 3.4, Eqs. (7)–(8)] The Neumann-series decomposition is algebraically incorrect. From the fixed point of Eq. (5), the latent Jacobian is Jz = η[I − Σ(I − ηW)]⁻¹Σ. On the active set, Σ = I, so the bracketed operator reduces to ηW_aa. Hence the active block of Jz is W_aa⁻¹, not Σₖ(ηΣW)ᵏΣ = (I − ηW_aa)⁻¹ as claimed in Eq. (8). The two expressions coincide only when ηW_aa = 0. The expansion in Eq. (8) is the Neumann series for a different operator, [I − ηΣWΣ]⁻¹Σ. Consequently, the mechanistic picture in §4.2 (sparse gating, first-order spread, second-order spread, and higher-order recurrent spreading), the interpretation of Fig. 3D, and the stimulus-dependent routing hypothesis in §4.4 do not follow from the actual Jacobian of the trained denoiser. This is not a matter of external interpretation; a direct algebraic check settles it.
  2. [Section 3.1, Eqs. (3)–(4)] The MAP approximation replaces the true score ∇x log p_{θ,σ}(x) = E_{p(z|x)}[(Φz − x)/σ²] with the single-point gradient (Φz*(x) − x)/σ². The concentration argument in the text is not justified for the high-noise, high-ambiguity regime that the paper studies; the likelihood term becomes flatter as σ grows, and no posterior-variance diagnostic or comparison with a sampling-based score is provided. Because Eq. (4) is exactly a deterministic reconstruction loss, the training procedure may yield a good denoiser without yielding a model of the marginal density p(x). The paper's central claim that the model is 'equivalent to a minimal diffusion model' and the Jacobian analysis of the 'implicit prior' depend on this equivalence. A proof or careful empirical validation of the MAP score approximation is needed before the diffusion-model interpretation can be accepted.
minor comments (4)
  1. [Figure 4] The text refers to panels B–C for dictionary/histogram results and E–F for ablations, but the figure labels are D, E, F, G, H. Please align all in-text references with the actual panel labels.
  2. [References] Reference [55] is misspelled as 'Aleimda'; it should be 'Almeida'. Reference [74] appears incomplete or malformed.
  3. [Appendix A.7] The NLL estimates reported in Fig. 1 inherit the MAP-score approximation. The appendix should state this caveat explicitly when presenting bits/dim values.
  4. [Section 3.4] Typo: 'perbutation' should be 'perturbation'.

Circularity Check

0 steps flagged

No significant circularity: the derivation is self-contained and the V1/diffusion interpretations are post hoc analyses of a trained model, not fitted predictions.

full rationale

I walked the derivation chain: the energy model (Eq. 2), MAP score approximation (Eq. 3), DSM training objective (Eq. 4 / Eq. 9), implicit differentiation (Eq. 6), and the Jacobian decomposition (Eqs. 7-8). At no point is a target quantity defined in terms of the quantity it is supposed to explain, nor is a fitted parameter renamed as a prediction. The interaction matrix M is learned from natural images via the DSM objective and then compared post hoc to V1 horizontal-connection structure; this is empirical interpretation of a trained model, not circular. The V1 hypotheses in Section 4.4 are generated by perturbing the trained model's latent variables, not fitted to V1 data. Self-citations (e.g., Olshausen & Field, Garrigues & Olshausen, Chen et al.) provide background and prior methodology but are not load-bearing for the central claim, and no uniqueness theorem is imported from the authors' prior work. The footnote's tight-frame approximation is an acknowledged modeling assumption rather than a circular reduction. One caveat, which is a correctness concern rather than a circularity: the Neumann-series step from Eq. 7 to Eq. 8 appears algebraically questionable, since [I - Sigma(I - eta W)]^{-1} is not generally expanded as sum (eta Sigma W)^k. But even if Eq. 8 is unsupported, that is an internal derivation error, not an equivalence-by-construction of the kind this pass flags. Overall, the paper's central derivation is self-contained against external benchmarks and does not reduce to its inputs.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 1 invented entities

The central derivation assumes MAP concentration and fixed-point stability, and the biological mapping is asserted rather than quantitatively tested. The learned parameters Φ, M, λ, and γ are all fitted to natural images, so the empirical claims about V1-like structure are observations about fitted values, not parameter-free predictions.

free parameters (5)
  • Interaction matrix M = learned from natural images, values not reported numerically
    Core pairwise prior; the central claim that M mirrors V1 horizontal connections depends on the fitted values.
  • Dictionary Φ = learned from natural images
    Sparse-coding dictionary; the Jacobian decomposition and denoising performance depend on it.
  • Per-neuron sparsity weights λ_i = learned from natural images
    Each latent has its own sparsity penalty in Eq. 5; fitted during training.
  • Noise-balance function γ(σ) (MLP ψ) = learned from natural images
    Controls the σ²γ(σ)M vs Gram drive in Eq. 5; central to the state-dependent routing claim in Section 4.4.
  • Inference step size η = not stated
    ISTA step size in Eq. 5; affects convergence and the Jacobian formula; chosen by hand and not reported.
axioms (4)
  • ad hoc to paper The posterior p(z|x) is highly concentrated around its MAP mode, so the score of the marginal density equals the energy gradient at z* (Eq. 3).
    Introduced without justification or error bounds; this is the bridge from the latent-variable energy to the DSM diffusion-model training objective.
  • domain assumption The ISTA recurrence converges to an isolated stable fixed point with non-singular Jacobian, so the implicit function theorem applies to a ReLU-gated dynamics whose active set is treated as fixed.
    Required for Eq. 6 and Eq. 7; not proven for the multiscale, asymmetric sweeping implementation in Appendix A.8.
  • domain assumption Natural images are well modeled as sparse signals on a continuous manifold, so local linear interpolation along Jacobian eigenvectors is the correct notion of generalization.
    Assumed in Section 5 and used to interpret eigen-decompositions as 'naturalistic deformations'.
  • ad hoc to paper The learned M's collinear facilitation corresponds to V1 horizontal connections and association fields.
    The biological mapping is asserted from visual similarity in Figure 2B and Figure S2 without quantitative physiological fitting or prediction.
invented entities (1)
  • Detached neurons no independent evidence
    purpose: Latent units with near-zero dictionary norms that purportedly form an emergent hierarchical 'second layer' enforcing global consistency across an image (Section 4.3).
    No evidence exists outside this model; ablations show local PSNR effects, but there is no independent falsifiable handle and the bifurcation threshold is descriptive rather than predicted.

pith-pipeline@v1.3.0-alltime-deepseek · 18933 in / 19721 out tokens · 169142 ms · 2026-08-01T22:34:49.452938+00:00 · methodology

0 comments
read the original abstract

We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be readily understood from its parameters. The model is based on sparse coding with a non-factorial prior over latent variables in the form of an unconstrained, pairwise interaction matrix, extending standard sparse coding inference to a general recurrent dynamical system. We efficiently train these recurrent dynamics using a denoising score-matching objective and implicit differentiation. After training on natural images, the learned interaction matrix mirrors the structure of horizontal connections in superficial layers of V1 that link neurons of similar orientation tuning. This model exhibits exceptionally good denoising performance, restoring image features such as extended contours amid extreme visual ambiguity, nearly matching the behavior of standard, black-box diffusion architectures in generalization regime. Owing to the model's simplicity, the network's Jacobian can be decomposed directly in terms of the interaction matrix between latent variables, revealing mechanistically how the recurrent dynamics assign high probability over a continuous family of natural structural deformations. Intriguingly, within this circuit, a large fraction of latent variables learn to disconnect from visual input altogether, essentially forming a hierarchical representation that appears to enforce global consistency among image features. Together, the model and results bridge two distinct domains: for neuroscience, it generates concrete, testable hypotheses regarding functional connectivity in recurrent neural circuits during perceptual inference tasks; for machine learning, it elucidates the internal mechanisms learned by diffusion models that allow them to generate infinitely many novel images from a finite training set.

Figures

Figures reproduced from arXiv: 2607.15693 by Alexander Belsten, Bruno A. Olshausen, Dasheng Bi, Yubei Chen, Zahra Kadkhodaie, Zeyu Yun.

Figure 1
Figure 1. Figure 1: Likelihood and Jacobian analysis of contour stimuli. (A) Estimated negative log￾likelihood (NLL), reported in bits/dim (see Appendix A.7), for three input images. Likelihood strictly increases (lower NLL) across the sequence: purely random lines (#1), random lines with hidden disconnected contours (#2), and the diffusion model’s output when the clean image #2 is evaluated at a nonzero noise embedding, deno… view at source ↗
Figure 2
Figure 2. Figure 2: Emergence of V1-like connectivity and denoising performance. (A) Denoising comparison across increasing noise levels (rows) showing clean, noisy, factorial, and non-factorial reconstructions. (B) Visualization of learned latent interactions for a subset of dictionary elements. Each column is centered on one reference basis function Φi shown in the top row. The yellow/black bar indicates its position and pr… view at source ↗
Figure 3
Figure 3. Figure 3: Mechanistic decomposition of the denoising Jacobian. (A) Dominant eigenvectors of the pixel-space Jacobian J(x) for the non-factorial sparse coding model (σ = 0.17). Boxes denote zoom regions for panel C. (B) The model’s denoised output D(x, σ) on a hidden contour stimulus. (C) Paired pixel-space and latent-space eigenvectors from the highlighted regions. Latent activations are shown as needle plots (orien… view at source ↗
Figure 4
Figure 4. Figure 4: Semantic generalization and the emergence of detached neurons. (A-B) Dominant eigenvectors of the pixel-space Jacobian J(x) for a standard U-Net (A) and our non-factorial sparse coding model (B), showing coordinated semantic deformations. (D) Learned dictionary elements (Φ) sorted by norm magnitude. (E) Dictionary norm histogram revealing a structural bifurcation: ∼30% of the population learns near-zero we… view at source ↗
Figure 5
Figure 5. Figure 5: State-dependent excitation and dynamic geometric routing. (A) The learned scal￾ing factor σ 2γ(σ) that balances prior vs likelihood contribution in the network’s recurrent drive (Φ ⊤Φ + σ 2γ(σ)M). As visual ambiguity (σ) increases, the network upweights the prior M, relying more heavily on learned co-occurrence statistics. (B) Context-dependent excitatory spread for indi￾vidual target neurons across varyin… view at source ↗
Figure 6
Figure 6. Figure 6: Generation Performance. (A) Denoising PSNR curves of our non-factorial sparse coding model versus a parameter-matched U-Net across varying noise levels. (B) Reverse diffusion sampling trajectories. (C) Final generated samples. biologically grounded architecture captures the core structural statistics of complex distributions just as effectively as standard black-box deep network-based models [66–69]. 5 Dis… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

90 extracted references · 12 linked inside Pith

  1. [1]

    Contour integration by the human visual system: evidence for a local “association field

    D. J. Field, A. Hayes, and R. F. Hess, “Contour integration by the human visual system: evidence for a local “association field”,”Vision research, vol. 33, no. 2, pp. 173–193, 1993. 10

  2. [2]

    Robust and inter- pretable blind image denoising via bias-free convolutional neural networks,

    S. Mohan, Z. Kadkhodaie, E. P. Simoncelli, and C. Fernandez-Granda, “Robust and inter- pretable blind image denoising via bias-free convolutional neural networks,”arXiv preprint arXiv:1906.05478, 2019

  3. [3]

    Synaptic physiology of horizontal connections in the cat’s visual cortex,

    J. A. Hirsch and C. D. Gilbert, “Synaptic physiology of horizontal connections in the cat’s visual cortex,”Journal of Neuroscience, vol. 11, no. 6, pp. 1800–1809, 1991

  4. [4]

    The functional organization of local circuits in visual cortex: insights from the study of tree shrew striate cortex,

    D. Fitzpatrick, “The functional organization of local circuits in visual cortex: insights from the study of tree shrew striate cortex,”Cerebral cortex, vol. 6, no. 3, pp. 329–341, 1996

  5. [5]

    Learning horizontal connections in a sparse coding model of natural images,

    P. Garrigues and B. Olshausen, “Learning horizontal connections in a sparse coding model of natural images,”Advances in neural information processing systems, vol. 20, 2007

  6. [6]

    Hierarchical bayesian inference in the visual cortex,

    T. S. Lee and D. Mumford, “Hierarchical bayesian inference in the visual cortex,”Journal of the Optical Society of America A, vol. 20, no. 7, pp. 1434–1448, 2003

  7. [7]

    Neural variability and sampling-based proba- bilistic representations in the visual cortex,

    G. Orbán, P. Berkes, J. Fiser, and M. Lengyel, “Neural variability and sampling-based proba- bilistic representations in the visual cortex,”Neuron, vol. 92, no. 2, pp. 530–543, 2016

  8. [8]

    Cortical-like dynamics in recurrent circuits optimized for sampling-based probabilistic inference,

    R. Echeveste, L. Aitchison, G. Hennequin, and M. Lengyel, “Cortical-like dynamics in recurrent circuits optimized for sampling-based probabilistic inference,”Nature neuroscience, vol. 23, no. 9, pp. 1138–1149, 2020

  9. [9]

    Trace inference, curvature consistency, and curve detection,

    P. Parent and S. Zucker, “Trace inference, curvature consistency, and curve detection,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 11, no. 8, pp. 823–839, 1989

  10. [10]

    Visual brain and visual perception: how does the cortex do perceptual grouping?,

    S. Grossberg, E. Mingolla, and W. D. Ross, “Visual brain and visual perception: how does the cortex do perceptual grouping?,”Trends in neurosciences, vol. 20, no. 3, pp. 106–111, 1997

  11. [11]

    A neural model of contour integration in the primary visual cortex,

    Z. Li, “A neural model of contour integration in the primary visual cortex,”Neural computation, vol. 10, no. 4, pp. 903–940, 1998

  12. [12]

    A detailed theory of thalamic and cortical microcircuits for predictive visual inference,

    D. George, M. Lázaro-Gredilla, W. Lehrach, A. Dedieu, G. Zhou, and J. Marino, “A detailed theory of thalamic and cortical microcircuits for predictive visual inference,”Science Advances, vol. 11, no. 6, p. eadr6698, 2025

  13. [13]

    Generative inference unifies feedback processing for learning and perception in natural and artificial vision,

    T. Toosi and K. D. Miller, “Generative inference unifies feedback processing for learning and perception in natural and artificial vision,”bioRxiv, pp. 2025–10, 2025

  14. [14]

    A generative vision model that trains with high data efficiency and breaks text-based captchas,

    D. George, W. Lehrach, K. Kansky, M. Lázaro-Gredilla, C. Laan, B. Marthi, X. Lou, Z. Meng, Y . Liu, H. Wang,et al., “A generative vision model that trains with high data efficiency and breaks text-based captchas,”Science, vol. 358, no. 6368, p. eaag2612, 2017

  15. [15]

    Emergence of simple-cell receptive field properties by learning a sparse code for natural images,

    B. A. Olshausen and D. J. Field, “Emergence of simple-cell receptive field properties by learning a sparse code for natural images,”Nature, vol. 381, no. 6583, pp. 607–609, 1996

  16. [16]

    Sparse coding with an overcomplete basis set: A strategy employed by V1?,

    B. A. Olshausen and D. J. Field, “Sparse coding with an overcomplete basis set: A strategy employed by V1?,”Vision Research, vol. 37, no. 23, pp. 3311–3325, 1997

  17. [17]

    Relations between the statistics of natural images and the response properties of cortical cells,

    D. J. Field, “Relations between the statistics of natural images and the response properties of cortical cells,”Journal of the Optical Society of America A, vol. 4, no. 12, pp. 2379–2394, 1987

  18. [18]

    What is the goal of sensory coding?,

    D. J. Field, “What is the goal of sensory coding?,”Neural Computation, vol. 6, no. 4, pp. 559– 601, 1994

  19. [19]

    What natural scene statistics can tell us about cortical representation,

    B. A. Olshausen and M. S. Lewicki, “What natural scene statistics can tell us about cortical representation,”The New Visual Neurosciences,(London), pp. 1247–1262, 2014

  20. [20]

    Modeling the joint statistics of images in the wavelet domain,

    E. P. Simoncelli, “Modeling the joint statistics of images in the wavelet domain,” inWavelet Applications in Signal and Image Processing VII, vol. 3813, pp. 188–195, SPIE, 1999

  21. [21]

    A parametric texture model based on joint statistics of complex wavelet coefficients,

    J. Portilla and E. P. Simoncelli, “A parametric texture model based on joint statistics of complex wavelet coefficients,”International journal of computer vision, vol. 40, no. 1, pp. 49–70, 2000. 11

  22. [22]

    Hyvärinen, J

    A. Hyvärinen, J. Hurri, and P. O. Hoyer,Natural image statistics: A probabilistic approach to early computational vision., vol. 39. Springer Science & Business Media, 2009

  23. [23]

    The nonlinear statistics of high-contrast patches in natural images,

    A. B. Lee, K. S. Pedersen, and D. Mumford, “The nonlinear statistics of high-contrast patches in natural images,”International Journal of Computer Vision, vol. 54, no. 1, pp. 83–103, 2003

  24. [24]

    On the local behavior of spaces of natural images,

    G. Carlsson, T. Ishkhanov, V . De Silva, and A. Zomorodian, “On the local behavior of spaces of natural images,”International journal of computer vision, vol. 76, no. 1, pp. 1–12, 2008

  25. [25]

    The sparse manifold transform,

    Y . Chen, D. M. Paiton, and B. A. Olshausen, “The sparse manifold transform,” inAdvances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pp. 10534– 10545, 2018

  26. [26]

    Bag of image patch embedding behind the success of self-supervised learning,

    Y . Chen, A. Bardes, Z. Li, and Y . LeCun, “Bag of image patch embedding behind the success of self-supervised learning,”arXiv preprint arXiv:2206.08954, 2022

  27. [27]

    Edge co-occurrence in natural images predicts contour grouping performance,

    W. S. Geisler, J. S. Perry, B. Super, and D. Gallogly, “Edge co-occurrence in natural images predicts contour grouping performance,”Vision research, vol. 41, no. 6, pp. 711–724, 2001

  28. [28]

    On a common circle: natural scenes and gestalt rules,

    M. Sigman, G. A. Cecchi, C. D. Gilbert, and M. O. Magnasco, “On a common circle: natural scenes and gestalt rules,”Proceedings of the National Academy of Sciences, vol. 98, no. 4, pp. 1935–1940, 2001

  29. [29]

    A fast iterative shrinkage-thresholding algorithm for linear inverse problems,

    A. Beck and M. Teboulle, “A fast iterative shrinkage-thresholding algorithm for linear inverse problems,”SIAM Journal on Imaging Sciences, vol. 2, no. 1, pp. 183–202, 2009

  30. [30]

    Emergence of phase-and shift-invariant features by decomposition of natural images into independent feature subspaces,

    A. Hyvärinen and P. Hoyer, “Emergence of phase-and shift-invariant features by decomposition of natural images into independent feature subspaces,”Neural computation, vol. 12, no. 7, pp. 1705–1720, 2000

  31. [31]

    Natural signal statistics and sensory gain control,

    O. Schwartz and E. P. Simoncelli, “Natural signal statistics and sensory gain control,”Nature neuroscience, vol. 4, no. 8, pp. 819–825, 2001

  32. [32]

    Topographic independent component analysis,

    A. Hyvärinen, P. O. Hoyer, and M. Inki, “Topographic independent component analysis,”Neural Computation, vol. 13, no. 7, pp. 1527–1558, 2001

  33. [33]

    Learning higher-order structures in natural images,

    Y . Karklin and M. S. Lewicki, “Learning higher-order structures in natural images,”Network: Computation in Neural Systems, vol. 14, no. 3, pp. 483–499, 2003

  34. [34]

    A hierarchical bayesian model for learning nonlinear statistical regularities in nonstationary natural signals,

    Y . Karklin and M. S. Lewicki, “A hierarchical bayesian model for learning nonlinear statistical regularities in nonstationary natural signals,”Neural computation, vol. 17, no. 2, pp. 397–423, 2005

  35. [35]

    Learning image representations from the pixel level via hierar- chical sparse coding,

    K. Yu, Y . Lin, and J. Lafferty, “Learning image representations from the pixel level via hierar- chical sparse coding,” inCVPR 2011, pp. 1713–1720, IEEE, 2011

  36. [36]

    A hierarchical statistical model of natural images explains tuning properties in v2,

    H. Hosoya and A. Hyvärinen, “A hierarchical statistical model of natural images explains tuning properties in v2,”Journal of Neuroscience, vol. 35, no. 29, pp. 10412–10428, 2015

  37. [37]

    Learning visual spatial pooling by strong pca dimension reduc- tion,

    H. Hosoya and A. Hyvärinen, “Learning visual spatial pooling by strong pca dimension reduc- tion,”Neural computation, vol. 28, no. 7, pp. 1249–1264, 2016

  38. [38]

    Effect of top-down connections in hierarchical sparse coding,

    V . Boutin, A. Franciosini, F. Ruffier, and L. Perrinet, “Effect of top-down connections in hierarchical sparse coding,”Neural Computation, vol. 32, no. 11, pp. 2279–2309, 2020

  39. [39]

    Subspace locally competi- tive algorithms,

    D. M. Paiton, S. Shepard, K. H. R. Chan, and B. A. Olshausen, “Subspace locally competi- tive algorithms,” inProceedings of the 2020 Annual Neuro-Inspired Computational Elements Workshop, pp. 1–8, 2020

  40. [40]

    Group sparse coding with a laplacian scale mixture prior,

    P. Garrigues and B. Olshausen, “Group sparse coding with a laplacian scale mixture prior,” Advances in neural information processing systems, vol. 23, 2010

  41. [41]

    Bilinear models of natural images,

    B. A. Olshausen, C. Cadieu, J. Culpepper, and D. K. Warland, “Bilinear models of natural images,” inHuman Vision and Electronic Imaging XII, vol. 6492, pp. 67–76, SPIE, 2007. 12

  42. [42]

    Topographic product models applied to natural scene statistics,

    S. Osindero, M. Welling, and G. E. Hinton, “Topographic product models applied to natural scene statistics,”Neural Computation, vol. 18, no. 2, pp. 381–414, 2006

  43. [43]

    Extracting and composing robust features with denoising autoencoders,

    P. Vincent, H. Larochelle, Y . Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” inProceedings of the 25th international conference on Machine learning, pp. 1096–1103, 2008

  44. [44]

    Deep unsupervised learning using nonequilibrium thermodynamics,

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” inInternational Conference on Machine Learning, pp. 2256–2265, PMLR, 2015

  45. [45]

    Neural empirical bayes,

    S. Saremi and A. Hyvärinen, “Neural empirical bayes,”Journal of Machine Learning Research, vol. 20, no. 181, pp. 1–23, 2019

  46. [46]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”Advances in Neural Information Processing Systems, vol. 33, pp. 6840–6851, 2020

  47. [47]

    Score-based generative modeling through stochastic differential equations,

    Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” inInternational Conference on Learning Representations, 2021

  48. [48]

    Generalization in diffusion models arises from geometry-adaptive harmonic representations,

    Z. Kadkhodaie, F. Guth, E. P. Simoncelli, and S. Mallat, “Generalization in diffusion models arises from geometry-adaptive harmonic representations,”arXiv preprint arXiv:2310.02557, 2023

  49. [49]

    An analytic theory of creativity in convolutional diffusion models,

    M. Kamb and S. Ganguli, “An analytic theory of creativity in convolutional diffusion models,” arXiv preprint arXiv:2412.20292, 2024

  50. [50]

    Towards a mechanistic explanation of diffusion model generalization,

    M. Niedoba, B. Zwartsenberg, K. Murphy, and F. Wood, “Towards a mechanistic explanation of diffusion model generalization,”arXiv preprint arXiv:2411.19339, 2024

  51. [51]

    Do diffusion models learn semantically meaningful and efficient representations?,

    Q. Liang, Z. Liu, and I. Fiete, “Do diffusion models learn semantically meaningful and efficient representations?,”arXiv preprint arXiv:2402.03305, 2024

  52. [52]

    A fast iterative shrinkage-thresholding algorithm for linear inverse problems,

    A. Beck and M. Teboulle, “A fast iterative shrinkage-thresholding algorithm for linear inverse problems,”SIAM journal on imaging sciences, vol. 2, no. 1, pp. 183–202, 2009

  53. [53]

    C. J. Rozell, D. H. Johnson, R. G. Baraniuk, and B. A. OlshausenNeural computation, vol. 20, no. 10, pp. 2526–2563, 2008

  54. [54]

    Generalization of back propagation to recurrent and higher order neural networks,

    F. Pineda, “Generalization of back propagation to recurrent and higher order neural networks,” inNeural information processing systems, 1987

  55. [55]

    A learning rule for asynchronous perceptrons with feedback in a combinatorial environment,

    L. B. Almeida, “A learning rule for asynchronous perceptrons with feedback in a combinatorial environment,” inArtificial neural networks: concept learning, pp. 102–111, 1990

  56. [56]

    Deep equilibrium models,

    S. Bai, J. Z. Kolter, and V . Koltun, “Deep equilibrium models,”Advances in neural information processing systems, vol. 32, 2019

  57. [57]

    On training implicit models,

    Z. Geng, X.-Y . Zhang, S. Bai, Y . Wang, and Z. Lin, “On training implicit models,”Advances in neural information processing systems, vol. 34, pp. 24247–24260, 2021

  58. [58]

    Implicit deep learning,

    L. El Ghaoui, F. Gu, B. Travacca, A. Askari, and A. Tsai, “Implicit deep learning,”SIAM Journal on Mathematics of Data Science, vol. 3, no. 3, pp. 930–958, 2021

  59. [59]

    Fixed point diffusion models,

    X. Bai and L. Melas-Kyriazi, “Fixed point diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9430–9440, 2024

  60. [60]

    Hierarchical reasoning model,

    G. Wang, J. Li, Y . Sun, X. Chen, C. Liu, Y . Wu, M. Lu, S. Song, and Y . A. Yadkori, “Hierarchical reasoning model,”arXiv preprint arXiv:2506.21734, 2025

  61. [61]

    Jfb: Jacobian-free backpropagation for implicit networks,

    S. W. Fung, H. Heaton, Q. Li, D. McKenzie, S. Osher, and W. Yin, “Jfb: Jacobian-free backpropagation for implicit networks,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 6648–6656, 2022. 13

  62. [62]

    Independent component filters of natural images compared with simple cells in primary visual cortex,

    J. H. Van Hateren and A. van der Schaaf, “Independent component filters of natural images compared with simple cells in primary visual cortex,”Proceedings of the Royal Society of London. Series B: Biological Sciences, vol. 265, no. 1394, pp. 359–366, 1998

  63. [63]

    Single-neuron perturbations reveal feature-specific competition in v1,

    S. N. Chettih and C. D. Harvey, “Single-neuron perturbations reveal feature-specific competition in v1,”Nature, vol. 567, pp. 334–340, Mar 2019

  64. [64]

    Circuits and mechanisms for surround modulation in visual cortex,

    A. Angelucci, M. Bijanzadeh, L. Nurminen, F. Federer, S. Merlin, and P. C. Bressloff, “Circuits and mechanisms for surround modulation in visual cortex,”Annu. Rev. Neurosci., vol. 40, pp. 425–451, July 2017

  65. [65]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,”CoRR, vol. abs/1505.04597, 2015

  66. [66]

    Deep learning,

    Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,”nature, vol. 521, no. 7553, pp. 436–444, 2015

  67. [67]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016

  68. [68]

    Visualizing and understanding convolutional networks,

    M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” inEuro- pean conference on computer vision, pp. 818–833, Springer, 2014

  69. [69]

    Deconvolutional networks,

    M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus, “Deconvolutional networks,” in2010 IEEE Computer Society Conference on computer vision and pattern recognition, pp. 2528–2535, IEEE, 2010

  70. [70]

    Residual connections encourage iterative inference,

    S. Jastrz˛ ebski, D. Arpit, N. Ballas, V . Verma, T. Che, and Y . Bengio, “Residual connections encourage iterative inference,”arXiv preprint arXiv:1710.04773, 2017

  71. [71]

    A simple early exiting framework for accelerated sampling in diffusion models,

    T. Moon, M. Choi, E. Yun, J. Yoon, G. Lee, J. Cho, and J. Lee, “A simple early exiting framework for accelerated sampling in diffusion models,”arXiv preprint arXiv:2408.05927, 2024

  72. [72]

    Odet (odel): Shortcutting the time and the length in diffusion and flow models for faster sampling,

    D. Gudovskiy, W. Zheng, T. Okuno, Y . Nakata, and K. Keutzer, “Odet (odel): Shortcutting the time and the length in diffusion and flow models for faster sampling,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 6111–6120, 2026

  73. [73]

    Interpretability in the wild: a circuit for indirect object identification in gpt-2 small,

    K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt, “Interpretability in the wild: a circuit for indirect object identification in gpt-2 small,”arXiv preprint arXiv:2211.00593, 2022

  74. [74]

    & batson, j.(2025). on the biology of a large language model,

    J. Lindsey, W. Gurnee, E. Ameisen, B. Chen, A. Pearce, and N. Turner, “& batson, j.(2025). on the biology of a large language model,”Transformer Circuits Thread

  75. [75]

    Minimalistic unsupervised representation learning with the sparse manifold transform,

    Y . Chen, Z. Yun, Y . Ma, B. Olshausen, and Y . LeCun, “Minimalistic unsupervised representation learning with the sparse manifold transform,” inThe Eleventh International Conference on Learning Representations, 2022

  76. [76]

    Nonlinear dimensionality reduction by locally linear embedding,

    S. T. Roweis and L. K. Saul, “Nonlinear dimensionality reduction by locally linear embedding,” science, vol. 290, no. 5500, pp. 2323–2326, 2000

  77. [77]

    A global geometric framework for nonlinear dimensionality reduction,

    J. B. Tenenbaum, V . De Silva, and J. C. Langford, “A global geometric framework for nonlinear dimensionality reduction,”science, vol. 290, no. 5500, pp. 2319–2323, 2000

  78. [78]

    Diffusion model’s generalization can be characterized by inductive biases toward a data-dependent ridge manifold,

    Y . He, Y . Qiu, and M. Tao, “Diffusion model’s generalization can be characterized by inductive biases toward a data-dependent ridge manifold,”arXiv preprint arXiv:2602.06021, 2026

  79. [79]

    What does the retina know about natural scenes?,

    J. J. Atick and A. N. Redlich, “What does the retina know about natural scenes?,”Neural Computation, vol. 4, no. 2, pp. 196–210, 1992

  80. [80]

    Efficient coding of natural images with a population of noisy linear-nonlinear neurons,

    Y . Karklin and E. Simoncelli, “Efficient coding of natural images with a population of noisy linear-nonlinear neurons,”Advances in neural information processing systems, vol. 24, 2011. 14

Showing first 80 references.