Pith. sign in

REVIEW 3 major objections 5 minor 242 references

Thermal equilibrium of a tunable physical system can be programmed to run machine-learning operations, with gradients read from measured covariances — a path to computing near the Landauer limit.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Tunable energy landscapes whose thermal averages equal sigmoid, softmax, and matrix-vector products can, in principle, form the basis of a low-energy analog computer, with a superconducting double-well device as a first experimental step.

T0 review reviewed 2026-08-01 challenge →

load-bearing objection A checkable theoretical blueprint for thermodynamic ML primitives; the hardware prototype is real but its validation does not yet confirm the Langevin model. the 3 major comments →

arxiv 2607.16183 v1 pith:HOLYO5WR submitted 2026-07-17 cs.LG cs.ETphysics.app-ph

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing

classification cs.LG cs.ETphysics.app-ph
keywords thermodynamic computingenergy-based modelsLangevin dynamicsGibbs distributionprobabilistic graphical modelssuperconducting circuitscontrastive divergenceLandauer limit
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a physical system evolving under underdamped Langevin dynamics settles into the Gibbs distribution, so any energy potential that can be physically built is automatically a programmable thermal sampler. The authors construct explicit potentials whose thermal averages implement the elementary operations of neural networks — sigmoid, softmax, matrix-vector product, and addition — and they derive a gradient rule, ∂E[y]/∂θ = −Cov(y, ∂E_θ/∂θ), that turns backpropagation into a covariance measurement on the same hardware. On these primitives they assemble a blueprint for differentiable machine learning, including factor-graph models and a thermodynamic transformer, and report a first experimental building block: a superconducting double-well 'thermodynamic neuron.' The payoff, if the framework holds, is sampling-based machine learning with energy consumption potentially orders of magnitude closer to the Landauer limit than deterministic digital computing. The authors themselves flag that the transformer's layer-norm block is not an equilibrium realization and that projected chip energies exclude cryogenic, control, and calibration overhead.

Core claim

The central claim: equilibrium physics is a sufficient substrate for differentiable machine learning. Underdamped Langevin dynamics (Eq. 8) has the Gibbs distribution π_θ(x) ∝ e^(−βU_θ(x)) as its steady state, so a tunable potential U_θ turns hardware into an energy-based model whose samples come from the physics itself. Specific potentials make thermal expectations equal the sigmoid (Eq. 21), softmax (Eq. 29), matrix-vector products (Eq. 27), and addition (Eq. 28); the identity ∂E[y]/∂θ = −Cov(y, ∂E_θ/∂θ) (Eq. 34) turns training into covariance estimation. These blocks compose as factor graphs into mixture models, HMMs, continuous Ising machines, and a thermodynamic transformer. A fabricate

What carries the argument

The load-bearing object is the Gibbs distribution as the exact steady state of the underdamped Langevin equation: fluctuation-dissipation-matched damping and noise guarantee that the equilibrium of any physically built potential U_θ is the Boltzmann weight e^(−βU_θ), however complicated the landscape. Riding on it are the potential 'gadgets' that translate ML operations into energy landscapes (double-well → sigmoid, simplex-constrained wells → softmax, tilted Gaussians → matrix-vector products and sums), the covariance gradient identity ∂E[y]/∂θ = −Cov(y, ∂E_θ/∂θ) that converts backpropagation into a sampling task, and the factor-graph rule that composes gadgets into larger models. On the ex

Load-bearing premise

The load-bearing premise is that the fabricated superconducting neuron is exactly the idealized underdamped Langevin system, with a single parallel resistance modeling both loss and noise so that escape obeys the Arrhenius law with E_esc = k_B T; the paper's own Fig. 13(c) shows a low-temperature plateau equivalent to ~193 mK against a predicted crossover of 46-76 mK and a thermal slope of ~12 GHz/K against k_B/h = 20.8 GHz/K, discrepancies the authors attribute to the very l

What would settle it

Measure escape energy E_esc versus temperature on a device whose parameters C, L, I_c are pinned down by direct spectroscopy of the neuron itself, not inferred from resonator fits. If the low-temperature plateau still corresponds to ~193 mK rather than the predicted T_cross of 46-76 mK, or the thermal slope still falls short of k_B/h = 20.8 GHz/K, the single-resistor Langevin description of the device is falsified. A complementary numerical check: simulate the circuit with temperature-dependent quasiparticle resistance and see whether any parameter set within the stated uncertainties reproduce

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If equilibrium sampling is as programmable as claimed, energy-based models become trainable on hardware that draws each sample in roughly one thermalization time, removing the sampling bottleneck that makes EBMs intractable digitally.
  • Gradient training of a computation graph reduces to estimating a cross-covariance per parameter block — a native analog operation — so backpropagation can be implemented without digital differentiation.
  • Precision becomes a continuously tunable knob: error scales as N^(−1/2) in sample count, with relative error uniform across magnitudes, unlike floating-point's fixed mantissa.
  • Idealized work per coupling operation is of order k_B T (against a Landauer floor of ln 2·k_B T), the paper's quantitative argument that thermodynamic ML could run orders of magnitude below digital energy per operation.
  • The same primitives compose as factor graphs, so any model expressible as a probabilistic graphical model — transformers, HMMs, mixtures — inherits the paradigm.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The energy projections exclude cryogenic cooling, control electronics, and calibration overhead; if those dominate, as they do in today's superconducting systems, the practical energy advantage over digital hardware remains open even if the equilibrium-computation claim is correct.
  • The covariance-gradient identity is substrate-agnostic: any fluctuating physical system whose steady state is Boltzmann — optomechanical, CMOS, photonic — could run the same training rule, so the blueprint's core could outlive superconducting hardware.
  • The roughly 3x low-temperature escape-energy excess is directly testable: independent characterization of C, L, and I_c by neuron spectroscopy would separate a wrong device model from wrong fitted parameters — a decisive experiment the current data cannot yet adjudicate.
  • A natural next test is whether the sub-k_B/h slope above crossover (12 vs 20.8 GHz/K) is caused by temperature-dependent quasiparticle resistance; a device with a normal-metal shunt of known resistance would make the damping model checkable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a framework for energy-based thermodynamic computing based on sampling from the equilibrium distribution of underdamped Langevin dynamics. It introduces elemental potentials whose thermal expectations approximate sigmoid (Eq. 21), softmax (Eq. 29), matrix-vector products (Eq. 27), and addition (Eq. 28), and derives a covariance-based gradient estimator for parameter gradients (Eq. 34). The framework is then embedded in factor graphs, extended to on-chip self-learning via a Born-Oppenheimer force, and illustrated with numerical demos (PGM, GMM, HMM, continuous Ising, thermoformer). A superconducting double-well 'thermodynamic neuron' is presented as a preliminary experimental realization, with escape-rate measurements intended to confirm the Kramers/Arrhenius prediction E_esc = k_B T (Sec. VIII).

Significance. The theoretical core of the paper is largely sound and useful: the potential constructions are explicit and concrete, the gradient identity of Eq. (34) is derived from first principles and is potentially hardware-friendly, and the factor-graph assembly is a coherent organizing principle. The numerical demonstrations are illustrative, and the framework makes falsifiable predictions for thermal expectations that could be tested in simulation or hardware. However, the experimental section, which is the paper's claim of a hardware proof-of-principle, does not quantitatively validate the thermal-activation model: the data in Fig. 13(c) show a low-temperature plateau corresponding to ~193 mK against a predicted T_cross of 46–76 mK, and a high-temperature slope of ~12 GHz/K against the expected k_B/h = 20.8 GHz/K. The authors' own explanation relies on parameter uncertainty and unmodeled quasiparticle damping, but no independent calibration is provided. Thus the hardware claim is not established, while the framework-level derivations stand.

major comments (3)
  1. [Sec. VIII C, Fig. 13(c), Eqs. (60)-(62)] The escape-energy data do not confirm the claimed thermal-activation regime. The low-temperature plateau E_esc/h = 4.02 GHz corresponds to ~193 mK, while T_cross estimated from Eq. (60) is 46–76 mK; the slope above 100 mK is ~12 GHz/K versus k_B/h = 20.8 GHz/K. The authors attribute this to 'very large uncertainty' in C, L, I_c and to quasiparticle damping. However, the same fitted parameters determine the ω_b used in T_cross and the ω_p and ΔU used in the Arrhenius fit of Eq. (65), so the comparison is not an independent test. With binary readout only and no direct measurement of R(T) or the thermodynamic neuron's own frequency, the experiment cannot distinguish a wrong device model from wrong fitted parameters.
  2. [Sec. VIII C, Eq. (60), Appendix H] The identification of the low-temperature plateau with k_B T_cross is asserted without derivation. For macroscopic quantum tunneling, the escape rate is not generally of the Arrhenius form with E_esc = k_B T_cross; the effective activation energy depends on the action and damping. Moreover, the missing data between 60 and 100 mK (Appendix J) means the crossover region is never directly measured, so the transition from plateau to linear rise is inferred, not observed. A conclusive test requires resolving the crossover or fitting the full escape-rate expression with independently calibrated parameters.
  3. [Sec. VIII B 2, Appendix H] The extraction of E_esc via Eq. (65) relies on ω_p and ΔU computed from the Hamiltonian fit. According to Appendix H, the fit uses 10 free parameters after fixing C and is performed only against readout resonator frequencies, not against the thermodynamic neuron's dynamics. The authors state the fit has 'very large uncertainty' in C, L, I_c, yet they do not propagate this uncertainty into E_esc or report confidence intervals for the slopes in Fig. 13(b). The claim that the measured slope is 40% below k_B/h is therefore not statistically grounded. Please propagate the fit covariance into E_esc and provide an independent estimate of ω_p and ΔU.
minor comments (5)
  1. [Eq. (59)] The displayed Hamiltonian is corrupted by stray tokens ("⌟⟨⟨⟪rl⟫l⟩⟩...") and is unreadable. The equation must be reset to the intended expression.
  2. [Sec. VII E, Eq. (58)] The variance potential is described as computing the variance of z, but the formula involves auxiliary x_i and y with the term (x_i - (z_i - μ))^2. The role of these auxiliary variables and the domain of validity should be clarified, especially since the text states this potential is not meant to be an equilibrium realization.
  3. [References] Several core components are cited only as pending patent applications by the same authors (e.g., Refs. [100-102], [114-115], [134], [147-149], [189], [202-203], [205]). For a journal publication, please provide public references or full technical descriptions in the text so that reviewers and readers can verify these components.
  4. [General] No code or data repository is provided for the numerical experiments (PGM, GMM, HMM, Ising, thermoformer) or for the experimental escape-rate data. The manuscript would be stronger if the authors released code/data, or at least provided the raw fit parameters and uncertainties for the key figures.
  5. [Sec. VII, Fig. 9] The tokens-per-joule projections exclude cryogenic cooling, control electronics, and calibration overhead, as stated in the text. The comparison with H100 GPUs in Fig. 9 should be interpreted with care; a clear caveat in the caption would prevent over-generalization.

Circularity Check

0 steps flagged

No significant circularity: the Gibbs steady state, potential constructions, and covariance-gradient identity are derived in-text or from standard external results; the experimental discrepancy is an acknowledged validation limitation, not a constructional fit.

full rationale

The paper's claimed derivation chain is self-contained. The steady-state Gibbs distribution of the underdamped Langevin equation (Eqs. 8-11) is the standard Klein-Kramers equilibrium solution, cited to Risken, so it does not reduce to the paper's own inputs. The sigmoid, matrix-vector, addition, and softmax operations (Eqs. 21, 27-29) are obtained by writing potentials whose hard-constraint equilibrium expectations are these functions; these are constructions, not fitted parameters, and the numerical checks use fixed stated potentials. The gradient rule Eq. (34) is derived in full in Sec. IV A from the quotient rule. The experimental section fits C, L, and I_c to the readout-resonator spectrum (Appendix H), then compares the extracted escape energy to the independently stated k_B T and T_cross predictions; the mismatches (low-temperature plateau ~193 mK vs T_cross 46-76 mK; slope ~12 GHz/K vs k_B/h = 20.8 GHz/K) are explicitly reported and attributed to the 'very large uncertainty' of Hamiltonian parameters and unmodeled quasiparticle damping. That is a validation weakness - the experiment cannot currently distinguish wrong device model from wrong fitted parameters - but it is not circular, because E_esc is not used to define the model parameters. Citations to the authors' pending patent applications (e.g., relay oscillators [100,101], mean-field propagation [114,115], natural-gradient readout [147], thermodynamic neuron [205]) name hardware components but are not load-bearing for the mathematical derivations, which are either given in the body or explicitly described as speculative. No load-bearing step reduces by definition to its input; therefore no circularity step is exhibited.

Axiom & Free-Parameter Ledger

5 free parameters · 7 axioms · 3 invented entities

The framework rests on standard statistical mechanics (Gibbs/Langevin/Fokker-Planck), so its axiom burden is modest. The experimental and projection content carries the free parameters: the device Hamiltonian parameters (fitted, with acknowledged large uncertainty), the lambda values that set approximation quality, and assumed hardware parameters for the energy projections. The invented entities are hardware concepts; only the thermodynamic neuron has been demonstrated.

free parameters (5)
  • Hamiltonian fit parameters (C, L, L1,2, Ic, chi, M) = C=120 fF (fixed), L=750 pH, L1,2=50 pH, Ic=0.997 uA, chi=0.0104, M=24.7 pH
    Fitted to readout-resonator frequency maps via Levenberg-Marquardt (Appendix H, Table I) with 11 free parameters; T_cross and barrier-height predictions are computed from them, and the paper assigns the E_esc discrepancy to their uncertainty.
  • lambda_1, lambda_2 potential strengths = lambda=10 for softmax convergence plots; lambda=1.0 for thermoformer; lambda_c=100 for work simulations
    Hand-chosen; they set the approximation-vs-speed tradeoff (well height vs Kramers time) and directly control the numeric results and the energy/time projections.
  • Hardware projection parameters = 100 fF, 100 pH, 20 kOhm/100 Ohm, 150 mK/50 mK
    Assumed superconducting device parameters used to produce the tokens/Joule projections (Secs. III and VII E); stated to be 'not unreasonable' but not measured.
  • Escape-energy extraction fits = A, B, tau in Eq. (64) per relaxation curve; slope of ln(Gamma/omega_p) vs Delta-U per temperature
    Per-curve and per-temperature fits that define the reported E_esc values; Delta-U and omega_p come from the Hamiltonian model, so the extracted 'escape energy' is model-dependent.
  • Damping prefactor a_t (Eq. 62) = regime-dependent: |omega_b|/eta, 1, or proportional to eta*sqrt(C*Delta-U/k_B T)
    Modeling choice; the paper says a_t may change with temperature through 1/R, which is not accounted for in the E_esc analysis.
axioms (7)
  • standard math Equilibrium of the Fokker-Planck equation for underdamped Langevin dynamics is the Gibbs distribution (Sec. II C, Eqs. 10-11)
    Textbook result (Risken); the backbone of the framework.
  • standard math Fluctuation-dissipation: noise amplitude sqrt(2*gamma_i/beta) in Eq. (8)
    Standard relation linking damping and thermal noise; assumed to hold for the superconducting device.
  • standard math Kramers/Arrhenius escape law Eq. (62) with k_B T << Delta-U validity and regime-dependent prefactor a_t
    Known result, but its application to this device over the measured range (with small barriers where k_B T ~ Delta-U) is an extrapolation the paper itself notes 'has been shown to be effective even when k_B T ~ Delta-U'.
  • domain assumption Born-Oppenheimer timescale separation for slow parameter dynamics (Sec. VI, Eqs. 44-45)
    Assumes parameter dynamics are much slower than variable dynamics so fast variables stay in equilibrium; not demonstrated.
  • domain assumption Device is described by underdamped Langevin dynamics with a single effective parallel resistance R for loss and noise (Sec. VIII A, Eq. 61)
    Load-bearing for the experiment; quasiparticle-induced temperature dependence of R is acknowledged but not modeled in the E_esc extraction.
  • domain assumption The simplified Hamiltonian Eq. (59) captures the potential landscape
    The paper itself uses the more complex 'no caps' model of Appendix H for fitting; the simplified model is stated in the main text and neglects junction asymmetry and intrinsic capacitances/inductances.
  • domain assumption Samples used in expectation estimates are effectively independent
    The O(N^{-1/2}) error scaling (Secs. II D-E) assumes independent samples; thermalized samples from a single device are correlated across tau_therm.
invented entities (3)
  • Thermodynamic neuron independent evidence
    purpose: Tunable double-well superconducting circuit implementing the core double-well potential of the framework
    Experimentally characterized in Sec. VIII (escape rates vs temperature); the measured escape energies do not yet match theory quantitatively.
  • Estimation (relay) oscillators no independent evidence
    purpose: Averaging component to convert samples into expected values without digital conversion (Sec. III B)
    No hardware demonstration; the paper states these are 'likely a difficult operation to perform on superconducting hardware' (Sec. IX A).
  • Born-Oppenheimer parameter oscillators (self-learning) no independent evidence
    purpose: Promoting model parameters to physical degrees of freedom whose early-time dynamics encode gradients (Sec. VI)
    Sketch with early-time expansions in Appendix B; no implementation.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing." pith.science (2026). https://pith.science/paper/HOLYO5WR

@misc{pith2026260716183,
  author       = {Pith},
  title        = {Pith review of: A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HOLYO5WR}},
  note         = {Machine review of arXiv:2607.16183}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

To address the escalating energy and latency demands of machine-learning workloads, we introduce a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical hardware. In this work, we focus on energy-based thermodynamic computing where the stochastic process is well described by Langevin dynamics with tunable energy potentials. The implementation of such potentials in physical hardware enables us to generate and sample from basic parameterized energy-based models. We demonstrate how to construct and train popular classes of machine learning models based on these hardware-native energy-based models, using the framework of probabilistic graphical models. We analyze the runtime and energy consumption of different models in this thermodynamic paradigm based on theoretical considerations and numerical studies. As a preliminary experimental realization of such hardware, we present our stochastic analog superconducting circuits driven by thermal noise. Together, these results outline a path toward energy-efficient thermodynamic hardware for probabilistic machine learning.

Figures

Figures reproduced from arXiv: 2607.16183 by Christopher Chamberland, Frank Sch\"afer, Guillaume Verdon, J\'er\'emy B\'ejanin, Joost Bus, Owen Lockwood, Patrick Huembeli.

Figure 1
Figure 1. Figure 1: Comparison of energy curves and probability dis [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Distributions of the work for the coupling in [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Convergence of the expectation of many trajecto [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Samples vs. data for a trained Gaussian PGM. [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Components of a (purely digital) Gaussian Mixture [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Training a 2D HMM on time series data. High [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Training a continuous Ising machine. Mean-squared [PITH_FULL_IMAGE:figures/full_fig_p016_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Diagram of transformer-decoder model. thermalize, the total time can be obtained by summing the thermalization times of each block. Similarly, the total energy can be computed by summing the energy contributions from each step. With values that are not unreasonable in current day superconducting hardware (100 fF, 100 pH, 100 Ω, operating at 50 mK, with λ parameters of 1), the resulting chip projections sug… view at source ↗
Figure 10
Figure 10. Figure 10: Training of a thermoformer with different num [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 9
Figure 9. Figure 9: Tokens per Joule for varying decoder depths and [PITH_FULL_IMAGE:figures/full_fig_p017_9.png] view at source ↗
Figure 11
Figure 11. Figure 11: Circuit diagram of the thermodynamic neuron [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Sigmoidal relationship between the expectation value of position and a linear force for the double-well per Eq. ( [PITH_FULL_IMAGE:figures/full_fig_p021_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Results of the relaxation experiments. (a) Trajectories of the left well population [PITH_FULL_IMAGE:figures/full_fig_p022_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Heatmap of ∣S21∣ data versus ϕbar and ϕtilt for readout frequency fprobe = 11.790 GHz. The mapping between control voltage and flux was done using Eq. (D1). The data on the corners is missing due to voltage amplitude limits of the AWG. φ 1 initialize (a) φ 2 confine φ 3 relax φ 4 readout ⋆ 1 1 2 4 0.5 0.6 0.7 0.8 0.9 −0.2 −0.1 0 0.1 0.2 φbar (Φ0) φtilt (Φ0) −53 −52 −51 −50 ∣S21∣ (dB) (b) [PITH_FULL_IMAGE… view at source ↗
Figure 15
Figure 15. Figure 15: (a) Experimental sequence of the relaxation experiment. The system is initialized in either the left or right well [PITH_FULL_IMAGE:figures/full_fig_p037_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Optical image of the thermodynamic neuron, with the tilt flux line on the left at the shorted end, readout resonator [PITH_FULL_IMAGE:figures/full_fig_p038_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Readout resonator frequency as function of [PITH_FULL_IMAGE:figures/full_fig_p040_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: (bottom) Visibility of the readout procedure at 120 mK for variable flux and tilt control voltages where the readout [PITH_FULL_IMAGE:figures/full_fig_p041_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: Non-affine voltage-flux relationship at 80 mK. [PITH_FULL_IMAGE:figures/full_fig_p041_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Experimental setup diagram of the Bluefors LD400 dilution refrigerator and the wiring. The refrigerator has a room [PITH_FULL_IMAGE:figures/full_fig_p042_20.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

242 extracted references · 42 linked inside Pith

  1. [1]

    de Vries, The growing energy footprint of artificial intelligence, Joule7, 2191 (2023)

    A. de Vries, The growing energy footprint of artificial intelligence, Joule7, 2191 (2023)

  2. [2]

    Reuther, P

    A. Reuther, P. Michaleas, M. Jones, V. Gadepally, S. Samsi, and J. Kepner, Ai and ml accelerator survey and trends, in2022 IEEE High Performance Extreme Computing Conference (HPEC)(IEEE, 2022) pp. 1–10

  3. [3]

    Luccioni, Y

    S. Luccioni, Y. Jernite, and E. Strubell, Power hungry processing: Watts driving the cost of ai deployment?, in The 2024 ACM Conference on Fairness, Accountability, and Transparency(2024) pp. 85–99

  4. [4]

    Masanet, A

    E. Masanet, A. Shehabi, N. Lei, S. Smith, and J. Koomey, Recalibrating global data center energy-use estimates, Science367, 984 (2020)

  5. [5]

    Mohseni, A

    M. Mohseni, A. Scherer, K. G. Johnson, O. Wertheim, M. Otten, N. A. Aadit, K. M. Bresniker, K. Y. Cam- sari, B. Chapman, S. Chatterjee,et al., How to build a quantum supercomputer: Scaling challenges and oppor- tunities, arXiv preprint arXiv:2411.10406 (2024)

  6. [6]

    N. P. De Leon, K. M. Itoh, D. Kim, K. K. Mehta, T. E. Northup, H. Paik, B. Palmer, N. Samarth, S. Sangtawesin, and D. W. Steuerman, Materials chal- lenges and opportunities for quantum computing hard- ware, Science372, eabb2823 (2021)

  7. [7]

    Wendin, Quantum information processing with su- perconducting circuits: a review, Reports on Progress in Physics80, 106001 (2017)

    G. Wendin, Quantum information processing with su- perconducting circuits: a review, Reports on Progress in Physics80, 106001 (2017)

  8. [8]

    D. H. Wolpert, The stochastic thermodynamics of com- putation, Journal of Physics A: Mathematical and The- oretical52, 193001 (2019)

  9. [9]

    D. H. Wolpert, J. Korbel, C. W. Lynn, F. Tasnim, J. A. Grochow, G. Karde¸ s, J. B. Aimone, V. Balasubrama- nian, E. De Giuli, D. Doty,et al., Is stochastic thermo- dynamics the key to understanding the energy costs of computation?, Proceedings of the National Academy of Sciences121, e2321112121 (2024)

  10. [10]

    Hooker, The hardware lottery, Communications of the ACM64, 58 (2021)

    S. Hooker, The hardware lottery, Communications of the ACM64, 58 (2021)

  11. [11]

    Conte, E

    T. Conte, E. DeBenedictis, N. Ganesh, T. Hylton, J. P. Strachan, R. S. Williams, A. Alemi, L. Altenberg, G. Crooks, J. Crutchfield,et al., Thermodynamic com- 25 puting, arXiv preprint arXiv:1911.01968 (2019)

  12. [12]

    Langevin, Sur la th´ eorie du mouvement brownien (1908)

    P. Langevin, Sur la th´ eorie du mouvement brownien (1908)

  13. [13]

    G. E. Crooks,Excursions in Statistical Dynamics, Ph.D. thesis, University of California at Berkeley, Berkeley, California, United States of America (1999)

  14. [14]

    A. A. Golubov, M. Y. Kupriyanov, and E. Il’Ichev, The current-phase relation in josephson junctions, Reviews of modern physics76, 411 (2004)

  15. [15]

    Saira, M

    O.-P. Saira, M. H. Matheny, R. Katti, W. Fon, G. Wim- satt, J. P. Crutchfield, S. Han, and M. L. Roukes, Nonequilibrium thermodynamics of erasure with su- perconducting flux logic, Physical Review Research2, 013249 (2020), publisher: American Physical Society

  16. [16]

    Huembeli, J

    P. Huembeli, J. M. Arrazola, N. Killoran, M. Mohseni, and P. Wittek, The physics of energy-based models, Quantum Machine Intelligence4, 1 (2022)

  17. [17]

    Lockwood, F

    O. Lockwood, F. Sch¨ afer, and P. Huembeli, Energy Based Models with Deep Neural Networks: A Review (2025), work in progress

  18. [18]

    LeCun, S

    Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, F. Huang,et al., A tutorial on energy-based learning, Predicting structured data1(2006)

  19. [19]

    Du and I

    Y. Du and I. Mordatch, Implicit generation and model- ing with energy based models, Advances in Neural In- formation Processing Systems32(2019)

  20. [20]

    Song and D

    Y. Song and D. P. Kingma, How to train your energy- based models, arXiv preprint arXiv:2101.03288 (2021)

  21. [21]

    Jelinˇ ciˇ c, O

    A. Jelinˇ ciˇ c, O. Lockwood, A. Garlapati, G. Verdon, and T. McCourt, An efficient probabilistic hardware architecture for diffusion-like models, arXiv preprint arXiv:2510.23972 (2025)

  22. [22]

    W. Moy, I. Ahmed, P.-w. Chiu, J. Moy, S. S. Sapatnekar, and C. H. Kim, A 1,968-node coupled ring oscillator circuit for combinatorial optimization problem solving, Nature Electronics5, 310 (2022)

  23. [23]

    Inagaki, Y

    T. Inagaki, Y. Haribara, K. Igarashi, T. Sonobe, S. Ta- mate, T. Honjo, A. Marandi, P. L. McMahon, T. Umeki, K. Enbutsu,et al., A coherent ising machine for 2000- node optimization problems, Science354, 603 (2016)

  24. [24]

    Mohseni, P

    N. Mohseni, P. L. McMahon, and T. Byrnes, Ising ma- chines as hardware solvers of combinatorial optimization problems, Nature Reviews Physics4, 363 (2022)

  25. [25]

    N. A. Aadit, A. Grimaldi, M. Carpentieri, L. Theogara- jan, J. M. Martinis, G. Finocchio, and K. Y. Camsari, Massively parallel probabilistic computing with sparse ising machines, Nature Electronics5, 460 (2022)

  26. [26]

    N. S. Singh, K. Kobayashi, Q. Cao, K. Selcuk, T. Hu, S. Niazi, N. A. Aadit, S. Kanai, H. Ohno, S. Fukami, et al., Cmos plus stochastic nanomagnets enabling het- erogeneous computers for probabilistic inference and learning, Nature Communications15, 2685 (2024)

  27. [27]

    Laydevant, D

    J. Laydevant, D. Markovi´ c, and J. Grollier, Training an ising machine with equilibrium propagation, Nature Communications15, 3671 (2024)

  28. [28]

    J. Chou, S. Bramhavar, S. Ghosh, and W. Herzog, Ana- log coupled oscillator based weighted ising machine, Sci- entific reports9, 14786 (2019)

  29. [29]

    K. Y. Camsari, R. Faria, B. M. Sutton, and S. Datta, Stochastic p-bits for invertible logic, Physical Review X 7, 031014 (2017)

  30. [30]

    K. Y. Camsari, B. M. Sutton, and S. Datta, P-bits for probabilistic spin logic, Applied Physics Reviews6 (2019)

  31. [31]

    Kaiser and S

    J. Kaiser and S. Datta, Probabilistic computing with p-bits, Applied Physics Letters119(2021)

  32. [32]

    Chowdhury, A

    S. Chowdhury, A. Grimaldi, N. A. Aadit, S. Niazi, M. Mohseni, S. Kanai, H. Ohno, S. Fukami, L. Theog- arajan, G. Finocchio,et al., A full-stack view of proba- bilistic computing with p-bits: devices, architectures, and algorithms, IEEE Journal on Exploratory Solid- State Computational Devices and Circuits9, 1 (2023)

  33. [33]

    Niazi, S

    S. Niazi, S. Chowdhury, N. A. Aadit, M. Mohseni, Y. Qin, and K. Y. Camsari, Training deep boltzmann networks with sparse ising machines, Nature Electronics , 1 (2024)

  34. [34]

    Freitas, G

    N. Freitas, G. Massarelli, J. Rothschild, D. Keane, E. Dawe, S. Hwang, A. Garlapati, and T. McCourt, Taming nonequilibrium thermal fluctuations in sub- threshold cmos circuits (2026)

  35. [35]

    M. W. Johnson, M. H. Amin, S. Gildert, T. Lanting, F. Hamze, N. Dickson, R. Harris, A. J. Berkley, J. Jo- hansson, P. Bunyk,et al., Quantum annealing with manufactured spins, Nature473, 194 (2011)

  36. [36]

    A. D. King, S. Suzuki, J. Raymond, A. Zucca, T. Lant- ing, F. Altomare, A. J. Berkley, S. Ejtemaee, E. Hoskin- son, S. Huang,et al., Coherent quantum annealing in a programmable 2,000 qubit ising chain, Nature Physics 18, 1324 (2022)

  37. [37]

    A. D. King, J. Raymond, T. Lanting, R. Harris, A. Zucca, F. Altomare, A. J. Berkley, K. Boothby, S. Ejtemaee, C. Enderud,et al., Quantum critical dy- namics in a 5,000-qubit programmable spin glass, Na- ture617, 61 (2023)

  38. [38]

    J. M. Shainline, S. M. Buckley, R. P. Mirin, and S. W. Nam, Superconducting optoelectronic circuits for neuro- morphic computing, Physical Review Applied7, 034013 (2017)

  39. [39]

    Kumar, U

    A. Kumar, U. S. Goteti, E. Cubukcu, R. C. Dynes, and D. Kuzum, Evaluation of fluxon synapse device based on superconducting loops for energy efficient neuromor- phic computing, Frontiers in Neuroscience19, 1511371 (2025)

  40. [40]

    Kudithipudi, C

    D. Kudithipudi, C. Schuman, C. M. Vineyard, T. Pan- dit, C. Merkel, R. Kubendran, J. B. Aimone, G. Or- chard, C. Mayr, R. Benosman,et al., Neuromorphic computing at scale, Nature637, 801 (2025)

  41. [41]

    J. B. Aimone, Neuromorphic computing: A theoretical framework for time, space, and energy scaling, arXiv preprint arXiv:2507.17886 (2025)

  42. [42]

    P. J. Coles, C. Szczepanski, D. Melanson, K. Donatella, A. J. Martinez, and F. Sbahi, Thermodynamic ai and the fluctuation frontier, in2023 IEEE International Conference on Rebooting Computing (ICRC)(IEEE,

  43. [43]

    Aifer, K

    M. Aifer, K. Donatella, M. H. Gordon, S. Duffield, T. Ahle, D. Simpson, G. Crooks, and P. J. Coles, Ther- modynamic linear algebra, npj Unconventional Com- puting1, 13 (2024)

  44. [44]

    Melanson, M

    D. Melanson, M. A. Khater, M. Aifer, K. Donatella, M. H. Gordon, T. Ahle, G. Crooks, A. J. Mar- tinez, F. Sbahi, and P. J. Coles, Thermodynamic computing system for ai applications, arXiv preprint arXiv:2312.04836 (2023)

  45. [45]

    Aifer, S

    M. Aifer, S. Duffield, K. Donatella, D. Melanson, P. Klett, Z. Belateche, G. Crooks, A. J. Martinez, and P. J. Coles, Thermodynamic bayesian inference, arXiv preprint arXiv:2410.01793 (2024). 26

  46. [46]

    Donatella, S

    K. Donatella, S. Duffield, M. Aifer, D. Melanson, G. Crooks, and P. J. Coles, Thermodynamic natu- ral gradient descent, arXiv preprint arXiv:2405.13817 (2024)

  47. [47]

    Osadchy, M

    M. Osadchy, M. Miller, and Y. Cun, Synergistic face de- tection and pose estimation with energy-based models, Advances in neural information processing systems17 (2004)

  48. [48]

    Ranzato, C

    M. Ranzato, C. Poultney, S. Chopra, and Y. Cun, Effi- cient learning of sparse representations with an energy- based model, Advances in neural information processing systems19(2006)

  49. [49]

    S. Zhai, Y. Cheng, W. Lu, and Z. Zhang, Deep struc- tured energy based models for anomaly detection, in International conference on machine learning(PMLR,

  50. [50]

    N. Liu, S. Li, Y. Du, A. Torralba, and J. B. Tenenbaum, Compositional visual generation with composable diffu- sion models, inEuropean Conference on Computer Vi- sion(Springer, 2022) pp. 423–439

  51. [51]

    G. E. Hinton, Training products of experts by mini- mizing contrastive divergence, Neural computation14, 1771 (2002)

  52. [52]

    M. A. Carreira-Perpinan and G. Hinton, On contrastive divergence learning, inInternational workshop on artifi- cial intelligence and statistics(PMLR, 2005) pp. 33–40

  53. [53]

    Bengio and O

    Y. Bengio and O. Delalleau, Justifying and generalizing contrastive divergence, Neural computation21, 1601 (2009)

  54. [54]

    Sutskever and T

    I. Sutskever and T. Tieleman, On the convergence prop- erties of contrastive divergence, inProceedings of the thirteenth international conference on artificial intelli- gence and statistics(JMLR Workshop and Conference Proceedings, 2010) pp. 789–795

  55. [55]

    Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, Score-based generative mod- eling through stochastic differential equations, arXiv preprint arXiv:2011.13456 (2020)

  56. [56]

    Cheng and P

    X. Cheng and P. Bartlett, Convergence of langevin mcmc in kl-divergence, inAlgorithmic Learning Theory (PMLR, 2018) pp. 186–211

  57. [57]

    Cheng, N

    X. Cheng, N. S. Chatterji, P. L. Bartlett, and M. I. Jor- dan, Underdamped langevin mcmc: A non-asymptotic analysis, inConference on learning theory(PMLR,

  58. [58]

    Zhang, X

    R. Zhang, X. Liu, and Q. Liu, A langevin-like sampler for discrete distributions, inInternational Conference on Machine Learning(PMLR, 2022) pp. 26375–26396

  59. [59]

    H. Sun, H. Dai, B. Dai, H. Zhou, and D. Schuur- mans, Discrete langevin samplers via wasserstein gra- dient flow, inInternational Conference on Artificial In- telligence and Statistics(PMLR, 2023) pp. 6290–6313

  60. [60]

    G. O. Roberts and R. L. Tweedie, Exponential conver- gence of Langevin distributions and their discrete ap- proximations, Bernoulli2, 341 (1996)

  61. [61]

    Balakrishnan, Fluctuation-dissipation theorems from the generalised langevin equation, Pramana12, 301 (1979)

    V. Balakrishnan, Fluctuation-dissipation theorems from the generalised langevin equation, Pramana12, 301 (1979)

  62. [62]

    S¨ arkk¨ a and A

    S. S¨ arkk¨ a and A. Solin,Applied stochastic differential equations, Vol. 10 (Cambridge University Press, 2019)

  63. [63]

    Risken,The Fokker-Planck Equation: Methods of So- lution and Applications, second edition ed., edited by H

    H. Risken,The Fokker-Planck Equation: Methods of So- lution and Applications, second edition ed., edited by H. Haken (Springer, Berlin, 1989)

  64. [64]

    C. H. Bennett, The thermodynamics of computation—a review, International Journal of Theoretical Physics21, 905 (1982)

  65. [65]

    J. A. Vaccaro and S. M. Barnett, Information erasure without an energy cost, Proceedings of the Royal Soci- ety A: Mathematical, Physical and Engineering Sciences 467, 1770 (2011)

  66. [66]

    Fredkin and T

    E. Fredkin and T. Toffoli, Conservative logic, Interna- tional Journal of theoretical physics21, 219 (1982)

  67. [67]

    S. Shankar, Energy estimates across layers of comput- ing: from devices to large-scale applications in machine learning for natural language processing, scientific com- puting, and cryptocurrency mining, in2023 IEEE High Performance Extreme Computing Conference (HPEC) (IEEE, 2023) pp. 1–6

  68. [68]

    Nakazato and S

    M. Nakazato and S. Ito, Geometrical aspects of en- tropy production in stochastic thermodynamics based on wasserstein distance, Physical Review Research3, 043093 (2021)

  69. [69]

    Reilly and S

    M. Reilly and S. Lloyd, Physical complexity and black hole quantum computers, inJournal of Physics: Con- ference Series, Vol. 3017 (IOP Publishing, 2025) p. 012010

  70. [70]

    Lahiri, J

    S. Lahiri, J. Sohl-Dickstein, and S. Ganguli, A universal tradeoff between power, precision and speed in phys- ical communication, arXiv preprint arXiv:1603.07758 (2016)

  71. [71]

    Ito, Stochastic thermodynamic interpretation of in- formation geometry, Physical review letters121, 030605 (2018)

    S. Ito, Stochastic thermodynamic interpretation of in- formation geometry, Physical review letters121, 030605 (2018)

  72. [72]

    Ito and A

    S. Ito and A. Dechant, Stochastic time evolution, infor- mation geometry, and the cram´ er-rao bound, Physical Review X10, 021056 (2020)

  73. [73]

    Ito, Geometric thermodynamics for the fokker–planck equation: stochastic thermodynamic links between in- formation geometry and optimal transport, Information Geometry7, 441 (2024)

    S. Ito, Geometric thermodynamics for the fokker–planck equation: stochastic thermodynamic links between in- formation geometry and optimal transport, Information Geometry7, 441 (2024)

  74. [74]

    Hnybida and S

    J. Hnybida and S. Verret, Minimal-dissipation learning for energy-based models, arXiv preprint arXiv:2510.03137 (2025)

  75. [75]

    Villaniet al.,Optimal transport: old and new, Vol

    C. Villaniet al.,Optimal transport: old and new, Vol. 338 (Springer, 2009)

  76. [76]

    Klinger and G

    J. Klinger and G. M. Rotskoff, Minimally dissi- pative multi-bit logical operations, arXiv preprint arXiv:2506.24021 (2025)

  77. [77]

    Rolandi, P

    A. Rolandi, P. Abiuso, P. Lipka-Bartosik, M. Aifer, P. J. Coles, and M. Perarnau-Llobet, Energy-time-accuracy tradeoffs in thermodynamic computing, arXiv preprint arXiv:2601.04358 (2026)

  78. [78]

    H. A. Kramers, Brownian motion in a field of force and the diffusion model of chemical reactions, physica7, 284 (1940)

  79. [79]

    B¨ uttiker, E

    M. B¨ uttiker, E. Harris, and R. Landauer, Thermal ac- tivation in extremely underdamped josephson-junction circuits, Physical Review B28, 1268 (1983)

  80. [80]

    Risken and K

    H. Risken and K. Voigtlaender, Eigenvalues and eigen- functions of the fokker-planck equation for the ex- tremely underdamped brownian motion in a double-well potential, Journal of statistical physics41, 825 (1985)

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.