REVIEW 3 major objections 5 minor 34 references
Ternary Stochastic Neuron -- Implemented with a Single Strained Magnetostrictive Nanomagnet
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A single strained nanomagnet can act as a ternary stochastic neuron.
desk verdict A plausible new nanomagnetic TSN mechanism undermined by an unexplained ~8x error in the stress-energy accounting; needs a serious referee and a substantial revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the stress-anisotropy energy E = -(3/2)λsσΩcos²θ together with the spin-transfer torque from the injected current. When λsσ<0, the energy minimum sits at θ=90° (magnetization along x, my=0), holding the neuron in the zero state until the current's torque pulls it toward ±y; the dynamics are simulated with the stochastic LLG equation including Slonczewski and field-like torques (relative weights A=1, B=0.3) and thermal noise.
What would settle it
A micromagnetic simulation of the same 100 nm diameter, 2 nm thick FeGa disk under -80 MPa uniaxial compressive stress and spin-polarized current, without the macrospin assumption, would settle whether the zero-current plateau survives in realistic nonuniform magnetization dynamics.
Extended reading notes
Core claim
The central claim is that a zero-energy-barrier (shape-isotropic) magnetostrictive nanomagnet under uniaxial compressive stress — the sign that makes λsσ negative for FeGa — gives a three-level activation function <my(t)> versus spin-polarized current Is, with a plateau around Is=0 that constitutes the stable 0 state of a ternary stochastic neuron. The same device under tensile stress (positive λsσ) instead pins the magnetization near ±y and produces an activation curve that depends on the initial state, which the paper shows is not useful for a TSN. The plateau width grows with stress magnitude, and the resulting activation function also acts as the threshold-based ternary function used in ternary neural networks.
Load-bearing premise
The prediction depends on the 100 nm disk behaving as a single macrospin and on treating the gate-induced biaxial strain as a stronger uniaxial strain; if either approximation is wrong, the plateau could shift, narrow, or disappear.
Editorial extensions
If this is right
- A TSN built this way occupies the same chip area as a binary stochastic neuron but encodes three states, increasing information density.
- Stress magnitude tunes the plateau width, giving a voltage-controlled window for the stable 0 state.
- The activation function also implements threshold-based ternary functions (Eq. 5), the building block for ternary neural networks that minimize distance between full-precision and ternary weights.
- Because the piezoelectric gate is a charged capacitor at steady state, holding the strain consumes no standby power, and a lattice-mismatched substrate could supply the strain without any voltage, at the cost of reconfigurability.
Reading between the lines
- Relaxing the macrospin assumption in a micromagnetic simulation could blur the plateau, since strain-induced fields vary across the disk; this is a concrete test of whether the TSN works outside the single-domain idealization.
- Ensemble stress non-uniformity will spread plateau widths across many neurons; the paper argues TSN function survives, but the effect on network-level training convergence is left open.
- The same strain-anisotropy trick could generalize to higher-radix stochastic neurons by engineering multiple energy minima, though the paper does not pursue that extension.
- The positive-λsσ activation curve resembles asymmetric neural-network activations such as ReLU-family functions, suggesting a possible separate use for strained nanomagnets as nonlinear transfer elements.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a design for a ternary stochastic neuron (TSN) based on a single circular magnetostrictive (FeGa) nanomagnet with zero in-plane shape anisotropy, subjected to uniaxial strain and injected with a spin-polarized current. The authors carry out stochastic Landau-Lifshitz-Gilbert (LLG) simulations and find that when the product of magnetostriction and stress (λ_s σ) is negative, the time-averaged y-component of magnetization <my(t)> versus current Is exhibits a plateau around Is=0, giving three stable states -1, 0, +1. For λ_s σ positive, the activation curve is asymmetric and initial-condition dependent. The paper concludes that this is the first nanomagnetic implementation of a TSN. The central claim is therefore that the strain-induced anisotropy creates a potential well along the x-axis, producing the zero-output plateau needed for ternary behavior.
Significance. If the simulation results are quantitatively reliable, the proposed device would be a simple, compact building block for ternary stochastic neural networks and related probabilistic computing architectures, with potential advantages in area and energy efficiency. The study uses a standard LLG model with material parameters taken from the literature (α=0.017, Ms=1.32×106 A/m, λ_s=266.6 ppm) and does not fit any parameter to a target activation function; the three-state behavior emerges from the physics. The paper also provides a clear qualitative explanation of the plateau mechanism in terms of stress-induced anisotropy energy. However, the quantitative inconsistency in the stress-energy accounting (Section 4.2 versus Eq. (4) and Fig. 3) and the unquantified biaxial-to-uniaxial approximation mean that the specific predicted plateau widths are not reproducible as written; these issues must be resolved before the results can be trusted.
major comments (3)
- [Section 4.2, Eq. (4), Fig. 3] The stress-anisotropy energy values reported in Section 4.2 are internally inconsistent with Eq. (4) and with the stated geometry and stress. For the d=100 nm, t=2 nm FeGa disk (volume Ω=1.57×10^-23 m^3), λ_s=266.6 ppm, and σ=50 MPa, Eq. (4) gives E=(3/2)λ_s σ Ω = 3.14×10^-19 J ≈ 76 kT at T=300 K, not the quoted 9.85 kT. For 100 MPa the correct value is about 152 kT, not 19.7 kT. The quoted values correspond to σ ≈ 6.5 and 13 MPa, roughly a factor of 7.7 smaller. Moreover, Section 4.2 refers to '50 MPa and 100 MPa' while Fig. 3 uses 20, 40, and 80 MPa. If Eq. (4) is correct, even 20 MPa gives a barrier of about 30 kT, which should be sufficient to produce a plateau on the 1 µs simulation time scale; the absence of a plateau at 20 MPa in Fig. 3 is then inconsistent with the stated equations. Since the plateau width is the defining property of the TSN activation function, this factor-of-eight energy-scale discrepancy makes the central simulation result quantitatively untrustworthy. The authors must correct the equations, the geometric/material parameters, or the stress labels, and ideally provide a version of Fig. 3 with the computed energy barriers for each stress value.
- [Section 3, biaxial-to-uniaxial approximation] The paper approximates the biaxial strain generated by the piezoelectric gate as a uniaxial strain along the y-axis with a 'larger' magnitude, but it never quantifies this replacement. For a biaxial stress state with σ_xx = -σ_yy, the magnetoelastic energy has the form E = -(3/2)λ_s(σ_xx cos²θ + σ_yy sin²θ)Ω, which is not equivalent to a simple uniaxial term E = -(3/2)λ_s σ_eff cos²θ with an unspecified σ_eff. The effective uniaxial constant depends on the ratio of the two stress components and on the assumed energy expression, and the difference is a factor of order 2 in energy. Because the plateau width is set by the barrier height, this unquantified approximation directly affects the central quantitative result. The authors should either specify the effective uniaxial stress value used in Eq. (3) and how it was derived from the biaxial strain, or implement the full biaxial energy in the simulations.
- [Section 3 and Section 6] The paper repeatedly claims to have 'implemented' a TSN and calls this 'the first and only nanomagnetic implementation of a TSN.' However, the manuscript presents only stochastic LLG simulations; no device is fabricated or measured. The word 'implementation' is thus an overstatement that misrepresents the contribution as an experimental realization. The authors should consistently describe their work as a simulation-based proposal or design, and moderate corresponding novelty claims in the abstract, Section 6, and the title if needed.
minor comments (5)
- [Section 4.1 and 4.2 headings] The headings 'Positive λsσ product or compressive stress' and 'Negative λsσ product or tensile stress' are reversed with respect to the standard relation: for FeGa with positive λ_s, compressive stress (σ<0) gives a negative λ_s σ product and tensile stress (σ>0) gives a positive product. The text within the sections is correct, but the headings should be swapped to avoid confusion.
- [Section 4.2] The sentence 'For the two stress values considered here, 50 MPa and 100 MPa' is not consistent with Fig. 3, which uses 20, 40, and 80 MPa. Please correct the stress values cited in the energy-barrier discussion.
- [References] Reference [14] lists the year as '2027' for 'Trained ternary quantization'; this appears to be a typo (likely 2017 or 2018). Please correct.
- [Abstract and text formatting] The abstract and some sections contain formatting issues such as 'CIF AR-10' (with a space) and '10 7' (instead of 10^7) for the sample count. Please fix these.
- [Section 6] The sentence 'We point out that that the contribution here is not just with respect to the activation function' contains a doubled 'that'. Please correct.
Circularity Check
No significant circularity; the activation plateau follows from a standard stochastic LLG simulation with literature parameters and no fit to the target curve.
full rationale
The paper's central claim is that a compressive-strained magnetostrictive nanomagnet, modeled by the stochastic Landau-Lifshitz-Gilbert equation, produces an activation function <my(t)> versus Is with a plateau around Is=0. This is a direct simulation result, not a quantity defined in terms of the claimed outcome. The stress-anisotropy energy Eq. (4) and the stress field in Eq. (3) are standard magnetoelastic expressions; the material parameters (Ms, lambda_s, alpha) are taken from the literature, and no parameter is fitted to the target activation function. The plateau is a physical consequence of the negative lambda_s-sigma product and emerges from the dynamics rather than being imposed. The paper does cite prior work by the same group for the numerical solver [21,22] and for the stress-field expression [26], but these citations supply simulation methodology or standard model inputs, not the conclusion itself. There is no self-citation chain that forces the plateau, no imported uniqueness theorem, and no ansatz smuggled in as external fact. The internal numerical discrepancy in Section 4.2 between the quoted stress-energy values (9.85 and 19.7 kT) and what Eq. (4) yields with the stated geometry and stress (roughly 76 and 152 kT) is a quantitative correctness concern, not a circularity: the derivation chain still runs from the stated model equations to the simulated curve. Overall, the derivation is self-contained against external benchmarks, so the circularity score is low; the minor self-citations are not load-bearing.
Assumptions & free parameters
free parameters (6)
- Gilbert damping alpha =
0.017 (from Ref [24])
- Saturation magnetization Ms =
1.32e6 A/m (from Ref [24])
- Magnetostriction coefficient lambda_s =
266.6 ppm (from Ref [24])
- Spin polarization fraction eta =
0.5
- Torque coefficients A and B =
A=1, B=0.3
- Applied stress magnitude =
20-80 MPa compressive and tensile
assumptions (4)
- domain assumption The 100 nm diameter, 2 nm thick FeGa nanomagnet is monodomain, so the macrospin approximation holds.
- standard math The stochastic Landau-Lifshitz-Gilbert equation with Gaussian white thermal noise and the given Slonczewski and field-like torques describes the magnetization dynamics.
- ad hoc to paper Biaxial strain generated by the piezoelectric gate can be approximated as uniaxial strain along the y-axis with a larger effective magnitude.
- domain assumption The desired TSN activation function requires a plateau near zero input, as shown in Fig. 1(b) and motivated by Ref [20].
Cite this review
Pith. "Pith review of Ternary Stochastic Neuron -- Implemented with a Single Strained Magnetostrictive Nanomagnet." pith.science (2026). https://pith.science/paper/T6YKWJ27
@misc{pith2026241204246,
author = {Pith},
title = {Pith review of: Ternary Stochastic Neuron -- Implemented with a Single Strained Magnetostrictive Nanomagnet},
year = {2026},
howpublished = {\url{https://pith.science/paper/T6YKWJ27}},
note = {Machine review of arXiv:2412.04246}
}
read the original abstract
Stochastic neurons are extremely efficient hardware for solving a large class of problems and usually come in two varieties -- "binary" where the neuronal statevaries randomly between two values of -1, +1 and "analog" where the neuronal state can randomly assume any value between -1 and +1. Both have their uses in neuromorphic computing and both can be implemented with low- or zero-energy-barrier nanomagnets whose random magnetization orientations in the presence of thermal noise encode the binary or analog state variables. In between these two classes is n-ary stochastic neurons, mainly ternary stochastic neurons (TSN) whose state randomly assumes one of three values (-1, 0, +1), which have proved to be efficient in pattern classification tasks such as recognizing handwritten digits from the MNIST data set or patterns from the CIFAR-10 data set. Here, we show how to implement a TSN with a zero-energy-barrier (shape isotropic) magnetostrictive nanomagnet subjected to uniaxial strain.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Rastegari M, Ordonez V, Redmon J. and Farhadi A 2016 Xnor-net: Imagenet classification using binary convolutional neural networks, European Conference on Computer Vision, Lecture Notes in Computer Science, vol 9908. Springer, Cham. DOI:10 .1007/978 − 3 − 319 − 46493 − 0 − 32
work page 2016
-
[2]
2016 Convolutional networks for fast, energy-efficient neuromorphic computing, Proc
Esser S K, et al. 2016 Convolutional networks for fast, energy-efficient neuromorphic computing, Proc. Nat. Acad. Sci. , 113, pp. 11441–11446
work page 2016
-
[3]
(published in 2016 International Conference on Learning Representations)
Han S, Mao H and Dally W J 2015 Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding, arXiv 1510.00149. (published in 2016 International Conference on Learning Representations)
arXiv 2015
-
[4]
Courbariaux M, Bengio Y and David J P 2015 Binaryconnect: Training deep neural networks with binary weights during propagations, in Advances in Neural Information Processing Systems , Association for Computing Machinery: New York, NY, USA, pp. pp. 3123–3131
work page 2015
-
[5]
Hubara I, Soudry D and Yaniv R E 2016 Binarized neural networks, in Advances in Neural Information Processing Systems, Association for Computing Machinery: New York, NY, USA,
work page 2016
-
[6]
Lin Z, Courbariaux M, Memisevic R and Y Bengio Y 2015 Neural networks with few multiplications, arXiv:1510.03009 (published in 2015 International Conference on Learning Representations)
arXiv 2015
-
[7]
Iandola F N, Moskewicz M W, Ashraf K, Han S, Dally W J and Keutzer K 2016 Squeezenet: Alexnet-level accuracy with 50 × fewer parameters and < 1mb model size, arXiv:1602.07360
arXiv 2016
-
[8]
Howard A G, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M and Adam H 2017 Mobilenets: Efficient convolutional neural networks for mobile vision applications, CoRR, vol. abs/1704.04861
arXiv 2017
Show all 34 references
-
[9]
Zhang X, Zhou X, Lin M and Sun J 2018 Shufflenet: An extremely efficient convolutional neural network for mobile devices, in IEEE Conference on Computer Vision and Pattern Recognition, 2018
2018
-
[10]
Ternary Stochastic Neuron – Implemented with a Single Strained Magnetostrictive Nanomagnet13
Liu H, Simonyan K and Yang Y 2019 DARTS: differentiable architecture search, in International Conference on Learning Representation. Ternary Stochastic Neuron – Implemented with a Single Strained Magnetostrictive Nanomagnet13
2019
-
[11]
3065–3072
Wang X, Xue C, Yan J, Yang X, Hu Y and Sun K 2020 Mergenas: Merge operations into one for differentiable architecture search, in International Joint Conference on Artificial Intelligence, pp. 3065–3072
2020
-
[12]
Wang X, Lin J, Zhao J, Yang X and Yan J 2022 Eautodet: Efficient architecture search for object detection, in European Conference on Computer Vision (Springer Nature, Switzerland)
2022
-
[13]
See also 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, DOI:10.1109/ICASSP49357.2023.10094626
Li F, Liu B, Wang X, Zhang B and Yan J 2016 Ternary weight networks, arXiv:1605.04711v3. See also 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, DOI:10.1109/ICASSP49357.2023.10094626
2016 arXiv
-
[14]
Also in International Conference on Learning Representation
Zhu C, Han S, Mao H and Dally W J 2027 Trained ternary quantization, arXiv:1612.01064v3. Also in International Conference on Learning Representation
2027 arXiv
-
[15]
Deng L, Jiao P, Wu Z and Li G 2018 GXNOR-Net: Training deep neural networks with ternary weights and activations without full-precision memory under a unified discretization framework, Neural Networks, 100, 49-58
2018
-
[16]
Zahoor F, Jaber R A, Isyaku U B, Sharma T, Bashir F, Abbas H, Alzahrani A S, Gupta S and Hanif M 2024 Design implementations of ternary logic systems: A critical review, Results in Engineering, 23, 102761
2024
-
[17]
Hassan O, Faria R, Camsari K Y, Sun J Z and Datta S 2019 Low barrier magnet design for efficient hardware binary stochastic neurons, IEEE Magn. Lett. , 10, 4502805
2019
-
[18]
Hassan O, Datta S and Camsari K Y 2021 Quantitative evaluation of hardware binary stochastic neuron, Phys. Rev. Appl. , 15, 064046
2021
-
[19]
Rahman R, Ganguly S and Bandyopadhyay S 2024 Reconfigurable stochastic neurons based on strain engineered low barrier nanomagnets, Nanotechnology, 35 , 325205
2024
-
[20]
Pitis S 2017 Beyond, binary, ternary and one-hot neurons, https : //r2rt.com/beyond − binary − ternary − and − one − hot − neurons.html
2017
-
[21]
Rahman R and Bandyopadhyay S 2023 Increasing flips per second and speed of p-computers by using dilute magnetic semiconductors, IEEE Magnetics Letters , 14, 4500604
2023
-
[22]
Rahman R and Bandyopadhyay S 2023 The strong sensitivity of the characteristics of binary stochastic neurons employing low barrier nanomagnets to small geometrical variations, IEEE Transactions on Nanotechnology, 22, 112
2023
-
[23]
Cui J Z, Liang C Y, Paisley E A, Sepulveda A, Ihlefeld J F, Carman G P and Lynch C S 2015 Generation of localized strain in a thin film piezoelectric to control individual magnetoelectric heterostructures, Appl. Phys. Lett. 107, 092903
2015
-
[24]
Bhattacharya D, Al-Rashid M A, D’Souza N, Bandyopadhyay S and Atulasimha J 2017 Incoherent magnetization dynamics in strain mediated switching of magnetostrictive nanomagnets, Nanotechnology 28, 015202
2017
-
[25]
Roy K, Bandyopadhyay S and Atulasimha J 2012 Metastable state in a shape-anisotropic single- domain nanomagnet subjected to spin-transfer-torque, Appl. Phys. Lett. , 101, 162405
2012
-
[26]
Salehi Fashami M, Roy K, Atulasimha J and Bandyopadhyay S 2011 Magnetization dynamics, Bennett clocking and associated energy dissipation in multiferroic logic, Nanotechnology, 22, 155201
2011
-
[27]
Datta S, Atulasimha J, Mudivarthi C and Flatau A B 2010 Stress and magnetic field-dependent Young’s modulus in single crystal iron–gallium alloys, J. Magn. Magn. Mater. 322 2135-2144
2010
-
[28]
1097–1105
Krizhevsky A, Sutskever I and Hinton G E 2012 Imagenet classification with deep convolutional neural networks, in Advances in Neural Information Processing Systems , Association for Computing Machinery: New York, NY, USA, pp. 1097–1105
2012
-
[29]
Ramachandran P, Zoph B and Le Q V 2017 Searching for activation functions, arXiv:1710.05941
2017 arXiv
-
[30]
Misra D 2020 Mish: A self regularized non-monotonic neural activation function, arXiv:1908.08681
2020 arXiv
-
[31]
Chai E, Yu W, Cui T, Ren J and Ding S 2022 An efficient asymmetric nonlinmear activation function for deep neural networks, Symmetry 14 1027
2022
-
[32]
Bandyopadhyay S 2022 Magnetic Straintronics: An energy-efficient hardware paradigm for digital and analog information processing, Synthesis Lectures on Engineering, Science and Technology, Ternary Stochastic Neuron – Implemented with a Single Strained Magnetostrictive Nanomagn...
2022
-
[33]
Bandyopadhyay S, Atulasimha J and Barman A 2021 Magnetic straintronics: Manipulating the magnetization of magnetostrictive nanomagnets with strain for energy-efficient applications, Appl. Phys. Rev. , 8, 041323
2021
-
[34]
2016 Nanomagnetic and Spintronic Devices for Energy- Efficient Memory and Computing , John Wiley & Sons, Chichester, West Sussex, UK
Atulasimha J and Bandyopadhyay S Eds. 2016 Nanomagnetic and Spintronic Devices for Energy- Efficient Memory and Computing , John Wiley & Sons, Chichester, West Sussex, UK
2016
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.