Pith. sign in

REVIEW 2 major objections 2 minor 1 cited by

Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics

T0 review · 2 major / 2 minor · reviewed 2026-05-24 · grok-4.3

Pith's one-line read Deep learning methods, both hybrid and pure, are reviewed for use in solid and fluid mechanics simulations.

desk verdict This review builds DL concepts from basics for mechanics readers and flags some AI misconceptions, but its value as state-of-the-art coverage rests on whether the citations are representative. read the letter →

arxiv 2212.08989 v3 pith:CAVQMFLS submitted 2022-12-18 cs.LG

classification cs.LG
keywords deeplearningcomputationalmechanicsphysics-informedneuralnetworkshybridmethodsLSTMfiniteelementmethodmodelorderreductionconstitutivemodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes a detailed survey of how artificial neural networks and deep learning are applied to computational mechanics problems involving solids, fluids, and finite-element technology. It distinguishes hybrid approaches that combine traditional PDE discretizations with machine learning from pure machine learning methods such as physics-informed neural networks. The review builds DL concepts from the basics for readers already familiar with mechanics, while also covering LSTM architectures, attention mechanisms, optimizers, and kernel methods like Gaussian processes. A sympathetic reader would care because the survey aims to bring newcomers quickly to the research frontier and to correct misconceptions found even in well-known references on the history and limits of AI. The positioning and control of a large-deformable beam serves as a concrete example throughout.

What carries the argument

Hybrid methods that augment traditional PDE discretizations with ML and pure ML methods such as physics-informed neural networks, with LSTM for constitutive modeling and model reduction and attention for discontinuities.

What would settle it

Discovery of a substantial number of peer-reviewed works on deep learning for finite-element or continuum mechanics problems that are omitted from the review would indicate the coverage is incomplete.

Watch

Extended reading notes

Core claim

The paper claims that recent deep learning developments relevant to computational mechanics can be organized into hybrid methods, which use LSTM networks to model nonlinear constitutive relations or reduce model order and convolutional networks to accelerate traditional integrators, and pure ML methods represented by physics-informed neural networks that may incorporate attention to handle discontinuous solutions; it further reviews LSTM and attention architectures along with stochastic optimizers and kernel machines to sufficient depth for advanced follow-on work.

Load-bearing premise

The chosen papers and methods accurately represent the current state of the art without significant selection bias or major omissions.

Editorial extensions

If this is right

  • Hybrid LSTM-based methods can capture complex nonlinear material behavior within existing finite-element frameworks.
  • Model-order reduction via LSTM can make turbulence simulations more efficient.
  • Convolutional networks can speed up specific steps inside conventional time-integration schemes.
  • PINNs, possibly augmented with attention, can solve nonlinear PDEs directly without traditional discretization.
  • Kernel machines including Gaussian processes provide a foundation for understanding infinite-width shallow networks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The review structure could serve as a template for similar surveys in related fields such as structural optimization or multiphysics coupling.
  • Explicit discussion of limitations in the classics may encourage more careful citation practices when referencing early AI work in engineering contexts.
  • The beam-positioning example suggests that the reviewed techniques are already close to practical control applications in deformable-body dynamics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript is a review paper surveying recent deep learning applications to computational mechanics. It covers hybrid methods that combine traditional PDE discretizations with LSTM (for constitutive modeling and model-order reduction) and CNN (for simulation acceleration), pure ML approaches such as PINNs with attention mechanisms for discontinuous solutions, reviews of LSTM/attention architectures, modern optimizers, and kernel machines (including Gaussian processes and infinite-width networks), plus discussion of AI history, limitations, and misconceptions. An example application to positioning/pointing control of a large-deformable beam is included. The target audience is computational-mechanics experts new to DL, with concepts built from the basics.

Significance. If the literature selection is representative and the coverage balanced, the review would provide a useful on-ramp for mechanics researchers entering DL, explicitly contrasting hybrid and pure-ML strategies and correcting common misconceptions about the classics. The inclusion of both modern architectures and kernel-machine background for advanced readers adds pedagogical value.

major comments (2)
  1. [Abstract] Abstract and opening sections: the central claim that the paper reviews 'many recent developments ... in detail' and supplies the 'state of the art' rests on the assumption of unbiased, comprehensive paper selection up to the 2022 cutoff. No explicit selection methodology, inclusion/exclusion criteria, or discussion of potential gaps (e.g., key LSTM turbulence papers or additional PINN variants) is provided, making it impossible to verify representativeness.
  2. [Introduction (implied by abstract)] The positioning statement that the review brings 'first-time learners quickly to the forefront of research' is load-bearing for the intended contribution, yet the manuscript does not compare its scope or depth against existing surveys in the same area, leaving the incremental value of this particular synthesis unclear.
minor comments (2)
  1. [Abstract] The three motivating AI breakthroughs cited in the abstract are not enumerated explicitly; listing them would strengthen the opening motivation.
  2. Ensure that every cited work is dated no later than the stated 2022 cutoff and that references to the 'classics' are accompanied by the specific misstatements being corrected.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments. We address each major comment below, agreeing that additional clarifications on scope and comparisons to prior surveys will strengthen the manuscript.

read point-by-point responses
  1. Referee: [Abstract] Abstract and opening sections: the central claim that the paper reviews 'many recent developments ... in detail' and supplies the 'state of the art' rests on the assumption of unbiased, comprehensive paper selection up to the 2022 cutoff. No explicit selection methodology, inclusion/exclusion criteria, or discussion of potential gaps (e.g., key LSTM turbulence papers or additional PINN variants) is provided, making it impossible to verify representativeness.

    Authors: We agree that an explicit discussion of literature selection would improve transparency. Although the review was compiled based on relevance to computational mechanics applications up to the 2022 cutoff, we will add a new paragraph in the Introduction describing the general search approach, inclusion focus on solid/fluid mechanics and finite-element contexts, and explicit acknowledgment of potential gaps (e.g., certain turbulence LSTM works or post-cutoff PINN variants). revision: yes

  2. Referee: [Introduction (implied by abstract)] The positioning statement that the review brings 'first-time learners quickly to the forefront of research' is load-bearing for the intended contribution, yet the manuscript does not compare its scope or depth against existing surveys in the same area, leaving the incremental value of this particular synthesis unclear.

    Authors: The manuscript's distinctive elements include the joint treatment of hybrid LSTM/CNN methods with pure PINN approaches, coverage of kernel machines and infinite-width networks, and discussion of AI history with corrections to common misconceptions. We nevertheless recognize the benefit of explicit positioning. We will revise the Introduction to include a concise comparison with related surveys (e.g., those focused primarily on PINNs or data-driven constitutive modeling) and to articulate the incremental synthesis provided here. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: review draws from external citations without internal derivations

full rationale

This is a literature review paper with no original mathematical derivations, predictions, or fitted models presented as results. The central content consists of summaries of external cited works on DL methods for mechanics (LSTM, PINN, etc.), built from basics for the reader. No steps match the enumerated circularity patterns, as there are no equations reducing to inputs by construction, no fitted parameters renamed as predictions, and no load-bearing self-citations that justify a uniqueness theorem or ansatz. The paper is self-contained as a survey against external benchmarks.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

As a review article the central content rests on the accuracy and completeness of the surveyed literature rather than new mathematical derivations or postulates.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics." pith.science (2026). https://pith.science/paper/CAVQMFLS

@misc{pith2026221208989,
  author       = {Pith},
  title        = {Pith review of: Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CAVQMFLS}},
  note         = {Machine review of arXiv:2212.08989}
}
read the original abstract

Three recent breakthroughs due to AI in arts and science serve as motivation: An award winning digital image, protein folding, fast matrix multiplication. Many recent developments in artificial neural networks, particularly deep learning (DL), applied and relevant to computational mechanics (solid, fluids, finite-element technology) are reviewed in detail. Both hybrid and pure machine learning (ML) methods are discussed. Hybrid methods combine traditional PDE discretizations with ML methods either (1) to help model complex nonlinear constitutive relations, (2) to nonlinearly reduce the model order for efficient simulation (turbulence), or (3) to accelerate the simulation by predicting certain components in the traditional integration methods. Here, methods (1) and (2) relied on Long-Short-Term Memory (LSTM) architecture, with method (3) relying on convolutional neural networks. Pure ML methods to solve (nonlinear) PDEs are represented by Physics-Informed Neural network (PINN) methods, which could be combined with attention mechanism to address discontinuous solutions. Both LSTM and attention architectures, together with modern and generalized classic optimizers to include stochasticity for DL networks, are extensively reviewed. Kernel machines, including Gaussian processes, are provided to sufficient depth for more advanced works such as shallow networks with infinite width. Not only addressing experts, readers are assumed familiar with computational mechanics, but not with DL, whose concepts and applications are built up from the basics, aiming at bringing first-time learners quickly to the forefront of research. History and limitations of AI are recounted and discussed, with particular attention at pointing out misstatements or misconceptions of the classics, even in well-known references. Positioning and pointing control of a large-deformable beam is given as an example.

Figures

Figures reproduced from arXiv: 2212.08989 by the authors.

Figure 1
Figure 1. AI-generated image won contest in the category of Digital Arts, Emerging Artists, on 2022.08.29 (Section 1). “Théâtre D’opéra Spatial” (Space Opera Theater) by “Jason M. Allen via Midjourney”, which is “an artificial intelligence program that turns lines of text into hyper￾realistic graphics” [4]. Colorado State Fair, 2022 Fine Arts First, Second & Third. (Permission of Jason M. Allen, CEO, Incarnate Games) [PITH_F… view at source ↗
Figure 2
Figure 2. Breakthroughs in AI (Section 2). Left: The journal Science 2021 Breakthough of the Year. Protein folded 3-D shape produced by the AI software AlphaFold compared to experiment with high accuracy [5]. The AlphaFold Protein Structure Database contains more than 200 million protein structure predictions, a holy grail sought after in the last 50 years. Right: The AI solfware AlphaGo, a runner-up in the journal Science 20… view at source ↗
Figure 3
Figure 3. ImageNet competitions (Section 2). Top (smallest) classification error rate versus competition year. A sharp decrease in error rate in 2012 sparked a resurgence in AI interest and research [13]. By 2015, the top classification error rate surpassed human classification error rate of 5.1% with Parametric Rectified Linear Unit [61]; see Section 5.3.3 and also [62]. Figure from [63]. (Figure reproduced with permission o… view at source ↗
Figures from the paper (159 more)
Figure 4
Figure 4. Figure 4: Handwritten equation 1 (Section 2.1) into this LaTeX code “p \times q = m \Rightarrow p = \frac { m } { q }” to yield the equation image: p × q = m ⇒ p = m q (1) Another example is the hand-written multiplication work below by the same pupil [PITH_FULL_IMAGE:figures/f…
Figure 5
Figure 5. Figure 5: Handwritten equation 2 (Section 2.1). Hand-written multiplication work of an eleven-year old pupil. 23“The World Health Organization declares COVID-19 a pandemic” on 2020 Mar 11, CDC Museum COVID-19 Timeline, Internet archive 2022.06.02. 24Krisher T., Teslas with Autop…
Figure 6
Figure 6. Figure 6: Artificial intelligence and subfields (Section 2.2). Three classes of methods— Artificial Intelligence (AI), Machine Learning (ML), and Deep Learning (DL)—and their rela￾tionship, with an example of method in each class. A knowledge-base method is an AI method, but is …
Figure 7
Figure 7. Figure 7: Feedforward neural network (Section 2.3.1). A feedforward neural network in [38], rotated clockwise by 90 degrees to compare to its equivalent in [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: Artificial neuron (Section 2.3.1). A neuron with its multiple inputs O p−1 i (which are outputs from the previous layer (p−1), and thus the variable name “O”), processing operations (multiply inputs with network weights w p−1 ji , sum weighted inputs, add bias θ p j , …
Figure 9
Figure 9. Figure 9: Cube and distorted cube elements (Section 2.3.1). Regular and distorted linear hexahedral elements [38]. (Figure reproduced with permission of the authors.) prescribed accuracy, and (2) corrections to the quadrature weights by trying one million randomly generated sets…
Figure 10
Figure 10. Figure 10: Distributions of error ratios defined by Eq. (21), when using correction factors estimated by deep learning. 3.5.2.3. Application phase [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: Dual-porosity single-permeability medium (Section 2.3.2). Left: Actual reservoir. Dual (or double) porosity indicates the presence of two types of porosity in naturally-fractured reservoirs (e.g., of oil): (1) Primary porosity in the matrix (e.g., voids in sands) with…
Figure 12
Figure 12. Figure 12: Pore structure of Majella limestone, dual porosity (Section 2.3.2), a carbonate rock with high total porisity at 30%. Backscattered SEM images of Majella limestone: (a)- (c) sequence of zoomed-ins; (d) zoomed-out. (a) The larger macropores (dark areas) have dimensions…
Figure 13
Figure 13. Figure 13: Majella limestone, nonlinear stress-strain relations (Section 2.3.2). Differential stress (i.e., the difference between the largest principal stress and the smallest one) vs axial strain (left) and vs volumetric strain (right) [90]. See Remark 11.7, Section 11.3.4, an…
Figure 7
Figure 7. Figure 7: Hierarchy of a multi-scale multi-physics poromechanics problem for fluid-infiltrating media. Black arrow represents a definition or a “universal principle”; red arrow represents either a phenomenological relation or an operator that is defined not based on first princi…
Figure 15
Figure 15. Figure 15: LSTM variant with “peephole” connections, block diagram (Sections 2.3.2, 7.2).43 Unlike the original LSTM unit (see Section 7.2), both the input gate and the forget gate in an LSTM unit with peephole connections receive the cell state as input. The above figure from W…
Figure 16
Figure 16. Figure 16: Coordination number CN (Section 2.3.2, 11.3.2). (a) Chemistry. Number of bonds to the central atom. Uranium borohydride U(BH4)4 has CN = 12 hydrogen bonds to uranium. (b, c) Photoelastic discs showing number of contact points (coordination number) on a particle. (b) R…
Figure 17
Figure 17. Figure 17: Network with LSTM and microstructure data (porosity ϕ, coordination number CN = Nc, [PITH_FULL_IMAGE:figures/full_fig_p028_17.png]
Figure 18
Figure 18. Figure 18: Reduced-order POD basis (Sections 2.3.3, 12.1). For each dataset (also Fig￾ure 116), which contained k snapshots, the full POD reconstruction of the flow-field dynamical quantity u(x, t), where x is a point in the 3-D flow field, consists of all k basis functions ϕi(x…
Figure 6
Figure 6. Figure 6: LSTM-ROM Methodology using the LSTM NN. An important assumption often made in ROM, including Galerkin-based ROM, is that the dominant POD modes for the training and test datasets are qualitatively similar [4]. For instance, flows within a narrow range of Reynolds numbe…
Figure 10
Figure 10. Figure 10: LSTM and BiLSTM predictions of Dominant POD ↵(t + t 0 ) for Isotropic turbulence test data [PITH_FULL_IMAGE:figures/full_fig_p031_10.png]
Figure 11
Figure 11. Figure 11: Mean Absolute Scaled Error (MASE) for LSTM predictions on all test samples in ISO dataset 5023 realizations. The results show that the MASE is generally low, except at samples where a sudden increase is observed. A similar trend is also observed for BiLSTM in [PITH_F…
Figure 22
Figure 22. Figure 22: Function mapping, graphical representation (Section 4.3.1): n inputs in x ∈ R n×1 (n × 1 column matrix of real numbers) are fed into function f to produce m outputs in y ∈ R m×1 . The multiple levels of compositions in Eq. (18) can then be represented by x = y (0) | {…
Figure 23
Figure 23. Figure 23: Feedforward network (Sections 4.3.1, 4.4.4): Multilevel composition in feedfor￾ward network with L layers represented as a sequential application of functions f (ℓ) , with ℓ = 1, · · · , L, to n inputs gathered in x = y (0) ∈ R n×1 (n × 1 column matrix of real num￾ber…
Figure 1.2
Figure 1.2. Figure 1.2: Figure1.2. See also Remark [PITH_FULL_IMAGE:figures/full_fig_p037_1_2.png]
Figure 24
Figure 24. Figure 24: Activation function (Section 4.4.2): Rectified linear function and its derivatives. See also Section 5.3.3 and [PITH_FULL_IMAGE:figures/full_fig_p040_24.png]
Figure 25
Figure 25. Figure 25: Current I versus voltage V (Section 4.4.2): Ideal diode, resistance, scaled rectified linear function as activation (transfer) function for the ideal diode and resistance in series. (Figure plotted with R = 2.) See also [PITH_FULL_IMAGE:figures/full_fig_p040_25.png]
Figure 26
Figure 26. Figure 26: Halfwave rectifier circuit (Section 4.4.2), with a primary alternative current z going in as input (left), passing through a transformer to lower the voltage amplitude, with the sec￾ondary alternative current out of the transformer being put through a closed circuit w…
Figure 27
Figure 27. Figure 27: FI curves (Sections 4.4.2, 13.2.2). Firing rate frequency (F) versus applied depo￾larizing current (I), thus FI curves. Three types of FI curves. The time histories of voltage Vm provide a visualization of the spikes, current threshold, and spike firing rates. The app…
Figure 28
Figure 28. Figure 28: FI or FV curves (Sections 3, 4.4.2, 13.2.2). Neuron firing rate (F) versus input current (I) (FI curves, a,b,c) or voltage (V). The Integrate-and-Fire model in SubFigure (c) can be used to replace the sigmoid function to fit the experimental data points in SubFigure (…
Figure 29
Figure 29. Figure 29: Halfwave rectifier (Sections 4.4.2, 5.3.2). Current I versus voltage V [red line in SubFigure (b)] in the halfwave rectifier circuit of [PITH_FULL_IMAGE:figures/full_fig_p044_29.png]
Figure 30
Figure 30. Figure 30: Logistic sigmoid function (Sections 4.4.2, 5.1.3, 5.3.1, 13.3.3): s(z) = [1 + exp(−z)]−1 = [tanh(z/2) + 1]/2 (red), with the tangent at the origin z = 0 (blue). See also Remark 5.3 and [PITH_FULL_IMAGE:figures/full_fig_p045_30.png]
Figure 31
Figure 31. Figure 31: Hyperbolic tangent function (Section 4.4.2): g(z) = tanh(z) = 2s(2z) − 1 (red) and its tangent g(z) = z at the coordinate origin (blue), showing that this activation function is identity for small signals. (2) Distributivity. Each feature of the data is represented di…
Figure 32
Figure 32. Figure 32: One-layer network (Section 4.4.3) representing the relation between the predicted output ye and the input x, i.e., ye = f(x) = a(W x + b) = a(z), with the weighted sum z := W x + b; see Eq. (26) and Eq. (35) with ℓ = 1. For a lower-level details of this one layer, see…
Figure 33
Figure 33. Figure 33: One-layer network (Section 4.4.3) in [PITH_FULL_IMAGE:figures/full_fig_p046_33.png]
Figure 35
Figure 35. Figure 35: Low-level details of layer (ℓ) (Sections 4.4.3, 4.4.4) of the multilayer neural net￾work in [PITH_FULL_IMAGE:figures/full_fig_p046_35.png]
Figure 36
Figure 36. Figure 36: Artificial neuron (Sections 2.3.1, 4.4.4, 13.1), row i in layer (ℓ) in [PITH_FULL_IMAGE:figures/full_fig_p046_36.png]
Figure 37
Figure 37. Figure 37: Representing XOR function (Sections 4.5, 13.2). This one-layer network (which is not the Rosenblatt perceptron in [PITH_FULL_IMAGE:figures/full_fig_p048_37.png]
Figure 38
Figure 38. Figure 38: Representing XOR function (Sections 4.5). This two-layer network can perform this task. The four points in the design matrix X = [x1, . . . , x4] ∈ R 2×4 (see [PITH_FULL_IMAGE:figures/full_fig_p049_38.png]
Figure 39
Figure 39. Figure 39: Two-layer network for XOR representation (Sections 4.5). Left: XOR function, with A = x (1) 1 = [0, 0]T , B = x (1) 2 = [0, 1]T , C = x (1) 3 = [1, 0]T , D = x (1) 4 = [1, 1]T ; see Eq. (52). The XOR value for the solid red dots is 1, and for the open blue dots 0. Rig…
Figure 40
Figure 40. Figure 40: Two-layer network for XOR representation (Sections 4.5). Left: Images of points A, B, C, D of Z(1) in Eq. (56), obtained after a translation by adding the bias b (1) = [0, −1]T in Eq. (51) to the same points A, B, C, D in the right subfigure of [PITH_FULL_IMAGE:figur…
Figure 41
Figure 41. Figure 41: Test accuracy versus network depth (Section 4.6.1), showing that test accuracy for this example increases monotonically with the network depth (number of layers). [78], p. 196. (Figure reproduced with permission of the authors.) But it is not clear where in [13] that …
Figure 42
Figure 42. Figure 42: Increasing network size over time (Section 4.6.1, 13.2). All networks before 2015 had their number of neurons smaller than that of a frog at 1.6 × 107 , and still far below that in a human brain at 8.6 × 1010; see “List of animals by number of neurons”, Wikipedia, ver…
Figure 43
Figure 43. Figure 43: Training/test error vs. iterations, depth (Sections 4.6.2, 6). The training error and test error of deep fully-connected networks increased when the number of layers (depth) increased [127]. (Figure reproduced with permission of the authors.) [PITH_FULL_IMAGE:figures…
Figure 44
Figure 44. Figure 44: Residual network (Sections 4.6.2, 6), basic building block having two layers with the rectified linear activation function (ReLU), for which the input is x, the output is H(x) = F(x) + x, where the internal mapping function F(x) = H(x) − x is called the residual. Chai…
Figure 45
Figure 45. Figure 45: Full residual network (Sections 4.6.2, 6) with 34 layers, made up from 16 building blocks with two layers each ( [PITH_FULL_IMAGE:figures/full_fig_p057_45.png]
Figure 46
Figure 46. Figure 46: Sofmax function for two classes, logistic sigmoid (Section 5.1.3, 5.3.1): s(z) = [1 + exp(−z)]−1 and s(−z) = [1 + exp(z)]−1 , such that s(z) + s(−z) = 17. See also [PITH_FULL_IMAGE:figures/full_fig_p061_46.png]
Figure 47
Figure 47. Figure 47: Backpropagation building block, typical layer (ℓ) (Section 5.2, Algorithm 1, Ap￾pendix 1). The forward propagation path is shown in blue, with the backpropagation path in red. The update of the parameters θ (ℓ) in layer (ℓ) is done as soon as the gradient ∂J/∂θ (ℓ) is…
Figure 48
Figure 48. Figure 48: Backpropagation in fully-connected network (Section 5.2, 5.3, Algorithm 1, Ap￾pendix 1). Starting from the predicted output ye = y (L) In the last layer (L) at the end of any forward propagation (blue arrows), and going backward (red arrows) to the first layer with ℓ …
Figure 49
Figure 49. Figure 49: Vanishing gradient problem (Section 5.3). Speed of learning of earlier layers is much slower than that of later layers. Here, after 400 epochs of training, the speed of learning of Layer (1) at 10−5 (blue line) is 100 times slower than that of Layer (4) at 10−3 (green…
Figure 50
Figure 50. Figure 50: Neural network with four layers (Section 5.3), one neuron per layer, scalar input x, scalar output y, cost function J(θ) = 1 2 (y − ye) 2 , with ye = y (4) being the target output and also the output of layer (4), such that f (ℓ) (y (ℓ−1)) = a(z (ℓ) ), with a(·) being…
Figure 51
Figure 51. Figure 51: Neural network with four layers in [PITH_FULL_IMAGE:figures/full_fig_p068_51.png]
Figure 52
Figure 52. Figure 52: Successive multiplications of these derivatives will result in smaller and smaller values along the back propagation path. If the weights w (ℓ) in Eq. (110) are also smaller than 1, then the gradient ∂J/∂b(1) will tend toward 0, i.e., vanish. The problem is further ex…
Figure 53
Figure 53. Figure 53: Cost-function cliff (Section 5.3.1). A cliff, or a sharp drop in the cost function. The parameter space is represented by a weight w and a bias b. The slope at the brink of the cliff leads to large-magnitude gradients, which when multiplied with each other several tim…
Figure 54
Figure 54. Figure 54: Rectified Linear Unit (ReLU, left) and Parametric ReLU (right) (Section 5.3.2), in which the slope s is a parameter to optimize; see Section 5.3.3. See also [PITH_FULL_IMAGE:figures/full_fig_p071_54.png]
Figure 55
Figure 55. Figure 55: Cost-function landscape (Section 6). Residual network with 56 layers (ResNet-56) on the CIFAR-10 training set. Highly non-convex, with many local minima, and deep, narrow valleys [132]. The training error and test error for fully-connected network increased when the n…
Figure 56
Figure 56. Figure 56: Training set, validation set, test set (Section 6.1). Partition of whole dataset. The examples are independent. The three subsets are identically distributed. 6.1 Training set, validation set, test set, stopping criteria The classical (old) thinking—starting in 1992 w…
Figure 57
Figure 57. Figure 57: Training and validation learning curves—Classical viewpoint (Section 6.1), i.e., plots of training error and validation errors versus epoch number (time). While the training cost decreased continuously, the validation cost reaches a minimum around epoch 20, then start…
Figure 58
Figure 58. Figure 58: Validation learning curve (Section 6.1, Algorithm 4). Validation error vs epoch number. Some validation error could oscillate wildly around the mean, resulting in an “ugly reality”. The global minimum validation error corresponded to epoch number τ ⋆ . Since the stopp…
Figure 59
Figure 59. Figure 59: Bias-variance trade-off (Section 6.1). Training error (cost) and test error versus model capacity. Two ways to change the model capacity: (1) change the number of network parameters, (2) change the values of these parameters (weight decay). The generalization gap is t…
Figure 60
Figure 60. Figure 60: Modern interpolation regime (Sections 6.1, 14.2). Beyond the interpolation thresh￾old, the test error goes down as the model capacity (e.g., number of parameters) increases, describing the observation that networks with high capacity beyond the interpolation threshold…
Figure 61
Figure 61. Figure 61: Empirical test error vs Number of paramesters (Sections 6.1, 14.2). Experiments using the MNIST handwritten digit database in [137] confirmed the modern interpolation regime in [PITH_FULL_IMAGE:figures/full_fig_p077_61.png]
Figure 62
Figure 62. Figure 62: Inexact line search, Goldstein’s rule (Section 6.2.4). acceptable step lengths would be such that a decrease in the cost function J, denoted by ∆J in Eq. (124), falls into an acceptable sector formed by an upper-bound line and a lower-bound line. the upper bound is gi…
Figure 63
Figure 63. Figure 63: SGD with momentum, small heavy sphere Section 6.3.2. The descent direction (negative gradient, black arrows) bounces back and forth between the steep slopes of a deep and narrow valley. The small-heavy-sphere method, or SGD with momentum, follows a faster descent (red…
Figure 64
Figure 64. Figure 64: Optimal minibatch size vs. training-set size (Section 6.3.5). For a given training￾set size, the smallest minibatch size that achieves the highest accuracy is optimal. Left figure: The optimal mimibatch size was moving to the right with increasing training-set size M.…
Figure 65
Figure 65. Figure 65: Minibatch-size increase vs. step-length decay, training schedules (Section 6.3.5). Left figure: Step length (learning rate) vs. number of epochs. Right figure: Minibatch size vs. number of epochs. Three learning-rate schedules167 were used for training: (1) The step l…
Figure 66
Figure 66. Figure 66: Minibatch-size increase, fewer parameter updates, faster comutation (Sec￾tion 6.3.5). For each of the three training schedules in [PITH_FULL_IMAGE:figures/full_fig_p098_66.png]
Figure 67
Figure 67. Figure 67: Weight decay (Section 6.3.6). Effects of magnitude of weight-decay parameter d. Adapted from [78], p. 116. (Figure reproduced with permission of the authors.) 6.3.7 Combining all add-on tricks To have a general parameter-update equation that combines all of the above …
Figure 68
Figure 68. Figure 68: Convergence of adaptive learning-rate algorithms (Section 6.3.2): AdaGrad, RM￾SProp, SGDNesterov, AdaDelta, Adam [170]. (Figure reproduced with permission of the authors.) 6.5.2 AdaGrad: Adaptive Gradient Starting the line of research on adaptive learning-rate algorit…
Figure 69
Figure 69. Figure 69: Dow Jones Industrial Average (DJIA, Section 6.5.3) stock index year-to-date (YTD) chart as from 2019.01.01 to 2019.11.30, Google Finance. “Exponential smoothing methods have been around since the 1950s, and are still the most popular fore￾casting methods used in busin…
Figure 70
Figure 70. Figure 70: Saudi Arabia oil production during 1996-2013 (Section 6.5.3). Piecewise linear data (black) and fitted curve (red), despite the name “smoothing”. From [207], Chap. 7. (Figure reproduced with permission of the authors.) For neural networks, early use of exponential smo…
Figure 71
Figure 71. Figure 71: AMSGrad vs Adam, numerical examples (Sections 6.1, 6.5.7). The MNIST dataset is used. The first two figures on the left were the results of using logistic regression (network with one layer with logistic sigmoid activation function), whereas the figure on the right is…
Figure 72
Figure 72. Figure 72: Overfitting (Section 6.5.9, 6.5.10). Left: Underfitting with 1st-order polynomial. Middle: Appropriate fitting with 2nd-order polynomial. Right: Overfitting with 9th-order poly￾nomial. See [78], p. 110, Figure5.2. (Figure reproduced with permission of the authors.) 6.…
Figure 73
Figure 73. Figure 73: Standard SGD and SGD with momentum vs AdaGrad, RMSProp, Adam on CIFAR￾10 dataset (Sections 6.1, 6.3.2, 6.5.9). From [55], where a method for step-size tuning and step-size decaying was proposed to achieve lowest training error and generalization (test) error for both …
Figure 74
Figure 74. Figure 74: AdamW vs Adam, SGD, and variants on CIFAR-10 dataset (Sections 6.1, 6.5.10). While AdamW achieved lowest training loss (error) after 1800 epochs, the results showed that SGD with weight decay (SGDW) and with warm restart (SGDWR) achieved lower test (generalization) er…
Figure 75
Figure 75. Figure 75: Cosine annealing (Sections 6.3.4, 6.5.10). Annealing factor ak as a function of epoch number. Four annealing cycles p = 1, . . . , 4, with the following schedule for Tp in Eq. (154): (1) Cycle 1, T1 = 100 epochs, epoch 0 to epoch 100, (2) Cycle 2, T2 = 200 epochs, epo…
Figure 76
Figure 76. Figure 76: CIFAR-100 test loss using Resnet-34 and DenseNet-121 (Section 6.5.10). Compar￾ison between various optimizers, including Adam and AdamW, showing that SGD achieved the lowest global minimum loss (blue line) compared to all adaptive methods tested as shown [168]. See al…
Figure 77
Figure 77. Figure 77: SGD frequently outperformed all adaptive methods (Section 6.5.10). The table contains the global minimum for each optimizer, for each of the two datasets CIFAR-10 and CIFAR-100, using two different networks. For each network, an error percentage and the loss (cost) we…
Figure 78
Figure 78. Figure 78: Stochastic Newton with Armijo-like 2nd order line search (Section 6.7). IJCNN1 dataset from the LIBSVM library. Three batch sizes were used (1%, 5%, 100%) for both SGD and ALAS (stochastic Newton Algorithm 7). The exact gradient norm for each of these six cases was pl…
Figure 79
Figure 79. Figure 79: Folded and unfolded discrete RNN (Section 7.1, 13.2.2). Left: Folded discrete RNN at configuration (or state) number [k], where k is an integer, with input x [k] to a multilayer neural network f(·) = f (1) ◦ f (2) ◦ · · · ◦ f (L) (·) as in Eq. (18), having a feedback …
Figure 80
Figure 80. Figure 80: RNN with two multilayer neural networks (MLNs), (Section 7.1) denoted by f1(·) and f2(·), whose outputs are fed into the loss function for optimization. This RNN is a gener￾alization of the RNN in [PITH_FULL_IMAGE:figures/full_fig_p128_80.png]
Figure 81
Figure 81. Figure 81: Folded Recurrent Neural Network (RNN) with Long Short-Term Memory (LSTM) cell (Section 7.2, 11.3.3). The cell state at [k] is denoted by z [k] s ≡ c [k] . Two feedback loops, one for cell state zs and one for hidden state h, with one-step delay [k − 1]. The key unifie…
Figure 82
Figure 82. Figure 82: Unfolded RNN with LSTM cells (Sections 2.3.2, 7.2, 12.1): In this unfolded RNN, the cell states are centered at the LSTM cell [k = n], preceded by the LSTM cell [k = n − 1], and followed by the LSTM cell [k = n+ 1]. See Eq. (290) for the recurring relation among the s…
Figure 83
Figure 83. Figure 83: Folded RNN with Gated Recurrent Unit (GRU) (Section 7.3). The cell state at [k − 1], i.e., (x [k−1] , h [k−1]) are inputs to produce the hidden state h [k] . One feedback loop for the hidden state h, with one-step delay [k − 1]. The key unified recurring relation is F…
Figure 84
Figure 84. Figure 84: Scaled dot-product attention and multi-head attention (Section 7.4.3). Scaled-dot product attention (left) is the elementary building block of the Transformer model. It compares query vectors (Q) against a set of key vectors (K) to produce a context vector by weightin…
Figure 85
Figure 85. Figure 85: Transformer architecture (Section 7.4.3). The Transformer is a sequence-to￾sequence model without recurrent connections. Encoder and decoder are entirely built upon scaled dot-product attention. Items of source and target sequences are numerically represented as vecto…
Figure 86
Figure 86. Figure 86: Gaussian process priors (Section 8.3). Left: Two samples with Gaussian kernel. Right: Two samples with Laplacian kernel. Parameters for both kernels: Kernel precision (inverse of variance) γ = σ −2 = 0.2 in Eq. (358), isotropic noise variance ν 2I = 10−6I added to cov…
Figure 87
Figure 87. Figure 87: Gaussian process prior and posterior samplings, Gaussian kernel (Section 8.3). Top left: Gaussian-prior samples (Section 8.3.1). The shaded red zones represent the predic￾tive density of at each input location. Top right: Gaussian-posterior samples with 1 data point. …
Figure 88
Figure 88. Figure 88: Gaussian process posterior samplings, noise effects (Section 8.3). Not all sampled curves in [PITH_FULL_IMAGE:figures/full_fig_p152_88.png]
Figure 89
Figure 89. Figure 89: Gaussian process posterior samplings, animation (Section 8.3). Interactive Gaus￾sian Process Visualization, Infinite curiosity. Click on the plot area to specify data points. See Figures 87 and 88. DL-related software framework, see [PITH_FULL_IMAGE:figures/full_fig_…
Figure 90
Figure 90. Figure 90: Top deep-learning libraries in 2018 by the “Power Score” in [249]. By 2022, using Google Trends, the popularity of different frameworks is significantly different; see [PITH_FULL_IMAGE:figures/full_fig_p154_90.png]
Figure 91
Figure 91. Figure 91: Google Trends of deep-learning software libraries (Section 9). The chart shows the popularity of five DL-related software libraries most “powerful” in 2018 over the last 5 years (as of July 2022). See also [PITH_FULL_IMAGE:figures/full_fig_p155_91.png]
Figure 92
Figure 92. Figure 92: Positioning and pointing control of large deformable beam (Section 9, Remark 9.1). Reinforcement learning. The agent is trained to align the tip of the flexible beam with the target position (red ball). For this purpose, the agent can move the base of the cantilever; …
Figure 93
Figure 93. Figure 93: DL-frameworks in nonlinear finite-element problems (Section 9.4). The computa￾tional efficiency of a PyTorch-based (Version 1.8) finite-element code implemented was com￾pared against the state-of-the-art general purpose Netgen/NGSolve [265] for a problem of non￾linear…
Figure 94
Figure 94. Figure 94: Physics-Informed Neural Networks (PINN) concept (Section 9.5). The goal is to find the optimal network parameters θ ⋆ (weights) and PDE parameters λ ⋆ that minimize the total weighted loss function L(θ, λ), which is a linear combination of four loss functions: (1) The…
Figure 95
Figure 95. Figure 95: Coupled nonlinear hyperbolic equations (Section 9.5). Analytical solution, pre￾dicted solution by NeuralPDE [275] and error for the coupled nonlinear hyperbolic equations in Eq. (383). Additional PINN software packages other than those in [PITH_FULL_IMAGE:figures/ful…
Figure 7
Figure 7. Figure 7: Nodes A, B and D of any 8-noded element are shifted to x-y plane by translation (a) and rotations (b), (c) and (d). E (±rd, ±rd, 1 ± rd), F (1 ± rd, ±rd, 1 ± rd), G (1 ± rd, 1 ± rd, 1 ± rd), and H (±rd, 1 ± rd, 1 ± rd). Here, the maximum amount of change in the coordin…
Figure 97
Figure 97. Figure 97: Creation of randomly distorted elements (Section 10). Hexahedra forming the train￾ing and validation sets are created by randomly displacing the nodes of a regular hexahedral. To comply with the normalization procedure, node A remains fixed, node B is shifted along th…
Figure 98
Figure 98. Figure 98: Method 1, Optimal number of integration points, feasibility (Section 10.2.1). Dis￾tribution of minimum numbers of integration points on a local coordinate axes for a maximum error of e tol = 10−3 among 10,000 elements generated randomly using the method in Fig￾ure 97.…
Figure 99
Figure 99. Figure 99: Method 1, Optimal network architecture for training (Section 10.2.2). The number of hidden layers varies from 1 to 5, keeping the number of neurons per hidden layer constant at 50. The network with 3 hidden layers provided the highest accuracy for both the training se…
Figure 100
Figure 100. Figure 100: Method 1, application phase (Section 10.2.3). The numbers of quadrature points predicted by the neural network was compared to the minimum numbers of quadrature points for maximum error e tol = 10−3 [38]. Table (a) shows the results for the training set (“pat￾terns”)…
Figure 101
Figure 101. Figure 101: Method 2, Quadrature weight correction, feasibility (Section 10.3.1). Each ele￾ment was tested 1 million times with randomly generated sets of quadrature weights. There were 4000 elements in each of the 5 groups with different degrees of maximum distortion, d. Quadra…
Figure 102
Figure 102. Figure 102: Method 2, training phase, classifier network (Section 10.3.2). The training and validation sets comprised 5000 elements each, of which 3707 and 3682, respectively, belonged to Category A (no improvements upon weight correction). A first neural network with 4 hidden l…
Figure 103
Figure 103. Figure 103: Method 2, training phase, regression network (Section 10.3.2). A second neu￾ral network estimated 8 correction factors {wi,j,k}, with i, j, k ∈ {1, 2}, to be multiplied by the standard quadrature weights for each element. Distribution of normalized errors, i.e., the …
Figure 10
Figure 10. Figure 10: Distributions of error ratios defined by Eq. (21), when using correction factors estimated by deep learning. to deduce the constitutive behavior on the macroscopic scale is evaluated at the quadrature points of the [PITH_FULL_IMAGE:figures/full_fig_p172_10.png]
Figure 104
Figure 104. Figure 104: Three scales in data-driven fault-reactivation simulations (Sections 2.3.2, 11.1, 11.3.5). Relative orientation of Representative Volume Elements (RVEs). Left: Microscale (µ) RVE using Discrete Element Method (DEM), [PITH_FULL_IMAGE:figures/full_fig_p173_104.png]
Figure 105
Figure 105. Figure 105: Single-physics block diagram (Section 11.2). Single physics is an easiest way to see the role of deep learning in modeling complex nonlinear constitutive behavior (stress￾strain relation, red arrow), as first realized in [23], where balance of linear momentum and str…
Figure 106
Figure 106. Figure 106: Microscale RVE (Sections 11.3.2, 11.3.3, 11.3.5). A 10 cm × 10 cm × 5 cm box of identical spheres of 0.5 cm diameter ( [PITH_FULL_IMAGE:figures/full_fig_p175_106.png]
Figure 107
Figure 107. Figure 107: Optimal RNN-LSTM architecture (Section 11.3.3). 5 different configurations of RNNs with LSTM units [25]. (Table reproduced with permission of the authors.) 11.3.3 Optimal RNN-LSTM architecture Using the same discrete element assembly of microscale RVE in [PITH_FULL_…
Figure 108
Figure 108. Figure 108: Optimal RNN-LSTM architecture (Section 11.3.3). Training error and test errors for 5 different configurations of RNN with LSTM units, see [PITH_FULL_IMAGE:figures/full_fig_p176_108.png]
Figure 109
Figure 109. Figure 109: Optimal RNN-LSTM architecture (Section 11.3.3). Training error (a) and testing error (b), close-up views of [PITH_FULL_IMAGE:figures/full_fig_p177_109.png]
Figure 110
Figure 110. Figure 110: Mesoscale RNN with LSTM units. Traction-separation law (Sections 11.3.3, 11.3.5). Left: Sequence of imposed displacement jumps on microscale RVE ( [PITH_FULL_IMAGE:figures/full_fig_p178_110.png]
Figure 111
Figure 111. Figure 111: Continuum with embedded strong discontinuity (Section 11.3.5). Domain B = B + ∪ B− with embedded discontinuity surface Γ, running through the middle of a narrow band (light blue) Bh = (B + h ∪ B− h ) ⊂ B between the parallel surfaces Γ + and Γ −. Objects behind Γ in …
Figure 112
Figure 112. Figure 112: Mesoscale RVE (Sections 11.3.3, 11.3.5). A 2-D domain of size 1 m × 1 m (Re￾mark 11.9). See [PITH_FULL_IMAGE:figures/full_fig_p180_112.png]
Figure 113
Figure 113. Figure 113: Mesoscale RVE (Section 11.3.3). Strains and displacement jumps [25] (Figure reproduced with permission of the authors.) where τ is the shear stress along the fault line, τp the critical shear stress for the onset of fault reactivation, C the cohesion strength, µ the …
Figure 114
Figure 114. Figure 114: Mesoscale RVE (Section 11.3.5). Validation of coupled FEM and RNN with LSTM units (FEM-LSTM, red dotted line) against coupled FEM and DEM (FEM-DEM, blue line) to analyze the mesoscale RVE in [PITH_FULL_IMAGE:figures/full_fig_p182_114.png]
Figure 26
Figure 26. Figure 26: Loading path of three selected training cases TR1, TR2, TR3 and three selected testing cases TE1, TE2, TE3 on the meso-scale RVE. un and us are the normal and tangential displacement jumps. The coordinate system is {M, N} (or {x, y}) depicted in [PITH_FULL_IMAGE:figu…
Figure 115
Figure 115. Figure 115: Macroscale RNN with LSTM units (Section 11.3.5). Normal traction (Tn) vs im￾posed displacement jumps (Un) on mesoscale RVE ( [PITH_FULL_IMAGE:figures/full_fig_p183_115.png]
Figure 28
Figure 28. Figure 28: Comparison of the meso-scale FEM–LSTM simulation data and the trained macro-scale data-driven model. Tangential traction against tangential displacement jump for the selected training and testing cases. The numbers mark the sequence of loading–unloading cycles. MSE re…
Figure 116
Figure 116. Figure 116: 2-D datasets for training neural networks (Sections 2.3.3, 12.1). Extract 2-D datasets from 3-D turbulent flow field evolving in time. From the 3-D flow field, extract N equidistant 2-D planes (slices). Within each 2-D plane, select a region (yellow square), and k te…
Figure 117
Figure 117. Figure 117: LSTM unit and BiLSTM unit (Sections 2.3.2, 2.3.3, 7.2, 12.2). Each blue dot is an original LSTM unit (in folded form [PITH_FULL_IMAGE:figures/full_fig_p186_117.png]
Figure 118
Figure 118. Figure 118: LSTM/BiLSTM training strategy (Sections 12.2.1, 12.2.2). From the 1-D time series αi(t) of each dominant mode ϕi , for i = 1, . . . , m, use a moving window to extract thousands of samples αi(t), t ∈ [tk, tspl k ], with tk being the time of snapshot k. Each sample is…
Figure 15
Figure 15. Figure 15: Training of a unified NN model for all POD dominant modes chaotic systems, they are outside the scope of this work. 4.2 Magnetohydrodynamic Turbulence The strategy in the previous section required a NN model to be trained for each POD mode - a multiple model approach.…
Figure 120
Figure 120. Figure 120: Hurst exponent vs POD-mode rank for Isotropic Turbulence (ISO) (Sections 12.3). POD modes with larger eigenvalues (Eq. (438)) are higher ranked, and have lower rank number, e.g., POD mode rank 7 has larger eigenvalue, and thus more dominant, than POD mode rank 50. Th…
Figure 121
Figure 121. Figure 121: Space-time solution of inviscid 1D-Burgers’ equation (Section 12.4.1). The solu￾tion shows a characteristic steep spatial gradient, which shifts and further steepens in the course of time. The FOM solution (left) and the solution of the proposed hyper-reduced ROM (ce…
Figure 122
Figure 122. Figure 122: Dense vs. shallow decoder networks (Section 12.4.3). Contributing neurons (or￾ange “nodes”) and connections (orange “edges”) lie in the “active” paths arriving at the selected outputs (solid orange “nodes”) from the decoder’s inputs. In dense networks as the one in (…
Figure 123
Figure 123. Figure 123: Sparsity masks (Section 12.4.3) used to realize sparse decoders in one- and two￾dimensional problems. The structure of the respective binary-valued mask matrices S is in￾spired by grid-points required in the finite-difference approximation of the Laplace operator in …
Figure 124
Figure 124. Figure 124: Subnet construction (Section 12.4.4). To reduce computational cost, a subnet representing the set of active paths, which comprise all neurons and connections needed for the evaluation of selected outputs (highlighted in orange), i.e., the reduced residual rb, is con￾…
Figure 125
Figure 125. Figure 125: 2-D Burger’s equation. Solution snapshots of full and reduced-order models (Section 12.4.5). From left to right, the components u (top row) and v (bottom row) of the ve￾locity field at time t = 2 are shown for the FOM, the hyper-reduced nonlinear-manifold-based ROM (…
Figure 126
Figure 126. Figure 126: 2-D Burger’s equation. Reynolds number vs. singular values (Section 12.4.5). Performing SVD on FOM solution snapshots, which were partitioned into x and y-components, the influence of the Reynolds number on the singular values is illustrated. In diffusion￾dominated p…
Figure 127
Figure 127. Figure 127: 2D-Burgers’ equation: relative errors of nonlinear manifold and linear subspace ROMs (Section 12.4.5). (Figure reproduced with permission of the authors.) [PITH_FULL_IMAGE:figures/full_fig_p207_127.png]
Figure 128
Figure 128. Figure 128: Machine-learning accelerated CFD (Section 12.4.5). Speed-up factor, compared to direct integration, was much higher than those obtained from nonlinear model-order reduc￾tion in [PITH_FULL_IMAGE:figures/full_fig_p208_128.png]
Figure 129
Figure 129. Figure 129: Machine-learning accelerated CFD (Section 12.4.5). Good accuracy and good generalization, devoiding of non-physical solutions [317]. Permission of NAS [PITH_FULL_IMAGE:figures/full_fig_p208_129.png]
Figure 130
Figure 130. Figure 130: Machine-learning accelerated CFD (Section 12.4.5). The neural network gener￾ates interpolation coefficients based on local-flow properties, while ensuring at least first-order accuracy relative to the grid spacing [317]. Permission of NAS. Remark 12.8. Machine-learni…
Figure 131
Figure 131. Figure 131: Biological Neuron and signal flow (Sections 4.4.4, 13.1, 13.2.2) along myelinated axon, with inputs at the synapses (input points) in the dendrites and with outputs at the axon terminals (output points,which are also the synapses for the next neuron). Each input curr…
Figure 132
Figure 132. Figure 132: The perceptron network (Sections 4.5, 13.2)—introduced by Rosenblatt (1958) [119], (1962) [120]—has a linear combination with weights and bias as expressed in z (1)(xi) = wxi + b ∈ R, but differs from the one-layer network in [PITH_FULL_IMAGE:figures/full_fig_p210_1…
Figure 133
Figure 133. Figure 133: Rosenblatt and the Mark I computer (Sections 4.6.1, 13.2) based on the percep￾tron, described in the New York Times article titled “New Navy device learns by doing” on 1958 July 8 (Internet archive), as a “computer designed to read and grow wiser”, and would be able …
Figure 134
Figure 134. Figure 134: Model of neocortical neurons in [118] as a simplification of the model in [322] (Section 13.2.2): A capacitor C with a potential V across its plates, in parallel with the equilib￾rium potentials ENa (sodium) and EK (potassium) in opposite direction. Two variable resi…
Figure 135
Figure 135. Figure 135: Continuous recurrent neural network with time-dependent delay d(t) (green feed￾back loop, Section 13.2.2), as expressed in Eq. (514), where f(·) is the operator with the first defivative term plus a standard static term—which is an activation function acting on linea…
Figure 136
Figure 136. Figure 136: Crayfish (Section 13.3.2), freshwater crustaceans. Anatomy. 13.3 Activation functions 13.3.1 Logistic sigmoid The use of the logistic sigmoid function ( [PITH_FULL_IMAGE:figures/full_fig_p220_136.png]
Figure 137
Figure 137. Figure 137: Crayfish giant motor synapse (Section 13.3.2). The (pre-synaptic) lateral giant fiber was connected to the (post-synaptic) giant motor fiber through a synapse where the two fibers cross each other at the location annotated by “Giant motor synapse” in the figure. This…
Figure 138
Figure 138. Figure 138: Crayfish Giant Motor Synapse (Section 13.3.2). The response in SubFigure (a) is similar to that of a rectifier circuit with leaky diode in [PITH_FULL_IMAGE:figures/full_fig_p222_138.png]
Figure 139
Figure 139. Figure 139: Swish function (Section 13.3.3) x · s(βx), with s(·) being the logistic sigmoid in [PITH_FULL_IMAGE:figures/full_fig_p223_139.png]
Figure 140
Figure 140. Figure 140: MIT COVID-19 diagnosis by cough recordings. Machine learning architecture. Audio Mel Frequency Cepstrum Coefficients (MFCC) as input. Each cough signal is split into 6 audio chunks, processed by the MFCC package, then passed through the Biomarker 1 to check on muscul…
Figure 141
Figure 141. Figure 141: Tesla Full-Self-Driving (FSD) controversy (Section 14.1). Left: Tesla in FSD mode hit a child-size mannequin, repeatedly in safety tests by The Dawn Project, a software competitor to Tesla, 2022.08.09 [376] [377]. Right: Tesla in FSD mode went around a child￾size man…
Figure 142
Figure 142. Figure 142: Tesla Full-Self-Driving (FSD) controversy (Section 14.1). The Tesla was about to run down the child-size mannequin at 23 mph, hitting it at 24 mph. The driver did not hold on, but only kept his hands close, to the driving wheel for safety, and did not put his foot on…
Figure 143
Figure 143. Figure 143: Tesla crash (Section 14.1). July 2020. Left: “Less than a half-second after [the Tesla driver] flipped on her turn signal, Autopilot started moving the car into the right lane and gradually slowed, video and sensor data showed.” Right: “Halfway through, the Tesla sen…
Figure 144
Figure 144. Figure 144: Tesla crash (Section 14.1). July 2020. “Less than a second after the Tesla has slowed to roughly 55 m.p.h. [Left], its rear camera shows a car rapidly approaching [Right]” [382]. There were no moving cars on both lanes in front of the Tesla for a long distance ahead …
Figure 145
Figure 145. Figure 145: Tesla crash (Section 14.1). July 2020. The fast-coming blue car rear-ended the Tesla, indented its own front bumper, with flying broken glass (or clear plastic) cover shards captured by the Tesla rear camera [382]. See also Figures 143, 144, 146. (Data and video prov…
Figure 146
Figure 146. Figure 146: Tesla crash (Section 14.1). After hitting the Tesla, the blue car “spun across the highway [Left] and onto the far shoulder [Right],” as another car was coming toward on the right lane (left in photo), but still at a safe distance so not to hit it. [382]. See also Fi…
Figure 147
Figure 147. Figure 147: Mayflower autonomous ship (Section 14.1) sailing from Plymouth, UK, planning to arrive at Plymouth, MA, U.S., like the original Mayflower 400 years ago, but instead arriving at Halifax, Nova Scotia, Canada, on 2022 Jun 05, due to mechanical problems [394]. (CC BY￾SA …
Figure 148
Figure 148. Figure 148: Network with infinite width (left) and Gaussian distribution (Right) (Section 6.1, 14.2). “A number of recent results have shown that DNNs that are allowed to become in￾finitely wide converge to another, simpler, class of models called Gaussian processes. In this lim…
Figure 149
Figure 149. Figure 149: Deepfake images (Section 14.4.1). AI-generated portraits using Generative Ad￾versarial Network (GAN) models. See also [397] [398], Chap. 8, “GAN Fingerprints in Face Image Synthesis.” (Images from ‘This Person Does Not Exist’ site.) 14.4.1 Deepfakes AI software avail…
Figure 150
Figure 150. Figure 150: DeepFake detection (Section 14.4.1). Violin plots. • Individual vs machine. The leading model had an accuracy of 65% on 4,000 videos (Col. 1). In Experiment 1 (E1), 5,524 participants were asked to identify a deepfake from each of 56 pairs of videos. The participants…
Figure 151
Figure 151. Figure 151: Lack of transparency and irreproducibility (Section 14.7). The table shows many missing pieces of information for the three networks—Lesion, Breast, and Case models—used to detect breast cancer. Learning rate, Section 6.2. Learning-rate schedule, Section 6.3.1, [PIT…
Figure 152
Figure 152. Figure 152: below [PITH_FULL_IMAGE:figures/full_fig_p268_152.png]
Figure 10
Figure 10. Figure 10: in the online book [PITH_FULL_IMAGE:figures/full_fig_p268_10.png]
Figure 153
Figure 153. Figure 153: The first two waves of AI, according to [78], p.13, showing the “cybernetics” wave (blue line) started in the 1940s peaked before 1970, then gradually declined toward 2006 and beyond. The results were based on a search for frequency of words in Google Books. It was m…
Figure 154
Figure 154. Figure 154: Cybernetics papers, (Appendix 4). Web of Science search on 2020.04.15, having more than 100 Web of Science categories. The first paper was [426]. There was no clear wave that crested before 1970, but actually the number of papers in Cybernetics continue to increase o…
Figure 155
Figure 155. Figure 155: Cybernetics papers, (Appendix 4). Web of Science search on 2020.04.17, ALL Computer-Science categories (3,555 papers): Cybernetics (2,666), Artificial Intelligence (602), Information Systems (432), Theory Methods (300), Interdisciplinary Applications (293), Soft￾ware…
Figure 156
Figure 156. Figure 156: Cybernetics papers, (Appendix 4). Web of Science search on 2020.04.15 (two days before [PITH_FULL_IMAGE:figures/full_fig_p274_156.png]
Figure 157
Figure 157. Figure 157: Cybernetics papers, (Appendix 4). Web of Science search on 2020.04.15 (two days before [PITH_FULL_IMAGE:figures/full_fig_p274_157.png]
Figure 158
Figure 158. Figure 158: Artificial Intelligence (AI), Machine Learning (ML), and Deep Learning (DL). Cybernetics is broad and encompasses many fields, including AI. See also [PITH_FULL_IMAGE:figures/full_fig_p275_158.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SLIDE: A machine-learning based method for forced dynamic response estimation of multibody systems

    cs.LG 2024-09 unverdicted novelty 6.0 of 10

    SLIDE is a deep learning estimator that truncates initial effects via complex eigenvalues of linearized equations to predict output sequences of damped multibody systems, reporting speedups up to several million times.

Reference graph

Works this paper leans on

286 extracted references · 286 canonical work pages · cited by 1 Pith paper

  1. [2]

    Rosenblatt, F. (1962). Principles of neurodynamics: Perceptrons and the theory of brain mechanisms. Spartan Books. 2, 11, 46, 55, 210, 212, 213, 214, 215, 271

  2. [3]

    Polyak, B. (1964). Some methods of speeding up the convergence of iteration methods . USSR Com- putational Mathematics and Mathematical Physics, 4(5), 1–17. DOI 10.1016/0041-5553(64)90137-5. 2, 10, 11, 85, 89, 90, 91

  3. [4]

    Roose, K. (2022). An A.I.-Generated Picture Won an Art Prize. Artists Aren’t Happy.New York Times, (Sep 2). Original website. 6, 7

  4. [5]

    Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873), 583–589. 7

  5. [6]

    J., Guez, A., Sifre, L., et al

    Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484+. Original website. 7, 12, 13

  6. [7]

    How Google’s AlphaGo Beat a Go World Champion

    Moyer, C. How Google’s AlphaGo Beat a Go World Champion. 2016 Mar 28, Original website. 7

  7. [8]

    Edwards, B. (2022). DeepMind breaks 50-year math record using AI; new record falls a week later. Ars Technica, (Oct 13). Original website, Internet archive. 7

  8. [9]

    Vu-Quoc, L., Humer, A. (2022). Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics. arXiv:2212.08989. 8

Show all 286 references
  1. [10]

    Roose, K. (2023). Bing (Yes, Bing) Just Made Search Interesting Again. New York Times, (Feb 8). Original website. 8

  2. [11]

    Knight, W. (2023). Meet Bard, Google’s Answer to ChatGPT. WIRED, (Feb 6). Original website. 8

  3. [12]

    Schmidhuber, J. (2015). Deep learning in neural networks: An overview. Neural Networks, 61, 87–

  4. [13]

    8, 36, 38, 52, 223, 224, 225, 272

  5. [14]

    LeCun, Y ., Bengio, Y ., Hinton, G. (2015). Deep learning.Nature, 521(7553), 436–444. 8, 12, 14, 38, 52, 53, 54, 129, 131

  6. [15]

    Khan, S., Yairi, T. (2018). A review on the application of deep learning in system health management. Mechanical Systems and Signal Processing, 107, 241–265. 8

  7. [16]

    Sanchez-Lengeling, B., Aspuru-Guzik, A. (2018). Inverse molecular design using machine learning: Generative models for matter engineering. Science, 361(6400, SI), 360–365. 8

  8. [17]

    S., Beaulieu-Jones, B

    Ching, T., Himmelstein, D. S., Beaulieu-Jones, B. K., Kalinin, A. A., Do, B. T., et al. (2018). Opportu- nities and obstacles for deep learning in biology and medicine. Journal of the Royal Society Interface, 15(141). 8

  9. [18]

    A., Nyhan, M

    Quinn, J. A., Nyhan, M. M., Navarro, C., Coluccia, D., Bromley, L., et al. (2018). Humanitarian applications of machine learning with remote-sensing data: review and case study in refugee settlement mapping. Philosophical Transactions of the Royal Society A-Mathematical Physic...

  10. [19]

    F., Higham, D

    Higham, C. F., Higham, D. J. (2019). Deep learning: An introduction for applied mathematicians. SIAM Review, 61(4), 860–891. 8

  11. [20]

    Dayan, P., Abbott, L. (2001). Theoretical Neuroscience: Computational and Mathematical Modeling of Neural Systems. MIT Press. 8, 9, 11, 30, 31, 38, 39, 40, 41, 43, 212, 215, 216, 217, 219

  12. [21]

    Sze, V ., Chen, Y .-H., Yang, T.-J., Emer, J. S. (2017). Efficient Processing of Deep Neural Networks: A Tutorial and Survey. Proceedings of the IEEE, 105(12), 2295–2329. 8, 17, 32, 38, 209

  13. [22]

    Nielsen, M. (2015). Neural Networks and Deep Learning . Determination Press. Original website. Internet archive. 8, 32, 38, 66, 67, 209, 210, 213

  14. [23]

    Rumelhart, D., Hinton, G., Williams, R. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. 8, 90, 215, 223, 224, 225, 271

  15. [24]

    Ghaboussi, J., Garrett, J., Wu, X. (1991). Knowledge-based modeling of material behavior with neural networks. Journal of Engineering Mechanics-ASCE, 117(1), 132–153. 8, 9, 26, 32, 173, 209, 272

  16. [26]

    Wang, K., Sun, W. C. (2018). A multiscale multi-permeability poroplasticity model linked by recursive homogenizations and deep learning. Computer Methods in Applied Mechanics and Engineering, 334, 337–380. 8, 9, 11, 22, 24, 25, 26, 27, 28, 172, 173, 174, 175, 176, 177, 178, 17...

  17. [27]

    Mohan, A., Gaitonde, D. (2018). A deep learning based approach to reduced order modeling for turbulent flow control using LSTM neural networks. arXiv:1804.09269 [physics.comp-ph]. Apr 24. 8, 9, 11, 28, 29, 30, 184, 185, 186, 187, 188, 189, 190, 191, 192

  18. [28]

    Zaman, M., Zhu, J. (1998). A neural network model for a cohesionless soilIn AttohOkine, NO. Arti- ficial Intelligence and Mathematical Methods in Pavement and Geomechanical Systems. International Workshop on Artificial Intelligence and Mathematical Methods in Pavement and Geom...

  19. [29]

    Su, H., Fan, L., Schlup, J. (1998). Monitoring the process of curing of epoxy/graphite fiber composites with a recurrent neural network as a soft sensor. Engineering Applications of Artificial Intelligence , 11(2), 293–306. 9

  20. [30]

    Li, C., Huang, T. (1999). Automatic structure and parameter training methods for modeling of me- chanical systems by recurrent neural networks. Applied Mathematical Modelling , 23(12), 933–944. 9

  21. [31]

    Waszczyszyn, Z. (2000). Neural networks in structural engineering: Some recent results and prospects for applicationsIn Topping, BHV. Computational Mechanics for the Twenty-First Century. 5th Inter- national Conference on Computational Structures Technology/2nd International C...

  22. [32]

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., et al. (2017). Attention Is All You Need. CoRR, abs/1706.03762v5. arXiv:1706.03762v5. See Footnote 337. 9, 11, 135, 138, 139, 140, 141, 142, 143, 248

  23. [33]

    Hahnloser, R., Sarpeshkar, R., Mahowald, M., Douglas, R., Seung, S. (2000). Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit (vol 405, pg 947, 2000). Nature, 408(6815), 1012–U24. 9, 39, 219, 221, 222

  24. [34]

    Jarrett, K., Kavukcuoglu, K., Ranzato, M., LeCun, Y . (2009). What is the Best Multi-Stage Architec- ture for Object Recognition?In 2009 IEEE 12th International Conference on Computer Vision (ICCV). IEEE International Conference on Computer Vision. IEEE; IEEE Comp Soc. 12th IE...

  25. [35]

    Nair, V ., Hinton, G. (2010). Rectified linear units improve restricted boltzmann machines.Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel. 9, 39

  26. [36]

    Little, W. (1974). The existence of persistent states in the brain. Mathematical Biosciences, 19, 101–

  27. [37]

    In Cabrera, B and Gutfreund, H and Kresin, V (eds), From High-Temperature Superconductivity to Microminiature Refrigeration, William Little Symposium on From High-Temperature Supercon- ductivity to Microminiature Refrigeration, Stanford Univ, Stanford, CA, Sep 30, 1995.336. 9, 220

  28. [38]

    Ramachandran, P., Barret, Z., Le, Q. (2017). Searching for Activation Functions. CoRR (Computing Research Repository), abs/1710.05941v2. arXiv:1710.05941v2. See Footnote 337. 9, 52, 219, 221, 222, 223

  29. [40]

    Oishi, A., Yagawa, G. (2017). Computational mechanics enhanced by deep learning. Computer Meth- ods in Applied Mechanics and Engineering, 327, 327–351. 9, 11, 18, 19, 20, 21, 32, 46, 53, 60, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 209

  30. [41]

    Zienkiewicz, O., Taylor, R., Zhu, J. (2013). The Finite Element Method: Its Basis and Fundamentals. Oxford: Butterworth-Heineman. 7th edition. 9, 35, 163, 164

  31. [42]

    Barlow, J. (1976). Optimal stress locations in finite-element models. International Journal for Numer- ical Methods in Engineering, 10(2), 243–251. 9

  32. [43]

    Barlow, J. (1977). Optimal stress locations in finite-element models - reply. International Journal for Numerical Methods in Engineering, 11(3), 604. 9

  33. [44]

    Theory Guide

    Abaqus 6.14. Theory Guide. Simulia Systems, Dassault Systèmes. Subsection 3.2.4 Solid isoparamet- ric quadrilaterals and hexahedra. (Website, go to Section Reference, Abaqus Theory Guide, Section 3 Elements, Section 3.2 Continuum elements, then Section 3.2.4.). 9

  34. [45]

    Ghaboussi, J., Garrett, J., Wu, X. (1990). Material Modeling with Neural NetworksIn Pande, GN and Middleton, J. Numerical Methods in Engineering : Theory and Applications, Vol 2. 3rd International Conf on Numerical Methods in Engineering : Theory and Applications ( NUMETA 90 )...

  35. [46]

    Chen, C. (1989). Applying and validating neural network technology for nondestructive evaluation of materialsIn 1989 IEEE International Conference on Systems, Man, and Cybernetics, Vols 1-3: Con- ference Proceedings. 1989 IEEE International Conf on Systems, Man, and Cybernetic...

  36. [47]

    Sayeh, M., Viswanathan, R., Dhali, S. (1990). Neural networks for assessment of impact and stress relief on composite-materialsIn Genisio, M. Sixth Annual Conference on Materials Technology: Com- posite Technology. 6th Annual Conf on Materials Technology : Composite Technology...

  37. [48]

    Chen, C., Leclair, S. (1991). A probability neural network (pnn) estimator for improved reliability of noisy sensor data. Journal of Reinforced Plastics and Composites, 10(4), 379–390. 9

  38. [49]

    Kim, Y ., Choi, Y ., Widemann, D., Zohdi, T. (2020). A fast and accurate physics-informed neural network reduced order model with shallow masked autoencoderer. ( Sep 28). Version 2, 2020.09.28: arXiv:2009.11990v2, 2009.11990. 9, 10, 11, 193, 194, 195, 196, 197, 198, 199, 200, ...

  39. [50]

    Kim, Y ., Choi, Y ., Widemann, D., Zohdi, T. (2020). Efficient nonlinear manifold reduced order model. (Nov 13). arXiv:2011.07727, 2011.07727. 9, 10, 11, 193

  40. [51]

    Robbins, H., Monro, S. (1951b). Stochastic approximation. Annals of Mathematical Statistics, 22(2),

  41. [52]

    Nesterov, I. (1983). A method of the solution of the convex-programming problem with a speed of convergence O(1/k2). Doklady Akademii Nauk SSSR, 269(3), 543–547. In Russian. 10, 89, 91

  42. [53]

    Nesterov, Y . (2018). Lecture on Convex Optimization. 2nd edition. Switzerland: Springer Nature. 10, 89, 91

  43. [54]

    Duchi, J., Hazan, E., Singer, Y . (2011). Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. Journal of Machine Learning Research, 12, 2121–2159. 10, 105

  44. [55]

    Tieleman, T., Hinton, G. (2012). Lecture 6e, rmsprop: Divide the gradient by a running average of its recent magnitude. Youtube video, time 5:54. Lecture notes, p.29: Original website, Internet archive. 10, 108

  45. [56]

    Zeiler, M. D. (2012). ADADELTA: An adaptive learning rate method. ( Dec 22). arXiv:1212.5701. 10, 106, 108, 109

  46. [58]

    Loshchilov, I., Hutter, F. (2019). Decoupled weight decay regularization. (Jan 4). arXiv:1711.05101v3. OpenReview. 10, 85, 87, 92, 93, 99, 106, 109, 115, 116, 117, 123

  47. [59]

    Bahdanau, D., Cho, K., Bengio, Y . (2015). Neural machine translation by jointly learning to align and translate. CoRR, abs/1409.0473. arXiv:1409.0473. 11, 135, 136, 137, 138

  48. [60]

    Furshpan, E., Potter, D. (1957). Mechanism of nerve-impulse transmission at a crayfish synapse. Nature, 180(4581), 342–343. 11, 222

  49. [61]

    Furshpan, E., Potter, D. (1959b). Slow post-synaptic potentials recorded from the giant motor fibre of the crayfish. Journal of Physiology-London, 145(2), 326–335. 11, 222

  50. [62]

    Gershgorn, D. (2017). The data that transformed AI research—and possibly the world. Quartz, (Jul 26). Original website. Internet archive (blurry images). 11, 13

  51. [63]

    He, K., Zhang, X., Ren, S., Sun, J. (2015). Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. CoRR, abs/1502.01852. arXiv:1502.01852, 1502.01852. 12, 40, 70, 206, 220

  52. [64]

    Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., et al. (2015). ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3), 211–252. 12, 13

  53. [65]

    Park, E., Liu, W., Russakovsky, O., Deng, J., Li, F., et al. (2017). ImageNet Large scale visual recogni- tion challenge (ILSVRC) 2017, Overview. ILSVRC 2017, (Jul 26). Original website Internet archive. 12, 13

  54. [66]

    Science’s 2021 Breakthrough: AI-powered Protein Prediction

    Beckwith, W. Science’s 2021 Breakthrough: AI-powered Protein Prediction. 2022 Dec 17, Original website. 11, 12

  55. [67]

    DeepMind, 2022 Jul 28, Original website, Internet archive

    AlphaFold reveals the structure of the protein universe. DeepMind, 2022 Jul 28, Original website, Internet archive. 12

  56. [68]

    DeepMind’s AI predicts structures for a vast trove of proteins

    Callaway, E. DeepMind’s AI predicts structures for a vast trove of proteins. 2021 Jul 21, Original website. 12

  57. [69]

    The Guardian view on the future of AI: Great power, great irresponsibility

    Editorial (2019). The Guardian view on the future of AI: Great power, great irresponsibility. The Guardian, (Jan 01). Original website. Internet archive. 12, 236, 237

  58. [70]

    Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., et al. (2018). A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play.Science, 362(6419), 1140+. 12

  59. [71]

    A., Veness, J., et al

    Mnih, V ., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. 13

  60. [72]

    P., Buesing, L., Guez, A., et al

    Racaniere, S., Weber, T., Reichert, D. P., Buesing, L., Guez, A., et al. (2017). Imagination-Augmented Agents for Deep Reinforcement Learning. In Guyon, I and Luxburg, UV and Bengio, S and Wallach, H and Fergus, R and Vishwanathan, S and Garnett, R, editor,Advances in Neural I...

  61. [73]

    Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., et al. (2017). Mastering the game of Go without human knowledge. Nature, 550(7676), 354+. 13

  62. [74]

    Artificial intelligence - hype, hope and fear

    Cellan-Jones, Rory (2017). Artificial intelligence - hype, hope and fear. BBC, (Oct 16). Original website. Internet archive. 13

  63. [75]

    Campbell, M. (2018). Mastering board games. A single algorithm can learn to play three hard board games. Science, 362(6419), 1118. 13

  64. [76]

    Why artificial intelligence is enjoying a renaissance

    The Economist (2016). Why artificial intelligence is enjoying a renaissance. ( Jul 15 ). (https://goo.gl/Grkofq). 13, 54, 226

  65. [77]

    From not working to neural networking

    The Economist (2016). From not working to neural networking. ( Jun 25). (https://goo.gl/z1c9pc). 13, 52, 54, 226

  66. [79]

    Hardesty, L. (2017). Explained: Neural networks. MIT News, (Apr 14). Original website. Internet archive. 13, 210

  67. [80]

    Goodfellow, I., Bengio, Y ., Courville, A. (2016). Deep Learning. Cambridge, MA: The MIT Press. 14, 16, 17, 27, 32, 34, 35, 36, 37, 38, 39, 40, 44, 46, 47, 48, 49, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 65, 67, 69, 70, 72, 73, 75, 76, 77, 78, 84, 85, 86, 87, 89, 90, 91, 9...

  68. [81]

    Ford, K. (2018). Architects of Intelligence: The truth about AI from the people building it . Packt Publishing. 14, 16, 221, 223, 224, 225, 235

  69. [82]

    E., Nocedal, J

    Bottou, L., Curtis, F. E., Nocedal, J. (2018). Optimization Methods for Large-Scale Machine Learning. SIAM Review, 60(2), 223–311. 14, 76, 78, 84, 85, 87, 93, 106, 108, 109

  70. [83]

    Khullar, D. (2019). A.I. Could Worsen Health Disparities. New York Times, (Jan 31). Original website. 14

  71. [84]

    Kornfield, M., Firozi, P. (2020). Artificial intelligence use is growing in the U.S. healthcare system. Washington Post, (Feb 24). Original website. 14

  72. [85]

    Lee, K. (2018a). AI Superpowers: China, Silicon Valley, and the New World Order. Houghton Mifflin Harcourt. 14

  73. [86]

    Lee, K. (2018b). How AI can save our humanity. TED2018, (Apr). Original website. 14

  74. [87]

    Dunjko, V ., Briegel, H. J. (2018). Machine learning & artificial intelligence in the quantum domain: a review of recent progress. Reports on Progress in Physics, 81(7), article no.074001. 16, 17

  75. [88]

    E., Osindero, S., Teh, Y .-W

    Hinton, G. E., Osindero, S., Teh, Y .-W. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18(7), 1527–1554. 16

  76. [89]

    A., Arthur, J

    Merolla, P. A., Arthur, J. V ., Alvarez-Icaza, R., Cassidy, A. S., Sawada, J., et al. (2014). A mil- lion spiking-neuron integrated circuit with a scalable communication network and interface. Science, 345(6197), 668–673. 17

  77. [90]

    K., Merolla, P

    Esser, S. K., Merolla, P. A., Arthur, J. V ., Cassidy, A. S., Appuswamy, R., et al. (2016). Convolutional networks for fast, energy-efficient neuromorphic computing. Proceedings of the National Academy of Sciences of the United States of America, 113(41), 11441–11446. 17

  78. [91]

    Warren, J., Root, P. (1963). The behavior of naturally fractured reservoirs. Society of Petroleum Engineers Journal, 3(03), 245–255. 22

  79. [92]

    A., Baud, P., Wong, T.-F

    Ji, Y ., Hall, S. A., Baud, P., Wong, T.-F. (2015). Characterization of pore structure and strain localiza- tion in Majella limestone by X-ray computed tomography and digital image correlation. Geophysical Journal International, 200(2), 701–719. 23, 24

  80. [93]

    Christensen, R. (2013). The Theory of Materials Failure. 1st edition. Oxford University Press. 22

  81. [95]

    Ho, C. K. (2000). Dual porosity vs. dual permeability models of matrix diffusion in fractured rock. Technical report. International High-Level Radioactive Waste Conference, Las Vegas, NV (US), 04/29/2001-05/03/2001. Sandia National Laboratories, Albuquerque, NM (US), Report No...

  82. [96]

    Datta-Gupta, A., King, M. J. (2007). Streamline simulation: Theory and practice, volume 11. Society of Petroleum Engineers Richardson. 22, 23, 24

  83. [97]

    Croizé, D., Renard, F., Gratier, J.-P. (2013). Chapter 3 - compaction and porosity reduction in carbonates: A review of observations, theory, and experiments. In R. Dmowska, editor, Advances in Geophysics, volume 54 of Advances in Geophysics. Elsevier, 181 – 238. 23, 24

  84. [98]

    Lu, J., Qu, J., Rahman, M. M. (2019). A new dual-permeability model for naturally fractured reser- voirs. Special Topics & Reviews in Porous Media: An International Journal, 10(5). 23

  85. [99]

    A., Schmidhuber, J

    Gers, F. A., Schmidhuber, J. (2000). Recurrent nets that time and countIn Proceedings of the IEEE- INNS-ENNS International Joint Conference on Neural Networks. IEEE. 24

  86. [100]

    Santamarina, J. C. (2003). Soil behavior at the microscale: particle forces. In Soil behavior and soft ground construction. 25–56. Proc. of the Symposium in honor of Charles C. Ladd, October 2001, MIT. 27

  87. [101]

    F., Haque, A., Ranjith, P

    Alam, M. F., Haque, A., Ranjith, P. G. (2018). A study of the particle-level fabric and morphology of granular soils under one-dimensional compression using insitu x-ray ct imaging. Materials, 11(6),

  88. [102]

    Karatza, Z., Andò, E., Papanicolopulos, S.-A., Viggiani, G., Ooi, J. Y . (2019). Effect of particle morphology and contacts on particle breakage in a granular assembly studied using x-ray tomography. Granular Matter, 21(3), 44. 26

  89. [103]

    Shire, T., O’Sullivan, C., Hanley, K., Fannin, R. J. (2014). Fabric and effective stress distribution in internally unstable soils. Journal of Geotechnical and Geoenvironmental Engineering , 140(12), 04014072. 26

  90. [104]

    Kanatani, K.-I. (1984). Distribution of directional data and fabric tensors. International journal of engineering science, 22(2), 149–164. 26, 174

  91. [105]

    Fu, P., Dafalias, Y . F. (2015). Relationship between void-and contact normal-based fabric tensors for 2d idealized granular materials. International Journal of Solids and Structures, 63, 68–81. 26

  92. [106]

    Graves, A., Schmidhuber, J. (2005). Framewise phoneme classification with bidirectional LSTM and other neural network architectures. Neural Networks, 18(5–6), 602–610. 29

  93. [107]

    Graham, J., Kanov, K., Yang, X., Lee, M., Malaya, N., et al. (2016). A web services accessible database of turbulent channel flow and its use for testing a new integral wall model for les. Journal of Turbulence, 17(2), 181–215. 29, 188

  94. [108]

    Rossant, C., Goodman, D. F. M., Fontaine, B., Platkiewicz, J., Magnusson, A. K., et al. (2011). Fitting neuron models to spike trains . Front. Neurosci., Feb 23. 31

  95. [109]

    Brillouin, L. (1964). Tensors in Mechanics and Elasticity. New York: Academic Press. 32, 34

  96. [110]

    Misner, C., Thorne, K., Wheeler, J. (1973). Gravitation. New York: W.H. Freeman and Company. 32

  97. [111]

    Malvern, L. (1969). Introduction to the Mechanics of a Continuous Medium. Englewood Cliffs, New Jersey: Prentice Hall. 34

  98. [112]

    Marsden, J., Hughes, T. (1994). Mathematical Foundation of Elasticity. New York: Dover. 34

  99. [113]

    Vu-Quoc, L., Li, S. (1995). Dynamics of sliding geometrically-exact beams - large-angle maneuver and parametric resonance. Computer Methods in Applied Mechanics and Engineering, 120(1-2), 65–

  100. [115]

    Glorot, X., Bordes, A., Bengio, Y . (2011). Deep Sparse Rectifier Neural Networks. In Proceedings of Machine Learning Research (PMLR), Vol.15, Fourteenth International Conference on Artificial In- telligence and Statistics (AISTATS), 11-13 April 2011, Fort Lauderdale, FL, USA ...

  101. [116]

    Drion, G., O’Leary, T., Marder, E. (2015). Ion channel degeneracy enables robust and tunable neuronal firing rates. Proceedings of the National Academy of Sciences of the United States of America, 112(38), E5361–E5370. 40, 42

  102. [117]

    van Welie, I., van Hooft, J., Wadman, W. (2004). Homeostatic scaling of neuronal excitability by synaptic modulation of somatic hyperpolarization-activated I-h channels. Proceedings of the National Academy of Sciences of the United States of America, 101(14), 5123–5128. 43

  103. [118]

    L., Steyn-Ross, D

    Steyn-Ross, M. L., Steyn-Ross, D. A. (2016). From individual spiking neurons to population behavior: Systematic elimination of short-wavelength spatial modes. Physical Review E, 93(2). 43, 218

  104. [119]

    R., Ganguly, U

    Dutta, S., Kumar, V ., Shukla, A., Mohapatra, N. R., Ganguly, U. (2017). Leaky Integrate and Fire Neuron by Charge-Discharge Dynamics in Floating-Body MOSFET. Scientific Reports, 7. 43

  105. [120]

    Wilson, H. (1999). Simplified dynamics of human and mammalian neocortical neurons. Journal of Theoretical Biology, 200(4), 375–388. 43, 217, 218

  106. [121]

    Rosenblatt, F. (1958). The perceptron - A probabilistic model for information-storage and organization in the brain. Psychological Review, 65(6), 386–408. 45, 46, 48, 49, 54, 55, 209, 210, 212, 213, 214, 215

  107. [122]

    Block, H. (1962a). Perceptron - A model for brain functioning .1. Reviews of Modern Physics, 34(1), 123–135. 45, 46, 210, 213, 214, 215

  108. [123]

    Minsky, M., Papert, S. (1969). Perceptrons: An introduction to computational geometry. MIT Press. 1988 expanded edition. 2017 edition with foreword by Leon Bottou, Facebook AI. 46, 47, 213, 214, 215

  109. [124]

    Herzberger, M. (1949). The normal equations of the method of least squares and their solution. Quar- terly of Applied Mathematics, 7(2), 217–223. (pdf). 48

  110. [125]

    Weisstein, E. W. Normal equation. From MathWorld–A Wolfram Web Resource. URL: http://mathworld.wolfram.com/NormalEquation.html. 48

  111. [126]

    Dyson, F. (2004). A meeting with Enrico Fermi - How one intuitive physicist rescued a team from fruitless research. Nature, 427(6972), 297. 52

  112. [127]

    Mayer, J., Khairy, K., Howard, J. (2010). Drawing an elephant with four complex parameters. Ameri- can Journal of Physics, 78(6), 648–649. 52

  113. [128]

    Hsu, J. (2015). Biggest Neural Network Ever Pushes AI Deep Learning. IEEE Spectrum. 54

  114. [129]

    He, K., Zhang, X., Ren, S., Sun, J. (2015). Deep Residual Learning for Image Recognition. CoRR (Computing Research Repository), abs/1512.03385v1. arXiv:1512.03385v1. See Footnote 337. 55, 56, 57, 67, 100, 102

  115. [130]

    Huang, G., Sun, Y ., Liu, Z., Sedra, D., Weinberger, K. (2016). Deep Networks with Stochastic Depth. CoRR (Computing Research Repository), abs/1603.09382v3. arXiv:1603.09382v3. See Footnote 337. 56

  116. [131]

    Zagoruyko, S., Komodakis, N. (2017). Wide residual networks. (Jun 17). CoRR (Computing Research Repository), arXiv:1605.07146v4. 56

  117. [132]

    Bishop, C. M. (2006). Pattern Recognition and Machine Learning . New York: Springer Sci- ence+Business Media. 59, 60, 61, 62, 72, 143, 147, 148, 149, 150, 151, 269, 270, 271

  118. [134]

    Li, H., Xu, Z., Taylor, G., Studer, C., Goldstein, T. (2018). Visualizing the loss landscape of neural nets. (Nov 7). arXiv:1712.09913v3. 71

  119. [135]

    Geman, S., Bienenstock, E., Doursat, R. (1992). Neural networks and the bias/variance dilemma. Neural computation, 4(1), 1–58. pdf, pdf. 72, 74

  120. [136]

    Hastie, T., Tibshirani, R., Friedman, J. H. (2001). The elements of statistical learning: Data mining, inference, prediction. 1st edition. Springer. 2nd edition, corrected, 12 printing, 2017 Jan 13. 72

  121. [137]

    Prechelt, L. (1998). Early Stopping — But When ? In G. Orr, K. Muller. Neural Networds: Tricks of the Trade . Springer . LLCS State-of-the-Art Survey. Paper pdf, Internet archive. 73, 74, 75

  122. [138]

    Belkin, M., Hsu, D., Ma, S., Mandal, S. (2019). Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences, 116(32), 15849– 15854. Original website, arXiv:1812.11118. 75, 76, 77

  123. [139]

    Geiger, M., Jacot, A., Spigler, S., Gabriel, F., Sagun, L., et al. (2020). Scaling description of gener- alization with number of parameters in deep learning. Journal of Statistical Mechanics: Theory and Experiment, 2020(2), 023401. Original website, arXiv:1901.01608. 75, 77

  124. [140]

    Sampaio, P. R. (2020). Deft-funnel: an open-source global optimization solver for constrained grey- box and black-box problems. (Jan 2020). arXiv:1912.12637. 76

  125. [141]

    Polak, E. (1971). Computational Methods in Optimization: A Unified Approach. Academic Press. 78, 79, 80, 81, 82, 83, 84, 90, 121

  126. [142]

    Lewis, R., Torczon, V ., Trosset, M. (2000). Direct search methods: then and now. Journal of Compu- tational and Applied Mathematics, 124(1-2), 191–207. 78, 84

  127. [143]

    Kolda, T., Lewis, R., Torczon, V . (2003). Optimization by direct search: New perspectives on some classical and modern methods. SIAM Review, 45(3), 385–482. 78, 84

  128. [144]

    Kafka, D., Wilke, D. (2018). Gradient-only line searches: An alternative to probabilistic line searches. (Mar 22). arXiv:1903.09383. 78, 125

  129. [145]

    Mahsereci, M., Hennig, P. (2017). Probabilistic line searches for stochastic optimization. Jour- nal of Machine Learning Research , 18. Article No.1. Also, CoRR, abs/1703.10034v2, Jun 30. arXiv:1703.10034v2, 1703.10034. 78, 84, 85, 123

  130. [146]

    Paquette, C., Scheinberg, K. (2018). A stochastic line search method with convergence rate analysis. (Jul 20). arXiv:1807.07994v1. 78, 81, 82, 83, 85, 87, 117, 119, 120, 121

  131. [147]

    Bergou, E., Diouane, Y ., Kungurtsev, V ., Royer, C. W. (2018). A subsampling line-search method with second-order results. ( Nov 21). arXiv:1810.07211v2. 78, 81, 82, 83, 85, 87, 117, 120, 121, 122, 123, 124

  132. [148]

    Wills, A., Schön, T. (2018). Stochastic quasi-newton with adaptive step lengths for large-scale problems. (Feb 22). arXiv:1802.04310v1. 78, 120, 123

  133. [149]

    Mahsereci, M., Hennig, P. (2015). Probabilistic line searches for stochastic optimization. CoRR, (Feb 10). Abs/1502.02846. arXiv:1502.02846. 78, 84

  134. [150]

    (2016).Linear and Nonlinear Programming

    Luenberger, D., Ye, Y . (2016).Linear and Nonlinear Programming. 4th edition. Springer. 79, 81, 90

  135. [151]

    Polak, E. (1997). Optimization: Algorithms and Consistent Approximations. Springer Verlag. 79, 80, 81, 84, 90

  136. [152]

    Goldstein, A. (1965). On steepest descent. SIAM Journal of Control, Series A, 3(1), 147–151. 79, 80, 81, 84, 85

  137. [153]

    Armijo, L. (1966). Minimization of functions having lipschitz continuous partial derivatives. Pacific Journal of Mathematics, 16(1), 1–3. 79, 80, 81, 85, 121

  138. [154]

    Wolfe, P. (1969). Convergence conditions for ascent methods. SIAM Review, 11(2), 226–235. 79, 81, 84, 85

  139. [156]

    Goldstein, A. (1967). Constructive Real Analysis. New York: Harper. 79, 84

  140. [157]

    Goldstein, A., Price, J. (1967). An effective algorithm for minimization. Numerische Mathematik, 10, 184–189. 79, 80, 81

  141. [158]

    Ortega, J., Rheinboldt, W. (1970). Iterative Solution of Nonlinear Equations in Several Variables. New York: Academic Press. Republished in 2000 by SIAM, Classics in Applied Mathematics, V ol.30. 79, 80, 81, 90

  142. [159]

    Nocedal, J., Wright, S. (2006). Numerical Optimization. Springer. 2nd edition. 81, 90

  143. [160]

    H., Nocedal, J

    Bollapragada, R., Byrd, R. H., Nocedal, J. (2019). Exact and inexact subsampled Newton methods for optimization. IMA Journal of Numerical Analysis, 39(2), 545–578. 81

  144. [161]

    S., Byrd, R

    Berahas, A. S., Byrd, R. H., Nocedal, J. (2019). Derivative-free optimization of noisy functions via quasi-newton methods. SIAM Journal on Optimization, 29(2), 965–993. 81

  145. [162]

    Larson, J., Menickelly, M., Wild, S. M. (2019). Derivative-free optimization methods. ( Jun 25). arXiv:1904.11585v2. 81

  146. [163]

    Shi, Z., Shen, J. (2005). Step-size estimation for unconstrained optimization methods. Computational and Applied Mathematics, 24(3), 399–416. 84

  147. [164]

    Sun, S., Cao, Z., Zhu, H., Zhao, J. (2019). A survey of optimization methods from a machine learning perspective. (Oct 23). arXiv:1906.06821v2. 85, 106, 108, 109

  148. [165]

    Kirkpatrick, S., Gelatt, C., Vecchi, M. (1983). Optimization by simulated annealing. Science, 220(4598), 671–680. 85, 96, 99

  149. [166]

    L., Kindermans, P.-J., Ying, C., Le, Q

    Smith, S. L., Kindermans, P.-J., Ying, C., Le, Q. V . (2018). Don’t decay the learning rate, increase the batch size. (Feb 2018). arXiv:1711.00489v2. OpenReview. 85, 93, 95, 96, 97, 98, 114

  150. [167]

    Schraudolph, N. (1998). Centering Neural Network Gradient Factors In G. Orr, K. Muller. Neural Networds: Tricks of the Trade . Springer . LLCS State-of-the-Art Survey. 85, 90, 107

  151. [168]

    Neuneier, R., Zimmermann, H. (1998). How to Train Neural Networks In G. Orr, K. Muller. Neural Networds: Tricks of the Trade . Springer . LLCS State-of-the-Art Survey. 85, 107

  152. [169]

    Robbins, H., Monro, S. (1951a). A stochastic approximation method. Annals of Mathematical Statis- tics, 22(3), 400–407. 85

  153. [170]

    Aitchison, L. (2019). Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods. (Jul 31). arXiv:1807.07540v4. 87, 107, 117, 118

  154. [171]

    Goudou, X., Munier, J. (2009). The gradient and heavy ball with friction dynamical systems: The qua- siconvex case. Mathematical Programming, 116(1-2), 173–191. 7th French-Latin American Congress in Applied Mathematics, Univ Chile, Santiago, CHILE, JAN, 2005. 89, 91

  155. [172]

    P., Ba, J

    Kingma, D. P., Ba, J. (2014). Adam: A method for stochastic optimization. ( Dec 22). Version 1, 2014.12.22: arXiv:1412.6980v1. Version 9, 2017.01.30: arXiv:1412.6980v9. 90, 102, 105, 106, 110, 111, 112, 113

  156. [173]

    Bertsekas, D., Tsitsiklis, J. (1995). Neuro-Dynamic Programming . Athena Scientific. 90, 91

  157. [174]

    Hinton, G. (2012). A Practical Guide to Training Restricted Boltzmann Machines In G. Montavon, G. Orr, K. Muller. Neural Networds: Tricks of the Trade . Springer . LLCS State-of-the-Art Survey. 90

  158. [175]

    Incerti, S., Parisi, V ., Zirilli, F. (1979). New method for solving non-linear simultaneous equations. SIAM Journal on Numerical Analysis, 16(5), 779–789. 90

  159. [176]

    V oigt, R. (1971). Rates of convergence for a class of iterative procedures.SIAM Journal on Numerical Analysis, 8(1), 127–&. 90

  160. [177]

    C., Nowlan, S

    Plaut, D. C., Nowlan, S. J., Hinton, G. E. (1986). Experiments on learning by back propagation. Technical Report Technical Report CMU-CS-86-126, June. Website. 90, 99

  161. [179]

    Hagiwara, M. (1992). Theoretical derivation of momentum term in back-propagation In Proceedings of the International Joint Conference on Neural Networks (IJCNN’92) . volume 1. Piscataway, NJ, IEEE . 90

  162. [180]

    Gill, P., Murray, W., Wright, M. (1981). Practical Optimization . Academic Press. 90

  163. [181]

    Snyman, J., Wilke, D. (2018). Practical Mathematical Optimization: Basic optimization theory and gradient-based algorithms . Springer. 90, 125

  164. [182]

    Priddy, K., Keller, P. (2005). Artificial neural network: An introduction . SPIE. 91

  165. [183]

    Sutskever, I., Martens, J., Dahl, G., Hinton, G. (2013). On the importance of initialization and momen- tum in deep learning. Proceedings of the 30th International Conference on Machine Learning, PMLR, 28(3). Original website. 91

  166. [184]

    J., Kale, S., Kumar, S

    Reddi, S. J., Kale, S., Kumar, S. (2019). On the convergence of Adam and beyond. ( Oct 23 ). arXiv:1904.09237. OpenReview. Best paper ICLR 2018. 92, 102, 104, 106, 110, 111, 112, 113, 123

  167. [185]

    T., Phong, L

    Phuong, T. T., Phong, L. T. (2019). On the convergence proof of AMSGrad and a new version. ( Oct 31). arXiv:1904.03590v4. 92, 112, 113

  168. [186]

    Li, X., Orabona, F. (2019). On the convergence of stochastic gradient descent with adaptive stepsizes. (Feb 26). arXiv:1805.08114v3. 93

  169. [187]

    Gardiner, C. (2004). Handbook of Stochastic Methods: for Physics, Chemistry and the Natural Sciences . Synergetics, 3rd edition. Springer. 94, 97, 98

  170. [188]

    L., Le, Q

    Smith, S. L., Le, Q. V . (2018). A bayesian perspective on generalization and stochastic gradient descent. (Feb 2018). arXiv:1710.06451v3. OpenReview. 95, 96

  171. [189]

    Li, Q., Tai, C., E, W. (2017). Stochastic modified equations and adaptive stochastic gradient algo- rithms. ( Jun 20). arXiv:1511.06251v3. Proceedings of Machine Learning Research, 70:2101-2110,

  172. [190]

    Lemons, D., Gythiel, A. (1997). Paul Langevin’s 1908 paper ‘’On the theory of Brownian motion”. American Journal of Physics, 65(11), 1079–1081. 98, 99

  173. [191]

    (2004).The Langevin Equation

    Coffey, W., Kalmikov, Y ., Waldron, J. (2004).The Langevin Equation . 2nd edition. World Scientific. 98

  174. [192]

    Lones, M. A. (2014). Metaheuristics in nature-inspired algorithmsIn Proceedings of the Companion Publication of the 2014 Annual Conference on Genetic and Evolutionary Computation. 99

  175. [193]

    Yang, X.-S. (2014). Nature-inspired optimization algorithms. Elsevier. 99

  176. [194]

    R., Fanany, M

    Rere, L. R., Fanany, M. I., Arymurthy, A. M. (2015). Simulated annealing algorithm for deep learning. Procedia Computer Science, 72(1), 137–144. 99

  177. [195]

    I., Arymurthy, A

    Rere, L., Fanany, M. I., Arymurthy, A. M. (2016). Metaheuristic algorithms for convolution neural network. Computational intelligence and neuroscience, 2016. 99

  178. [196]

    Fong, S., Deb, S., Yang, X.-s. (2018). How meta-heuristic algorithms contribute to deep learning in the hype of big data analytics. In Progress in Intelligent Computing Techniques: Theory, Practice, and Applications. Springer, 3–25. 99

  179. [197]

    Bozorg-Haddad, O. (2018). Advanced optimization by nature-inspired algorithms. Springer. 99

  180. [198]

    Al-Obeidat, F., Belacel, N., Spencer, B. (2019). Combining machine learning and metaheuristics algorithms for classification method proaftn. In Enhanced Living Environments. Springer, 53–79. 99

  181. [199]

    Bui, Q.-T. (2019). Metaheuristic algorithms in optimizing neural network: A comparative study for forest fire susceptibility mapping in Dak Nong, Vietnam.Geomatics, Natural Hazards and Risk, 10(1), 136–150. 99

  182. [201]

    S., Lewis, A

    Mirjalili, S., Dong, J. S., Lewis, A. (2020). Nature-Inspired Optimizers. Springer. 99

  183. [202]

    N., Topin, N

    Smith, L. N., Topin, N. (2018). Super-convergence: Very fast training of residual networks using large learning rates. (May 2018). arXiv:1708.07120v3. OpenReview. 99

  184. [203]

    Rögnvaldsson, T. S. (1998). A Simple Trick for Estimating the Weight Decay Parameter In G. Orr, K. Muller. Neural Networds: Tricks of the Trade. Springer . LLCS State-of-the-Art Survey. 99

  185. [204]

    Glorot, X., Bengio, Y . (2010). Understanding the difficulty of training deep feedforward neural net- worksIn Proceedings of the thirteenth international conference on artificial intelligence and statistics. JMLR Workshop and Conference Proceedings. 100

  186. [205]

    Bock, S., Goppold, J., Weiss, M. (2018). An improvement of the convergence proof of the ADAM- optimizer. (Apr 27). arXiv:1804.10587v1. 102, 112, 113

  187. [206]

    Huang, H., Wang, C., Dong, B. (2019). Nostalgic Adam: Weighting more of the past gradients when designing the adaptive learning rate. (Feb 23). arXiv:1805.07557v2. 102, 113

  188. [207]

    Chen, X., Liu, S., Sun, R., Hong, M. (2019). On the convergence of a class of Adam-type algorithms for non-convex optimization. (Mar 10). arXiv:1808.02941v2. OpenReview. 106

  189. [208]

    J., Koehler, A

    Hyndman, R. J., Koehler, A. B., Ord, J. K., Snyder, R. D. (2008). Forecasting with Exponential Smoothing: A state state approach. Springer. 106

  190. [209]

    J., Athanasopoulos, G

    Hyndman, R. J., Athanasopoulos, G. (2018). Forecasting: Principles and Practices . 2nd edition. OTexts: Melbourne, Australia. Original website, open online text. 107, 108

  191. [210]

    Dreiseitl, S., Ohno-Machado, L. (2002). Logistic regression and artificial neural network classification models: a methodology review . Journal of Biomedical Informatics, 35, 352–359. 112

  192. [211]

    Gugger, S., Howard, J. (2018). AdamW and Super-convergence is now the fastest way to train neural nets. Fast.AI, (Jul 02). Original website, Internet Archive. 113, 117

  193. [212]

    Xing, C., Arpit, D., Tsirigotis, C., Bengio, Y . (2018). A walk with sgd. ( May 2018 ). arXiv:1802.08770v4. OpenReview. 114

  194. [213]

    Prokhorov, D. (2001). IJCNN 2001 neural network competition. Slide presentation in IJCNN’01, Ford Research Laboratory, 2001 Internet Archive. 123

  195. [214]

    Chang, C.-C., Lin, C.-J. (2011). LIBSVM: A Library for Support Vector Machines.ACM Transactions on Intelligent Systems and Technology , 2(3). Article 27, April 2011. Original website for software (Version 3.24 released 2019.09.11), Internet Archive. 123

  196. [215]

    Brogan, W. L. (1990). Modern Control Theory. 3rd edition. Pearson. 125

  197. [216]

    Hopfield, J. J. (1984). Neurons with graded response have collective computational properties like those of two-state neurons. Proceedings of the National Academy of Sciences , 81(10), 3088–3092. Original website. 125

  198. [217]

    Pineda, F. J. (1987). Generalization of back-propagation to recurrent neural networks.Physical Review Letters, 59(19), 2229–2232. 125

  199. [218]

    Newmark, N. M. (1959). A Method of Computation for Structural Dynamics. Number 85 in A Method of Computation for Structural Dynamics. American Society of Civil Engineers. 126

  200. [219]

    M., Hughes, T

    Hilber, H. M., Hughes, T. J., Taylor, R. L. (1977). Improved numerical dissipation for time integration algorithms in structural dynamics. Earthquake Engineering & Structural Dynamics , 5(3), 283–292. Original website. 126

  201. [220]

    Chung, J., Hulbert, G. M. (1993). A Time Integration Algorithm for Structural Dynamics With Im- proved Numerical Dissipation: The Generalized-α Method. Journal of Applied Mechanics, 60(2), 371. Original website. 126

  202. [221]

    Olah, C. (2015). Understanding LSTM Networks. colah’s blog, (Aug 27). Original website. Internet archive. 131

  203. [223]

    Chung, J., Gulcehre, C., Cho, K., Bengio, Y . (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv:1412.3555. 134, 135

  204. [224]

    Kim, Y ., Denton, C., Hoang, L., Rush, A. M. (2017). Structured attention networks. International Conference on Learning Representations, OpenReview.net, arXiv:1702.00887. 135

  205. [225]

    Cho, K., van Merriënboer, B., Bahdanau, D., Bengio, Y . (2014). On the properties of neural ma- chine translation: Encoder–decoder approachesIn Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation. Doha, Qatar: Association for Co...

  206. [226]

    Schuster, M., Paliwal, K. K. (1997). Bidirectional recurrent neural networks. IEEE transactions on Signal Processing, 45(11), 2673–2681. 137

  207. [227]

    L., Kiros, J

    Ba, J. L., Kiros, J. R., Hinton, G. E. (2016). Layer normalization. arXiv:1607.06450. 141

  208. [228]

    B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., et al

    Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., et al. (2020). Language models are few-shot learners. arXiv:2005.14165v4. 143

  209. [229]

    H., Bai, S., Yamada, M., Morency, L.-P., Salakhutdinov, R

    Tsai, Y .-H. H., Bai, S., Yamada, M., Morency, L.-P., Salakhutdinov, R. (2019). Transformer dissection: A unified understanding of transformer’s attention via the lens of kernel. arXiv:1908.11775. 143

  210. [230]

    C., Friesen, T., et al

    Rodriguez-Torrado, R., Ruiz, P., Cueto-Felgueroso, L., Green, M. C., Friesen, T., et al. (2022). Physics-informed attention-based neural network for hyperbolic partial differential equations: appli- cation to the buckley–leverett problem. Scientific reports, 12(1), 1–12. Origi...

  211. [231]

    Bahri, Y . (2019). Towards an Understanding of Wide, Deep Neural Networks. Youtube. 143

  212. [232]

    Ananthaswamy, A. (2021). A New Link to an Old Model Could Crack the Mystery of Deep Learning. Quanta Magazine, (Oct 11). Original website. 143

  213. [233]

    S., Pennington, J., et al

    Lee, J., Bahri, Y ., Novak, R., Schoenholz, S. S., Pennington, J., et al. (2018). Deep neural networks as gaussian processes. arXiv:1711.00165. 143, 237

  214. [234]

    Jacot, A., Gabriel, F., Hongler, C. (2018). Neural tangent kernel: Convergence and generalization in neural networks. arXiv:1806.07572. 143, 144, 162

  215. [235]

    Quanta Magazine, 2021 Dec 31

    2021’s Biggest Breakthroughs in Math and Computer Science. Quanta Magazine, 2021 Dec 31. Youtube. 143, 236

  216. [236]

    E., Williams, C

    Rasmussen, C. E., Williams, C. K. (2006). Gaussian processes for machine learning . MIT press Cambridge, MA. MIT website, GaussianProcess.org. 143, 147, 148, 149, 151, 271

  217. [237]

    Belkin, M., Ma, S., Mandal, S. (2018). To understand deep learning we need to understand kernel learning. arXiv:1802.0139. 144, 147

  218. [238]

    S., Pennington, J., Adlam, B., Xiao, L., et al

    Lee, J., Schoenholz, S. S., Pennington, J., Adlam, B., Xiao, L., et al. (2020). Finite versus infinite neural networks: an empirical study. arXiv:2007.15801. 144

  219. [239]

    Aronszajn, N. (1950). Theory of reproducing kernels. Transactions of the American mathematical society, 68(3), 337–404. 144, 147

  220. [240]

    Hastie, T., Tibshirani, R., Friedman, Friedman, J. H. (2017). The elements of statistical learning: Data mining, inference, and prediction. 2 edition. Springer. Corrected, 12th printing, Jan 13. 145, 146, 147

  221. [241]

    Evgeniou, T., Pontil, M., Poggio, T. (2000). Regularization networks and support vector machines. Advances in computational mathematics, 13(1), 1–50. Semantic Scholar. 145, 146, 147

  222. [242]

    Berlinet, A., Thomas-Agnan, C. (2004). Reproducing kernel Hilbert spaces in probability and statis- tics. New York: Springer Science & Business Media. 146, 147, 148

  223. [243]

    Girosi, F. (1998). An equivalence between sparse approximation and support vector machines. Neural computation, 10(6), 1455–1480. Original website, Semantic Scholar. 146, 147

  224. [244]

    Wahba, G. (1990). Spline Models for Observational Data . Philadelphia, Pennsylvania: SIAM. 4th printing 2002. 147

  225. [246]

    Schaback, R., Wendland, H. (2006). Kernel techniques: From machine learning to meshless methods. Acta numerica, 15, 543–639. 147

  226. [247]

    Yaida, S. (2020). Non-gaussian processes and neural networks at finite widthsIn Mathematical and Scientific Machine Learning. Proceedings of Machine Learning Research. PMLR site. 148

  227. [248]

    Sendera, M., Tabor, J., Nowak, A., Bedychaj, A., Patacchiola, M., et al. (2021). Non-gaussian gaussian processes for few-shot regression. Advances in Neural Information Processing Systems , 34, 10285– 10298. arXiv:2110.13561. 148

  228. [249]

    Duvenaud, D. (2014). Automatic model construction with Gaussian processes. Ph.D. thesis, University of Cambridge. PhD dissertation. Thesis repository, CC BY-SA 2.0 UK. 151, 152

  229. [250]

    von Mises, R. (1964). Mathematical theory of probability and statistics. Elsevier. Book site. 151, 271

  230. [251]

    Hale, J. (2018). Deep Learning Framework Power Scores 2018. Towards Data Science, (Sep 19). Original website. Internet archive. 152, 154

  231. [252]

    Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., et al. (2015). TensorFlow: Large-scale machine learning on heterogeneous systems. Whitepaper pdf, Software available from tensorflow.org. 154

  232. [253]

    Google supercharges machine learning tasks with TPU custom chip

    Jouppi, N. Google supercharges machine learning tasks with TPU custom chip. Original website. 154

  233. [254]

    Chollet, F., et al. (2015). Keras. Original website. 155

  234. [255]

    Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., et al. (2019). Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, R. Garnett, editors,Advances in Neural Information Processing S...

  235. [256]

    Chintala, S. (2022). Decisions and pivots on pytorch. 2022 Jan 19, Original website Internet archive. 155

  236. [257]

    PyTorch Turns 5! 2022 Jan 20, Youtube. 155

  237. [258]

    P., Littman, M

    Kaelbling, L. P., Littman, M. L., Moore, A. W. (1996). Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4, 237–285. 155

  238. [259]

    P., Brundage, M., Bharath, A

    Arulkumaran, K., Deisenroth, M. P., Brundage, M., Bharath, A. A. (2017). Deep reinforcement learn- ing: A brief survey. IEEE Signal Processing Magazine, 34, 26–38. 155

  239. [260]

    Sünderhauf, N., Brock, O., Scheirer, W., Hadsell, R., Fox, D., et al. (2018). The limits and potentials of deep learning for robotics. The International Journal of Robotics Research, 37, 405–420. 156

  240. [261]

    C., Vu-Quoc, L

    Simo, J. C., Vu-Quoc, L. (1988). On the dynamics in space of rods undergoing large motions–a geometrically exact approach. Computer Methods in Applied Mechanics and Engineering , 66, 125–

  241. [262]

    Humer, A. (2013). Dynamic modeling of beams with non-material, deformation-dependent boundary conditions. Journal of sound and vibration, 332(3), 622–641. 156

  242. [263]

    Steinbrecher, I., Humer, A., Vu-Quoc, L. (2017). On the numerical modeling of sliding beams: A comparison of different approaches. Journal of Sound and Vibration, 408, 270–290. 156

  243. [264]

    Humer, A., Steinbrecher, I., Vu-Quoc, L. (2020). General sliding-beam formulation: A non-material description for analysis of sliding structures and axially moving beams. Journal of Sound and Vibra- tion, 480, 115341. Original website. 156

  244. [265]

    J., Leary, C., et al

    Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., et al. (2018). JAX: composable transformations of Python+NumPy programs. Original website. 156

  245. [266]

    Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., et al. (2020). Flax: A neural network library and ecosystem for JAX. Original website. 156

  246. [267]

    Schoeberl, J. (2014). C++11 Implementation of Finite Elements in NGSolve. Scientific report. 157, 158

  247. [269]

    Lavin, A., Zenil, H., Paige, B., Krakauer, D., Gottschlich, J., et al. (2021). Simulation Intelligence: Towards a New Generation of Scientific Methods. arXiv:2112.03235. 157

  248. [270]

    Cai, S., Mao, Z., Wang, Z., Yin, M., Karniadakis, G. E. (2021). Physics-informed neural networks (PINNs) for fluid mechanics: A review.Acta Mechanica Sinica, 37(12), 1727–1738. Original website, arXiv:2105.09506. 157, 158, 159

  249. [271]

    S., Giampaolo, F., Rozza, G., Raissi, M., et al

    Cuomo, S., di Cola, V . S., Giampaolo, F., Rozza, G., Raissi, M., et al. (2022). Scientific Machine Learning through Physics-Informed Neural Networks: Where we are and What’s next. Journal of Scientific Computing, 92(3). Article No. 88, Original website, arXiv:2201.05624. 158,...

  250. [272]

    E., Kevrekidis, I

    Karniadakis, G. E., Kevrekidis, I. G., Lu, L., Perdikaris, P., Wang, S., et al. (2021). Physics-informed machine learning. Nature Reviews Physics, 3(6), 422–440. Original website. 158, 159, 160

  251. [273]

    Lu, L., Meng, X., Mao, Z., Karniadakis, G. E. (2021). DeepXDE: A deep learning library for solving differential equations. SIAM Review, 63(1), 208–228. Original website, pdf, arXiv:1907.04502. 158, 160

  252. [274]

    SimNet” has been changed to “Modulus

    Hennigh, O., Narasimhan, S., Nabian, M. A., Subramaniam, A., Tangsali, K., et al. (2020). NVIDIA SimNet (tm): an AI-accelerated multi-physics simulation framework. arXiv:2012.07938. The software name “SimNet” has been changed to “Modulus”; see NVIDIA Modulus. 160

  253. [275]

    Koryagin, A., Khudorozkov, R., Tsimfer, S. (2019). PyDEns: a Python Framework for Solving Dif- ferential Equations with Neural Networks. arXiv:1909.11544. 160

  254. [276]

    Chen, F., Sondak, D., Protopapas, P., Mattheakis, M., Liu, S., et al. (2020). NeuroDiffEq: A python package for solving differential equations with neural networks. Journal of Open Source Software , 5(46), 1931. Original website. 159, 160

  255. [277]

    Rackauckas, C., Nie, Q. (2017). DifferentialEquations. jl–a performant and feature-rich ecosystem for solving differential equations in Julia. Journal of Open Research Software , 5(1). Original website. 159, 160

  256. [278]

    Haghighat, E., Juanes, R. (2021). Sciann: A keras/tensorflow wrapper for scientific computations and physics-informed deep learning using artificial neural networks. Computer Methods in Applied Mechanics and Engineering, 373, 113552. 160

  257. [279]

    Xu, K., Darve, E. (2020). ADCME: Learning Spatially-varying Physical Fields using Deep Neural Networks. arXiv:2011.11955. 160

  258. [280]

    R., Pleiss, G., Bindel, D., Weinberger, K

    Gardner, J. R., Pleiss, G., Bindel, D., Weinberger, K. Q., Wilson, A. G. (2018). Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration. [v6] Tue, 29 Jun 2021 arXiv:1809.11165. 160

  259. [281]

    S., Novak, R

    Schoenholz, S. S., Novak, R. Fast and Easy Infinitely Wide Networks with Neural Tangents. Google AI Blog, 2020 Mar 13, Original website. 160, 237

  260. [282]

    He, J., Li, L., Xu, J., Zheng, C. (2020). ReLU deep neural networks and linear finite elements. Journal of Computational Mathematics, 38(3), 502–527. arXiv:1807.03973. 160, 162

  261. [283]

    Arora, R., Basu, A., Mianjy, P., Mukherjee, A. (2016). Understanding deep neural networks with rectified linear units. arXiv:1611.01491. 160

  262. [284]

    Raissi, M., Perdikaris, P., Karniadakis, G. E. (2019). Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational physics, 378, 686–707. Original website. 161, 162

  263. [285]

    Kharazmi, E., Zhang, Z., Karniadakis, G. E. (2019). Variational physics-informed neural networks for solving partial differential equations. arXiv:1912.00873. 161, 162

  264. [286]

    Kharazmi, E., Zhang, Z., Karniadakis, G. E. (2021). hp-vpinns: Variational physics-informed neural networks with domain decomposition. Computer Methods in Applied Mechanics and Engineering , 374, 113547. See also arXiv:1912.00873. 161, 162 258 First online at CMES on 2023.03.0...

  265. [287]

    Berrone, S., Canuto, C., Pintore, M. (2022). Variational physics informed neural networks: the role of quadratures and test functions. Journal of Scientific Computing, 92(3), 1–27. Original website. 161

  266. [288]

    Wang, S., Yu, X., Perdikaris, P. (2020). When and why pinns fail to train: A neural tangent kernel perspective. arXiv:2007.14527. 162

  267. [289]

    M., Posch, S., Gössnitzer, C., Geiger, B

    Rohrhofer, F. M., Posch, S., Gössnitzer, C., Geiger, B. C. (2022). Understanding the difficulty of training physics-informed neural networks on dynamical systems. arXiv:2203.13648. 162

  268. [290]

    B., Muehlebach, M., Mahoney, M

    Erichson, N. B., Muehlebach, M., Mahoney, M. W. (2019). Physics-informed Autoencoders for Lyapunov-stable Fluid Flow Prediction. arXiv:1905.10866. 162

  269. [291]

    Raissi, M., Perdikaris, P., Karniadakis, G. E. (2021). Physics informed learning machine. US Patent 10,963,540, Mar 30. Google Patents, pdf. 162, 163

  270. [292]

    E., Likas, A., Fotiadis, D

    Lagaris, I. E., Likas, A., Fotiadis, D. I. (1998). Artificial neural networks for solving ordinary and partial differential equations. IEEE transactions on neural networks, 9(5), 987–1000. Original website. 162

  271. [293]

    E., Likas, A

    Lagaris, I. E., Likas, A. C., Papageorgiou, D. G. (2000). Neural-network methods for boundary value problems with irregular boundaries. IEEE Transactions on Neural Networks, 11(5), 1041–1049. Orig- inal website. 162

  272. [294]

    Raissi, M., Perdikaris, P., Karniadakis, G. E. (2017). Physics Informed Deep Learning (Part I): Data- driven Solutions of Nonlinear Partial Differential Equations. arXiv:1711.10561. 162

  273. [295]

    Raissi, M., Perdikaris, P., Karniadakis, G. E. (2017). Physics Informed Deep Learning (Part II): Data- driven Discovery of Nonlinear Partial Differential Equations. arXiv:1711.10566. 162

  274. [296]

    Gupta, S., Agrawal, A., Gopalakrishnan, K., Narayanan, P. (2015). Deep learning with limited numer- ical precisionIn International Conference on Machine Learning. arXiv:1502.02551. 171

  275. [297]

    Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., Bengio, Y . (2016). Binarized neural networks: Training deep neural networks with weights and activations constrained to+ 1 or-1. arXiv:1602.02830. 171

  276. [298]

    De Sa, C., Feldman, M., Ré, C., Olukotun, K. (2017). Understanding and optimizing asynchronous low-precision stochastic gradient descentIn Proceedings of the 44th Annual International Symposium on Computer Architecture. https://dl.acm.org/doi/abs/10.1145/3079856.3080248. 171

  277. [299]

    Borja, R. I. (2000). A finite element model for strain localization analysis of strongly discontinu- ous fields based on standard galerkin approximation. Computer Methods in Applied Mechanics and Engineering, 190(11-12), 1529–1549. 179, 182, 183, 184

  278. [300]

    Sibson, R. H. (1985). A note on fault reactivation. Journal of Structural Geology, 7(6), 751–754. 180

Pith tools

Reviewed May 24, 2026 · model on record in the stance chip above.