REVIEW 3 cited by
Performance Portable Solid Mechanics via Matrix-Free $p$-Multigrid
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Finite element analysis of solid mechanics is a foundational tool of modern engineering, with low-order finite element methods and assembled sparse matrices representing the industry standard for implicit analysis. We use performance models and numerical experiments to demonstrate that high-order methods greatly reduce the costs to reach engineering tolerances while enabling effective use of GPUs; these data structures also offer up to 2x benefit for linear elements. We demonstrate the reliability, efficiency, and scalability of matrix-free $p$-multigrid methods with algebraic multigrid coarse solvers through large deformation hyperelastic simulations of multiscale structures. We investigate accuracy, cost, and execution time on multi-node CPU and GPU systems for moderate to large models (millions to billions of degrees of freedom) using AMD MI250X (OLCF Crusher), NVIDIA A100 (NERSC Perlmutter), and V100 (LLNL Lassen and OLCF Summit), resulting in order of magnitude efficiency improvements over a broad range of model properties and scales. We discuss efficient matrix-free representation of Jacobians and demonstrate how automatic differentiation enables rapid development of nonlinear material models without impacting debuggability and workflows targeting GPUs. The methods are broadly applicable and amenable to common workflows, presented here via open source libraries that encapsulate all GPU-specific aspects and are accessible to both new and legacy code, allowing application code to be GPU-oblivious without compromising end-to-end performance on GPUs.
Forward citations
Cited by 3 Pith papers
-
A Portable and Versatile Limited-Memory BFGS Implementation in PETSc/TAO
An intermediate dense L-BFGS that applies H with one base-H0 solve and avoids Q/Z recomputation is implemented in PETSc/TAO and beats recursive and compact-dense variants on variable-metric CPU/GPU benchmarks.
-
Matrix-free phase-field modeling of fracture in micromechanical testing simulations of inelastic materials
A matrix-free GPU implementation of phase-field fracture coupled in series with a Perić–Dettmer visco-elastoplastic rheology reproduces qualitative 3D crack patterns in particle-matrix microstructures.
-
Matrix-Free Methods for Finite-Strain Elasticity: Automatic Code Generation with No Performance Overhead
Code generated by automatic differentiation runs as fast or faster than hand-written quadrature kernels in matrix-free finite-strain elasticity, with lower development effort.
Discussion (0). Continue with ORCID to comment.