Pith. sign in

REVIEW 3 cited by

Performance Portable Solid Mechanics via Matrix-Free $p$-Multigrid

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.01722 v3 pith:V3I4O37H submitted 2022-04-04 cs.MS cs.CEcs.DCcs.NAmath.NA

classification cs.MScs.CEcs.DCcs.NAmath.NA
keywords methodsdemonstrategpusmatrix-freemodelsmultigridperformanceanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Finite element analysis of solid mechanics is a foundational tool of modern engineering, with low-order finite element methods and assembled sparse matrices representing the industry standard for implicit analysis. We use performance models and numerical experiments to demonstrate that high-order methods greatly reduce the costs to reach engineering tolerances while enabling effective use of GPUs; these data structures also offer up to 2x benefit for linear elements. We demonstrate the reliability, efficiency, and scalability of matrix-free $p$-multigrid methods with algebraic multigrid coarse solvers through large deformation hyperelastic simulations of multiscale structures. We investigate accuracy, cost, and execution time on multi-node CPU and GPU systems for moderate to large models (millions to billions of degrees of freedom) using AMD MI250X (OLCF Crusher), NVIDIA A100 (NERSC Perlmutter), and V100 (LLNL Lassen and OLCF Summit), resulting in order of magnitude efficiency improvements over a broad range of model properties and scales. We discuss efficient matrix-free representation of Jacobians and demonstrate how automatic differentiation enables rapid development of nonlinear material models without impacting debuggability and workflows targeting GPUs. The methods are broadly applicable and amenable to common workflows, presented here via open source libraries that encapsulate all GPU-specific aspects and are accessible to both new and legacy code, allowing application code to be GPU-oblivious without compromising end-to-end performance on GPUs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. A Portable and Versatile Limited-Memory BFGS Implementation in PETSc/TAO

    cs.DC 2026-07 conditional novelty 6.0 of 10

    An intermediate dense L-BFGS that applies H with one base-H0 solve and avoids Q/Z recomputation is implemented in PETSc/TAO and beats recursive and compact-dense variants on variable-metric CPU/GPU benchmarks.

  2. Matrix-free phase-field modeling of fracture in micromechanical testing simulations of inelastic materials

    physics.comp-ph 2026-07 conditional novelty 6.0 of 10

    A matrix-free GPU implementation of phase-field fracture coupled in series with a Perić–Dettmer visco-elastoplastic rheology reproduces qualitative 3D crack patterns in particle-matrix microstructures.

  3. Matrix-Free Methods for Finite-Strain Elasticity: Automatic Code Generation with No Performance Overhead

    math.NA 2025-05 conditional novelty 6.0 of 10

    Code generated by automatic differentiation runs as fast or faster than hand-written quadrature kernels in matrix-free finite-strain elasticity, with lower development effort.

Pith tools