Pith. sign in

REVIEW 2 cited by

Coulomb, Landau and Maximally Abelian Gauge Fixing in Lattice QCD with Multi-GPUs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1212.5221 v2 pith:Y6SGJP6F submitted 2012-12-20 hep-lat

classification hep-lat
keywords latticegaugegpusperformanceabelianalgorithmsapplicationscoulomb
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A lattice gauge theory framework for simulations on graphic processing units (GPUs) using NVIDIA's CUDA is presented. The code comprises template classes that take care of an optimal data pattern to ensure coalesced reading from device memory to achieve maximum performance. In this work we concentrate on applications for lattice gauge fixing in 3+1 dimensional SU(3) lattice gauge field theories. We employ the overrelaxation, stochastic relaxation and simulated annealing algorithms which are perfectly suited to be accelerated by highly parallel architectures like GPUs. The applications support the Coulomb, Landau and maximally Abelian gauges. Moreover, we explore the evolution of the numerical accuracy of the SU(3) valued degrees of freedom over the runtime of the algorithms in single (SP) and double precision (DP). Therefrom we draw conclusions on the reliability of SP and DP simulations and suggest a mixed precision scheme that performs the critical parts of the algorithm in full DP while retaining 80-90% of the SP performance. Finally, multi-GPUs are adopted to overcome the memory constraint of single GPUs. A communicator class which hides the MPI data exchange at the boundaries of the lattice domains, via the low bandwidth PCI-Bus, effectively behind calculations in the inner part of the domain is presented. Linear scaling using 16 NVIDIA Tesla C2070 devices and a maximum performance of 3.5 Teraflops on lattices of size down to 64^3 x 256 is demonstrated.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Non-perturbative quark production and transport in the evolving glasma: the WAGASHI event generator

    hep-ph 2026-08 conditional novelty 6.0 of 10

    Schwinger pair production in the evolving glasma generates hundreds of quarks and antiquarks per unit rapidity in central Pb-Pb collisions, with momentum broadening and spin randomization before hydrodynamics begins.

  2. Cartan Fluxes in $SU(3)$ Lattice Gauge Theory

    hep-lat 2026-04 unverdicted novelty 5.0 of 10

    A root-lattice based DeGrand-Toussaint algorithm detects SU(3) monopoles with a density about 15% lower than the standard U(1)^3 method and yields Weyl-symmetric charge counts.

Pith tools