Pith. sign in

REVIEW 1 cited by

A systolic update scheme to overcome memory bandwidth limitations in GPU-accelerated FDTD simulations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.20610 v1 pith:T2SPSDYT submitted 2025-02-28 physics.optics

classification physics.optics
keywords photonicfdtdneedschemesimulationalgorithmcomputeengines
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The exponential growth of artificial intelligence has fueled the development of high-bandwidth photonic interconnect fabrics as a critical component of modern AI supercomputers. As the demand for ever-increasing AI compute and connectivity continues to grow, the need for high-throughput photonic simulation engines to accelerate and even revolutionize photonic design and verification workflows will become an increasingly indispensable capability for the integrated photonics industry. Unfortunately, the mainstay and workhorse of photonic simulation algorithms, the finite-difference time-domain (FDTD) method, because it is a memory-intensive but computationally-lightweight algorithm, is fundamentally misaligned with modern computational platforms which are equipped to deal with compute intensive workloads instead. This paper introduces a systolic update scheme for the FDTD method, which circumvents this mismatch by reducing the need for global synchronization while also relegating the need to access global memory to the case of boundary values between neighboring subdomains only. We demonstrate a practical implementation of our scheme as applied to the full three-dimensional FDTD algorithm that achieves a performance of roughly 0.15 trillion cell updates per second (TCUPS) on a single Nvidia H100 GPU. Our work paves the way for the increasingly efficient, cost-effective, and high-throughput photonic simulation engines needed to continue powering the AI era.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploiting nonlinear incoherent image formation through linear volume metaoptics for inference

    physics.optics 2025-08 conditional novelty 5.0 of 10

    A depth map of an opaque scene is encoded into the imaging response, letting a linear optics element plus linear readout perform nonlinear inference on depth.

Pith tools