REVIEW 5 cited by
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads and unlocking new possibilities in fields like AI and climate modeling. With many workloads becoming memory-bound, improving the communication latency and bandwidth within the system has become a main driver in the development of new architectures. The Grace Hopper Superchip (GH200) is a significant step in the direction of tightly coupled heterogeneous systems, in which all CPUs and GPUs share a unified address space and support transparent fine grained access to all main memory on the system. We characterize both intra- and inter-node memory operations on the Quad GH200 nodes of the new Swiss National Supercomputing Centre Alps supercomputer, and show the importance of careful memory placement on example workloads, highlighting tradeoffs and opportunities.
Forward citations
Cited by 5 Pith papers
-
SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference
SiFAR cuts All-Reduce latency up to 52% and end-to-end decode throughput up to 18.6% at TP=8 by dual buffering, in-switch redundant pull, and speculative reduction with validation.
-
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...
-
Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks
A microbenchmark study maps memory hierarchy, execution pipelines, and FP4/FP6 tensor-core behavior on Nvidia's Blackwell RTX 5080 and compares it with Hopper's H100.
-
Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction
A distributed equivariant GNN with a neighbor-minimizing graph partitioner scales electronic-structure (Hamiltonian) prediction to 512 GPUs and 190,000 atoms, with an 87% weak-scaling efficiency.
-
Observational Analysis of Multi-thermal Counter-streaming Flows in a Forming Filament and Their Relationship with Local Heating at Filament Footpoints
An abstract on solar filament flows and footpoint heating is attached to a manuscript body about AMD MI300A unified physical memory, leaving the solar analysis completely unsupported.
Discussion (0). Continue with ORCID to comment.