Pith. sign in

REVIEW 5 cited by

Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11556 v2 pith:KSIFQSWB submitted 2024-08-21 cs.DC

classification cs.DC
keywords heterogeneousmemoryworkloadsbecomecoupledgh200gpusgrace
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Heterogeneous supercomputers have become the standard in HPC. GPUs in particular have dominated the accelerator landscape, offering unprecedented performance in parallel workloads and unlocking new possibilities in fields like AI and climate modeling. With many workloads becoming memory-bound, improving the communication latency and bandwidth within the system has become a main driver in the development of new architectures. The Grace Hopper Superchip (GH200) is a significant step in the direction of tightly coupled heterogeneous systems, in which all CPUs and GPUs share a unified address space and support transparent fine grained access to all main memory on the system. We characterize both intra- and inter-node memory operations on the Quad GH200 nodes of the new Swiss National Supercomputing Centre Alps supercomputer, and show the importance of careful memory placement on example workloads, highlighting tradeoffs and opportunities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference

    cs.DC 2026-07 accept novelty 6.5 of 10

    SiFAR cuts All-Reduce latency up to 52% and end-to-end decode throughput up to 18.6% at TP=8 by dual buffering, in-switch redundant pull, and speculative reduction with validation.

  2. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  3. Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks

    cs.DC 2025-07 conditional novelty 6.0 of 10

    A microbenchmark study maps memory hierarchy, execution pipelines, and FP4/FP6 tensor-core behavior on Nvidia's Blackwell RTX 5080 and compares it with Hopper's H100.

  4. Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A distributed equivariant GNN with a neighbor-minimizing graph partitioner scales electronic-structure (Hamiltonian) prediction to 512 GPUs and 190,000 atoms, with an 87% weak-scaling efficiency.

  5. Observational Analysis of Multi-thermal Counter-streaming Flows in a Forming Filament and Their Relationship with Local Heating at Filament Footpoints

    astro-ph.SR 2025-08 unverdicted novelty 5.0 of 10

    An abstract on solar filament flows and footpoint heating is attached to a manuscript body about AMD MI300A unified physical memory, leaving the solar analysis completely unsupported.

Pith tools