Pith. sign in

REVIEW 3 major objections 3 minor 46 references

Enhanced Sensitivity and Noise Resilience in Two-Qubit Quantum Magnetometers

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper claims an optimized two-qubit Hamiltonian delivers a magnetic-field sensor with higher sensitivity and better noise resilience than existing two-qubit designs, backed by analytical QFI and SNR derivations that are not present in t

desk verdict The submission is not an evaluable paper: the abstract describes a two-qubit magnetometer, but the full text is an unrelated systems paper on GPU all-reduce, so the claimed QFI/SNR derivation is missing. read the letter →

arxiv 2508.13400 v1 pith:PZXERPRG submitted 2025-08-18 quant-ph

classification quant-ph
keywords two-qubitmagnetometerQuantumFisherInformationsignal-to-noiseratioHamiltonianoptimizationentanglement-enhancedsensingnoiseresiliencemagneticfieldmetrology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's stated project is a two-qubit quantum magnetometer: a Hamiltonian engineered so that a magnetic field leaves a sharper, more noise-resistant imprint on two coupled qubits than existing two-qubit designs. It claims that analytical derivations of the Quantum Fisher Information ($\mathcal{F}_Q$) and the signal-to-noise ratio (SNR) demonstrate gains in accuracy, robustness against noise, and entanglement dynamics, and that an alternate entangled initial state makes the entanglement benefit explicit. If correct, this would give a concrete route to better small-scale magnetic-field sensors without brute-force numerical search. The submitted full text, however, does not contain the promised magnetometer derivation; the body is an unrelated manuscript on GPU all-reduce communication optimization. The magnetometer claim therefore rests on the abstract alone, and the reader cannot yet check the derivation from the material provided.

What carries the argument

The load-bearing object is the optimized two-qubit Hamiltonian together with the Quantum Fisher Information ($\mathcal{F}_Q$) computed from its evolution: $\mathcal{F}_Q$ sets the quantum Cramér–Rao bound on how well the magnetic-field parameter can be estimated, and the SNR converts that bound into a practical sensitivity figure. The argument is meant to show that the Hamiltonian and the initial (possibly entangled) state steer the two-qubit system to a point where a small field change produces a large, distinguishable change in the output. This machinery is announced but not exhibited in the submitted text.

What would settle it

Prepare the proposed two-qubit Hamiltonian with both the product and entangled initial states on a two-qubit device, inject calibrated dephasing, and measure the minimum detectable field or SNR; if the entangled protocol does not beat the comparison designs under the same noise, the claimed sensitivity advantage is falsified. At the document level, the absence of any QFI or SNR derivation in the submitted text means the claim is currently uncheckable.

Watch

Extended reading notes

Core claim

The central claim is that a specifically chosen two-qubit Hamiltonian—with its coupling and field-dependent terms arranged to maximize the information a two-qubit state carries about a magnetic field—outperforms existing two-qubit magnetometers in precision, noise resilience, and entanglement behavior. The paper asserts that this is established by analytically deriving the Quantum Fisher Information, which sets the best possible precision for estimating the field, and the signal-to-noise ratio, which connects that bound to a practical readout. It further asserts that starting from a different entangled initial state reveals how entanglement boosts sensitivity. In the manuscript as submitted,

Load-bearing premise

The claimed advantage stands on the assumption that the noise model used in the promised QFI/SNR calculation matches how a real two-qubit magnetic sensor loses coherence, and that the entangled initial state can be prepared without destroying the advantage; the submitted text gives no evidence for either.

Editorial extensions

If this is right

  • If the Hamiltonian performs as claimed, two-qubit magnetometers could resolve smaller magnetic-field increments at a fixed interrogation time.
  • If the noise-resilience claim is correct, the sensor could maintain its precision advantage over longer measurement windows or under stronger decoherence.
  • If the alternate entangled initial state raises $\mathcal{F}_Q$ and SNR, the paper would give a concrete experimental prescription: prepare entanglement rather than a product state.
  • If the comparison against existing models is accurate, this Hamiltonian becomes a reference baseline for future two-qubit magnetic sensing work.
  • The derived QFI and SNR curves, if supplied, would let experimentalists choose operating points (coupling strength, field range, initial state) without exhaustive numerical scanning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A complete version of this work would likely follow the standard quantum-metrology route: evolve the two-qubit state under the proposed Hamiltonian, add a decoherence channel, and evaluate the symmetric logarithmic derivative to get $\mathcal{F}_Q$; the key unstated question is which noise channel was used.
  • The abstract's promise to 'discuss limitations' suggests the authors are aware their noise model may not cover all realistic decoherence; a version that specifies whether bit-flip, phase-flip, or amplitude-damping noise was treated would make the claim testable.
  • A natural experimental check would be to implement both the product and entangled initial states on a superconducting or trapped-ion two-qubit register with calibrated injected dephasing and compare measured SNR; if the entangled advantage is smaller than the paper claims, the noise model is too optimistic.
  • Because the supplied body text is unrelated to the abstract, any reader relying on this submission should treat the magnetometer result as an unverified assertion until the derivation is actually published or posted.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The abstract announces a novel two-qubit quantum magnetometer Hamiltonian 'optimized for enhanced sensitivity and noise resilience,' with analytical derivations of the Quantum Fisher Information (QFI) and Signal-to-Noise Ratio (SNR), a comparative analysis with prior research, an analysis of an entangled initial state, and a discussion of limitations. The full text, however, is an entirely different manuscript—'Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU' (arXiv:2508.13397v2)—containing no quantum magnetometry content whatsoever. No Hamiltonian, no QFI, no SNR, no noise model, no magnetometer baselines, and no limitations discussion appear anywhere in the document. The submission is internally inconsistent; the claimed scientific content is absent.

Significance. If the claimed analytical QFI/SNR derivation and noise-resilience advantage were actually present and correct, a two-qubit quantum magnetometer with these properties could constitute a useful contribution to quantum sensing. However, the submitted document provides no such derivation, no equations, no numerical predictions, no code, and no machine-checked artifacts. The allreduce body contains reproducible benchmarks and speedup measurements, but these bear no relation to the magnetometer claim. Consequently, the significance of the contribution cannot be assessed in its current form: there is no verifiable result to evaluate.

major comments (3)
  1. [Abstract vs. Full Text] The central claim is unverifiable. The abstract (paragraphs 1–2) promises a Hamiltonian 'optimized for enhanced sensitivity and noise resilience' and analytical QFI/SNR derivations. The body contains neither: Sections I–VI discuss MPI all-reduce algorithms and GPU communication only. There is no equation defining a Hamiltonian or a quantum Fisher information; no section, equation, or table can be checked. This is not a local omission but the absence of the paper's claimed content.
  2. [Abstract, paras. 1 and 3] No falsifiable predictions or baselines are provided. The claimed 'advantages in accuracy, robustness against noise, and entanglement dynamics' and the 'comparative analysis with leading research' are unsupported by any quantitative comparison to existing two-qubit magnetometers. The body's only empirical results (Section V) are allreduce speedups, which are unrelated. Thus the asserted 'enhanced sensitivity' cannot be confirmed or refuted from this document.
  3. [Abstract, para. 3] The limitations statement is non-functional. The abstract states that the paper will 'discuss the limitations of our current study,' but no such discussion appears anywhere in the body. The manuscript also fails to specify the Hamiltonian, the entangled initial state, or the noise model, so even the abstract's own program is incomplete.
minor comments (3)
  1. [Metadata] The body carries arXiv identifier 2508.13397v2, whereas the manuscript identifier is 2508.13400. These identifiers should be reconciled, and the correct manuscript should be submitted.
  2. [Title/Body Mismatch] The title 'Enhanced Sensitivity and Noise Resilience in Two-Qubit Quantum Magnetometers' has no counterpart in the body; the body is a high-performance computing systems paper with unrelated figures and references.
  3. [Acknowledgments/References] The acknowledgments and reference list pertain to allreduce optimization and GPU communication (e.g., NCCL, Cray MPICH, MI300A partitioning), which are irrelevant to the quantum magnetometer claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identifiable: the claimed QFI/SNR derivation is entirely absent from the supplied text, so there is no derivation chain to reduce.

full rationale

The supplied document is internally inconsistent: the abstract describes a two-qubit quantum magnetometer Hamiltonian, QFI/SNR derivations, and a comparative analysis (arXiv:2508.13400), while the full text is an unrelated computer-science paper on optimizing Allreduce operations (arXiv:2508.13397v2). The abstract's claim, 'Using analytical methods, we derive the Quantum Fisher Information (QFI) and the Signal-to-Noise Ratio (SNR),' cannot be checked because the full text contains no Hamiltonian, no QFI or SNR equations, no noise model, and no magnetometer baselines. Under the hard rule requiring quotation and exhibition of a specific reduction (e.g., Eq. X = Eq. Y by construction, or a fitted parameter renamed as a prediction), no circular step can be identified: there is no derivation present to walk. It is possible that a Hamiltonian chosen to maximize QFI would later be presented as yielding high QFI by construction, but that pattern cannot be confirmed or excluded from the available text. The appropriate finding is therefore unverifiability—a completeness/correctness failure—not circularity. Accordingly, the circularity score is 0, with no circular steps listed.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The ledger is necessarily sparse because the available evidence is sparse: the abstract names no fitted numbers, and the body text belongs to a different paper. The central claim therefore rests on standard metrology background (QFI and the Cramer-Rao bound) plus three domain assumptions: that the noise model reflects real device physics, that the entangled state is preparable, and that the unnamed comparison baselines are representative. If the actual paper contains hand-tuned coefficients in the Hamiltonian or noise model, they are not visible here and must be audited against the true full text.

assumptions (4)
  • standard math Quantum Fisher information and SNR, as computed from the two-qubit state evolution, are the correct figures of merit for magnetometric sensitivity (quantum Cramer-Rao bound framework).
    The entire sensitivity analysis is framed through QFI and SNR, presupposing standard quantum metrology; stated implicitly in abstract paragraph 1.
  • domain assumption The chosen two-qubit Hamiltonian and decoherence model represent a real magnetic-field sensor well enough that derived 'noise resilience' implies practical viability.
    Abstract claims 'practical viability for magnetic field sensing' without any experimental or numerical validation; this is the load-bearing modeling premise.
  • domain assumption The 'different initial entangled state' used in the second analysis is assumed to be physically preparable and its preparation not to destroy the claimed advantage.
    Abstract paragraph 2 reports an entanglement-state comparison but gives no preparation scheme or state description.
  • domain assumption The 'leading research' baselines used for comparison are representative of the state of the art.
    Abstract paragraph 2 announces a comparative analysis but names no baselines; the improvement claim is only meaningful relative to those unnamed prior models.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enhanced Sensitivity and Noise Resilience in Two-Qubit Quantum Magnetometers." pith.science (2026). https://pith.science/paper/PZXERPRG

@misc{pith2026250813400,
  author       = {Pith},
  title        = {Pith review of: Enhanced Sensitivity and Noise Resilience in Two-Qubit Quantum Magnetometers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PZXERPRG}},
  note         = {Machine review of arXiv:2508.13400}
}
read the original abstract

We present a novel two-qubit quantum magnetometer Hamiltonian optimized for enhanced sensitivity and noise resilience. Compared to existing models, our formulation offers advantages in accuracy, robustness against noise, and entanglement dynamics. Using analytical methods, we derive the Quantum Fisher Information (QFI) and the Signal-to-Noise Ratio (SNR), highlighting its practical viability for magnetic field sensing. Our approach bridges theoretical insights with real-world applicability. We further analyze the performance of the magnetometer with a different initial entangled state, revealing the benefits of entanglement for sensitivity. A comparative analysis with leading research in the field underscores the advancements offered by our proposed design. Finally, we discuss the limitations of our current study and suggest potential avenues for future research.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 39 canonical work pages

  1. [1]

    Optimization of collective reduction operations,

    R. Rabenseifner, “Optimization of collective reduction operations,” in Computational Science - ICCS 2004, M. Bubak, G. D. van Albada, P. M. A. Sloot, and J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2004, pp. 1–9

  2. [2]

    Improving performance models for irregular point-to-point communication,

    A. Bienz, W. D. Gropp, and L. N. Olson, “Improving performance models for irregular point-to-point communication,” inProceedings of the 25th European MPI Users’ Group Meeting, ser. EuroMPI ’18. New York, NY , USA: Association for Computing Machinery, 2018. [Online]. Available: https://doi.org/10.1145/3236367.3236368

  3. [3]

    Decomposing mpi collectives for exploiting multi-lane communication,

    J. L. Traff and S. Hunold, “Decomposing mpi collectives for exploiting multi-lane communication,”2020 IEEE International Conference on Cluster Computing (CLUSTER), 2020

  4. [4]

    Optimization of collective communication operations in mpich,

    R. Thakur, R. Rabenseifner, and W. Gropp, “Optimization of collective communication operations in mpich,”The International Journal of High Performance Computing Applications, vol. 19, no. 1, pp. 49–66,

  5. [5]

    Bandwidth optimal all-reduce algorithms for clusters of workstations,

    P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,”J. Parallel Distrib. Comput., vol. 69, no. 2, p. 117–124, Feb. 2009. [Online]. Available: https://doi.org/10. 1016/j.jpdc.2008.09.002

  6. [6]

    Bringing hpc techniques to deep learning,

    A. Gibiansky, “Bringing hpc techniques to deep learning,”Baidu Re- search, Tech. Rep., 2017

  7. [7]

    Horovod: fast and easy distributed deep learning in tensorflow,

    A. Sergeev and M. D. Balso, “Horovod: fast and easy distributed deep learning in tensorflow,”ArXiv, vol. abs/1802.05799, 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:3398835

  8. [8]

    Partitioned reduction for heteroge- neous environments,

    A. De Rango, G. Utrera, M. Gil, X. Martorell, A. Giordano, D. D’Ambrosio, and G. Mendicino, “Partitioned reduction for heteroge- neous environments,” in2024 32nd Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP), 2024, pp. 285–289

Show all 46 references
  1. [9]

    Modeling data movement performance on heterogeneous architectures,

    A. Bienz, L. N. Olson, W. D. Gropp, and S. Lockhart, “Modeling data movement performance on heterogeneous architectures,” in2021 IEEE High Performance Extreme Computing Conference (HPEC), 2021, pp. 1–7

  2. [10]

    Finepoints: Partitioned multithreaded mpi commu- nication,

    R. E. Grant, M. G. F. Dosanjh, M. J. Levenhagen, R. Brightwell, and A. Skjellum, “Finepoints: Partitioned multithreaded mpi commu- nication,” inHigh Performance Computing, M. Weiland, G. Juckeland, C. Trinitis, and P. Sadayappan, Eds. Cham: Springer International Publishing, 2...

  3. [11]

    AMD Instinct MI300A APU Overview,

    AMD, “AMD Instinct MI300A APU Overview,” https://instinct. docs.amd.com/projects/amdgpu-docs/en/latest/gpu-partitioning/mi300a/ overview.html, 2025, accessed: 2026-01-19

  4. [12]

    A high-performance mpi implementation on a shared-memory vector supercomputer,

    W. Gropp and E. Lusk, “A high-performance mpi implementation on a shared-memory vector supercomputer,”Parallel Computing, vol. 22, no. 11, pp. 1513–1526, 1997. [Online]. Available: https: //www.sciencedirect.com/science/article/pii/S0167819196000622

  5. [13]

    MPI support for multi-core architec- tures: Optimized shared memory collectives,

    R. L. Graham and G. Shipman, “MPI support for multi-core architec- tures: Optimized shared memory collectives,” inRecent Advances in Parallel Virtual Machine and Message Passing Interface, A. Lastovetsky, T. Kechadi, and J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin He...

  6. [14]

    Framework for scalable intra-node collective operations using shared memory,

    S. Jain, R. Kaleem, M. G. Balmana, A. Langer, D. Durnov, A. Sannikov, and M. Garzaran, “Framework for scalable intra-node collective operations using shared memory,” inProceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis...

  7. [15]

    Hierarchical collectives in mpich2,

    H. Zhu, D. Goodell, W. Gropp, and R. Thakur, “Hierarchical collectives in mpich2,” inProceedings of the 16th European PVM/MPI Users’ Group Meeting on Recent Advances in Parallel Virtual Machine and Message Passing Interface. Berlin, Heidelberg: Springer-Verlag, 2009, pp. 325–3...

  8. [16]

    Exploiting hierarchy in parallel computer networks to optimize collective operation performance,

    N. Karonis, B. de Supinski, I. Foster, W. Gropp, E. Lusk, and J. Bres- nahan, “Exploiting hierarchy in parallel computer networks to optimize collective operation performance,” inProceedings 14th International Parallel and Distributed Processing Symposium. IPDPS 2000, 2000, pp...

  9. [17]

    Cheetah: A framework for scalable hierarchical collec- tive operations,

    R. Graham, M. G. Venkata, J. Ladd, P. Shamis, I. Rabinovitz, V . Filipov, and G. Shainer, “Cheetah: A framework for scalable hierarchical collec- tive operations,” in2011 11th IEEE/ACM International Symposium on Cluster , Cloud and Grid Computing, 2011, pp. 73–83

  10. [18]

    Efficient allgather for regular smp-clusters,

    J. L. Träff, “Efficient allgather for regular smp-clusters,” inProceedings of the 13th European PVM/MPI User’s Group Conference on Recent Advances in Parallel Virtual Machine and Message Passing Interface, ser. EuroPVM/MPI’06. Berlin, Heidelberg: Springer-Verlag, 2006, p. 58–6...

  11. [19]

    Designing multi-leader-based allgather algorithms for multi- core clusters,

    K. Kandalla, H. Subramoni, G. Santhanaraman, M. Koop, and D. K. Panda, “Designing multi-leader-based allgather algorithms for multi- core clusters,” in2009 IEEE International Symposium on Parallel & Distributed Processing, 2009, pp. 1–8

  12. [20]

    Scalable reduction collectives with data partitioning-based multi-leader design,

    M. Bayatpour, S. Chakraborty, H. Subramoni, X. Lu, and D. K. D. Panda, “Scalable reduction collectives with data partitioning-based multi-leader design,” inProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’17...

  13. [21]

    Node-aware improvements to allreduce,

    A. Bienz, L. Olson, and W. Gropp, “Node-aware improvements to allreduce,” in2019 IEEE/ACM Workshop on Exascale MPI (ExaMPI), 2019, pp. 19–28

  14. [23]

    Applying on node aggregation methods to mpi alltoall collectives: Matrix block aggregation algorithm,

    G. Chochia, D. Solt, and J. Hursey, “Applying on node aggregation methods to mpi alltoall collectives: Matrix block aggregation algorithm,” inProceedings of the 29th European MPI Users’ Group Meeting, ser. EuroMPI/USA ’22. New York, NY , USA: Association for Computing Machiner...

  15. [24]

    Mapping applications with collectives over sub-communicators on torus networks,

    A. Bhatele, T. Gamblin, S. H. Langer, P.-T. Bremer, E. W. Draeger, B. Hamann, K. E. Isaacs, A. G. Landge, J. A. Levine, V . Pascucci, M. Schulz, and C. H. Still, “Mapping applications with collectives over sub-communicators on torus networks,” inProceedings of the Internationa...

  16. [25]

    Massively distributed sgd: Imagenet/resnet-50 training in a flash,

    H. Mikami, H. Suganuma, P. U-chupala, Y . Tanaka, and Y . Kageyama, “Massively distributed sgd: Imagenet/resnet-50 training in a flash,” 2018

  17. [26]

    Mpi applications on grids: A topology aware approach,

    C. Coti, T. Herault, and F. Cappello, “Mpi applications on grids: A topology aware approach,” inEuropean Conference on Parallel Processing. Springer, 2009, pp. 466–477

  18. [27]

    Topology-aware rank reordering for mpi collectives,

    S. H. Mirsadeghi and A. Afsahi, “Topology-aware rank reordering for mpi collectives,” in2016 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), 2016, pp. 1759–1768

  19. [28]

    Process distance- aware adaptive mpi collective communications,

    T. Ma, T. Herault, G. Bosilca, and J. J. Dongarra, “Process distance- aware adaptive mpi collective communications,” in2011 IEEE Interna- tional Conference on Cluster Computing, 2011, pp. 196–204

  20. [29]

    Bandwidth efficient all-reduce operation on tree topologies,

    P. Patarasuk and X. Yuan, “Bandwidth efficient all-reduce operation on tree topologies,” in2007 IEEE International Parallel and Distributed Processing Symposium, March 2007, pp. 1–8

  21. [30]

    Scalable hierarchical aggregation and reduc- tion protocol (sharp)tm streaming-aggregation hardware design and evaluation,

    R. L. Graham, L. Levi, D. Burredy, G. Bloch, G. Shainer, D. Cho, G. Elias, D. Klein, J. Ladd, O. Maor, A. Marelli, V . Petrov, E. Romlet, Y . Qin, and I. Zemah, “Scalable hierarchical aggregation and reduc- tion protocol (sharp)tm streaming-aggregation hardware design and eval...

  22. [31]

    Hierarchical distributed- memory multi-leader mpi-allreduce for deep learning workloads,

    T. Thao Nguyen, M. Wahib, and R. Takano, “Hierarchical distributed- memory multi-leader mpi-allreduce for deep learning workloads,” in 2018 Sixth International Symposium on Computing and Networking Workshops (CANDARW), 2018, pp. 216–222

  23. [32]

    Hiccl: A hierarchical collective communi- cation library,

    M. Hidayetoglu, S. G. de Gonzalo, E. Slaughter, P. Surana, W.-m. Hwu, W. Gropp, and A. Aiken, “Hiccl: A hierarchical collective communi- cation library,” in2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2025, pp. 950–961

  24. [33]

    Dma-assisted, intranode communication in gpu accelerated systems,

    F. Ji, A. M. Aji, J. Dinan, D. Buntinas, P. Balaji, R. Thakur, W.- c. Feng, and X. Ma, “Dma-assisted, intranode communication in gpu accelerated systems,” in2012 IEEE 14th International Conference on High Performance Computing and Communication & 2012 IEEE 9th International Co...

  25. [34]

    Optimizing mpi communication on multi-gpu systems using cuda inter- process communication,

    S. Potluri, H. Wang, D. Bureddy, A. Singh, C. Rosales, and D. K. Panda, “Optimizing mpi communication on multi-gpu systems using cuda inter- process communication,” in2012 IEEE 26th International Parallel and Distributed Processing Symposium Workshops & PhD F orum, 2012, pp. 1848–1857

  26. [35]

    Mvapich2-gpu: optimized gpu to gpu communication for infiniband clusters,

    H. Wang, S. Potluri, M. Luo, A. K. Singh, S. Sur, and D. K. Panda, “Mvapich2-gpu: optimized gpu to gpu communication for infiniband clusters,”Comput. Sci., vol. 26, no. 3–4, p. 257–266, Jun. 2011. [Online]. Available: https://doi.org/10.1007/s00450-011-0171-3

  27. [36]

    Cuda kernel based collective reduction operations on large-scale gpu clusters,

    C.-H. Chu, K. Hamidouche, A. Venkatesh, A. A. Awan, and D. K. Panda, “Cuda kernel based collective reduction operations on large-scale gpu clusters,” in2016 16th IEEE/ACM International Symposium on Cluster , Cloud and Grid Computing (CCGrid), 2016, pp. 726–735

  28. [37]

    Understanding gpu-based lossy compression for extreme- scale cosmological simulations,

    S. Jin, P. Grosset, C. M. Biwer, J. Pulido, J. Tian, D. Tao, and J. Ahrens, “Understanding gpu-based lossy compression for extreme- scale cosmological simulations,” in2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, May 2020, p. 105–115. [On...

  29. [38]

    Accelerating mpi all-to-all communication with online compression on modern gpu clusters,

    Q. Zhou, P. Kousha, Q. Anthony, K. Shafie Khorassani, A. Shafi, H. Sub- ramoni, and D. K. Panda, “Accelerating mpi all-to-all communication with online compression on modern gpu clusters,” inHigh Performance Computing, A.-L. Varbanescu, A. Bhatele, P. Luszczek, and B. Marc, Ed...

  30. [39]

    Accelerating mpi allreduce communication with efficient gpu-based compression schemes on modern gpu clusters,

    Q. Zhou, B. Ramesh, A. Shafi, M. Abduljabbar, H. Subramoni, and D. K. Panda, “Accelerating mpi allreduce communication with efficient gpu-based compression schemes on modern gpu clusters,” inISC High Performance 2024 Research Paper Proceedings (39th International Conference), ...

  31. [40]

    Zccl: Significantly improving collective communication with error-bounded lossy compression,

    J. Huang, S. Di, X. Yu, Y . Zhai, Z. Zhang, J. Liu, X. Lu, K. Raffenetti, H. Zhou, K. Zhao, K. Alharthi, Z. Chen, F. Cappello, Y . Guo, and R. Thakur, “Zccl: Significantly improving collective communication with error-bounded lossy compression,” 2025. [Online]. Available: http...

  32. [41]

    NCCL (NVIDIA Collective Communications Library),

    NVIDIA, “NCCL (NVIDIA Collective Communications Library),” https://github.com/NVIDIA/nccl, 2025, accessed: 2025-08-07

  33. [42]

    Rocm communication collectives library (rccl),

    AMD, “Rocm communication collectives library (rccl),” https://github. com/ROCm/rccl, 2025, accessed: 2025-08-07

  34. [43]

    Does a single AMD GPU can be shared among containers? Things like MPS of Nvidia,

    Felix Kuehling, “Does a single AMD GPU can be shared among containers? Things like MPS of Nvidia,” https://github.com/ROCm/ ROCm-docker/issues/62, 2019, accessed: 2026-01-13

  35. [44]

    Inter-apu communication on amd mi300a systems via infinity fabric: a deep dive,

    G. Schieffer, J. Wahlgren, R. Shi, E. A. León, R. Pearce, M. Gokhale, and I. Peng, “Inter-apu communication on amd mi300a systems via infinity fabric: a deep dive,” 2025. [Online]. Available: https://arxiv.org/abs/2508.11298

  36. [45]

    Roofline analysis of tightly-coupled cpu-gpu superchips: A study on mi300a and gh200,

    O. Antepara, L. Oliker, and S. Williams, “Roofline analysis of tightly-coupled cpu-gpu superchips: A study on mi300a and gh200,” in Proceedings of the SC ’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC Wor...

  37. [2005]

    Available: https://doi.org/10.1177/1094342005051521

    [Online]. Available: https://doi.org/10.1177/1094342005051521

  38. [2017]

    Available: https://doi.org/10.1145/3126908.3126954

    [Online]. Available: https://doi.org/10.1145/3126908.3126954

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.