REVIEW 5 major objections 5 minor 64 references
Scalable parallel simulation of quantum circuits on CPU and GPU systems
T0 review · 5 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that Q2Chemistry, a full-amplitude quantum circuit simulator, consistently outperforms nine open-source simulators across QFT, VQE-HEA, and QAOA circuits by combining three parallel optimizations, achieving up to 4.52x CPU
desk verdict Plausible engineering optimizations, but the 'consistently outperforms' claim outruns the evidence—no correctness check, and benchmark methodology is too thin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is hybrid-level parallel scheduling built on the statevector's bit-index structure: amplitudes that differ only in one target-qubit bit form update pairs, and the position of that bit relative to each MPI rank's local qubits decides whether a gate is local or needs communication. Three named optimizations carry the argument: BBOP (batch-buffered overlap processing) pipelines MPI sends and receives through double buffering to hide communication latency; SMGP (staggered multi-gate parallelism) maps amplitude pairs onto two-dimensional GPU thread blocks so different gates work on different segments in a staggered cycle, avoiding conflicts and raising throughput; and DAGC (
What would settle it
Run the same 20-30 qubit QFT, VQE-HEA, and QAOA circuits on the same hardware with absolute wall-clock timings, repeated runs, and error bars, using each competing simulator's own recommended settings (for example, Qiskit with its optimized Aer backend and QuEST with its suggested thread/MPI configuration). If any competitor matches or beats Q2Chemistry on QAOA or QFT at 30 qubits, or if UCCSD-style circuits show the speedups largely vanish, the claim that Q2Chemistry consistently outperforms current state-of-the-art simulators would be falsified.
Extended reading notes
Core claim
The paper's central claim is that the bottleneck in distributed full-amplitude quantum simulation is data movement and sequential gate execution, not raw floating-point throughput, and that these can be attacked at three levels simultaneously. BBOP splits the local state vector into batches and uses multi-buffered non-blocking MPI so communication overlaps with computation, reducing communication-bound gate time by up to 76.5%. SMGP reorganizes GPU thread blocks from one dimension to two, assigning different gates to different amplitude segments in a cyclic stagger so memory-level parallelism rises and write conflicts fall, yielding 3.35x average speedup on QAOA and 4.96x on VQE-HEA GPU kern
Load-bearing premise
The central performance claim rests on the cross-simulator benchmarks in Section 4.4 being fair and representative: each competing simulator must have been run in a comparably optimized configuration, and the three circuit families must adequately stand in for 'various circuit types'.
Editorial extensions
If this is right
- Q2Chemistry can run a 30-qubit VQE-HEA simulation in about 20 seconds on a 64-thread CPU cluster, compared with roughly 91 seconds for its unoptimized baseline.
- On a four-A100 GPU node, the paper reports up to 13.44x speedup over QuEST for 30-qubit QFT circuits, the largest multi-GPU gap among tested simulators.
- The optimizations are complementary: BBOP gives the largest gains on CPU clusters, SMGP on GPUs, and DAGC helps both by shrinking gate counts and communication events.
- The batched communication scheme reduces the working-memory overhead of distributed simulation, potentially allowing more qubits per fixed memory budget at the cost of some wall-clock time.
Reading between the lines
- Because DAGC's compression grows with the density of single-qubit gates, circuits even richer in rotations than VQE-HEA—such as UCCSD ansätze—may show larger fusion gains than the 52-63% reported here.
- The paper's own ablation shows BBOP contributes little in GPU environments where communication is about 99% of time, suggesting the optimization recipe should be selected automatically from the machine's compute-to-bandwidth ratio rather than applied uniformly.
- A natural extension is to test whether these three passes, described in hardware-agnostic terms, transfer to other statevector simulators that use the same amplitude-pair update pattern, and whether they preserve numerical fidelity on deeper circuits.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript describes a parallelization and optimization effort for the Q2Chemistry full-amplitude quantum circuit simulator. It introduces three techniques: Batch-Buffered Overlap Processing (BBOP) to overlap MPI communication with computation via multi-buffering; Staggered Multi-Gate Parallelism (SMGP) to execute multiple gates concurrently on GPU blocks using 2D thread mappings; and Dependency-Aware Gate Contraction (DAGC), a DAG-based greedy algorithm for fusing independent gates. Benchmarks are reported for QFT, QAOA, and VQE-HEA circuits on a CPU cluster and A100 GPU nodes, with claims of up to 4.52x CPU and 3.57x GPU speedups for 30-qubit VQE-HEA, and 3.01x CPU and 2.66x GPU speedups for 30-qubit QAOA, as well as favorable comparisons with nine open-source simulators.
Significance. If the results hold, the main value is a practical, open-source engineering contribution with a combination of distributed-memory and GPU optimizations that are not all present in a single existing package. The paper's transparency is strengthened by the Zenodo code deposit and the breadth of benchmark circuits and simulators. However, the paper's central quantitative claims -- 'consistently outperforms' and specific speedups -- are supported only by single measurements with no correctness validation, so the significance cannot be assessed at present. The work is not circular: no parameters are fitted to benchmarks, and the techniques are direct algorithmic modifications. The correctness and methodology gaps below are the main obstacles.
major comments (5)
- [Sec. 4.3 and Sec. 4.4] The benchmark sections report only execution times for the optimized configurations. Nowhere is the final statevector (or any derived observable) compared against the unoptimized baseline, an exact reference, or another simulator. Since SMGP and DAGC change the order and grouping of gate operations, a bug in either could yield large speedups while producing incorrect output. The claim that Q2Chemistry 'consistently outperforms' is a performance claim about a correct simulator, so a validation subsection (e.g., L2 norm of the difference between optimized and reference statevectors for random circuits from 20 to 30 qubits) is load-bearing and must be added.
- [Sec. 3.3 (Table 3)] Table 3 illustrates staggered execution of gates G0-G3 on overlapping segments of the same statevector. The schedule may be conflict-free, but the manuscript provides no proof or empirical check. If two gates write to the same amplitude pair in different time steps without synchronization, write-after-read hazards corrupt the state. The authors should either give a formal argument for conflict-freedom (based on index arithmetic) or provide a randomized test that compares SMGP output with sequential gate order.
- [Sec. 3.4 (Eq. 7)] Eq. (7) defines the fused two-qubit gate as M = M2 ⊗ M1. This is only valid when the target qubits appear in the order used by the Kronecker product (t2 more significant than t1). For t1 > t2, or when t1 and t2 are not adjacent in the qubit ordering, the 4x4 matrix must be permuted. The text does not state how DAGC handles this case. If the implementation follows Eq. (7) literally, the fused gate is not the simultaneous application of the two gates. Please clarify the permutation rule or adjust the equation.
- [Sec. 4.3 (Figs. 15, 20)] The ablation numbers are internally inconsistent. Sec. 4.3 (Fig. 15) reports SMGP-only average speedups of 3.35x (QAOA) and 4.96x (VQE-HEA) on GPU, while the combined SMGP+DAGC results in Fig. 20 give 30-qubit speedups of 2.66x and 3.57x for the same circuit types. A combined optimization should not be slower than one of its components unless the baselines, hardware counts, or measurement procedures differ between the figures. The authors should provide absolute execution times and ensure the baseline is identical across all ablation and final-benchmark runs.
- [Sec. 4.4 (Figs. 21-23)] Sec. 4.4 compares normalized runtimes only; no absolute wall-clock times, number of repetitions, variance, or error bars are reported for any simulator. It is also unclear whether each competing package was run in its recommended optimized configuration (e.g., Qiskit's optimization level, Qulacs's OpenMP settings, QuEST's MPI+OpenMP, etc.) with the same node/thread allocation. The claim of 'consistent' outperformance is stronger than the data shown. Report absolute times (or at least a supplementary table), the optimization/compiler flags used for each package, and the run-to-run variability.
minor comments (5)
- [Abstract and Sec. 1] 'comprise accuracy' should be 'compromise accuracy'.
- [Sec. 3.4, Eq. (6)] The second component of the output vector in Eq. (6) is typeset identically to the first; it should be α′_{*0_{t2}*1_{t1}*}.
- [Sec. 4.3, BBOP GPU paragraph] 'the 2.66× speedup observed in CPU clusters' seems misstated; 2.66× is the combined GPU QAOA speedup in Fig. 20. Clarify which number is meant.
- [Fig. 17] The secondary axis label 'Compression Ratio' is plotted on the same panel; make the axis assignment explicit for readability.
- [References [10], [56]] Reference [10] has a duplicated URL string and reference [56] has garbled author formatting. These need cleanup.
Circularity Check
No circularity: all speedups are direct measurements of engineering optimizations; no fitted parameter is renamed as a prediction.
full rationale
The paper's claims are performance measurements, not derived predictions. BBOP, SMGP, and DAGC are described as implementation strategies whose effects are measured before/after (Figures 14-20), not computed from a model fitted to the benchmark data. The cross-software benchmark simply reports normalized runtimes against other simulators; normalizing to Q2Chemistry = 1 is a plotting convention, not a construction of the speedup. The paper explicitly reports negative/limited results for BBOP on GPU (Table 4), which is incompatible with a circular argument that optimizations must succeed. The only self-citation ([55], the Q2Chemistry package) identifies the software being optimized and is not invoked as evidence for any optimization's validity. No uniqueness theorem, ansatz, or fitted constant is imported from prior work. Potential concerns about correctness verification (e.g., no statevector comparison against a reference) are correctness risks, not circularity, since they do not make the measured speedup equal to an input by construction.
Assumptions & free parameters
free parameters (2)
- BBOP batch size exponent b =
not reported
- Number of pipeline buffers in BBOP =
not reported
assumptions (4)
- standard math The amplitude-pair update formulas for single-qubit and controlled gates correctly implement unitary evolution.
- domain assumption The distributed statevector partition, with high qubits encoding MPI rank, makes the peer process formula Eq. (5) valid.
- ad hoc to paper Uniform time overhead across arithmetic, data copy, and memory operations in the DAGC cost comparison.
- domain assumption The compared open-source simulators were run in fair, comparably optimized configurations.
Cite this review
Pith. "Pith review of Scalable parallel simulation of quantum circuits on CPU and GPU systems." pith.science (2026). https://pith.science/paper/MTVMH3M4
@misc{pith2026250904955,
author = {Pith},
title = {Pith review of: Scalable parallel simulation of quantum circuits on CPU and GPU systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/MTVMH3M4}},
note = {Machine review of arXiv:2509.04955}
}
abstract
Quantum computing enables parallelism through superposition and entanglement and offers advantages over classical computing architectures. However, due to the limitations of current quantum hardware in the noisy intermediate-scale quantum (NISQ) era, classical simulation remains a critical tool for developing quantum algorithms. In this research, we present a comprehensive parallelization solution for the Q$^2$Chemistry software package, delivering significant performance improvements for the full-amplitude simulator on both CPU and GPU platforms. By incorporating batch-buffered overlap processing, dependency-aware gate contraction and staggered multi-gate parallelism, our optimizations significantly enhance the simulation speed compared to unoptimized baselines, demonstrating the effectiveness of hybrid-level parallelism in HPC systems. Benchmark results show that Q$^2$Chemistry consistently outperforms current state-of-the-art open-source simulators across various circuit types. These benchmarks highlight the capability of Q$^2$Chemistry to effectively handle large-scale quantum simulations with high efficiency and high portability.
Figures
Figures from the paper (20 more)
Reference graph
Works this paper leans on
-
[1]
P. Hohenberg and W. Kohn. Inhomogeneous electron gas.Phys. Rev., 136:B864–B871, 1964. DOI: 10.1103/PhysRev.136.B864. URLhttps: //link.aps.org/doi/10.1103/PhysRev.136.B864
-
[2]
W. Kohn and L. J. Sham. Self-consistent equations including exchange and correlation effects.Phys. Rev., 140:A1133–A1138, 1965. DOI: 10.1103/PhysRev.140.A1133. URLhttps://link.aps.org/doi/10. 1103/PhysRev.140.A1133
-
[3]
W. Kohn, A. D. Becke, and R. G. Parr. Density functional theory of electronic structure.J. Phys. Chem., 100:12974–12980, 1996. DOI: 10.1021/jp960669l. URLhttps://doi.org/10.1021/jp960669l
-
[4]
A. J. Cohen, P. Mori-Sanchez, and W. Yang. Challenges for den- sity functional theory.Chem. Rev., 112:289–320, 2012. DOI: 10.1021/cr200107z. URLhttps://doi.org/10.1021/cr200107z
-
[5]
Vogiatzis, Dongxia Ma, Jeppe Olsen, Laura Gagliardi, and Wibe A
Konstantinos D. Vogiatzis, Dongxia Ma, Jeppe Olsen, Laura Gagliardi, and Wibe A. de Jong. Pushing configuration-interaction to the limit: Towards massively parallel mcscf calculations.The Journal of Chemical Physics, 147(18):184111, 11 2017. ISSN 0021-9606. DOI: 10.1063/1.4989858. URLhttps://doi.org/10.1063/1.4989858
-
[6]
Richard P. Feynman. Simulating physics with computers.Int. J. Theor. Phys., 21(6):467–488, Jun 1982. ISSN 1572-9575. DOI: 10.1007/BF02650179. URLhttps://doi.org/10.1007/BF02650179
-
[7]
Peter W. Shor. Polynomial-time algorithms for prime factorization and discrete logarithms on a quantum computer.SIAM Journal on Com- puting, 26(5):1484–1509, 1997. DOI: 10.1137/S0097539795293172. URL https://doi.org/10.1137/S0097539795293172
-
[8]
Jiangfeng Du, Nanyang Xu, Xinhua Peng, Pengfei Wang, Sanfeng Wu, and Dawei Lu. Nmr implementation of a molecular hydrogen quan- tum simulation with adiabatic state preparation.Phys. Rev. Lett., 104:030502, Jan 2010. DOI: 10.1103/PhysRevLett.104.030502. URL https://link.aps.org/doi/10.1103/PhysRevLett.104.030502
Show all 64 references
-
[9]
Bardin, Rami Barends, Rupak Biswas, Sergio Boixo, Fernando G
Frank Arute, Kunal Arya, Ryan Babbush, Dave Bacon, Joseph C. Bardin, Rami Barends, Rupak Biswas, Sergio Boixo, Fernando G. S. L. Brandão, David A. Buell, Brian Burkett, Yu Chen, Zijun Chen, Ben- jamin Chiaro, Roberto Collins, William Courtney, Andrew Dunsworth, Edward Farhi, B...
2019
-
[10]
The varia- tional quantum eigensolver: a review of methods and best practices,
Jules Tilly, Hongxiang Chen, Shuxiang Cao, et al. The varia- tional quantum eigensolver: a review of methods and best practices,
-
[11]
Cerezo, Andrew Arrasmith, Ryan Babbush, et al
M. Cerezo, Andrew Arrasmith, Ryan Babbush, et al. Variational quan- tum algorithms.Nat. Rev. Phys., 3(9):625–644, Sep 2021. ISSN 2522-
2021
-
[12]
Magann, Christian Arenz, Matthew D
Alicia B. Magann, Christian Arenz, Matthew D. Grace, et al. From pulses to circuits and back again: A quantum optimal control per- spective on variational quantum algorithms.Phys. Rev. X. Quan- tum, 2:010101, Jan 2021. DOI: 10.1103/PRXQuantum.2.010101. URL https://link.aps.org...
2021 doi
-
[13]
Fedorov, Bo Peng, Niranjan Govind, et al
Dmitry A. Fedorov, Bo Peng, Niranjan Govind, et al. Vqe method: a short survey and recent developments.Materials Theory, 6(1):2, Jan
-
[14]
S. B. Bravyi and A. Y. Kitaev. Fermionic quantum computation.Ann. Phys., 298:210–226, 2002. DOI: https://doi.org/10.1006/aphy.2002.6254. URLhttps://www. sciencedirect.com/science/article/pii/S0003491602962548
2002
-
[15]
Quantum computational chemistry.Rev
Sam McArdle, Suguru Endo, Alán Aspuru-Guzik, et al. Quantum computational chemistry.Rev. Mod. Phys., 92:015003, 2020. DOI: 10.1103/RevModPhys.92.015003. URLhttps://link.aps.org/doi/ 10.1103/RevModPhys.92.015003
2020 doi
-
[16]
Olson, et al
Yudong Cao, Jonathan Romero, Jonathan P. Olson, et al. Quantum chemistry in the age of quantum computing.Chem. Rev., 119:10856– 10915, 2019. DOI: 10.1021/acs.chemrev.8b00803. URLhttps://doi. org/10.1021/acs.chemrev.8b00803
2019 doi
-
[17]
Quantum computing in the nisq era and beyond.Quan- tum, 2:79, August 2018
John Preskill. Quantum computing in the nisq era and beyond.Quan- tum, 2:79, August 2018. ISSN 2521-327X. DOI: 10.22331/q-2018-08-06- 79
2018 doi
-
[18]
I. M. Georgescu, S. Ashhab, and Franco Nori. Quantum simulation. Rev. Mod. Phys., 86:153–185, 2014. DOI: 10.1103/RevModPhys.86.153. URLhttps://link.aps.org/doi/10.1103/RevModPhys.86.153
2014 doi
-
[19]
Aspuru-Guzik, A
A. Aspuru-Guzik, A. D. Dutoi, P. J. Love, et al. Simulated quan- tum computation of molecular energies.Science, 309:1704–1707, 2005. DOI: 10.1126/science.1113479. URLhttps://science.sciencemag. org/content/309/5741/1704
2005 doi
-
[20]
Quantum algo- rithm for obtaining the energy spectrum of molecular systems.Phys
Hefeng Wang, Sabre Kais, Alán Aspuru-Guzik, et al. Quantum algo- rithm for obtaining the energy spectrum of molecular systems.Phys. Chem. Chem. Phys., 10:5388–5393, 2008. DOI: 10.1039/B804804E. URLhttp://dx.doi.org/10.1039/B804804E
2008 doi
-
[21]
Peruzzo, J
A. Peruzzo, J. McClean, P. Shadbolt, et al. A variational eigen- value solver on a photonic quantum processor.Nat. Commun., 5:4213,
-
[22]
Hempel, C
C. Hempel, C. Maier, J. Romero, et al. Quantum chemistry calculations on a trapped-ion quantum simulator.Phys. Rev. X, 8:031022, 2018. DOI: 10.1103/PhysRevX.8.031022. URLhttps://link.aps.org/doi/ 10.1103/PhysRevX.8.031022
2018 doi
-
[23]
Pisenti, et al
Yunseong Nam, Jwo Sy Chen, Neal C. Pisenti, et al. Ground-state en- ergy estimation of the water molecule on a trapped-ion quantum com- puter.Npj Quantum Inf., 6:33, 2020. DOI: 10.1038/s41534-020-0259-3. URLhttps://doi.org/10.1038/s41534-020-0259-3
2020 doi
-
[24]
Y. Shen, X. Zhang, S. Zhang, et al. Quantum implementation of the unitary coupled cluster for simulating molecular electronic struc- ture.Phys. Rev. A: At., Mol., Opt. Phys., 95:020501, 2017. DOI: 10.1103/PhysRevA.95.020501. URLhttps://link.aps.org/doi/10. 1103/PhysRevA.95.020501
2017 doi
-
[25]
P. J. J. O’ Malley, R. Babbush, I. D. Kivlichan, et al. Scalable quantum simulation of molecular energies.Phys. Rev. X, 6:031007, 2016. DOI: 10.1103/PhysRevX.6.031007. URLhttps://link.aps.org/doi/10. 1103/PhysRevX.6.031007
2016 doi
-
[26]
Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets.Nature, 549(7671):242–246, Sep 2017
Abhinav Kandala, Antonio Mezzacapo, Kristan Temme, et al. Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets.Nature, 549(7671):242–246, Sep 2017. ISSN 1476-4687. DOI: 10.1038/nature23879. URLhttps://doi.org/10. 1038/nature23879
2017 doi
-
[27]
J. I. Colless, V. V. Ramasesh, D. Dahlen, et al. Computation of molec- ular spectra on a quantum processor with an error-resilient algorithm. Phys. Rev. X, 8:011021, 2018. DOI: 10.1103/PhysRevX.8.011021. URL https://link.aps.org/doi/10.1103/PhysRevX.8.011021
2018 doi
-
[28]
J. R. McClean, J. Romero, R. Babbush, et al. The theory of varia- tional hybrid quantum-classical algorithms.New J. Phys., 18:023023,
-
[29]
B. P. Lanyon, J. D. Whitfield, G. G. Gillett, et al. Towards quantum chemistry on a quantum computer.Nat. Chem., 2:106–111, 2010. DOI: 10.1038/nchem.483. URLhttps://doi.org/10.1038/nchem.483
2010 doi
-
[31]
Variational Quantum Computation of Excited States.Quantum, 3:156, 2019
Oscar Higgott, Daochen Wang, and Stephen Brierley. Variational Quantum Computation of Excited States.Quantum, 3:156, 2019. DOI: 10.22331/q-2019-07-01-156. URLhttps://doi.org/10.22331/ q-2019-07-01-156
2019 doi
-
[32]
McClean, Mollie E
Jarrod R. McClean, Mollie E. Kimchi-Schwartz, Jonathan Carter, et al. Hybrid quantum-classical hierarchy for mitigation of decoherence and determination of excited states.Phys. Rev. A, 95:042308, Apr 2017. DOI: 10.1103/PhysRevA.95.042308. URLhttps://link.aps.org/ doi/10.1103/P...
2017 doi
-
[33]
Casanova, A
Man Hong Yung, J. Casanova, A. Mezzacapo, et al. From transistor to trapped-ion computers for quantum chemistry.Sci. Rep., 4(1):3589, Jan 2014. DOI: 10.1038/srep03589. URLhttps://doi.org/10.1038/ srep03589
2014 doi
-
[34]
A quantum ap- proximate optimization algorithm, 2014
Edward Farhi, Jeffrey Goldstone, and Sam Gutmann. A quantum ap- proximate optimization algorithm, 2014. URLhttps://arxiv.org/ abs/1411.4028
2014 arXiv
-
[35]
R. J. Bartlett, S. A. Kucharski, and J. Noga. Alternative coupled-cluster ansätze ii. the unitary coupled-cluster method.Chem. Phys. Lett., 155: 133–140, 1989. DOI: https://doi.org/10.1016/S0009-2614(89)87372-
1989 doi
-
[36]
A. G. Taube and R. J. Bartlett. New perspectives on unitary coupled- cluster theory.Int. J. Quantum Chem., 106:3393–3401, 2006. DOI: https://doi.org/10.1002/qua.21198. URLhttps://onlinelibrary. wiley.com/doi/abs/10.1002/qua.21198
2006 doi
-
[37]
Differentiable matrix product states for simulating variational quantum computa- tional chemistry.Quantum, 7:1192, December 2023
Chu Guo, Yi Fan, Zhiqian Xu, and Honghui Shang. Differentiable matrix product states for simulating variational quantum computa- tional chemistry.Quantum, 7:1192, December 2023. ISSN 2521-327X. DOI: 10.22331/q-2023-12-04-1192. URLhttps://doi.org/10.22331/ q-2023-12-04-1192
2023 doi
-
[38]
G. Vidal. Classical simulation of infinite-size quantum lattice systems in one spatial dimension.Phys. Rev. Lett., 98:070201, Feb 2007. DOI: 10.1103/PhysRevLett.98.070201. URLhttps://link.aps.org/doi/ 10.1103/PhysRevLett.98.070201
2007 doi
-
[39]
M. B. Hastings. Light-cone matrix product.Journal of Mathe- matical Physics, 50(9):095207, 06 2009. ISSN 0022-2488. DOI: 10.1063/1.3149556. URLhttps://doi.org/10.1063/1.3149556
2009 doi
-
[40]
Qibo: a framework for quantum sim- ulation with hardware acceleration.Quantum Science and Technology, 7(1):015018, December 2021
Stavros Efthymiou, Sergi Ramos-Calderer, Carlos Bravo-Prieto, Adrián Pérez-Salinas, Diego García-Martín, Artur Garcia-Saez, José Ignacio Latorre, and Stefano Carrazza. Qibo: a framework for quantum sim- ulation with hardware acceleration.Quantum Science and Technology, 7(1):01...
2021 doi
-
[41]
URLhttps://www.sciencedirect.com/science/article/pii/ S0009261489873725
-
[42]
Putra, Takashi Imamichi, and Hiroshi Horii
Jun Doi, Hitomi Takahashi, Raymond H. Putra, Takashi Imamichi, and Hiroshi Horii. Quantum computing simulator on a heterogenous hpc system.Proceedings of the 16th ACM International Conference on Computing Frontiers, 2019. DOI: 10.1145/3310273.3323053. URL https://api.semantics...
2019
-
[43]
Trenas, and Emilio L
Eladio Gutierrez, Sergio Romero, Maria A. Trenas, and Emilio L. Zap- ata. Simulation of quantum gates on a novel gpu architecture. InPro- ceedings of the 7th WSEAS International Conference on Systems The- ory and Scientific Computation, pages 121–126, Stevens Point, Wiscon- si...
2007 doi
-
[44]
Trenas, and Emilio L
Eladio Gutiérrez, Sergio Romero, María A. Trenas, and Emilio L. Zapata. Parallel quantum computer simulation on the cuda ar- chitecture. InInternational Conference on Conceptual Structures,
-
[45]
Thomas Häner and Damian S. Steiger. 0.5 petabyte simulation of a 45- qubit quantum circuit. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’17, New York, NY, USA, 2017. Association for Computing Ma- chinery. I...
2017
-
[46]
De Raedt, K
K. De Raedt, K. Michielsen, H. De Raedt, B. Trieu, G. Arnold, M. Richter, Th. Lippert, H. Watanabe, and N. Ito. Massively paral- lel quantum computer simulator.Computer Physics Communications, 176(2):121–136, 2007. ISSN 0010-4655. DOI: 10.1016/j.cpc.2006.08.007. URLhttp://dx.d...
2007 doi
-
[47]
Gpu-aware distributed quantum simulation
Anderson Avila, Adriano Maron, Renata Reiser, Maurício Pilla, and Adenauer Yamin. Gpu-aware distributed quantum simulation. InPro- ceedings of the 29th Annual ACM Symposium on Applied Comput- ing, pages 893–900, New York, NY, USA, March 2014. ACM. DOI: 10.1145/2554850.2554892
2014
-
[48]
Advanced simulation of quantum computations, 2018
Alwin Zulehner and Robert Wille. Advanced simulation of quantum computations, 2018. URLhttps://arxiv.org/abs/1707.00865
2018 arXiv
-
[49]
Xin-Chuan Wu, Sheng Di, Emma Maitreyee Dasgupta, Franck Cap- pello, Hal Finkel, Yuri Alexeev, and Frederic T. Chong. Full-state quantum circuit simulation by using data compression. InProceedings of the International Conference for High Performance Computing, Net- working, Sto...
2019
-
[50]
Chow, Antonio D
Gadi Aleksandrowicz, Thomas Alexander, Panagiotis Barkoutsos, Lu- ciano Bello, Yael Ben-Haim, David Bucher, Francisco Jose Cabrera- Hernández, Jorge Carballo-Franquis, Adrian Chen, Chun-Fu Chen, Jerry M. Chow, Antonio D. Córcoles-Gonzales, Abigail J. Cross, An- drew Cross, Jua...
2019
-
[51]
Quantum computer simu- lation on multi-gpu incorporating data locality
Pei Zhang, Jiabin Yuan, and Xiangwen Lu. Quantum computer simu- lation on multi-gpu incorporating data locality. InAlgorithms and Ar- chitectures for Parallel Processing (ICA3PP 2015), volume 9528, pages 241–256, 11 2015. ISBN 978-3-319-27118-7. DOI: 10.1007/978-3-319- 27119-4_17
2015 doi
-
[52]
Cirq, April 2022
Cirq Developers. Cirq, April 2022. URLhttps://zenodo.org/doi/ 10.5281/zenodo.4062499. Version v0.14.1; accessed 2022-08-01
2022 doi
-
[53]
Nakanishi, Kosuke Mitarai, Ryosuke Imai, Shiro Tamiya, Takahiro Yamamoto, Tennin Yan, Toru Kawakubo, Yuya O
Yasunari Suzuki, Yoshiaki Kawase, Yuya Masumura, Yuria Hiraga, Masahiro Nakadai, Jiabao Chen, Ken M. Nakanishi, Kosuke Mitarai, Ryosuke Imai, Shiro Tamiya, Takahiro Yamamoto, Tennin Yan, Toru Kawakubo, Yuya O. Nakagawa, Yohei Ibe, Youyuan Zhang, Hirotsugu Yamashita, Hikaru Yos...
2021 doi
-
[54]
Mikhail Smelyanskiy, Nicolas P. D. Sawaya, and Alán Aspuru-Guzik. qhipster: The quantum high performance software testing environment, 2016
2016
-
[55]
Q2chemistry: A quantum computation platform for quantum chemistry.JUSTC, 52(12):2, 2022
Yi Fan, Jie Liu, Xiongzhi Zeng, Zhiqian Xu, Honghui Shang, Zhenyu Li, and Jinlong Yang. Q2chemistry: A quantum computation platform for quantum chemistry.JUSTC, 52(12):2, 2022. DOI: 10.52396/JUSTC- 2022-0118
2022 doi
-
[56]
Optimiza- tion of quantum computing simulation with gate fusion.Infor- mation Processing Society of Japan, Mar 2021
Horii Hiroshi, Doi Jun, Horii Hiroshi, and Doi Jun. Optimiza- tion of quantum computing simulation with gate fusion.Infor- mation Processing Society of Japan, Mar 2021. ISSN 0167-9260. DOI: https://doi.org/10.1016/j.vlsi.2019.10.004. URLhttps://www. sciencedirect.com/science/a...
2021 doi
-
[57]
SCNet: Supercomputing network.https: //www.scnet.cn/ui/mall/detail/goods?type=software& shopId=1788135137565712385&common1=APP_SOFTWARE&id= 1834155267946004482, 2025
SCNet. SCNet: Supercomputing network.https: //www.scnet.cn/ui/mall/detail/goods?type=software& shopId=1788135137565712385&common1=APP_SOFTWARE&id= 1834155267946004482, 2025. Accessed: 2025-08-26, Software available under Apache License 2.0
2025
-
[58]
G. Zhong. Quantumsimulatoroptimizer, 2025. URLhttps://doi.org/ 10.5281/zenodo.17079328. Published September 8, 2025 | Version v1
2025 doi
-
[61]
Benjamin
Tyson Jones, Anna Brown, Ian Bush, and Simon C. Benjamin. QuEST and high performance simulation of quantum computers.Scientific Re- ports, 9(1):10736, July 2019. ISSN 2045-2322. DOI: 10.1038/s41598- 019-47174-9
2019 doi
-
[2008]
URLhttps://api
DOI: 10.1007/978-3-540-69384-0_75. URLhttps://api. semanticscholar.org/CorpusID:37760018
-
[2014]
URLhttps://doi.org/10.1038/ ncomms5213
DOI: 10.1038/ncomms5213. URLhttps://doi.org/10.1038/ ncomms5213
-
[2016]
URLhttps://doi.org/ 10.1088/1367-2630/18/2/023023
DOI: 10.1088/1367-2630/18/2/023023. URLhttps://doi.org/ 10.1088/1367-2630/18/2/023023
-
[2021]
org/abs/2111.05176, Accessed August 1, 2022
URLhttps://arxiv.org/abs/2111.05176.https://arxiv. org/abs/2111.05176, Accessed August 1, 2022
2022 arXiv
-
[2022]
DOI: 10.1186/s41313-021-00032-6
ISSN 2509-8012. DOI: 10.1186/s41313-021-00032-6. URLhttps: //doi.org/10.1186/s41313-021-00032-6
-
[5820]
URLhttps://doi.org/10
DOI: 10.1038/s42254-021-00348-9. URLhttps://doi.org/10. 1038/s42254-021-00348-9
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.