Pith. sign in

REVIEW 3 major objections 6 minor 61 references

COSMA: Communication-aware Optimization of Fermionic Simulation Kernels for Modular Quantum Architectures

T0 review · 3 major / 6 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read Jointly optimizing fermion-to-qubit mapping, Pauli scheduling, and multi-core allocation cuts inter-core state transfers by up to 2.5× for molecular quantum simulation.

desk verdict Joint F2Q–scheduling–allocation co-design that delivers real communication cuts on molecular Hamiltonians under the usual distance model, with the numbers and the heuristics both tied to that model. read the letter →

arxiv 2607.09381 v1 pith:5CPNYI3I submitted 2026-07-10 quant-ph

classification quant-ph
keywords quantumsimulationmodulararchitecturesfermion-to-qubitmappingPaulischedulingmulti-corequbitallocationinter-corecommunicationTrotterizationchemistry
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Quantum simulation of molecules is a leading target for quantum computers, yet realistic systems need far more qubits than a single chip can hold. Modular machines stitch many smaller cores together, but moving quantum states between cores is slow and noisy and can dominate runtime. COSMA is a compilation stack that attacks that bottleneck by treating three stages as one joint problem: how fermionic modes are encoded onto qubits, the order in which Pauli terms of the Trotter step are executed, and how parity trees and qubit layouts are chosen so that interacting qubits stay co-located. On fourteen molecular benchmarks the resulting circuits require far fewer inter-core transfers than pipelines that fix the mapping first and only then schedule and place. The practical claim is that communication-aware co-design is not an optional polish but a necessary condition for scalable chemistry simulation on multi-core hardware.

What carries the argument

COSMA's co-optimization loop: a genetic search over ternary-tree mappings whose fitness is the inter-core transfer cost of a Gray-inspired Pauli schedule fed into a lookahead misplacement-score heuristic that jointly synthesizes each Pauli gadget's CNOT parity tree and the sequence of core layouts.

What would settle it

Re-run the same molecular suite on a hardware model that charges for congestion or link capacity and check whether the 1.7 imes median communication advantage of COSMA over the Hungarian baselines disappears or reverses.

Watch

Extended reading notes

Core claim

For Trotterized fermionic simulation on modular architectures, the total number of inter-core state transfers can be substantially reduced by simultaneously searching over product-preserving ternary-tree fermion-to-qubit mappings, support-smoothing Pauli orderings, and parity-tree-aware multi-core layouts, rather than optimizing each stage in isolation. On the evaluated molecules the full pipeline yields up to a 2.5 imes reduction and a median 1.7 imes reduction relative to fixed-mapping baselines that use Hungarian allocation.

Load-bearing premise

All reported savings rest on counting shortest-path hops between cores and ignore real link capacity, congestion, and asymmetric latencies.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. COSMA is a compilation framework for Trotterized fermionic simulation on modular multi-core quantum architectures. It jointly optimizes three coupled stages: (i) fermion-to-qubit mapping over product-preserving ternary trees via a genetic algorithm, (ii) Pauli-term scheduling (magnitude, lexicographic, and Gray-inspired orderings), and (iii) a lookahead heuristic that co-designs parity-tree synthesis with multi-core qubit allocation and routing. The objective is the total inter-core transfer cost under a shortest-path distance model (Eqs. 13–14). On 14 STO-3G molecular Hamiltonians (14–90 modes), the full pipeline reports up to ~2.5× and median ~1.7× communication-cost reductions versus fixed mappings (JW/BK/PE/JKMN) plus Hungarian assignment after index-ordered CNOT chains. Ablations attribute gains to mapping, scheduling, and allocation; an open-source GPU-accelerated implementation is provided.

Significance. Inter-core communication is a first-order bottleneck for modular quantum architectures, and fermionic simulation is a primary application driver. Prior work largely optimizes F2Q mappings, Pauli ordering, or multi-core placement in isolation; a joint treatment for chemistry kernels is timely and useful. Strengths include evaluation on real molecular Hamiltonians (not synthetic graphs), controlled ablations that isolate mapping and scheduling, honest reporting that mapping gains shrink under a fixed GA budget as N grows, and a released C++/CUDA framework. Relative comparisons are fair within the same cost model used by the cited multi-core baselines. If the reported gains hold under that model, the paper makes a solid case that cross-layer co-design materially reduces communication for modular chemistry simulation.

major comments (3)
  1. The abstract and conclusion state that COSMA achieves up to 2.5× and a median 1.7× communication reduction and that cross-layer co-design is “essential.” All quantitative claims rest on the pure shortest-path transfer cost of Eqs. (13)–(14). Section VI correctly notes that this model omits link capacity, congestion, and latency asymmetry, and that a hardware-specific model is future work. Because the allocation heuristic (misplacement score Eq. 20, meeting core Eq. 22, merge scores Ψ and G) is driven by the same distance function, absolute factors and possibly rankings vs HQA could change under a different cost model. The relative comparison to HQA is still valid within the shared literature model, but the abstract’s “essential” language and headline factors should be scoped explicitly to the distance-based objective, with a short discussion of how congestion-aware costs might affect the
  2. Section V.A reports a median 59.7% reduction under Gray-inspired scheduling (~2.48×) and 42.8% under magnitude ordering (~1.75×) versus the best fixed-mapping HQA baseline. The abstract’s “up to 2.5× … median improvement of 1.7×” is not tied to a single, clearly defined comparison (which baseline family, which scheduling policy, max vs median over molecules). Please state the exact comparison that produces each number (e.g., median and max of cost_baseline/cost_COSMA over the 14 molecules for a named baseline) so the headline claim is reproducible from the text and figures.
  3. Section II-C and the problem formulation treat reordering of non-commuting Pauli terms as free for the communication objective and defer Trotter-error co-optimization (also flagged in Section VI). That choice is reasonable when communication dominates, but the paper still presents COSMA as a compilation path for quantum simulation kernels. A brief quantitative check—e.g., Trotter error or effective simulation accuracy for one small molecule under Gray vs magnitude ordering at fixed r—would show whether the communication-optimal schedule is accuracy-neutral or requires compensating larger r. Without that, the claim that joint optimization is ready for “efficient and scalable quantum simulation” should be tempered to communication cost under product formulas.
minor comments (6)
  1. Free parameters of the allocator (lookahead window W, decay γ, meeting-core bias λ=1/2) and GA settings (population 50, 25 generations) are stated but not ablated. A short sensitivity note for W/γ on one mid-size molecule would strengthen confidence that gains are not brittle to these choices.
  2. Fig. 8–11 use normalized communication cost without stating the normalization baseline in the captions. Please specify (e.g., relative to JW+MAG+HQA or to the best baseline per molecule) so the plots are self-contained.
  3. In Section IV-C, “F orward merge phase” appears to be a typo for “Forward merge phase.”
  4. Abstract and introduction use “up to 2.5× reduction … median improvement of 1.7×”; body text often reports percentage reductions. Prefer a single convention (multiplicative factor or percent) throughout for consistency.
  5. Related work covers F2Q, Pauli compilation, and modular mapping well; a one-sentence contrast with concurrent multi-core chemistry-specific compilers (if any) or with Treespilation’s cost functions would further clarify novelty of the communication objective.
  6. Eq. (17) and the super-exponential search-space argument motivate heuristics well; a rough wall-clock comparison of one GA generation vs a single HQA baseline run on the largest molecule would help readers judge practicality.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: COSMA reports empirical cost reductions under an explicit, non-fitted inter-core distance objective on independent molecular benchmarks.

full rationale

The paper defines communication cost explicitly as the sum of shortest-path inter-core distances across slices (Eqs. 13–14), then uses a genetic algorithm whose fitness is exactly that measured cost after running the downstream scheduling and allocation heuristics. The reported 1.7–2.5× reductions are obtained by comparing the resulting costs against fixed-mapping + HQA baselines on the same 14 PubChem molecules and the same grid architectures; nothing is fitted to a subset of the data and then re-presented as a prediction, nor is any uniqueness theorem or ansatz imported via self-citation to force the result. Self-citations appear only as related-work baselines or prior multi-core allocation methods, not as load-bearing premises. The idealized nature of the distance model is a modeling limitation (acknowledged in Sec. VI), not a circular derivation. The evaluation chain is therefore self-contained and non-circular.

Assumptions & free parameters 4 free parameters · 4 assumptions · 1 invented entities

The central empirical claim rests on a standard modular cost model, the PPTT family of encodings, and a handful of hand-chosen algorithmic hyper-parameters; no new physical entities are postulated. The free parameters control search effort and heuristic bias rather than being fitted to produce the claimed speed-ups after the fact.

free parameters (4)
  • GA population size / generations = 50 / 25
    Fixed at 50 / 25 for all molecules; the authors note that this fixed budget explores a smaller fraction of the mapping space for larger N, directly affecting the reported mapping-only gains.
  • lookahead window W and decay γ
    Control the misplacement score (Eq. 20); values are not exhaustively justified and influence which qubits are preferred for displacement.
  • meeting-core bias λ = 1/2
    Set to 1/2 in the merge-direction score (Eq. 23); a free algorithmic constant that tilts the heuristic.
  • core capacity κ = 8
    Fixed at 8 qubits per core for all experiments; changes the feasible layout space and therefore absolute transfer counts.
assumptions (4)
  • domain assumption Two-qubit gates may execute only when both logical qubits reside in the same core; inter-core interaction is realized solely by state transfer.
    Stated in Section II-D and used to define every allocation constraint and cost term.
  • domain assumption Total communication cost equals the sum of shortest-path inter-core distances of every logical qubit between consecutive slices (Eq. 13).
    Adopted from prior multi-core literature and used as the sole optimization objective (Eq. 14); congestion and link capacity are ignored.
  • domain assumption Product-preserving ternary-tree (PPTT) mappings are a sufficient and convenient search space for fermion-to-qubit encodings.
    Section II-B and IV-A restrict the genetic search to PPTT; non-PPTT encodings are left to future work.
  • ad hoc to paper Reordering non-commuting Pauli terms does not need to be co-optimized with Trotter error for the communication objective.
    Explicitly chosen in Section II-C and IV-B; the authors acknowledge the accuracy–communication trade-off is future work.
invented entities (1)
  • misplacement score / meeting-core / forward-merge heuristic
    purpose: Drive the co-synthesis of parity CNOT trees and multi-core layouts for each Pauli gadget.
    Algorithmic constructs introduced in Section IV-C; they have no independent physical existence outside the compiler.

how reviews work

0 comments
Cite this review

Pith. "Pith review of COSMA: Communication-aware Optimization of Fermionic Simulation Kernels for Modular Quantum Architectures." pith.science (2026). https://pith.science/paper/5CPNYI3I

@misc{pith2026260709381,
  author       = {Pith},
  title        = {Pith review of: COSMA: Communication-aware Optimization of Fermionic Simulation Kernels for Modular Quantum Architectures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5CPNYI3I}},
  note         = {Machine review of arXiv:2607.09381}
}
abstract

Quantum simulation is a leading application of quantum computing, but scaling to chemically relevant problems requires modular architectures composed of interconnected quantum processing units. In such systems, inter-core quantum communication becomes a major performance bottleneck. In this work, we present COSMA, a communication-aware compilation framework for fermionic simulation kernels targeting modular quantum architectures. Our approach jointly optimizes fermion-to-qubit mapping, Pauli scheduling, and qubit allocation to minimize inter-core state transfers. Evaluated on molecular benchmarks, COSMA achieves up to $2.5\times$ reduction in communication cost compared to state-of-the-art baselines, with a median improvement of $1.7\times$. These results demonstrate that cross-layer co-design is essential for efficient and scalable quantum simulation on multi-core quantum hardware.

Figures

Figures reproduced from arXiv: 2607.09381 by the authors.

Figure 1
Figure 1. Overview of the compilation flow for quantum simulation kernels on modular architectures. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. JKMN-like [17] ternary-tree mapping example for 10 modes with identity mode-order bijection f(u) = u. More generally, a Pauli-string F2Q mapping assigns a Pauli string Pk to each Majorana operator mk such that the anticommutation algebra is preserved: {mi , mj} = 2δij 1 7→ {Pi , Pj} = 2δij 1. (11) Jordan–Wigner belongs to the broader class of product￾preserving ternary-tree (PPTT) F2Q mappings [16]. Mappings in this… view at source ↗
Figure 5
Figure 5. Synthesized circuit for one Trotter step after decom [PITH_FULL_IMAGE:figures/full_fig_p004_5.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Gate-level implementation of the Pauli gadget [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 6
Figure 6. Figure 6: Modular quantum-architecture model and allocation [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: Overview of the COSMA framework. parity trees and the qubit-interaction structure seen by the allocator. The ordering π controls the smoothness of the sup￾port sequence across consecutive gadgets, potentially reducing qubit displacement between slices. Each tree Tt fix…
Figure 8
Figure 8. Figure 8: Comparison against baselines considering magnitude [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Comparison against baselines considering Gray-like [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 12
Figure 12. Figure 12: Single-thread (ST), multi-thread (MT) and CUDA [PITH_FULL_IMAGE:figures/full_fig_p009_12.png]
Figure 11
Figure 11. Figure 11: Impact of scheduling on communication cost. [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 9 linked inside Pith

  1. [1]

    Simulating fermions with a digital quantum computer,

    R. W. Chien, M. Chiew, B. Harrison, J. Necaise, W. Wang, M. Mudassar, C. McLauchlan, T. M. Henderson, G. E. Scuseria, S. Strelchuket al., “Simulating fermions with a digital quantum computer,”Nature Reviews Physics, pp. 1–15, 2026

  2. [2]

    I. I. Manin,Mathematics as metaphor: Selected essays of Yuri I. Manin. American Mathematical Soc., 2007, vol. 20

  3. [3]

    Simulating physics with computers,

    R. P. Feynman, “Simulating physics with computers,” inFeynman and computation. cRc Press, 2018, pp. 133–153

  4. [4]

    Quantum chemistry in the age of quantum computing,

    Y . Cao, J. Romero, J. P. Olson, M. Degroote, P. D. Johnson, M. Kieferov ´a, I. D. Kivlichan, T. Menke, B. Peropadre, N. P. Sawaya et al., “Quantum chemistry in the age of quantum computing,”Chemical reviews, vol. 119, no. 19, pp. 10 856–10 915, 2019

  5. [5]

    Computational approaches stream- lining drug discovery,

    A. V . Sadybekov and V . Katritch, “Computational approaches stream- lining drug discovery,”Nature, vol. 616, no. 7958, pp. 673–685, 2023

  6. [6]

    Eluci- dating reaction mechanisms on quantum computers,

    M. Reiher, N. Wiebe, K. M. Svore, D. Wecker, and M. Troyer, “Eluci- dating reaction mechanisms on quantum computers,”Proceedings of the national academy of sciences, vol. 114, no. 29, pp. 7555–7560, 2017

  7. [7]

    Surface codes: Towards practical large-scale quantum computation,

    A. G. Fowler, M. Mariantoni, J. M. Martinis, and A. N. Cleland, “Surface codes: Towards practical large-scale quantum computation,”Physical Review A—Atomic, Molecular , and Optical Physics, vol. 86, no. 3, p. 032324, 2012

  8. [8]

    Large-scale modular quantum-computer architecture with atomic memory and photonic interconnects,

    C. Monroe, R. Raussendorf, A. Ruthven, K. R. Brown, P. Maunz, L.-M. Duan, and J. Kim, “Large-scale modular quantum-computer architecture with atomic memory and photonic interconnects,”Physical Review A, vol. 89, no. 2, p. 022317, 2014

Show all 61 references
  1. [9]

    On double full-stack communication-enabled architectures for multicore quantum computers,

    S. Rodrigo, S. Abadal, E. Alarcon, M. Bandic, H. Van Someren, and C. G. Almud ´ever, “On double full-stack communication-enabled architectures for multicore quantum computers,”IEEE micro, vol. 41, no. 5, pp. 48–56, 2021

  2. [10]

    Telesabre: Heuristic layout synthesis in multi-core quantum systems with teleport interconnect,

    E. Russo, E. Vinciguerra, M. Palesi, D. Patti, G. Ascia, and V . Catania, “Telesabre: Heuristic layout synthesis in multi-core quantum systems with teleport interconnect,” in2025 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1. IEEE, 2025, pp. 749–758

  3. [11]

    Distributed quantum computing: a survey,

    M. Caleffi, M. Amoretti, D. Ferrari, J. Illiano, A. Manzalini, and A. S. Cacciapuoti, “Distributed quantum computing: a survey,”Computer Networks, vol. 254, p. 110672, 2024

  4. [12]

    Term grouping and travelling salesperson for digital quantum simulation,

    K. Gui, T. Tomesh, P. Gokhale, Y . Shi, F. T. Chong, M. Martonosi, and M. Suchara, “Term grouping and travelling salesperson for digital quantum simulation,”arXiv preprint arXiv:2001.05983, 2020

  5. [13]

    Szabo and N

    A. Szabo and N. S. Ostlund,Modern quantum chemistry: introduction to advanced electronic structure theory. Courier Corporation, 2012

  6. [14]

    Simulation of electronic structure hamiltonians using quantum computers,

    J. D. Whitfield, J. Biamonte, and A. Aspuru-Guzik, “Simulation of electronic structure hamiltonians using quantum computers,”Molecular Physics, vol. 109, no. 5, pp. 735–750, 2011

  7. [15]

    Hamiltonian simulation by qubitization,

    G. H. Low and I. L. Chuang, “Hamiltonian simulation by qubitization,” Quantum, vol. 3, p. 163, 2019

  8. [16]

    Bonsai algorithm: Grow your own fermion-to-qubit mappings,

    A. Miller, Z. Zimbor ´as, S. Knecht, S. Maniscalco, and G. Garc ´ıa-P´erez, “Bonsai algorithm: Grow your own fermion-to-qubit mappings,”PRX quantum, vol. 4, no. 3, p. 030314, 2023

  9. [17]

    Optimal fermion- to-qubit mapping via ternary trees with applications to reduced quantum states learning,

    Z. Jiang, A. Kalev, W. Mruczkiewicz, and H. Neven, “Optimal fermion- to-qubit mapping via ternary trees with applications to reduced quantum states learning,”Quantum, vol. 4, p. 276, 2020

  10. [18]

    ¨Uber das paulische ¨aquivalenzverbot,

    P. Jordan and E. Wigner, “ ¨Uber das paulische ¨aquivalenzverbot,” Zeitschrift f ¨ur Physik, vol. 47, no. 9, pp. 631–651, 1928

  11. [19]

    Taper- ing off qubits to simulate fermionic hamiltonians,

    S. Bravyi, J. M. Gambetta, A. Mezzacapo, and K. Temme, “Taper- ing off qubits to simulate fermionic hamiltonians,”arXiv preprint arXiv:1701.08213, 2017

  12. [20]

    Fermionic quantum computation,

    S. B. Bravyi and A. Y . Kitaev, “Fermionic quantum computation,” Annals of Physics, vol. 298, no. 1, pp. 210–226, 2002

  13. [21]

    Treespilation: architecture-and state-optimised fermion-to-qubit mappings,

    A. Miller, A. Glos, and Z. Zimbor ´as, “Treespilation: architecture-and state-optimised fermion-to-qubit mappings,”npj Quantum Information, 2026

  14. [22]

    An adaptive variational algorithm for exact molecular simulations on a quantum computer,

    H. R. Grimsley, S. E. Economou, E. Barnes, and N. J. Mayhall, “An adaptive variational algorithm for exact molecular simulations on a quantum computer,”Nature communications, vol. 10, no. 1, p. 3007, 2019

  15. [23]

    Redefining lexicographical ordering: Optimizing pauli string decompo- sitions for quantum compiling,

    Q. Huang, D. Winderl, A. Meijer-Van De Griend, and R. Yeung, “Redefining lexicographical ordering: Optimizing pauli string decompo- sitions for quantum compiling,” in2024 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1. IEEE, 2024, pp. 885–896

  16. [24]

    A comparison of the bravyi–kitaev and jordan–wigner transformations for the quantum simulation of quantum chemistry,

    A. Tranter, P. J. Love, F. Mintert, and P. V . Coveney, “A comparison of the bravyi–kitaev and jordan–wigner transformations for the quantum simulation of quantum chemistry,”Journal of chemical theory and computation, vol. 14, no. 11, pp. 5617–5630, 2018

  17. [25]

    Improving quan- tum algorithms for quantum chemistry,

    M. B. Hastings, D. Wecker, B. Bauer, and M. Troyer, “Improving quan- tum algorithms for quantum chemistry,”arXiv preprint arXiv:1403.1539, 2014

  18. [26]

    A generic compila- tion strategy for the unitary coupled cluster ansatz,

    A. Cowtan, W. Simmons, and R. Duncan, “A generic compila- tion strategy for the unitary coupled cluster ansatz,”arXiv preprint arXiv:2007.10515, 2020

  19. [27]

    Pauli- hedral: a generalized block-wise compiler optimization framework for quantum simulation kernels,

    G. Li, A. Wu, Y . Shi, A. Javadi-Abhari, Y . Ding, and Y . Xie, “Pauli- hedral: a generalized block-wise compiler optimization framework for quantum simulation kernels,” inProceedings of the 27th ACM Interna- tional Conference on Architectural Support for Programming Languages...

  20. [28]

    Well-conditioned multiprod- uct hamiltonian simulation,

    G. H. Low, V . Kliuchnikov, and N. Wiebe, “Well-conditioned multiprod- uct hamiltonian simulation,”arXiv preprint arXiv:1907.11679, 2019

  21. [29]

    Modelling short-range quantum teleportation for scalable multi-core quantum com- puting architectures,

    S. Rodrigo, S. Abadal, C. G. Almud ´ever, and E. Alarc ´on, “Modelling short-range quantum teleportation for scalable multi-core quantum com- puting architectures,” inACM International Conference on Nanoscale Computing and Communication, 2021, pp. 1–7

  22. [30]

    Optimized compiler for distributed quantum computing,

    D. Cuomo, M. Caleffi, K. Krsulich, F. Tramonto, G. Agliardi, E. Prati, and A. S. Cacciapuoti, “Optimized compiler for distributed quantum computing,”ACM Transactions on Quantum Computing, vol. 4, no. 2, pp. 1–29, 2023

  23. [31]

    Mapping quantum circuits to modular architectures with qubo,

    M. Bandic, L. Prielinger, J. N ¨ußlein, A. Ovide, S. Rodrigo, S. Abadal, H. Van Someren, G. Vardoyan, E. Alarcon, C. G. Almudeveret al., “Mapping quantum circuits to modular architectures with qubo,” in2023 IEEE International Conference on Quantum Computing and Engineer- ing (...

  24. [32]

    Optimal qubit assignment and routing via integer programming,

    G. Nannicini, L. S. Bishop, O. G ¨unl¨uk, and P. Jurcevic, “Optimal qubit assignment and routing via integer programming,”ACM Transactions on Quantum Computing, vol. 4, no. 1, pp. 1–31, 2022

  25. [33]

    Time-sliced quantum circuit partitioning for modular architectures,

    J. M. Baker, C. Duckering, A. Hoover, and F. T. Chong, “Time-sliced quantum circuit partitioning for modular architectures,” inProceedings of the 17th ACM International Conference on Computing Frontiers, 2020, pp. 98–107

  26. [34]

    Revisiting the mapping of quantum circuits: Entering the multi-core era,

    P. Escofet, A. Ovide, M. Bandic, L. Prielinger, H. van Someren, S. Feld, E. Alarc ´on, S. Abadal, and C. G. Almud ´ever, “Revisiting the mapping of quantum circuits: Entering the multi-core era,”ACM Transactions on Quantum Computing, 2024

  27. [35]

    Hungarian qubit assignment for optimized mapping of quantum cir- cuits on multi-core architectures,

    P. Escofet, A. Ovide, C. G. Almudever, E. Alarc ´on, and S. Abadal, “Hungarian qubit assignment for optimized mapping of quantum cir- cuits on multi-core architectures,”IEEE Computer Architecture Letters, vol. 22, no. 2, pp. 161–164, 2023

  28. [36]

    Mitchell,An introduction to genetic algorithms

    M. Mitchell,An introduction to genetic algorithms. MIT press, 1998

  29. [37]

    Applying adaptive algorithms to epistatic domains

    L. Daviset al., “Applying adaptive algorithms to epistatic domains.” in IJCAI, vol. 85, 1985, pp. 162–164

  30. [38]

    Ordering of trotterization: Impact on errors in quantum simulation of electronic structure,

    A. Tranter, P. J. Love, F. Mintert, N. Wiebe, and P. V . Coveney, “Ordering of trotterization: Impact on errors in quantum simulation of electronic structure,”Entropy, vol. 21, no. 12, p. 1218, 2019

  31. [39]

    Pubchem 2025 update,

    S. Kim, J. Chen, T. Cheng, A. Gindulyte, J. He, S. He, Q. Li, B. A. Shoemaker, P. A. Thiessen, B. Yuet al., “Pubchem 2025 update,” Nucleic acids research, vol. 53, no. D1, pp. D1516–D1525, 2025

  32. [40]

    Recent developments in the pyscf program package,

    Q. Sun, X. Zhang, S. Banerjee, P. Bao, M. Barbry, N. S. Blunt, N. A. Bogdanov, G. H. Booth, J. Chen, Z.-H. Cuiet al., “Recent developments in the pyscf program package,”The Journal of chemical physics, vol. 153, no. 2, 2020

  33. [41]

    Openfermion: the electronic structure package for quantum computers,

    J. R. McClean, N. C. Rubin, K. J. Sung, I. D. Kivlichan, X. Bonet- Monroig, Y . Cao, C. Dai, E. S. Fried, C. Gidney, B. Gimbyet al., “Openfermion: the electronic structure package for quantum computers,” Quantum Science & Technology, vol. 5, no. 3, p. 034014, 2020

  34. [42]

    Optimizing qubit assignment in modular quantum systems via attention-based deep reinforcement learning,

    E. Russo, M. Palesi, D. Patti, G. Ascia, and V . Catania, “Optimizing qubit assignment in modular quantum systems via attention-based deep reinforcement learning,” in2025 Design, Automation & Test in Europe Conference (DATE). IEEE, 2025, pp. 1–7

  35. [43]

    [Online]

    CCCL Development Team,CCCL: CUDA C++ Core Libraries, 2023. [Online]. Available: https://github.com/NVIDIA/cccl

  36. [44]

    Thrust: A productivity-oriented library for cuda,

    N. Bell and J. Hoberock, “Thrust: A productivity-oriented library for cuda,” inGPU computing gems Jade edition. Elsevier, 2012, pp. 359– 371

  37. [45]

    Quantum data centres: why entanglement changes everything,

    A. S. Cacciapuoti, C. Pellitteri, J. Illiano, L. d’Avossa, F. Mazza, S. Chen, and M. Caleffi, “Quantum data centres: why entanglement changes everything,”Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, vol. 384, no. 2315, 2026

  38. [46]

    Fast simulation of fermions with reconfigurable qubits,

    N. Maskara, M. Kalinowski, D. Gonzalez-Cuadra, and M. D. Lukin, “Fast simulation of fermions with reconfigurable qubits,”arXiv preprint arXiv:2509.08898, 2025

  39. [47]

    Stabilizer-based quantum simula- tion of fermion dynamics with local qubit encodings,

    A. Gandon, S. Piccinelli, M. Rossmannek, F. Tacchino, A. Ba- iardi, J. Nys, and I. Tavernelli, “Stabilizer-based quantum simula- tion of fermion dynamics with local qubit encodings,”arXiv preprint arXiv:2512.11418, 2025

  40. [48]

    Optimal fermion- qubit mappings via quadratic assignment,

    M. Chiew, C. Ibrahim, I. Safro, and S. Strelchuk, “Optimal fermion- qubit mappings via quadratic assignment,” in2025 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1. IEEE, 2025, pp. 560–571

  41. [49]

    Hatt: Hamiltonian adaptive ternary tree for optimizing fermion- to-qubit mapping,

    Y . Liu, K. Yao, J. Hong, J. Froustey, E. Rrapaj, C. Iancull, G. Li, and Y . Shi, “Hatt: Hamiltonian adaptive ternary tree for optimizing fermion- to-qubit mapping,” in2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 2025, pp. 143– 157

  42. [50]

    Optimised fermion-qubit encodings for quantum simulation with reduced transpiled circuit depth,

    M. W. de la Bastida, T. M. Bickley, and P. V . Coveney, “Optimised fermion-qubit encodings for quantum simulation with reduced transpiled circuit depth,”arXiv preprint arXiv:2512.13580, 2025

  43. [51]

    Fermihedral: On the optimal compilation for fermion-to-qubit encoding,

    Y . Liu, S. Che, J. Zhou, Y . Shi, and G. Li, “Fermihedral: On the optimal compilation for fermion-to-qubit encoding,” inProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, V olume 3, 2024, pp. 382–397

  44. [52]

    A theory of trotter error,

    A. M. Childs, Y . Su, M. C. Tran, N. Wiebe, and S. Zhu, “A theory of trotter error,”arXiv preprint arXiv:1912.08854, 2019

  45. [53]

    Optimized quantum program execution ordering to miti- gate errors in simulations of quantum systems,

    T. Tomesh, K. Gui, P. Gokhale, Y . Shi, F. T. Chong, M. Martonosi, and M. Suchara, “Optimized quantum program execution ordering to miti- gate errors in simulations of quantum systems,” in2021 International Conference on Rebooting Computing (ICRC). IEEE, 2021, pp. 1–13

  46. [54]

    Pauliforest: Connectivity-aware synthesis and pauli-oriented qubit mapping for near- term quantum simulation,

    Y . Li, Y . Zhang, H. Deng, M. Chen, and Z. Li, “Pauliforest: Connectivity-aware synthesis and pauli-oriented qubit mapping for near- term quantum simulation,”IEEE Transactions on Computer-Aided De- sign of Integrated Circuits and Systems, vol. 44, no. 6, pp. 2119–2129, 2024

  47. [55]

    Tetris: A compilation framework for vqa applications in quantum computing,

    Y . Jin, Z. Li, F. Hua, T. Hao, H. Zhou, Y . Huang, and E. Z. Zhang, “Tetris: A compilation framework for vqa applications in quantum computing,” in2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 2024, pp. 277–292

  48. [56]

    Kernpiler: Compiler opti- mization for quantum hamiltonian simulation with partial trotterization,

    E. Decker, L. Goetz, E. McKinney, E. Gustafson, J. Zhou, Y . Liu, A. K. Jones, A. Li, A. Schuckert, S. Steinet al., “Kernpiler: Compiler opti- mization for quantum hamiltonian simulation with partial trotterization,” arXiv preprint arXiv:2504.07214, 2025

  49. [57]

    Multicore quantum computing,

    H. Jnane, B. Undseth, Z. Cai, S. C. Benjamin, and B. Koczor, “Multicore quantum computing,”Physical Review Applied, vol. 18, no. 4, 2022

  50. [58]

    Route-forcing: Scalable quantum circuit mapping for scalable quantum computing architectures,

    P. Escofet, A. Gonzalvo, E. Alarc ´on, C. G. Almud ´ever, and S. Abadal, “Route-forcing: Scalable quantum circuit mapping for scalable quantum computing architectures,” in2024 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1. IEEE, 2024, pp. 909–920

  51. [59]

    Optimized quantum circuit partitioning across multiple quantum processors,

    E. Kaur, S. Pouryousef, H. Shapourian, J. Zhao, M. Kilzer, R. Kompella, and R. Nejabati, “Optimized quantum circuit partitioning across multiple quantum processors,”IEEE Transactions on Quantum Engineering, 2025

  52. [60]

    Distributed quantum simulation,

    T. Feng, J. Xu, W. Yu, Z. Ye, P. Yao, and Q. Zhao, “Distributed quantum simulation,”arXiv preprint arXiv:2411.02881, 2024

  53. [61]

    Simulating time evolution on distributed quantum computers,

    F. L. Buessen, D. Segal, and I. Khait, “Simulating time evolution on distributed quantum computers,”Physical Review Research, vol. 5, no. 2, p. L022003, 2023

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.