Pith. sign in

REVIEW 3 cited by

A Reinforcement Learning Environment for Directed Quantum Circuit Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07054 v1 pith:KPFHNJH7 submitted 2024-01-13 quant-ph cs.AI

classification quant-phcs.AI
keywords quantumcircuitsenvironmentcircuitlearningstatesqubitreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With recent advancements in quantum computing technology, optimizing quantum circuits and ensuring reliable quantum state preparation have become increasingly vital. Traditional methods often demand extensive expertise and manual calculations, posing challenges as quantum circuits grow in qubit- and gate-count. Therefore, harnessing machine learning techniques to handle the growing variety of gate-to-qubit combinations is a promising approach. In this work, we introduce a comprehensive reinforcement learning environment for quantum circuit synthesis, where circuits are constructed utilizing gates from the the Clifford+T gate set to prepare specific target states. Our experiments focus on exploring the relationship between the depth of synthesized quantum circuits and the circuit depths used for target initialization, as well as qubit count. We organize the environment configurations into multiple evaluation levels and include a range of well-known quantum states for benchmarking purposes. We also lay baselines for evaluating the environment using Proximal Policy Optimization. By applying the trained agents to benchmark tests, we demonstrated their ability to reliably design minimal quantum circuits for a selection of 2-qubit Bell states.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RubriQ: Rubric-Guided Group Relative Policy Optimization for Constraint-Aware Quantum Circuit Synthesis

    quant-ph 2026-07 conditional novelty 6.0 of 10

    A rubric-guided GRPO pipeline fine-tunes a 7B LLM to synthesize quantum circuits achieving 3.31x T-gate compression with <1% hardware-constraint violations, validated on IBM and IonQ processors.

  2. Quantum Architecture Search for Solving Quantum Machine Learning Tasks

    quant-ph 2025-09 conditional novelty 5.0 of 10

    A reinforcement learning framework (RL-QAS) discovers compact variational quantum circuit architectures for Iris and binary MNIST classification, outperforming a simple strongly-entangling-layer baseline.

  3. Unitary Synthesis with AlphaZero via Dynamic Circuits

    quant-ph 2025-08 conditional novelty 5.0 of 10

    An AlphaZero-like RL agent can synthesize exact Clifford+T circuits for up to three qubits with ancilla, recovering known optimal decompositions and a 4-T Toffoli implementation.

Pith tools