REVIEW 3 major objections 6 minor 59 references
Reinforcement Learning Enhancing Entanglement for Two-Photon-Driven Rabi Model
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A reinforcement-learning agent learns temporal sequences of two-photon drive pulses that raise the qubit–cavity entanglement witness above the uncontrolled baseline in a dissipative Rabi model.
desk verdict Useful numerical proof-of-principle for RL-based entanglement enhancement, but the missing trivial baselines undercut the central claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the combination of a phase diagram that identifies the drive regime worth controlling and a deep Q-network that searches that regime. The phase diagram is built from the energy gap $E_1 - E_0$, the entanglement witness $E_T = \sum_{\lambda_i<0}|\lambda_i|$ of the partially transposed ground-state density matrix, the equal-time second-order correlation $g^{(2)}$, the Wigner-function negativity $S_W$, and a Kullback–Leibler truncation metric that pinpoints the critical point $\Omega_c = \delta_c/2$. The control agent observes six expectation values of the driven system, chooses one of three pulse levels each time step, and receives a reward $R_t = 10\langle \Delta E_T > 0\rangle - \langle \Delta E_T \le 0\rangle$ that heavily rewards increases of the entanglement witness. The learned pulse sequence is applied in an open-loop manner, and the open-system dynamics are simulated with a Lindblad master equation whose relaxation rates, built from dressed-state couplings $C^\chi_{jk}$, assume a flat bath spectral density and $\alpha^2_\chi(\Delta)\propto \Delta$, as realizable in circuit QED.
What would settle it
Retrain the agent using the identical reward and hyperparameters but with an Ohmic spectral density $d(\Delta)\propto \Delta^s$ for $s\neq0$ (holding the total damping rate fixed) and compare the final-time entanglement witness $E_T$ under the learned pulses; a clear drop or disappearance of the enhancement for $s$ slightly above or below zero would show the flat-spectrum assumption is load-bearing.
Extended reading notes
Core claim
The central discovery is that reinforcement learning can reliably find open-loop control fields that enhance entanglement in the two-photon-driven Rabi model under dissipation. Concretely, the agent outputs a sequence of 30 square pulses with amplitudes in $\{0, 0.5\,\Omega_{\max}/\delta_c, \Omega_{\max}/\delta_c\}$ that steer the system so that the partial-transpose entanglement witness $E_T$ at the final time exceeds the uncontrolled value, for initial ground and coherent-product states, and for maximum drive amplitudes $\Omega_{\max}/\delta_c = 0.3$ and $0.5$. The learned policy is then executed open-loop, demonstrating that the enhancement comes from the pulse design itself rather than from measurement feedback. The control effect degrades monotonically with increasing dissipation rate and thermal population, but remains positive throughout the tested range.
Load-bearing premise
The simulations assume a flat bath spectral density and a linear-in-frequency system-bath coupling in the relaxation rates of Eq. (9); if a real circuit-QED bath deviates from this, the learned pulses may not enhance entanglement as strongly.
Editorial extensions
If this is right
- The same RL protocol can be lifted to any amplitude-tunable two-photon drive in cavity or circuit QED, without requiring real-time feedback during the pulse.
- Setting the drive amplitude around the phase boundary $\Omega_c = \delta_c/2$ yields the most pronounced entanglement gain, so the phase diagram can be used to preselect operating points.
- The learned enhancement persists under moderate dissipation and thermal noise, suggesting the pulses could work in realistic devices where $\gamma$ and $\bar{n}$ are not perfectly zero.
- The approach is modular: the DQN agent can be replaced by other RL or optimization modules, and the controlled system can be replaced by other tunable quantum systems, broadening the applicability.
Reading between the lines
- Because the reward only tracks changes in the entanglement witness, the same scheme should work with other entanglement measures such as concurrence; a cheap test is to retrain with concurrence as the reward and compare the resulting final-time entanglement.
- The learned pulse statistics show the agent favoring nonzero drive values, which suggests the optimal strategy is close to bang-bang control; if true, even simpler square-pulse sequences might achieve most of the gain, and the RL step could be bypassed by a parameter search.
- The robustness claim is tied to the flat-spectral-density assumption. A direct numerical check with an Ohmic bath ($d(\Delta)\propto \Delta^s$, $s\neq0$) would reveal whether the enhancement survives a frequency-dependent environment.
- The same training pipeline could be adapted to enhance other resource quantifiers, such as squeezing or steering, since the observation vector and reward function are the only problem-specific parts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the two-photon-driven quantum Rabi model and proposes a deep Q-network (DQN) control scheme that modulates the time-dependent two-photon drive amplitude as a sequence of three-level square pulses to enhance the qubit-cavity entanglement witness E_T. The authors first characterize the phase diagram of the model via the energy gap, the entanglement witness, the second-order correlation, and the Wigner-function negativity, identifying a critical point at Ω_c = δ_c/2. They then train a DQN agent with a reward function that rewards increases of E_T, and apply the learned open-loop pulse sequence under a Lindblad master equation with dressed-state relaxation rates. They report that the controlled dynamics enhance E_T for two initial states, at several dissipation strengths and thermal photon numbers, and they claim the scheme is generalizable. No comparison with trivial control baselines is provided, and the phase-transition evidence is based on a possibly ill-defined divergence that may be a truncation artifact.
Significance. If the central claim is validated, the paper offers a concrete demonstration that RL-designed open-loop pulse sequences can preserve or enhance entanglement in a strongly coupled open quantum system, which is a valuable step for quantum control. The paper is clearly written and the numerical workflow is presented in detail. However, the claim relies on the absence of a baseline comparison; because the reward function directly optimizes E_T, the result that the learned policy increases E_T is not surprising. The paper's value would be significantly strengthened by showing that the learned policy outperforms constant and random controls at matched amplitude and dissipation. The multi-indicator phase diagram (energy gap, E_T, g^(2), Wigner negativity) is a useful consistency check, but the truncation-related divergence warrants care. No code or data are provided, limiting reproducibility.
major comments (3)
- [§IV.D, Figs. 5–7] The central claim that the learned pulse sequence 'enhances the entanglement' is not substantiated because no comparison with trivial control baselines is provided. In Sec. IV.D, the text states that 'the application of the control field designed by the RL agent has a positive effect on enhancing the entanglement' (paragraph after Fig. 7), but all dynamics in Figs. 5–7 are either controlled dynamics or uncontrolled dynamics at different parameters; there is no same-panel curve for a constant drive at the same Ω_max, a zero-amplitude drive, or a random pulse sequence drawn from the same action set {0, 0.5Ω_max/δ_c, Ω_max/δ_c}. Since the reward function (8) explicitly gives +10 for any increase of E_T and -1 for any decrease, an agent is expected to discover some sequence that raises E_T almost by construction. To establish the RL-specific advantage, please add baselines: (i) constant Ω(t) = Ω_max, (ii) Ω(t) = 0, (iii) a random sequence of the same three action values, and (iv) the best-of-N random sequences with the same number of episodes, all at the same total pulse energy (or same average amplitude) and same dissipation rates. Report the time-averaged E_T and its standard error over at least 20 independent training runs for the RL agent and over the same number of random realizations for the baselines.
- [§III, Fig. 1, Eq. (2)] The phase-transition claim is insufficiently supported because the KL-divergence quantity defined in Eq. (2) is not well defined as written, and its divergence at Ω/δ_c > 0.5 may be a truncation artifact. The expression 'KL^q_M = ρ_{M+q}| log ρ_{M+q} − log ρ_M |' lacks a trace (or other operation) and the replacement of zero matrix elements by 1 before taking logarithms is an uncontrolled approximation. More importantly, the ground state of the two-photon-driven Rabi model is expected to have a large photon-number support near the spectral collapse point, so for fixed M=60 the difference between ρ_M and ρ_{M+1} can become large simply because the state is not converged. The paper itself marks Ω/δ_c ≥ 0.5 as an area where 'numerical simulations might fail' (white shadow in Fig. 1), yet it later uses Ω_max/δ_c = 0.5 as a control point in Figs. 5 and 7. Please either (i) provide a convergence study showing that the divergence point is stable as M increases (e.g., plot KL^q_M vs Ω for M = 20, 40, 60, 80, 100) and define KL properly, or (ii) weaken the phase-transition claims and rely on the energy-gap and other indicators that are less sensitive to truncation.
- [§IV.C, Eq. (9)] The robustness-to-dissipation result relies on a specific microscopic model for the bath. Eq. (9) assumes a flat spectral density and α^2_χ(Δ) ∝ Δ, giving Γ^χ_jk = γ_χ (Δ_{kj}/ω_0)|C^χ_jk|^2. The paper justifies this by circuit-QED realizability, but no experimental parameters are given, and no sensitivity analysis is performed. Because the main positive message of the paper is that the RL control 'exhibits robustness against dissipation' (abstract and Sec. VI), this assumption is load-bearing. Please add a brief sensitivity check, e.g., repeating the control optimization with a different spectral density (such as an ohmic or Lorentzian form) or with bare (non-dressed) collapse operators, to verify that the enhancement persists. If such a check cannot be done within the manuscript's scope, please explicitly state this limitation in the conclusions.
minor comments (6)
- [Fig. 7 caption] In the caption, '(a2) Ω_max = 0.3/δ_c' should read 'Ω_max = 0.3δ_c' to match the other panels and the text.
- [§IV.B] The sentence 'output neurons provide the probability of choosing which action' is inaccurate: DQN outputs Q-values, and action selection is typically ε-greedy with respect to those Q-values. Please rephrase.
- [§IV.D] The statement that 'the probability of Ω_max/δ_c = 0 decreases in the distribution of the control pulses' is not quantified. Please provide a histogram or a plot of the action frequencies versus training epoch to support this claim.
- [§V] The generalization claims in Sec. V are not demonstrated; the text only lists alternative RL algorithms without showing results. Please either add supporting simulations or clearly label this section as an outlook.
- [Figures 1–6] Most figures lack error bars or an indication of the number of independent runs. Only Fig. 7 shows error zones. Please add error bars or at least a sentence stating the number of seeds used for all RL results.
- [Global] The manuscript contains numerous typographical errors (e.g., 'Two-Ph oton' in the title, 'con trol' in the abstract) and missing axis labels or color scales in some figures (e.g., Fig. 1(a)). A careful proofreading pass is needed.
Circularity Check
No significant circularity: the RL reward encodes the stated objective, and the claimed enhancement is a verified optimization result rather than an input.
full rationale
The paper's central claim is that DQN-designed square pulses increase the entanglement witness E_T under dissipative dynamics. Although the reward in Eq. (8) is constructed from increments of E_T, this is the standard role of a reward function in reinforcement learning: it incentivizes the desired behavior but does not by itself guarantee that the agent will find pulses that increase E_T in the Lindblad dynamics of Eq. (9). The reported enhancement is therefore a numerical optimization outcome, not a quantity equal to its input by construction. The phase diagram is cross-checked by four independent indicators (energy gap, E_T, g^(2), Wigner negativity), and the critical point Omega_c = delta_c/2 is identified by a convergence diagnostic and then corroborated rather than assumed. Self-citations [29,44] are background references for the RL control methodology and are not load-bearing for the novel result. No fitted parameter is renamed as a prediction, and no uniqueness theorem is imported from the authors' prior work. The absence of a random- or constant-pulse baseline is a legitimate experimental-design criticism, but it is not circularity: it affects the strength of the claim that RL specifically outperforms trivial controls, not whether the reported E_T values are derived from the inputs by definition. Overall, the derivation chain is self-contained with respect to circularity.
Assumptions & free parameters
free parameters (4)
- Truncation dimension M =
60 (and larger for convergence checks)
- Maximum control amplitude Omega_max =
0.3 delta_c (also 0.5 delta_c)
- Reward function coefficients =
10 (positive reward) and -1 (negative reward)
- DQN hyperparameters =
batch size 128, discount 0.99, epsilon 0.9 to 0.05, learning rate 1e-4, etc.
assumptions (4)
- domain assumption The Hamiltonian (1) with the two-photon drive and optional Kerr term correctly describes the two-photon-driven Rabi system.
- domain assumption The Markovian Lindblad master equation (9) with the dressed-state jump operators and relaxation rates Gamma_jk^chi = gamma_chi (Delta_kj/omega_0) |C^chi_jk|^2 is a valid description of the open-system dynamics.
- domain assumption The numerical truncation of the Hilbert space at dimension M yields converged ground states and dynamics for the parameters studied.
- ad hoc to paper The RL reward based on the entanglement witness E_T is a suitable proxy for the control objective.
Cite this review
Pith. "Pith review of Reinforcement Learning Enhancing Entanglement for Two-Photon-Driven Rabi Model." pith.science (2026). https://pith.science/paper/22L5SPQ6
@misc{pith2026241115841,
author = {Pith},
title = {Pith review of: Reinforcement Learning Enhancing Entanglement for Two-Photon-Driven Rabi Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/22L5SPQ6}},
note = {Machine review of arXiv:2411.15841}
}
read the original abstract
A control scheme is proposed that leverages reinforcement learning to enhance entanglement by modulating the two-photon-driven amplitude in a Rabi model. The quantum phase diagram versus the amplitude of the two-photon process and the coupling between the cavity field and the atom in the Rabi model, is indicated by the energy spectrum of the hybrid system, the witness of entanglement, second order correlation, and negativity of Wigner function. From a dynamical perspective, the behavior of entanglement can reflect the phase transition and the reinforcement learning agent is employed to produce temporal sequences of control pulses to enhance the entanglement in the presence of dissipation. The entanglement can be enhanced in different parameter regimes and the control scheme exhibits robustness against dissipation. The replaceability of the controlled system and the reinforcement learning module demonstrates the generalization of this scheme. This research paves the way of positively enhancing quantum resources in non-equilibrium systems.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
3 and Ω /δ c = 0. 1, 0. 5, 0. 503 with the corresponding Wigner functions, respectively. The partial transpose criterion involves analyzing the eigenvalues of the partially transposed density matrix of the composite system, and it offers a sufficient (but not necessary) criterion for witnessing entanglement in bi- partite quantum systems [40, 41]. The appear...
-
[2]
I. I. Rabi, Space quantization in a gyrating magnetic field, Phys. Rev. 51, 652 (1937)
1937
-
[3]
H. Walther, B. T. H. Varcoe, B. G. Englert, and T. Becker, Cavity quantum electrodynamics, Rep. Prog. Phys. 69, 1325 (2006)
work page 2006
-
[4]
S. Haroche, and J. -M. Raimond, Exploring the Quan- tum: Atoms, Cavities, and Photons, Oxford University Press (2006)
work page 2006
-
[5]
P. Forn-D ´ ıaz, L. Lamata, E. Rico, J. Kono, and E. Solano, Ultrastrong coupling regimes of light-matter in- teraction, Rev. Mod. Phys. 91, 025005 (2019)
work page 2019
-
[6]
A. F. Kockum, A. Miranowicz, S. De Liberato, S. Savasta, and F. Nori, Ultrastrong coupling between light and matter, Nat. Rev. Phys. 1, 19 (2019)
2019
-
[7]
F. Yoshihara, T. Fuse, Z. Ao, S. Ashhab, K. Kakuyanagi, S. Saito, T. Aoki, K. Koshino, and K. Semba, Inversion of Qubit Energy Levels in Qubit-Oscillator Circuits in the Deep-Strong-Coupling Regime, Phys. Rev. Lett. 120, 183601 (2018)
work page 2018
-
[8]
Ashhab, Superradiance transition in a system with a single qubit and a single oscillator, Phys
S. Ashhab, Superradiance transition in a system with a single qubit and a single oscillator, Phys. Rev. A 87, 013826 (2013)
2013
Show all 59 references
-
[9]
M. J. Hwang, R. Puebla, and M. B. Plenio, Quantum Phase Transition and Universal Dynamics in the Rabi Model, Phys. Rev. Lett. 115, 180404 (2015)
2015
-
[10]
Ashhab, and F
S. Ashhab, and F. Nori, Qubit-oscillator systems in the ultrastrong-coupling regime and their potential for preparing nonclassical states, Phys. Rev. A 81, 042311 (2010)
2010
-
[11]
Leroux, L
C. Leroux, L. C. G. Govia, and A.A. Clerk, Simple vari- ational ground state and pure-cat-state generation in the quantum Rabi model, Phys. Rev. A 96, 043834 (2017)
2017
-
[12]
Gietka, C
K. Gietka, C. Hotter, and H. Ritsch, Unique Steady-Stat e Squeezing in a Driven Quantum Rabi Model, Phys. Rev. Lett. 131, 223604 (2023)
2023
-
[13]
Niemczyk, F
T. Niemczyk, F. Deppe, H. Huebl, E. P. Menzel, F. Hocke, M. J. Schwarz, J. J. Garcia-Ripoll, D. Zueco, T. H¨ummer, E. Solano, A. Marx and R. Gross, Circuit quan- tum electrodynamics in the ultrastrong-coupling regime, Nat. Phys. 6, 772 (2010)
2010
-
[14]
Yoshihara, T
F. Yoshihara, T. Fuse, S. Ashhab, K. Kakuyanagi, S. Saito, and K. Semba, Superconducting qubit-oscillator circuit beyond the ultrastrong-coupling regime, Nat. Phys. 13, 44 (2017)
2017
-
[15]
Braum¨ uller, M
J. Braum¨ uller, M. Marthaler, A. Schneider, A. Stehli, H. Rotzinger, M. Weides, and A. V. Ustinov, Analog quan- tum simulation of the Rabi model in the ultra-strong coupling regime, Nat. Commun. 8, 779 (2017)
2017
-
[16]
Ballester, G
D. Ballester, G. Romero, J. J. Garc ´ ıa-Ripoll, F. Deppe , and E. Solano, Quantum Simulation of the Ultrastrong- Coupling Dynamics in Circuit Quantum Electrodynam- ics, Phys. Rev. X 2, 021007 (2012)
2012
-
[17]
Wallraff, D
A. Wallraff, D. I. Schuster, A. Blais, L. Frunzio, R.S. Huang, J. Majer, S. Kumar, S. M. Girvin, and R. J. 8 Schoelkopf, Strong coupling of a single photon to a super- conducting qubit using circuit quantum electrodynamics, Nature 431, 162 (2004)
2004
-
[18]
G. J. Milburn, and C. A. Holmes, Quantum coherence and classical chaos in a pulsed parametric oscillator with a Kerr nonlinearity, Phys. Rev. A 44, 4704 (1991)
1991
-
[19]
W. Qin, A. Miranowicz, P. B. Li, X. Y. L¨ u, J.Q. You, and F. Nori, Exponentially enhanced light-matter inter- action, cooperativities, and steady-state entanglement using parametric amplification, Phys. Rev. Lett. 120, 093601 (2018)
2018
-
[20]
Leroux, L
C. Leroux, L. C. G. Govia, and A. A. Clerk, Enhanc- ing Cavity Quantum Electrodynamics via Antisqueezing: Synthetic Ultrastrong Coupling, Phys. Rev. Lett. 120, 093602 (2018)
2018
-
[21]
C. J. Zhu, L. L. Ping, Y. P. Yang, and G. S. Agarwal, Squeezed Light Induced Symmetry Breaking Superra- diant Phase Transition, Phys. Rev. Lett. 124, 073602 (2020)
2020
-
[22]
S. Puri, S. Boutin, and A. Blais, Engineering the quan- tum states of light in a Kerr-nonlinear resonator by two- photon driving, npj Quantum Inf. 3, 18 (2017)
2017
-
[23]
Grimm, N
A. Grimm, N. E. Frattini, S. Puri, S. O. Mundhada, S. Touzard, M. Mirrahimi, S. M. Girvin, S. Shankar, and M. H. Devoret, Stabilization and operation of a Kerr-cat qubit, Nature, 584, 205 (2020)
2020
-
[24]
W. Qin, A. F. Kockum, C. S. Munoz, A. Miranowicz, and F. Nori, Quantum amplification and simulation of strong and ultra-strong coupling of light and matter, Physics Reports 1078, 1 (2024)
2024
-
[25]
K. P. Murphy, Machine Learning: A Probabilistic Per- spective, MIT Press (2012)
2012
-
[26]
R. S. Sutton, and A. G. Barto, Reinforcement learning: An introduction, MIT Press (2018)
2018
-
[27]
F¨osel, P
T. F¨osel, P. Tighineanu, T. Weiss, and F. Marquardt, Re- inforcement learning with neural networks for quantum feedback, Phys. Rev. X 8, 031084 (2018)
2018
-
[28]
Dunjko, and H
V. Dunjko, and H. J. Briegel, Machine learning artifi- cial intelligence in the quantum domain: a review of re- cent progress, Reports on Progress in Physics 81, 074001 (2018)
2018
-
[29]
Q. S. Tan, M. Zhang, Y. Chen, J. Q. Liao, and J. Liu, Generation and storage of spin squeezing via learning- assisted optimal control, Phys. Rev. A 103, 032601 (2021)
2021
-
[30]
X. L. Zhao, Y. M. Zhao, M. Li, T. T. Li, Q. Liu, S. Guo,and X. X. Yi, A Strategy for Preparing Quantum Squeezed States Using Reinforcement Learning, Annalen der Physik 536. 2400056 (2024)
2024
-
[31]
Bukov, Reinforcement learning for autonomous preparation of Floquet-engineered states: Inverting the quantum Kapitza oscillator, Phys
M. Bukov, Reinforcement learning for autonomous preparation of Floquet-engineered states: Inverting the quantum Kapitza oscillator, Phys. Rev. B 98, 224305 (2018)
2018
-
[32]
Jaouadi, E
A. Jaouadi, E. Mangaud, M. Desouter-Lecomte, Re- exploring control strategies in a non-Markovian open quantum system by reinforcement learning, Phys. Rev. A 109, 013104 (2024)
2024
-
[33]
Miller, T
R. Miller, T. E. Northup, K. M. Birnbaum, A. D. B. A. Boca, A. D. Boozer, and H. J. Kimble, Trapped atoms in cavity QED: coupling quantized light and matter, J. Phys. B 38, S551 (2005)
2005
-
[34]
J. M. Raimond, M. Brune, and S. Haroche, Manipulat- ing quantum entanglement with atoms and photons in a cavity, Rev. Mod. Phys. 73, 565 (2001)
2001
-
[35]
Puebla, M
R. Puebla, M. J. Hwang, J. Casanova, and M. B. Plenio, Probing the Dynamics of a Superradiant Quantum Phase Transition with a Single Trapped Ion, Phys. Rev. Lett. 118, 073001 (2017)
2017
-
[36]
Q. X. Mei, B. W. Li, Y. K. Wu, M. L. Cai, Y. Wang, L. Yao, Z. C. Zhou, and L. M. Duan, Experimental Real- ization of the Rabi-Hubbard Model with Trapped Ions, Phys. Rev. Lett. 128, 160504 (2022)
2022
-
[37]
Ciuti, and I
C. Ciuti, and I. Carusotto, Input-output theory of cavi - ties in the ultrastrong coupling regime: The case of time- independent cavity parameters, Phys. Rev. A 74, 033811 (2006)
2006
-
[38]
S. Seah, S. Nimmrichter, and V. Scarani, Refrigeration beyond weak internal coupling, Phys. Rev. E 98, 012131 (2018)
2018
-
[39]
A. F. Kockum, A. Miranowicz, V. Macr ` ı, S. Savasta, and F. Nori, Deterministic quantum nonlinear optics with sin- gle atoms and virtual photons, Phys. Rev. A 95, 063849 (2017)
2017
-
[40]
Kullback, and R
S. Kullback, and R. A. Leibler, On Information and Suffi- ciency, Annals of Mathematical Statistics, 22, 79 (1951)
1951
-
[41]
Vidal, and R
G. Vidal, and R. F. Werner, Computable measure of en- tanglement, Phys. Rev. A 65, 032314 (2002)
2002
-
[42]
Peres, Separability Criterion for Density Matrices , Phys
A. Peres, Separability Criterion for Density Matrices , Phys. Rev. Lett. 77, 1413 (1996)
1996
-
[43]
Garziano, A
L. Garziano, A. Ridolfo, R. Stassi, O. Di Stefano, and S. Savasta, Switching on and off of ultrastrong light-matter interaction: Photon statistics of quantum vacuum radia- tion, Phys. Rev. A 88, 063829 (2013)
2013
-
[44]
Garziano, R
L. Garziano, R. Stassi, V. Macr ` ı, A. F. Kockum, S. Savasta, and F. Nori, Multiphoton quantum Rabi os- cillations in ultrastrong cavity QED, Phys. Rev. A 92, 063830 (2015)
2015
-
[45]
X. L. Zhao, Y. L. Ma, H. Y. Ma, T. H. Qiu, and X. X. Yi, Prepare non-classical collective spin state by designing control fields, Phys. Lett. A 425, 127874 (2022)
2022
-
[46]
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, Ma. G. Bellemare, et al., Human-level control through deep reinforcement learning, Nature 518, 529 (2015)
2015
-
[47]
Bellman, On the theory of dynamic programming, Proc
R. Bellman, On the theory of dynamic programming, Proc. Natl. Acad. Sci. 38, 716 (1952)
1952
-
[48]
LeCun, Y
Y. LeCun, Y. Bengio, and G. Hinton, Deep learning, Na- ture 521, 436 (2015)
2015
-
[49]
G. A. Rummery and M. Niranjan, On-line Q-learning using connectionist systems, vol. 37, University of Cambridge, Department of Engineering Cambridge, UK, https://api.semanticscholar.org/CorpusID:59872172 (1994)
1994
-
[50]
Silver, G
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, Proceedings of the 31st International Conference on Machine Learning, PMLR 32, 387 (2014)
2014
-
[51]
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, arXiv:1602.01783 (2016)
2016 arXiv
-
[52]
Z. Y. Wang, T. Schaul, M. Hessel, H. v. Hasselt, M. Lanc- tot, and N. d. Freitas, Dueling Network Architectures for Deep Reinforcement Learning, Proceedings of The 33rd International Conference on Machine Learning, PMLR 48, 1995 (2016)
2016
-
[53]
P. J. Huber, Robust Estimation of a Location Parameter, Annals of Mathematical Statistics 35, 73 (1964)
1964
-
[54]
Loshchilov, F
I. Loshchilov, F. Hutter, Decoupled weight decay regu- larization, arXiv:1711.05101 (2017). 9
2017 arXiv
-
[55]
Beaudoin, J
F. Beaudoin, J. M. Gambetta, and A. Blais, Dissipation and ultrastrong coupling in circuit QED, Phys. Rev. A 84, 043832 (2011)
2011
-
[56]
Viola, and S
L. Viola, and S. Lloyd, Dynamical suppression of deco- herence in two-state quantum systems, Phys. Rev. A 58, 2733 (1998)
1998
-
[57]
S. C. Hou, M. A. Khan, X. X. Yi, D. Dong, and I. R. Petersen, Optimal Lyapunov-based quantum control for quantum systems, Phys. Rev. A 86, 022321 (2012)
2012
-
[58]
G¨ unter, A
G. G¨ unter, A. A. Anappara, J. Hees, A. Sell, G. Bia- siol, L. Sorba, S. De Liberato, C. Ciuti, A. Tredicucci, A. Leitenstorfer, and R. Huber, Sub-cycle switch-on of ultra- strong light-matter interaction, Nature 458, 178 (2009)
2009
-
[59]
X. Y. L¨ u, Y. Wu, J. R. Johansson, H. Jing, J. Zhang, and F. Nori, Squeezed Optomechanics with Phase-Matched Amplification and Dissipation, Phys. Rev. Lett. 114, 093602(2015)
2015
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.