Pith. sign in

REVIEW 4 major objections 4 minor 52 references

Resolving the HOM coincidence spectrum into a tensor readout yields, the paper claims, an optical actor-critic that beats parameter-matched neural baselines and restores drifted two-qubit gates to above 99.8% of reference fidelity.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A spectrum-resolved HOM interference readout, mapped into an actor-critic, is claimed to outperform matching MLP agents on continuous-control benchmarks and to restore drifted transmon-gate fidelities in simulation.

T0 review reviewed 2026-08-01 challenge →

load-bearing objection The SR-HOM actor-critic idea is worth a serious look, but the paper never specifies the numerical coincidence tensor it actually optimizes, and the natural discretization would make the action readout constant—so the benchmark story is unverifiable as written. the 4 major comments →

arxiv 2607.26438 v1 pith:QQVAGGQA submitted 2026-07-29 quant-ph

Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference

classification quant-ph PACS 42.50.-p03.67.-a
keywords spectrum-resolved Hong-Ou-Mandel interferencequantum optical neural networkreinforcement learningactor-criticcontinuous controlphoton frequency encodingtunable-coupler gate calibrationtwo-photon interference
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to establish that the spectral structure discarded by standard Hong-Ou-Mandel (HOM) coincidence counting can be promoted into a trainable computational resource. It introduces a spectrum-resolved HOM (SR-HOM) readout that outputs a frequency-binned coincidence tensor, uses grouped diagonal entries to generate continuous actions and the full tensor to estimate action values, and reports that this optical actor-critic learns faster and more stably than parameter-matched multilayer-perceptron baselines on five continuous-control benchmarks. The same architecture is then applied to online calibration of drifted tunable-coupler two-qubit gates, restoring CZ and iSWAP fidelities to 0.9917 and 0.9952, over 99.8% of their drift-free references. The broader claim is that photonic interference can act as a nonlinear feature map embedded in a learning architecture, not merely as a similarity measurement. A sympathetic reader would care because it points toward compact optical hardware that outputs structured continuous signals directly from photon statistics.

Core claim

The central claim is that replacing the scalar HOM visibility with a spectrum-resolved coincidence tensor turns the interferometer into a trainable feature map for reinforcement learning. Concretely, an input photon whose spectral mode encodes the environment state interferes with a trainable probe photon; the frequency-resolved coincidence density is binned into a k×k tensor C_ij(x,λ). Grouped diagonal sums, passed through a fixed monotone transformation, produce the continuous action vector, while the tensor's spectral correlations are linearly read out to estimate Q-values. The paper reports that on five continuous-control benchmarks this SR-HOM agent reaches reward thresholds faster than

What carries the argument

The central object is the spectrum-resolved coincidence tensor C_ij(x,λ), the frequency-binned joint probability that two interfering photons are detected in bins i and j, normalized to sum to one. It replaces the scalar coincidence probability of standard HOM detection and serves as a structured optical feature map: grouped diagonal sums define the action through a fixed monotone rescaling, and the tensor's spectral correlations are passed through a trainable linear readout to produce Q-values. The paper also proves a sampling bound — O(ε^-2 log k) coincidence samples suffice to estimate all diagonal entries to accuracy ε with high confidence, independent of the state dimension — which is t

Load-bearing premise

All reported learning curves and the claimed advantage over the neural baseline assume the spectrum-resolved coincidence tensor is read out as an exact, noiseless probability; the supplemental material explicitly states that finite-shot photon-counting fluctuations are not included in the reported learning curves.

What would settle it

Re-run the five benchmarks with the exact tensor replaced by empirical estimates from a finite number M of coincidence samples (e.g., M chosen from the paper's bound with ε=0.05, δ=0.01). If the SR-HOM learning curves collapse to or below the parameter-matched MLP baseline once shot noise is present, the claimed advantage is falsified. A hardware experiment that measures the binned coincidence tensor via frequency-resolved detection on a photonic chip would directly test whether the exact-probability oracle is physically achievable.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If SR-HOM is correct, a compact photonic device (state-encoded single photon, reconfigurable probe, grating demultiplexing) could implement an actor-critic agent whose action generation cost grows only logarithmically in spectral resolution and not with state dimension.
  • The same tensor readout can be repurposed beyond RL: the paper argues that assigning distinct computational roles to diagonal, grouped, and off-diagonal components extends naturally to classification, prediction, and other decision-making tasks.
  • For superconducting processors, the online calibration policy could extend intervals between full recalibrations by tracking slow control-line distortions with low-overhead diagnostics, since the true fidelity is excluded from observations and reward.
  • The reported stability in a strongly capacity-limited regime suggests physical interference primitives may act as well-behaved function approximators where equally sized classical networks fail to converge.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If finite-shot photon-counting noise is introduced at the level allowed by Eq. (8), the performance gap versus the MLP baseline may narrow; a natural next step is to re-run the same benchmarks with empirical coincidence samples and measure the degradation.
  • The diagonal-for-action/off-diagonal-for-value assignment is a general design principle that could transfer to other multi-photon interference platforms, where tensor-shaped readouts might provide different inductive biases than scalar overlaps.
  • The calibration result hints at a broad recipe: any experimental device with a slow drift and a cheap diagnostic signal could be regulated by an RL agent reading a physical interference tensor, without a model of the underlying drift.
  • The authors leave open whether the learned calibration policy generalizes to drift excursions larger than those sampled during training; a testable extension is to evaluate on out-of-distribution drift magnitudes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a spectrum-resolved Hong-Ou-Mandel (SR-HOM) architecture for reinforcement learning, in which the frequency-resolved coincidence tensor C_ij of two interfering single photons is used as a feature map: grouped diagonal entries produce continuous actions, while the tensor (or a related readout) supplies state-action value features. The authors report that a parameter-matched MLP actor-critic is outperformed on five Gymnasium continuous-control benchmarks, with a 4.4x improvement in episodes-to-threshold on LunarLanderContinuous-v3, and that the same architecture learns an online calibration policy that partly recovers drifted CZ and iSWAP gate fidelities. The central numerical experiments use an exact-probability limit for C_ij, as explicitly stated in the Supplemental Material.

Significance. The underlying idea — using the spectral structure of HOM interference as a trainable tensor feature map rather than a scalar overlap — is original and could be of interest to the quantum-optical machine-learning community. The paper is also honest in disclosing that the reported learning curves exclude finite-shot photon-counting fluctuations, and the tunable-coupler calibration environment is described in considerable detail. However, the central numerical object of the paper, the finite-bin C_ij actually computed in the simulations, is never specified. As written, the paper does not provide enough information to verify the benchmark results, and at least one natural implementation of the stated equations makes the actor readout identically constant. Until the numerical model is fully specified and the reported results are reproduced with that model, the empirical claims cannot be interpreted as evidence for the SR-HOM architecture.

major comments (4)
  1. [SM 'Numerical Implementation', Eq. (S2); main Eqs. (3)-(5)] The numerical formula for C_ij is never given. The SM says only that C_ij is 'evaluated directly from the normalized spectral-mode coefficients in the exact-probability limit,' with one complex amplitude per frequency bin. If the spectral amplitude is piecewise constant within each bin, then for i=j the integrand in Eq. (2) is ψ_i φ_i − ψ_i φ_i = 0, so every diagonal entry C_ii is exactly zero. Eq. (5) then makes all actions constant, contradicting every reported learning curve. If a sub-bin integration, a different basis, or a modified tensor was used, that formula must be stated explicitly. This is a load-bearing omission: the central benchmarks depend on the value of C_ii, and the paper currently leaves that object undefined.
  2. [Main 'Complexity Analysis', Eqs. (8)-(9); SM 'Numerical Implementation'] The reported 4.4x 'sample efficiency' is episodes-to-threshold in an exact-probability oracle, not a hardware sample-efficiency result. The SM explicitly says finite-shot photon-counting fluctuations are not included in the learning curves. Eq. (8) is a Hoeffding bound on estimating the diagonal entries p_i from M accepted coincidence samples; it is not an RL sample-complexity bound and does not imply that 666 training episodes correspond to a comparable physical resource cost. A fair comparison would need to count photon-pair samples per action or provide a rigorous bound converting estimation error into policy performance. As it stands, the headline efficiency claim concerns only a noiseless oracle.
  3. [SM 'Critic readout', Eq. (S8); main Eqs. (6)-(7)] The critic actually implemented in the SM uses F^Q_i ∝ |α_i|^2 |u_i|^2 |⟨v,β⟩|^2, a separable product of spectral intensities and an overlap. This is not the spectrum-resolved HOM tensor C_ii defined by Eq. (3); it contains no two-photon interference and no coincidence-tensor structure. Consequently, the benchmark results do not evaluate the 'same SR-HOM architecture' claimed for both actor and critic. The paper should either implement the critic with the actual tensor readout or explicitly acknowledge that the critic is a classical surrogate feature map.
  4. [Main Eqs. (3)-(5); Table S2 (Ant-v5)] Because Σ_{i,j} C_ij = 1 and all entries are nonnegative, Σ_r z_r = trace(C) ≤ 1. Thus the deterministic mean actions are confined to a simplex-constrained subset of the action box. For Ant-v5 (m=8, action bounds [−1,1]), even with an affine map g sending [0,1] to [−1,1], the feasible mean actions satisfy Σ_r (a_r+1)/2 ≤ 1, so most of the action space is inaccessible to the deterministic policy. The paper does not discuss this restriction or its effect on the comparison with an unconstrained MLP. This is especially relevant because the SR-HOM advantage is claimed on exactly these benchmarks.
minor comments (4)
  1. [Throughout] The text contains a recurring typo 'iSW AP' where 'iSWAP' is intended (e.g., in the abstract and Section 'Online RL calibration'). Please correct.
  2. [Benchmark protocol, SM] The paper does not report the number of SR-HOM seeds or the distribution of SR-HOM results, while it does show multi-seed MLP curves. Reporting only the best MLP run and a single (or unstated) SR-HOM trajectory makes the comparison difficult to assess. Error bars or multiple SR-HOM seeds should be provided.
  3. [SM 'Numerical Implementation'] No code or complete hyperparameter list (K, group sizes, learning rates, network widths, seed counts) is provided. Given the missing C_ij formula, a code release or a fully specified implementation appendix is essential for reproducibility.
  4. [Online RL calibration] The gate-calibration results are reported as absolute improvements over the uncorrected baseline, but no MLP or classical controller baseline is included. This makes it difficult to attribute the calibration success specifically to the SR-HOM architecture.

Circularity Check

0 steps flagged

No significant circularity; the reported results are empirical simulation outcomes plus a standard Hoeffding bound, with verifiability gaps but no definitional reduction.

full rationale

The paper's central claims are empirical simulation results, not derived predictions. The SR-HOM tensor is defined by Eq. (2)/(S2); actions and value outputs are explicit functions of diagonal entries (Eqs. (4)-(7) and (S4)-(S11)); and the complexity bound (Eqs. (8)-(9)) is a standard Hoeffding/union-bound estimate for a categorical coincidence distribution. None of these steps fits a parameter to a target and then re-predicts that target. The benchmark comparison (Table S2) reports actual returns from training runs under the same PPO protocol; the MLP baseline is hyperparameter-selected and best-seed reported, so the comparison is not constructed to force the SR-HOM advantage. The gate-calibration experiments train a policy against a diagnostic-based reward and then evaluate fidelity with a separately computed propagator-based metric; while the reward is a surrogate for the components of fidelity, this is ordinary reward design, not a definitional identity or a fitted-input prediction. The SM explicitly states 'Finite-shot photon-counting fluctuations are therefore not included in the reported learning curves' and gives no explicit finite-bin formula for C_ij beyond Eq. (S2), so the exact-probability oracle and the relation of the numerical tensor to Eq. (2) are unverified; these are correctness/verifiability gaps, not circularity. Self-citations [17,18] are contextual and not load-bearing. Therefore no circular step is identified.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 1 invented entities

Ledger summary: the central 'sample efficiency' claim rests on an exact-probability oracle; the calibration numbers rest on a coherent-only drift model with unstated reward weights; the MLP comparison rests on an unstated parameter budget. None of these are fatal by themselves, but they shift the center of gravity of the work from 'measured improvement' to 'plausible simulated improvement'.

free parameters (4)
  • Δf_c (coupler frequency-excursion scale) = selected by scanning 251 values to maximize CZ fidelity (SM, 'Flux-to-frequency model')
    The frequency excursion of the flux pulse is not derived from a device model; it is numerically calibrated to the clean-pulse fidelity before the RL experiments, and all CZ results inherit this calibration.
  • decay times τ_1, τ_2 of transient modes = 47.83 ns, 528.10 ns (taken from Ref. [34])
    These are empirical constants from the cited C78 coupler measurement; not fitted here, but the calibration environment depends on them.
  • reward weights w_φ, w_L, w_Z, w_a, w_Δa = not stated numerically in the manuscript
    The reward shaping coefficients are load-bearing for the calibration result (they define the proxy the policy optimizes) and are not given in the main text or supplement.
  • action update scale s_a, bounds p_max = not stated numerically
    These bound the corrective control parameterization; the trained policy is sensitive to them through Eq. (S18).
axioms (3)
  • domain assumption The coincidence tensor C_ij is evaluated exactly from normalized spectral-mode coefficients in all learning-curve simulations (SM, 'Spectrum-resolved coincidence probability matrix').
    This idealization removes photon-counting noise; all benchmark and calibration training curves assume it.
  • domain assumption A tunable-coupler gate is adequately modeled by three-level modes with the Hamiltonian (S12) and the clipping flux-to-frequency map (S13) without energy relaxation or dephasing.
    The claimed fidelity recovery is defined only within this coherent model; the paper states relaxation and dephasing are not included.
  • domain assumption The MLP baseline comparison is parameter-matched in the sense the authors define (same PPO protocol, same parameter budget).
    The actual MLP architecture budget is not listed in the manuscript; the claim 'parameter-matched' cannot be checked from the text.
invented entities (1)
  • SR-HOM tensor readout as an actor-critic feature map no independent evidence
    purpose: To let HOM interference emit continuous actions and value estimates from one physical module
    This is a new readout construction, not a new physical entity. It has no falsifiable handle outside the paper's own simulations.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference." pith.science (2026). https://pith.science/paper/QQVAGGQA

@misc{pith2026260726438,
  author       = {Pith},
  title        = {Pith review of: Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QQVAGGQA}},
  note         = {Machine review of arXiv:2607.26438}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Hong-Ou-Mandel (HOM) interference-based optical neural networks can offer complexity advantages on benchmark learning tasks, but conventional readout compresses the coincidence spectrum into a single scalar, limiting its use in complex settings such as continuous-action reinforcement learning. Here we introduce a spectrum-resolved HOM (SR-HOM) architecture that promotes the photons' spectral degrees of freedom to a trainable computational resource and use it to construct a compact optical actor-critic agent. Diagonal spectral responses generate continuous actions, while higher-order spectral correlations provide nonlinear state-action features for value estimation. Across five continuous-control benchmarks, SR-HOM outperforms parameter-matched multilayer-perceptron baselines, including a \(4.4\times\) improvement in sample efficiency and a \(74.0\%\) increase in best 100-episode moving-average return for LunarLanderContinuous-v3. Applied to online calibration of drifted tunable-coupler CZ and iSWAP gates for transmon qubits, simulations show it restores fidelities to \(0.9917\) and \(0.9952\) respectively, exceeding \(99.8\%\) of their drift-free calibrated values.

Figures

Figures reproduced from arXiv: 2607.26438 by Chang-Ling Zou, Chenglong You, Guangwei Deng, Jiahua Xu, Luyan Sun, Shan Jin, Shaojun Wu, Xiaoting Wang, Yifang Xu, Zhen Yang.

Figure 1
Figure 1. Figure 1: FIG. 1. Quantum optical reinforcement learning with SR-HOM modules. (a) Actor-critic architecture. The environment [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Smoothed training-return comparison between the parameter-matched MLP baseline and the proposed SR-HOM agent [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 6 linked inside Pith

  1. [1]

    HOM interference 3

    Photon preparation 2. HOM interference 3. Spectrum-resolved detection Input photon 1𝜑𝑥 encodes data 𝑥 SPDC source Trainable probe 1𝜑𝜆 encodes optical parameters λ 50:50 BS Actor and critic share the same SR-HOM structure Diffraction grating Frequency-resolved detector coincidence counting Post-processing SPDC source 1𝜑(𝑥,𝑎) SR-HOM Actor 1𝜑𝑥 1𝜑𝜆 1𝜑𝜅 Action...

  2. [2]

    B. J. Shastri, A. N. Tait, T. Ferreira de Lima, W. H. P. Pernice, H. Bhaskaran, C. D. Wright, and P. R. Prucnal, Nature Photonics15, 102 (2021)

  3. [3]

    P. L. McMahon, Nature Reviews Physics5, 717 (2023)

  4. [4]

    X. Lin, Y. Rivenson, N. T. Yardimci, M. Veli, Y. Luo, M. Jarrahi, and A. Ozcan, Science361, 1004 (2018)

  5. [5]

    Carolan, C

    J. Carolan, C. Harrold, C. Sparrow,et al., Science349, 711 (2015)

  6. [6]

    Slussarenko and G

    S. Slussarenko and G. J. Pryde, Applied Physics Reviews 6, 041303 (2019)

  7. [7]

    Wetzstein, A

    G. Wetzstein, A. Ozcan, S. Gigan, S. Fan, D. Englund, M. Soljaˇ ci´ c, C. Denz, D. A. B. Miller, and D. Psaltis, Nature588, 39 (2020)

  8. [8]

    Miscuglio and V

    M. Miscuglio and V. J. Sorger, Applied Physics Reviews 7, 031404 (2020)

  9. [9]

    Y. Shen, N. C. Harris, S. Skirlo, M. Prabhu, T. Baehr- Jones, M. Hochberg, X. Sun, S. Zhao, H. Larochelle, D. Englund, and M. Soljaˇ ci´ c, Nature Photonics11, 441 (2017)

  10. [10]

    Bowie, S

    C. Bowie, S. Shrapnel, and M. J. Kewming, Quantum Science and Technology9, 015001 (2024)

  11. [11]

    Roncallo, A

    S. Roncallo, A. R. Morgillo, C. Macchiavello, L. Maccone, and S. Lloyd, Communications Physics8, 147 (2025)

  12. [12]

    Roncallo, A

    S. Roncallo, A. R. Morgillo, S. Lloyd, C. Macchi- avello, and L. Maccone, arXiv preprint arXiv:2507.21036 (2025), arXiv:2507.21036 [quant-ph]

  13. [13]

    Minati, S

    G. Minati, S. Roncallo, S. Scrofana, A. R. Morgillo, N. Spagnolo, C. Macchiavello, L. Maccone, V. Ci- mini, and F. Sciarrino, arXiv preprint arXiv:2603.28879 (2026), arXiv:2603.28879 [quant-ph]

  14. [14]

    Roncallo, A

    S. Roncallo, A. R. Morgillo, S. Lloyd, C. Macchi- avello, and L. Maccone, arXiv preprint arXiv:2604.08094 (2026), arXiv:2604.08094 [quant-ph]

  15. [15]

    C. K. Hong, Z. Y. Ou, and L. Mandel, Physical Review Letters59, 2044 (1987)

  16. [16]

    Bouchard, A

    F. Bouchard, A. Sit, Y. Zhang, R. Fickler, F. M. Miatto, Y. Yao, F. Sciarrino, and E. Karimi, Reports on Progress in Physics84, 012402 (2021)

  17. [17]

    Legero, T

    T. Legero, T. Wilk, M. Hennrich, G. Rempe, and A. Kuhn, Physical Review Letters93, 070503 (2004)

  18. [18]

    S. Wu, S. Jin, D. Wen, D. Han, and X. Wang, Quantum 9, 1660 (2025)

  19. [19]

    S. Wu, S. Jin, and X. Wang, in2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (2023) pp. 390–395

  20. [20]

    M. Y. Niu, S. Boixo, V. N. Smelyanskiy, and H. Neven, npj Quantum Information5, 1 (2019)

  21. [21]

    Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, inProceedings of the 33rd International Con- ference on Machine Learning - Volume 48, ICML’16 (New York, NY, USA, 2016) pp. 1329–1338

  22. [22]

    V. V. Sivak, A. Eickbusch, H. Liu, B. Royer, I. Tsioutsios, and M. H. Devoret, Phys. Rev. X12, 011059 (2022)

  23. [23]

    R. S. Sutton and A. G. Barto,Reinforcement Learning: An Introduction, 2nd ed. (The MIT Press, 2018)

  24. [24]

    Silver, G

    D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, inProceedings of the 31st International Conference on Machine Learning, Proceedings of Ma- chine Learning Research, Vol. 32 (PMLR, Beijing, China,

  25. [25]

    T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, Contin- uous control with deep reinforcement learning (2015), arXiv:1509.02971

  26. [26]

    Fujimoto, H

    S. Fujimoto, H. van Hoof, and D. Meger, inProceedings of the 35th International Conference on Machine Learn- ing, Proceedings of Machine Learning Research, Vol. 80 (2018) pp. 1587–1596

  27. [27]

    Haarnoja, A

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, inPro- ceedings of the 35th International Conference on Machine 6 Learning, Proceedings of Machine Learning Research, Vol. 80 (2018) pp. 1861–1870

  28. [28]

    Brecht, D

    B. Brecht, D. V. Reddy, C. Silberhorn, and M. G. Raymer, Physical Review X5, 041017 (2015)

  29. [29]

    J. M. Lukens and P. Lougovski, Optica4, 8 (2017)

  30. [30]

    M. G. Raymer and I. A. Walmsley, Physica Scripta95, 064002 (2020)

  31. [31]

    R.-B. Jin, T. Gerrits, M. Fujiwara, R. Wakabayashi, T. Yamashita, S. Miki, H. Terai, R. Shimizu, and M. Sasaki, Optics Express23, 28836 (2015)

  32. [32]

    F. Yan, P. Krantz, Y. Sung, M. Kjaergaard, D. L. Campbell, J. I. J. Wang, T. P. Orlando, S. Gustavsson, and W. D. Oliver, Physical Review Applied10, 054062 (2018)

  33. [33]

    J. Koch, T. M. Yu, J. Gambetta, A. A. Houck, D. I. Schuster, J. Majer, A. Blais, M. H. Devoret, S. M. Girvin, and R. J. Schoelkopf, Physical Review A76, 042319 (2007)

  34. [34]

    Y. Sung, L. Ding, J. Braum¨ uller, A. Veps¨ al¨ ainen, B. Kan- nan, M. Kjaergaard, A. Greene, G. O. Samach, C. Mc- Nally, D. Kim, A. Melville, B. M. Niedzielski, M. E. Schwartz, J. L. Yoder, T. P. Orlando, S. Gustavsson, and W. D. Oliver, Physical Review X11, 021058 (2021)

  35. [35]

    Li, J.-C

    T.-M. Li, J.-C. Zhang, B.-J. Chen, K. Huang, H.-T. Liu, Y.-X. Xiao, C.-L. Deng, G.-H. Liang, C.-T. Chen, Y. Liu, H. Li, Z.-T. Bao, K. Zhao, Y. Xu, L. Li, Y. He, Z.-H. Liu, Y.-H. Yu, S.-Y. Zhou, Y.-J. Liu, X. Song, D. Zheng, Z. Xiang, Y.-H. Shi, K. Xu, and H. Fan, Physical Review Applied23, 024059 (2025)

  36. [36]

    H. Xu, J. Han, S. Ou, C. Ye, Z. Shen, J. Gao, Y. Wang, T. Che, Y. Song, W. Liu, L. Wang, L.-F. Zhang, P. Zhang, and H.-F. Yu, arXiv: 2606.22376 (2026), arXiv:2606.22376

  37. [37]

    Mohseni, A

    M. Mohseni, A. Scherer, K. G. Johnson, O. Wertheim, M. Otten, N. Anand, N. A. Aadit, Y. Alexeev, G. Ben- Shach, K. M. Bresniker, K. Y. Camsari, B. Chapman, S. Chatterjee, S. Chowdhury,et al., arXiv: 2411.10406 (2024), arXiv:2411.10406

  38. [38]

    Mandel and E

    L. Mandel and E. Wolf,Optical Coherence and Quantum Optics(Cambridge University Press, Cambridge, 1995)

  39. [39]

    Z.-Y. J. Ou,Multi-Photon Quantum Interference (Springer, 2007)

  40. [40]

    Gerrits, F

    T. Gerrits, F. Marsili, V. B. Verma, L. K. Shalm, M. D. Shaw, R. P. Mirin, and S. W. Nam, Physical Review A 91, 013830 (2015)

  41. [41]

    Yepiz-Graciano, A

    P. Yepiz-Graciano, A. M. Angulo Mart ´ ınez, D. Lopez- Mago, H. Cruz-Ramirez, and A. B. U’Ren, Photonics Research8, 1023 (2020)

  42. [42]

    Lavoie, J

    J. Lavoie, J. M. Donohue, L. G. Wright, A. Fedrizzi, and K. J. Resch, Nature Photonics7, 363 (2013)

  43. [43]

    Karpi´ nski, M

    M. Karpi´ nski, M. Jachura, L. J. Wright, and B. J. Smith, Nature Photonics11, 53 (2017)

  44. [44]

    So´ snicki, M

    F. So´ snicki, M. Miko lajczyk, A. Golestani, and M. Karpi´ nski, Nature Photonics17, 761 (2023)

  45. [45]

    H.-H. Lu, M. Liscidini, A. L. Gaeta, A. M. Weiner, and J. M. Lukens, Optica10, 1655 (2023)

  46. [46]

    Buddhiraju, A

    S. Buddhiraju, A. Dutt, M. Minkov, I. A. D. Williamson, and S. Fan, Nature Communications12, 2401 (2021)

  47. [47]

    D. Zhu, C. Chen, M. Yu, L. Shao, Y. Hu, C. J. Xin, M. Yeh, S. Ghosh, L. He, C. Reimer, N. Sinclair, F. N. C. Wong, M. Zhang, and M. Lonˇ car, Light: Science & Ap- plications11, 327 (2022)

  48. [48]

    R. Yang, W. Zhou, D.-J. Guo, H.-M. Ke, L. Tao, Y. Wei, J.-C. Duan, Y. Cui, K. Jia, Z. Xie, Z. Lin, X. Cai, Y.- X. Gong, and S.-N. Zhu, arXiv: 2603.11471 (2026), arXiv:2603.11471

  49. [49]

    Hoeffding, Journal of the American Statistical Asso- ciation58, 13 (1963)

    W. Hoeffding, Journal of the American Statistical Asso- ciation58, 13 (1963)

  50. [50]

    Brockman, V

    G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, Openai gym (2016), arXiv:1606.01540

  51. [51]

    D. C. McKay, C. J. Wood, S. Sheldon, J. M. Chow, and J. M. Gambetta, Phys. Rev. A96, 022330 (2017)

  52. [52]

    Episodes to threshold

    M. Schuld and N. Killoran, Physical Review Letters122, 040504 (2019). 7 Supplemental Material for Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference NUMERICAL IMPLEMENT A TION OF THE SR-HOM AGENT This section describes how the spectrum-resolved Hong-Ou-Mandel (SR-HOM) readout is implemented numerically using a finite-...

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.