REVIEW 2 major objections 4 minor 3 cited by
Enhanced Image Recognition Using Gaussian Boson Sampling
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A GBS-based classifier on an 8176-mode photonic processor reaches 95.86% accuracy on MNIST and 85.95% on Fashion-MNIST, beating linear-kernel SVC and prior physical ELM experiments.
desk verdict A genuinely large-scale GBS device wired into an ELM for MNIST is worth a serious look, but the paper overclaims what GBS uniquely buys you, since the feature map rests on classically simulable 16-mode marginals. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a GBS-based feature map: a fixed random nonlinear transformation in which image-derived PCA features select spatial-temporal modes of an 8176-mode interferometer, and the empirical photon-number counts on selected computational bases become the feature vector for a pseudo-inverse-trained linear classifier. ELM uses only these transformed features, while RVFL concatenates the original data with the transformed features. The scheme's key property is that programmability is reduced to choosing which modes to count, while the expensive nonlinearity is supplied by the quantum device.
What would settle it
Take the same PCA preprocessing, mode-selection rule, computational-basis selection, and linear readout, but replace the GBS device with a classical sampler that draws from the 16-mode marginal photon-number distribution of the same Gaussian state (or computes those probabilities exactly); if the accuracy reaches or exceeds 95.86% on MNIST and 85.95% on Fashion-MNIST, the claim that GBS is the enabling resource is falsified.
Extended reading notes
Core claim
The central claim is that GBS can be used as the hidden layer of an ELM/RVFL classifier and that this yields better accuracy than a classical linear SVM and prior optical ELM baselines on standard benchmarks. The classifier's only trainable part is a linear readout; the nonlinear transformation is produced by mapping each PCA feature to a spatial-temporal mode pair in the 8176-mode interferometer, counting samples on selected computational bases, and using renormalized counts as features. The authors demonstrate this on Jiuzhang, with approximately 2200 average photon clicks, and report GRVFL accuracies of 95.86% on MNIST and 85.95% on Fashion-MNIST. They further show that coherent-state input performs worse, that accuracy improves with efficiency and sample count, and that combining GBS features with original data beats using either alone.
Load-bearing premise
The accuracy claim rests on the assumption that the photon-count statistics from the selected modes are a useful feature representation that a classical computer cannot practically reproduce, a comparison the paper does not make.
Editorial extensions
If this is right
- A GBS device can act as a fixed random feature layer for image classification, with training reduced to a pseudo-inverse readout and no iterative optimization.
- The GRVFL variant, which concatenates GBS features with raw pixels, reaches 95.86% on MNIST and 85.95% on Fashion-MNIST, above the linear-kernel SVC baselines of 92.9% and 83.9%.
- Squeezed-state GBS features outperform coherent-state features with the same mean photon number, indicating that the non-classical photon statistics contribute to the transformation.
- Classification accuracy increases with GBS system efficiency and with the number of samples, so hardware improvements should translate directly into better machine-learning performance.
- Because the same mode-selection and counting scheme requires little programmability, the approach can scale to larger mode counts without redesigning the readout.
Reading between the lines
- The feature vector is built from the marginal photon-number statistics on just 16 of 8176 modes per group, and the input Gaussian state is fully characterized; a classical program can therefore sample or compute those marginals directly. Comparing the same feature map fed by classical samples would test whether the quantum device is the actual source of the accuracy gain, a comparison the paper do
- The hyperparameter trends suggest the scheme saturates around 5 million samples; if that saturation is generic, cheaper or faster samplers could reproduce the features, and the practical bottleneck shifts from sampling to the classical readout training.
- The same mode-selection encoding could be applied to other stochastic photonic devices, and the 16-mode group structure leaves room for larger groups or multiple selections per image, which may raise accuracy without retraining the device.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a scheme for image classification using Gaussian boson sampling (GBS) as the nonlinear feature extractor in an extreme learning machine (ELM) and random vector functional-link (RVFL) classifier. Input images are reduced by PCA to M features, each feature selects a spatial-temporal mode pair in an 8176-mode GBS device, and the photon-count statistics on the selected computational bases form the feature vector. The output layer is trained analytically by pseudo-inverse. On MNIST and Fashion-MNIST, the authors report testing accuracies of 95.86% and 85.95% for GRVFL, outperforming linear-kernel SVC and three physical ELM implementations. They also study the effect of the number of computational bases, the number of PCA features, the number of GBS samples, and the device efficiency.
Significance. The experimental execution is a strength: a large-scale photonic device is used in a concrete machine-learning pipeline with standard benchmarks and reproducible analytical training. The hyperparameter and efficiency studies are useful. However, the scientific significance depends on whether GBS provides an advantage over classically available nonlinear feature maps, and this is not established. The feature map uses only the marginal photon-number distribution on a small number of modes, which is classically simulable for Gaussian input states. Without a classical simulation of the same feature map, the results are best viewed as a demonstration of a GBS-implemented feature extractor, not as evidence of quantum-enhanced recognition.
major comments (2)
- [Fourth step (computational basis counting) and Table I] The feature vector is constructed by counting GBS samples on the computational bases of the selected 16 modes per group (32 modes for M=32). For a Gaussian input state and a known linear interferometer, the marginal photon-number distribution on these modes is fully characterized by the corresponding covariance submatrix; sampling from this marginal or computing the 2^m pattern probabilities (m ≤ 32) is classically efficient. The paper does not compare against a classical simulation of this exact feature map. The coherent-state comparison in Table I uses a different marginal distribution, so it cannot isolate whether the squeezed-state feature map itself has any quantum advantage. The authors should either implement a classical simulation of the same feature map (for example, by sampling from the Gaussian marginal or using their MPS samples) and compare accuracies, or restrict the claims to the statement that a GBS device can implement this feature map. As written, the claim that 'GBS is a better choice than coherent states' does not support the title's implication that GBS is the enabling resource.
- [Table I and Fig. 2] The headline accuracies in Table I are reported without error bars or confidence intervals, although the hyperparameter study (Fig. 3) reports double standard error for the same models. The reported margins over the linear-kernel SVC baseline (1.56% on MNIST, 2.05% on Fashion-MNIST) may be sensitive to the random selection of computational bases and to the particular 5-million-sample subset. Additionally, linear-kernel SVC is a weak baseline for these datasets; without stronger classical baselines (for example, RBF-kernel SVM, random forest, or a classical ELM with random features of comparable dimension), the statement that the results 'surpass classical methods' is not well supported.
minor comments (4)
- [First step and Fig. 3(c,g)] The text states that 9 million samples are generated initially, while the main results and hyperparameter default use 5 million samples; the relationship between these two numbers should be clarified explicitly.
- [Reference [7]] Reference [7] contains the informal sentence 'this article will be uploaded to arXiv soon,' which should be replaced with a proper citation or removed.
- [Third step (mode mapping description)] The description of how PCA feature values are mapped to temporal mode indices would benefit from a concrete example or pseudo-code, as the current text is ambiguous about the grouping and partition assignment.
- [Throughout] The abbreviations 'GELM' and 'GRVFL' are used without explicit expansion in the text; they should be defined at first occurrence for readers outside the quantum information community.
Circularity Check
No circularity: the GBS image-recognition results are empirical measurements with a clear training/test separation, not derivations from fitted inputs or from self-citation.
full rationale
The paper reports an experimental machine-learning pipeline: PCA-reduced images are mapped to mode selections of a GBS device, counts on selected computational bases form a feature vector, and a linear classifier is trained by pseudo-inverse. The reported accuracies (95.86% MNIST, 85.95% Fashion-MNIST) are held-out test results on fixed splits, not predictions forced by construction. Hyperparameters such as the number of computational bases, PCA features, and GBS samples are varied and evaluated with 7-fold cross-validation, so they are not fitted to the test labels. The coherent-state comparison is an empirical control with matched average photon number, not a derivation that assumes the conclusion. The prior GBS device papers, including the self-cited Jiuzhang 4.0 reference, supply the experimental apparatus rather than the classification claim. Even though Ref. [7] is an unpublished self-citation, the central accuracy numbers are directly measured and reported in the confusion matrices and Table I, so no load-bearing argument reduces to that citation. The concern that the 16-mode marginal photon-count statistics may be classically simulable is a correctness or interpretation risk, not a circularity: the paper does not claim a formal quantum advantage proof for this task. Thus no step in the claimed derivation chain is equivalent to its own input by construction.
Assumptions & free parameters
free parameters (6)
- Number of selected computational bases N =
3000
- Number of PCA features M =
32
- Number of GBS samples =
5,000,000
- Training images for basis selection n1 =
1000
- Top counts per image n2 =
2000
- Temporal partition boundary =
256th temporal mode
assumptions (4)
- standard math The linear classifier trained by pseudo-inverse is sufficient for the transformed features.
- domain assumption PCA with 32 features retains enough information for classification.
- domain assumption Empirical frequencies over 5 million GBS samples approximate the true GBS probabilities on the selected modes.
- ad hoc to paper Mapping each PCA feature value to a single temporal mode index is a faithful encoding of the input.
Cite this review
Pith. "Pith review of Enhanced Image Recognition Using Gaussian Boson Sampling." pith.science (2026). https://pith.science/paper/JWSOZEZR
@misc{pith2026250619707,
author = {Pith},
title = {Pith review of: Enhanced Image Recognition Using Gaussian Boson Sampling},
year = {2026},
howpublished = {\url{https://pith.science/paper/JWSOZEZR}},
note = {Machine review of arXiv:2506.19707}
}
read the original abstract
Gaussian boson sampling (GBS) has emerged as a promising quantum computing paradigm, demonstrating its potential in various applications. However, most existing works focus on theoretical aspects or simple tasks, with limited exploration of its capabilities in solving real-world practical problems. In this work, we propose a novel GBS-based image recognition scheme inspired by extreme learning machine (ELM) to enhance the performance of perceptron and implement it using our latest GBS device, Jiuzhang. Our approach utilizes an 8176-mode temporal-spatial hybrid encoding photonic processor, achieving approximately 2200 average photon clicks in the quantum computational advantage regime. We apply this scheme to classify images from the MNIST and Fashion-MNIST datasets, achieving a testing accuracy of 95.86% on MNIST and 85.95% on Fashion-MNIST. These results surpass those of classical method SVC with linear kernel and previous physical ELM-based experiments. Additionally, we explore the influence of three hyperparameters and the efficiency of GBS in our experiments. This work not only demonstrates the potential of GBS in real-world machine learning applications but also aims to inspire further advancements in powerful machine learning schemes utilizing GBS technology.
Figures
Forward citations
Cited by 3 Pith papers
-
Robust quantum computational advantage with programmable 3050-photon Gaussian boson sampling
Jiuzhang 4.0 performs Gaussian boson sampling with up to 3050 detected photons and 1024 input squeezed states, claiming a quantum speedup exceeding 10^54 over the best classical simulation.
-
The trainability of photonic quantum circuits
Fixed-order photon-number polynomial observables make passive linear-optical variational circuits trainable with polynomially many samples, while output-probability and high-order observables are exponentially hard to train.
-
Gaussian Boson Sampling for Asset Clustering in Statistical Arbitrage Portfolios
GBS-based clustering (GBS Roots and adapted GBS Boost) produced higher StatArb portfolio returns than classical Spectral/SPONGE clustering in simulated S&P 500 backtests, with the advantage shrinking outside high-vola...
Reference graph
Works this paper leans on
-
[1]
Randomly select n1 images from the training set
-
[2]
For each image, count the number of samples on the computational bases for the selected modes
-
[3]
Select the computational bases with the largest n2 counts
-
[4]
Here, we set n1 = 1000 and n2 = 2000
Select the most frequent N computational bases as the final selected bases. Here, we set n1 = 1000 and n2 = 2000. Considering the structure of our interferometer, we set M = 16k,k ∈ Z+ and partition the features into k groups. For the fixed selected computational bases, the odd-th groups count samples in partition A, while the even-th groups count samples...
work page 2000
-
[5]
C. S. Hamilton, R. Kruse, L. Sansoni, S. Barkhofen, C. Silberhorn, and I. Jex, Gaussian Boson Sampling, Physical Review Letters 119, 170501 (2017)
2017
- [6]
-
[7]
H.-S. Zhong, H. Wang, Y.-H. Deng, M.-C. Chen, L.-C. 6 Peng, Y.-H. Luo, J. Qin, D. Wu, X. Ding, Y. Hu, P. Hu, X.-Y. Yang, W.-J. Zhang, H. Li, Y. Li, X. Jiang, L. Gan, G. Yang, L. You, Z. Wang, L. Li, N.-L. Liu, C.-Y. Lu, and J.-W. Pan, Quantum computational advantage using photons, Gaussian Boson Sampling, Science 370, 1460 (2020)
work page 2020
-
[8]
H.-S. Zhong, Y.-H. Deng, J. Qin, H. Wang, M.-C. Chen, L.-C. Peng, Y.-H. Luo, D. Wu, S.-Q. Gong, H. Su, Y. Hu, P. Hu, X.-Y. Yang, W.-J. Zhang, H. Li, Y. Li, X. Jiang, L. Gan, G. Yang, L. You, Z. Wang, L. Li, N.-L. Liu, J. J. Renema, C.-Y. Lu, and J.-W. Pan, Phase-Programmable Gaussian Boson Sampling Using Stimulated Squeezed Light, Physical Review Letters ...
work page 2021
Show all 41 references
-
[9]
L. S. Madsen, F. Laudenbach, M. F. Askarani, F. Rortais, T. Vincent, J. F. F. Bulmer, F. M. Miatto, L. Neuhaus, L. G. Helt, M. J. Collins, A. E. Lita, T. Gerrits, S. W. Nam, V. D. Vaidya, M. Menotti, I. Dhand, Z. Vernon, N. Quesada, and J. Lavoie, Quantum computational ad- van...
2022
-
[10]
Deng, Y.-C
Y.-H. Deng, Y.-C. Gu, H.-L. Liu, S.-Q. Gong, H. Su, Z.- J. Zhang, H.-Y. Tang, M.-H. Jia, J.-M. Xu, M.-C. Chen, J. Qin, L.-C. Peng, J. Yan, Y. Hu, J. Huang, H. Li, Y. Li, Y. Chen, X. Jiang, L. Gan, G. Yang, L. You, L. Li, H.-S. Zhong, H. Wang, N.-L. Liu, J. J. Renema, C.-Y. Lu,...
2023
-
[11]
H.-L. Liu, H. Su, S.-Q. Gong, Y.-C. Gu, H.-Y. Tang, M.- H. Jia, Q. Wei, D. Wang, M. Zheng, F. Chen, L. Li, S. Ren, X. Zhu, L. Song, P. Yang, J. Chen, H. An, L. Zhang, L. Gan, Y.-H. Deng, H. Wang, H.-S. Zhong, M.-C. Chen, X. Jiang, N.-L. Liu, X.-L. Su, Q. Zhang, C.- Y. Lu, and ...
2025
-
[12]
C. Oh, M. Liu, Y. Alexeev, B. Fefferman, and L. Jiang, Classical algorithm for simulating experimental Gaussian boson sampling, Nature Physics , 1 (2024)
2024
-
[13]
J. M. Arrazola, T. R. Bromley, and P. Rebentrost, Quan- tum approximate optimization with Gaussian boson sam- pling, Physical Review A 98, 012322 (2018)
2018
-
[14]
J. M. Arrazola and T. R. Bromley, Using Gaussian Bo- son Sampling to Find Dense Subgraphs, Physical Review Letters 121, 030503 (2018)
2018
-
[15]
Br´ adler, P.-L
K. Br´ adler, P.-L. Dallaire-Demers, P. Rebentrost, D. Su, and C. Weedbrook, Gaussian boson sampling for perfect matchings of arbitrary graphs, Physical Review A 98, 032310 (2018)
2018
-
[16]
Deng, S.-Q
Y.-H. Deng, S.-Q. Gong, Y.-C. Gu, Z.-J. Zhang, H.-L. Liu, H. Su, H.-Y. Tang, J.-M. Xu, M.-H. Jia, M.-C. Chen, H.-S. Zhong, H. Wang, J. Yan, Y. Hu, J. Huang, W.-J. Zhang, H. Li, X. Jiang, L. You, Z. Wang, L. Li, N.-L. Liu, C.-Y. Lu, and J.-W. Pan, Solving Graph Problems Using G...
2023
-
[17]
J. Huh, G. G. Guerreschi, B. Peropadre, J. R. McClean, and A. Aspuru-Guzik, Boson sampling for molecular vi- bronic spectra, Nature Photonics 9, 615 (2015)
2015
-
[18]
Jahangiri, J
S. Jahangiri, J. Miguel Arrazola, N. Quesada, and A. Del- gado, Quantum algorithm for simulating molecular vibra- tional excitations, Physical Chemistry Chemical Physics 22, 25528 (2020)
2020
-
[19]
Jahangiri, J
S. Jahangiri, J. M. Arrazola, and A. Delgado, Quan- tum Algorithm for Simulating Single-Molecule Electron Transport, The Journal of Physical Chemistry Letters 12, 1256 (2021)
2021
-
[20]
Shang, H.-S
Z.-X. Shang, H.-S. Zhong, Y.-K. Zhang, C.-C. Yu, X. Yuan, C.-Y. Lu, J.-W. Pan, and M.-C. Chen, Boson sampling enhanced quantum chemistry, 2403.16698
-
[21]
J. Shi, Y. Tang, Y. Lu, Y. Feng, R. Shi, and S. Zhang, Quantum Circuit Learning With Parameterized Boson Sampling, IEEE Transactions on Knowledge and Data Engineering 35, 1965 (2023)
2023
-
[22]
Sakurai, A
A. Sakurai, A. Hayashi, W. Munro, and K. Nemoto, Quantum Optical Reservoir Computing powered by Boson Sampling, Optica Quantum 10.1364/OPTI- CAQ.541432 (2025)
2025 doi
-
[23]
Yu, Z.-P
S. Yu, Z.-P. Zhong, Y. Fang, R. B. Patel, Q.-P. Li, W. Liu, Z. Li, L. Xu, S. Sagona-Stophel, E. Mer, S. E. Thomas, Y. Meng, Z.-P. Li, Y.-Z. Yang, Z.-A. Wang, N.- J. Guo, W.-H. Zhang, G. K. Tranmer, Y. Dong, Y.-T. Wang, J.-S. Tang, C.-F. Li, I. A. Walmsley, and G.-C. Guo, A uni...
2023
-
[24]
Fujiyoshi, T
H. Fujiyoshi, T. Hirakawa, and T. Yamashita, Deep learning-based image recognition for autonomous driv- ing, IATSS Research 43, 244 (2019)
2019
-
[25]
L. Cai, J. Gao, and D. Zhao, A review of the applica- tion of deep learning in medical image classification and segmentation, Annals of Translational Medicine 8, 713 (2020)
2020
-
[26]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cour- napeau, M. Brucher, M. Perrot, and ´E. Duchesnay, Scikit-learn: Machine Learning in Python, J. Mach. Learn. Res. 12,...
2011
-
[27]
Lecun, L
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, Gradient-based learning applied to document recogni- tion, Proceedings of the IEEE 86, 2278 (1998)
1998
-
[28]
Huang, Q.-Y
G.-B. Huang, Q.-Y. Zhu, and C.-K. Siew, Extreme learn- ing machine: Theory and applications, Neurocomputing Neural Networks, 70, 489 (2006)
2006
-
[29]
Pao, G.-H
Y.-H. Pao, G.-H. Park, and D. J. Sobajic, Learning and generalization characteristics of the random vector functional-link net, Neurocomputing Backpropagation, Part IV, 6, 163 (1994)
1994
-
[30]
Deng, The MNIST Database of Handwritten Digit Im- ages for Machine Learning Research [Best of the Web], IEEE Signal Processing Magazine 29, 141 (2012)
L. Deng, The MNIST Database of Handwritten Digit Im- ages for Machine Learning Research [Best of the Web], IEEE Signal Processing Magazine 29, 141 (2012)
2012
-
[31]
H. Xiao, K. Rasul, and R. Vollgraf, Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms (2017), cs.LG/1708.07747
2017 arXiv
-
[32]
K. P. and, Liii. on lines and planes of closest fit to systems of points in space, The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science 2, 559 (1901)
1901
-
[33]
Hotelling, Analysis of a complex of statistical vari- ables into principal components, Journal of Educational Psychology 24, 417 (1933)
H. Hotelling, Analysis of a complex of statistical vari- ables into principal components, Journal of Educational Psychology 24, 417 (1933)
1933
-
[34]
Suprano, D
A. Suprano, D. Zia, L. Innocenti, S. Lorenzo, V. Ci- mini, T. Giordani, I. Palmisano, E. Polino, N. Spagnolo, F. Sciarrino, G. M. Palma, A. Ferraro, and M. Paternos- tro, Experimental Property Reconstruction in a Photonic 7 Quantum Extreme Learning Machine, Physical Review Let...
2024
-
[35]
Cimini, M
V. Cimini, M. M. Sohoni, F. Presutti, B. K. Malia, S.- Y. Ma, R. Yanagimoto, T. Wang, T. Onodera, L. G. Wright, and P. L. McMahon, Large-scale quantum reser- voir computing using a Gaussian Boson Sampler (2025), 2505.13695 [quant-ph]
2025 arXiv
-
[36]
Azam and R
P. Azam and R. Kaiser, Optically accelerated extreme learning machine using hot atomic vapors, Physical Re- view Applied 22, 034041 (2024)
2024
-
[37]
Pierangeli, G
D. Pierangeli, G. Marcucci, and C. Conti, Photonic ex- treme learning machine by free-space optical propaga- tion, Photonics Research 9, 1446 (2021)
2021
-
[38]
Sakurai, M
A. Sakurai, M. P. Estarellas, W. J. Munro, and K. Nemoto, Quantum Extreme Reservoir Computation Utilizing Scale-Free Networks, Physical Review Applied 17, 064044 (2022)
2022
-
[39]
De Lorenzis, M
A. De Lorenzis, M. Casado, M. Estarellas, N. Lo Gullo, T. Lux, F. Plastina, A. Riera, and J. Settino, Harnessing quantum extreme learning machines for image classifica- tion, Physical Review Applied 23, 044024 (2025)
2025
-
[40]
C. L. P. Chen and Z. Liu, Broad Learning System: An Ef- fective and Efficient Incremental Learning System With- out the Need for Deep Architecture, IEEE Transactions on Neural Networks and Learning Systems 29, 10 (2018)
2018
-
[41]
Kohavi, A study of cross-validation and bootstrap for accuracy estimation and model selection, in Ijcai, Vol
R. Kohavi, A study of cross-validation and bootstrap for accuracy estimation and model selection, in Ijcai, Vol. 14 (Montreal, Canada, 1995) pp. 1137–1145
1995
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.