{"id":"815a6e81-2180-4093-9b1b-97007e324a51","arxiv_id":"2507.19041","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"PGKET replaces softmax attention with a photonic Gaussian kernel score, claiming accuracy gains over two quantum transformer baselines on small subsets of MedMNIST v2 and CIFAR-10.","lead":"This paper proposes a Transformer variant whose attention scores are meant to be computed by a photonic circuit using coherent states, and reports small accuracy gains on tiny image classification subsets. A generalist might read it as a test of whether photonic hardware ideas can be bolted onto deep learning, but the paper's central derivation and experiments are too thin to support the claimed advantage.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The photonic circuit in Fig. 1 is not shown to produce Eq. (14): photon counting measures the squared amplitude exp(-|Xi-Xj|^2), not exp(-|Xi-Xj|^2/2), and the beamsplitter network is only asserted to supply exp(Theta,f).","rationale":"The reader's weakest assumption identifies exactly the load-bearing gap: the paper jumps from the displacement-gate amplitude to the PGKSAS formula without specifying a measurement that yields that amplitude as a real score. This is not a stylistic issue; it is the hinge of the paper. If the measured quantity is exp(-|Xi-Xj|^2) rather than exp(-|Xi-Xj|^2/2), then the photonic circuit does not implement the stated kernel attention, and the comparison against HQViT and Quixer is not evidence for the claimed mechanism. The beamsplitter-as-scalar issue reinforces the same conclusion: Eq. (14) contains a trainable scalar exp(Theta,f), but Step 3 gives no unitary construction that multiplies every attention score by that scalar. A beamsplitter network is unitary and norm-preserving; it cannot selectively scale amplitudes without changing the state in ways not accounted for. The experimental sections are too small (150 training samples per dataset, only two quantum baselines, no code) to rescue the claim even if the circuit worked. No ad hominem is intended; the concern is on the derivation and evidence. The reader's REJECT verdict with moderate confidence is appropriate, and no adjustment is needed.","tokens_in":14447,"tokens_out":4533,"duration_ms":49758,"concrete_test":"Simulate the exact Fig. 1 circuit in Strawberry Fields with photon-number-resolving detectors. For two inputs Xi and Xj, compute the vacuum probability P(0) = |<0|D(Xi-Xj)|0>|^2. If P(0) = exp(-|Xi-Xj|^2) instead of exp(-|Xi-Xj|^2/2), then Eq. (14) is not the measured observable. Then specify a POVM (homodyne, heterodyne, or PNR) whose expectation value equals exp(Theta,f) exp(-|Xi-Xj|^2/2), or amend Eq. (14) and the circuit description.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the circuit in Fig. 1 computes PGKSAS(i,j) = exp(Theta,f) exp(-|Xi-Xj|^2/2). Section III-A derives only the state D(Xi-Xj)|0>, whose vacuum amplitude is exp(-|Xi-Xj|^2/2) by Eq. (12). That is a probability amplitude, not a measured score. The gray-box measurement is never specified. With photon-number-resolving detection, the vacuum probability is |<0|D(z)|0>|^2 = exp(-|z|^2), not exp(-|z|^2/2); heterodyne detection gives the Husimi Q function, again exponential in -|z|^2. No standard measurement directly yields the amplitude as a real-valued attention score, so Eq. (14) is not derived from the circuit. Separately, Step 3 asserts that the beamsplitter stack implements exp(Theta,f), but Eq. (8) defines a unitary operator on qumodes, not a scalar multiplication. A unitary network cannot produce an input-independent scalar factor exp(Theta,f) for each pair (i,j) without a derivation that is absent. If Eq. (14) is not what the photonic hardware outputs, the claimed PGKSAM reduces to a classical Gaussian-kernel attention with an added trainable scalar, and the main claimed contribution does not hold.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PGKET, a transformer architecture whose encoder replaces softmax attention with a Photonic Gaussian Kernel Self-Attention Mechanism (PGKSAM). The central object is the Photonic Gaussian Kernel Self-Attention Score PGKSAS(i,j) = exp(Theta,f) exp(-|Xi-Xj|^2/2), claimed to be computed by a photonic circuit using displacement gates and a stacked beamsplitter network. The paper reports noise-free and noisy 5-classification experiments on subsets of MedMNIST v2 and CIFAR-10, comparing PGKET against HQViT and Quixer on accuracy, loss, and convergence epoch.","tokens_in":14773,"tokens_out":4135,"duration_ms":41390,"significance":"If the photonic implementation were rigorously established, using displacement and interferometric operations to compute kernel attention scores in parallel could be a worthwhile direction for photonic machine learning. The paper clearly specifies the intended score formula and includes a systematic small-scale benchmark with a noise-robustness experiment, which is commendable. However, the central derivation connecting the photonic circuit to Eq. (14) is incomplete, and the role of the beamsplitter network as a scalar multiplier is asserted rather than derived. Because these issues concern the main claimed contribution and not merely presentation, the significance of the paper in its current form is limited. The paper also does not provide code, error bars, or statistical tests, so the empirical claims are not yet supported.","major_comments":[{"comment":"The derivation of PGKSAS from the circuit is incomplete: Step 2 obtains |phi> = D(Xi-Xj)|0>, and Eq. (12) gives the vacuum amplitude exp(-|Xi-Xj|^2/2). That quantity is a probability amplitude, not a directly measured attention score. With photon-number-resolving detection, the only measurement described in Section II-B, the probability of the vacuum outcome is |<0|D(z)|0>|^2 = exp(-|z|^2), i.e., exp(-|Xi-Xj|^2), not exp(-|Xi-Xj|^2/2). The gray-box measurement is never specified, so Eq. (14) is not shown to be computable by the proposed circuit.","section":"III-A, Eq. (14), Fig. 1"},{"comment":"The paper states that the beamsplitter stack implements exp(Theta,f), but Eq. (8) defines BS(theta,phi) as a unitary operator on qumodes. A composition of unitary beamsplitter gates remains unitary and cannot multiply every pair's score by an input-state-independent scalar factor exp(Theta,f) unless additional non-unitary components are introduced; no such derivation is provided. Consequently, the second factor in Eq. (14) is asserted rather than derived.","section":"III-A, Step 3, Eq. (8)"},{"comment":"The empirical claim rests on single-run comparisons on very small subsets: 30 training and 10 test images per class for MedMNIST v2 and CIFAR-10. No standard deviations, random seeds, or statistical significance tests are reported, and the reported differences of 1-5% in accuracy are within plausible random variation for these sample sizes. The claim that PGKET outperforms HQViT and Quixer is therefore not supported by the evidence as presented.","section":"IV-A, IV-B, Tabs. III-V"},{"comment":"PGKSAS in Eq. (14) removes the full covariance matrix Gamma of the GKSAM in Eq. (3) and replaces it with identity covariance plus a trainable scalar exp(Theta,f). Since no classical GKSAM baseline is included in the experiments, the reported improvements cannot be attributed to the kernel choice or to photonic processing; a purely classical model using Eq. (14) would be a natural control that is absent.","section":"II-A vs. III-A"}],"minor_comments":[{"comment":"The references to HQViT [27] and Quixer [28] are incorrect; the bibliography lists HQViT as [29] and Quixer as [30].","section":"IV-A"},{"comment":"There are typos and inconsistent terms, including 'Quixe' for Quixer, 'summarry' in Table III, 'muti-head' in Fig. 2, and 'PGKAM' in Definition III.1 where PGKSAM is intended.","section":"Throughout"},{"comment":"The notation BS(theta,phi) = exp(theta,phi) = exp[...] is ambiguous because the left side is a gate name while the argument of the exponential is an operator; the later use of exp(Theta,f) in Eq. (14) is undefined.","section":"II-B, Eq. (8)"}],"recommendation":"reject","confidential_remarks":"The central issue is not a matter of presentation: the paper does not establish that its photonic circuit outputs the quantity in Eq. (14). Standard photon counting gives the squared amplitude, and the beamsplitter network cannot act as a scalar multiplier without additional non-unitary elements. The empirical section is far too small-scale and lacks statistical support, and the absence of a classical GKSAM baseline makes it impossible to isolate the claimed photonic benefit. I see no path to acceptance without a genuinely new measurement scheme and a realistic implementation of exp(Theta,f), together with a properly powered experimental comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The central claim is not supported. The photonic circuit in Fig. 1 produces the coherent state D(Xi-Xj)|0>, whose vacuum amplitude is exp(-|Xi-Xj|^2/2), but the paper never specifies how that amplitude is read out as a real-valued attention score. Photon counting gives the probability exp(-|Xi-Xj|^2); heterodyne detection gives the Husimi Q function. No standard measurement directly yields the amplitude, so Eq. (14) is not derived from the circuit. The beamsplitter network is similarly asserted to implement the trainable scalar exp(Theta, f), but a unitary network cannot produce an input-independent scalar multiplier without a derivation.\n\nWhat the paper does well: it is clearly organized, the displacement-operator identities in Eqs. (10)-(12) are correct, and the authors are upfront about simulating on classical hardware rather than claiming a real photonic experiment (Section IV-A). They also cite the GKSAM that their score turns out to be a special case of.\n\nThe soft spots are proportionally large. The PGKSAS in Eq. (14) is just the GKSAM of [14] with Gamma=I and a learned scalar, so the novelty is thin. The experiments use 150 training samples total (30 per class for 5 classes) and compare only against two quantum transformers, HQViT and Quixer. No error bars, no multiple seeds, no classical baselines, no code, no data. Claims of 1-4% accuracy gains are not convincing at that scale.\n\nI side with the reader's rejection. The derivation gap is the main issue: if the circuit cannot output Eq. (14) directly, the 'photonic' mechanism reduces to a classical Gaussian kernel with a trainable scalar, which is already known.\n\nWho this is for: people working on photonic quantum kernels might find the specific question interesting, but they would need a corrected measurement scheme. The paper is a reasonable starting point, not a finished result.\n\nRecommendation: if an editor is willing to spend a referee on a clearly written but incomplete idea, this could go to review and come back with specific requests (specify the measurement, derive the beamsplitter action, add classical baselines). But a desk reject is also defensible given the low novelty and tiny experiments. I would not cite it in its current form.","headline":"The photonic circuit does not demonstrably compute Eq. (14) because the measurement step is missing, and with tiny experiments and low novelty the paper's central claim fails, though it is clearly written and honest about its limits.","tokens_in":15302,"tokens_out":4150,"would_cite":false,"duration_ms":41445,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PGKET claims a photonic circuit of displacement and beamsplitter gates can compute a Gaussian-kernel self-attention score, and that the resulting transformer outperforms HQViT and Quixer on five-class MedMNIST v2 and CIFAR-10.","keywords":["photonic computing","Gaussian kernel self-attention","displacement gate","beamsplitter network","transformer","image classification","MedMNIST v2","CIFAR-10"],"falsifier":"Simulate the single-qumode circuit $D(X_i)D^\\dagger(X_j)|0\\rangle$ followed by photon-number-resolving detection: the vacuum probability is $e^{-|X_i-X_j|^2}$, not $e^{-|X_i-X_j|^2/2}$, so comparing measured vacuum counts with Eq. (14) settles whether the claimed score is actually what the circuit outputs. A phase-sensitive heterodyne measurement would be needed to recover the amplitude $e^{-|X_i-X_j|^2/2}$.","tokens_in":14187,"feed_emoji":"💡","tokens_out":9662,"duration_ms":82742,"temperature":0.7,"pith_summary":"The paper proposes a transformer variant whose self-attention score is a Gaussian kernel computed by photon displacement gates and interference, replacing the usual softmax attention. The aim is to combine kernel-method attention with photonic parallelism, so attention scores can be processed in parallel and the attention computation can shed the softmax energy term while keeping accuracy. The paper claims this PGKET outperforms the quantum transformers HQViT and Quixer in five-class image classification on MedMNIST v2 and CIFAR-10, and that the advantage persists under 0.4 Gaussian noise. A sympathetic reader would care because it offers a concrete route from an optical displacement identity to a working attention mechanism with measured accuracy gains.","feed_headline":"Photonic circuit computes Gaussian attention scores","feed_subtitle":"A transformer using this photon-based score beats two quantum baselines on small image benchmarks.","key_machinery":"The load-bearing object is the PGKSAS photonic circuit: $d$ qumodes in vacuum, displacement gates $D(X_i)$ and $D^\\dagger(X_j)$ that load two input features, a staircase of beamsplitter gates implementing $\\exp(\\Theta,f)$, and photon detection at the output. The identity carrying the argument is the displacement-operator composition rule, which turns the two displacements into a single $D(X_i-X_j)$, together with the vacuum amplitude formula $D(z)|0\\rangle = e^{-|z|^2/2}\\sum_b z^b/\\sqrt{b!}\\,|b\\rangle$. That pair of relations is what converts a feature difference into the Gaussian weight, so softmax can be replaced by a quantity computed through photon interference and superposition.","core_discovery":"The central claim is that the attention score $PGKSAS(i,j)=\\exp(\\Theta,f)\\exp(-|X_i-X_j|^2/2)$ can be realized on a photonic processor by starting in vacuum, applying displacement gates $D(X_i)$ and $D^\\dagger(X_j)$, reducing the pair to $D(X_i-X_j)$ through the displacement-composition rule, and reading the Gaussian factor off the displacement vacuum amplitude. The stacked beamsplitter network writes the trainable shared weight $\\exp(\\Theta,f)$, and the resulting PGKSAM replaces the softmax attention score with this kernel attention. The paper reports that PGKET achieves higher final accuracy than HQViT and Quixer on all five MedMNIST v2 subsets tested, average gains of about 2.09 and 3.64 percentage points respectively, and on CIFAR-10 reaches 0.7494 accuracy versus 0.7345 and 0.6957, with lower final loss in the CIFAR-10 experiment and better robustness under added Gaussian noise.","pith_inferences":["The paper's Step 2 derives the Gaussian factor from the displacement amplitude, but a photon counter measures $|e^{-|X_i-X_j|^2/2}|^2=e^{-|X_i-X_j|^2}$; recovering Eq. (14) would require a phase-sensitive measurement such as homodyne or heterodyne detection, which the text does not specify.","If the beamsplitter parameters $\\Theta$ and $f$ are trainable, the circuit can learn more than one shared scalar, effectively generalizing the Gaussian kernel to an anisotropic or data-dependent form $\\exp(-\\frac12 (X_i-X_j)^T\\Gamma(X_i-X_j))$ with learned $\\Gamma$.","The same displacement-and-interference construction could be reused for other translation-invariant kernels by changing the gate sequence or readout, so the mechanism is not tied to the specific Gaussian form.","The reported experiments use only 30 training and 10 test images per class, so a natural next test is whether the accuracy advantage persists at full-scale MedMNIST subsets or larger natural-image datasets."],"forward_implications":["Attention scores of the Gaussian-kernel form can be generated in parallel by photonic hardware, so self-attention no longer requires the explicit softmax normalization step at each pairwise comparison.","A transformer using this score reports higher final accuracy than HQViT and Quixer on all five MedMNIST v2 subsets tested: PathMNIST, DermaMNIST, RetinaMNIST, BloodMNIST, and OrganAMNIST.","On CIFAR-10 the reported test accuracy is 0.7494, above HQViT's 0.7345 and Quixer's 0.6957, with the final loss also lower than both baselines.","Under 0.4 Gaussian noise on MedMNIST v2, PGKET reaches 95% of peak accuracy in fewer epochs than both baselines on every subset, indicating noise robustness."],"supporting_citations":[{"why":"Defines the Gaussian kernel self-attention score that PGKSAM replaces softmax with.","marker":"[14]"},{"why":"Supplies the displacement-operator algebra including the composition rule used to derive the Gaussian factor.","marker":"[25]"},{"why":"Supplies the displacement and beamsplitter gate conventions used in the photonic circuit.","marker":"[26]"},{"why":"Provides the MedMNIST v2 benchmark datasets for the five-class image classification experiments.","marker":"[27]"},{"why":"Provides the CIFAR-10 dataset for the natural-image classification experiments.","marker":"[28]"},{"why":"HQViT is the quantum vision transformer baseline that PGKET is compared against.","marker":"[29]"},{"why":"Quixer is the quantum transformer baseline that PGKET is compared against.","marker":"[30]"}],"fun_headline_variants":["Photonic Gaussian kernel attention boosts transformer accuracy","Transformer with photon-based attention beats two quantum baselines","Photonic kernel attention improves transformers on image benchmarks","Gaussian kernel attention via photons outperforms prior transformers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The construction depends on the circuit's measured output being the Gaussian amplitude $e^{-|X_i-X_j|^2/2}$ rather than its squared absolute value, and on the beamsplitter network acting as a simple multiplier $\\exp(\\Theta,f)$; the text specifies neither the measurement scheme nor a derivation of the multiplier.","fun_headline_variants_meta":{"raw":{"variants":["Photonic Gaussian kernel attention boosts transformer accuracy","Transformer with photon-based attention beats two quantum baselines","Photonic kernel attention improves transformers on image benchmarks","Gaussian kernel attention via photons outperforms prior transformers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000309,"raw_usage":{"total_tokens":1728,"prompt_tokens":869,"completion_tokens":859,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":799}},"tokens_in":485,"tokens_out":859,"duration_ms":7557,"temperature":1.0,"reasoning_tokens":799,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:01:59.684733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the single-qumode circuit $D(X_i)D^\\dagger(X_j)|0\\rangle$ followed by photon-number-resolving detection: the vacuum probability is $e^{-|X_i-X_j|^2}$, not $e^{-|X_i-X_j|^2/2}$, so comparing measured vacuum counts with Eq. (14) settles whether the claimed score is actually what the circuit outputs. A phase-sensitive heterodyne measurement would be needed to recover the amplitude $e^{-|X_i-X_j|^2/2}$.","supporting_citations":[{"cited_title":"Gaussian kernelized self-attent ion for long sequence data and its application to CTC-based speech recognition,","cited_arxiv_id":null,"evidence_quote":"Defines the Gaussian kernel self-attention score that PGKSAM replaces softmax with."},{"cited_title":"Quantum mechanical pure states with Gaussian wave functions,","cited_arxiv_id":null,"evidence_quote":"Supplies the displacement-operator algebra including the composition rule used to derive the Gaussian factor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the displacement and beamsplitter gate conventions used in the photonic circuit."},{"cited_title":"MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classiﬁca- tion,","cited_arxiv_id":null,"evidence_quote":"Provides the MedMNIST v2 benchmark datasets for the five-class image classification experiments."}],"review_version":2}