{"id":"ae45d3bb-6d43-4618-8ba5-fb4a5dc3b705","arxiv_id":"2507.01235","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A quantum neural network reached 55% test accuracy on 20 pedestrian stress samples, the best score reported, but the difference from classical baselines is not statistically supported.","lead":"This paper applies quantum machine learning models, a quantum support vector machine and a quantum neural network, to classify pedestrian stress from skin conductance data collected in a virtual reality crossing experiment. On a 100-sample dataset split 80/20, the quantum neural network reached 55% test accuracy, the best of the four models tested, but the small test set and missing statistical checks make the claimed advantage unreliable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline comparison rests on a single 80/20 split of 100 events; with 20 test samples, 55% vs 45% is a two-prediction difference and within sampling noise, so the claim that QNN is a 'better classification model' is unsupported.","rationale":"I agree with the reader's verdict and with their identified weakest assumption. The single 80/20 split of 100 events is the load-bearing point: the comparison is a case study of 20 test examples, and the reported gap is two cases. Even granting every implementation detail, no statistical evidence distinguishes the models. This alone justifies retaining the reader's REJECT verdict. I also note an auxiliary internal red flag: the QNN's reported 55% accuracy on a balanced set with precision and recall both 'around 0.75' is arithmetically inconsistent, since balanced data with equal precision and recall r forces accuracy r; this further indicates that the underlying measurements should be audited. However, the sample-size issue is sufficient to decide the verdict.","tokens_in":8577,"tokens_out":11483,"duration_ms":133346,"concrete_test":"Re-run all four models on 200 stratified random 80/20 splits of the same 100-event sample (or 10-fold cross-validation), retraining each model from scratch on every split and recording test accuracy. Compute the paired accuracy difference between QNN and each alternative with a 95% bootstrap or permutation interval; if the QNN advantage does not persist in at least 95% of splits, or if the interval includes 50% accuracy, the headline 'better classification model' claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV-A takes a random sample of 100 SCR events and splits it 80/20; Section V reports all accuracies from that one split. With n_test=20, QNN accuracy 55% is 11/20 and QSVM/classical NN 45% is 9/20, so the entire claimed advantage is two test predictions. The 95% binomial intervals for these counts overlap substantially, and the paper gives no repeated splits, confidence intervals, or paired significance test such as McNemar. Therefore the central comparative claim in the abstract, Section V, and the conclusion is not supported by the evidence reported. The same fragility affects the parameter-efficiency comparison, since it is anchored to these same point estimates.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a case study applying quantum support vector machines (QSVM) and quantum neural networks (QNN) to classify skin conductance response (SCR) events from a virtual-reality pedestrian stress experiment. Three quantum feature maps are compared; the ZZ feature map is selected. The QSVM and QNN are then compared against classical SVM and neural network baselines on a random sample of 100 SCR events with an 80/20 train/test split. The paper claims the QNN achieved a higher test accuracy (55%) than the QSVM and classical versions (45%), and that both quantum models used far fewer parameters, concluding that quantum approaches are more effective for this task. The central evidence is a single split with 20 test samples, and no uncertainty quantification is provided.","tokens_in":8701,"tokens_out":6524,"duration_ms":67865,"significance":"If the comparative claims were statistically supported, this would be a useful data point on parameter-efficient quantum classifiers for small physiological datasets and a concrete comparison of quantum kernels versus variational circuits on a real transportation-related problem. The paper also makes an effort to compare three feature maps and to align quantum and classical model families. However, the headline claim that the QNN is a 'better classification model' is not supported by the evidence as reported: the entire advantage rests on a two-prediction difference in a 20-sample test set, and additional methodological inconsistencies and apparent reporting errors further undermine the conclusions.","major_comments":[{"comment":"The central comparative claim rests on a single 80/20 split of a random sample of 100 SCR events, giving only 20 test samples. With n_test=20, the reported QNN accuracy of 55% (11/20) versus QSVM/classical NN accuracy of 45% (9/20) is a two-prediction difference, well within sampling noise. The paper reports no confidence intervals, repeated splits, cross-validation results, or significance tests such as McNemar's test. The abstract and conclusion therefore overstate the evidence: the claim that the QNN is 'a better classification model' is not supported by the point estimates alone.","section":"V and IV-A"},{"comment":"The class definitions are inconsistent across the compared models. Section III defines a four-class target variable amp_class (classes 0–3), while Sections IV-E and IV-F explicitly reduce the task to binary low/high classification for the classical neural network and the QNN. It is not stated whether the QSVM and classical SVMs are trained on the four-class or binary target. If the models are evaluated on different tasks, the accuracies in Figures 4 and 5 are not directly comparable, and the 'better classification model' claim is ill-posed.","section":"III, IV-E, IV-F"},{"comment":"The metrics reported for the QNN are internally inconsistent. For a balanced binary test set, precision and recall of 0.75 for the 'high amplitude' class imply an accuracy of approximately 75%, not the reported 55%. Specifically, with equal positive and negative examples, accuracy = (1 + 2*recall - recall/precision) / 2 = (1 + 1.5 - 1) / 2 = 0.75. This discrepancy suggests an error in metric computation, class labeling, or the stated class balance, and it casts doubt on the quantitative comparisons in Figures 4 and 5.","section":"V"},{"comment":"The selection of the ZZ feature map is post hoc and biased. Figure 3 shows that the ZZ map has the largest generalization gap (33.75%) and the worst test performance compared to angle encoding (8.75% gap) and amplitude encoding (2.50% gap). The authors state they 'decided to use ZZ map in the rest of the analysis' because it has the highest training accuracy. Choosing the model with the highest training accuracy rewards overfitting and undermines the subsequent conclusion that quantum models overfit; a principled feature-map selection based on validation performance or a held-out set is needed.","section":"V"},{"comment":"The qubit count is inconsistent. The abstract describes an 'eight-qubit ZZ feature map', and Sections IV-D and IV-F refer to an eight-qubit system, but Section IV and Figure 2(a) describe a four-qubit circuit with pairwise interactions only between x0–x1 and x2–x3. Since the number of qubits directly affects the feature space and the circuit complexity, this inconsistency must be resolved for the reported models to be reproducible.","section":"Abstract and IV"}],"minor_comments":[{"comment":"Reference [11] contains corrupted text ('J. M. Gambett that like 6 years to people a') and reference [2] lists 'P. yeah' as an author; both appear to be placeholder or OCR errors and should be corrected.","section":"References"},{"comment":"There are typos in Section V: 'QSVN' should be 'QSVM' in the paragraph comparing QNN and QSVM, and 'overfiting' should be 'overfitting'.","section":"V"},{"comment":"The phrase 'traveller scales person problem' in the background section is garbled; it likely refers to the 'traveling salesman problem' and should be rewritten.","section":"II"},{"comment":"The paper does not report the total size of the original dataset, the class distribution before sampling, or the composition of the 100-sample random draw. This information is necessary for reproducibility and for assessing how representative the sampled subset is.","section":"IV-A"},{"comment":"The precision/recall and F1 values for the classical neural network (0.42/0.24, F1=0.31) are reported without discussion of how they relate to the stated balanced class distribution; the low recall suggests a possible mismatch between the reported accuracy and the class balance.","section":"V"}],"recommendation":"reject","confidential_remarks":"The paper has several load-bearing problems beyond the statistical fragility flagged in the reader's report: the class definitions differ across models, the QNN metrics appear mathematically inconsistent with the reported accuracy, and the qubit count is contradictory between the abstract and the methods. The feature-map selection is also post hoc and biased. While the underlying idea is a reasonable case study, these issues are too central to the main claim and would require substantial recomputation and reanalysis to address. I concur with the reader's reject verdict."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a legitimate application of standard QML building blocks (ZZ feature map, TTN ansatz) to a real SCR pedestrian-stress dataset, and the authors are transparent about their simulator constraints. The new element is the dataset, not the method. The paper's central claim—that the QNN is a better classifier than QSVM and classical versions—is not supported by the evidence as reported.\n\nWhat it does well: the feature-map comparison (angle vs amplitude vs ZZ) is a sensible ablation, and the authors honestly report generalization gaps and overfitting. They also flag the emulator-speed limitation and the inability to do repeated runs. That is the kind of candor you want.\n\nThe soft spots are serious. The whole comparison rests on one 80/20 split of 100 events. With 20 test samples, 55% vs 45% is 11 vs 9 correct—two predictions. The 95% binomial intervals overlap almost completely, and there is no McNemar test or repeated cross-validation. The authors even mention the emulator restricted repeated runs, but that does not make the claim supportable; it means the claim should be downgraded to a single-run observation. On top of that, the ZZ feature map was chosen after seeing training accuracy, so the test comparison inherits selection bias. There are also internal inconsistencies: one section says a four-qubit circuit encoding two features; another says eight qubits with four features; and it is never made explicit whether the SVM and QNN are solving the same binary task, since the NN section describes reducing four amplitude classes to two while the SVM section just says 'binary.' The parameter-efficiency comparison is interesting but anchored to the same noisy point estimates.\n\nNone of this means the raw measurements are fabricated. The paper is a reasonable case study that would be more honest if it presented the numbers as exploratory. For a serious referee, I would condition acceptance on repeated stratified splits (or at least bootstrap CIs) and a clear statement of the classification task for every model.\n\nRecommendation: send to peer review, but expect major revision. The dataset and the QML-in-ITS angle are worth one round of proper statistical work.","headline":"A sincere but statistically under-powered case study; the comparative claim about quantum advantage rests on two test predictions.","tokens_in":9160,"tokens_out":3275,"would_cite":false,"duration_ms":36411,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that a quantum neural network with a tree-tensor-network ansatz classifies pedestrian stress events from skin-conductance data more accurately than a quantum kernel SVM or the classical neural network, while using far…","keywords":["quantum machine learning","quantum neural network","quantum support vector machine","skin conductance response","pedestrian stress","ZZ feature map","tree tensor network","intelligent transportation systems"],"falsifier":"Run the same comparison under repeated random 80/20 splits or k-fold cross-validation and compute confidence intervals for test accuracy; if the QNN is not consistently more accurate than the classical neural network by more than the split-to-split noise, the paper's central performance claim collapses.","tokens_in":8366,"feed_emoji":"🚶","tokens_out":10062,"duration_ms":105454,"temperature":0.7,"pith_summary":"The paper tries to establish that quantum machine learning can be applied to a real transportation problem: classifying pedestrian stress from skin-conductance response (SCR) events captured during a virtual-reality road-crossing experiment. Its main result is that a quantum neural network built on an eight-qubit ZZ feature map and a tree-tensor-network ansatz reaches 55% test accuracy, outperforming both a quantum support vector machine (45%) and a classical neural network (45%) on the same binary high/low stress task. The same report shows the QSVM overfits sharply (78.75% training accuracy versus 45% test), which it attributes to a fixed high-dimensional quantum kernel without regularization. This matters because it is a concrete, real-data comparison of how quantum classifiers behave on a small physiological dataset, including a parameter-efficiency observation: 24 trainable parameters for the QNN versus 145 for the classical network.","feed_headline":"Quantum neural net edges out classical models on pedestrian-stress data","feed_subtitle":"The quantum model scored 55% test accuracy with 24 parameters; the classical neural network scored 45% with 145.","key_machinery":"The central machinery is the pairing of an eight-qubit ZZ feature map with a tree-tensor-network variational ansatz. The ZZ feature map encodes normalized features as $R_z$ rotations on Hadamard-initialized qubits and introduces pairwise CNOT-$R_z$-CNOT interactions, producing the quantum state $|\\phi(x)\\rangle$ whose overlap defines the quantum kernel $k(x,y)=|\\langle\\phi(x)|\\phi(y)\\rangle|^2$. The QNN then feeds this encoded state through a binary-tree circuit of two entangling layers and 24 trainable $R_y$ rotations, measures the Pauli-$Z$ expectation on an output qubit, and maps that value through a sigmoid to a stress-class probability. This trainable structure is what lets the QNN adjust its decision boundary during optimization, whereas the QSVM's kernel remains fixed.","core_discovery":"On the SCR dataset, with identical inputs and preprocessing, the variational QNN reaches 55% test accuracy for the binary classification of high versus low stress amplitude, higher than the QSVM (45%) and the classical neural network (45%), while using 24 trainable parameters instead of 145. The QSVM attains 78.75% training accuracy but only 45% test accuracy, a generalization gap the paper ascribes to the fixed ZZ-feature-map kernel and the absence of regularization. The paper's conclusion is that variational quantum circuits can adapt their decision boundary during training and are therefore more effective on this small, noisy physiological dataset than fixed quantum kernel methods.","pith_inferences":["A natural extension is to match the classical network's capacity to the QNN's 24 parameters; if the classical model then reaches the same accuracy, the reported advantage is about capacity rather than quantum mechanics.","Because the SCR dataset has only four input features, the exponential Hilbert-space advantage of the quantum feature map is not actually exercised; a more informative test would apply the same QNN to higher-dimensional transportation data.","The overfitting pattern suggests quantum kernel methods will need classical-style regularization strategies before they can be trusted on small physiological datasets, a point the paper raises but does not test.","For crossing-safety applications, the practical value of a 55% accuracy classifier depends on the relative cost of missing a high-stress event versus raising a false alarm, so cost-sensitive evaluation would be the next step."],"forward_implications":["If the reported accuracies are stable, variational quantum classifiers can match or beat classical neural networks on small physiological datasets while using far fewer trainable parameters.","The QSVM's large generalization gap implies that expressive quantum feature maps such as the ZZ map need regularization or more training data before they generalize on small intelligent-transportation datasets.","The comparison suggests that quantum-classical benchmarks in transportation should report training-test generalization gaps and parameter counts alongside classification accuracy.","The authors' expectation is that moving from an eight-qubit emulator to 20-30 qubits on actual hardware could improve the accuracy of both quantum models."],"supporting_citations":[{"why":"Supplies the VR road-crossing skin-conductance dataset that all models are trained and evaluated on.","marker":"[3]"},{"why":"Supplies the quantum circuit emulator and automatic differentiation used to implement the QSVM and QNN.","marker":"[4]"},{"why":"Introduces the entangling ZZ feature map and quantum kernel construction the QSVM builds on.","marker":"[11]"},{"why":"Provides the quantum feature Hilbert space formulation used to encode classical inputs into quantum states.","marker":"[16]"},{"why":"Establishes the quantum support vector machine classification approach that the QSVM implementation follows.","marker":"[18]"},{"why":"Supplies the tree tensor network structure used as the variational ansatz in the QNN.","marker":"[20]"},{"why":"Provides the variational quantum neural network classification scheme with expectation-value measurement and thresholding.","marker":"[21]"}],"fun_headline_variants":["Quantum net wins pedestrian-stress test: 55% vs 45%","Variational quantum circuit outperforms fixed kernel on stress data","QNN tops QSVM and classical in pedestrian stress classification","Small quantum model beats classical and kernel on stress data","Pedestrian stress: quantum neural net hits 55% test accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a single random sample of 100 SCR events, split 80/20 into training and test sets, produces stable test accuracies; with 20 test samples the reported 55% versus 45% difference is just two misclassified events and could invert under another split.","fun_headline_variants_meta":{"raw":{"variants":["Quantum net wins pedestrian-stress test: 55% vs 45%","Variational quantum circuit outperforms fixed kernel on stress data","QNN tops QSVM and classical in pedestrian stress classification","Small quantum model beats classical and kernel on stress data","Pedestrian stress: quantum neural net hits 55% test accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000839,"raw_usage":{"total_tokens":3608,"prompt_tokens":850,"completion_tokens":2758,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":2671}},"tokens_in":466,"tokens_out":2758,"duration_ms":21833,"temperature":1.0,"reasoning_tokens":2671,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:56:10.368946+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same comparison under repeated random 80/20 splits or k-fold cross-validation and compute confidence intervals for test accuracy; if the QNN is not consistently more accurate than the classical neural network by more than the split-to-split noise, the paper's central performance claim collapses.","supporting_citations":[{"cited_title":"De- coding pedestrian stress on urban streets using electrodermal activity monitoring in virtual immersive reality,","cited_arxiv_id":null,"evidence_quote":"Supplies the VR road-crossing skin-conductance dataset that all models are trained and evaluated on."},{"cited_title":"Supervised learning with quantum-enhanced feature spaces,","cited_arxiv_id":null,"evidence_quote":"Introduces the entangling ZZ feature map and quantum kernel construction the QSVM builds on."},{"cited_title":"Quantum support vector machine for big data classification,","cited_arxiv_id":null,"evidence_quote":"Establishes the quantum support vector machine classification approach that the QSVM implementation follows."},{"cited_title":"Tree tensor networks for generative modeling,","cited_arxiv_id":null,"evidence_quote":"Supplies the tree tensor network structure used as the variational ansatz in the QNN."}],"review_version":1}