{"id":"4d7598aa-0758-4caa-b27e-7c13bdc14bf1","arxiv_id":"2412.08010","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A quantum-tunnelling neural network is claimed to reproduce human-like classification uncertainty and train 50 times faster than a classical MLP, but the supporting evidence is absent.","lead":"This paper applies a recently proposed quantum-tunnelling neural network to Fashion MNIST and compares its classification outputs with a classical network. It claims the model mimics human-like uncertainty and trains up to 50 times faster, but the evidence is largely qualitative and lacks human or training-time data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'human-like decision-making' claim is never tested against human judgments; Sections 3.1 and 4.1 supply only post-hoc narrative, and the 50x faster claim is asserted without timing data.","rationale":"The reader's strongest claim is two-pronged: the QT-NN is supposed to be both human-like and faster/outperforming. The human-like prong is the more distinctive and the least secure. For it to hold, there must be a measurable correspondence between QT-NN confidence or uncertainty and human confidence or uncertainty on ambiguous images. The paper's Sections 3.1 and 4.1 supply only qualitative interpretations of category confusion and an appeal to quantum-cognition theory; they do not provide a behavioural benchmark. This is an evidentiary gap, not merely a disagreement with the scientific consensus on QCT: the paper's own claim is about observable output behaviour, so a direct comparison to human responses is the appropriate test. I am not asking the authors to settle the philosophical status of QCT; a positive result on the proposed test would support the observed-similarity claim, while a negative result would collapse the human-like claim. The speed claim is also unsupported, but it is secondary to the human-cognition motivation, and it could only matter if the human-like claim is established. I credit the paper for a reproducible architecture and for the concrete near-zero JSD observation in Figure 5c.ii, which is a genuine empirical report about weight-space dynamics; however, it does not by itself support human-likeness. Because both headline prongs are load-bearing and neither is established, the reader's REJECT verdict remains appropriate.","tokens_in":16001,"tokens_out":8182,"duration_ms":85492,"concrete_test":"On the same 50 Fashion MNIST test images used in Figure 3 (preferably a larger balanced sample), collect 10-class forced-choice classifications and a confidence or uncertainty rating from at least 50 human participants; per image, form an empirical human response distribution. Average the Jensen-Shannon divergence between these human distributions and (a) the QT-NN softmax outputs and (b) the classical softmax outputs over all images. If the QT-NN's mean JSD to humans is not significantly lower than the classical model's (using a paired test or bootstrap), then the abstract's 'replicate human-like decision-making' is not supported. This test directly targets the missing behavioural validation; a negative result would also make the faster-training claim irrelevant to the paper's stated human-cognition motivation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To make the central claim of 'replicat[ing] human-like decision-making' (Abstract) hold, the QT-NN's output distributions must be shown to correspond to human perceptual uncertainty on the same stimuli. The paper contains no such evidence: no human participants, no human judgement dataset, no behavioural comparison. Section 3.1 infers human-likeness only from category confusion (e.g., T-Shirt vs Dress; Pullover vs Coat vs Shirt) and Section 4.1 appeals to quantum-cognition theory. These are post-hoc interpretations, not tests; the classical model is sometimes the more uncertain model (Coat, Sandal), so higher Shannon entropy is not uniquely human-like. The 'up to 50 times faster' assertion in Section 3.2 is likewise unsupported by any wall-clock or epoch-to-target measurement. The paper's own statements (Institutional Review Board Statement: Not applicable; Data Availability: no additional data) confirm that no human experiment was run. The near-zero JSD=2.4e-7 for W1 is an interesting weight-space observation, but it is invoked as support for the cognition narrative rather than as evidence about human behaviour. Thus the most distinctive prong of the headline claim lacks a falsifiable basis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a quantum-tunnelling neural network (QT-NN) whose hidden-layer activation is derived from the Schrödinger transmission coefficient, and applies it to Fashion MNIST classification. It compares the QT-NN's output distributions and trained weight distributions with those of a classical multi-layer perceptron using Shannon entropy and Jensen–Shannon divergence, and claims that the QT-NN's output patterns replicate human-like decision-making and that it can be trained up to 50 times faster. The discussion connects the model to quantum cognition theory through a qualitative wave-packet tunnelling simulation and proposes future hybrid quantum-Bayesian architectures.","tokens_in":16241,"tokens_out":4302,"duration_ms":41546,"significance":"If the human-likeness and training-efficiency claims were established, this would be a notable contribution: a qubit-free quantum-inspired neural network with uncertainty-aware outputs and a claimed practical training speed advantage. The paper's use of Shannon entropy and Jensen–Shannon divergence as quantitative descriptors, and its choice of Fashion MNIST as a benchmark, are sensible and align with current interests in uncertainty quantification. However, the headline claims are not supported by the evidence presented: no human behavioural data are used, the 50-times-faster assertion is made without timing measurements, and the connection between Schrödinger tunnelling and human cognition is assumed from prior theory rather than tested. As it stands, the paper is a qualitative exposition with suggestive figures rather than a validated demonstration of human-like decision-making.","major_comments":[{"comment":"The central claim that the QT-NN 'replicat[es] human-like decision-making' (Abstract) is not tested against any human data. Section 3.1 interprets category-level confusion and Shannon-entropy values as evidence of human-likeness, but there are no human participants, no human judgement dataset, and no behavioural comparison; the Institutional Review Board Statement 'Not applicable' and the Data Availability Statement 'no additional data' are consistent with this absence. Higher output entropy is not uniquely human-like: Figure 4 shows that for 'Coat' and 'Sandal' the classical model has higher SE than the QT-NN, so the argument that greater QT-NN uncertainty mimics human ambiguity is selective. A falsifiable test would be to compare the models' output distributions with human perceptual judgements on the same Fashion MNIST images or on stimuli with known ambiguity labels.","section":"Abstract; Section 3.1; Figure 4"},{"comment":"The claim that 'the QT-NN can be trained up to 50 times faster than the classical model' is unsupported. No wall-clock training time, number of epochs to a target accuracy, or any other timing or convergence criterion is reported. The only quantitative evidence offered is JSD = 2.4e-7 between the initial and trained W1 weight distributions for the QT-NN, which shows that the weight distribution barely changes; this does not entail faster training. A direct speed comparison would require identical hardware, implementation, and stopping rules, together with measured epochs or seconds to a specified accuracy.","section":"Section 3.2"},{"comment":"The manuscript does not give the algebraic form of the QT activation function \\phi_QT, instead referring to Ref. [12] for the final ML-adopted forms. Since this activation is the defining ingredient of the model and the paper claims to assess its behaviour, the relevant expressions and the values of all free parameters (barrier thickness and height, learning rate, batch size, hidden-layer size, training schedule) should be specified. Without these, the experiments are not reproducible and the sensitivity of the conclusions to barrier geometry cannot be assessed.","section":"Section 2.1; Section 2.2"},{"comment":"The wave-packet simulation in Figure 7 and the 'ball climbing a wall' analogy are not quantitatively linked to the classification results. The text asserts that the barrier thickness controls model confidence and that the QT-NN is 'inherently more adept at managing ambiguity', but no experiment varies the barrier or compares predicted ambiguity with human perception. The physical model is therefore illustrative rather than evidential for the human-likeness claim.","section":"Section 4.1; Figures 6-7"}],"minor_comments":[{"comment":"The sentence 'Such a configuration has been shown to effectively predict MNIST images with an accuracy of less than 2% [66]' is self-contradictory; 'error rate of less than 2%' is presumably intended.","section":"Section 2.1"},{"comment":"There is a typo in 'Relevant information can be found in the the prior publications [12,53].'","section":"Section 2.1"},{"comment":"The panels do not appear to have axis labels in the text description; the caption should clarify what the bar heights represent (e.g., mean softmax probability over the 50 test images) and how many independent training runs were averaged.","section":"Figure 3"},{"comment":"The training protocol is described as '32 batches of training image sets, with 100 training epochs for each batch', but the batch size and whether this constitutes multiple passes over the full training set are not stated; please clarify.","section":"Section 3.1"},{"comment":"Several references contain typos: Ref. [47] spells 'Neural Networks' as 'Neural Newt.', Ref. [72] spells 'Springer' as 'Spriger', and Ref. [9] should be checked for its volume and article-number formatting.","section":"References"}],"recommendation":"reject","confidential_remarks":"The manuscript relies heavily on prior work by one of the authors (Refs. [12,53,92]) for the activation function and for the quantum-cognition interpretation. For this venue, the incremental contribution is a re-application of an existing model to Fashion MNIST with qualitative uncertainty analysis, and the two headline claims (human-like behaviour and 50x faster training) lack the required evidence. Adding a human-judgement study and direct timing measurements would be substantial new work rather than a minor revision, so I do not see a path to acceptance within the current manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things up front. First, there is a genuinely interesting observation in this paper: the QT-NN's input-layer weights barely change during training (JSD ≈ 2.4e-7), while a classical MLP's weights narrow noticeably. That is a concrete, reproducible finding about an unconventional activation function, and it is not in the earlier QT-NN papers. Second, the paper's headline claims—replicating human-like decision-making, outperforming traditional ML, training up to 50x faster—are not supported by the evidence presented. The gap between the kernel and the claims is large.\n\nWhat the paper does well: it runs a clean comparison of QT-NN and a classical MLP on Fashion MNIST with identical architecture and initialization, and it reports entropy and JSD figures for both output probabilities and weight distributions. The distribution analysis is standard and clearly described. It also honestly shows that the classical model is sometimes more confident (e.g., Ankle Boot, Coat, Sandal), so it is not cherry-picking in the obvious direction. The weight-stability result is the closest thing to a novel empirical contribution, and the paper identifies it as such.\n\nThe soft spots are serious but localized. The 'human-like' claim is never tested against any human behavioral data. The paper says IRB not applicable and no additional data, which confirms no human experiment was run. Sections 3.1 and 4.1 offer post-hoc interpretations of category confusions and an analogy to the Necker cube, but those are not tests. The '50x faster' claim has no wall-clock or epoch-to-target measurements; it is inferred from the weight distributions, which is suggestive but not demonstrated. The 'outperform' claim rests on a single untuned baseline, and the accuracy numbers are close enough that no meaningful superiority is established. The paper also does not consider alternative explanations for the weight stability, such as the QT activation function compressing gradients or the network behaving like a random feature map. The heavy self-citation is not itself a problem—the QT-NN architecture is legitimately prior work—but it should not be mistaken for independent validation.\n\nWho is this for? A reader curious about unusual activation functions and weight dynamics in small networks could get value from the JSD observation. A reader looking for evidence of quantum cognition in AI will not find it here.\n\nRecommendation: this deserves a serious referee, but with the expectation of heavy revision. The weight-stability finding is interesting enough to warrant scrutiny, and a reviewer could push the authors to either test the human-likeness claim directly or drop it. As it stands, the paper overreaches, but it is not empty.","headline":"There is a real empirical kernel (QT-NN weights barely move during training) buried under unsupported headline claims about human-like decision-making and 50x faster training.","tokens_in":16759,"tokens_out":1749,"would_cite":false,"duration_ms":20680,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A quantum-tunnelling activation function yields a neural network that classifies images with human-like confidence and trains up to 50 times faster than a classical network.","keywords":["quantum tunnelling neural networks","quantum cognition theory","uncertainty quantification","Shannon entropy","image classification","Fashion MNIST","human decision-making simulation","confidence estimation"],"falsifier":"Present human participants with the same Fashion MNIST images and record their category choices and confidence; if human uncertainty patterns (e.g., confusion between T-shirt and dress) do not match the QT-NN's softmax entropy ordering across categories, the human-likeness claim is falsified. Alternatively, test whether a classical network with a stochastic activation of similar shape reproduces the same entropy and speedup; if it does, tunnelling is not the active ingredient.","tokens_in":15783,"feed_emoji":"⚛️","tokens_out":6285,"duration_ms":54538,"temperature":0.7,"pith_summary":"This paper argues that replacing a classical neural network's activation function with the quantum-tunnelling transmission coefficient yields a network that classifies images with human-like confidence and uncertainty. On the Fashion MNIST dataset, the quantum-tunnelling neural network (QT-NN) produces higher Shannon entropy for visually ambiguous categories such as T-shirt versus dress, where a human observer would hesitate, while a classical network is overconfident. The authors also report that the QT-NN trains up to 50 times faster than an otherwise identical classical network, because its probabilistic tunnelling activation barely moves the randomly initialised weights during training. If correct, this points to a cheaper way to build uncertainty-aware classifiers that signal doubt instead of failing silently.","feed_headline":"Quantum-tunnelling net trains 50x faster, mirrors human doubt","feed_subtitle":"On Fashion MNIST, the tunnelling network flags ambiguous items with human-like uncertainty and trains far faster.","key_machinery":"The central object is the quantum-tunnelling neural network (QT-NN), whose hidden-layer activation function $\\phi_{QT}$ is the algebraic transmission coefficient $T$ of an electron penetrating a potential barrier, i.e., the solution of the Schrödinger equation for a particle incident on a barrier. This probabilistic activation, combined with white-noise weight initialisation, lets the network keep its weights spread across the full range of possible values instead of converging to fixed points, which the authors identify as the source of both faster training and human-like uncertainty. The analysis tools are Shannon entropy $H(p) = -\\sum_i p(x_i)\\log p(x_i)$ for measuring output uncertainty and the Jensen–Shannon divergence for comparing initial versus trained weight distributions. The barrier thickness acts as a hyperparameter that can be adjusted to control the model's confidence level.","core_discovery":"The paper's central claim is that a neural network whose hidden-layer activation is the transmission coefficient of an electron tunnelling through a potential barrier, computed from the Schrödinger equation, reproduces key features of human perception and judgment in image classification. On Fashion MNIST, the QT-NN and a classical MLP with identical architecture achieve broadly similar accuracies, but their output entropies differ: the QT-NN is markedly less certain on categories that look alike (T-shirt versus dress, shirt versus coat), which the authors take as evidence of ambiguity handling akin to human cognition. The supporting quantitative finding is that training barely changes the QT-NN's input-to-hidden weights (Jensen–Shannon divergence of $2.4\\times10^{-7}$ from the initial random distribution), whereas the classical weights narrow into a tighter Gaussian; from this the authors conclude the QT-NN can be trained up to 50 times faster than the classical model.","pith_inferences":["A direct test of the human-likeness claim would present human participants with the same Fashion MNIST images and compare their category confusions and confidence ratings to the QT-NN's softmax entropy; the paper reports no such behavioural data, so this test is still open.","The 50x speedup is stated relative to a fixed training schedule (32 batches, 100 epochs each); a fairer comparison would train both models to convergence or to a target accuracy, which could change the magnitude of the advantage.","If the tunnelling activation is the active ingredient, a classical stochastic activation with a similarly shaped saturating curve should reproduce the entropy and speedup; if it does, the quantum-cognitive interpretation would not be needed to explain the results.","The barrier-thickness hyperparameter suggests a tunable confidence mechanism: varying it across a test set could reveal a calibration curve connecting the physics parameter to human agreement rates."],"forward_implications":["If the central claim holds, the QT-NN offers a drop-in replacement for the activation layer of a feedforward network that yields uncertainty-aware classifications without Bayesian inference or ensembles.","A network that barely moves its initial weights challenges the standard assumption that extensive weight updates are necessary for learning, implying that task-relevant structure can be exploited through stochastic activations alone.","The reported 50-fold training speedup would make quantum-cognition-inspired models practical for resource-constrained or real-time classification settings.","The entropy patterns on ambiguous Fashion MNIST categories suggest that QT-NN softmax outputs could be used directly as a confidence signal to trigger human review in safety-critical systems.","The framework extends naturally to hybrid quantum-Bayesian architectures, which the authors propose as a route to combine uncertainty quantification with fast tunnelling-based feature extraction."],"supporting_citations":[{"why":"It defines the QT-NN architecture and the algebraic transmission-coefficient activation function that the paper builds on.","marker":"[12]"},{"why":"It supplies quantum cognition theory, the assumed mapping from quantum mechanics to human decision-making.","marker":"[1]"},{"why":"It introduces white-noise weight initialisation and the quantum-inspired model of optical illusions that motivates the ambiguity handling.","marker":"[53]"},{"why":"It provides the superposition-enhanced quantum neural network baseline for multi-class image classification that the QT-NN is compared with.","marker":"[42]"},{"why":"It supplies the Fashion MNIST dataset used for all classification and uncertainty experiments.","marker":"[68]"},{"why":"It grounds the Necker-cube perceptual-ambiguity analogy in a quantum two-level system, tying the tunnelling picture to human perception.","marker":"[60]"}],"fun_headline_variants":["Quantum tunneling net doubts like humans, trains 50x faster","Tunneling neural net flags ambiguous items, learns 50x faster","Quantum net hesitates on lookalikes, trains 50x faster than classic","Human-like uncertainty in quantum tunneling net, 50x faster training","Tunneling model mirrors human judgment, 50x training speedup"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the network's behaviour mimics human cognition rests on the assumption that quantum tunnelling, as described by the Schrödinger equation, is a valid model of human perception and decision-making; the paper never measures human judgements, so this mapping is assumed rather than tested.","fun_headline_variants_meta":{"raw":{"variants":["Quantum tunneling net doubts like humans, trains 50x faster","Tunneling neural net flags ambiguous items, learns 50x faster","Quantum net hesitates on lookalikes, trains 50x faster than classic","Human-like uncertainty in quantum tunneling net, 50x faster training","Tunneling model mirrors human judgment, 50x training speedup"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000678,"raw_usage":{"total_tokens":3022,"prompt_tokens":825,"completion_tokens":2197,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":2102}},"tokens_in":441,"tokens_out":2197,"duration_ms":17240,"temperature":1.0,"reasoning_tokens":2102,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:18:55.433222+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Present human participants with the same Fashion MNIST images and record their category choices and confidence; if human uncertainty patterns (e.g., confusion between T-shirt and dress) do not match the QT-NN's softmax entropy ordering across categories, the human-likeness claim is falsified. Alternatively, test whether a classical network with a stochastic activation of similar shape reproduces the same entropy and speedup; if it does, tunnelling is not the active ingredient.","supporting_citations":[{"cited_title":"Superposition-enhanced quantum neural network for multi-class image classification","cited_arxiv_id":null,"evidence_quote":"It provides the superposition-enhanced quantum neural network baseline for multi-class image classification that the QT-NN is compared with."},{"cited_title":"Graphics and Quantum Mechanics–The Necker Cube as a Quantum-like Two-Level System","cited_arxiv_id":null,"evidence_quote":"It grounds the Necker-cube perceptual-ambiguity analogy in a quantum two-level system, tying the tunnelling picture to human perception."}],"review_version":1}