Pith. sign in

REVIEW 1 major objections 8 minor 58 references

Representation Learning with Parameterised Quantum Circuits for Advancing Speech Emotion Recognition

T0 review · 1 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that a hybrid quantum-classical CNN, with a parameterised quantum circuit inserted between convolutional features and the classifier, outperforms an identical classical CNN on four speech-emotion benchmarks while using…

desk verdict A legitimate first application of hybrid PQC-CNN to SER, with honest reporting, but the headline accuracy gain is statistically unsupported and the grid-search protocol needs tightening. read the letter →

arxiv 2501.12050 v3 pith:35FXRMMQ submitted 2025-01-21 cs.LG cs.SDeess.AS

classification cs.LGcs.SDeess.AS
keywords parameterisedquantumcircuitsmachinelearningspeechemotionrecognitionhybridquantum-classicalmodelrepresentationvalenceclassificationunweightedaveragerecall
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that inserting a parameterised quantum circuit (PQC) into a convolutional neural network improves speech emotion recognition on three benchmark datasets while cutting trainable parameters by more than 50 percent. The claim is that quantum superposition and entanglement, realised through the PQC layer, enrich the learned feature representation beyond what the same classical CNN can do. If true, it would point toward smaller, more parameter-efficient models for affective computing without sacrificing accuracy on valence and emotion-class tasks. The evidence is simulated, not run on quantum hardware, so the claim is about the PQC layer's representational contribution in simulation.

What carries the argument

The load-bearing component is the quantum representation-learning block, a PQC layer made of three modules: a quantum embedding (angle, amplitude, or IQP) that maps classical CNN features into an eight-qubit Hilbert space; a circuit layer of random or strongly entangling gates whose rotation angles are trained; and a measurement step (PauliZ, PauliX, Z, or probability) that projects the processed state back to classical values for the classifier. The strongly entangling circuit uses cascaded CNOT gates to build correlations among qubits, which is the paper's concrete mechanism for capturing dependencies between acoustic features.

What would settle it

Re-running the four best grid-search configurations on the same folds with many random seeds and comparing the UAR distributions, or computing a paired test across folds, would settle whether the hybrid's edge is reproducible; if the distributions overlap heavily, the central claim would reduce to a point-estimate artefact.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a CNN whose feature maps are passed through a trainable eight-qubit PQC block outperforms the identical classical CNN on all four evaluated tasks: IEMOCAP binary valence (64.68 UAR vs 61.36), RECOLA binary valence (80.85 vs 74.42), IEMOCAP four-class emotion (55.93 vs 52.54), and MSP-Improv four-class emotion (34.60 vs 32.78). The best hybrid configurations use angle or amplitude embedding, random or strongly entangling circuit layers, and zero weight decay, and the parameter reduction is 50.34 percent. The authors interpret this as evidence that the quantum block contributes to feature representation rather than merely replacing parameters.

Load-bearing premise

The claim collapses if the reported UAR gaps are just random variation, since every hybrid-versus-classical difference in the main table falls within one reported standard deviation and the paper gives no significance test or repeated-seed analysis.

Editorial extensions

If this is right

  • On the paper's results, a PQC layer can replace roughly half the classical parameters of a simple SER CNN while matching or exceeding its UAR on the same data.
  • The consistent selection of zero weight decay suggests the quantum layer may supply its own regularisation, so classical L2 penalties may be unnecessary in hybrid models.
  • Because the optimal embedding and circuit choice differs across datasets, the quantum layer's contribution is configuration-dependent, not automatic.
  • If the parameter reduction transfers to real devices, hybrid SER models would need less memory and energy at inference time than their classical counterparts.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A fair follow-up would re-run the winning configurations across many seeds and report paired differences; the paper's mean gaps all sit within one reported standard deviation, so the improvement may reduce to a point-estimate artefact.
  • The same hybrid block could be tested as a drop-in adapter for stronger SER backbones, such as attention-based or transformer models, to see whether the parameter savings persist.
  • One could ablate the PQC layer against a random fixed nonlinear feature map; if the gain disappears, the advantage may come from extra capacity rather than from quantum-specific correlations.
  • The zero-weight-decay finding could be probed directly by switching L2 regularisation on only for the classical branch; this would test whether the quantum layer truly regularises the whole model.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 8 minor

Summary. The paper proposes a hybrid classical-quantum architecture for speech emotion recognition (SER). A parameterised quantum circuit (PQC) block, comprising a quantum embedding, a variational circuit layer, and a quantum measurement, is inserted between a classical CNN feature extractor and a fully connected classifier. Experiments on IEMOCAP, RECOLA, and MSP-Improv cover binary valence classification and four-class emotion classification. The reported results (Table III) show that the best hybrid configuration selected by grid search achieves higher unweighted average recall (UAR) than a classical CNN baseline on all four tasks while using roughly half the trainable parameters. The authors also recount earlier unsuccessful attempts with static circuits and fusion-based models in Section VI-B.

Significance. If the reported accuracy improvements were statistically reliable, this would be a useful early demonstration that simulated PQC layers can be integrated into a simple SER pipeline while substantially reducing parameter count. The paper has several strengths: the code is publicly available, the authors transparently document unsuccessful earlier designs, and the parameter-reduction claim is an architectural fact verifiable from the reported counts. However, the central empirical claim of improved classification performance is not currently supported, because every reported UAR gain in Table III falls within one reported standard deviation and no significance tests, confidence intervals, or repeated-seed analyses are provided. The model-selection protocol also lacks clarity, which compounds the uncertainty. These issues must be addressed before the performance claim can be accepted.

major comments (1)
  1. [Section V, Figure 5] The classical baseline has roughly twice the trainable parameters of the hybrid model (approximately 2.26 million versus 1.12 million in all four experiments). It is not clear whether this capacity difference is intrinsic to the architectures or a result of different hyperparameter choices, nor whether the classical model received comparable tuning effort. The text says the classical model 'closely mirror[s]' the hybrid architecture, but Figure 5 shows a large parameter gap. Please clarify the architectural differences beyond the presence of the quantum block, and ensure both models are compared under equal tuning effort; otherwise the claimed improvement could reflect capacity or tuning asymmetry rather than the quantum contribution.
minor comments (8)
  1. [Table II] The Learning Rate row lists '0.001, 0.001, 0.00001', duplicating 0.001 and omitting the value 0.0001 that appears in Table III for MSP-Improv; the list should be corrected.
  2. [Section II-C] The sentence 'This paper aims to investigate the of integration QML techniques' is ungrammatical and should be rephrased.
  3. [Section III-A] The text contains '![43]' in the Z Measurement paragraph, where the exclamation mark appears to be a typographical error; the citation should be plain [43].
  4. [Section V-B-2] The dataset name is written as 'MSP-Improve' in one place and 'MSP-Improv' elsewhere; use a consistent spelling.
  5. [Section VI-A and VII] The claim that quantum layers 'inherently provide sufficient regularisation' is an over-interpretation of the observation that all best grid-search configurations had zero weight decay. To support this, the authors should perform controlled experiments varying weight decay for a fixed quantum configuration, or at least explicitly label this as a hypothesis rather than a conclusion.
  6. [Figure 5] The label 'No.Parameters (10^6)' is abbreviated; consider spelling out 'Number of Parameters' for clarity.
  7. [Table II and Section V] The measurement method 'Probability' is offered in the grid search but never appears in a best configuration, and the combined 'Z + PauliZ' measurement used in the best IEMOCAP models is not defined in Section III-A; a one-sentence explanation of the combination and why Probability underperformed would help.
  8. [Section IV-C] The paper states the grid search was 'computationally intensive' without giving run counts or compute time; a brief quantitative description would help readers judge reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the hybrid-vs-classical comparison is an empirical benchmark, not a derivation that reduces to its own inputs.

full rationale

The paper's central claim is that a hybrid CNN+PQC model achieves higher UAR than a purely classical CNN with fewer trainable parameters (Table III, Figure 5). This is an empirical result produced by training both models on the same datasets and comparing the resulting point estimates, so it does not reduce by construction to its inputs: the PQC layer is not defined in terms of the claimed improvement, and the parameter-count reduction is an independent architectural fact. The self-citations present (e.g., Latif et al. [21] for a representation-learning survey, and [50] also by coauthors) are used only as background and are not load-bearing for the hybrid model's performance claim. The quantum components are drawn from external prior work (Havlíček et al. [47], Schuld et al. [48]) and are not justified by an author-specific uniqueness theorem. The main methodological weaknesses are statistical: all four UAR margins in Table III fall within one reported standard deviation, no significance test or repeated-seed analysis is reported, and the grid search selects the best configuration on the same data used for reporting. These are correctness and robustness concerns about whether the improvement is real, not evidence that the derivation is circular. Because no prediction is a renamed fit and no load-bearing step is equivalent to its own input by definition or by self-citation, the circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical entities. Its central claim rests on domain assumptions about feature choice, simulation fidelity, metric validity, and the fairness of the baseline, plus hyperparameters selected by grid search on the same data. None of these are independently verified within the paper.

free parameters (3)
  • Hyperparameter configurations selected by grid search per dataset = IEMOCAP binary: lr=1e-5, Adam, wd=0, Angle, Random, Z+PauliZ; RECOLA: lr=1e-5, SGD, wd=0, Amplitude, Strongly…
    These choices were tuned against the same datasets whose performance is reported, so they are fitted values rather than fixed design constraints.
  • Number of qubits n = 8
    Chosen by the authors and not varied or justified; it bounds the quantum feature space and affects all results.
  • Quantum circuit depth / number of layers = Not specified in text
    The paper specifies embedding, circuit type, and measurement, but does not report the number of layers or repetitions, leaving the capacity of the quantum block underspecified.
assumptions (4)
  • domain assumption Mel-spectrograms with the stated resampling, truncation, and zero-padding settings are adequate input representations for SER.
    Section IV-B motivates this choice but does not compare against raw audio or other feature sets.
  • domain assumption Simulated noiseless PQCs in PennyLane capture the relevant behavior of quantum representation learning for SER.
    Section VI-C acknowledges that the study relies on simulation; the central empirical claim rests on this premise.
  • domain assumption UAR is the appropriate evaluation metric, and trainable parameter count is a meaningful proxy for computational efficiency.
    Section IV-E and V use UAR and parameter count; the parameter count ignores simulation time, energy, and hardware overhead.
  • domain assumption The classical CNN baseline is a fair control for isolating the quantum block's contribution.
    Section III-B acknowledges the baseline omits LSTM and attention layers; the attribution of gains to quantum properties depends on this control being appropriate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Representation Learning with Parameterised Quantum Circuits for Advancing Speech Emotion Recognition." pith.science (2026). https://pith.science/paper/35FXRMMQ

@misc{pith2026250112050,
  author       = {Pith},
  title        = {Pith review of: Representation Learning with Parameterised Quantum Circuits for Advancing Speech Emotion Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/35FXRMMQ}},
  note         = {Machine review of arXiv:2501.12050}
}
read the original abstract

Quantum machine learning (QML) offers a promising avenue for advancing representation learning in complex signal domains. In this study, we investigate the use of parameterised quantum circuits (PQCs) for speech emotion recognition (SER) a challenging task due to the subtle temporal variations and overlapping affective states in vocal signals. We propose a hybrid quantum classical architecture that integrates PQCs into a conventional convolutional neural network (CNN), leveraging quantum properties such as superposition and entanglement to enrich emotional feature representations. Experimental evaluations on three benchmark datasets IEMOCAP, RECOLA, and MSP-IMPROV demonstrate that our hybrid model achieves improved classification performance relative to a purely classical CNN baseline, with over 50% reduction in trainable parameters. This work provides early evidence of the potential for QML to enhance emotion recognition and lays the foundation for future quantum-enabled affective computing systems.

Figures

Figures reproduced from arXiv: 2501.12050 by the authors.

Figure 1
Figure 1. By combining the expressive power of quantum cir [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. A simplified architecture of the proposed Hybrid [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of the model architecture used in this study. Input features pass through a CNN representation learning [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figures from the paper (5 more)
Figure 3
Figure 3. Figure 3: Quantum Circuit Diagram of a Strongly Entangling [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 4
Figure 4. Figure 4: Architecture of the classical model used in this study. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Comparison of UAR (%) and Number of Parameters with Hybrid classical-quantum model and Classical Model in the [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: 2-CNOT, 3-CNOT, and 4-CNOT quantum circuits we [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Two of the model architectures employed in our experiments involving static quantum circuits. [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 52 canonical work pages

  1. [1]

    A survey of speech emotion recognition in natural environment,

    M. Shah Fahad, A. Ranjan, J. Yadav, and A. Deepak, “A survey of speech emotion recognition in natural environment,” Digital Signal Processing, vol. 110, p. 102951, 3 2021

  2. [2]

    Speech Emotion Recognition Us- ing Multi-Layer Sparse Auto-Encoder Extreme Learning Machine and Spectral/Spectro-Temporal Features with New Weighting Method for Data Imbalance,

    F. Daneshfar and S. J. Kabudian, “Speech Emotion Recognition Us- ing Multi-Layer Sparse Auto-Encoder Extreme Learning Machine and Spectral/Spectro-Temporal Features with New Weighting Method for Data Imbalance,” ICCKE 2021 - 11th International Conference on Computer Engineering and Knowledge , pp. 419–423, 2021

  3. [3]

    A review on speech emotion recognition: A survey, recent advances, challenges, and the influence of noise,

    S. M. George and P. Muhamed Ilyas, “A review on speech emotion recognition: A survey, recent advances, challenges, and the influence of noise,” Neurocomputing, vol. 568, p. 127015, 2 2024

  4. [4]

    The role of entanglement for enhancing the efficiency of quantum kernels towards classification,

    D. Sharma, P. Singh, and A. Kumar, “The role of entanglement for enhancing the efficiency of quantum kernels towards classification,” Physica A: Statistical Mechanics and its Applications , vol. 625, p. 128938, 9 2023

  5. [5]

    Transition role of entangled data in quantum machine learning,

    X. Wang, Y . Du, Z. Tu, Y . Luo, X. Yuan, and D. Tao, “Transition role of entangled data in quantum machine learning,” Nature Communications 2024 15:1 , vol. 15, no. 1, pp. 1–8, 5 2024. [Online]. Available: https://www.nature.com/articles/s41467-024-47983-1

  6. [6]

    Exploring the Power of Entangled Data in Quantum Machine Learning,

    X. W ANG, Y . DU, Z. TU, Y . LUO, X. YUAN, and D. TAO, “Exploring the Power of Entangled Data in Quantum Machine Learning,” Wuhan University Journal of Natural Sciences , vol. 29, no. 3, pp. 193–194, 6 2024. 12

  7. [7]

    Entanglement- enhanced Quantum Reinforcement Learning: an Application using Single-Photons,

    J. M. Gaspar, A. Bergerault, V . Apostolou, and A. Ricou, “Entanglement- enhanced Quantum Reinforcement Learning: an Application using Single-Photons,” 2024 IEEE International Conference on Quantum Computing and Engineering (QCE) , pp. 329–334, 9 2024. [Online]. Available: https://ieeexplore.ieee.org/document/10821127/

  8. [8]

    Classification with Quantum Neural Networks on Near Term Processors,

    E. Farhi and H. Neven, “Classification with Quantum Neural Networks on Near Term Processors,” 2 2018. [Online]. Available: https://arxiv.org/abs/1802.06002v2

Show all 58 references
  1. [9]

    Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network,

    G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network,” ICASSP , IEEE International Conference on Acoustics, Speech and Signal Proc...

  2. [10]

    Real time speech emotion recognition using RGB image classification and transfer learning,

    M. N. Stolar, M. Lech, R. S. Bolia, and M. Skinner, “Real time speech emotion recognition using RGB image classification and transfer learning,” 2017, 11th International Conference on Signal Processing and Communication Systems, ICSPCS 2017 - Proceedings , vol. 2018- January, ...

  3. [11]

    GFRN-SEA: Global-Aware Feature Representa- tion Network for Speech Emotion Analysis,

    L. Pan and Q. Wang, “GFRN-SEA: Global-Aware Feature Representa- tion Network for Speech Emotion Analysis,” IEEE Access, 2024

  4. [12]

    Multimodal Emotion Recognition from Raw Audio with Sinc-convolution,

    X. Zhang, W. Fu, and M. Liang, “Multimodal Emotion Recognition from Raw Audio with Sinc-convolution,” 2 2024. [Online]. Available: https://arxiv.org/abs/2402.11954v1

  5. [13]

    CNN+LSTM Architecture for Speech Emotion Recognition with Data Augmentation,

    C. Etienne, G. Fidanza, A. Petrovskii, L. Devillers, and B. Schmauch, “CNN+LSTM Architecture for Speech Emotion Recognition with Data Augmentation,” in Workshop on Speech, Music and Mind (SMM 2018) . ISCA: ISCA, 9 2018

  6. [14]

    An ensemble 1D-CNN-LSTM-GRU model with data augmentation for speech emotion recognition,

    M. Rayhan Ahmed, S. Islam, A. K. Muzahidul Islam, and S. Shatabda, “An ensemble 1D-CNN-LSTM-GRU model with data augmentation for speech emotion recognition,” Expert Systems with Applications, vol. 218, p. 119633, 5 2023

  7. [15]

    Efficient Emotion Recognition from Speech Using Deep Learning on Spectrograms,

    A. Satt, S. Rozenberg, and R. Hoory, “Efficient Emotion Recognition from Speech Using Deep Learning on Spectrograms,” Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, vol. 2017-August, pp. 1089–1093, 2017

  8. [16]

    Speech emotion recognition using deep 1D & 2D CNN LSTM networks,

    J. Zhao, X. Mao, and L. Chen, “Speech emotion recognition using deep 1D & 2D CNN LSTM networks,” Biomedical Signal Processing and Control, vol. 47, pp. 312–323, 1 2019

  9. [17]

    Speech Emotion Classification Using Attention-Based LSTM,

    Y . Xie, R. Liang, Z. Liang, C. Huang, C. Zou, and B. Schuller, “Speech Emotion Classification Using Attention-Based LSTM,” IEEE/ACM Transactions on Audio Speech and Language Processing, vol. 27, no. 11, pp. 1675–1685, 11 2019

  10. [18]

    Automatic speech emo- tion recognition using recurrent neural networks with local attention,

    S. Mirsamadi, E. Barsoum, and C. Zhang, “Automatic speech emo- tion recognition using recurrent neural networks with local attention,” ICASSP , IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, pp. 2227–2231, 6 2017

  11. [19]

    HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

    W. N. Hsu, B. Bolte, Y . H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021. [O...

  12. [20]

    Towards discrimi- native representation learning for speech emotion recognition,

    R. Li, Z. Wu, J. Jia, Y . Bu, S. Zhao, and H. Meng, “Towards discrimi- native representation learning for speech emotion recognition,” IJCAI International Joint Conference on Artificial Intelligence , vol. 2019- August, pp. 5060–5066, 2019

  13. [21]

    Sur- vey of Deep Representation Learning for Speech Emotion Recognition,

    S. Latif, R. Rana, S. Khalifa, R. Jurdak, J. Qadir, and B. Schuller, “Sur- vey of Deep Representation Learning for Speech Emotion Recognition,” IEEE Transactions on Affective Computing , vol. 14, no. 2, pp. 1634– 1654, 4 2023

  14. [22]

    Deep learning approaches for speech emotion recognition: state of the art and research challenges,

    R. Jahangir, Y . W. Teh, F. Hanif, and G. Mujtaba, “Deep learning approaches for speech emotion recognition: state of the art and research challenges,” Multimedia Tools and Applications , vol. 80, no. 16, pp. 23 745–23 812, 7 2021. [Online]. Available: https://link.springer.co...

  15. [23]

    Quantum convolutional neural network based on variational quantum circuits,

    L. H. Gong, J. J. Pei, T. F. Zhang, and N. R. Zhou, “Quantum convolutional neural network based on variational quantum circuits,” Optics Communications, vol. 550, p. 129993, 1 2024

  16. [24]

    Quantum Support Vector Machine for Big Data Classification,

    P. Rebentrost, M. Mohseni, and S. Lloyd, “Quantum Support Vector Machine for Big Data Classification,” Physical Review Letters, vol. 113, no. 13, 9 2014

  17. [25]

    Quantum machine learning,

    J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, “Quantum machine learning,” Nature 2017 549:7671 , vol. 549, no. 7671, pp. 195–202, 9 2017. [Online]. Available: https://www.nature.com/articles/nature23474

  18. [26]

    Variational quantum algorithms,

    M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, “Variational quantum algorithms,” Nature Reviews Physics 2021 3:9 , vol. 3, no. 9, pp. 625–644, 8 2021. [Online]. Available: https://www.natu...

  19. [27]

    Quantum Machine Learning in Feature Hilbert Spaces,

    M. Schuld and N. Killoran, “Quantum Machine Learning in Feature Hilbert Spaces,” Physical Review Letters , vol. 122, no. 4, p. 040504, 2 2019

  20. [28]

    QFSM: A Novel Quantum Federated Learning Algorithm for Speech Emotion Recognition With Minimal Gated Unit in 5G IoV,

    Z. Qu, Z. Chen, S. Dehdashti, and P. Tiwari, “QFSM: A Novel Quantum Federated Learning Algorithm for Speech Emotion Recognition With Minimal Gated Unit in 5G IoV,” IEEE Transactions on Intelligent Vehicles, vol. Early Access, pp. 1–12, 2024. [Online]. Available: https://doi.or...

  21. [29]

    Hybrid Quantum-Classical Convolutional Neural Networks,

    J. Liu, K. H. Lim, K. L. Wood, W. Huang, C. Guo, and H.-L. Huang, “Hybrid Quantum-Classical Convolutional Neural Networks,” Science China: Physics, Mechanics and Astronomy , vol. 64, no. 9, 11 2019. [Online]. Available: http://dx.doi.org/10.1007/s11433-021-1734-3

  22. [30]

    Quantum classical hybrid convolutional neural networks for breast cancer diagnosis,

    Q. Xiang, D. Li, Z. Hu, Y . Yuan, Y . Sun, Y . Zhu, Y . Fu, Y . Jiang, and X. Hua, “Quantum classical hybrid convolutional neural networks for breast cancer diagnosis,” Scientific Reports 2024 14:1, vol. 14, no. 1, pp. 1–13, 10 2024. [Online]. Available: https://www.nature.com...

  23. [31]

    Speech Recognition Using Quantum Convolutional Neural Network,

    B. Thejha, S. Yogeswari, A. Vishalli, and J. Jeyalakshmi, “Speech Recognition Using Quantum Convolutional Neural Network,” Proceed- ings of 8th IEEE International Conference on Science, Technology, Engineering and Mathematics, ICONSTEM 2023 , 2023

  24. [32]

    Quantum Machine Learn- ing for Audio Classification with Applications to Healthcare,

    M. Esposito, G. Uehara, and A. Spanias, “Quantum Machine Learn- ing for Audio Classification with Applications to Healthcare,” 13th International Conference on Information, Intelligence, Systems and Applications, IISA 2022 , 2022

  25. [33]

    Quantum AI in Speech Emotion Recognition,

    M. Norval and Z. Wang, “Quantum AI in Speech Emotion Recognition,” PREPRINT (Version 1) , 9 2024. [Online]. Available: https://www.researchsquare.com/article/rs-4894795/v1

  26. [34]

    Quantum convolutional neural networks,

    I. Cong, S. Choi, and M. D. Lukin, “Quantum convolutional neural networks,” Nature Physics 2019 15:12 , vol. 15, no. 12, pp. 1273– 1278, 8 2019. [Online]. Available: https://www.nature.com/articles/ s41567-019-0648-8

  27. [35]

    DiCOV A Challenge: Dataset, Task, and Baseline System for COVID-19 Diagnosis Using Acoustics,

    A. Muguli, L. Pinto, R. Nirmala, N. Sharma, P. Krishnan, P. K. Ghoshy, R. Kumar, S. Bhat, S. R. Chetupalli, S. Ganapathy, S. Ramoji, and V . Nanda, “DiCOV A Challenge: Dataset, Task, and Baseline System for COVID-19 Diagnosis Using Acoustics,” Proceedings of the Annual Confere...

  28. [36]

    Design of Speech Corpus for Mandarin Text to Speech,

    J. Tao, F. Liu, M. Zhang, and H. Jia, “Design of Speech Corpus for Mandarin Text to Speech,” 2008. [Online]. Available: https://api.semanticscholar.org/CorpusID:15860480

  29. [37]

    The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,

    S. R. Livingstone and F. A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PLOS ONE, vol. 13, no. 5, p. e0196391, 5 2018. [Online]. Available: https: //jou...

  30. [38]

    A Database of German Emotional Speech,

    F. Burkhardt, A. Paeschke, M. Rolfes, W. Sendlmeier, and B. Weiss, “A Database of German Emotional Speech,” in Interspeech, Lisbona, 2005, pp. 1517–1520

  31. [39]

    IEMOCAP: interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: interactive emotional dyadic motion capture database,” Language Resources and Evaluation , vol. 42, no. 4, p. 335, 2008

  32. [40]

    Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions,

    F. Ringeval, A. Sonderegger, J. Sauer, and D. Lalanne, “Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions,” 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition, FG 2013 , 2013

  33. [41]

    MSP-IMPROV: An Acted Corpus of Dyadic Inter- actions to Study Emotion Perception,

    C. Busso, S. Parthasarathy, A. Burmania, M. AbdelWahab, N. Sadoughi, and E. M. Provost, “MSP-IMPROV: An Acted Corpus of Dyadic Inter- actions to Study Emotion Perception,” IEEE Transactions on Affective Computing, vol. 8, no. 1, pp. 67–80, 2017

  34. [42]

    A Multi- Classification Hybrid Quantum Neural Network Using an All-Qubit Multi-Observable Measurement Strategy,

    Y . Zeng, H. Wang, J. He, Q. Huang, and S. Chang, “A Multi- Classification Hybrid Quantum Neural Network Using an All-Qubit Multi-Observable Measurement Strategy,” Entropy, vol. 24, no. 3, p. 394, 3 2022

  35. [43]

    QuaLITi: Quantum Machine Learning Hardware Selection for Inferencing with Top-Tier Performance,

    K. Phalak and S. Ghosh, “QuaLITi: Quantum Machine Learning Hardware Selection for Inferencing with Top-Tier Performance,” 5

  36. [44]

    Zur Quantenmechanik der Stoßvorg ¨ange,

    M. Born, “Zur Quantenmechanik der Stoßvorg ¨ange,” Zeitschrift f ¨ur Physik, vol. 37, no. 12, pp. 863–867, 12 1926. [Online]. Available: https://link.springer.com/article/10.1007/BF01397477

  37. [45]

    Speech emotion recognition with deep convolutional neural networks,

    D. Issa, M. Fatih Demirci, and A. Yazici, “Speech emotion recognition with deep convolutional neural networks,” Biomedical Signal Processing and Control, vol. 59, p. 101894, 5 2020. 13

  38. [46]

    The Impact of Attention Mechanisms on Speech Emotion Recognition,

    S. Chen, M. Zhang, X. Yang, Z. Zhao, T. Zou, and X. Sun, “The Impact of Attention Mechanisms on Speech Emotion Recognition,” Sensors, vol. 21, no. 22, p. 7530, 11 2021

  39. [47]

    Supervised learning with quantum-enhanced feature spaces,

    V . Havl ´ıˇcek, A. D. C ´orcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, “Supervised learning with quantum-enhanced feature spaces,” Nature 2019 567:7747 , vol. 567, no. 7747, pp. 209–212, 3 2019. [Online]. Available: https: //www.nature.com/artic...

  40. [48]

    Circuit-centric quantum classifiers,

    M. Schuld, A. Bocharov, K. M. Svore, and N. Wiebe, “Circuit-centric quantum classifiers,” Physical Review A , vol. 101, no. 3, p. 032308, 3 2020

  41. [49]

    Speech based human emotion recognition using MFCC,

    M. S. Likitha, S. R. R. Gupta, K. Hasitha, and A. U. Raju, “Speech based human emotion recognition using MFCC,” Proceedings of the 2017 International Conference on Wireless Communications, Signal Processing and Networking, WiSPNET 2017 , vol. 2018-January, pp. 2257–2260, 7 2017

  42. [50]

    Direct Modelling of Speech Emotion from Raw Speech,

    S. Latif, R. Rana, S. Khalifa, R. Jurdak, and J. Epps, “Direct Modelling of Speech Emotion from Raw Speech,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 2019, pp. 3920–3924

  43. [51]

    Speech Emotion Recognition using MFCC, GFCC, Chromagram and RMSE features,

    H. Patni, A. Jagtap, V . Bhoyar, and A. Gupta, “Speech Emotion Recognition using MFCC, GFCC, Chromagram and RMSE features,” Proceedings of the 8th International Conference on Signal Processing and Integrated Networks, SPIN 2021 , pp. 892–897, 2021

  44. [52]

    Speech emotion recognition using ANN on MFCC features,

    H. Dolka, M. V . Arul Xavier, and S. Juliet, “Speech emotion recognition using ANN on MFCC features,” 2021 3rd International Conference on Signal Processing and Communication, ICPSC 2021 , pp. 431–435, 5 2021

  45. [53]

    A tutorial on adaptive design optimization,

    J. I. Myung, D. R. Cavagnaro, and M. A. Pitt, “A tutorial on adaptive design optimization,” Journal of Mathematical Psychology , vol. 57, no. 3-4, pp. 53–67, 6 2013

  46. [54]

    Domain adaptation for speech emotion recognition by sharing priors between related source and target classes,

    Q. Mao, W. Xue, Q. Rao, F. Zhang, and Y . Zhan, “Domain adaptation for speech emotion recognition by sharing priors between related source and target classes,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 3 2016, pp. 2608–2612

  47. [55]

    Cross-Corpus Speech Emotion Recognition Based on Few-Shot Learning and Domain Adaptation,

    Y . Ahn, S. J. Lee, and J. W. Shin, “Cross-Corpus Speech Emotion Recognition Based on Few-Shot Learning and Domain Adaptation,” IEEE Signal Processing Letters , vol. 28, pp. 1190–1194, 2021

  48. [56]

    TC-Net: A Modest & Lightweight Emotion Recognition System Using Temporal Convolution Network,

    M. Ishaq, M. Khan, and S. Kwon, “TC-Net: A Modest & Lightweight Emotion Recognition System Using Temporal Convolution Network,” Computer Systems Science and Engineering , vol. 46, no. 3, pp. 3355–3369, 4 2023. [Online]. Available: https://www.techscience.com/ csse/v46n3/52204

  49. [57]

    MSER: Multimodal speech emotion recognition using cross-attention with deep fusion,

    M. Khan, W. Gueaieb, A. El Saddik, and S. Kwon, “MSER: Multimodal speech emotion recognition using cross-attention with deep fusion,” Expert Systems with Applications , vol. 245, p. 122946, 7 2024

  50. [2024]

    Available: https://arxiv.org/pdf/2405.11194

    [Online]. Available: https://arxiv.org/pdf/2405.11194

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.