REVIEW 4 major objections 7 minor 50 references
Learning Complex Word Embeddings in Classical and Quantum Spaces
T0 review · 4 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper shows that complex-valued word embeddings, including quantum states prepared by parameterised circuits, can match classical Skip-gram on similarity benchmarks via a two-stage pipeline that scales to a 426k vocabulary.
desk verdict Useful, honest scaling pipeline for quantum word embeddings, but the headline claim of parity with classical is a few points overstated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The two-stage pipeline is the load-bearing mechanism. Stage one is a complex-valued Skip-gram with negative sampling, implemented by modifying the C word2vec code to store real and imaginary parts separately and to feed the scaled fidelity $F_D(v_f,v_c) = D(2|\langle v_f|v_c\rangle|^2 - 1)$ into a sigmoid loss; for the PyTorch models the overlap $|\langle v_f|v_c\rangle|^2$ is also used directly as a probability in the cross-entropy loss. Stage two fits, for each word, a parameterised quantum circuit—a unitary built from single-qubit Y-rotations and controlled X-rotations, with three layers of Ansatz 5 or Ansatz 14—to the normalised complex vector from stage one, using the overlap as the fitting loss. The fidelity overlap is the object that connects classical complex geometry to quantum states: normalised complex vectors are quantum states, and the squared absolute inner product is the Born-rule probability between pure states.
What would settle it
Compute the final per-word fidelity loss for all 426,507 fitted PQCs and inspect the tail; if a non-negligible number of words, or a distinct semantic class of words, have loss far above zero, the claim that the fitted PQCs reproduce the complex embeddings across the vocabulary fails. A simpler check is to hold out a random set of words from the fitting stage and see whether their PQC embeddings still match the target vectors.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that parameterised quantum circuits can serve as word embeddings without a quality penalty, provided they are fitted in a second stage to complex embeddings produced by an efficient classical Skip-gram variant. Directly training PQCs as part of the Skip-gram objective degrades performance, especially when both focal and context words are circuits, but the two-stage procedure—first learn arbitrary complex vectors from the corpus, then fit a PQC to each normalised vector—yields PQC embeddings whose WordSim353, MEN, and RG-65 correlations match the complex vectors exactly, and SCWS within 0.2 points. The paper attributes this to the expressivity of the ansatz: three layers of Ansatz 5 or Ansatz 14 fit the arbitrary complex states with the loss 'effectively going to zero.' Because the fitting stage only touches each vocabulary item and not each token, the route scales with vocabulary size rather than corpus size.
Load-bearing premise
The claim that the whole 426k-vocabulary PQC set is high quality rests on the assumption that three layers of the chosen ansatz can fit every word's complex embedding to effectively zero loss; the paper reports that loss 'effectively went to zero' but verifies it mainly through identical WordSim353 scores rather than a per-word fitting-loss distribution.
Editorial extensions
If this is right
- If the central claim holds, a 426k-word vocabulary of PQC word embeddings is available for quantum natural language processing models at inference time, without needing a quantum device during training.
- Complex-valued Skip-gram with the fidelity overlap can be trained as efficiently as classical word2vec on multi-billion-word corpora, since the overhead is a constant factor and gradients are computed explicitly.
- Direct PQC training is not the route to good embeddings; the two-stage fitting procedure is, so future work on quantum embeddings should treat the circuit as a compression of a classically learned complex vector.
- On the additional similarity datasets, the fitted PQC matches the complex embedding scores on three of four datasets and stays within 0.2 points on the fourth, so the result is not specific to WordSim353.
Reading between the lines
- If the per-word fit is truly lossless, the quantum circuits inherit the full information of the complex embeddings, making the PQC set a drop-in replacement for complex vectors in any downstream compositional model, not just similarity benchmarks.
- The closeness of the 64-dimension complex model to the 100-dimension real baseline, and the lack of gain from 64 to 128 complex dimensions, suggests the fidelity geometry rather than raw dimension is what drives performance; a controlled test with matched parameter counts across more dimensions would clarify this.
- A natural extension is to use the same two-stage idea for GloVe-style objectives, where the overlap would predict log co-occurrence counts; success there would broaden the method beyond Skip-gram.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces complex-valued extensions of the Skip-gram word embedding model, replacing real-valued vectors with complex vectors and the inner product with a fidelity-based overlap. The authors describe a PyTorch implementation for small corpora and a modified C implementation of word2vec for a 3.8B-word corpus with a vocabulary of over 400k. They then train parameterized quantum circuits (PQCs) in two ways: directly, with the PQC producing focal or both focal and context embeddings, and through a two-stage procedure that fits a PQC to each already-trained complex embedding. The models are evaluated on WordSim353, MEN, RG-65, and SCWS. The main empirical findings are that the arbitrary complex embeddings are competitive with the classical Skip-gram baseline and that the two-stage fitted PQCs essentially reproduce the evaluation scores of the complex embeddings. The paper argues that this provides a scalable route to producing quantum word embeddings for large vocabularies.
Significance. If the results are reproducible, the two-stage pipeline is a meaningful practical contribution: it avoids corpus-scale PQC training, scales with vocabulary size rather than corpus size, and yields PQC-parameterized states for hundreds of thousands of words that can be used in QNLP inference or compositional models. The paper also provides a useful empirical comparison of circuit ansatze and layer counts for fitting arbitrary complex vectors. The C implementation that enables training on a 3.8B-word corpus is a nontrivial engineering contribution, and the authors are transparent about the limitations of direct PQC training. However, the absence of code or data release and the lack of per-word fitting evidence currently temper the significance of the claims.
major comments (4)
- [Section 3.4] The claim that the fitted PQCs reproduce the arbitrary complex embeddings 'perfectly,' with the loss going to 'effectively zero,' is not backed by per-word evidence. WordSim353 contains 353 word pairs and Spearman correlation is an aggregate rank measure, so identical WordSim353 scores do not rule out a heavy tail of poorly fit words; Table 7 itself shows a 0.2-point gap on SCWS (65.7 vs 65.9), so the fit is not exactly lossless for every word. Please report the distribution of the final fitting loss or fidelity over the full 426,507-word vocabulary (or a random sample), including the fraction of words below a stated fidelity threshold. This is load-bearing for the abstract's claim that the quantum embeddings 'perform as well' for the whole vocabulary.
- [Section 3.1 and Section 3.5] WordSim353 is used both for hyperparameter selection (e.g., the D scaling factor, ansatz, number of layers, learning rate) and as the main evaluation set; Section 3.5 further states that 'WordSim353 is used as a validation set to choose the best-performing model' before reporting the additional datasets. No held-out split or significance tests are provided, and several of the headline differences are small relative to the reported standard deviations (e.g., Table 2: 64.6 vs 63.0 at dimension 64, with SDs of 0.33 and 0.50). Please add paired significance tests (e.g., bootstrap over runs) or a properly separated validation set, and report the scores of all models without selection on the test set.
- [Sections 3.2-3.4] The paper introduces a custom C implementation, a newly created 3.8B-word corpus, and per-word PQC fitting code, but no code, embeddings, or fitted parameters are released. This prevents independent verification of both the large-scale training results and the central per-word fitting claim. At minimum, please release the trained complex embeddings and the fitted PQC parameters for the evaluation vocabulary, along with a script that recomputes the evaluation scores, or provide a clear statement about any restrictions.
- [Abstract and Section 3.4] The two-stage PQC result is partly by construction: once a PQC is fitted to a complex embedding, its evaluation score is, by design, nearly identical to that of the complex embedding. The abstract's wording could be read as an independent demonstration that quantum embeddings match classical ones, which is not what the experiment shows. The paper should explicitly state that the value of the two-stage method is the scalable production of PQC-prepared states that inherit the semantic quality of the complex embeddings, and that the direct PQC-training results (Tables 3, 4, 6) are the appropriate tests of PQC-based learning.
minor comments (7)
- [Section 2.1] The phrase 'to the predict the context word' should be 'to predict the context word'.
- [Table 4 header] 'Anzatze' should be 'Ansätze'.
- [Section 3.4] 'in the worse case' should be 'in the worst case'.
- [Section 3.5] 'WordsSim353' should be 'WordSim353'.
- [Table 6] The WordSim353 score for the 1-way PQC model (64.0) does not match the corresponding dimension-64 entry in Table 3 (63.5); please clarify whether these are the same best-run weights or different selections.
- [Section 2.1 vs Section 3.1] Section 2.1 says 'a value of around D = 3 works well in practice,' but Section 3.1 states D = 3.5 for all experiments using loss (4); please reconcile this or comment on the sensitivity to D.
- [Table 7] The abstract's phrase 'comparable numbers of parameters' should be quantified: a 100-dimensional classical embedding has 100 real parameters per word, a 64-dimensional complex embedding has 128, and a 3-layer A5 PQC on 6 qubits has (log_2^2(64)+3 log_2(64))*3 = 162 parameters per word.
Circularity Check
The two-stage PQC result inherits its WordSim score from the fitted complex vectors by construction, so the claim that quantum embeddings 'perform as well' is not an independent prediction.
-
fitted input called prediction
[Section 3.4, Table 5 discussion; echoed in the Abstract]
"The PQC for each word was 3 layers of A5 and, since the loss effectively went to zero, the performance of the PQCs on WordSim353 was identical to the arbitrary complex embeddings on which the PQCs were trained."
The PQC embeddings are produced by minimizing a fitting loss against the stage-1 complex embeddings; with the loss at 'effectively zero', the PQC's WordSim353 score is identical to the target's score by construction. The abstract's 'quantum word embeddings from the two-stage process perform as well as the classical Skip-gram embeddings' therefore reduces to the stage-1 complex model's already-reported score, not to any independent quality of the PQC representation. The only genuinely new content is that the ansatz can approximate the target vectors, but that is verified by the fitting loss, not by the downstream evaluation.
full rationale
The paper is largely self-contained and its pipeline is clearly described: a complex Skip-gram model is trained on a corpus, and then PQCs are fitted to those complex vectors. The ansatz choice and expressivity discussion rely on external prior work (Sim et al. [SJAG19]), not on a self-citation chain. No uniqueness theorem is imported from the authors, and no known result is merely renamed. The central circularity is limited to the two-stage PQC evaluation: because the fitting loss is reported as effectively zero, the 'quantum embedding' scores on WordSim353, MEN, RG-65, and SCWS are, up to small fitting error, the same as the stage-1 complex embedding scores by construction. The paper is transparent about this, but the abstract's framing that the two-stage quantum embeddings 'perform as well' presents an inherited result as a finding about the PQC model. There is also a mild selection effect in that WordSim353 is used as a validation set for choosing hyperparameters and the best run, but that is an overfitting concern rather than a definitional circularity. Overall, the non-redundant contribution—scalable fitting of PQCs to complex embeddings—does not itself reduce to the evaluation, so the circularity score is moderate rather than severe.
Assumptions & free parameters
free parameters (3)
- D scaling factor =
3.5
- word frequency cutoff =
17 (small), 100 (large)
- PQC layers and ansatz =
3 layers of A5 or A14
assumptions (4)
- domain assumption The distributional hypothesis: words appearing in similar contexts should have similar embeddings.
- domain assumption The fidelity measure |<v|w>|^2, scaled by D, is a suitable objective for learning semantic similarity in complex spaces.
- domain assumption The chosen PQCs (3 layers of A5 or A14) are expressive enough to fit arbitrary complex embedding vectors.
- domain assumption Human similarity and relatedness ratings in WordSim353, MEN, RG-65, and SCWS are a valid evaluation of embedding quality.
Cite this review
Pith. "Pith review of Learning Complex Word Embeddings in Classical and Quantum Spaces." pith.science (2026). https://pith.science/paper/LZ2CS4DJ
@misc{pith2026241213745,
author = {Pith},
title = {Pith review of: Learning Complex Word Embeddings in Classical and Quantum Spaces},
year = {2026},
howpublished = {\url{https://pith.science/paper/LZ2CS4DJ}},
note = {Machine review of arXiv:2412.13745}
}
read the original abstract
We present a variety of methods for training complex-valued word embeddings, based on the classical Skip-gram model, with a straightforward adaptation simply replacing the real-valued vectors with arbitrary vectors of complex numbers. In a more "physically-inspired" approach, the vectors are produced by parameterised quantum circuits (PQCs), which are unitary transformations resulting in normalised vectors which have a probabilistic interpretation. We develop a complex-valued version of the highly optimised C code version of Skip-gram, which allows us to easily produce complex embeddings trained on a 3.8B-word corpus for a vocabulary size of over 400k, for which we are then able to train a separate PQC for each word. We evaluate the complex embeddings on a set of standard similarity and relatedness datasets, for some models obtaining results competitive with the classical baseline. We find that, while training the PQCs directly tends to harm performance, the quantum word embeddings from the two-stage process perform as well as the classical Skip-gram embeddings with comparable numbers of parameters. This enables a highly scalable route to learning embeddings in complex spaces which scales with the size of the vocabulary rather than the size of the training corpus. In summary, we demonstrate how to produce a large set of high-quality word embeddings for use in complex-valued and quantum-inspired NLP models, and for exploring potential advantage in quantum NLP models.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
A study on similarity and relatedness using distributional and W ord N et-based approaches
Eneko Agirre, Enrique Alfonseca, Keith Hall, Jana Kravalova, Marius Pa s ca, and Aitor Soroa. A study on similarity and relatedness using distributional and W ord N et-based approaches. In Mari Ostendorf, Michael Collins, Shri Narayanan, Douglas W. Oard, and Lucy Vanderwende, editors, Proceedings of Human Language Technologies: The 2009 Annual Conference ...
work page 2009
-
[2]
Quantum Computing since Democritus
Scott Aaronson. Quantum Computing since Democritus . Cambridge University Press, Cambridge, 2013
work page 2013
-
[3]
Huggins, Ramis Movassagh, Dar Gilboa, and Jarrod McClean
Amira Abbas, Robbie King, Hsin-Yuan Huang, William J. Huggins, Ramis Movassagh, Dar Gilboa, and Jarrod McClean. On quantum backpropagation, information reuse, and cheating measurement collapse. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems . Curran Associates, Inc., 2023
work page 2023
-
[4]
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio. Unitary evolution recurrent neural networks. In Proceedings of the International Conference on Machine Learning , page 1120–1128, New York, USA, 2016
work page 2016
-
[5]
Recurrent quantum neural networks
Johannes Bausch. Recurrent quantum neural networks. In Advances in Neural Information Processing Systems , volume 33, pages 1368--1379. Curran Associates, Inc., 2020
work page 2020
-
[6]
D. Bankova, B. Coecke, M. Lewis, and D. Marsden. Graded hyponymy for compositional distributional semantics. Journal of Language Modelling , 6(2):225–260, 2019
work page 2019
-
[7]
Harry Buhrman, Richard Cleve, John Watrous, and Ronald de Wolf. Quantum fingerprinting. Phys. Rev. Lett. , 87:167902, Sep 2001
work page 2001
-
[8]
A quantum-theoretic approach to distributional semantics
William Blacoe, Elham Kashefi, and Mirella Lapata. A quantum-theoretic approach to distributional semantics. In Lucy Vanderwende, Hal Daum \'e III, and Katrin Kirchhoff, editors, Proceedings of the 2013 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 847--857, Atlanta, Georgia...
work page 2013
Show all 50 references
-
[9]
Parameterized quantum circuits as machine learning models
Marcello Benedetti, Erika Lloyd, Stefan Sack, and Mattia Fiorentini. Parameterized quantum circuits as machine learning models. Quantum Science and Technology , 4(4), 2019
2019
-
[10]
Vector space models of lexical meaning
Stephen Clark. Vector space models of lexical meaning. In The Handbook of Contemporary Semantic Theory , chapter 16, pages 493--522. Wiley Blackwell, 2nd edition, 2015
2015
-
[11]
Mathematical foundations for a compositional distributional model of meaning
Bob Coecke, Mehrnoosh Sadrzadeh, and Stephen Clark. Mathematical foundations for a compositional distributional model of meaning. Linguistic Analysis , 36(1-4):345--384, 2010
2010
-
[12]
Scalable and interpretable quantum natural language processing: an implementation on trapped ions
Tiffany Duneau, Saskia Bruhn, Gabriel Matos, Tuomas Laakkonen, Katerina Saiti, Anna Pearson, Konstantinos Meichanetzidis, and Bob Coecke. Scalable and interpretable quantum natural language processing: an implementation on trapped ions. https://arxiv.org/abs/2409.08777, 2024
2024 arXiv
-
[13]
M. P. da Silva, C. Ryan-Anderson, J. M. Bello-Rivas, A. Chernoguzov, J. M. Dreiling, C. Foltz, F. Frachon, J. P. Gaebler, T. M. Gatterman, L. Grans-Samuelsson, D. Hayes, N. Hewitt, J. Johansen, D. Lucchetti, M. Mills, S. A. Moses, B. Neyenhuis, A. Paz, J. Pino, P. Siegfried, J...
2024 arXiv
-
[14]
J. R. Firth. A synopsis of linguistic theory 1930-1955. In Studies in Linguistic Analysis , pages 1--32. Oxford: Philological Society, 1957
1930
-
[15]
The Geometry of Meaning
Peter G \"a rdenfors. The Geometry of Meaning . The MIT Press, 2014
2014
-
[16]
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvarinen. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. Proceedings of Machine Learning Research , 9:297--304, 2010
2010
-
[17]
Georgiou and C
G.M. Georgiou and C. Koutsougeras. Complex domain backpropagation. IEEE Transactions on Circuits and Systems II: Analog and Digital Signal Processing , 39(5):330--334, 1992
1992
-
[18]
Quantum linear algebra is all you need for transformer architectures
Naixu Guo, Zhan Yu, Matthew Choi, Aman Agrawal, Kouhei Nakaji, Alan Aspuru-Guzika, and Patrick Rebentrost. Quantum linear algebra is all you need for transformer architectures. arXiv preprint arXiv:2402.16714 , 2024
2024
-
[19]
Z. Harris. Distributional structure. Word , 10:146--162, 1954
1954
-
[20]
Birgitta Whaley, and E
William Huggins, Piyush Patil, Bradley Mitchell, K. Birgitta Whaley, and E. Miles Stoudenmire. Towards quantum machine learning with tensor networks. Quantum Science and Technology , 4, 2019
2019
-
[21]
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. Long Short-Term Memory . Neural Computation , 9(8):1735--1780, 11 1997
1997
-
[22]
Improving word representations via global context and multiple word prototypes
Eric Huang, Richard Socher, Christopher Manning, and Andrew Ng. Improving word representations via global context and multiple word prototypes. In Haizhou Li, Chin-Yew Lin, Miles Osborne, Gary Geunbae Lee, and Jong C. Park, editors, Proceedings of the 50th Annual Meeting of th...
2012
-
[23]
Sequence processing with quantum tensor networks
Carys Harvey, Richie Yeung, and Konstantinos Meichanetzidis. Sequence processing with quantum tensor networks. https://arxiv.org/abs/2308.07865, 2023
2023 arXiv
-
[24]
Quixer: A quantum transformer model
Nikhil Khatri, Gabriel Matos, Luuk Coopmans, and Stephen Clark. Quixer: A quantum transformer model. arXiv preprint arXiv:2406.04305 , 2024
2024 arXiv
-
[25]
Towards logical negation for compositional distributional semantics
Martha Lewis. Towards logical negation for compositional distributional semantics. IfCoLoG Journal of Logics and their Applications , 7, 2020
2020
-
[26]
QNLP in practice: Running compositional models of meaning on a quantum computer
Robin Lorenz, Anna Pearson, Konstantinos Meichanetzidis, Dimitri Kartsaklis, and Bob Coecke. QNLP in practice: Running compositional models of meaning on a quantum computer. Journal of Artificial Intelligence Research , 76, 2023
2023
-
[27]
Coles, Lukasz Cincio, Jarrod R
Martin Larocca, Supanut Thanasilp, Samson Wang, Kunal Sharma, Jacob Biamonte, Patrick J. Coles, Lukasz Cincio, Jarrod R. McClean, Zoe Holmes, and M. Cerezo. A review of barren plateaus in variational quantum computing. https://arxiv.org/abs/2405.00781, 2024
2024 arXiv
-
[28]
Quantum-inspired complex word embedding
Qiuchi Li, Sagar Uprety, Benyou Wang, and Dawei Song. Quantum-inspired complex word embedding. In Isabelle Augenstein, Kris Cao, He He, Felix Hill, Spandana Gella, Jamie Kiros, Hongyuan Mei, and Dipendra Misra, editors, Proceedings of the Third Workshop on Representation Learn...
2018
-
[29]
S. A. Moses, C. H. Baldwin, M. S. Allman, R. Ancona, L. Ascarrunz, C. Barnes, J. Bartolotta, B. Bjork, P. Blanchard, M. Bohn, J. G. Bohnet, N. C. Brown, N. Q. Burdick, W. C. Burton, S. L. Campbell, J. P. Campora III au2, C. Carron, J. Chambers, J. W. Chan, Y. H. Chen, A. Chern...
2023
-
[30]
McClean, Sergio Boixo, Vadim N
Jarrod R. McClean, Sergio Boixo, Vadim N. Smelyanskiy, Ryan Babbush, and Hartmut Neven. Barren plateaus in quantum neural network training landscapes. Nature Communications , 9(1), 2018
2018
-
[31]
Efficient estimation of word representations in vector space
Tom \' a s Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. In 1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings , 2013
2013
-
[32]
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In C.J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K.Q. Weinberger, editors, Advances in Neural Information Processing Sy...
2013
-
[33]
Nielsen and Isaac L
Michael A. Nielsen and Isaac L. Chuang. Quantum Computation and Quantum Information . Cambridge University Press, 2000
2000
-
[34]
Poincare embeddings for learning hierarchical representations
Maximillian Nickel and Douwe Kiela. Poincare embeddings for learning hierarchical representations. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems , volume 30. Curran Ass...
2017
-
[35]
Peruzzo, J
A. Peruzzo, J. McClean, P. Shadbolt, Man-Hong Yung, Xiao-Qi Zhou, Peter J. Love, Alan Aspuru-Guzik, and Jeremy L. O'Brien. A variational eigenvalue solver on a photonic quantum processor. Nat. Commun. , 5, 2014
2014
-
[36]
Quantum C omputing in the NISQ era and beyond
John Preskill. Quantum C omputing in the NISQ era and beyond. Quantum , 2:79, August 2018
2018
-
[37]
G lo V e: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. G lo V e: Global vectors for word representation. In Alessandro Moschitti, Bo Pang, and Walter Daelemans, editors, Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pa...
2014
-
[38]
Conversational negation using worldly context in compositional distributional semantics
Benjamin Rodatz, Razin Shaikh, and Lia Yeh. Conversational negation using worldly context in compositional distributional semantics. In Martha Lewis and Mehrnoosh Sadrzadeh, editors, Proceedings of the 2021 Workshop on Semantic Spaces at the Intersection of NLP, Physics, and C...
2021
-
[39]
Johnson, and Alan Aspuru-Guzik
Sukin Sim, Peter D. Johnson, and Alan Aspuru-Guzik. Expressibility and entangling capability of parameterized quantum circuits for hybrid quantum-classical algorithms. Advanced Quantum Technologies , 2(12), 2019
2019
-
[40]
Modeling term dependencies with quantum language models for IR
Alessandro Sordoni, Jian-Yun Nie, and Yoshua Bengio. Modeling term dependencies with quantum language models for IR . In Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval , pages 653--662, Dublin, Ireland, 2013. Ass...
2013
-
[41]
Supervised learning with tensor networks
Edwin Stoudenmire and David J Schwab. Supervised learning with tensor networks. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems , volume 29. Curran Associates, Inc., 2016
2016
-
[42]
Chiheb Trabelsi, Olexa Bilaniuk, Ying Zhang, Dmitriy Serdyuk, Sandeep Subramanian, Joao Felipe Santos, Soroush Mehri, Negar Rostamzadeh, Yoshua Bengio, and Christopher J. Pal. Deep complex networks. In Sixth International Conference on Learning Representations ( ICLR ) , Vanco...
2018
-
[43]
Shaikh, Sara Sabrina Zemljic, and Stephen Clark
Sean Tull, Razin A. Shaikh, Sara Sabrina Zemljic, and Stephen Clark. From conceptual spaces to quantum concepts: Formalising and learning structured conceptual models. Quantum Machine Intelligence , 6, 2024
2024
-
[44]
The Geometry of Information Retrieval
Keith van Rijsbergen. The Geometry of Information Retrieval . Cambridge University Press, 2004
2004
-
[45]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems , pages 5998--6008, 2017
2017
-
[46]
Quantum natural language processing
Dominic Widdows, Willie Aboumrad, Dohun Kim, Sayonee Ray, and Jonathan Mei. Quantum natural language processing. https://arxiv.org/abs/2403.19758, 2024
2024 arXiv
-
[47]
Orthogonal negation in vector spaces for modelling word-meanings and document retrieval
Dominic Widdows. Orthogonal negation in vector spaces for modelling word-meanings and document retrieval. In Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics , pages 136--143, Sapporo, Japan, July 2003. Association for Computational Linguistics
2003
-
[48]
Natural language processing meets quantum physics: A survey and categorization
Sixuan Wu, Jian Li, Peng Zhang, and Yue Zhang. Natural language processing meets quantum physics: A survey and categorization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages 3172--3182, Online and Punta Cana, Dominican Republi...
2021
-
[49]
Quantum recurrent architectures for text classification
Wenduan Xu, Stephen Clark, Douglas Brown, Gabriel Matos, and Konstantinos Meichanetzidis. Quantum recurrent architectures for text classification. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural ...
2024
-
[50]
Quantum state preparation with optimal circuit depth: Implementations and applications
Xiao-Ming Zhang, Tongyang Li, and Xiao Yuan. Quantum state preparation with optimal circuit depth: Implementations and applications. Phys. Rev. Lett. , 129:230504, Nov 2022
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.