Pith. sign in

REVIEW 4 major objections 5 minor 82 references

Quantum AIXI: Universal Intelligence via Quantum Information

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Quantum AIXI is a universal Bayesian agent defined over quantum environments, reducing to classical AIXI when observables commute.

desk verdict A promising but formally incomplete quantum generalization of AIXI; the mixture's normalization is the load-bearing gap. read the letter →

arxiv 2505.21170 v2 pith:7OOSUMJO submitted 2025-05-27 quant-ph cs.AI

classification quant-phcs.AI MSC 68Q3081P6868T05
keywords QuantumAIXIKolmogorovcomplexityuniversalinductionreinforcementlearningcontextualityno-cloningtheoreminformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Quantum AIXI (QAIXI) is proposed as the quantum analogue of AIXI, the classical ideal Bayesian agent for universal intelligent decision-making. The paper's central claim is that replacing the classical mixture over computable environments with an operator-valued mixture weighted by quantum Kolmogorov complexity, and replacing the agent–environment loop with quantum registers, channels, and measurement instruments, yields a coherent model of a universal agent in a quantum universe. If the framework holds together, classical AIXI appears as the commuting special case, and the remaining obstructions—measurement back-action, contextuality, no-cloning, and uncomputability—are not implementation details but intrinsic features of optimal quantum agency. The work therefore matters because it turns a question about the physical basis of intelligence into a concrete mathematical structure that can be examined, modified, and tested.

What carries the argument

The load-bearing object is the operator-valued universal mixture $\Xi_Q(a_{1:m})=\sum_{Q\in\mathcal{Q}_{\rm sol}}2^{-K_Q(Q)}\,\rho_Q^E(a_{1:m})$, a semi-density operator (a positive operator with trace at most $1$) that takes the role of the classical universal prior inside the decision process. It is built from two substitutions: classical environments become chronological semi-computable quantum channels with complexity $K_Q(Q)$, and probability measures become environment states $\rho_Q^E(a_{1:m})$ generated by those channels. The QAIXI policy is the argmax of the mixture-averaged discounted value functional $V^\pi_{\Xi_Q}$, with value-function updates that include the post-measurement state, which is what makes measurement back-action part of the optimisation. Projection onto a POVM recovers ordinary probabilities, and the commuting limit returns the classical prior and policy.

What would settle it

Exhibit two universal quantum Turing machines $U_1,U_2$ and an infinite family of channels $Q$ for which $|K_Q^{U_1}(Q)-K_Q^{U_2}(Q)|$ grows without bound; that would falsify the invariance assumption that carries the quantum universal prior.

Watch

Extended reading notes

Core claim

On the paper's own terms, QAIXI is an agent whose private state is a quantum register $\mathcal{H}_A$ interacting with an environment register $\mathcal{H}_E$; an action is either a unitary channel (coherent control) or a quantum instrument (measurement), and percepts are classical outcomes plus rewards computed from the instrument. The environment class $\mathcal{Q}_{\rm sol}$ consists of chronological, semi-computable quantum channels identified by a purification vector of their channel map. Quantum Kolmogorov complexity $K_Q(Q)$ is the length of the shortest classical program that makes a universal quantum Turing machine reproduce the channel to within $\varepsilon$; the universal prior is the semi-density operator $\Xi_Q(a_{1:m}) = \sum_Q 2^{-K_Q(Q)}\,\rho_Q^E(a_{1:m})$, whose trace is at most one. Projecting $\Xi_Q$ onto the agent's measurement POVM returns a scalar probability; when every environment state is diagonal and the agent's instruments are diagonal in the same basis, $K_Q(Q)$ agrees with classical Kolmogorov complexity up to an additive constant, $\Xi_Q$ reduces to the classical universal prior, and the optimal policy reduces to the classical AIXI policy. The paper also states a convergence theorem for this quantum induction under ergodicity, informational completeness, and a finite complexity gap, and proves that no universal predictor built from commuting projectors can reproduce the statistics of a contextually uncolourable environment, so a QAIXI history must record the whole future instrument schedule.

Load-bearing premise

The framework depends on quantum Kolmogorov complexity being invariant across universal quantum Turing machines up to an additive constant; without that, the prior weights are machine-dependent and the mixture is not universal.

Editorial extensions

If this is right

  • In the commuting limit the quantum formulation reproduces the classical AIXI agent, making AIXI a special case rather than a competing model.
  • QAIXI agents can choose coherent unitary actions or measurement instruments, so the agent's decisions directly control how much quantum information is preserved versus decohered.
  • Under the stated conditions, the posterior mixture is claimed to converge to the true environment state at $O(t^{-1/2})$ in trace distance, giving an inductive basis for quantum agency.
  • Contextuality rules out any predictor that assigns fixed outcomes to all measurements, so the posterior cannot be a function of a classical history string alone.
  • No-cloning makes each observation single-use, so the sample complexity of quantum learning is tied to state-preparation resources rather than to replay of past data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the invariance assumption for $K_Q$ is the first place to probe; if it fails, the prior's weights are machine-dependent and the universality claim needs repair.
  • Editorial inference: classical shadow tomography is a natural candidate to approximate $\Xi_Q$ without full state tomography, but the paper's convergence conditions would need to be re-derived for shadow-based updates.
  • Editorial inference: the boson-sampling-style argument implies an empirical litmus test: if a classical agent can estimate QAIXI's value in such an environment, then the cited hardness assumptions would force a polynomial-hierarchy collapse.
  • Editorial inference: the requirement to record the future instrument schedule suggests that any practical QAIXI should optimise over measurement policies, not just action sequences, which could change how optimality is defined.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Quantum AIXI (QAIXI), a quantum generalization of Hutter's classical AIXI agent. It models the agent and environment as quantum registers interacting through unitary channels and quantum instruments, defines a quantum Kolmogorov complexity K_Q for quantum environments, introduces an operator-valued universal mixture Ξ_Q over semi-computable quantum environments, and formulates a quantum Bellman equation and a QAIXI policy. It claims that in the commuting limit K_Q reduces to classical Kolmogorov complexity and the mixture reduces to the classical Solomonoff prior, so that QAIXI generalizes AIXI. The paper also states a convergence theorem for quantum Solomonoff induction (QSI), discusses its limitations including measurement back-action, contextuality, Bell non-locality and no-cloning, and gives a Kochen-Specker-based corollary about the impossibility of a non-contextual predictive quantum Turing machine. Appendices cover entanglement in the interaction loop, quantum advantage for specific environment classes, and a detailed discussion of the QSI proof sketch and its unresolved steps.

Significance. If the construction can be made rigorous, QAIXI would be a valuable conceptual bridge between algorithmic information theory and quantum AGI, and the commuting-limit reduction to classical AIXI is a genuine consistency check. The paper is unusually honest about the status of its claims: it explicitly states that proving QSI convergence is an open question, and Appendix C.4 lists the martingale, semi-density-operator, and contextuality hurdles that a full proof would need to overcome. The paper also makes falsifiable structural claims, such as the recovering of the classical Solomonoff prior in the commuting limit and the obstruction posed by Kochen-Specker contextuality. However, the central construction currently rests on a load-bearing normalization gap in the definition of Ξ_Q and on an unproven convergence theorem, so the significance is contingent on substantial further mathematical work.

major comments (4)
  1. [Sec. 3.3, Eqs. (4)-(5), and Sec. 4, Eq. (9)] The universal mixture Ξ_Q(a1:m) = Σ_Q 2^{-K_Q(Q)} ρ_Q^E(a1:m) uses the plain (non-prefix-free) quantum Kolmogorov complexity K_Q defined by a minimum over all binary strings p in Eq. (4). For plain complexity the Kraft inequality does not apply, and classically Σ_x 2^{-K(x)} diverges. Since the commuting limit of Qsol contains the classical semi-computable environments, the same divergence afflicts the quantum mixture unless K_Q is replaced by a prefix (self-delimiting) complexity or the sum is otherwise shown to converge. The assertion 0 < Tr(Ξ_Q(a1:m)) ≤ 1 is therefore unsupported, and the posterior update in Eq. (10), the value functional in Eqs. (7)-(8), and the claimed reduction to the classical Solomonoff prior in Sec. 3.3 are not well-defined. Appendix C's appeal to 'prefix-free codes with binary word-lengths' does not repair Eq. (4), because that equation imposes no self-delimiting constraint on p. The paper needs either a definition of a prefix-free quantum Kolmogorov complexity, a restricted enumerable code for Qsol with a proven convergence bound, or a proof that Eq. (5) converges despite using plain complexity.
  2. [Sec. 4.1, Theorem 1 and Eq. (12)] Theorem 1 is stated as a result, but the text immediately says 'Proving QSI convergence is an open question.' The proof sketch relies on the E_{Q*}-martingale property of the divergence difference D(ρ*_E || Ξ_Q) - D(Λ*_k || Λ^Ξ_k), on monotonicity of Umegaki relative entropy under branch maps, and on applying the chain rule to semi-density operators; Appendix C.4 explicitly lists all of these as requiring rigorous justification. Additionally, the initial-divergence bound D_0 ≤ K_Q(Q*) ln 2 + ln(1+g) depends on the normalization of Ξ_Q, which is not established because of the plain-complexity issue raised above. As it stands, Eq. (12) is a conjecture, not a theorem. Please either provide a complete proof, including the normalization of the prior and the martingale convergence for operator-valued mixtures, or reclassify the statement as a conjecture/open problem and remove the theorem label.
  3. [Sec. 4.2, Corollary 1 and Eq. (14)] The proof of Corollary 1 assumes an action string a†_1:m that instructs the agent to measure, at the final cycle, every projector in a Kochen-Specker set. But the projectors in a KS set are generally noncommuting and cannot be measured simultaneously, so no single quantum instrument can produce outcomes for all of them. The corollary's hypothesis also says the QTM outputs a commuting family {Q_{a1:m}(e1:m)}, and such a commuting family cannot reproduce the Born probabilities of noncommuting observables that Eq. (14) would require. Consequently, the claimed contradiction with KS uncolourability is not established as stated. The corollary needs to be reformulated in a way that respects the noncommutativity of the projectors, for example by explicitly considering separate instruments for each context, rather than a single final-cycle measurement of all projectors.
  4. [Sec. 3.3, paragraph after Eq. (4)] The assumption that quantum Kolmogorov complexity is invariant up to an additive constant for any two universal QTMs is asserted without proof, and the paper itself notes that this is 'technically subtler to prove than in the classical case.' Since the universality of the mixture Ξ_Q, the convergence bound, and the claimed equivalence with classical Kolmogorov complexity in the commuting limit all inherit this assumption, it should be either proved with a precise invariance theorem or explicitly stated as a conjecture with the necessary qualifications.
minor comments (5)
  1. [Sec. 4.1] The text says 'by the quantum Pinsker inequality [17,32]', but reference [17] is the classical Csiszár-Körner book; the quantum Pinsker inequality is normally attributed to Lindblad or to the quantum information literature, so the citation should be corrected or clarified.
  2. [Appendix C, Eqs. (25)-(26)] The chain rule for quantum relative entropy is applied to the semi-density operator Ξ_Q without an explicit statement of where the normalization constant Z is inserted. Appendix C.4 already acknowledges this issue, but the main derivation should make the normalization explicit at each step to avoid confusion.
  3. [Appendix D.1] The text refers to 'Theorem 1 provides that no quantum Turing machine Q can output...' but the result being described is Corollary 1 in Section 4.2. The cross-reference is inconsistent and should be corrected.
  4. [Sec. 4.2, Bell non-locality paragraph] The statement that Bell violations 'break the martingale structure assumed in Theorem 1' is conceptually reasonable, but since Theorem 1 is itself unproven, this should be phrased as a limitation of the proposed proof strategy rather than as a failure of an established theorem.
  5. [Appendix C, text near Eq. (3)] A displayed equation is numbered '(3)' in the middle of the derivation, while the main-text equations are numbered differently; this numbering should be aligned with the rest of the manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: QAIXI is a definitional framework, and its commuting-limit reduction to classical AIXI is a consistency check rather than a derived prediction.

full rationale

The paper builds QAIXI by definition: Eq. (5)/(9) define the operator-valued mixture xi_Q, Eq. (7) defines the mixture value functional, and Eq. (8) defines the QAIXI policy as the argmax of that functional. These are stipulative constructions, not predictions, so there is no fitted-input-called-prediction pattern. The commuting-limit statements in Sec. 3.3 ('K_Q(Q) coincides up to a constant with K(nu), xi_Q reduces to the classical Solomonoff prior xi_U and the AIXI policy') are consistency checks of the definitions: when the environment states are diagonal and measurements commute, the formal expressions reduce by construction to the classical ones. That is a special-case check, not a circular derivation of the quantum framework from classical AIXI. Theorem 1 is explicitly presented as a sketch ('Proving QSI convergence is an open question. One potential avenue is as follows.'), with (C1)-(C3) stated as assumptions; the bound K_Q(Q*) ln 2 + ln(1+g) follows algebraically from the definitions omega_Q = 2^{-K_Q(Q)} and g = sum_{Q!=Q*} 2^{-(K_Q(Q)-K_Q(Q*))}, so it does not assume the theorem's conclusion. The paper's self-references ([53] Quantum Geometric Machine Learning and [54] the author's Zenodo appendices) are contextual or duplicate material already present in the arXiv version; neither is load-bearing for the central construction. The invariant-complexity assumption after Eq. (4) is explicitly labeled an assumption rather than imported from prior work. The only serious defect visible in the derivation chain is the normalization claim 0 < Tr(xi_Q) <= 1 in Eq. (9): Eq. (4) uses plain (non-prefix-free) program lengths, while Appendix C justifies sub-normalization by appealing to 'prefix-free codes with binary word-lengths (see [28])'. This is a mathematical gap in the construction, but it is not a circularity: no equation is being reused as its own premise, and no result is forced by a self-citation. Accordingly the circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper is a definitional framework with no fitted constants. The only free parameters are technical tolerances and conditions used in the definitions and the convergence sketch; none are calibrated to data. No new physical entities are postulated.

free parameters (3)
  • ε (KQ approximation tolerance) = not specified (<1)
    Chosen by hand in Eq. (4) to define quantum Kolmogorov complexity; affects the constant in convergence bounds but not the qualitative rate.
  • δ (ergodicity constant in C1) = not specified (>0)
    Condition (C1) for Theorem 1; the convergence bound depends on the assumption that time-averaged state variation is bounded.
  • ϵ (informational completeness error in C2) = not specified (>0)
    Condition (C2) for Theorem 1; the paper notes an extra O(ϵ) term appears in the Pinsker bound as ϵ tends to 0.
assumptions (4)
  • domain assumption Invariance of quantum Kolmogorov complexity up to an additive constant across universal QTMs
    Assumed after Eq. (4); needed for the universality of the quantum Solomonoff prior and for the convergence bound. The paper admits this is 'technically subtler' in the quantum setting.
  • domain assumption Well-definedness of the class Qsol and convergence of the mixture sum with Tr Ξ_Q ≤ 1
    Used in Eq. (5) and Eq. (9); without this the prior is not a semi-density operator.
  • ad hoc to paper Conditions C1-C3 (ergodicity, informational completeness, finite complexity gap) for Theorem 1
    Imposed in Section 4.1 to make the convergence sketch go through; the paper does not prove they hold for generic quantum environments.
  • domain assumption Standard (Copenhagen-like) measurement formalism
    The framework assumes projective or instrument-based measurement with back-action; the paper notes alternative interpretations would change the model (Section 5).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Quantum AIXI: Universal Intelligence via Quantum Information." pith.science (2026). https://pith.science/paper/7OOSUMJO

@misc{pith2026250521170,
  author       = {Pith},
  title        = {Pith review of: Quantum AIXI: Universal Intelligence via Quantum Information},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7OOSUMJO}},
  note         = {Machine review of arXiv:2505.21170}
}
read the original abstract

AIXI is a widely studied model of artificial general intelligence (AGI) based upon principles of induction and reinforcement learning. However, AIXI is fundamentally classical in nature - as are the environments in which it is modelled. Given the universe is quantum mechanical in nature and the exponential overhead required to simulate quantum mechanical systems classically, the question arises as to whether there are quantum mechanical analogues of AIXI. To address this question, we extend the framework to quantum information and present Quantum AIXI (QAIXI). We introduce a model of quantum agent/environment interaction based upon quantum and classical registers and channels, showing how quantum AIXI agents may take both classical and quantum actions. We formulate the key components of AIXI in quantum information terms, extending previous research on quantum Kolmogorov complexity and a QAIXI value function. We discuss conditions and limitations upon quantum Solomonoff induction and show how contextuality fundamentally affects QAIXI models.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

82 extracted references · 70 canonical work pages

  1. [1]

    Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences463(2088), 3089–3114 (2007)

    Aaronson, S.: The learnability of quantum states. Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences463(2088), 3089–3114 (2007)

  2. [2]

    Cambridge University Press (2013)

    Aaronson, S.: Quantum computing since Democritus. Cambridge University Press (2013)

  3. [3]

    Nature Physics p

    Aaronson, S.: Quantum machine learning algorithms: Read the fine print. Nature Physics p. 5 (2014)

  4. [4]

    In: Proceedings of the 50th annual ACM SIGACT symposium on theory of computing

    Aaronson, S.: Shadow tomography of quantum states. In: Proceedings of the 50th annual ACM SIGACT symposium on theory of computing. pp. 325–338 (2018)

  5. [5]

    In: Proceedings of the forty-third annual ACM symposium on Theory of computing

    Aaronson, S., Arkhipov, A.: The computational complexity of linear optics. In: Proceedings of the forty-third annual ACM symposium on Theory of computing. pp. 333–342 (2011)

  6. [6]

    Cambridge University Press, Cambridge, 2nd edn

    Bell, J.: Speakable and Unspeakable in Quantum Mechanics. Cambridge University Press, Cambridge, 2nd edn. (2004)

  7. [7]

    In: Artificial General Intelligence

    Bennett, M.T.: The optimal choice of hypothesis is the weakest, not the shortest. In: Artificial General Intelligence. Springer Nature (2023)

  8. [8]

    Bennett, M.T.: Technical appendices (2024).https://doi.org/10.5281/zenodo.7641741,https:// github.com/ViscousLemming/Technical-Appendices

Show all 82 references
  1. [9]

    In: Artificial General Intelligence

    Bennett, M.T., Maruyama, Y.: The artificial scientist: Logicist, emergentist, and universalist approaches to artificial general intelligence. In: Artificial General Intelligence. Springer (2022)

  2. [10]

    Journal of Computer and System Sciences63(2), 201–221 (2001)

    Berthiaume, A., Van Dam, W., Laplante, S.: Quantum kolmogorov complexity. Journal of Computer and System Sciences63(2), 201–221 (2001)

  3. [11]

    Nature549(7671), 195–202 (2017)

    Biamonte, J., Wittek, P., Pancotti, N., Rebentrost, P., Wiebe, N., Lloyd, S.: Quantum machine learning. Nature549(7671), 195–202 (2017)

  4. [12]

    i and ii

    Bohm, D.: A suggested interpretation of the quantum theory in terms of "hidden" variables. i and ii. Physical Review85(2), 166–193 (1952)

  5. [13]

    Quantum6, 882 (2022)

    Bostanci, J., Watrous, J.: Quantum game theory and the complexity of approximating quantum nash equilibria. Quantum6, 882 (2022)

  6. [14]

    Nature Physics10(4), 259–263 (2014)

    Brukner, Č.: Quantum causality. Nature Physics10(4), 259–263 (2014)

  7. [15]

    Catt, E., Hutter, M.: A gentle introduction to quantum computing algorithms with applications to universal prediction (2020)

  8. [16]

    Nature Computational Science (2022)

    Cerezo, M., Verdon, G., Huang, H.Y., Cincio, L., Coles, P.J.: Challenges and opportunities in quantum machine learning. Nature Computational Science (2022)

  9. [17]

    Cambridge University Press (2011)

    Csiszár, I., Körner, J.: Information theory: coding theorems for discrete memoryless systems. Cambridge University Press (2011)

  10. [18]

    Pro- ceedings of the Royal Society of London

    Deutsch, D.: Quantum theory, the church-turing principle and the universal quantum computer. Pro- ceedings of the Royal Society of London. A. Mathematical and Physical Sciences400(1818), 97–117 (1985)

  11. [19]

    Proceedings of the Royal Society of London

    Deutsch, D.: Quantum theory of probability and decisions. Proceedings of the Royal Society of London. Series A: Mathematical, Physical and Engineering Sciences455(1988), 3129–3137 (1999)

  12. [20]

    IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)38(5), 1207–1220 (2008)

    Dong, D., Chen, C., Chen, H., Tarn, T.J.: Quantum reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)38(5), 1207–1220 (2008)

  13. [21]

    Reports on Progress in Physics81(7), 074001 (2018)

    Dunjko, V., Briegel, H.J.: Machine learning & artificial intelligence in the quantum domain: a review of recent progress. Reports on Progress in Physics81(7), 074001 (2018)

  14. [22]

    Physical Review E111(1), 014118 (2025)

    Ebtekar, A., Hutter, M.: Foundations of algorithmic thermodynamics. Physical Review E111(1), 014118 (2025)

  15. [23]

    In: Interna- tional Conference on Artificial General Intelligence

    Fallenstein, B., Soares, N., Taylor, J.: Reflective variants of solomonoff induction and aixi. In: Interna- tional Conference on Artificial General Intelligence. pp. 60–69. Springer (2015)

  16. [24]

    Physical review letters124(10), 100501 (2020)

    Fang, K., Fawzi, O., Renner, R., Sutter, D.: Chain rule for the quantum relative entropy. Physical review letters124(10), 100501 (2020)

  17. [25]

    American Journal of Physics82(8), 749–754 (2014) 12 E

    Fuchs, C.A., Mermin, N.D., Schack, R.: An introduction to qbism with an application to the locality of quantum mechanics. American Journal of Physics82(8), 749–754 (2014) 12 E. Perrier

  18. [26]

    Goertzel, B.: The general theory of general intelligence: A pragmatic patternist perspective. Tech. rep., Singularity Net (2021)

  19. [27]

    Goertzel, B., et al.: Opencog hyperon: A framework for agi at the human level and beyond. Tech. rep., OpenCog (2023)

  20. [28]

    Handbook of the Philosophy of Information pp

    Grünwald, P.D., Vitányi, P., et al.: Algorithmic information theory. Handbook of the Philosophy of Information pp. 281–320 (2008)

  21. [29]

    Decision Support Systems46(1), 318–332 (2008)

    Guo, H., Zhang, J., Koehler, G.J.: A survey of quantum games. Decision Support Systems46(1), 318–332 (2008)

  22. [30]

    Gutoski,G.,Watrous,J.:Towardageneraltheoryofquantumgames.In:Proceedingsofthethirty-ninth annual ACM symposium on Theory of computing. pp. 565–574 (2007)

  23. [31]

    Quantum Studies: Mathematics and Foundations8(3), 351–373 (2021)

    Haapasalo, E.: The choi–jamiołkowski isomorphism and covariant quantum channels. Quantum Studies: Mathematics and Foundations8(3), 351–373 (2021)

  24. [32]

    arXiv preprint arXiv:2005.04553 (2020)

    Hirota, O.: Application of quantum pinsker inequality to quantum communications. arXiv preprint arXiv:2005.04553 (2020)

  25. [33]

    California Institute of Technology (2024)

    Huang, H.Y.: Learning in the Quantum Universe. California Institute of Technology (2024)

  26. [34]

    Science376(6598), 1182–1186 (2022)

    Huang, H.Y., Broughton, M., Cotler, J., Chen, S., Li, J., Mohseni, M., Neven, H., Babbush, R., Kueng, R., Preskill, J., et al.: Quantum advantage in learning from experiments. Science376(6598), 1182–1186 (2022)

  27. [35]

    Nature Physics16(10), 1050–1057 (2020)

    Huang, H.Y., Kueng, R., Preskill, J.: Predicting many properties of a quantum system from very few measurements. Nature Physics16(10), 1050–1057 (2020)

  28. [36]

    Springer Science & Business Media (2004)

    Hutter, M.: Universal artificial intelligence: Sequential decisions based on algorithmic probability. Springer Science & Business Media (2004)

  29. [37]

    Hutter, M.: Universal Algorithmic Intelligence: A Mathematical Top→Down Approach, pp. 227–290. Springer Berlin Heidelberg, Berlin, Heidelberg (2007)

  30. [38]

    Springer, Heidelberg (2010)

    Hutter, M.: Universal Artificial Intelligence: Sequential Decisions Based on Algorithmic Probability. Springer, Heidelberg (2010)

  31. [39]

    Advances in Neural Information Processing Systems34, 28362–28375 (2021)

    Jerbi, S., Gyurik, C., Marshall, S., Briegel, H., Dunjko, V.: Parametrized quantum policies for rein- forcement learning. Advances in Neural Information Processing Systems34, 28362–28375 (2021)

  32. [40]

    arXiv preprint quant-ph/9508006 (1995)

    Knill, E.: Approximation by quantum circuits. arXiv preprint quant-ph/9508006 (1995)

  33. [41]

    Journal of Mathe- matics and Mechanics17(1), 59–87 (1967)

    Kochen, S., Specker, E.P.: The problem of hidden variables in quantum mechanics. Journal of Mathe- matics and Mechanics17(1), 59–87 (1967)

  34. [42]

    COLT (2015)

    Leike, J., Hutter, M.: Bad universal priors and notions of optimality. COLT (2015)

  35. [43]

    Springer, 4 edn

    Li, M., Vitányi, P.M.B.: An Introduction to Kolmogorov Complexity and Its Applications. Springer, 4 edn. (2019)

  36. [44]

    Nature Physics pp

    Liu, Y., Arunachalam, S., Temme, K.: A rigorous and robust quantum speed-up in supervised machine learning. Nature Physics pp. 1–5 (2021)

  37. [45]

    New Journal of Physics26(5), 053029 (2024)

    Lupu-Gladstein, N., Brodutch, A., Ferretti, H., Tham, W.K., Pang, A.O.T., Bonsma-Fisher, K., Stein- berg, A.M.: Do qubits dream of entangled sheep? quantum measurement without classical output. New Journal of Physics26(5), 053029 (2024)

  38. [46]

    Communications Biology7(1), 378 (Mar 2024)

    McMillen, P., Levin, M.: Collective intelligence: A unifying concept for integrating biology across scales and substrates. Communications Biology7(1), 378 (Mar 2024)

  39. [47]

    arXiv preprint arXiv:2211.03464 (2022)

    Meyer, N., Ufrecht, C., Periyasamy, M., Scherer, D.D., Plinge, A., Mutschler, C.: A survey on quantum reinforcement learning. arXiv preprint arXiv:2211.03464 (2022)

  40. [48]

    Cambridge University Press, 10th anniversary edn

    Nielsen, M.A., Chuang, I.L.: Quantum Computation and Quantum Information. Cambridge University Press, 10th anniversary edn. (2010)

  41. [49]

    arXiv preprint arXiv:1504.03303 (2015)

    Özkural, E.: Ultimate intelligence part ii: Physical measure and complexity of intelligence. arXiv preprint arXiv:1504.03303 (2015)

  42. [50]

    In: Artificial General Intelligence: 9th International Conference, AGI 2016, New York, NY, USA, July 16-19, 2016, Proceedings 9

    Özkural, E.: Ultimate intelligence part ii: physical complexity and limits of inductive inference systems. In: Artificial General Intelligence: 9th International Conference, AGI 2016, New York, NY, USA, July 16-19, 2016, Proceedings 9. pp. 33–42. Springer (2016)

  43. [51]

    CONCLUSION AND OPEN QUESTIONS 13

  44. [52]

    arXiv preprint arXiv:2402.08060 (2024)

    Pang, A.O., Lupu-Gladstein, N., Yilmaz, Y.B., Brodutch, A., Steinberg, A.M.: Information gain and measurement disturbance for quantum agents. arXiv preprint arXiv:2402.08060 (2024)

  45. [53]

    Journal of Physics A: Mathematical and General24(4), L175 (1991)

    Peres, A.: Two simple proofs of the kochen-specker theorem. Journal of Physics A: Mathematical and General24(4), L175 (1991)

  46. [54]

    arXiv:2409.04955 (2024)

    Perrier, E.: Quantum Geometric Machine Learning. arXiv:2409.04955 (2024)

  47. [55]

    15645658,https://zenodo.org/records/15645658

    Perrier, E.: Quantum AIXI - technical appendices (2025).https://doi.org/10.5281/zenodo. 15645658,https://zenodo.org/records/15645658

  48. [56]

    In: International Confer- ence on Artificial General Intelligence

    Sarkar, A., Al-Ars, Z., Bertels, K.: Qksa: Quantum knowledge seeking agent. In: International Confer- ence on Artificial General Intelligence. pp. 384–393. Springer (2022)

  49. [57]

    Physical Review A99(3) (mar 2019)

    Schuld, M., Bergholm, V., Gogolin, C., Izaac, J., Killoran, N.: Evaluating analytic gradients on quantum hardware. Physical Review A99(3) (mar 2019)

  50. [58]

    Springer (2018)

    Schuld, M., Petruccione, F.: Supervised Learning with Quantum Computers. Springer (2018)

  51. [59]

    Springer (2021)

    Schuld, M., Petruccione, F.: Machine Learning with Quantum Computers. Springer (2021)

  52. [60]

    In: Artificial General Intelligence: 8th International Conference, AGI 2015, AGI 2015, Berlin, Germany, July 22-25, 2015, Proceedings 8

    Soares, N., Fallenstein, B.: Two attempts to formalize counterpossible reasoning in deterministic set- tings. In: Artificial General Intelligence: 8th International Conference, AGI 2015, AGI 2015, Berlin, Germany, July 22-25, 2015, Proceedings 8. pp. 156–165. Springer (2015)

  53. [61]

    Part I & II

    Solomonoff, R.J.: A Formal Theory of Inductive Inference. Part I & II. Information and Control7(1–2), 1–22, 224–254 (1964)

  54. [62]

    Philosophical Transactions of the Royal Society B: Biological Sciences374(1774), 20190040 (2019)

    Solé, R., Moses, M., Forrest, S.: Liquid brains, solid brains. Philosophical Transactions of the Royal Society B: Biological Sciences374(1774), 20190040 (2019)

  55. [63]

    In: Artificial General Intelligence: 9th International Conference, AGI 2016, New York, NY, USA, July 16-19, 2016, Proceedings 9

    Thórisson, K.R., Bieger, J., Thorarensen, T., Sigurðardóttir, J.S., Steunebrink, B.R.: Why artificial intelligence needs a task theory: and what it might look like. In: Artificial General Intelligence: 9th International Conference, AGI 2016, New York, NY, USA, July 16-19, 2016...

  56. [64]

    Physical review letters91(14), 147902 (2003)

    Vidal, G.: Efficient classical simulation of slightly entangled quantum computations. Physical review letters91(14), 147902 (2003)

  57. [65]

    IEEE Transactions on Information Theory47(6), 2464–2479 (2002)

    Vitányi, P.M.: Quantum kolmogorov complexity based on classical descriptions. IEEE Transactions on Information Theory47(6), 2464–2479 (2002)

  58. [66]

    Ox- ford University Press (2012)

    Wallace, D.: The Emergent Multiverse: Quantum Theory according to the Everett Interpretation. Ox- ford University Press (2012)

  59. [67]

    Wang,P.,Hammer, P.: Assumptionsof decision-makingmodelsin agi.In: ArtificialGeneral Intelligence: 8th International Conference, AGI 2015, AGI 2015, Berlin, Germany, July 22-25, 2015, Proceedings 8. pp. 197–207. Springer (2015)

  60. [68]

    Cambridge University Press (2018)

    Watrous, J.: The Theory of Quantum Information. Cambridge University Press (2018)

  61. [69]

    In: Quantum [Un] Speakables II: Half a Century of Bell’s Theorem, pp

    Wiseman, H.M., Cavalcanti, E.G.: Causarum investigatio and the two bell’s theorems of john bell. In: Quantum [Un] Speakables II: Half a Century of Bell’s Theorem, pp. 119–142. Springer (2016) 14 E. Perrier A Entanglement in the Interaction Loop In §3.2, the post-measurement jo...

  62. [70]

    (9)) is a semi-density operator representing the agent’s belief about the environment state before them- th observation, given actionsa1:m

    Likelihood Operators and Initial Divergence Bound:The QSI mixtureΞ Q(a1:m)(Eq. (9)) is a semi-density operator representing the agent’s belief about the environment state before them- th observation, given actionsa1:m. The posteriorΞ (t) Q is obtained via updates based on obse...

  63. [71]

    Monotonicity of Quantum Relative Entropy and Martingale Argument:The core of the con- vergence argument hinges on the data-processing inequality for quantum relative entropy. This principle states that for any quantum operation (a completely-positive trace-preserving map, CPTP...

  64. [72]

    (30), andDt =D(ρ ⋆ E(a1:t)∥Ξ (t) Q (a1:t))is the divergence at stept

    Average Divergence:By summing these non-negative expected drops overtinteraction cycles, we get: tX k=1 EQ⋆,hist[∆k] =EQ⋆,hist[D0 −D t]≤D 0, whereD 0 is the initial divergence bounded as in Eq. (30), andDt =D(ρ ⋆ E(a1:t)∥Ξ (t) Q (a1:t))is the divergence at stept. This implies ...

  65. [73]

    Applying this to the result from step 3: EQ⋆ 1 2 ρ⋆ E(a1:t)−Ξ (t) Q (a1:t) 2 1 ≤E Q⋆ [Dt]≤ KQ(Q⋆) ln 2 + ln(1 +g) t

    Quantum Pinsker Inequality for Trace Distance Convergence:Finally, the quantum Pinsker in- equality provides a link between the trace distance∥ρ−σ∥1 (a measure of distinguishability) and the relative entropyD(ρ∥σ): 1 2 ∥ρ−σ∥ 2 1 ≤D(ρ∥σ). Applying this to the result from step 3...

  66. [74]

    (2)) from environments that are, by their classical nature, local

    Classical AIXI, relying on Solomonoff induction over classical Turing machinesMsol, constructs its universal priorξU (Eq. (2)) from environments that are, by their classical nature, local. Such a prior would assign zero probability to observing correlations that violate Bell i...

  67. [75]

    The term2 −KQ(Q)ρQ E(a1:m)must includeQs whoseρ Q E(a1:m)can lead to Bell- violating statistics upon appropriate measurements

    QAIXI, by summing over quantum environmentsQsol, can, in principle, learn and adapt to a non-localQ ⋆. The term2 −KQ(Q)ρQ E(a1:m)must includeQs whoseρ Q E(a1:m)can lead to Bell- violating statistics upon appropriate measurements. For example, consider a QAIXI agent designed wi...

  68. [76]

    The post-measurement state is different, and the specific instance ofρQ E(a1:t)is consumed in yielding the outcomeot

    When QAIXI performs an actionat (especially if it’s a measurement) on the environment state ρQ E(a1:t), this interaction alters the state. The post-measurement state is different, and the specific instance ofρQ E(a1:t)is consumed in yielding the outcomeot

  69. [77]

    Classical information can be copied and reused

    Classical AIXI can, in principle, take a historyh<t and test many hypothetical continuations with a given modelνwithout altering the datah <t. Classical information can be copied and reused. 26 E. Perrier

  70. [78]

    To know howQ ⋆ would have responded to a different actiona ′ t from theexact sameinstance ofρ ⋆ E(a1:t−1), is impossible

    QAIXI cannot do this with quantum states. To know howQ ⋆ would have responded to a different actiona ′ t from theexact sameinstance ofρ ⋆ E(a1:t−1), is impossible. It would needQ⋆ to produce that state (or an identical one) again. Sample ComplexityThis one-shot nature of quant...

  71. [79]

    To distinguish between different hypothesesQabout the environment, or to estimate the ex- pectedoutcome/rewardfordifferentactions,QAIXIneedstoobservetheenvironment’sresponse multiple times

  72. [80]

    Since each observation of a specific state instance is unique and unrepeatable, QAIXI effectively needsQ ⋆ to prepare a new instance of a comparable state for each piece of information it wants to gather about a particular type of situation or action

  73. [81]

    Full tomography of ann-qubit state generally requires a number of measurement settings and repetitions that scales exponentially withn

    If QAIXI wants to learn the full characteristics ofρ⋆ E(a1:t)through measurements, it is per- forming a form of quantum state tomography. Full tomography of ann-qubit state generally requires a number of measurement settings and repetitions that scales exponentially withn. Whi...

  74. [82]

    As noted above, this means the learning rate is limited by state-preparation resources. IfQ⋆ itself has computational costs or time delays associated with preparing states, this directly translates into a slower learning rate for QAIXI in terms of real-time or computational st...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.