Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T15:00:44.145471Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2512.18763.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T15:00:44.145471Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-30T21:23:27.894346Z
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cccccf50-b56f-46a0-a4af-9b92bbb9f1e5 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Bertsekas,Reinforcement Learning and Optimal Control
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8b8b4b-c79d-4b77-9c8d-62f8a1846ec4 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e33159-f696-47b5-8a06-47aee17f0c0c · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Q-learning,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538eac2f-24f4-4606-bc6e-41ac1b2b4d77 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Convergence results for single-step on-policy reinforcement-learning algorithms,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1923cd9a-ca93-42dd-8144-06fcdc64dcbe · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Kernel-based reinforcement learning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e1b635c-5caa-4c83-af7e-65dabe41ad6d · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Kernel-based reinforcement learning in average-cost problems,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6085959e-d528-487c-8c1f-b659e5943ad8 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Stochastic kernel temporal difference for reinforcement learning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800e07e6-e3f2-4c7a-8bb0-b8602b0f40a4 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Learning to predict by the methods of temporal differ- ences,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d186437-b2a4-478e-aaa4-f711b997671c · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Least-squares policy iteration,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b082f3d-7235-4fc0-bf83-0d76c3ac3e68 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Regularized policy iteration with nonparametric function spaces,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122b5260-d903-4da2-b62a-8355c89fdfc9 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Kernel-based least squares policy iteration for reinforcement learning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e01dbce6-101e-4ce9-8287-40eb7d39c3e7 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Online Bellman residual and temporal difference algorithms with predictive error guarantees,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa52abcf-19fd-439a-9829-4646586fa9a6 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Dynamic selection of p- norm in linear adaptive filtering via online kernel-based reinforcement learning,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16d4662-b278-4b2a-9eb6-f5db820748dc · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Proximal Bellman mappings for rein- forcement learning and their application to robust adaptive filtering,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe271ff-fb7e-4a15-ba2e-fe620db329fa · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Nonparametric Bellman map- pings for reinforcement learning: Application to robust adaptive filter- ing,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8355b67a-2076-4a1a-966c-419ac7be4d1f · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Theory of reproducing kernels,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64edac9f-38ac-4f7a-8ed2-2d8dd09b363d · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Schölkopf and A
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1df23fc-7010-44a7-8e20-25959e1f79bc · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Playing Atari with Deep Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4ee15a5-c33f-4700-95f4-5bdba6824f16 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Deep reinforcement learning with double Q-learning,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9356b85-9b00-4db6-8eae-84ddc859700d · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Reinforcement learning for robots using neural networks,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae98b34-96ee-4f01-978d-cf256a871126 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Reinforcement learning based on on-line EM algorithm,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6440d4f-45bc-4ef7-981b-274b00adca48 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Reinforcement learning with Gaussian processes,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3ac583-5915-4746-96ae-a69f00277375 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Reinforcement learning with a Gaussian mixture model,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf0cbf36-bf79-4a11-a559-3dab74876e60 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Online reinforcement learning using a probability density estimation,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a6e18fd-5c8e-48d7-9a55-d43fb261ac90 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Distributional deep reinforcement learning with a mixture of Gaussians,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd0886f-550f-4e81-a439-d071fa951ad9 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Gaussian mixture models,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2890067a-993b-464f-abfc-6c15bd6fc123 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning McLachlan and D
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75c20cd-f5d4-4155-975e-fe69a9d7b030 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Maximum likelihood from incomplete data via the EM algorithm,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f83b8f9d-cd82-4272-a975-2eb15470fcee · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Unsupervised learning of finite mixture models,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33994f8a-b74d-48a4-8b2a-30324f61a1e1 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Absil, R
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bce3830-30bc-433a-9a26-5895c20fae36 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Rie- mannian proximal policy optimization,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61fabfe3-29af-44f3-8604-fb0d6711790c · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Policy gradient methods for reinforcement learning with function approximation,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87b11ec-d658-4798-9d20-1193cf6bd2d6 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Approximation and radial-basis-function networks,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a34a34e6-1758-4b25-8245-c31e18a4b42b · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Riemannian Q-functions for policy iteration in reinforcement learning,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ceed7cb-9bdc-4899-b6c0-de888843039d · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 299d50ca-68bc-4bf7-8b48-5f25f371a6b8 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Tight performance bounds on greedy policies based on imperfect value functions,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a478d8-e42e-4032-9dfd-7eb9999ae22c · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning A correspondence between Bayesian estimation on stochastic processes and smoothing by splines,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1044a165-ffc6-4346-9b63-74a4326bb042 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Universal approximation capability of EBF neural networks with arbitrary activation functions,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb26320-5c77-4a12-b487-70a56b7d1fb9 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32153aa7-e41c-4793-a8ea-64a7c885a306 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Williams,Probability with Martingales
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcf3dbe5-4f35-4d4c-85f1-42e801afae75 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 190f0118-1b9d-4d9a-bf30-c465cafb6e52 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Rudin,Real and Complex Analysis, 3rd ed
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a42f425f-f0e7-4518-807f-819e714a0b7a · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Wasserstein Riemannian geometry of Gaussian densities,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a32366-ddf7-4e9a-a8fa-cb5ca0bb7bf5 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning A sampling approach to finding Lyapunov functions for nonlinear discrete-time systems,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb3f4e8-2daf-48ef-8b40-fb6155c7b56b · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Boumal,An Introduction to Optimization on Smooth Manifolds
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0c3880-bc09-43f4-a4fe-b2edee3979d9 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Early stopping—but when?
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10912af-6da8-4f2e-97ba-1343143a56f5 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Learning near-optimal poli- cies with bellman-residual minimization based fitted policy iteration and a single sample path,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65a8e296-ff23-403a-942e-ee2cbd1726d3 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Deterministic Bellman residual minimization,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0b9e33-f2f4-4e64-ad0c-6e5d7e05978e · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Big Omicron and big Omega and big Theta,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259417fb-cbfb-4951-b447-39aa192f71d2 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Reinforcement learning in continuous time and space,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99ac932-6bb8-435e-8319-4bcf81e86838 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Efficient memory-based learning for robot control,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a065fd9-f530-4c43-a003-d986e25616de · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning The swing up control problem for the acrobot,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe09ffb-b129-46ca-85df-6ce287e156c6 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Dueling network architectures for deep reinforcement learning,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb698973-f46e-4815-8eed-74d5a33ce3cd · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd934d2d-3976-49cb-8954-c613fe2fbaac · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Julia: A fresh approach to numerical computing,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31ccf1c-9e64-4f23-b859-fa24d7b7b9c6 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Robust reinforcement learning using least squares policy iteration with provable performance guarantees,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3fc58a1-2470-4994-a10b-ca3c41af1a21 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Geron,Tiny-dqn, https://github.com/ageron/tiny-dqn, 2017
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b819367f-3f15-46ff-ae46-43a413d06ed2 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f12b05-1702-4d05-8162-2e080e7b0234 · outbound
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning Density in approximation theory,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef344002-af4d-4b95-a275-c1b983409b7e · inbound
Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.