Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:37:43.138992Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2509.01432.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:37:43.138992Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T09:24:21.047027Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 7017f336-dd9e-45ff-8230-88b646aeb530 · outbound
The Geometry of Nonlinear Reinforcement Learning Maximum a Posteriori Policy Optimisation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a833dc7-94fc-4146-b840-4d5c48f00440 · outbound
The Geometry of Nonlinear Reinforcement Learning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00258e82-80e4-45ea-8e95-57830cc71bbb · outbound
The Geometry of Nonlinear Reinforcement Learning Embedding Safety into RL: A New Take on Trust Region Methods
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a1ef6a6-4859-4cd6-84de-2a19d6439ba4 · outbound
The Geometry of Nonlinear Reinforcement Learning Central Path Proximal Policy Optimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da82822d-3ba2-44a5-917e-2123bfae2ef0 · outbound
The Geometry of Nonlinear Reinforcement Learning Challenging Common Assumptions in Convex Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b485471-8f66-456a-bb8a-ff4990223dc6 · outbound
The Geometry of Nonlinear Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd88364a-c138-46fb-9d54-6033d95ee3aa · outbound
The Geometry of Nonlinear Reinforcement Learning ∇θ log π(a′|s′) X s,a Mπ(s, a|s′, a′)rπ(s, a) # (30) = (1 − γ)Es′,a′∼ω
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aeb9da49-ed6c-420a-af72-5dde93a64e1d · outbound
The Geometry of Nonlinear Reinforcement Learning Mirror descent policy optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b3b05ba-0956-46ea-af58-317195781842 · outbound
The Geometry of Nonlinear Reinforcement Learning Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9aeef1b-a7fe-40fa-8600-2d5e8c36dc54 · outbound
The Geometry of Nonlinear Reinforcement Learning Interestingly, it also appears in the differential of the map between policy and state-action spaces
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c10c0bf0-b641-4ffe-83d7-6b67e25d83fe · outbound
The Geometry of Nonlinear Reinforcement Learning The reward isrπ(s, a) = P i zi[log pπ(i|s) − log zi]
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf70a841-3077-48e5-9eeb-61078b2c0de7 · outbound
The Geometry of Nonlinear Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2bf02d1-78b0-4fe9-b301-dc9c0ef871f9 · outbound
The Geometry of Nonlinear Reinforcement Learning Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e82d5808-1bfc-4614-9f81-0debb65398b8 · outbound
The Geometry of Nonlinear Reinforcement Learning This results in anintractable policy divergence with Hessian 16 GTML 2025 HC(θ) = Es∼ωπ F (θ) + X i βiϕ′′(bi − Vci (θ))∇2 θVci (θ) θ=θk
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0433396-15d7-4f67-b278-5d0b262b1b41 · outbound
The Geometry of Nonlinear Reinforcement Learning 17 GTML 2025 When this standard geometry is restricted to the manifoldΩ, i.e
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69823b71-98d0-4259-bac2-dcf6352681f2 · outbound
The Geometry of Nonlinear Reinforcement Learning Definition C.1(Successor Representation)
Reference 1993
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3426bb7-af16-429d-8bb2-cc5bfcfd3252 · outbound
The Geometry of Nonlinear Reinforcement Learning Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints
Reference 1994
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1fc9aea-a97a-4876-811b-c0960063ca8f · outbound
The Geometry of Nonlinear Reinforcement Learning Diversity is All You Need: Learning Skills without a Reward Function
Reference 1999
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ad5aa8-9218-485c-9508-bfe56b4e9f54 · outbound
The Geometry of Nonlinear Reinforcement Learning Motivation for a General Hessian FrameworkThe existence of at least two distinct, natural geometries on the same occupancy manifold is a key insight
Reference 2001
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 785880e8-847f-4a4d-b36e-ed8247e1eb0f · outbound
The Geometry of Nonlinear Reinforcement Learning Discovering Diverse Nearly Optimal Policies with Successor Features
Reference 2006
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 577f2040-8e0c-4fc8-84f9-032c966a58e6 · outbound
The Geometry of Nonlinear Reinforcement Learning Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning
Reference 2007
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2546c89-beba-4e33-9484-133f4faf63b9 · outbound
The Geometry of Nonlinear Reinforcement Learning ∞X t=0 γtf (st, at) # = Es,a∼dµ π [f (s, a)] (5) 8 GTML 2025 Proof. (1 − γ)Eτ ∼π,µ
Reference 2008
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e23415ff-ead5-4ac1-be5c-778fe12aa0a4 · outbound
The Geometry of Nonlinear Reinforcement Learning Unresolved cited work
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b258de-e9bb-4761-8430-f7f02e90b5ad · outbound
The Geometry of Nonlinear Reinforcement Learning Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5e69283-05ea-44f4-a252-07ef7c669a48 · outbound
The Geometry of Nonlinear Reinforcement Learning Variational Intrinsic Control
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b51656a-4d66-476c-bf2c-b372f13c1876 · outbound
The Geometry of Nonlinear Reinforcement Learning cc/paper_files/paper/2019/file/873be0705c80679f2c71fbf4d872df59-Paper.pdf
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe7a6825-0ebf-4860-99c5-caa4cffe3680 · outbound
The Geometry of Nonlinear Reinforcement Learning A unified view of entropy-regularized Markov decision processes
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba514d3-52d7-4f4c-b52e-f0152f3eb758 · outbound
The Geometry of Nonlinear Reinforcement Learning Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0003b65f-ac2a-4d90-bd5a-7a38a7dc51b4 · outbound
The Geometry of Nonlinear Reinforcement Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbecf1e9-e25b-42d0-a701-e49ea4ca7953 · outbound
The Geometry of Nonlinear Reinforcement Learning Fast Task Inference with Variational Intrinsic Successor Features
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b588e5e9-9603-43e3-81cf-0272262e8a24 · outbound
The Geometry of Nonlinear Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca580af9-5ac7-4c3c-93c9-5f55179017f4 · inbound
Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs The Geometry of Nonlinear Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.