Pith. sign in

Paper Citation Record · LEDGER

The Geometry of Nonlinear Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2509.01432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01432 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:37:43.138992Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T09:24:21.047027Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7017f336-dd9e-45ff-8230-88b646aeb530 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

The Geometry of Nonlinear Reinforcement Learning Maximum a Posteriori Policy Optimisation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.058718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.058718Z digest=sha256:49eeb044f3da7fa8bb8c0bbb577bb0fdb14fb77d35e729d6b8fbb5ec1824342b

Observation 4a833dc7-94fc-4146-b840-4d5c48f00440 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:37:45.041818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.672881Z digest=sha256:cb42fdaa06cf6a4a3a6d4593ce1b01c8a03ea43ccf556a90c2ad82ad01152c9d

Observation 00258e82-80e4-45ea-8e95-57830cc71bbb · outbound

This paper cites Embedding Safety into RL: A New Take on Trust Region Methods.

The Geometry of Nonlinear Reinforcement Learning Embedding Safety into RL: A New Take on Trust Region Methods

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:43.877379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:41.538793Z digest=sha256:1da070d085b1985b966be3327c32032cca67a5459869c57527a39d5588df3c09

Observation 7a1ef6a6-4859-4cd6-84de-2a19d6439ba4 · outbound

This paper cites Central Path Proximal Policy Optimization.

The Geometry of Nonlinear Reinforcement Learning Central Path Proximal Policy Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.630698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.630698Z digest=sha256:8abc165b1052c09973d4d61e97a2c00073e6fd22d6c068395fa9ab25a80fa15c

Observation da82822d-3ba2-44a5-917e-2123bfae2ef0 · outbound

This paper cites Challenging Common Assumptions in Convex Reinforcement Learning.

The Geometry of Nonlinear Reinforcement Learning Challenging Common Assumptions in Convex Reinforcement Learning

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:37:43.566193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:41.781429Z digest=sha256:a7ebeb345c21a0be7764372188b966aa45c192b7b024a443f6e2665ecf998901

Observation 9b485471-8f66-456a-bb8a-ff4990223dc6 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

The Geometry of Nonlinear Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.925676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.925676Z digest=sha256:f9e119aef236a10f42925b9d74a17592594fcf79a48c80843e2e6375e9790e5d

Observation cd88364a-c138-46fb-9d54-6033d95ee3aa · outbound

This paper cites ∇θ log π(a′|s′) X s,a Mπ(s, a|s′, a′)rπ(s, a) # (30) = (1 − γ)Es′,a′∼ω.

The Geometry of Nonlinear Reinforcement Learning ∇θ log π(a′|s′) X s,a Mπ(s, a|s′, a′)rπ(s, a) # (30) = (1 − γ)Es′,a′∼ω

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.051249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.618804Z digest=sha256:7ae849e2f837ae4bc69922699cbdca76110480a5cbab8fc6b151c2e8e870845a

Observation aeb9da49-ed6c-420a-af72-5dde93a64e1d · outbound

This paper cites Mirror descent policy optimization.

The Geometry of Nonlinear Reinforcement Learning Mirror descent policy optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.091767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.255450Z digest=sha256:95f0ba3889fa4075a2b3b652a88fa639d6cfcc9a0422739470734bc0e4884fd5

Observation 0b3b05ba-0956-46ea-af58-317195781842 · outbound

This paper cites Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence.

The Geometry of Nonlinear Reinforcement Learning Policy Mirror Descent for Regularized Reinforcement Learning: A Generalized Framework with Linear Convergence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:42.367138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:42.367138Z digest=sha256:6a2d40db48da42ee9dc57c1fc1f605cef50d62bd1c6b45765116f044ffc86526

Observation a9aeef1b-a7fe-40fa-8600-2d5e8c36dc54 · outbound

This paper cites Interestingly, it also appears in the differential of the map between policy and state-action spaces.

The Geometry of Nonlinear Reinforcement Learning Interestingly, it also appears in the differential of the map between policy and state-action spaces

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.061885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.537762Z digest=sha256:ceed4a383c185e7a9991dcf0579a8f5b3a8e5d8f01505b1787e570151754351f

Observation c10c0bf0-b641-4ffe-83d7-6b67e25d83fe · outbound

This paper cites The reward isrπ(s, a) = P i zi[log pπ(i|s) − log zi].

The Geometry of Nonlinear Reinforcement Learning The reward isrπ(s, a) = P i zi[log pπ(i|s) − log zi]

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.031379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.757521Z digest=sha256:559ada9237fcd01143200c21e51cd53a8caab1d8b4a8a130f639309a104b1243

Observation bf70a841-3077-48e5-9eeb-61078b2c0de7 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:37:45.020228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.843902Z digest=sha256:561497fb2f1dae7f91201447cecfb0ec48591e989d97d5caee72d5035b8f565e

Observation e2bf02d1-78b0-4fe9-b301-dc9c0ef871f9 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:37:45.010369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.902934Z digest=sha256:a8f72d7b79583e7654f885cd478dbed3e2bc3d2655dfba75acd67e0165ddd66a

Observation e82d5808-1bfc-4614-9f81-0debb65398b8 · outbound

This paper cites This results in anintractable policy divergence with Hessian 16 GTML 2025 HC(θ) = Es∼ωπ F (θ) + X i βiϕ′′(bi − Vci (θ))∇2 θVci (θ) θ=θk.

The Geometry of Nonlinear Reinforcement Learning This results in anintractable policy divergence with Hessian 16 GTML 2025 HC(θ) = Es∼ωπ F (θ) + X i βiϕ′′(bi − Vci (θ))∇2 θVci (θ) θ=θk

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:44.999360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.993772Z digest=sha256:25cfb3164e3a43fb9e190fa7c6728d38aa38343f7daa02bd04472117aa46d677

Observation c0433396-15d7-4f67-b278-5d0b262b1b41 · outbound

This paper cites 17 GTML 2025 When this standard geometry is restricted to the manifoldΩ, i.e.

The Geometry of Nonlinear Reinforcement Learning 17 GTML 2025 When this standard geometry is restricted to the manifoldΩ, i.e

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:44.988359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:43.077827Z digest=sha256:c676252c662be774028d7d7d42518adff4953fd117b049dcbc6201c173e1013c

Observation 69823b71-98d0-4259-bac2-dcf6352681f2 · outbound

This paper cites Definition C.1(Successor Representation).

The Geometry of Nonlinear Reinforcement Learning Definition C.1(Successor Representation)

Reference 1993

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.071461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.484428Z digest=sha256:5eb45a9f113f56136bca824a92c59a9bbf9900fb789ccd2b4bd9d8609c0407ea

Observation d3426bb7-af16-429d-8bb2-cc5bfcfd3252 · outbound

This paper cites Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints.

The Geometry of Nonlinear Reinforcement Learning Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints

Reference 1994

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:44.142127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:41.477513Z digest=sha256:385230183fd68c29a2ce27f0020ef2549ce78430c5224a27e1c920c18fd01a52

Observation b1fc9aea-a97a-4876-811b-c0960063ca8f · outbound

This paper cites Diversity is All You Need: Learning Skills without a Reward Function.

The Geometry of Nonlinear Reinforcement Learning Diversity is All You Need: Learning Skills without a Reward Function

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.228201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.228201Z digest=sha256:09a4b37b6a7b8cc36dede8f7a0f20aeb5ac2f99245afcd63af9a930d1093eb00

Observation 14ad5aa8-9218-485c-9508-bfe56b4e9f54 · outbound

This paper cites Motivation for a General Hessian FrameworkThe existence of at least two distinct, natural geometries on the same occupancy manifold is a key insight.

The Geometry of Nonlinear Reinforcement Learning Motivation for a General Hessian FrameworkThe existence of at least two distinct, natural geometries on the same occupancy manifold is a key insight

Reference 2001

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:44.977412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:43.138992Z digest=sha256:962390ae63f3dd28d25f4c7aca44de94725d10d5975ea6a44609aff8088b2756

Observation 785880e8-847f-4a4d-b36e-ed8247e1eb0f · outbound

This paper cites Discovering Diverse Nearly Optimal Policies with Successor Features.

The Geometry of Nonlinear Reinforcement Learning Discovering Diverse Nearly Optimal Policies with Successor Features

Reference 2006

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:43.327076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.318169Z digest=sha256:13bbb429632ae46228538070ccbc00d25bfdc058f293c95fd2038fd175f1010a

Observation 577f2040-8e0c-4fc8-84f9-032c966a58e6 · outbound

This paper cites Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning.

The Geometry of Nonlinear Reinforcement Learning Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning

Reference 2007

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.100499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.172013Z digest=sha256:bc26e6afda9e44c242c9a7fc151e79f199eb3f5ab74f51362eb93a23c5cb2a28

Observation e2546c89-beba-4e33-9484-133f4faf63b9 · outbound

This paper cites ∞X t=0 γtf (st, at) # = Es,a∼dµ π [f (s, a)] (5) 8 GTML 2025 Proof. (1 − γ)Eτ ∼π,µ.

The Geometry of Nonlinear Reinforcement Learning ∞X t=0 γtf (st, at) # = Es,a∼dµ π [f (s, a)] (5) 8 GTML 2025 Proof. (1 − γ)Eτ ∼π,µ

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.081576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:42.447186Z digest=sha256:b2b82a17e06397f762cce0a583f8f38776305203f6a8b12c423f9e3c60636820

Observation e23415ff-ead5-4ac1-be5c-778fe12aa0a4 · outbound

This paper cites an unresolved cited work.

The Geometry of Nonlinear Reinforcement Learning Unresolved cited work

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:42.043518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:42.043518Z digest=sha256:0bd5a2635424621910649fe4b7ff8d8fb320a6030ae7a09b405f0ad5a9cb59ec

Observation 40b258de-e9bb-4761-8430-f7f02e90b5ad · outbound

This paper cites Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics.

The Geometry of Nonlinear Reinforcement Learning Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics

Reference 2016

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:37:44.503773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:41.368134Z digest=sha256:36bc6473cb8234ff133db234ec5d1a89ff1cd4ff145a1f9cb418c4edad11d89e

Observation b5e69283-05ea-44f4-a252-07ef7c669a48 · outbound

This paper cites Variational Intrinsic Control.

The Geometry of Nonlinear Reinforcement Learning Variational Intrinsic Control

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.314647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.314647Z digest=sha256:89c1e7c81caae01b29d281b9ec8c3b35e674fc2a0901b4418d2c40db4bfd8660

Observation 6b51656a-4d66-476c-bf2c-b372f13c1876 · outbound

This paper cites cc/paper_files/paper/2019/file/873be0705c80679f2c71fbf4d872df59-Paper.pdf.

The Geometry of Nonlinear Reinforcement Learning cc/paper_files/paper/2019/file/873be0705c80679f2c71fbf4d872df59-Paper.pdf

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:37:45.109888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:41.701622Z digest=sha256:2175af9966de077047193ada491c4e8d9ebaea13e70aa6f0b5d2fd09d0e24cd2

Observation fe7a6825-0ebf-4860-99c5-caa4cffe3680 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

The Geometry of Nonlinear Reinforcement Learning A unified view of entropy-regularized Markov decision processes

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.845185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.845185Z digest=sha256:2ba5fb5834f21150bd151dad1543acbed528fd92ca6b732b3ae590111d8ebb99

Observation 8ba514d3-52d7-4f4c-b52e-f0152f3eb758 · outbound

This paper cites Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint.

The Geometry of Nonlinear Reinforcement Learning Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:37:44.810808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:37:41.266986Z digest=sha256:a6c1234a9609ec0898a0d2125a66cfaacf786152e6f039ef4b18590c4da2d2e3

Observation 0003b65f-ac2a-4d90-bd5a-7a38a7dc51b4 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

The Geometry of Nonlinear Reinforcement Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.141219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.141219Z digest=sha256:df0d835ea31d0c49a80b6d466b53f184f461fae47ba27a70c94c5cf25254e7f7

Observation cbecf1e9-e25b-42d0-a701-e49ea4ca7953 · outbound

This paper cites Fast Task Inference with Variational Intrinsic Successor Features.

The Geometry of Nonlinear Reinforcement Learning Fast Task Inference with Variational Intrinsic Successor Features

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.422027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.422027Z digest=sha256:f1dac7a376474e6a7c52f7ce4cb637a1fc08efd52569c398b375c4e4f9f9ecc4

Observation b588e5e9-9603-43e3-81cf-0272262e8a24 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Geometry of Nonlinear Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:42.090426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:42.090426Z digest=sha256:bf2c94a521fed05f111cf3b6a83f536aec02dda24c2b897536f353838c494524

Pith citing papers

Observation ca580af9-5ac7-4c3c-93c9-5f55179017f4 · inbound

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs cites this paper.

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs The Geometry of Nonlinear Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:24:31.577549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:24:21.047027Z digest=sha256:fa94fee98fddea9f9ac05e4a5d77a5d83f2d0437057acaf985709cd588f70273