Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:26:02.942258Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.11052.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:26:02.942258Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 28a544c9-ce1c-46a3-9b4c-2e9a0bb31ec3 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Therefore, αDKL(epπθ ∥epϕ) =E τ∼epπθ " ∞X t=1 (αlogπ θ(at |s t)−r ϕ(st, at)) # +αlogZ ϕ
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8d26906d-e9d7-4351-9676-5ae542388b98 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning ∞X t=1 γt−1 logπ θ(at |s t) # =−E τ∼p expert
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 99918331-970a-4806-b099-f850cac23993 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Natural hypergradient de- scent: Algorithm design, convergence analysis, and parallel implementation.arXiv preprint arXiv:2602.10905,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a6300e82-9013-4f85-b8f3-27690ad68883 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Natural Policy Gradients In Reinforcement Learning Explained
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a2d1dcd2-a8ef-4ae1-8586-5ac449da702a · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning ∂g ∂θ θ⋆(ϕ),ϕ #−1 ∂g ∂ϕ θ⋆(ϕ),ϕ . Since ∂g ∂θ = ∂2Linner ∂θ 2 , ∂g ∂ϕ = ∂2Linner ∂θ∂ϕ , we get dθ⋆ dϕ ϕ =−
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 835c4e51-022e-47ac-b652-c035985b099a · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning ∞X t=1 gt(τ)g t(τ) ⊤ # . Finally, applying Lemma 4.1 componentwise gives Eτ∼epπθ
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77b36f46-a84c-4f6c-8c3f-4d103b32b54e · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Souradip Chakraborty, Amrit Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, and Furong Huang
Reference 2014
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c70af340-f658-41db-9ef7-2cf63e31a3b2 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Rank-1 ap- proximation of inverse fisher for natural policy gradients in deep reinforcement learning.arXiv preprint arXiv:2601.18626,
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d5f7b85f-c790-4f6b-be2c-172610441a12 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f09e1a8-f956-4cb0-b6ec-e6f769f93771 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Frequent Directions : Simple and Deterministic Matrix Sketching
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53da9989-9480-4f50-bfb5-83c2cef374ac · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444421fd-0641-4f09-ab05-109e196003cc · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 239f0717-51cb-4e91-9462-21904e9e9e82 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Then the gradient of the induced outer objective eLouter(ϕ) :=L outer(θ⋆(ϕ)) is given by ∇ϕ eLouter ϕ =− ∂2Linner ∂ϕ∂θ θ⋆(ϕ),ϕ " ∂2Linner ∂θ 2 θ⋆(ϕ),ϕ #−1 ∇θLouter|θ⋆(ϕ)
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e58c82b5-d24e-49e6-9a5d-35eefb8e4fc7 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Scalable linucb: Low-rank design matrix updates for recommenders with large action spaces.arXiv preprint arXiv:2510.19349,
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b77cffe1-f9a4-4743-bc39-6893c3c913f5 · outbound
Efficient Hypergradient Descent for Inverse Reinforcement Learning Approximation Methods for Bilevel Programming
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.