Pith. sign in

Paper Citation Record · LEDGER

Efficient Hypergradient Descent for Inverse Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.11052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11052 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:26:02.942258Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact5
  • verified fuzzy5
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28a544c9-ce1c-46a3-9b4c-2e9a0bb31ec3 · outbound

This paper cites Therefore, αDKL(epπθ ∥epϕ) =E τ∼epπθ " ∞X t=1 (αlogπ θ(at |s t)−r ϕ(st, at)) # +αlogZ ϕ.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Therefore, αDKL(epπθ ∥epϕ) =E τ∼epπθ " ∞X t=1 (αlogπ θ(at |s t)−r ϕ(st, at)) # +αlogZ ϕ

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.543094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.934443Z digest=sha256:198ca0b13dee07c1fddca6f46290bf4b2853e813213d50c43167b4cfbf6d0bf2

Observation 8d26906d-e9d7-4351-9676-5ae542388b98 · outbound

This paper cites ∞X t=1 γt−1 logπ θ(at |s t) # =−E τ∼p expert.

Efficient Hypergradient Descent for Inverse Reinforcement Learning ∞X t=1 γt−1 logπ θ(at |s t) # =−E τ∼p expert

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.530082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.938246Z digest=sha256:e5698c215b21a09b56dacfbc17259ad37e87007cdb42ccf29c96d49ae287d09f

Observation 99918331-970a-4806-b099-f850cac23993 · outbound

This paper cites Natural hypergradient de- scent: Algorithm design, convergence analysis, and parallel implementation.arXiv preprint arXiv:2602.10905,.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Natural hypergradient de- scent: Algorithm design, convergence analysis, and parallel implementation.arXiv preprint arXiv:2602.10905,

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-12T11:26:03.319693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.906229Z digest=sha256:14b12d481e16494c8a35b995b0d136df84e07b5d6af6194ca91d845adb5864ad

Observation a6300e82-9013-4f85-b8f3-27690ad68883 · outbound

This paper cites Natural Policy Gradients In Reinforcement Learning Explained.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Natural Policy Gradients In Reinforcement Learning Explained

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:26:03.105051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.918211Z digest=sha256:78d8c1c3539f6377532802a064a6478b1eae25eb55fa803f71a9ae29f5fa0914

Observation a2d1dcd2-a8ef-4ae1-8586-5ac449da702a · outbound

This paper cites ∂g ∂θ θ⋆(ϕ),ϕ #−1 ∂g ∂ϕ θ⋆(ϕ),ϕ . Since ∂g ∂θ = ∂2Linner ∂θ 2 , ∂g ∂ϕ = ∂2Linner ∂θ∂ϕ , we get dθ⋆ dϕ ϕ =−.

Efficient Hypergradient Descent for Inverse Reinforcement Learning ∂g ∂θ θ⋆(ϕ),ϕ #−1 ∂g ∂ϕ θ⋆(ϕ),ϕ . Since ∂g ∂θ = ∂2Linner ∂θ 2 , ∂g ∂ϕ = ∂2Linner ∂θ∂ϕ , we get dθ⋆ dϕ ϕ =−

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.554667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.930040Z digest=sha256:d1429b272e27cda72d07bd223ffb486b94f85b1ad26f1c822aabacc80bd38fc9

Observation 835c4e51-022e-47ac-b652-c035985b099a · outbound

This paper cites ∞X t=1 gt(τ)g t(τ) ⊤ # . Finally, applying Lemma 4.1 componentwise gives Eτ∼epπθ.

Efficient Hypergradient Descent for Inverse Reinforcement Learning ∞X t=1 gt(τ)g t(τ) ⊤ # . Finally, applying Lemma 4.1 componentwise gives Eτ∼epπθ

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.516484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.942258Z digest=sha256:aacaf6988145d5844ac2eb89700a97e69f6cdabf6be8677009757f0cb4109bc4

Observation 77b36f46-a84c-4f6c-8c3f-4d103b32b54e · outbound

This paper cites Souradip Chakraborty, Amrit Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, and Furong Huang.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Souradip Chakraborty, Amrit Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, and Furong Huang

Reference 2014

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T11:26:03.505054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.885600Z digest=sha256:7f5aedbc0aaa38de8465c11d6651e343a59e40c240feb75d35fd8f14650aa5fb

Observation c70af340-f658-41db-9ef7-2cf63e31a3b2 · outbound

This paper cites Rank-1 ap- proximation of inverse fisher for natural policy gradients in deep reinforcement learning.arXiv preprint arXiv:2601.18626,.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Rank-1 ap- proximation of inverse fisher for natural policy gradients in deep reinforcement learning.arXiv preprint arXiv:2601.18626,

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-12T11:26:03.397415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.902323Z digest=sha256:d115a79a2db29327a8f407153309ecd028dc70ed2b7aa786bb6b9fcc3a648a99

Observation d5f7b85f-c790-4f6b-be2c-172610441a12 · outbound

This paper cites Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.922372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.922372Z digest=sha256:9a8932f738a1356430731c96e6c8317c98e28a783a02e15c2bbbe30ef40e6d71

Observation 9f09e1a8-f956-4cb0-b6ec-e6f769f93771 · outbound

This paper cites Frequent Directions : Simple and Deterministic Matrix Sketching.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Frequent Directions : Simple and Deterministic Matrix Sketching

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.897810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.897810Z digest=sha256:29c29b9036a355ab4f75699644c24e453a0eb2215a3dc34af717dc1e97c6938a

Observation 53da9989-9480-4f50-bfb5-83c2cef374ac · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.909932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.909932Z digest=sha256:3fcb49875fd4adec8208f7931838d5fcfdff2a11f6288ccd73c325913061f4c0

Observation 444421fd-0641-4f09-ab05-109e196003cc · outbound

This paper cites Explaining and Preventing Alignment Collapse in Iterative RLHF.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:26:03.437865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.889788Z digest=sha256:7e12d6696cb818343da5e8ca226d9366caa44d3f0b9bf0f4e9468b46ceeee8f6

Observation 239f0717-51cb-4e91-9462-21904e9e9e82 · outbound

This paper cites Then the gradient of the induced outer objective eLouter(ϕ) :=L outer(θ⋆(ϕ)) is given by ∇ϕ eLouter ϕ =− ∂2Linner ∂ϕ∂θ θ⋆(ϕ),ϕ " ∂2Linner ∂θ 2 θ⋆(ϕ),ϕ #−1 ∇θLouter|θ⋆(ϕ).

Efficient Hypergradient Descent for Inverse Reinforcement Learning Then the gradient of the induced outer objective eLouter(ϕ) :=L outer(θ⋆(ϕ)) is given by ∇ϕ eLouter ϕ =− ∂2Linner ∂ϕ∂θ θ⋆(ϕ),ϕ " ∂2Linner ∂θ 2 θ⋆(ϕ),ϕ #−1 ∇θLouter|θ⋆(ϕ)

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:26:03.568289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.926453Z digest=sha256:1ae14f537a3317a93f5d0513b7117ef2f96e76845720aa662f9ce7629b9f0ed5

Observation e58c82b5-d24e-49e6-9a5d-35eefb8e4fc7 · outbound

This paper cites Scalable linucb: Low-rank design matrix updates for recommenders with large action spaces.arXiv preprint arXiv:2510.19349,.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Scalable linucb: Low-rank design matrix updates for recommenders with large action spaces.arXiv preprint arXiv:2510.19349,

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-12T11:26:03.214880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T11:26:02.914384Z digest=sha256:73a767e7b4af079fa7d15def0b03456c71587b051e5fe4893dccfc0cf84a0f51

Observation b77cffe1-f9a4-4743-bc39-6893c3c913f5 · outbound

This paper cites Approximation Methods for Bilevel Programming.

Efficient Hypergradient Descent for Inverse Reinforcement Learning Approximation Methods for Bilevel Programming

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T11:26:02.893846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:26:02.893846Z digest=sha256:e2b6b64b6ddb9d2b20f0c2d2c8dcf7f6623a854900d3dc6b8fce3825fa3d9608

Pith citing papers

No inbound Pith citation observations are available.