Pith. sign in

Paper Citation Record · LEDGER

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning

As of 25 July 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2601.16399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.16399 v6

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T11:36:29.670003Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact5
  • verified fuzzy14
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation adef480a-570d-47bf-a625-58295ecd19ec · outbound

This paper cites On the sample complexity bounds in bilevel reinforcement learning.arXiv preprint arXiv:2503.17644.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning On the sample complexity bounds in bilevel reinforcement learning.arXiv preprint arXiv:2503.17644

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:37:48.710780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:9d40d07ec7a7f90b98241b5647a504b71cf329e2472c2f4600691ae207310a06

Observation 89a21b52-0f2a-416b-85f2-8f8b344bb7fe · outbound

This paper cites Approximation Methods for Bilevel Programming.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Approximation Methods for Bilevel Programming

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:37:48.714049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:17dc50b0c153d4f9c5d8d5bcec4c9e0058d42947b57248a091817327fb05d26b

Observation 4dae121b-13fd-45aa-af4e-7832f379d756 · outbound

This paper cites Using Synthetic Data to Mitigate Unfairness and Preserve Privacy in Collaborative Machine Learning.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Using Synthetic Data to Mitigate Unfairness and Preserve Privacy in Collaborative Machine Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:37:48.717263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:b82e7673e73096f3745cf67f17cc99d278822d55f706738b6277e67163edbecc

Observation f9d736c9-7058-4291-8ee5-1deef649af6a · outbound

This paper cites Unlocking Global Optimality in Bilevel Optimization: A Pilot Study.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unlocking Global Optimality in Bilevel Optimization: A Pilot Study

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:37:48.707143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:05ed6480dd3d484077bc130aafaadd7599bb034667bc24693223c63d3c10003f

Observation 832a16cd-8b5a-49e2-b2d3-141cd30ea64f · outbound

This paper cites A First-order Generative Bilevel Optimization Framework for Diffusion Models.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning A First-order Generative Bilevel Optimization Framework for Diffusion Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:37:48.720397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:29461b0664a7b5c829f93e4c4ba420ed297b1bbe218d9596c81221794d4a2519

Observation dbe74d4e-3258-4936-8d81-71e298b05384 · outbound

This paper cites samples drawn from the stationary distribution, instead of continuously generated Markovian samples.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning samples drawn from the stationary distribution, instead of continuously generated Markovian samples

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.334005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:5143da75e34e1ca04c524ff804197b866beaa50ab6d49db2f31fedd4b01cba24

Observation d00c2528-db7a-41b6-9dc8-6faead6554b2 · outbound

This paper cites an unresolved cited work.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.361030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:9c1a4747b5a997637795269e0d827231445baabd9be30bbdb347c58043024fbc

Observation 0e26ac44-9db9-4719-bc4e-1e79df3b33d2 · outbound

This paper cites We defer the proof of the lemma to Appendix E.15.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning We defer the proof of the lemma to Appendix E.15

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.352671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:2bff8fd7f67ed86e15ca29366779f672e71f24c4483f9c2a85617810f720c5ac

Observation af1f6328-adbb-4abb-afb2-374ce51bcf54 · outbound

This paper cites The bound onE[ε V,L k+1]can be derived using an identical argument.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning The bound onE[ε V,L k+1]can be derived using an identical argument

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.355006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:9b94aa8512f1c5d00216c0e9847cbda174cc670b144b12bba558bc20513c21ff

Observation 3008e215-0487-45ad-bf9d-82daddfcb7a0 · outbound

This paper cites dπ′ ρ = 1 2 dπ1 ρ + 1 2 dπ2 ρ .(93) We use ˆdπ ρ to denote the extend discounted visitation distribution over state and action such that ˆdπ ρ(s, a) =d π ρ(s)π(a|s).

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning dπ′ ρ = 1 2 dπ1 ρ + 1 2 dπ2 ρ .(93) We use ˆdπ ρ to denote the extend discounted visitation distribution over state and action such that ˆdπ ρ(s, a) =d π ρ(s)π(a|s)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.348635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:0b867c148c130b430cb48e965e508d2b36408d098f715aa4b4929e93a2ff7c50

Observation a68be447-d618-411f-a2bf-70406579b392 · outbound

This paper cites an unresolved cited work.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.358849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:21059b9cecd6c38bda7798ae0aa3f0d3ecb92e0778ca9d1d50f588482660f737

Observation 070d11a7-e6c1-4848-8c1b-6be95407178b · outbound

This paper cites an unresolved cited work.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.346332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:cac6d8d16430a7cf454e8d74bcb4fab8516f19f580cb28830bd0e0e63e5105d2

Observation e3f8a955-3391-455f-aa0e-ecd769f02bdd · outbound

This paper cites Adapting the result from Shen et al.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Adapting the result from Shen et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.350782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:b6e257a87a358a2ea98e46d6b6072903efb66e0d760c13530b12a071aba81566

Observation 11785769-1754-42ce-9fd0-bbd9cb90af09 · outbound

This paper cites an unresolved cited work.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.356926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:a3a1304539973d85f782026df7cd80d3b402dfc245422fe9eba80368d64b4e52

Observation 76feface-f4b8-4422-acea-fe3b4cddcc08 · outbound

This paper cites an unresolved cited work.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.344349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:029846b24c99d793b0036d1128603c857c23092ad67f0696c28488d80676ccbd

Observation 00dd948d-1eb9-4bca-894e-6bfa89214a90 · outbound

This paper cites Therefore, ∇2 τ,θ Jτ(x, πθ) = 1 1−γ ∇θEs∼d πθρ , a∼πθ(·|s)[E(πθ, s)].(126) Zeng et al.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Therefore, ∇2 τ,θ Jτ(x, πθ) = 1 1−γ ∇θEs∼d πθρ , a∼πθ(·|s)[E(πθ, s)].(126) Zeng et al

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.340200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:f759621c25a881732a980c25468f6c9371a3dbb717c89f992ca990fb98e1f1ac

Observation da26f9de-4db2-459c-95d4-c04607036254 · outbound

This paper cites We next show the smoothness ofΦ w,τ.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning We next show the smoothness ofΦ w,τ

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.342077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:e46cec9a3b000753bc8479db1cf8f514f2a9aeaab056f361251acb62e3b12b29

Observation b3a24f6b-eef7-47aa-8e05-5b18de37ac11 · outbound

This paper cites Kwon et al.

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Kwon et al

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.336075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:7af24757b28394ca015198b59c28cd06e1ebe4728ed4435c1956cf8c4e0b254c

Observation a82a87b7-ce37-474d-90c8-842842db5674 · outbound

This paper cites This implies∇ 2 x,θJ(x, πθ) =∇ 2 x,θJτ(x, πθ).

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning This implies∇ 2 x,θJ(x, πθ) =∇ 2 x,θJτ(x, πθ)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:37:49.338025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T11:36:29.670003Z digest=sha256:a25ff3a10a03a1fbeaa54e0e43b20968ba72ec9ab6d007699b517bde64898d64

Pith citing papers

No inbound Pith citation observations are available.