Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T11:36:29.670003Z
Paper Citation Record · LEDGER
As of 25 July 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2601.16399.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T11:36:29.670003Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation adef480a-570d-47bf-a625-58295ecd19ec · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning On the sample complexity bounds in bilevel reinforcement learning.arXiv preprint arXiv:2503.17644
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 89a21b52-0f2a-416b-85f2-8f8b344bb7fe · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Approximation Methods for Bilevel Programming
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 4dae121b-13fd-45aa-af4e-7832f379d756 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Using Synthetic Data to Mitigate Unfairness and Preserve Privacy in Collaborative Machine Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation f9d736c9-7058-4291-8ee5-1deef649af6a · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unlocking Global Optimality in Bilevel Optimization: A Pilot Study
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 832a16cd-8b5a-49e2-b2d3-141cd30ea64f · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning A First-order Generative Bilevel Optimization Framework for Diffusion Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation dbe74d4e-3258-4936-8d81-71e298b05384 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning samples drawn from the stationary distribution, instead of continuously generated Markovian samples
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation d00c2528-db7a-41b6-9dc8-6faead6554b2 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 0e26ac44-9db9-4719-bc4e-1e79df3b33d2 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning We defer the proof of the lemma to Appendix E.15
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation af1f6328-adbb-4abb-afb2-374ce51bcf54 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning The bound onE[ε V,L k+1]can be derived using an identical argument
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 3008e215-0487-45ad-bf9d-82daddfcb7a0 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning dπ′ ρ = 1 2 dπ1 ρ + 1 2 dπ2 ρ .(93) We use ˆdπ ρ to denote the extend discounted visitation distribution over state and action such that ˆdπ ρ(s, a) =d π ρ(s)π(a|s)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation a68be447-d618-411f-a2bf-70406579b392 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 070d11a7-e6c1-4848-8c1b-6be95407178b · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation e3f8a955-3391-455f-aa0e-ecd769f02bdd · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Adapting the result from Shen et al
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 11785769-1754-42ce-9fd0-bbd9cb90af09 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 76feface-f4b8-4422-acea-fe3b4cddcc08 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 00dd948d-1eb9-4bca-894e-6bfa89214a90 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning Therefore, ∇2 τ,θ Jτ(x, πθ) = 1 1−γ ∇θEs∼d πθρ , a∼πθ(·|s)[E(πθ, s)].(126) Zeng et al
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation da26f9de-4db2-459c-95d4-c04607036254 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning We next show the smoothness ofΦ w,τ
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation b3a24f6b-eef7-47aa-8e05-5b18de37ac11 · outbound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation a82a87b7-ce37-474d-90c8-842842db5674 · outbound
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning This implies∇ 2 x,θJ(x, πθ) =∇ 2 x,θJτ(x, πθ)
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
No inbound Pith citation observations are available.