Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:51.203051Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2505.18433.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:51.203051Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 46a68390-421d-49d1-9370-1eac56f40bb0 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef881c85-c3de-4cea-8d50-00d201ca72a2 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning how do I threaten others?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2f8f0faf-52cd-4820-8589-704c567821b1 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning ∥E h δt,l2 (W ∗ t )ψi t,l2 Ft,l1 i − E ˆAdv(s, a; W ∗ t )ψθi t (s, ai) ∥ Ft # =2 ( 1 + γ 1 − γ + 1)rmax + 2ϵcritic · E
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4f2d4312-4f5a-4c2d-b944-14296e41961b · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning On the Global Convergence of Natural Actor-Critic with Two-layer Neural Network Parametrization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 489fd0ce-9c06-4eb2-913e-fbdd859e8013 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning To handle token inputs of variable length, we apply zero-padding to each critic input
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0119de7c-346a-44c3-9f95-bc1992000288 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Consensus-based Decentralized Multi-agent Reinforcement Learning for Random Access Network Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bea09f7e-dc5b-4e1d-a4b5-97fd6b65d0a2 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e4e62907-22b9-41d1-9c2b-62daafb75e67 · outbound
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 58a7b01a-3f0e-4346-9f4e-ff821e6e8c61 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning This shows that µ(·) is the stationary distribution ν(·) under kernel ePπθ and the proof is complete
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 04730589-14b8-4750-a1e4-9422399e38ec · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation feb4dd64-dd48-4c20-9cc1-d061b7f7e99b · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 80b18d47-1fa3-434a-94ef-4d9db26a8185 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Convergent Actor-Critic Algorithms Under Off-Policy Training and Function Approximation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeed37a9-e4e7-42af-9f55-a2ad88a69fba · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning P., Littman, M
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 31e4b1dc-e949-4db6-ae92-ac91011c0ef7 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f8e9cf-80a4-4bcb-9073-2f4fe6170e35 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a0b930-2f27-460e-9304-dc823c6fe3be · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6c4beec6-2389-4943-9809-ad2e33beae03 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Deep Reinforcement Learning framework for Autonomous Driving
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 953aad2d-8336-4120-adc9-57ec058fc692 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dcf521a0-b708-4b1f-8818-2aa335a7b53d · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning TinyLlama: An Open-Source Small Language Model
Reference 748
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4963354b-2683-465d-a8aa-e67098507826 · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cb628c24-4853-460c-a3d4-056fa4cffffd · outbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 4344
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c98cd0-fecb-4ba1-90e3-a3bc7e013b32 · outbound
Reference 9869
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.