Pith. sign in

Paper Citation Record · LEDGER

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2505.18433.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18433 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:51.203051Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46a68390-421d-49d1-9370-1eac56f40bb0 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.595626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.595626Z digest=sha256:6c16132b6d58296a94b0c678a051c6b2e136669317ee3184b46f11ee4bbdb95e

Observation ef881c85-c3de-4cea-8d50-00d201ca72a2 · outbound

This paper cites how do I threaten others?.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning how do I threaten others?

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.871455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:51.203051Z digest=sha256:113f260e7a131d02b7782f12e946aefd53ed4eaa8995c7c6de1aef9d8d7b99c4

Observation 2f8f0faf-52cd-4820-8589-704c567821b1 · outbound

This paper cites ∥E h δt,l2 (W ∗ t )ψi t,l2 Ft,l1 i − E ˆAdv(s, a; W ∗ t )ψθi t (s, ai) ∥ Ft # =2 ( 1 + γ 1 − γ + 1)rmax + 2ϵcritic · E.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning ∥E h δt,l2 (W ∗ t )ψi t,l2 Ft,l1 i − E ˆAdv(s, a; W ∗ t )ψθi t (s, ai) ∥ Ft # =2 ( 1 + γ 1 − γ + 1)rmax + 2ϵcritic · E

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:52.462204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.788819Z digest=sha256:144e2fd31c07a2c53b45af9c110ba36bac765226ff69f6649161ea34db1afa07

Observation 4f2d4312-4f5a-4c2d-b944-14296e41961b · outbound

This paper cites On the Global Convergence of Natural Actor-Critic with Two-layer Neural Network Parametrization.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning On the Global Convergence of Natural Actor-Critic with Two-layer Neural Network Parametrization

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:37:51.686242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:49.978040Z digest=sha256:c05c563c1a64ebcb4c84295c03ed5b25da1c9fc83e19106f220073ed217bbc6f

Observation 489fd0ce-9c06-4eb2-913e-fbdd859e8013 · outbound

This paper cites To handle token inputs of variable length, we apply zero-padding to each critic input.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning To handle token inputs of variable length, we apply zero-padding to each critic input

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:52.026075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:51.107253Z digest=sha256:90153ab098e0566e081f99df5ec05c23dec347a89bc2d3432a0054a8a518b6d6

Observation 0119de7c-346a-44c3-9f95-bc1992000288 · outbound

This paper cites Consensus-based Decentralized Multi-agent Reinforcement Learning for Random Access Network Optimization.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Consensus-based Decentralized Multi-agent Reinforcement Learning for Random Access Network Optimization

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:37:51.427489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.157860Z digest=sha256:88ae9151396d4389351cdd868dd3dd750feea86d00154250a3613be1a0a91bee

Observation bea09f7e-dc5b-4e1d-a4b5-97fd6b65d0a2 · outbound

This paper cites an unresolved cited work.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:53.176273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.312276Z digest=sha256:21098cfb9cb4818b72d1da08e8408e880329622bd0235264578c9bc252fdc0e1

Observation e4e62907-22b9-41d1-9c2b-62daafb75e67 · outbound

This paper cites and Hu, M.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning and Hu, M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:53.033616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.386417Z digest=sha256:88136d2082d7523db6aa0855c3d444e2fdc743f7ced5233b64d6a30d798aab74

Observation 58a7b01a-3f0e-4346-9f4e-ff821e6e8c61 · outbound

This paper cites This shows that µ(·) is the stationary distribution ν(·) under kernel ePπθ and the proof is complete.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning This shows that µ(·) is the stationary distribution ν(·) under kernel ePπθ and the proof is complete

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:52.636746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.713670Z digest=sha256:cd17c53627a325a0cb24e8e6b3484f6fdb004f575b78262b0f0e9e7b26f0d0d0

Observation 04730589-14b8-4750-a1e4-9422399e38ec · outbound

This paper cites an unresolved cited work.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:52.369370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.903345Z digest=sha256:1a48d2a99208daca82892a619f3f211dcaaa407bb8e31595699d4d0d22fab85f

Observation feb4dd64-dd48-4c20-9cc1-d061b7f7e99b · outbound

This paper cites an unresolved cited work.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:53.493998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:49.880124Z digest=sha256:31ca43e693507aa0ae4ccaec07843b6e7dc39358b251ff29b20d210b5ed602f5

Observation 80b18d47-1fa3-434a-94ef-4d9db26a8185 · outbound

This paper cites Convergent Actor-Critic Algorithms Under Off-Policy Training and Function Approximation.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Convergent Actor-Critic Algorithms Under Off-Policy Training and Function Approximation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.128557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.128557Z digest=sha256:7e1fbbe87406802a7b2d4ed743a9518356afdbc74ee127f1da58dccfbb4e16f6

Observation eeed37a9-e4e7-42af-9f55-a2ad88a69fba · outbound

This paper cites P., Littman, M.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning P., Littman, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:53.324972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.067698Z digest=sha256:563c5315c2035f8e97ee64a5c9fc19db26ac34b81a2132941a4f269d39a9ac9d

Observation 31e4b1dc-e949-4db6-ae92-ac91011c0ef7 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.719868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.719868Z digest=sha256:c8e400a81e6e47ff1f9c6bd6bd546db36ed772ae9569abea13a70c201a1242d2

Observation 35f8e9cf-80a4-4bcb-9073-2f4fe6170e35 · outbound

This paper cites Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.547041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.547041Z digest=sha256:ee33277f3385a02d76cd44f41ddec93cc33c5f00794d5abe80c5db82ffbce797

Observation 39a0b930-2f27-460e-9304-dc823c6fe3be · outbound

This paper cites an unresolved cited work.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:52.834453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:50.672319Z digest=sha256:1da9e57f85a3c1b5ed56857b5becd227a79a620773bd3fbc468c322d4d18ef7b

Observation 6c4beec6-2389-4943-9809-ad2e33beae03 · outbound

This paper cites Deep Reinforcement Learning framework for Autonomous Driving.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Deep Reinforcement Learning framework for Autonomous Driving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.230383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.230383Z digest=sha256:d333b1be20f6ecc00ba67cee5630fdb56a67ba26c6b5dfe5c8fecab146c60eef

Observation 953aad2d-8336-4120-adc9-57ec058fc692 · outbound

This paper cites an unresolved cited work.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:53.866601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:49.649145Z digest=sha256:93e16f5cfa7b4726cfdb2327d1be93861c76c4a40314ac5cc0e0fa390808226e

Observation dcf521a0-b708-4b1f-8818-2aa335a7b53d · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning TinyLlama: An Open-Source Small Language Model

Reference 748

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.634794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.634794Z digest=sha256:d7061d00145870e9d61db68154a0b22776ca1cf285f8971f749d0baee10c1082

Observation 4963354b-2683-465d-a8aa-e67098507826 · outbound

This paper cites an unresolved cited work.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Unresolved cited work

Reference 2017

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:52.183943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:51.000974Z digest=sha256:3a0f85d171053176c2fc35fdbdcd329a3975895bdae7153e8e2b53423791a769

Observation cb628c24-4853-460c-a3d4-056fa4cffffd · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 4344

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.431221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.431221Z digest=sha256:751b811ea65c0f3232f38633e1b703650ab7f5d6553fc9cce3b7da331260c837

Observation a9c98cd0-fecb-4ba1-90e3-a3bc7e013b32 · outbound

This paper cites Feriani, A.

Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Feriani, A

Reference 9869

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:53.684537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:37:49.803146Z digest=sha256:5491173fbf8cf7e084210ba38ac3297a35287f426083ac528242eae6c1e22cee

Pith citing papers

No inbound Pith citation observations are available.