Pith. sign in

Paper Citation Record · LEDGER

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2506.01261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01261 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:26.316448Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T11:44:53.211503Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:36:25.635046Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f866580d-3a6a-4ac7-b845-022d9dddbfc3 · outbound

This paper cites Fitted q-iteration in continuous action-space mdps.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Fitted q-iteration in continuous action-space mdps

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:32.085703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:23.122203Z digest=sha256:e53777b00ad2dd85a9278be36d78507d32947b449e2e1715868b144a57ab3c91

Observation 7ad5d9ca-845d-4ea8-8c06-04e49da4417b · outbound

This paper cites S., and Guin, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning S., and Guin, S

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.284017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.284017Z digest=sha256:034c51779dd8c1ad2f241407de622497059b092eb8789726ee551ce0175160a7

Observation c672ef05-d8e8-4300-b67b-2cfaae3b0fc8 · outbound

This paper cites OpenAI Gym.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.368455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.368455Z digest=sha256:557674752524375e79e99a1760bbb40cfdba451de3a87af8e1b07b7704c3169d

Observation c7cacfe5-bcd0-4e18-b0f7-f160f1699d2f · outbound

This paper cites D., and Wang, Z.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning D., and Wang, Z

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.888111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:23.514345Z digest=sha256:8f87b4d7fa2a82bdc34559f9f49e94d91f6bb088504c254bec4e3f9689f70246

Observation c8deb5e1-1c7a-4d43-b690-9ae93ad6d974 · outbound

This paper cites W., Hilton, J., Klimov, O., and Schulman, J.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning W., Hilton, J., Klimov, O., and Schulman, J

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.641212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:23.606041Z digest=sha256:0ee0e6bacb658cd18e7cf1443cfde9c0feead5703a769b10e0c1af5ae0af4427

Observation 300c19b8-a938-4320-b4dd-c41e32cb8cff · outbound

This paper cites Linear off-policy actor-critic.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Linear off-policy actor-critic

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.431771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:23.695877Z digest=sha256:2b4c035d3e27654c192c8b011a1cf839189ff803b82e2a601a502c991360131b

Observation c0ff3ad0-e46a-4700-a7cc-47aed292b5ad · outbound

This paper cites A theoretical analysis of deep q-learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A theoretical analysis of deep q-learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.246319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:23.790666Z digest=sha256:c5dedc416b859a0c849b08250079f12311294c629fa2d770df0d6a898e9f05f8

Observation 3a257f89-2ad1-41ae-9a14-5468bd6c9ba5 · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:31.047177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:23.873294Z digest=sha256:cd347717302abf832be858b64653d351edaed381cba6137c0428e64d12cdf292

Observation 2e2979de-4976-4bf6-87ac-ef722abf575a · outbound

This paper cites Error propagation for approximate policy and value iteration.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Error propagation for approximate policy and value iteration

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.864069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:23.956416Z digest=sha256:29b5776b33d796dbaf6733126ae0223467bb7d45ec5479aa7070aef0bf57fa0a

Observation 8dbfb1cd-09cf-482d-acd3-d80dfdbb919e · outbound

This paper cites Reinforcement learning with deep energy- based policies.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Reinforcement learning with deep energy- based policies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.668546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.009912Z digest=sha256:8deab9560202e2c3ec3ce92571237c49431cda84d743b89c309b1c2c5d9f0579

Observation 358e5bf4-13f5-43e1-9ef0-c75e466949be · outbound

This paper cites Federated reinforcement learning with environment heterogeneity.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated reinforcement learning with environment heterogeneity

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.469003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.059749Z digest=sha256:51b2a91d9ac4e1c2cf5c2f0d1fc664a6de73bb7a9f161272649d674e2e9e8444

Observation 12327596-a491-4a38-8cb4-3196ee54b005 · outbound

This paper cites P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.225430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.138911Z digest=sha256:11f4f708d278d03855240b9a84bc92ac00f26d8760822ac5385331f406861fe5

Observation 66c651a9-28fd-4962-84c3-d6b1b7bb6311 · outbound

This paper cites and Tsitsiklis, J.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Tsitsiklis, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.016740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.191908Z digest=sha256:cf45df55460e2242a34d2d2cafefef555ca2072311e33d80fb41fda3e79f6b03

Observation ba962be4-b9d1-4ccc-86b7-75adc426ddac · outbound

This paper cites K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.781855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.283564Z digest=sha256:f0d69e8f092c1ecfcda3f8196d1dbba8af39c3ca8f6057cefd29d03ee0a42251

Observation a6420aa0-3851-418a-b662-b8c9ab552015 · outbound

This paper cites On the convergence of fedavg on non-iid data.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning On the convergence of fedavg on non-iid data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.555582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.398859Z digest=sha256:01ec7f5f8185bc36307b5f048bb8bf719f6ed291e3b5d24783bdcd0d07021245

Observation 34a3dda1-b350-45d5-bab2-6d89dd56ba48 · outbound

This paper cites Neural trust region/proximal policy optimization attains globally optimal policy.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Neural trust region/proximal policy optimization attains globally optimal policy

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.346041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.532711Z digest=sha256:9f025efddaddb5472a195bf2ca3e8b22552a304b770e2ccbb43169f537c564a5

Observation 65f24c31-4e0e-473a-ab9e-ebeb3b62479b · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:29.152442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.593013Z digest=sha256:4ed93922a1318d21fc2e81c71145299093b204e5a7f6263f682efb67522a7494

Observation 709cd9a4-1602-451f-85c3-1e542b3155fc · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning On the global convergence rates of softmax policy gradient methods

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.968600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.703231Z digest=sha256:84dbc59f4f958bcdbdc9ba56cbdea06a256acd95b784389a2b9275e837badf5e

Observation b5d387ed-4cd7-43e2-8d98-86763ca50cc6 · outbound

This paper cites A., Veness, J., Bellemare, M.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A., Veness, J., Bellemare, M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.764589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.807776Z digest=sha256:937d94feadceb549d84309384ae5d98d9c36112c8716d65994275f814be20462

Observation 29a9c553-d78e-438d-bf62-611abe8ae074 · outbound

This paper cites and Szepesvári, C.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Szepesvári, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.584788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:24.918651Z digest=sha256:e07f88e89a65bdd719e70daa44ac344fc87f7c2f9d4a6dad1a6a9e7c860396a6

Observation 9718f0b3-1123-4839-9bce-329c0def2396 · outbound

This paper cites Planet dump retrieved from https://planet.osm.org.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Planet dump retrieved from https://planet.osm.org

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.421351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:25.028428Z digest=sha256:8db0f4370337faa6a03cae910e735bc935fd9871ed1feb3d027642405fc86c05

Observation 12f480a1-5c6c-4313-b891-fa47167bf3b7 · outbound

This paper cites Trust Region Policy Optimization.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Trust Region Policy Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.098661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.098661Z digest=sha256:89eb492b5ecd22fc787cf9dd65f7eccc596c8a3ee2e4143abe2a0064cfe37713

Observation 9292f581-f4b7-4a79-858f-3cf4e5080240 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.165690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.165690Z digest=sha256:b0ccdaddf5d045349839677049bb6b9707c166e737a3feb5812f2ca76b5e0a2c

Observation 12536ad8-a079-4510-876b-58edb0fc1fb0 · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.208016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:25.271960Z digest=sha256:4756b6da3f1a1261c8f2bf063af729ac95813d51082fee7e1c5ef606c4477738

Observation 54b8813e-da05-4c67-8138-8a79fcc5c049 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning S., McAllester, D., Singh, S., and Mansour, Y

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.978909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:25.369311Z digest=sha256:117ee0c5f055aebe6351c8cfd965eb7d1cf1b285aa04f88792d64f1b04bf6195

Observation fd1f72d0-98e4-4f32-9d04-6696078872f9 · outbound

This paper cites Boosted fitted q-iteration.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Boosted fitted q-iteration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.778136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:25.449432Z digest=sha256:c99291e3e1c0eab94a9141ab0c7f40fd4250289d44b7e30a1d266a06e25cbe54

Observation 5eb2fc9f-e882-4ecd-927b-9ba6abec0656 · outbound

This paper cites Deep reinforcement learning with double q-learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Deep reinforcement learning with double q-learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.514852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.514852Z digest=sha256:fc4208bc096119fcd30471f0602e84c6cdd35cc4b6e686405f7cf778ff401f7b

Observation 014e2b7f-8963-441d-95b9-07541ce6162d · outbound

This paper cites E., Srivastava, S., Tuia, D., and Falcão, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning E., Srivastava, S., Tuia, D., and Falcão, A

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T11:54:26.944638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:25.614183Z digest=sha256:41c9e4e1133a431bb53cb6421f3e8eef564fae48bb0a736d1573499f79d57fd1

Observation 27b574ad-db67-4b56-b86b-4083bd2da289 · outbound

This paper cites L., Kheterpal, N., Jang, K., Wu, C., Wu, F., Liaw, R., Liang, E., and Bayen, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning L., Kheterpal, N., Jang, K., Wu, C., Wu, F., Liaw, R., Liang, E., and Bayen, A

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.583825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:25.722659Z digest=sha256:5a80483986abd7edb945d87ba3282d9d0ae4da3c244f3528a8b989a56c58d250

Observation 935a9aa8-71f1-4b99-8138-d8b4689fdcc5 · outbound

This paper cites Neural Policy Gradient Methods: Global Optimality and Rates of Convergence.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Neural Policy Gradient Methods: Global Optimality and Rates of Convergence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.833975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.833975Z digest=sha256:e564f588e761d90aa13a29d95bbd680df972da0527578a9a16d0e77294264f5d

Observation 98bad912-d5dd-449b-b01d-99871ff1812b · outbound

This paper cites P., and Kakade, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning P., and Kakade, S

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.391913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:25.909375Z digest=sha256:e90e10271fcd75fe3f0be8dcf9841ffa56ee576d6fc20fc4d407e21d1d4935a4

Observation 39660861-fed3-4826-b7ea-26838d4e7be3 · outbound

This paper cites and Song, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Song, S

Reference 32

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T11:54:26.697142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:26.039203Z digest=sha256:0d17588d220e9c44f2f5a1519973ae9f21c6f401675dd1a1d0e9ebb86ae0b8f9

Observation 8b3badc4-4a1d-4f43-83be-6a93508fab2d · outbound

This paper cites A General Approach to Adding Differential Privacy to Iterative Training Procedures.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A General Approach to Adding Differential Privacy to Iterative Training Procedures

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:26.115849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:26.115849Z digest=sha256:2efc77093180c6ab2b1feebb0a0c31c6fa976d19cf5c7bb6018c7f9489b88e95

Observation 852b693b-da84-4ba9-8f6e-ee8146aed99f · outbound

This paper cites Federated natural policy gradient and actor critic methods for multi-task reinforcement learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated natural policy gradient and actor critic methods for multi-task reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.168863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:54:26.220188Z digest=sha256:ebcbd9beb4b71a3963cd5acfa8d7b90643d8b91d18f2669fc1116e29e6b359b1

Observation 722dd00c-e413-424a-b058-98232ef9a157 · outbound

This paper cites Federated Learning with Non-IID Data.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated Learning with Non-IID Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:26.316448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:26.316448Z digest=sha256:f5e8b79c190d1d2be23b4030ba471dfc14e4c997c5aec2b749df6f141c07f4b7

Pith citing papers

Observation b32d63e9-5e95-4638-a392-534a5ecc3069 · inbound

Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic cites this paper.

Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:08:29.340835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:07:23.862809Z digest=sha256:afa8b593db2c99651e48614f81cc7360af6a490cfa190e520ed6d052c6848e77

Observation 1377c879-6903-497f-b020-dba0efa9ead2 · inbound

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data cites this paper.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:36:25.637603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:78d6f949f2543568ae35c1e3b9a5ef9aed8d537d542b00d7025b57ab3e77f6ae