Pith. sign in

Paper Citation Record · LEDGER

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data

As of 5 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2606.03094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03094 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T11:44:53.211503Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact12
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0fc8a006-bb48-4b0d-b3ea-f57ac166742d · outbound

This paper cites Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:b4bf5ea82eaf09ae72c7ccf9f947fdee997a90ed5f427f1eb46828b4484bf0c7

Observation e89c9149-0451-4116-980e-a3433ca75164 · outbound

This paper cites Qwen3-VL Technical Report.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.604790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:7341651058019c640163e56402d8ca07285b49c25c857b8c461738f29d62aa79

Observation a779786c-2f63-4f26-b1d5-a989133ebe6f · outbound

This paper cites Qwen2.5-VL Technical Report.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.601956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:3ac3a3a2a695aeec89cbe946ac02d588b7db4e43af97298e17f1b54921098ed6

Observation a0d50fa7-dac5-4f28-a173-baba1f3522f8 · outbound

This paper cites R1-V: Reinforcing Super Gen- eralization Ability in Vision-Language Models with Less Than $3.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data R1-V: Reinforcing Super Gen- eralization Ability in Vision-Language Models with Less Than $3

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:d08e7bed82cec759becd369f396e7bd51dee28edee9b84c226762271122ac595

Observation 7416aa03-8007-4fb4-a17f-6c2dc8e53bbf · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.608391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:0818ead7430634d4d9d4001f58358ec4cf82a3a54eba25405715025e3c9e2d44

Observation 0d726caa-0f21-4f58-8b51-9130d531a41e · outbound

This paper cites Open-R1-Multimodal: A Fork to Add Multimodal Model Training to Open-R1.https://github.com/EvolvingLMMs-Lab/open-r1-multimodal, 2025.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Open-R1-Multimodal: A Fork to Add Multimodal Model Training to Open-R1.https://github.com/EvolvingLMMs-Lab/open-r1-multimodal, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:23b03a771aa606c41f44adaf7b1178171a54586cfaedbff2d1f20aecf9ddf985

Observation 5974101c-81ec-45ae-a144-ed5b981e3d63 · outbound

This paper cites Fault-Tolerant Federated Reinforcement Learning with Theoretical Guarantee.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Fault-Tolerant Federated Reinforcement Learning with Theoretical Guarantee

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:b42507a91534c6e6edbc8ba0c405619fa29decda563fbc227bdeae0735d7ebbd

Observation b3b2f1de-f8c2-46c7-b9fa-75e52e5f87f9 · outbound

This paper cites FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Per- sonalized RLHF.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Per- sonalized RLHF

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:b749de4380038d774da0cff28a41225b1430b79dcc908323c3e896525d811eeb

Observation ee0e91a9-93e9-4529-b3ff-e6513d3cf512 · outbound

This paper cites Provably Robust Federated Rein- forcement Learning.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Provably Robust Federated Rein- forcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:011307775dfaa03e18df7737a503b6588efd0674679986a00edfaf675e1bbb57

Observation 1b05db06-156e-437d-aa56-7bbeee9a4236 · outbound

This paper cites The Llama 3 Herd of Models.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data The Llama 3 Herd of Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.644592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:679afbb4776809bb0310bd3b9a53c920a39d5b96038f80ca5da33d169c4bf2ab

Observation de057067-a543-4ac3-8bdd-c36a497989a6 · outbound

This paper cites DeepSeek-R1 Incentivizes Reasoning in LLMs through Reinforcement Learning.Nature, 645(8081):633–638, 2025.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data DeepSeek-R1 Incentivizes Reasoning in LLMs through Reinforcement Learning.Nature, 645(8081):633–638, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:b96df53453e8e50c72893c9b93ba1ab3fd4db9a8992a0c6b9b60b82141b12ec0

Observation cdca963c-f562-4e70-993e-4c9dbe1bb52d · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:dc86824230424e7c3d77022efc7531ebb4525ed05db856ccfbbf58d1d824ed76

Observation 613ce512-b219-45a1-998f-5b473b38679a · outbound

This paper cites OpenAI o1 System Card.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data OpenAI o1 System Card

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.630277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:d9135302bafdcb58e7a5ec1721a74922c5cb93a52e964e2901824e6918d2e32f

Observation 0402d5f8-b096-492c-8b93-57f19fff116a · outbound

This paper cites FedHPD: Heterogeneous Federated Reinforcement Learning via Policy Distillation.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data FedHPD: Heterogeneous Federated Reinforcement Learning via Policy Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:52dd27ca6818f45c96ddb4ed03b669f80ceba55f70964ff08508db0d5ec52740

Observation ca2cf0b4-ad49-4277-bbdb-dd7015fb87f5 · outbound

This paper cites Federated Reinforcement Learning with Environment Heterogeneity.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Federated Reinforcement Learning with Environment Heterogeneity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:4da1ea362dd602f3d12ae6226747db5ce9323024c58f416a171fba519646b0f5

Observation 401a60ea-5ab6-49c7-af02-d62849219d02 · outbound

This paper cites Decentralized Federated Policy Gradient with Byzantine Fault-Tolerance and Provably Fast Convergence.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Decentralized Federated Policy Gradient with Byzantine Fault-Tolerance and Provably Fast Convergence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:f908a551c2d158ca49a5e0da225b331aa7c274223a6223c67c6813b0e4d56a43

Observation c316229a-e24f-42fd-afed-482d1d08da9a · outbound

This paper cites Reddi, Sebastian U.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Reddi, Sebastian U

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:9d3611360e31e628a7fb6b2ae1b35645945a61b3e6264a078eeb85d3885cd370

Observation 44ca493a-0f6f-429d-af78-420305798e2c · outbound

This paper cites Federated Reinforce- ment Learning: Linear Speedup Under Markovian Sampling.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Federated Reinforce- ment Learning: Linear Speedup Under Markovian Sampling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:b0b105c63530a80a11ba4f7a62dc8504c02f5dc7e0650a6955ae18c50b8f7843

Observation 82bf0832-ef63-443e-8c11-01985362c274 · outbound

This paper cites Konda and John N.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Konda and John N

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:cdc961838337b247e18c6f133bb2a0f6a1e5f3d747fa436913cfc3c3564cea84

Observation 3ce7a08f-8bd1-4439-85be-5d2d287d9f99 · outbound

This paper cites Asynchronous Federated Reinforcement Learning with Policy Gradient Updates: Algorithm Design and Convergence Analysis.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Asynchronous Federated Reinforcement Learning with Policy Gradient Updates: Algorithm Design and Convergence Analysis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:961b1d9ae37fe09c5c8f94cc25bfdbac166df786434fee45ed26968a3f3af9d4

Observation 8cf678f8-6f16-4867-9e46-297fe616b151 · outbound

This paper cites Federated Optimization in Heterogeneous Networks.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Federated Optimization in Heterogeneous Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:af733bc534112ad865b7c9d7552b36ee76045e7d79a276f633fefba435d5a2a5

Observation 2ca3e619-eea7-4b4e-a3e0-028441013a65 · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.627126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:82883b1e23e82972ddaab8ffa8ecd0a2e94249051855c95b61795e47e19c96dd

Observation bcbd946c-3474-4863-9e84-688380b64805 · outbound

This paper cites Communication-Efficient Learning of Deep Networks from Decentralized Data.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Communication-Efficient Learning of Deep Networks from Decentralized Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:4dcd1c1c9d020ecd49f68374288b959591d7c9d93b1f3b0d105ce7fd73153488

Observation b3fd9138-3dd4-43e3-8f17-f61d06bd8139 · outbound

This paper cites arXiv preprint arXiv:2508.02833 , year=.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data arXiv preprint arXiv:2508.02833 , year=

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.640889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:1bee99ea394bfeaab56f88198af554c086471575bc3152338e6e54f43035abd4

Observation b9851825-ceeb-4f5f-8c5d-7d1ac1b4c0ca · outbound

This paper cites Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Koneˇcný, Sanjiv Kumar, and Hugh Brendan McMahan.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Koneˇcný, Sanjiv Kumar, and Hugh Brendan McMahan

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:6a85a88a3c129c58377ff5bfb073b122afbece9c29d87e2c88ecb0af7f8547e2

Observation 9616074f-fa5a-4ffa-a361-14939d56222a · outbound

This paper cites Federated Ensemble-Directed Offline Reinforcement Learning.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Federated Ensemble-Directed Offline Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:ef75c02f238ccdd1e11a62bb15c1f7e2da1e95660e9cf5b4e2e2b0b57a0811cd

Observation 931c6ec8-8f3b-4f51-9162-7e28cfb80ed3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Proximal Policy Optimization Algorithms

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.617700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:193b6feadc9183cf4191f5f37c2deb0d7b7e0fe77dd760237d5375121e36599f

Observation 189e0812-88db-4cda-adf6-e3750250d072 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.623448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:8ac9157298abe578c20a36dd8b2a49aae06dcf72d8ca7775da9f119d8288d2e5

Observation 895a0291-6346-404b-9ddf-b27ce5ec45cc · outbound

This paper cites Momentum for the Win: Collaborative Federated Reinforcement Learning across Heterogeneous Environments.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Momentum for the Win: Collaborative Federated Reinforcement Learning across Heterogeneous Environments

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:40dd42d0c51143835f10815d9c1a56744b0006551a902e03b87e4b4ed09a7827

Observation 6cc4499f-e160-4191-b3a6-063c61465d71 · outbound

This paper cites The Blessing of Heterogeneity in Federated Q-Learning: Linear Speedup and Beyond.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data The Blessing of Heterogeneity in Federated Q-Learning: Linear Speedup and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:aa6a7fb96489ae0b2c6a4ec5fe852c50e231b9a17db1df2510206cee0defe0e4

Observation 2473d4ac-4869-423a-893c-51a8a9695d1d · outbound

This paper cites Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage Suffices.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage Suffices

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:e71c85448e6b2387259fe877bfb38beb94d3f0d7e7bf2cfb540fbb324c919f3b

Observation 1377c879-6903-497f-b020-dba0efa9ead2 · outbound

This paper cites The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:36:25.637603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:a43b8b5970969136b82c976bce87ab884f47d57221b504d28123a97c04207ab1

Observation 4a966b9b-9663-4304-8db4-c2506a92a7c0 · outbound

This paper cites On the Linear Speedup of Person- alized Federated Reinforcement Learning with Shared Representations.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data On the Linear Speedup of Person- alized Federated Reinforcement Learning with Shared Representations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:f998ac84c8d598e5ba884c297835e3086b18e49a67a9c22bf084618b51f0d689

Observation 2b672abb-9b7b-49e9-9ccf-eaac0810f9d7 · outbound

This paper cites Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:f8f67419546f49a637308fa59e0963e94a370b641ac8074ef745b53250facb7c

Observation 2075e713-5ec0-4798-a7ec-baea2701e6bb · outbound

This paper cites On Classes of Summable Functions and Their Fourier Series.Proc.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data On Classes of Summable Functions and Their Fourier Series.Proc

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:fa9b16e82ff87432b1c450bd772740449753f6ef01d0efd36a6c794a7a33e0c4

Observation 6227c12d-1497-40c1-9404-c6d1676e7e24 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.620697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:d7ae931f6e689be662c9e46530776522d49f096afa746b55446577cc36d117e0

Observation c3def03c-36a7-4de1-85bd-9f66f43636c5 · outbound

This paper cites Finite-Time Analysis of On- Policy Heterogeneous Federated Reinforcement Learning.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Finite-Time Analysis of On- Policy Heterogeneous Federated Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:e9945e9a1e8421663d331dbe03dd21ae4300e1826b0f756d809f7ead878afd73

Observation 163796d0-0c2b-42ce-ad3d-b44af6b9d091 · outbound

This paper cites and Zuo, C.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data and Zuo, C

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:36:25.611991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:d7a9467a890626ccac46a7ac255a77585882c6a5c2f4960ebcd4f797135f5e70

Observation e6979688-09c7-4088-a6f5-3712c954d57f · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data A Survey of Reinforcement Learning for Large Reasoning Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.614991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:f80a8b35fef260856f2f1aa03c6780b59caf430d3204d93535c19a0abde3ab34

Observation fec8d277-6c01-42bb-b628-d45fa8947f96 · outbound

This paper cites Group Sequence Policy Optimization.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Group Sequence Policy Optimization

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.633440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:6b39793bab7ebaa6be8fec1f454f2f4623f6a6b229fdde2c0ec729dcf6f6f164

Observation 4a5abb1e-3902-4d64-8583-86407890b6eb · outbound

This paper cites Recent theoretical research has focused on establishing rigorous convergence guarantees under the unique constraints of sequential decision-making.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Recent theoretical research has focused on establishing rigorous convergence guarantees under the unique constraints of sequential decision-making

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:7dcbd86c2208c56dc79ef735615b033df728e7a6e3c2e146593c15f1110d4c39

Observation a2661299-5acd-46d9-8ffe-5770b2399223 · outbound

This paper cites collaborative single-policy coverage.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data collaborative single-policy coverage

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:d0be93cf3c5c3dade3880366ab7bc04f5af358df681e78ed2f22dcfbdc148b25

Observation f5ef04ab-4657-42fa-bf47-c4803efb1d77 · outbound

This paper cites Similarly, DAPO [36] introduces a decoupled and dynamic sampling system designed to stabilize long Chain-of-Thought (CoT) reasoning.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data Similarly, DAPO [36] introduces a decoupled and dynamic sampling system designed to stabilize long Chain-of-Thought (CoT) reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T11:44:53.211503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:8b21b8e116c8dc4f09765af1440a49a6cf590c99f4397d0c1cfcc07ee327272e

Pith citing papers

No inbound Pith citation observations are available.