Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:11:38.252026Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2509.08721.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:11:38.252026Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-22T07:40:35.113858Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T07:41:14.649759Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 04965135-5b5a-4c16-b64c-7e8a866b9a66 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f31821-5267-445e-be6d-d62de77433d4 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93dd3fd5-e281-4625-bf24-5b7927551f3a · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Improving Factuality and Reasoning in Language Models through Multiagent Debate
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c60293-83df-4036-8035-987740f8026c · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6810192-2a98-4fcb-ada7-0e1d88d533be · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing On Designing Effective RL Reward at Training Time for LLM Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc77fef-8e5d-4786-a540-594480ad10e0 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Introducing rl swarm’s new backend: Genrl
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 269bfd1b-dbed-46e9-81af-3aacf9f19b3d · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Gensyn rl swarm
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d71e7bb5-96ed-4c46-9e88-5fcde5ef24ec · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46e0d65-e37e-4aa3-8edc-0ecfa851d540 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Bowman, Tim Rockt\" a schel, and Ethan Perez
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0c69c874-c4a3-4b7e-9290-54a0a856b72d · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae964b9-d577-476e-948f-b091cd0b449e · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 704f084e-e67f-4ce9-a5ef-2fca2056a526 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Camel: Communicative agents for "mind" exploration of large language model society
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5be3c361-0515-4534-9677-6fc7350b755b · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Improving Multi-Agent Debate with Sparse Communication Topology
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef425cb0-e61f-4267-b996-c693242287e0 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e39dd7-55d2-400f-aab7-650c5e6df32d · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MARFT: Multi-Agent Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b55754cc-3787-4499-a329-4b63d9320414 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Llm collaboration with multi-agent reinforcement learning, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa691390-2f9f-4629-b252-83b5634efa0d · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ba52a9bb-f820-45e3-a236-89c70f7539ee · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Magistral
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3afbe756-0010-4c42-9b30-c292b97b8845 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70bd771-4080-42a6-aafc-d4ffcc97d08f · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing A Survey of Small Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6475cc0-fe77-43e3-9e92-4992aa9c8d6d · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Aligning language models to follow instructions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b1967b71-42e4-4d3f-8704-939d8f2cd846 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Learning to reason with llms
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 827aaa37-c583-4756-b3cb-817dc88ce216 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad40bdd9-db83-4635-ab1b-fb9e691d24f2 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d045c633-4a48-4240-b5f8-3e00400c50c2 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Red Teaming Language Models with Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a97264f-265c-4f92-852d-c6e3ed522a5a · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Qwen2.5: A party of foundation models, September 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8febbaa1-a426-41cb-9ff9-5e68eb4c2459 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d47a817-b740-4b5b-9a05-0c50a2b5f7ed · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f9b2df-4f6a-43d8-9b18-9758b989a4a4 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b384e7-e0b4-4da2-ae8b-8f5aa93894fc · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1add689b-884f-42f3-88f8-c32f9affdb43 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Fine-tuning language models for factuality
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 875b28fa-e440-4d6d-bbee-76b50c82aa86 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f1c6ee-957a-4375-85b5-6dba009facdf · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fb83e6f-e9ff-4ebf-83be-ad000124da5c · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ee3926-cce7-4f2f-82a2-6936c1ba95ef · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151856ee-cfed-4593-950e-6a8c2fb8bf71 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5ad3ed-b20d-409a-9eed-2dff7c548699 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324bde60-eac4-4c01-8fcc-467217f76eda · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Fine-Tuning Language Models from Human Preferences
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc2580f9-6482-49b6-8f05-4750c704c76d · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing @esa (Ref
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c599f75f-d57a-4a96-a53b-4c927e5bde66 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f4ac86-69ee-4962-8a27-59534893ae92 · outbound
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing HDEE: Heterogeneous Domain Expert Ensemble
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eedf2711-0005-4d73-b17b-6169657a0b6f · inbound
Backdoor Attacks on Decentralised Post-Training Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2d9cc47d-bec0-4e20-8335-f2074985d647 · inbound
F-TIS: Harnessing Diverse Models in Collaborative GRPO Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.