Pith. sign in

Paper Citation Record · LEDGER

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

As of 17 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2509.08721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08721 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:11:38.252026Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T07:40:35.113858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T07:41:14.649759Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04965135-5b5a-4c16-b64c-7e8a866b9a66 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.023666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.023666Z digest=sha256:cf2cf84a7cc89f41f5a1d8606ba7f7774448435afdd380aa9bcf6e7413e0c61f

Observation 28f31821-5267-445e-be6d-d62de77433d4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.029405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.029405Z digest=sha256:4d37b06586ed37fd7b6adc8d0144b7c62bd49d00488777a422273e63c6fbd025

Observation 93dd3fd5-e281-4625-bf24-5b7927551f3a · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.035959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.035959Z digest=sha256:f92c38ea8041c8080700e906dd0f3da9c6c73e04435b8b6a272459b8a9aef2ca

Observation e6c60293-83df-4036-8035-987740f8026c · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.042144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.042144Z digest=sha256:b5d0d5611d07d5a4efd9ff8a320e9eb87f4ff8ef10c5f7ca785eb242cf7f3adc

Observation d6810192-2a98-4fcb-ada7-0e1d88d533be · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.047737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.047737Z digest=sha256:08d6b5fdee6cdf4e7704799acaec25be6f036817e1d48e5e0af5f48e8128c187

Observation 0dc77fef-8e5d-4786-a540-594480ad10e0 · outbound

This paper cites Introducing rl swarm’s new backend: Genrl.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Introducing rl swarm’s new backend: Genrl

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.271650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.053251Z digest=sha256:c2447a83cea1e85fe111b6ca4a2f012323fc9e4e05a1ccc8fece4dd3eafcdd5c

Observation 269bfd1b-dbed-46e9-81af-3aacf9f19b3d · outbound

This paper cites Gensyn rl swarm.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Gensyn rl swarm

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.255993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.058693Z digest=sha256:13e1ca7525382e2a19c44f9ac7e893c5a76da4c9468e683293a409e5d6ee0898

Observation d71e7bb5-96ed-4c46-9e88-5fcde5ef24ec · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.064488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.064488Z digest=sha256:028487f13cef561874689920c8384c210532e4033ea573d0a498e38d6271fcc4

Observation a46e0d65-e37e-4aa3-8edc-0ecfa851d540 · outbound

This paper cites Bowman, Tim Rockt\" a schel, and Ethan Perez.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Bowman, Tim Rockt\" a schel, and Ethan Perez

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.239713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.071245Z digest=sha256:4d3b8b1a4265f288c71a62e2f46edd50ce0fc64dc83464af3751fcd5d016363f

Observation 0c69c874-c4a3-4b7e-9290-54a0a856b72d · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.076287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.076287Z digest=sha256:e5bf1c30eb1959a386ad99c37778f37748cd82b1ad7d4caa84cdfa23eac41ecc

Observation eae964b9-d577-476e-948f-b091cd0b449e · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Coderl: Mastering code generation through pretrained models and deep reinforcement learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.082881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.082881Z digest=sha256:6cfff1959aba1f67c9a0cc46822005ba2bf5e4fba7c5fcaaa0eff32f880df81e

Observation 704f084e-e67f-4ce9-a5ef-2fca2056a526 · outbound

This paper cites Camel: Communicative agents for "mind" exploration of large language model society.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Camel: Communicative agents for "mind" exploration of large language model society

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.223623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.087984Z digest=sha256:2757cac94d3403e36cf9bc4d9893c300cddcdbc84bfacdbc5a8e0ba5d02c2899

Observation 5be3c361-0515-4534-9677-6fc7350b755b · outbound

This paper cites Improving Multi-Agent Debate with Sparse Communication Topology.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Improving Multi-Agent Debate with Sparse Communication Topology

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.093375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.093375Z digest=sha256:3073443deb88c91ead10f5db605a97ab7af6a93e25e3dbab074267252499d793

Observation ef425cb0-e61f-4267-b996-c693242287e0 · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.098724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.098724Z digest=sha256:8bdc88dbe6f315c5b504ca7788007c364fef22b5687bd2cb6d01fd4588f354ef

Observation 42e39dd7-55d2-400f-aab7-650c5e6df32d · outbound

This paper cites MARFT: Multi-Agent Reinforcement Fine-Tuning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.103829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.103829Z digest=sha256:8f9a205cfa4cfafcbe07f3ee17747437e137de3d3e0d9e12b72206701245f8fa

Observation b55754cc-3787-4499-a329-4b63d9320414 · outbound

This paper cites Llm collaboration with multi-agent reinforcement learning, 2025.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Llm collaboration with multi-agent reinforcement learning, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.109179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.109179Z digest=sha256:e84142d54f748344fa7051ff45227f334c626e278f52f9f71ab5e63123ca9bf3

Observation fa691390-2f9f-4629-b252-83b5634efa0d · outbound

This paper cites Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.206463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.116384Z digest=sha256:d84e6e64f7eb368b6b29e4e3f9893b47260cb265dd18ed6a20fe86c333caca03

Observation ba52a9bb-f820-45e3-a236-89c70f7539ee · outbound

This paper cites Magistral.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Magistral

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.122268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.122268Z digest=sha256:89e0e7de3fc91ec349d56cc62cc2ff88fc0cf9c120633ce07b474a5d62b8ad4f

Observation 3afbe756-0010-4c42-9b30-c292b97b8845 · outbound

This paper cites an unresolved cited work.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.127762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.127762Z digest=sha256:904a224409e160af35d46d0f9cb8b36f1214a9ba8fbf77208cdc59ffd634cc2d

Observation d70bd771-4080-42a6-aafc-d4ffcc97d08f · outbound

This paper cites A Survey of Small Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing A Survey of Small Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.132668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.132668Z digest=sha256:cb4a0f617cbc133af3df9ef8a0a902984abd1d086392530d67b638488a27dc29

Observation a6475cc0-fe77-43e3-9e92-4992aa9c8d6d · outbound

This paper cites Aligning language models to follow instructions.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Aligning language models to follow instructions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.137824Z digest=sha256:86bd399f8520ae2ac06f3a188e84809528894bc34208ec02a62bff4487196f1e

Observation b1967b71-42e4-4d3f-8704-939d8f2cd846 · outbound

This paper cites Learning to reason with llms.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Learning to reason with llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.160352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.143047Z digest=sha256:a7431c7b08160790743d58a6dabced66f9ec27957b1b3a94100bf1b8e14c63bd

Observation 827aaa37-c583-4756-b3cb-817dc88ce216 · outbound

This paper cites Training language models to follow instructions with human feedback.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.147977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.147977Z digest=sha256:a598884977bf98f1bbf9a4f154e955d3667ef8384674076f33296395467bce3a

Observation ad40bdd9-db83-4635-ab1b-fb9e691d24f2 · outbound

This paper cites MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.152836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.152836Z digest=sha256:5fc8aadc0e325505fc20f77b49f158e4a50f37f96f17baee68e2c413cea5ad6e

Observation d045c633-4a48-4240-b5f8-3e00400c50c2 · outbound

This paper cites Red Teaming Language Models with Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Red Teaming Language Models with Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.158686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.158686Z digest=sha256:c74eaf2195f417b21e6c210e918a889087d80a24eae2f9dfcce933915a1fa84f

Observation 5a97264f-265c-4f92-852d-c6e3ed522a5a · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Qwen2.5: A party of foundation models, September 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.164045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.164045Z digest=sha256:824cc186b6281d4361acb906ae0e3ad771b30bcc76039933e69d2dae0d742a4a

Observation 8febbaa1-a426-41cb-9ff9-5e68eb4c2459 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.169427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.169427Z digest=sha256:cbae64222983369f99e89f3b69eedcb19502ddc813908037c9e115d4250e13e1

Observation 9d47a817-b740-4b5b-9a05-0c50a2b5f7ed · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.175033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.175033Z digest=sha256:d6ef2995a282ce2f0cb23c4c4f8e8af8c8036cbfb6536dc87917bed9a97349f1

Observation 21f9b2df-4f6a-43d8-9b18-9758b989a4a4 · outbound

This paper cites Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.180073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.180073Z digest=sha256:c20b01d8196f300b72f39c49cab782adad16115b61a3ca5f5bc30e1801270002

Observation b6b384e7-e0b4-4da2-ae8b-8f5aa93894fc · outbound

This paper cites Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.192344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.192344Z digest=sha256:70e4c4defcdca94963b6e1a5a325cccde905160826ef5853f23ab97eb7820ce1

Observation 1add689b-884f-42f3-88f8-c32f9affdb43 · outbound

This paper cites Fine-tuning language models for factuality.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Fine-tuning language models for factuality

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.119425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.197585Z digest=sha256:62779486bc2331e593cedd42a00a6f2c11bf206fe202088f7000e2d7379bd2cd

Observation 875b28fa-e440-4d6d-bbee-76b50c82aa86 · outbound

This paper cites LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.202370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.202370Z digest=sha256:d0d9d6afec931acd6cbd12a3c346598e80909357788ffa0754bb0694d3dd446b

Observation 83f1c6ee-957a-4375-85b5-6dba009facdf · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.207162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.207162Z digest=sha256:e2e48e131a80ea57026c0cb4142bfbc91144bc604e744c6215c4571582420344

Observation 9fb83e6f-e9ff-4ebf-83be-ad000124da5c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.212687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.212687Z digest=sha256:e3dea0c0573bc85dbd5a2a3ee8d6ea6860b623d5e6cfc69ace7447b2effdc8d0

Observation 43ee3926-cce7-4f2f-82a2-6936c1ba95ef · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.218726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.218726Z digest=sha256:bbc891e3391cfa8880e846174677e680dac698342aa2ad69d7ffc2c7f658f3af

Observation 151856ee-cfed-4593-950e-6a8c2fb8bf71 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.223789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.223789Z digest=sha256:169e06a274fda29579950032414cf5afddc828a6df3f17c140223f5df632e6bd

Observation 2d5ad3ed-b20d-409a-9eed-2dff7c548699 · outbound

This paper cites SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.229066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.229066Z digest=sha256:8372a73b8f3389ac1e39e3cc0ac3b0d8e72e3afdc89d1dd3cf12be2578e5117c

Observation 324bde60-eac4-4c01-8fcc-467217f76eda · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Fine-Tuning Language Models from Human Preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.234186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.234186Z digest=sha256:00f80f5891368c80cbb8857094a654c88af0dccd94f5fe4bad7c5e7c5a15a6d0

Observation cc2580f9-6482-49b6-8f05-4750c704c76d · outbound

This paper cites @esa (Ref.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing @esa (Ref

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.239822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.239822Z digest=sha256:e2050dbddf85d463a96591f1736f52356987b818adc55efec02a342e34d83206

Observation c599f75f-d57a-4a96-a53b-4c927e5bde66 · outbound

This paper cites an unresolved cited work.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.245234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.245234Z digest=sha256:13bb20cf323abdb28c225a6c0cf9afae2e1e1dd103a793eefdc70343522d5b02

Observation a5f4ac86-69ee-4962-8a27-59534893ae92 · outbound

This paper cites HDEE: Heterogeneous Domain Expert Ensemble.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing HDEE: Heterogeneous Domain Expert Ensemble

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:11:38.305753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.252026Z digest=sha256:e2fcd658c104c27cf1eec022160a07249274a3d12f7f88a5f0bf8efeddcfe570

Pith citing papers

Observation eedf2711-0005-4d73-b17b-6169657a0b6f · inbound

Backdoor Attacks on Decentralised Post-Training cites this paper.

Backdoor Attacks on Decentralised Post-Training Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:23:26.221641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T23:23:22.658691Z digest=sha256:3794b441a248dc0120a06d73fd12e2d74a0f62f7a610f443838c74ca88d8c4f8

Observation 2d9cc47d-bec0-4e20-8335-f2074985d647 · inbound

F-TIS: Harnessing Diverse Models in Collaborative GRPO cites this paper.

F-TIS: Harnessing Diverse Models in Collaborative GRPO Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:41:14.654011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T07:40:35.113858Z digest=sha256:55bda8e233f46aba607b80c5b44c0a195647f0b2f8f13e40e01f2879bdf9213d