Pith. sign in

Paper Citation Record · LEDGER

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training

As of 23 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 1 inbound Pith citation observation for arXiv:2601.12784.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.12784 v2

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:25:56.310135Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T08:35:16.435272Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:58:08.764089Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b8c4b7e-93e3-48ba-b5b7-468fa1717476 · outbound

This paper cites Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:45.902248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:45.902248Z digest=sha256:3c44c68425a1fe917b5a529cc8ec9c720e049667b48058086d7aa966c82a6142

Observation 907c093c-b1ae-4d7f-8fc0-47545e2e0262 · outbound

This paper cites LongAlign: A Recipe for Long Context Alignment of Large Language Models.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training LongAlign: A Recipe for Long Context Alignment of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.082302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.082302Z digest=sha256:421d29aa0af534ef4ea68b4129575675515247e688085b898cd40f52faf45a72

Observation 50e8d2b0-f64a-4992-97c6-e3d020b4a30e · outbound

This paper cites A survey on mixture of experts in large language models.IEEE Transactions on Knowledge and Data Engineering, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training A survey on mixture of experts in large language models.IEEE Transactions on Knowledge and Data Engineering, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.235143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.235143Z digest=sha256:c6e2fb6f795d3d9cc82715fba1e63e9e95f0bbe1fe1d022a2cc9672fd7c2c53a

Observation 5f8c24d9-c3b9-4544-8674-9d055a58ceda · outbound

This paper cites Respec: Towards optimizing speculative decoding in reinforcement learning systems, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Respec: Towards optimizing speculative decoding in reinforcement learning systems, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.402975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.402975Z digest=sha256:f7d1f11538d2cd06485a7f5e00ff19447630c0e11e541bf35ac3cc0e89648eea

Observation d972a8dc-1106-4639-b004-12dd92cb13a4 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.529447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.529447Z digest=sha256:263559bd756fa757e49e4ee64945db664a4bddba034feedfeb77c15ee5c5ded2

Observation aa62aa5c-4434-43ff-90ea-d3e69ecc663c · outbound

This paper cites UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.677712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.677712Z digest=sha256:b3603ff4e54cb233b182069efc955558b8a5a578ccfb8d466ab6aa07cacaca32

Observation 225f3113-8c15-4110-8143-e0388d866bc8 · outbound

This paper cites Advances in importance sampling.Wiley Stat- sRef: Statistics Reference Online, 2021.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Advances in importance sampling.Wiley Stat- sRef: Statistics Reference Online, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.835994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.835994Z digest=sha256:8f5da91e7c144343ebe851ddca90415ec2b8f0d2c8e023f642a15bb4b8014886

Observation fc8debaf-3b4e-4486-a0b3-32dd89500f20 · outbound

This paper cites Aime24 dataset, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Aime24 dataset, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.009776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.009776Z digest=sha256:d7abc3f144e6af620eb2be20aab03aa1a019480312f482c1729158386cd66766

Observation a17c8c52-63a4-41c4-aa48-bff89400eafb · outbound

This paper cites Dapo-math-17k dataset, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Dapo-math-17k dataset, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.120583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.120583Z digest=sha256:1fe2ce3e9b9f1c316915c2f1ee26b70453ec74daf756a0d4b2888380544dd076

Observation 7d2d5ef5-381b-41af-b18b-44f0ae5fc23f · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.347729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.347729Z digest=sha256:4d9c39998c9cd764913c2a6f256c9415db8e95da416e4ab732eae85543578dbb

Observation 38d80a14-fb98-42a9-96b4-219a77a70d4f · outbound

This paper cites Apt-serve: Adaptive request scheduling on hybrid cache for scalable llm inference serving.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Apt-serve: Adaptive request scheduling on hybrid cache for scalable llm inference serving.Proc

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.508488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.508488Z digest=sha256:c37a663293e21e365549d490b3ed294a0cf9c7bb3ca2da65d9634ae66aeffc0b

Observation 6dc25872-f2f3-4214-a7fe-0dfde2e71a0a · outbound

This paper cites Rollpacker: Mitigating long-tail rollouts for fast, synchronous rl post-training, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Rollpacker: Mitigating long-tail rollouts for fast, synchronous rl post-training, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.514874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.514874Z digest=sha256:6dcff17cb045cd90791318163f5c563b5fb6977d3a9283d23b70d1e210a6c741

Observation 6a730f45-6758-40d8-a961-4c2314fea167 · outbound

This paper cites Enabling parallelism hot switching for efficient training of large language models.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Enabling parallelism hot switching for efficient training of large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.520984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.520984Z digest=sha256:7c4f1a16b7d2925af0c08947a3d61b1a66abd5dd3ddb98af0c9059c9517c7a29

Observation 5eebe9cf-6b28-48fb-8597-d82a73133610 · outbound

This paper cites Search and score-based waterfall auction optimization.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Search and score-based waterfall auction optimization

Reference 14

Resolution
verified exact
doi, observed 2026-08-04T06:28:25.768871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-04T06:25:47.562419Z digest=sha256:75435e8556bc47d90deb68ff9875b94e744a8cc969c17bda871b541982fb6058

Observation b50643ea-e6e5-4f20-be01-0a8656750e66 · outbound

This paper cites AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.686864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.686864Z digest=sha256:4246537524a047c84c9665fc7171328fa28005fb18c0f97e5b11dd0e071647ea

Observation a46dc5c9-5719-4b7b-b5d7-0787bff3f378 · outbound

This paper cites History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.766365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.766365Z digest=sha256:9a9d066bc88fb9e21f31e3dcdebd1b65d50953f9c223aa2b9f8f7b5af05d2fde

Observation cfdb1be1-1e9b-42bd-baf7-83a32845d875 · outbound

This paper cites Verl recipe: Fully async policy trainer, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Verl recipe: Fully async policy trainer, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.924312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.924312Z digest=sha256:cad8e5d7ec9988132ba0cc811191bb14cfe3003bc858b0eb134c552eb5e887dc

Observation 1d2fe562-2015-4b76-b712-e7ba8f7aa53f · outbound

This paper cites Verl recipe: One step off policy async trainer,.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Verl recipe: One step off policy async trainer,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:48.129022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:48.129022Z digest=sha256:9c85d7cb9b468e49dea795bbce89934895d25c1d139db1e5f6686f06651c7782

Observation e53b63b9-fc9c-4724-8b3f-55c1464a02da · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:48.406504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:48.406504Z digest=sha256:1a293cde67917650315d3c526fc3cb0e6e7de5fb01253c53f57006e7897b9eba

Observation eb9fcff0-e087-408a-94f2-21c73d46cd85 · outbound

This paper cites Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:48.539167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:48.539167Z digest=sha256:f0da3dada228401240dc35f548754abbc4d2fb3af5209ce255a3fb740caec756

Observation 9941557c-3dd0-4160-addf-4dd5e4369839 · outbound

This paper cites Qerl: Beyond efficiency – quantization-enhanced reinforcement learning for llms, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Qerl: Beyond efficiency – quantization-enhanced reinforcement learning for llms, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:48.677239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:48.677239Z digest=sha256:cd99830233d1dc9b2241c9f62199ebfccab7011e5a1ecf9848a0e252acebe875

Observation 0e7f2b6d-022e-405e-be6b-acbfa248bba4 · outbound

This paper cites Le, and Yonghui Chen.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Le, and Yonghui Chen

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:48.849801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:48.849801Z digest=sha256:2bb1de2c2ec2e3b417a1ef7257574abec406b055aec81dfb69d79f3922e12e14

Observation 1ad5ed8c-dd73-47b6-a113-3f07a8738e18 · outbound

This paper cites System optimizations for enabling training of extreme long sequence transformer models.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training System optimizations for enabling training of extreme long sequence transformer models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.004334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.004334Z digest=sha256:6946f5655ece4dbb175259c2912dc46cde49f8c2eeafebbdc57f0cbf43ecc40b

Observation b10af087-2781-4102-9693-97c692522776 · outbound

This paper cites Dynapipe: Optimizing multi-task training through dynamic pipelines.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Dynapipe: Optimizing multi-task training through dynamic pipelines

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.171345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.171345Z digest=sha256:91cd489b06c52dbc65dc11ebe284cc3a5703a160e94c7a008337421fced53258

Observation 60fdfe69-e7a5-42ec-929f-df86bb10a6d7 · outbound

This paper cites The Art of Scaling Reinforcement Learning Compute for LLMs.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.318391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.318391Z digest=sha256:b3c54bcfc6c6713faaac3720df205dbb7300cfacfc1fadea592078c0ab33e8e3

Observation 759b0824-ebb4-4f4c-91be-fca8c99cb1c3 · outbound

This paper cites Waterfall Bandits: Learning to Sell Ads Online.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Waterfall Bandits: Learning to Sell Ads Online

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.508315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.508315Z digest=sha256:effad3a2de485bf8ef0af8b38f01096b184af8637de629c1eae6e55d1714f9a0

Observation 9ed6135a-d992-47f3-a2ab-9d8f18e04021 · outbound

This paper cites Efficient mem- ory management for large language model serving with pagedattention.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Efficient mem- ory management for large language model serving with pagedattention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.682524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.682524Z digest=sha256:a2533535c7b8e7e95fc3fe6dc577c6aa2f7ff185a3109d6fda66e9cb6730cbfe

Observation b6776c5a-93e5-464a-bd42-290c099a750b · outbound

This paper cites Puzzle: efficiently aligning large language models through light-weight context switch.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Puzzle: efficiently aligning large language models through light-weight context switch

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.806586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.806586Z digest=sha256:1cb6bcfa4b6ea8bbe8f1e00f09a63f134b42e62a9721f62f0fe6f5a6d31a470d

Observation 971c45f1-5afc-4bae-9e6d-b00f6c7c220c · outbound

This paper cites {GS}hard: Scaling giant models with conditional computation and automatic sharding.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training {GS}hard: Scaling giant models with conditional computation and automatic sharding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:49.925465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:49.925465Z digest=sha256:34d31090574b14f2444bfee084ab60eb00fc0a460ce86ddcfec5452b56dab976

Observation bdefbd6a-bae6-493f-aeee-3cce72f33190 · outbound

This paper cites Fast inference from transform- ers via speculative decoding.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Fast inference from transform- ers via speculative decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.012227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.012227Z digest=sha256:7a330bca65c06b89016c0019c4443320dcedcb2959297f6d73125c5898a1feb3

Observation b38a508e-bf0e-42df-b69c-a927050b39e0 · outbound

This paper cites Hetu v2: A general and scalable deep learning system with hierarchical and heterogeneous single program multiple data annotations,.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Hetu v2: A general and scalable deep learning system with hierarchical and heterogeneous single program multiple data annotations,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.100852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.100852Z digest=sha256:0f21535e31b83b69fd6b88a5c53e296be5c81272a0ab3db1ef7d9ce0992ae0f5

Observation 42e09a2b-fdf6-4be4-91ec-89f98a815f44 · outbound

This paper cites Malleus: Straggler-resilient hybrid parallel training of large-scale models via malleable data and model parallelization.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Malleus: Straggler-resilient hybrid parallel training of large-scale models via malleable data and model parallelization.Proc

Reference 32

Resolution
verified exact
doi, observed 2026-08-04T06:28:25.646537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-04T06:25:50.255157Z digest=sha256:19be5cc5dee07c9fe70cf5b18f517d537d0a18057a59925f3d0f8c656a2b5927

Observation 24802f64-f13b-4f1c-8a60-73b766f0c820 · outbound

This paper cites Hydraulis: Balancing large transformer model training via co-designing parallel strategies and data assignment.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Hydraulis: Balancing large transformer model training via co-designing parallel strategies and data assignment.Proc

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.348565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.348565Z digest=sha256:8ae249552754cd1f9b693fee70073b8311789db59ad7ba06e22cd567ef59949b

Observation 4e5f30a5-0322-405c-a3e0-dd3bf5e9586a · outbound

This paper cites Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.166691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.166691Z digest=sha256:49069504a2a72d947034ceba8fc74831450aa6f15c904eaf246582c995319437

Observation 562a772b-9e7f-402d-a5e7-21576d326cad · outbound

This paper cites Let’s verify step by step.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Let’s verify step by step

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.509055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.509055Z digest=sha256:67b358537b1e0a406878d7b0438172a87507952f38a0ed63d980a508ba429f48

Observation 4c74e100-221f-47cc-a6b4-c7d9d0dcc07e · outbound

This paper cites Lobra: Multi-tenant fine-tuning over heterogeneous data.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Lobra: Multi-tenant fine-tuning over heterogeneous data.Proc

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.724092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.724092Z digest=sha256:e5fc9f1b49657e49a9502fb59279e9fc84358d5f9dd7c939ede1893e1dc76e8e

Observation f253949f-05d5-43be-a2bf-d406335d44d2 · outbound

This paper cites Hotprefix: Hotness-aware kv cache scheduling for efficient prefix sharing in llm inference systems.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Hotprefix: Hotness-aware kv cache scheduling for efficient prefix sharing in llm inference systems.Proc

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.415563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.415563Z digest=sha256:12f2d5cb368fdbf2077380b95f5fae4ed7a1307eb2296966448f82b5423e094e

Observation da905cd3-48b8-4494-8335-61b8e644890f · outbound

This paper cites Ringattention with blockwise trans- formers for near-infinite context.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Ringattention with blockwise trans- formers for near-infinite context

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.889465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.889465Z digest=sha256:a1b755a7671b2c38a0f636fabf3d3f0e92dd11b71d3f4e647dd2d4c71ce912cc

Observation 650fdb5a-4bf4-43ef-8ddd-ce6476a26ee5 · outbound

This paper cites an unresolved cited work.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.606722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.606722Z digest=sha256:8078f80f43f9b86e9d870af90a31953ea71c243bf5bfc7142c9ac7e926d70ab5

Observation 920e84a1-2abf-46e8-a3ce-4ba31dab7334 · outbound

This paper cites Flashrl: 8bit rollouts, full power rl, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Flashrl: 8bit rollouts, full power rl, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.048688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.048688Z digest=sha256:dc34e37744bbb71121392ac93a2b86987fce4165515143a4d8a0a0a2f48eb114

Observation 854ae4ba-1565-41a9-88c3-8bfcbaff2279 · outbound

This paper cites Spec-rl: Accelerating on-policy reinforce- ment learning with speculative rollouts, 2026.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Spec-rl: Accelerating on-policy reinforce- ment learning with speculative rollouts, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.785172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.785172Z digest=sha256:fe23694e4da12c0f48ee9cc08c147fec4f313dd97b156e2f13b5c51210a127c8

Observation 4562a217-4d99-4aab-ba01-2804c3aa3179 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3- mini level, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Deepcoder: A fully open-source 14b coder at o3- mini level, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.207314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.207314Z digest=sha256:ad2c7e9817893c9c16a4fa228eb61a28101399333e324974df43d07cbaeb5cfe

Observation 13681cab-9ead-46f5-9a41-d593f8293177 · outbound

This paper cites When speed kills stability: Demystifying rl collapse from the inference- training mismatch, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training When speed kills stability: Demystifying rl collapse from the inference- training mismatch, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:50.976916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:50.976916Z digest=sha256:478d31f65ccd24ebcadb18ae6db8ff7e85f6b5261d254b9fa3384b3656f9d76f

Observation 8f33d82f-ada1-4c8e-a12b-351bb4ddaf7a · outbound

This paper cites Galvatron: Efficient transformer training over multiple gpus using automatic parallelism.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Galvatron: Efficient transformer training over multiple gpus using automatic parallelism.Proc

Reference 44

Resolution
verified exact
doi, observed 2026-08-04T06:28:25.468395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-04T06:25:51.364088Z digest=sha256:9ccdfeccb02a1c248036fe2ff3b8dadd9393d20e8f7c53a82cd561fd351acb86

Observation d6481e17-341b-4ad4-8d04-d40f5bc6e55f · outbound

This paper cites Part ii: Roll flash – accelerating rlvr and agentic training with asynchrony, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Part ii: Roll flash – accelerating rlvr and agentic training with asynchrony, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.130997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.130997Z digest=sha256:206422d739caf6ed3953571f872a2da770b76bcf75eb7628de10a332b7e11066

Observation 96516336-859e-4d2d-9fca-43c12dca1d7b · outbound

This paper cites Devanur, Gregory R.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Devanur, Gregory R

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-04T06:25:51.491354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.491354Z digest=sha256:c4ca8f7da27d7f281508b5e98e32517c5b2c9addd2506c9ee43ae3ae49034db1

Observation 82e09f00-0196-4482-8a9a-f59041cd9a3a · outbound

This paper cites Real: Efficient RLHF training of large language models with parameter reallocation.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Real: Efficient RLHF training of large language models with parameter reallocation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.286478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.286478Z digest=sha256:90cada3bab5d5032a4d75cdab5aeae2014a019b46d8222d36d8a2fc1cd12af79

Observation 29d17927-1ebc-4160-a6e4-eacac569bc0b · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.694220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.694220Z digest=sha256:0116dc18a081c7de750e7b6355c02921c83b31b6f81bdb0958d20c96b38aa236

Observation 412432b0-32ad-419b-b4b4-7f8657f4e42d · outbound

This paper cites A comprehensive survey of mixture-of-experts: Algo- rithms, theory, and applications, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training A comprehensive survey of mixture-of-experts: Algo- rithms, theory, and applications, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.422548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.422548Z digest=sha256:eeac78c6e9e114a1472ae33a9d345be7f1225a2ebbfebca583d808c5836dd426

Observation cbcafe16-774f-4030-a824-b49b56b8e2e4 · outbound

This paper cites Nvidia inference xfer library (nixl), 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Nvidia inference xfer library (nixl), 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.842717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.842717Z digest=sha256:e317cf31bee337d31130cb66d0797890163d3ec301396f57be031000880296f6

Observation cd67b4a5-c2d9-4d92-8846-b060a70a52be · outbound

This paper cites Effi- cient large-scale language model training on gpu clusters using megatron-lm.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Effi- cient large-scale language model training on gpu clusters using megatron-lm

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.562805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.562805Z digest=sha256:92a4b42ee2f4d5be75624243e1b3be078bc80a572c2d2298048ef04c1a13d618

Observation 3e190654-6ff3-499f-9558-835da6982878 · outbound

This paper cites Unified communication x, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Unified communication x, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.126406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.126406Z digest=sha256:c40676bb84dbcb7d182e88b9df994140bfa8efa350c9ef738eed6d2073e28571

Observation 4c1d3651-d1eb-47b4-b34e-246125ea5d41 · outbound

This paper cites Nvidia collective communication library (nccl) documentation, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Nvidia collective communication library (nccl) documentation, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.767667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.767667Z digest=sha256:2e5b886bb0204b5129a95e49c75c4641405523ce4d31c9ae9068b36353a70953

Observation 2db48e94-488e-49d9-9b9f-5b2deaf7f653 · outbound

This paper cites Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.278491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.278491Z digest=sha256:ab5f8a8135ff54815a8c2d8701e8714903f4fca83a1c741165e31a571fc48647

Observation eddd806c-6695-4a0f-96ee-c77cc7de796b · outbound

This paper cites Openai o1 system card,.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Openai o1 system card,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:51.950676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:51.950676Z digest=sha256:4593ea569c5f6473c57ad75f02fd09b9c9add197837abe75e03784f8efabc822

Observation cac33af7-6f70-45f9-936c-7400690331c7 · outbound

This paper cites OpenAI o1 System Card.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training OpenAI o1 System Card

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.052018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.052018Z digest=sha256:95296e627f87d759fc92942e793687e31f1f6d263b3dd98f3bcf315dfce0daf2

Observation f84817e8-d325-4fd4-82d0-f84b5d17d5c5 · outbound

This paper cites Rosberg and I.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Rosberg and I

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.582254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.582254Z digest=sha256:9fc6cf12a624f136f5609a8f37d8a29ea3b64efcaec263a899c9e630a2c06101

Observation 9b11963d-20af-4e91-87d2-9c703e6a8a05 · outbound

This paper cites Multi-step reasoning with large language models, a survey.ACM Comput.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Multi-step reasoning with large language models, a survey.ACM Comput

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.209334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.209334Z digest=sha256:8115059da2878d35837388120b64349dfa483379711d29a946213fe7b76490f9

Observation 84380506-18cc-4997-8e56-5b6a75f96676 · outbound

This paper cites Proximal Policy Optimization Algorithms.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Proximal Policy Optimization Algorithms

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.769791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.769791Z digest=sha256:9430e4bdefae725e6c06233a860fa3e691a271d2dd4edda3343d0acfa29610d4

Observation b0dba959-26dd-4795-9aba-fde77b0fd286 · outbound

This paper cites Qwen2.5 technical report,.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Qwen2.5 technical report,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.382715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.382715Z digest=sha256:4a73e77cec07edb36018b0e56f00addc6eb8a5fdc28c98ea20ec8b7209196b69

Observation b00ef343-ed7e-4553-bf8f-68de283ebfeb · outbound

This paper cites Qwen2.5 Technical Report.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Qwen2.5 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.466334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.466334Z digest=sha256:3860b500b38a9a17b9453ccbdf145bffc4f7835299dcb19132785131404c8f04

Observation 61360b29-568e-4bbf-83af-e58566a5617d · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.512470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.512470Z digest=sha256:350fc70264ccd962dce82ae3463adc2be56988b252b02ea5c80498ccb44dcd23

Observation 0c014a9e-22d8-49ff-9bbb-972a2e8c9436 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Hybridflow: A flexible and efficient rlhf framework

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.226125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.226125Z digest=sha256:1453afca60d15f82c2ccdc16dd29b4b033cbad63f03bf8d5fa2366ea325aa969

Observation 6ff19108-eee3-4297-a4dd-95b62144ea25 · outbound

This paper cites Trust region policy optimization.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Trust region policy optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.665305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.665305Z digest=sha256:0061d47de836739216fabd69d8b9b099ca7ef638ba375007972d920b366c0b06

Observation 79d20292-6995-4f4a-918c-5155b8232b6f · outbound

This paper cites Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.415049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.415049Z digest=sha256:bc860c381930996ed6537f1c597c571a241fac95011676ae11dac4231482096d

Observation e695934d-434f-4b28-a7b5-a0305e3a12f5 · outbound

This paper cites Beat the long tail: Distribution-aware speculative decoding for rl training, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Beat the long tail: Distribution-aware speculative decoding for rl training, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.897121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.897121Z digest=sha256:9249ed5ee1e0d46f32f69dc20d6ba57737d892fc27d68c0ec46b7dfcabed92da

Observation 955ac5a8-e82a-44c4-b9c0-288a9f769fd6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:52.994986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:52.994986Z digest=sha256:1870b11df0c89fd6cf30f2cc899895794960fc51c142c8164a53e4baeef8b3dc

Observation 07b5031f-76a9-4ff2-908e-fb6a77d7e4ae · outbound

This paper cites Laminar: A scalable asynchronous rl post-training framework, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Laminar: A scalable asynchronous rl post-training framework, 2025

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.075946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.075946Z digest=sha256:d1c3bc92e039d6c0bffb8bb1030d1f58e33e9e3c96f55b856c8f76f19b533954

Observation d97d48ad-1e62-438c-90d9-480af97718d7 · outbound

This paper cites A survey on large language models for mathematical reasoning.ACM Comput.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training A survey on large language models for mathematical reasoning.ACM Comput

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.844641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.844641Z digest=sha256:0e6fc02c66f7e19770087e16949fb9f4f7c6878926464a7212d9e1e38c94e386

Observation 4612ab4b-409e-436f-918f-638a801a6d25 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.315719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.315719Z digest=sha256:4b7b590acca436d74d8ff8497c3d2476d72dedc60ec313046617f6ccd8b78b90

Observation 50fb9431-bd93-4066-b33a-9e69d465bb74 · outbound

This paper cites Improving automatic parallel training via balanced memory workload optimization.IEEE Transactions on Knowledge and Data Engineering, August 2024.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Improving automatic parallel training via balanced memory workload optimization.IEEE Transactions on Knowledge and Data Engineering, August 2024

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.057071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.057071Z digest=sha256:86209ee86ff9374ff6f43ecdd535ff02188cc3df44d194fb2080fcdfef227f42

Observation 28c2d01e-1ce4-4c4e-b6e8-1786e2da8467 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Kimi K2: Open Agentic Intelligence

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.541884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.541884Z digest=sha256:cb6000f2f6b0c808ea27087dc0091b1e454d8bf4ea97130b928cea21561d7bbb

Observation d5bdd606-2987-4e21-b619-42fbc18bb306 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.654690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.654690Z digest=sha256:cb6f39644f5f31af21761e5215844fd96d0e80ec9546881101b4153a60f856ab

Observation 17f4edf4-c35e-4fa6-8617-3f4c431f76c1 · outbound

This paper cites Tokdar and Robert E.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Tokdar and Robert E

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.746745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.746745Z digest=sha256:1f6c31fa4eafa733d687780dea8fe5deb1292d4afebd7503df5e4cfb6846edc1

Observation 2ad77e30-4bf1-42a9-8208-ebeed40cd770 · outbound

This paper cites Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.456232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.456232Z digest=sha256:07a8d01ade599e80195ad8e27929a031c6fe8da45b725e95ac6daaff73930a69

Observation 1acb3228-e099-4eb7-9fee-cabf3f8bd5a7 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:53.985066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:53.985066Z digest=sha256:ccaa3c36ea02e0d54872485c500bfa59c2ec974577cc6c78c61d65da867884fd

Observation 1ea59b38-0d3e-4802-8a51-36cf7a286158 · outbound

This paper cites An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.634282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.634282Z digest=sha256:5e49816901db4c184075e378a127137cd1c44d28e18fc74c8cfb61cd7d5980bf

Observation 2a1928ea-f039-40a9-85ac-a0841bbe5f13 · outbound

This paper cites Flexsp: Accelerating large language model training via flexible sequence parallelism.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Flexsp: Accelerating large language model training via flexible sequence parallelism

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.151703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.151703Z digest=sha256:9a2ed3c8e4d2107ad3a3576a653436a3edaef585814bd22b2660f51729b14a96

Observation 0fca62d3-9fac-4363-924b-df5ea3a77cc7 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.268990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.268990Z digest=sha256:90e8013e53e20823f3dc60f0eae0b3f3da8be6b6c748e1bf35f0c8b834ef381c

Observation e4c14c07-922f-412e-ac39-730f9a9a07cc · outbound

This paper cites The multiqueue: A simple and fast relaxed concurrent priority queue.ACM Trans.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training The multiqueue: A simple and fast relaxed concurrent priority queue.ACM Trans

Reference 80

Resolution
verified exact
doi, observed 2026-08-04T06:28:25.315086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-04T06:25:54.359978Z digest=sha256:65245fd7ee81fbf2193d7aff6626988e0861045ca7815e6165ad0fef8765072b

Observation eec37935-3da4-49bb-a1f5-921389fd9f31 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for LLM inference serving.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Flashinfer: Efficient and customizable attention engine for LLM inference serving

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.987759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.987759Z digest=sha256:a21a179970862ec99b62468e4a00925f8f0042538c84be46fcfd8fe2f28ac3b6

Observation 4e20ecf2-a14a-4eec-9edc-ab8728858d75 · outbound

This paper cites LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.516468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.516468Z digest=sha256:bb7e131ef28044b75ad6e79c22fe4ac8215b7860d5771c149efea9a2143a423d

Observation 5835fd35-7ffa-4c75-9c17-5d29d42d2d18 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? In2nd AI for Math Workshop @ ICML 2025, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? In2nd AI for Math Workshop @ ICML 2025, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.133042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.133042Z digest=sha256:63c435453ef7a36d3af982e9624b5edd009682f0ee62273fbc46ab14c7b1a4b4

Observation 02adb91b-d507-43bf-aba5-8937b5c69b74 · outbound

This paper cites Qwen3 Technical Report.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Qwen3 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.718668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.718668Z digest=sha256:b616aaf8d89b21e76fdf067517adc49c92641c78604c9d9468a1d2c9ba0990b8

Observation 4fbb9027-e47f-40b2-bf72-933a041c3656 · outbound

This paper cites Survey on knowledge distillation for large language models: Methods, evaluation, and application.ACM Trans.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Survey on knowledge distillation for large language models: Methods, evaluation, and application.ACM Trans

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.811385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.811385Z digest=sha256:9e117cf5260d37696ce6fc944e972becf0a245c4dc2bc4e229ce8ff5472c0ca1

Observation 78a1a2b6-cff0-4e93-9370-d5d938d93c0c · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Your efficient rl framework secretly brings you off-policy rl training, 2025

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.901122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.901122Z digest=sha256:b9bc45ba1fc84ddbad38699c24c0576c61ca13464c78e5210c3fa7e5bcb98e7f

Observation 6811364e-e69c-4297-bf0e-ee924a2a275b · outbound

This paper cites Small leak can sink a great ship–boost rl training on moe with icepop!, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Small leak can sink a great ship–boost rl training on moe with icepop!, 2025

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.483112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.483112Z digest=sha256:0d67c9371fccdca2b7c7e7fc447b4b54cc6de2d1696e5f1921be8f8f6f0e92fd

Observation bdd10f48-5fa4-4737-a8bd-bbd3f600476f · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.044439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.044439Z digest=sha256:207f2d3e693555ca8af5305dda106ce8ee67e5f918f3b90622ced9803509989a

Observation 425d4484-cecb-4b70-9866-6173253b653c · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Stabilizing reinforcement learning with llms: Formulation and practices, 2025

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.712023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.712023Z digest=sha256:1d1d49ee5ee43bae7c72dcd3e69309c890ebfd2ef63391a5b22cb91655022029

Observation 243bd583-69b6-4d5b-b20a-d6b730dd79ba · outbound

This paper cites Pqcache: Product quantization-based kvcache for long context llm inference.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Pqcache: Product quantization-based kvcache for long context llm inference.Proc

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.249282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.249282Z digest=sha256:273e9180772b5773b8f816689c9e316164895603bb3db195e3b85579af836d98

Observation b0f89fda-f521-49ad-a589-cecfaddff8b7 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training A Survey of Reinforcement Learning for Large Reasoning Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.329841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.329841Z digest=sha256:fc2b371ed3ce7961c6adee314a15b1c0e812d5536c7159d903f958e7447d22a5

Observation 9bfca4bf-a9f5-4f54-8a30-aea5556b1d13 · outbound

This paper cites SortedRL: Accelerating RL training for LLMs through online length-aware scheduling.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training SortedRL: Accelerating RL training for LLMs through online length-aware scheduling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.423450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.423450Z digest=sha256:bbe72e1d0704e66a7fca74f6943b6512786009cad445042ee41fe345707b51bc

Observation 8449252e-88cc-4f8c-a898-5018835a4ac2 · outbound

This paper cites Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:56.071060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:56.071060Z digest=sha256:a63a68f114a80d15bfaa001a922ad7ae2372b9a03d183e98c72eef1bee0bcd02

Observation 38a0cd2c-1b2a-4242-823b-8d1da3555989 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel.Proc.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Pytorch fsdp: Experiences on scaling fully sharded data parallel.Proc

Reference 94

Resolution
verified exact
doi, observed 2026-08-04T06:28:25.140551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-04T06:25:55.629911Z digest=sha256:6e04a998c0f21c4d2eab8b757ecb8657860a5d9cf5441dcb258cb35724d8b096

Observation c447d8ad-1637-49c6-8bfc-04edd57a3944 · outbound

This paper cites Optimizing rlhf training for large language models with stage fusion.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Optimizing rlhf training for large language models with stage fusion

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:56.265063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:56.265063Z digest=sha256:32c313f3be19b9269209fc01e24018853e9b4ad529a53fd0f6b9401911e60498

Observation 5f7de689-6ae7-496a-a0ec-3fba33b4d0e7 · outbound

This paper cites Group Sequence Policy Optimization.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Group Sequence Policy Optimization

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.821502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.821502Z digest=sha256:8cd4ea1928c3a4c1bd50becbbf5c5f9a0a448daffbab248ddda9c4a916554c86

Observation e76a52d9-9bc9-4be7-b84f-d0dd6b5f1783 · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms?, 2025.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Prosperity before collapse: How far can off-policy rl reach with stale data on llms?, 2025

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.909032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.909032Z digest=sha256:e27369428d7e086a6fd7eedd1c6890a8757cc1836f885e8cb9f2a0cd036b1f13

Observation 5fc58b4c-c25b-4b5f-90bd-1a1a04a51e7e · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training Gonzalez, Clark Barrett, and Ying Sheng

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:55.987217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:55.987217Z digest=sha256:7a5e789a30d773d5da925c2a71c57bd809be5e829e1e437e3381c7d01c673e06

Observation 940e6faa-2206-42ea-b86d-21bf3ee3ce7c · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:56.161461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:56.161461Z digest=sha256:7ff34a86cf89513243b6739c05179cd044bcf4999bd4a7e9b20ef7bcdfc565c0

Observation 227b2cd2-da42-4d57-82fe-30fc210e4840 · outbound

This paper cites April: Active partial rollouts in reinforcement learning to tame long-tail generation,.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training April: Active partial rollouts in reinforcement learning to tame long-tail generation,

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:56.310135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:56.310135Z digest=sha256:cdfaebf2f906f28a023d88da933812d3b9ab630e62d1220251a9b5894f0ea2e6

Pith citing papers

Observation df03747c-978f-42c3-b4cc-e1521ed2cd38 · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:39:57.610477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:37cc8b3d45e7d21fd7816d88531ef15fd30ee6ccb0b5206fc426b6df633188c2