Pith. sign in

Paper Citation Record · LEDGER

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2602.06422.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.06422 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:58:50.730090Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T20:38:28.328436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:35:47.308416Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd4823fd-3f9a-4bc0-8804-05f6c0a19a57 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Training Diffusion Models with Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.478031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.478031Z digest=sha256:e2544dbff9672086f827ddc0af296e627a846c10dc140fe7271b428a78548f1d

Observation 97819a72-f58a-435f-a44d-2f374c6a8b0a · outbound

This paper cites an unresolved cited work.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.730090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.730090Z digest=sha256:2905bd87ea05479eea5d108a9e143687be4498290ba785847c6b90de05c9c305

Observation 2e78b1ca-bb65-4826-9c13-564fee38b66d · outbound

This paper cites Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.835551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.835551Z digest=sha256:9d6fdb706f3b738790d5daa1724e90ee0d020bb4fe08e089e7fe76109c5a51f7

Observation 90896383-cba7-480c-8ca7-4346902e1dca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.943935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.943935Z digest=sha256:d0441456441c4f337ab88c6f7c4dfef64f632caf45c7e8be6300b6fd48b45f62

Observation f2b5f785-2e51-4051-a603-adbd2bd6ba1d · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.067234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.067234Z digest=sha256:8b10c4babce3cdf4f4cda5ad16245e5d800e0f42e2b25975771e43fd22a0ed86

Observation 5be4e67f-1567-4acd-9a94-175afaf311c8 · outbound

This paper cites an unresolved cited work.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.688606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.688606Z digest=sha256:2e1b69987988de94fa1e054197304490367ea8d1f3d494ce9390e076bfceffe7

Observation cde24e99-d1f6-4383-bea0-86b457ff8e3d · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.354792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.354792Z digest=sha256:9b8e91b779bcbc24eb22d512ea5ce537393ce12062939ffcd406048315e523c2

Observation d3bd812a-b04d-4bc6-ae04-74353ab868a8 · outbound

This paper cites Flow Matching for Generative Modeling.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow Matching for Generative Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.455396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.455396Z digest=sha256:3ebd8487174f4bd7eb65e04e7429585a4ef4d1276ad86de5a04793f934591e6b

Observation d961a994-2c86-4e5e-9bca-f7390645d529 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow-GRPO: Training Flow Matching Models via Online RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.630032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.630032Z digest=sha256:82d2106f17b3eb9d28509dc086c6e98cced1eac63ea7b49dcd724f21c51379f9

Observation 6e4b85ed-65c1-4eec-8794-4dff5e214a8d · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.747831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.747831Z digest=sha256:7adc317ac7dff5485b83de6ec6f2284b8f58d91ce57ffce70e18c453af15e437

Observation 639bd713-4120-44cb-adc5-b7d95bd6d50b · outbound

This paper cites Grpo-guard: Mitigating implicit over-optimization in flow matching via regulated clipping.arXiv preprint arXiv:2510.22319, 2025a.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Grpo-guard: Mitigating implicit over-optimization in flow matching via regulated clipping.arXiv preprint arXiv:2510.22319, 2025a

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.127534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.127534Z digest=sha256:17947d3acda5c4b35f6a35e7d049e3d9c4c8e9a54af020767cb3a59f365e5f3e

Observation a50eaa08-f5fb-4fdf-bef9-62c3a9ba542a · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DanceGRPO: Unleashing GRPO on Visual Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.281427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.281427Z digest=sha256:5c51dcae1545ebe86ada874614a5b669f495bc4c99b7df5cea0004073e735843

Observation 33790b14-018f-4f69-ada6-8934bee8ee34 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.427811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.427811Z digest=sha256:8efbe86d8913ccccc170dbdeb5bba6304ac2c6dffc9681510d88acbc8d6c590d

Observation 9f97412d-913d-4f1d-a1f2-cb5ea11b10c7 · outbound

This paper cites Group Sequence Policy Optimization.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Group Sequence Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.564346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.564346Z digest=sha256:239c58bbba64ca99079fdebb7695557411b5ed43c68b81ed8cb2f920ffde54a3

Observation 5fe96f48-1dd4-4a36-9acf-bbd97a657803 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.055891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.055891Z digest=sha256:3a26e0549e845e9491e165eca24a2d2cbf3a5f6ea9a9c595433483ff5d800f9c

Observation cdac7abf-3952-4199-803f-d6e647d6d7c0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Proximal Policy Optimization Algorithms

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.865910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.865910Z digest=sha256:41bfffed9a94994b6633a33cf5e0417e6a8986a56234e60b5fb8e15d47293fe6

Observation 9b7e307a-80b3-4b1a-b673-c4319ec040da · outbound

This paper cites OpenAI o1 System Card.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO OpenAI o1 System Card

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.120237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.120237Z digest=sha256:07affb29fbb17a932bd51f7c6ce78118fb12061996012a5f310b082b59be89bb

Observation c77a963d-838a-47de-9571-b24da31e2d72 · outbound

This paper cites PaddleOCR 3.0 Technical Report.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO PaddleOCR 3.0 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.541300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.541300Z digest=sha256:9a4fa75f06cc624fcc7079801f7865d8c33b09763e7c56b2fcc0a1c9cd56aa6d

Observation 777f14c1-41b3-445c-ac38-471d792bea4d · outbound

This paper cites Guiding a Diffusion Model with a Bad Version of Itself.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Guiding a Diffusion Model with a Bad Version of Itself

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.195604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.195604Z digest=sha256:f7d846d673f076650ed05e810f6d15de69e68b19d5d8fc2db2a1c9745b6d2a63

Observation 89b3d90f-74de-4370-8b2b-36f1f8542440 · outbound

This paper cites Densegrpo: From sparse to dense re- ward for flow matching model alignment.arXiv preprint arXiv:2601.20218,.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Densegrpo: From sparse to dense re- ward for flow matching model alignment.arXiv preprint arXiv:2601.20218,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.719811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.719811Z digest=sha256:07d541104dd1b28d003ee97be41d16af100512e4f930c35336ceef622287ab1d

Pith citing papers

Observation 72c1ca81-5a6e-46b9-912c-239b718f6730 · inbound

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO cites this paper.

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:12.085101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T20:38:28.328436Z digest=sha256:879ab65943016f2545794ae567e2ab414d8136f78c90d0fd93399c0915f89d24