Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:58:50.730090Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2602.06422.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:58:50.730090Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T20:38:28.328436Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T14:35:47.308416Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd4823fd-3f9a-4bc0-8804-05f6c0a19a57 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Training Diffusion Models with Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97819a72-f58a-435f-a44d-2f374c6a8b0a · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e78b1ca-bb65-4826-9c13-564fee38b66d · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90896383-cba7-480c-8ca7-4346902e1dca · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b5f785-2e51-4051-a603-adbd2bd6ba1d · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be4e67f-1567-4acd-9a94-175afaf311c8 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde24e99-d1f6-4383-bea0-86b457ff8e3d · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3bd812a-b04d-4bc6-ae04-74353ab868a8 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow Matching for Generative Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d961a994-2c86-4e5e-9bca-f7390645d529 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow-GRPO: Training Flow Matching Models via Online RL
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4b85ed-65c1-4eec-8794-4dff5e214a8d · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639bd713-4120-44cb-adc5-b7d95bd6d50b · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Grpo-guard: Mitigating implicit over-optimization in flow matching via regulated clipping.arXiv preprint arXiv:2510.22319, 2025a
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a50eaa08-f5fb-4fdf-bef9-62c3a9ba542a · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DanceGRPO: Unleashing GRPO on Visual Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33790b14-018f-4f69-ada6-8934bee8ee34 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f97412d-913d-4f1d-a1f2-cb5ea11b10c7 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Group Sequence Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe96f48-1dd4-4a36-9acf-bbd97a657803 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdac7abf-3952-4199-803f-d6e647d6d7c0 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Proximal Policy Optimization Algorithms
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b7e307a-80b3-4b1a-b673-c4319ec040da · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO OpenAI o1 System Card
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77a963d-838a-47de-9571-b24da31e2d72 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO PaddleOCR 3.0 Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 777f14c1-41b3-445c-ac38-471d792bea4d · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Guiding a Diffusion Model with a Bad Version of Itself
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b3d90f-74de-4370-8b2b-36f1f8542440 · outbound
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Densegrpo: From sparse to dense re- ward for flow matching model alignment.arXiv preprint arXiv:2601.20218,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c1ca81-5a6e-46b9-912c-239b718f6730 · inbound
RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.