Pith. sign in

Paper Citation Record · LEDGER

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.05600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05600 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:54:32.517155Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 74a6c8e8-bd8b-4f2a-982b-e23e62c00a24 · outbound

This paper cites gummy bear zone.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction gummy bear zone

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:33.556213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:32.517155Z digest=sha256:0f74fb2a02e06e2d272e52c50964e0f7b9d0388e61448e3fd11d6e07c58c2de9

Observation 78cdb6ac-7061-47d1-9ffd-515270f00c1b · outbound

This paper cites Densegrpo: From sparse to dense reward for flow matching model alignment.arXiv preprint arXiv:2601.20218,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Densegrpo: From sparse to dense reward for flow matching model alignment.arXiv preprint arXiv:2601.20218,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.442334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.442334Z digest=sha256:32e361352763ee99f9c6fadbf3214af336f1e558c709cae7d0d3be3c42dc0b13

Observation 2fbb39f5-cce6-4cdc-a516-4e2205c87a32 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Aligning Text-to-Image Models using Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.463629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.463629Z digest=sha256:9e861d24db0c44c46a2a93fe67e09a9351ab0b5d226542460862f2b14864b959

Observation 450e6ffc-4d27-4f86-a1e8-d3b8f04101f0 · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.466927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.466927Z digest=sha256:caec389c5aa43737a41417c6a6d2508f22e6ea5f0f7ae4b132b389050770d765

Observation 39601ea0-7e1a-4235-846f-40869afebf2c · outbound

This paper cites The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:33.179122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:32.470552Z digest=sha256:913757c25c329d3c1df9205d687f0d0ae92a8c465681f09b9778d20976865183

Observation d6aae682-450c-4d2e-ba33-86b016b218f8 · outbound

This paper cites Flow Matching for Generative Modeling.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Flow Matching for Generative Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.473890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.473890Z digest=sha256:b1b835b089e5f083e7f71224573e8cf3cd9dda7ab9c6027b143b8f673df17204

Observation 4a6fcdf9-3da1-4cfb-be82-b81ba54369fa · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.477285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.477285Z digest=sha256:1dde48081d265a5f90ccdbb404bac05d533a99317e5ef75be1bac5286aaa7ff2

Observation a895c95e-2f21-4848-8a64-1fc62b7c9d19 · outbound

This paper cites De- feating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction De- feating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.480506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.480506Z digest=sha256:a34fcc484026db10a22b641e681ef7082595903a69395e3870aa1f13b0d1da72

Observation 01be9f98-9d40-457e-a09f-925b05f97ef0 · outbound

This paper cites FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.483595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.483595Z digest=sha256:bafb2bae16757e638cc03d43edc725fa0150b955c815becab1ad9a74a490c79d

Observation df2b5a68-f02e-4db3-8dcd-f24c8368982e · outbound

This paper cites Denoising Diffusion Implicit Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Denoising Diffusion Implicit Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.486891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.486891Z digest=sha256:c7693228e5f49a4c20e660c7e0aa2589af9cbd724300ca8722352c67f2536b0d

Observation ffec1e08-447d-4820-9eec-013ba5787803 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Wan: Open and Advanced Large-Scale Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.490246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.490246Z digest=sha256:fce7b2f02bfc0d43b3f8787fa9d0db5b9c81a400646c52c8db9c3dfed45d0364

Observation b1c50b50-82fb-400b-8672-fe18ad1f0ac3 · outbound

This paper cites Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.493575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.493575Z digest=sha256:d3267e75d48f215603e859fa71815507437fd2e311a0173d1509fa9e6cbb10d8

Observation b1244178-d0ff-4c5a-ba0b-e271f9ac3452 · outbound

This paper cites Qwen-Image Technical Report.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Qwen-Image Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.497114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.497114Z digest=sha256:6cf825b7d807f11bd46430499ab27b193b63255e47c8d36fd220a7ef63e0e356

Observation 4770d7fa-9cce-4846-bb3b-8e84ac3e9d15 · outbound

This paper cites Human preference score: Better aligning text-to-image models with human preference.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Human preference score: Better aligning text-to-image models with human preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.500748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.500748Z digest=sha256:8743e91547d44f9b54167eff802b79a836fa1dfd38dc1c9dbbd4fe6c39d59ff9

Observation e0be788b-bf4b-4f9e-9ba9-d24a224a7464 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction DanceGRPO: Unleashing GRPO on Visual Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.503814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.503814Z digest=sha256:9af658d63522f6d78ee16d5ed2b396120ff4286ea492a49eca163e67b13c9ec6

Observation fbb1a55c-c91a-40e9-b85b-18c7e94a1d49 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:33.566447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:32.507119Z digest=sha256:4dd4d2a3c338c967d6162758cd64b357ae963fb5e0b3802d6fed455b099fa20a

Observation 2ccab37e-fedb-4b16-9eb7-fd2b44e3b0df · outbound

This paper cites DiffusionNFT: Online Diffusion Reinforcement with Forward Process.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.510791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.510791Z digest=sha256:fdf8109b42fcf0a13c8c5787d4058b0173c82f489c603c1c608c46ec7260628a

Observation 01ea36d9-80b1-401b-84a9-299eca4b0bae · outbound

This paper cites Manifold-aware exploration for reinforcement learning in video generation.arXiv preprint arXiv:2603.21872,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Manifold-aware exploration for reinforcement learning in video generation.arXiv preprint arXiv:2603.21872,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.514207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.514207Z digest=sha256:9132b7ff8a96f60e281b9360e5790ecf6df652d9fd42289d67eba5351f893d7b

Observation 6d317bed-9a7a-45c7-b002-0db3326cc03c · outbound

This paper cites Training diffusion models with reinforcement learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Training diffusion models with reinforcement learning

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.430935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.430935Z digest=sha256:472ab827071d16064b24a2c85737960a57b5e5514717f9b6efdd263586ea7132

Observation 566e7546-1251-4838-9300-bbc11ddd74d9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.459894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.459894Z digest=sha256:0578b8c9a0d3a0e64823b325a65b8fb0b158e3f466c709fd97fe89e47e65ab44

Observation 16dac972-a554-4a30-9948-9036c2d111f4 · outbound

This paper cites Treegrpo: Tree-advantage grpo for online rl post-training of diffusion models.arXiv preprint arXiv:2512.08153,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Treegrpo: Tree-advantage grpo for online rl post-training of diffusion models.arXiv preprint arXiv:2512.08153,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.445788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.445788Z digest=sha256:27dabdf3c6778d1d69f46d18a8179e2b036299d79e7fdfc9120897cf84e27f85

Observation 5c179457-288b-4b1c-a484-f93c36261d2e · outbound

This paper cites Directly fine-tuning diffusion models on differentiable rewards.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Directly fine-tuning diffusion models on differentiable rewards

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.438659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.438659Z digest=sha256:4786255ec29eae527f97ef7e436451fa89fff7b6915655600ccbad2ebb2a3013

Observation 8f03720c-298e-4372-b1a4-20b858522778 · outbound

This paper cites Gradient Guidance for Diffusion Models: An Optimization Perspective.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Gradient Guidance for Diffusion Models: An Optimization Perspective

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.453132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.453132Z digest=sha256:e2b0256374ac3f1dbb1425d50f60785465db2957d6b33e5946cdea80e1cc6d13

Observation e3072153-392b-4d9c-9309-5c258da56e75 · outbound

This paper cites Diffusion Posterior Sampling for General Noisy Inverse Problems.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Diffusion Posterior Sampling for General Noisy Inverse Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.434813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.434813Z digest=sha256:57521807892bf97e110c19e5cf124fc3b539adcda705992a360199316e2cd7be

Observation 25164c7e-f160-4a02-9bfe-9ea5ab4f88b4 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.449343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.449343Z digest=sha256:b307bba81d79d4a104eded563ee3a8db7952ff6f5cf2691aa6555ad9c2948734

Observation 357a86b0-f2f5-4465-b63e-091143c5efd5 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Clipscore: A reference-free evaluation metric for image captioning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.456777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.456777Z digest=sha256:28d6883c1e5251e6c1a4c76e653e0efbb91234385e32d4490c800182ba05c15a

Pith citing papers

No inbound Pith citation observations are available.