Pith. sign in

Paper Citation Record · LEDGER

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 89 inbound Pith citation observations for arXiv:2507.21802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21802 v6

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T13:27:50.031781Z

measured 148 of 148 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 89 of 89 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:42:14.452986Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:05:55.323624Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact23
  • verified fuzzy34
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a6e645e7-5263-405f-afda-4269c5cfde73 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.123522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:a13fb14d15833acee6f4486f68683328f1afd227091475f29f7e70143ab3bf6a

Observation 68d8ec45-ae93-4df7-9ce5-0aba8080296b · outbound

This paper cites Discount factor as a regularizer in reinforcement learning.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Discount factor as a regularizer in reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.152526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:39b567c953e09bda7677a695361400a9bb8e6903019789cd86b22f6cecac253d

Observation ab25c8a0-027a-4ba0-b7c7-66fa228ce327 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Training Diffusion Models with Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.100880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:43e849f87df90f56acca7a29986134e07cbcb2edbc03b293eebf910563607308

Observation 9e45b561-18ee-4719-b1f5-419c6aa564c8 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.154396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:edd700bb8fb85f7dd2452bb0b44f29452865dd8058a02a2725cea8b45d72f0d6

Observation 9fcc7767-544d-4831-bc52-de7205cd323a · outbound

This paper cites Optimizing DDPM Sampling with Shortcut Fine-Tuning.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Optimizing DDPM Sampling with Shortcut Fine-Tuning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.075708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:cf3d87300c42325176c6c9e5aee89f4fe9965b662ef1a2cad16313919e2b30b3

Observation 9b3836f6-0050-4b5a-a985-4c8975a6be23 · outbound

This paper cites Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.Advances in Neural Information Processing Systems, 36:79858–79885.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.Advances in Neural Information Processing Systems, 36:79858–79885

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.156549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:230c483a899147ecd13beb80f893f4ade680e79b341cc4e25f10ff338eadb562

Observation e6eb13b6-122a-4bba-ae32-163f0b45b68b · outbound

This paper cites Murphy, and Tim Salimans.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Murphy, and Tim Salimans

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.158107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:3b38a4c937b2f8b1bfbbd8c3e0c3ea34a230a06f696a3b9cbc0e9db3d2eb36a3

Observation d0b89f61-0c5c-47b7-b0e6-93fd71225cee · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.159961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:735384faaab7bd9d5483eacca5b03e5d222b830e35785994c0dfbb6f4eb33c1f

Observation 13f1ca0b-fe66-427b-9e49-ab08d0e12a3c · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Lora: Low-rank adaptation of large language models.ICLR, 1(2):3

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.161597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:90eceb91f5e501fd9a2515b50cc62caa33b09a9ef26c41cfcec1c62662715591

Observation 01590baa-ad40-4044-a802-d4b4ffc91cfe · outbound

This paper cites On the role of discount factor in offline reinforcement learn- ing.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE On the role of discount factor in offline reinforcement learn- ing

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.163557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:4c1febc9fb693b781811fee288602c049be0cfe0f8cac2a41a6b20285980775a

Observation 7802b059-b72e-454c-9e67-8e9c86a27f21 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Muon: An optimizer for hidden layers in neural networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.165632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:54538ce4f26dadfae10f5abd60fb8e22935a03ff7f8160b5887cdd779fff0d62

Observation 44cef111-8c63-4cee-bda8-972882025dd2 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.Advances in neural information processing systems, 35:26565–26577.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Elucidating the design space of diffusion-based generative models.Advances in neural information processing systems, 35:26565–26577

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.167439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:c6d1d9cd693da686a5f47d8732bd7fa9d7ee3cea84e6f7f9c51269dd7b651971

Observation 02e805f8-7857-48f4-9d6e-eec3f99e65e3 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.169322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:01b4f98313cf2dc52c7075595efe9dca04060e1bd8457499df6e71375f903004

Observation c0f2589a-66b5-43d3-9a22-5e409fac795c · outbound

This paper cites Flux.https://github.com/ black-forest-labs/flux.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Flux.https://github.com/ black-forest-labs/flux

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.171071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:b66315ff53db330f681bd4d5d77fc20ccd03914b10408f8d9c8d95e196cd3311

Observation f71c6218-2be7-42b8-a2f0-96400e795153 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Aligning Text-to-Image Models using Human Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:15.590122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:96696e3964179b660ce18b7d85d3c4f2623db67e7a6db2c9890e9ac448d2f54f

Observation a8026d3a-7ea5-4952-86c1-209faa2ec33f · outbound

This paper cites Aes- thetic post-training diffusion models from generic prefer- ences with step-by-step preference optimization.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Aes- thetic post-training diffusion models from generic prefer- ences with step-by-step preference optimization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.172848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:4bbbd2ac67b5e0447118ff19aa1e06c28d6ddd7ddb1c5edcd0fd5543c8b96d55

Observation e68b3415-a2e0-496a-ae47-3e278e920962 · outbound

This paper cites Flow Matching for Generative Modeling.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Flow Matching for Generative Modeling

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.084181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:cd7025c56d321aae0aa35ff542215da1f44df9b5d418e6b835068f078727f0d1

Observation 5bcd5663-6cd7-4c61-948b-8f4f6f33f526 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Flow-GRPO: Training Flow Matching Models via Online RL

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.089562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:371f7d82083e4d25f58554c179ed9187b343ecc66032e63f7e0318e7b7a654c3

Observation a7e2b88a-e714-48b1-8b17-154a4680fed5 · outbound

This paper cites Improving Video Generation with Human Feedback.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Improving Video Generation with Human Feedback

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:03.116566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:58c65bb7acd2c5f4ad0114a4d5d1b98eac3d7fb0f5f3f2cb8dd87b3d08ef31e6

Observation 56a82ff0-842a-45c4-9364-a2eca16dc7ee · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.095685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:b32c144805e41e43852220c0a769d91ee21e0bffab8b0167ecdf3bb36da4eea9

Observation 08338e1e-cf23-4266-9b4b-ab7decdcb2e6 · outbound

This paper cites Decoupled Weight Decay Regularization.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Decoupled Weight Decay Regularization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.098208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:5b3f152f286d267b2622a05335d38fe89c0ab969fdf2a72dcdf5e598506f2eb0

Observation bd3371fa-7d6f-432a-becc-3febbbf189ab · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.174992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:f579d734a9b1e81d32b419073650ab16dd6174f0c466ed7052bae1bda85c1042

Observation 69152a4d-ba6f-4b73-88b5-c0a5b3cc97f8 · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:54:11.877839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:f4238ae187a6525357a87f4ff57f62211ca7740d42408f9ebe33db99e1d0e3b5

Observation 57c13cb5-538d-4f6d-9f2d-53728ea26d6a · outbound

This paper cites Hpsv3: Towards wide-spectrum human preference score.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Hpsv3: Towards wide-spectrum human preference score

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.176708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:82982eab9bd38de759fb7dc244fc848d2db4fca574c072150951599f5b43673b

Observation bac5c7bc-42b5-42cc-a21f-5ea0ab4b7120 · outbound

This paper cites Reward hacking behavior can generalize across tasks—ai alignment forum.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Reward hacking behavior can generalize across tasks—ai alignment forum

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.178438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:ddc826b25d64339357caed3d3c2278c082ac12cc53c2407801fdb3d40f36d0f2

Observation 1539f7b2-1a2f-469e-bf57-d1dbd7a65d11 · outbound

This paper cites Stochastic differential equations.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Stochastic differential equations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.180067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:b99ce7cdf2b9374b2cf001e30ec3773c5a8d093bc94d4dcd992fe6035c6f15c0

Observation 2d55c5fc-224f-42b9-a012-97c979ab902d · outbound

This paper cites Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:27730– 27744.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:27730– 27744

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.181876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:6ee7652e06614d5c9c00179708c18eace92f6bb333067cfd60b325a8c12c0dd1

Observation 1cef30f6-08f4-453e-80ea-e904bd6b59e4 · outbound

This paper cites Rethinking the discount factor in reinforcement learning: A decision theoretic approach.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Rethinking the discount factor in reinforcement learning: A decision theoretic approach

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.183701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:e47f4799cb51183134ebc0b63e1948503f0c60b5d0fb5411205411c736d2963a

Observation 1be115c9-f082-4ce1-b519-cbb91234fd68 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Learning transferable visual models from natural language supervi- sion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.185632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:d1e026d88a8ec66ba58c0a03fac02f90d91ff6f35cfe37fc7ec6595392227d2f

Observation 5190fe21-3b18-430d-938c-4dea76ad2816 · outbound

This paper cites Springer.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Springer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.187402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:b80d79e2bbce6b5d751bac74a47b94ed54a3d5a4a8b8769a5165c04da8f38b1d

Observation bb6058f2-cf9f-4daf-a806-cda332d2e49e · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.189065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:0f24ad403a2341b3eea4a01ba24a3fa9202749e4a3cf126794c396957b4dc77a

Observation 4fabe8d6-152b-4079-b166-4af97248795c · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Progressive Distillation for Fast Sampling of Diffusion Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.107271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:acb497b58fccb6cbffe1f41a3f443127e19d788561ab9c81293cf9105662d07a

Observation 5b2e0eaf-324d-406c-8563-d5df48d7859f · outbound

This paper cites Proximal Policy Optimization Algorithms.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Proximal Policy Optimization Algorithms

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.110053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:49b1974fb97295745a1d1cce66a0667997be5b2f1cd966dbe2c47ea54ef9d6fc

Observation 79b08765-5bbd-4c67-8f7e-e078500d19f6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.112867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:ff8e712a0040e455905ce9ebae6914d453b565307a385a3d6fad51c75f081858

Observation ddab1155-8cac-45b3-9b25-31fd705afceb · outbound

This paper cites Denoising Diffusion Implicit Models.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Denoising Diffusion Implicit Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.115280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:6421f485f1f5b2dd1a0100880199fadeece01f53ded772220f3b78988a2d0f03

Observation ba742b2c-17a6-4821-91f3-907da8cd859d · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Score-Based Generative Modeling through Stochastic Differential Equations

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.118466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:7670bbbbbf1955a96fbb009859be61d065354ec4b916b50900cccc5463528834

Observation f281d133-e530-455f-b482-af198a50448d · outbound

This paper cites Hunyuanvideo 1.5 technical report.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Hunyuanvideo 1.5 technical report

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.190802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:b9ecca3327f841d0d69205a78f3d6e9f02010dbbdbcc5d3a7c6365c16379146e

Observation f6720920-3ec0-48c8-ae9e-3d68f885258f · outbound

This paper cites Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.062590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:655600def0232e181538059747d47d12b9f199d2cc46a3cd0cc4716cc3e2d7b9

Observation b1c37ea5-0fd9-4777-8225-44f2c5a59d27 · outbound

This paper cites Diffusion model align- ment using direct preference optimization.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Diffusion model align- ment using direct preference optimization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.125663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:6cce47a30bb14d0a2302051d96117ef8ebdec55f583228df5b9990325d983068

Observation d0d7be2e-637a-4acb-b18b-4f9005da4c64 · outbound

This paper cites Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.069638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:d6e8924e81d2e0595f17cd0cbeb443ffefa7804ed65fbf4de33c38e441e239dc

Observation f232023b-6c13-46b4-9a5b-7626c418d429 · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Unified Reward Model for Multimodal Understanding and Generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:44:31.060111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:e267464d2aad4c74d5e6f194428371ea740060c545b77d2a8c1a4fdf45fae2d2

Observation 56045798-c251-4190-9854-cecb6b809622 · outbound

This paper cites Reward hacking in reinforcement learning.lil- ianweng.github.io.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Reward hacking in reinforcement learning.lil- ianweng.github.io

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.127551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:f971bcc7cdfdb70f8abe3353f329af116d7c6249ae8520d27dab95e527278ff8

Observation aa41ecd4-f187-471c-8c5d-4737894c39c5 · outbound

This paper cites RewardDance: Reward Scaling in Visual Generation.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE RewardDance: Reward Scaling in Visual Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.078729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:33b6571595de5f567d7325d307c175454d4aa669aa8b602c0d1658720ab1885c

Observation f15c3b90-4822-4e3d-8280-d9a7e5301ab8 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.081665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:b443d1d54d309dbcc4436d35a8866610fc2140433ee57fc50658c7f2f5e57318

Observation a3aafaa2-2b7b-412b-bb95-778be32df14c · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.129615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:20c3dc7f5854e709995500b264c291c556cf03bf16a057038b9158b03879b0fd

Observation b523e086-e4e2-4522-80f8-1bcec5691206 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE DanceGRPO: Unleashing GRPO on Visual Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:27:50.087125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:697536308d0bf6c97606bf7ec5b04e47c85be35eb48d6e60408f3575377f5783

Observation d60bd1e5-c1e0-42d3-a2d0-5cc42091e832 · outbound

This paper cites Discussion on flow-grpo issue 7.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Discussion on flow-grpo issue 7

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.131677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:b9bf16d7e8259327d78cddaab0bc24789268eee2a541cf726bacb785d5e434fc

Observation 00afcc67-13fa-46d3-a932-a552faefd466 · outbound

This paper cites One-step diffusion with distribution matching distillation.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE One-step diffusion with distribution matching distillation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.133598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:dbd33340b99a7ae2c4e6a6abfaa2a4efd529d6f102160cf09547295f0f5dcd8a

Observation 42ca2359-5bf1-43c3-9a3d-057d3c58c765 · outbound

This paper cites Self-play fine-tuning of diffusion models for text-to-image generation.Advances in Neural Information Processing Sys- tems, 37:73366–73398.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Self-play fine-tuning of diffusion models for text-to-image generation.Advances in Neural Information Processing Sys- tems, 37:73366–73398

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.135487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:3e2149ab60b1bfac93e6165eb4c08356e6857088734232f5f641caecb835dffa

Observation c12589b4-64f3-4f19-a2dd-c1c6407e1ad4 · outbound

This paper cites Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Unipc: A unified predictor-corrector framework for fast sampling of diffusion models.Advances in Neural Information Processing Systems, 36:49842–49869

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.137623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:5e9c9668e9caac5ab7a4adc0aa468edde6e42a6436d08d85cb66b1c07f555a09

Observation 1b81e192-be19-4181-9d68-9bb09547c4ed · outbound

This paper cites Dpm- solver-v3: Improved diffusion ode solver with empirical model statistics.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Dpm- solver-v3: Improved diffusion ode solver with empirical model statistics

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.139412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:746f106e595fed599fb297c63a44cccc02ba751348e83274c4c2bee8cb70f7df

Observation ac18995e-0f18-415e-8f98-746c89367d39 · outbound

This paper cites Drivinggen: A comprehensive benchmark for generative video world models in autonomous driving.arXiv preprint arXiv:2601.01528.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Drivinggen: A comprehensive benchmark for generative video world models in autonomous driving.arXiv preprint arXiv:2601.01528

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.104527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:eb1194217ced3be85cf2c24a71867c43252f0d75e0ea087fec65b7fe713e9616

Observation 82c93157-c4d1-4944-a924-e379b49fb646 · outbound

This paper cites (7) has the same convergence as Eq.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE (7) has the same convergence as Eq

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.141448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:ff513e9c1c76fb71ad7158ff86552da437701a75f26e8ed8f117dfc8721d2e6b

Observation b047cdc8-b596-44ab-8bfc-8596ca35bb50 · outbound

This paper cites We denote the discrete time steps by an index i∈ {0,1.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE We denote the discrete time steps by an index i∈ {0,1

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.143652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:3e7752a3e5c95c332e10d76aff4717c58cee1d3c9e9a11ccead744c248e4e3c5

Observation 77cf1c3d-6c5e-4b67-ac12-9dcf247e591d · outbound

This paper cites an unresolved cited work.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-13T13:27:50.145255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:bf75f3a2fc6178712404363d72d8b3da3b50547ef21af68b59170850ef197a7d

Observation 8799ed57-928a-4b13-b04d-09ecfcfa136d · outbound

This paper cites an unresolved cited work.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-13T13:27:50.147017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:af6382be7fbc0439b0aa824c0073d84dcc4a725d5a75f7b64d3f1aa16074943c

Observation 360e075c-a856-46e1-bd30-fa0574ac969f · outbound

This paper cites We established two reciprocal settings to evaluate both in-domain (ID) and out-of-domain (OOD) performance.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE We established two reciprocal settings to evaluate both in-domain (ID) and out-of-domain (OOD) performance

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.148841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:bf58ff18548cf47c47e4320ea6c05a0db14c9d033275393817aa373a81b4de1a

Observation 3b000dbf-fad0-443b-bc8b-4c528505c820 · outbound

This paper cites an unresolved cited work.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Unresolved cited work

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.121066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:8e4f14e856eeb2fbef2339fb5053288893811895f27dc8d96d76987dc598aea3

Observation 63ddb5cc-a635-437f-999b-d36468793732 · outbound

This paper cites PROMPT: 16-year-old teenager wearing a white bear-ear hat with a smirk on their face.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE PROMPT: 16-year-old teenager wearing a white bear-ear hat with a smirk on their face

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:27:50.150581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:568e8b120821f202722042c4feeb18f68a616e3665a3e9c3048450abba7f574e

Pith citing papers

Observation 8868ede2-1a2b-4b50-80fc-f0d65d767f45 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 277

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:24.718091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:c966c661421993d7fcf0d78f5c9e52e2ff4d55e46c63ce8373b93c33d94544c5

Observation cd369139-8524-440d-b00a-1c51216cf953 · inbound

Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling cites this paper.

Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:30:39.200519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T21:30:24.268903Z digest=sha256:4ac446b565102c160fcafca29f2723fd3fdbb371c9a4f7f80876f43ddaab13c4

Observation 9245787c-c16d-41c4-8bd6-20dbe174cbb2 · inbound

HunyuanImage 3.0 Technical Report cites this paper.

HunyuanImage 3.0 Technical Report MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:02:32.872717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:02:32.806844Z digest=sha256:6cc9d04d3b887ffe5fc024d3b03a5070e71e051ed93056b756a8edfdc192fec8

Observation a398429a-adbb-4bf3-9d43-b4c12e721b91 · inbound

HunyuanImage 3.0 Technical Report cites this paper.

HunyuanImage 3.0 Technical Report MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:14.452986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:14.452986Z digest=sha256:42e604cbbebfda985e8b5ea1ac94d0cae37eae4129b6e691f32f6fc8cba81fc2

Observation a334ff87-f36a-48ed-88b5-8e58dd71a754 · inbound

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization cites this paper.

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:50:33.861376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T19:47:48.545820Z digest=sha256:83e1c5cb8ec65081489d24c0a8f912b6441b5b3101ce9906deea3e04febc5632

Observation 3e628352-3919-484a-8e72-2c47f290b740 · inbound

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation cites this paper.

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:44:18.933610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T18:42:29.327151Z digest=sha256:628b0d107d69e0db2653a10557dbc1f563e9398004d6f3340d036641dc6ab58d

Observation a53b66a3-d14b-4c2e-9068-81b22b044828 · inbound

HunyuanVideo 1.5 Technical Report cites this paper.

HunyuanVideo 1.5 Technical Report MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:31:35.086217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:29fc4460c9a1cf325c9d17e7b4aa89dc1d5b988c47ef7525ed34d9c6eb9e22a9

Observation 5db676d6-2ab4-4175-bdba-39f7c6d3850f · inbound

Rethinking Reward Signals in Video GRPO: When Scores Become Targets cites this paper.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.500523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.500523Z digest=sha256:524551c1ba1db7f7cd086392565e96726b6a7633b096bba03c65b8a7b5034245

Observation cde24e99-d1f6-4383-bea0-86b457ff8e3d · inbound

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO cites this paper.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.354792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.354792Z digest=sha256:9b8e91b779bcbc24eb22d512ea5ce537393ce12062939ffcd406048315e523c2

Observation 3debb3f2-0a17-447f-868d-d5c1937d0c32 · inbound

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing cites this paper.

Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T19:14:08.102419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:14:08.102419Z digest=sha256:4d32e6311ce99bae2d815e7bcb56acee0c2233c3f1f3dc3512c2661e665f3c2e

Observation b53cd30e-aeaf-490c-a020-e02ce5c46acd · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:55.487831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:55.487831Z digest=sha256:7a45a28c2fea85874a747687adaa8ff1137d2306e0dbd9e77be6cafce10417d8

Observation 6a73b0f2-fc5c-481e-929e-63ce7e7e413e · inbound

AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization cites this paper.

AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T23:09:50.674516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:09:50.674516Z digest=sha256:040ac54e594c47569e056e84f8d1ffdbe4fc0386dd69480b77d49116d0554c97

Observation 72b55ceb-f0de-40cf-ad23-fbfb41e5554c · inbound

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning cites this paper.

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:08:25.877291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:05:45.781048Z digest=sha256:92313f0b1a773b7b401aa2bc42772c1b7829b79d2bc0dbd8a80fa19f564b705f

Observation 8806594b-59f8-4b3f-bcaa-ac4c1f1187f3 · inbound

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning cites this paper.

CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:44:47.566144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:41:37.140709Z digest=sha256:6244ee5aa313392e3839f8c188c3df0f379fbfe9ae8832538ceb3c1e6b3710fb

Observation 058c0b06-edd5-4b43-bf49-2860a935b513 · inbound

YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance cites this paper.

YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:23:22.732115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:21:26.758816Z digest=sha256:a6966da2a6962fe6b964648f430906029cf4464f77daa27b236f36a18cebeb8b

Observation 2769c7b5-5aa3-4a4e-a357-fdf1b99b5d2c · inbound

YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance cites this paper.

YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T18:44:48.222670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:44:48.222670Z digest=sha256:0c41b5f635d668522645d32aa535d42bc10716a5fcb64b82bc90bfc920b994e8

Observation aefb9526-ba03-48b6-b838-d3c4a1c12c5f · inbound

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models cites this paper.

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T16:48:02.901980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T16:46:30.674244Z digest=sha256:3f6ca450ad5378ebf422680f2a35774022c928002dc37cb8198f29f43329e680

Observation d6128819-ff9e-4a21-872c-b4bddfb0b596 · inbound

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling cites this paper.

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:10:06.994557Z digest=sha256:48073d8529d5ca5ab6fce7a3c203900f392bb9556f6eb1c01517b0a8b333dbbe

Observation b1195fd9-08fd-4bf0-bb35-39b6fc714e87 · inbound

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation cites this paper.

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:00:50.105629Z digest=sha256:a61391ce9cc45265379e638d6785d5e1ad0a349964314cc5e4f19843b2e43585

Observation 4657ff26-0101-4a86-8ad3-22f841af5385 · inbound

Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing cites this paper.

Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:56:09.398673Z digest=sha256:fe29fd4263f51194c8700a9448a280390cb601af58f7ea3ad1548750250593d4

Observation b77e20c1-86dd-4a3d-9708-aaa4beea0bd4 · inbound

Reward-Aware Trajectory Shaping for Few-step Visual Generation cites this paper.

Reward-Aware Trajectory Shaping for Few-step Visual Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:43:55.499474Z digest=sha256:04a167f4e945bb1fde7985a743db745eceb0e454de2dc2d74c39a600da92b73f

Observation 2f7709ec-8759-42de-b918-7f75fbf299da · inbound

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories cites this paper.

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:23:28.424453Z digest=sha256:47f939e4b510f4c5c608ae1b26b7f6ae4666c5b5626a8a2a02f29c53b7ce99e0

Observation c6338d41-723b-47c8-84b2-19ed48105d2e · inbound

Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models cites this paper.

Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:34:01.064008Z digest=sha256:7fd96459694764668646de72c74631a5cf2b4e454215d92a497ad2d86bcad32f

Observation 7977d59e-bc2c-48ea-9515-d235b0431179 · inbound

Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models cites this paper.

Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.864977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-05T17:45:49.335925Z digest=sha256:92e45c9f3e100702b4b64ae425d27fa5490e2de111b447754c65aad7fe23e8d1

Observation 88b04ae8-eea4-4a3f-ad7b-760a37799735 · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:32:54.235578Z digest=sha256:2fd8bc1f99f9b2130facf3e83aaf6281036dc3b1328368ecc09ee01e69d01b50

Observation b81c3fca-2bca-4437-979b-38122cd35747 · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.741472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-05T11:32:38.636335Z digest=sha256:f4d04dbfc2b584dbb7b6ac629d68d1191fa890fe6f5ae297dd46e79afe5ca345

Observation e3c64190-649a-461e-bfc7-1c22d6f3af1d · inbound

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning cites this paper.

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:45:35.600729Z digest=sha256:8c4175752067a97751bd95751cb7930591c37a58add64223349a36340b2d3622

Observation 90700c28-8da6-4ce7-a47a-60532904e19c · inbound

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation cites this paper.

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:10:26.580964Z digest=sha256:54746769b3ff7d2303be9d0b8536e9b01ea8237e2603fd8d5855b691473ef01e

Observation bba641cb-c427-41bf-9519-cd9c2fdfe80a · inbound

ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control cites this paper.

ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:20:07.566207Z digest=sha256:9d552313a0988670c77a479e0b83fb284530b09bb1e2396de518b6ce475a87ba

Observation bfd9da75-e88f-4a31-adab-0d2c1ca98e54 · inbound

V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think cites this paper.

V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:27:17.629238Z digest=sha256:3a244ea925bd87c4c62ca039b1997910d47092c2d62bc54f52c47e23bd686772

Observation 3f1a5258-6228-419a-84c8-3d07a7c769f9 · inbound

POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation cites this paper.

POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:37:41.629935Z digest=sha256:c9228f56442c601ceafa17bdfcd5b0a96f11d9479cd4ae7f1ffbfb014cec7dda

Observation 736447ae-cdb4-46ed-abd0-37fdbfa90f61 · inbound

A Systematic Post-Train Framework for Video Generation cites this paper.

A Systematic Post-Train Framework for Video Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:58:33.014401Z digest=sha256:0f2a3e5c5242c9749b01356aefe2d62341229566d04b5a45a87f4966060b53f8

Observation a338c0d7-c605-4c99-b51b-1e3dca32e9af · inbound

Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline cites this paper.

Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T09:45:00.759474Z digest=sha256:ded350dbdd897556bdaa86cd25bc334655a5022ce60823d073b3cd79a5e34067

Observation bcdb6f22-f8a1-47fe-9494-6e1a8e4d7ab4 · inbound

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers cites this paper.

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:58:25.291448Z digest=sha256:f62697d41bfc2c110d527ff9607687d556b21a81f8c3df2b37fc766085e2845e

Observation 8fc19dc2-52cf-42ed-bf8a-91bfd7807841 · inbound

From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data cites this paper.

From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:22:55.682373Z digest=sha256:795ec96404ace365d870e729e2e1c469f662e8aacfa3774d99ed32f63ceb53a3

Observation 93266165-4672-47ae-b828-907094e9e052 · inbound

From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data cites this paper.

From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T14:39:01.164974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:39:01.164974Z digest=sha256:d593b867b73b5f68310c157adbae1b2a44eaae16069c24a323dc113a2118d605

Observation d57b77dc-2fd1-4040-9b6e-0da1eb9e5bc7 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:04:16.335479Z digest=sha256:5e8d06ac50d449e0a7cf1108fbff0df834e4dd2bd71f20ed96114bfa8d18af95

Observation 2c9671d1-f1f9-4855-a4e7-cdf8b1d98f26 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:17:08.955601Z digest=sha256:83e7e20430f95f9f7902ebb191ff66370b128eaf9d3f94e08137be400300faa2

Observation 5597444f-8c92-4e6a-b515-67b915947f29 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:55:05.190840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T05:52:26.736304Z digest=sha256:e379d912432df12d6405050e45fbced82bf0402eeffb5a6aa40b3d09cd10c0e6

Observation 6c153ea5-bd1d-4d84-b42e-757fcff11932 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:43:51.732869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:39:32.496303Z digest=sha256:d45c38c58cc04b5ba4b6249383bf34384b44c36940049810f87735b97c946ce0

Observation 468b2dac-b30a-4bec-88eb-4ca16c29c1ce · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:05:07.197061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:02:29.120150Z digest=sha256:3aa4ec5dea85e83c0bbb8c1a411142ad08b38cae0eefde03579863a6f6f87770

Observation ce58009f-87c4-4e77-8b9a-01e38797136e · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:388c70a0d19cf9d53d7ab5f6e1a5bd0e4dbded805fad55bcfb8d56d863484e50

Observation f9cf8ee5-bc33-4dcd-afc6-d767ea10d7d3 · inbound

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment cites this paper.

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:36:16.811765Z digest=sha256:b9211a503ce1ab1b763c925bcdd088d932ace16483101e558ff7640eb76d3129

Observation 73ca1cca-241d-4049-b55e-7618a2156398 · inbound

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment cites this paper.

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:03:02.895835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:01:21.695270Z digest=sha256:146d2faa937e461eaba346398c9884c5e25dec90e98b6ff898a21367d171ead9

Observation 5f30d301-1923-4535-a409-a8a487eebc22 · inbound

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy cites this paper.

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:50:13.653022Z digest=sha256:4f7aadfb5a7000ce9dcf64237f2abc5420408f0db93628e3d29c967e729cd7e9

Observation 32ae3a70-7185-4e87-a0fc-c041f32ff129 · inbound

Cutting rules in strong field QED with application to trident pair production cites this paper.

Cutting rules in strong field QED with application to trident pair production MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T16:57:23.214654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:57:23.214654Z digest=sha256:51af7fdd1f8d94a9024d2ce3decca5304630ee7b6195ac18d12c63c32af9d507

Observation 855ab3cb-01ab-4590-87a2-c7f8f7a414ae · inbound

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation cites this paper.

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.191417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:16:20.612291Z digest=sha256:cc92f4984d5fdbe45825054716ff9f046727575c366ab57bd71f17d6dbae921a

Observation 468f52a0-de5f-4cb4-b2c3-dccf3828645c · inbound

CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL cites this paper.

CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:43:33.349146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:42:08.454469Z digest=sha256:1c1aa719136051915c5889645ebf34942e8942698d94d8b0f630c7447b2b0ec6

Observation 55647549-b283-424b-9c35-f1bfbd389a0b · inbound

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models cites this paper.

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.847896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:20:23.420424Z digest=sha256:03d340964b250b4d09dac698bfff668232f7054345a3e60c0fb7616d045c42b2

Observation d70f820d-e498-43bc-a274-9d435c689212 · inbound

Embedding-perturbed Exploration Preference Optimization for Flow Models cites this paper.

Embedding-perturbed Exploration Preference Optimization for Flow Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:38:53.005025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:33:52.933672Z digest=sha256:c6714ee89c44b7ba83ffcd921c58c3c2cefc1301876497886be9c4a1c1f31bee

Observation efb5db7c-1705-4114-aaba-935d92d56471 · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:28:52.878331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:28:06.253200Z digest=sha256:0d4660af84403fd1b512ce68ad164740af939d16d797b65962e62ef79ab370e6

Observation a05b6ca9-d166-4486-bcb5-78a9f618a840 · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:45:00.963959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:37:18.563704Z digest=sha256:de1d1bbe04e3ed308b33437d7a86a3801270e83c5fac304b56e92967b1d09469

Observation 9cc0f2aa-b08e-4eae-a45c-7badd3bee755 · inbound

Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing cites this paper.

Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:47:45.812492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T20:44:54.356658Z digest=sha256:5094b2d6bddb20a01a90db4530bd9bab79f4f9369445563d43ad4118a23e63b7

Observation 63515aaa-78d0-4b25-a5ce-13a7999f49db · inbound

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cites this paper.

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:08:13.618209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:06:09.367559Z digest=sha256:fa251287bd195a2ec581bb073208899a74aeffc83b3e8c0f88fe2f12292f14e7

Observation ba26cc4b-fcee-4cf5-b73b-35389681885b · inbound

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data cites this paper.

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:13:05.210482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:13:01.821585Z digest=sha256:b6d47c0478028ca303437f4c95e2c8e5b7aceb167a619ad0d57cf45f9c29144f

Observation f8a0a477-8269-4383-a70f-699545078023 · inbound

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data cites this paper.

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:14:59.853209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:12:11.972394Z digest=sha256:74e1bfa6661d6866ab169a763be6cd4fb6ed59394b0d245c2c367f6697f7d388

Observation 1752c659-afff-4df8-939c-1aec8eda54e4 · inbound

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models cites this paper.

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:34:46.735661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:34:14.596976Z digest=sha256:b6ddb07b5271caa5d3166df0f672f4f6b8c535cb9a7e1aa18ebc971587cf11e8

Observation df2e6ac4-d499-4423-8150-6227bcdf5474 · inbound

Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models cites this paper.

Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:05:22.231265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:05:09.178674Z digest=sha256:1cde3c5269148ea72a112af35bf085f914320eba70c5396b90b4cd4b3d54f722

Observation c528e01b-bc56-4c96-a91f-4d0b9bd8f1ff · inbound

Geo-Align: Video Generation Alignment via Metric Geometry Reward cites this paper.

Geo-Align: Video Generation Alignment via Metric Geometry Reward MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:16:35.919678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:15:58.667672Z digest=sha256:14a7dcca4f89f97c54eb63b76d0bdcabee6f986c0778c1316598c989f46f069d

Observation cd05b501-623e-4bdb-a0c9-479db4f63265 · inbound

DRM: Diffusion-based Reward Model With Step-wise Guidance cites this paper.

DRM: Diffusion-based Reward Model With Step-wise Guidance MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.368093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:12:39.225557Z digest=sha256:61fe9b2fa12d0bf773630d5f04416998526b650301a81951a91e9652b0efce69

Observation b4459e72-fced-4a42-a3fc-3296ee2c0b17 · inbound

AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models cites this paper.

AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:01.609833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:44:25.152451Z digest=sha256:0b3438feee8c3d948ae89e49fb75bb8ff6934366a0184bcabd4828f3247511f4

Observation bd0106ab-4678-4d3f-8003-e469bb976eaa · inbound

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching cites this paper.

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:24:00.289147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:21:06.005125Z digest=sha256:1e6a77cbf29f76e1c971ed9781de04f087b3904313f56198f24c8a88cc0b26e9

Observation a0129766-898d-45d3-8dfc-f094de0ae7c7 · inbound

Explicit Critic Guidance for Aligning Diffusion Models cites this paper.

Explicit Critic Guidance for Aligning Diffusion Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.784073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:18:46.456767Z digest=sha256:1af64da3ccb720447fba2591d0ad9f74f176ae659c07f330940e3f46142361e8

Observation 2d926ce6-48ee-481b-9013-ece376fcfbb9 · inbound

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning cites this paper.

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:13:27.508475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:06:02.201607Z digest=sha256:6ea2df149111f4700281eed275c9d2bfb78306eb40cdf2aedf32ce1b1914e9d8

Observation 37a9e35d-a975-4f49-bfbc-c02194b2263b · inbound

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition cites this paper.

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:16:16.516711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:33:30.182150Z digest=sha256:f2d8bb22801ff9cf926fc3159d39cbdaacaa55ca85e5f54953b58131c8a1a376

Observation 437e2a9d-64fa-4988-af77-6536d4439b9e · inbound

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching cites this paper.

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.571746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:07:29.897089Z digest=sha256:8d502c5b9660674c74e1201c06c3eb71ed2e85da9e32ee3bc599e1eab5bc95d1

Observation a72b6175-15ee-47cc-90b7-7df0a9ad454c · inbound

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO cites this paper.

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:17:08.889278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:54:21.740892Z digest=sha256:44bfa2e6116bad1eca6fe69b380b8847ed1ee329caec48e2a7b3103ded91a8ad

Observation 0e5f2a0e-f460-406a-97ab-d3c6796ba186 · inbound

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models cites this paper.

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:57:38.720346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T14:20:02.171646Z digest=sha256:0bed2e225ba77fa02a82fca8333c3b8c291b61a7be8b516644196c1e70443b3e

Observation 61854866-0511-4415-a68f-919c5c15d9ac · inbound

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models cites this paper.

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:54:36.260111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:52:23.199230Z digest=sha256:652294666e1cd8b24d18fb3f1e3588a10740733ec7b9b07ea18814ef718e1d0a

Observation 609fe4fd-6705-4c20-9a2e-e2f9d73ebfe8 · inbound

Exploring the Design Space of Reward Backpropagation for Flow Matching cites this paper.

Exploring the Design Space of Reward Backpropagation for Flow Matching MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:07:36.933061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T14:10:05.247255Z digest=sha256:0a3200fcf56d2f98cba9e7e3ee911cbbb3e7f2300188417ab3a1c98f87eb37da

Observation 48ba4b77-2f01-49bd-af98-c50553121878 · inbound

World Model Self-Distillation: Training World Models to Solve General Tasks cites this paper.

World Model Self-Distillation: Training World Models to Solve General Tasks MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:07:55.841847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:16:35.511426Z digest=sha256:ca81ae1153deec2d74c701d9a118dafcf7117f4fb362c6deb54f82df7c62a2dd

Observation 70a60581-f4d6-40d4-a357-69d25c55e494 · inbound

NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment cites this paper.

NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.221211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:26:38.142556Z digest=sha256:4f2b73830b1468ab9ecf7d287cf7e324b51f9b20966e61ffabf06b14a5be4c52

Observation cad39056-3b68-460c-baba-dbdddb955a15 · inbound

NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment cites this paper.

NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:54:36.441526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:50:59.642305Z digest=sha256:157bf17575048344f47ff4fafad16739cb95277f4229518ba5522a682e1e8797

Observation 9fae1f0b-4bcc-49b0-9fd9-b6bee223b777 · inbound

ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL cites this paper.

ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.588700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:17:04.368521Z digest=sha256:3a4532687fb677ecb5bd83aa663cdeb780d2d141e8a5c3899b2c095a0bec5026

Observation 5062f946-8457-42e9-ae6b-1b72b58035e6 · inbound

ReFPO: Reflow Regularization for Flow Matching Policy Gradients cites this paper.

ReFPO: Reflow Regularization for Flow Matching Policy Gradients MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.820856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T14:38:54.754049Z digest=sha256:69e52c40d06dee6a18a0c3b161a85087143d031c0f09fe70d8b9dfb35ec16507

Observation 5624ce66-2b0c-4222-982d-456237860547 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:19:49.613140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:b09498eae46772ffa90286134c2c2a0e52c8975295375e101b294ea12881a1a2

Observation 76eea894-a3c4-41fc-a8b5-5126fd13f79d · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:8958b7af7a7b3154ee3b82a2efcb418e7f810c0a061cc68909925a12838de691

Observation a2707351-81d3-434d-8039-c8d1d9778a81 · inbound

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation cites this paper.

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:39:45.613142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:37:11.052296Z digest=sha256:0e54c9e345a067695df6e87117d36bffa356cfb3c1355f7cbec382416817ade8

Observation e3627ee1-7ffe-420d-aa21-474501174cef · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 200

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:59:58.067786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:6aceaba58283c2e5349606c71326810709b8fd1507e43848adbc0070644b8b08

Observation b9e3b56f-6cd3-4e37-bcec-ca53fad28eb4 · inbound

PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction cites this paper.

PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:13:49.499985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T05:11:09.189141Z digest=sha256:8c028f9da3dec7991b4d4f0a1bda19f8e03364a07a08f1278ed75b6606720e62

Observation 1fc2db18-0be0-43bf-964d-710ba3fadc86 · inbound

TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL cites this paper.

TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:13:53.597642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:46:48.414652Z digest=sha256:bbd93dbe55d85701ff5b5d8e977deb5ec169425bc5d2a63ba049802315dd36e9

Observation 6ab7d8e5-3320-4773-a77c-4dec285beb5d · inbound

TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL cites this paper.

TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:39:01.000105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T22:38:33.557109Z digest=sha256:4cd4e5978730999b22a1a4a83435ae3cb7417aa2aa330119e828fb10f0e38cc5

Observation d3310fa1-5aeb-4d6b-936a-80a7e48e8dd3 · inbound

Dual-Flow Reinforcement Learning with State-Aware Exploration cites this paper.

Dual-Flow Reinforcement Learning with State-Aware Exploration MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:24:21.809233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:16:11.130271Z digest=sha256:b5453db41e3b838b5f0819717db1a3d35ec5ce4f6cb4f9a371da62fdd50bc7a8

Observation 76415a70-29e8-403c-af69-5e4bde012122 · inbound

FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification cites this paper.

FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:14:20.929455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:10:44.244790Z digest=sha256:9a92144357f28822098aa9623e20d06f36d3e96144e17430b9eb8577b4fad216

Observation 8eb68b51-b172-464a-badd-d9190ae1d604 · inbound

Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition cites this paper.

Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:07:08.231157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T16:04:19.371174Z digest=sha256:41a4b5aadb46cc46a97e4b494d940215d93465ee79019347a06dfe6be49ddfb4

Observation 12a2d7a1-f1aa-40e8-8727-6e7ffa67f43f · inbound

Optimizing Visual Generative Models via Distribution-wise Rewards cites this paper.

Optimizing Visual Generative Models via Distribution-wise Rewards MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:48:39.749308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T16:39:12.711424Z digest=sha256:85600b511756705e6acf7967e38d090616387d42d692503170e58466f53ede3e

Observation eeef1473-3eaf-4837-88a0-1ef8999344ec · inbound

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence cites this paper.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.325140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:94f9cbd60c5fc28928fef1eede052ea7acad09ff7e713875af71790e40c4cde1

Observation 9a09be2e-4f18-419a-82e3-e1b14bb4f540 · inbound

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs cites this paper.

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:58.829878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:58.829878Z digest=sha256:10406ae6cd78fe34344bd122b263b531b802eeec6b6de3e238c0a5931b044e1b

Observation d44aa620-02de-4478-9c35-374604cc2efd · inbound

JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models cites this paper.

JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T17:41:00.353583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:41:00.353583Z digest=sha256:8d1599aaf0fdce7e8246cdc7b903c24d3b0d526d30cfb42d526906f893869a52