Pith. sign in

Paper Citation Record · LEDGER

Flow Matching Policy Gradients

As of 11 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 37 inbound Pith citation observations for arXiv:2507.21053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21053 v2

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:07:09.738381Z

measured 121 of 121 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:50:11.728877Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:19:37.772072Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved59
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d48a4f4-865f-4fbc-80ec-061b0d8d4127 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Flow Matching Policy Gradients Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.474916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.474916Z digest=sha256:be4dd6ea8e4626d51356639c5179be86ea41600e3d4f387ada417a35eecb9f16

Observation 1cf0b7e5-1090-468e-92ec-5c429fda7fe1 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Flow Matching Policy Gradients Photorealistic text-to-image diffusion models with deep language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.479145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.479145Z digest=sha256:3733551e7c9578deb2325141e71b602de388e688708870551d957f77149d1b89

Observation 17e532ea-88af-460f-b415-c19b93531e5b · outbound

This paper cites Video generation models as world simulators.

Flow Matching Policy Gradients Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.485898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.485898Z digest=sha256:037ba38767747414133acd044205334f0f73f1da0ae58225bfd4efdfe5308677

Observation 73329125-5b0e-4757-a629-9c0b07f1bb92 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Flow Matching Policy Gradients Movie Gen: A Cast of Media Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.488679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.488679Z digest=sha256:7ab92527eaaf607dd11729980c6187486aa33c4c608a538eaa581332a8901696

Observation 080afe53-3d90-4311-80d9-af33d5eb7f3a · outbound

This paper cites an unresolved cited work.

Flow Matching Policy Gradients Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.492400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.492400Z digest=sha256:fe16940a78a7703c692873432d5380ae0a876281d88e5cbbc013cf5073f6b036

Observation 511d80a0-aa90-46af-93dc-954e0950af30 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Flow Matching Policy Gradients AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.495265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.495265Z digest=sha256:e0082c7011eff5a981061f005ccb3ceadd91053dcd655c3c808358b08302d337

Observation 9bb4e091-57f8-4c23-8170-c196d80b0b7d · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

Flow Matching Policy Gradients DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.498444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.498444Z digest=sha256:79340de704c5e28f92de5c8bd67cfc30f3658f11eac284dfcc7996e8d32d853e

Observation 1b9295c3-0ca0-42bb-bc66-3eeb1c3d8dbc · outbound

This paper cites Action-Minimization Meets Generative Modeling: Efficient Transition Path Sampling with the Onsager-Machlup Functional.

Flow Matching Policy Gradients Action-Minimization Meets Generative Modeling: Efficient Transition Path Sampling with the Onsager-Machlup Functional

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.505195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.505195Z digest=sha256:6b62cf622fdabbd71fd4dfee1221e8543c6904d4af9d2d9ff0e4a0801d5c5ec8

Observation 87d264a7-87e9-4ca4-9188-a11c629ce240 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Flow Matching Policy Gradients SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.508045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.508045Z digest=sha256:c368e45f347a1d1b29dda66ca53f97b7cac5266410a2d0ef0cca27305bc8acff

Observation 2e2ad645-3fe4-42bc-8729-bdd904b89985 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Flow Matching Policy Gradients DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.511165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.511165Z digest=sha256:659449c80244c165cbb1b9201f9836108377bb2e44049ee58bbb8491904c016a

Observation 859eba96-90f8-4d58-a110-993b68a50282 · outbound

This paper cites Flow Matching for Generative Modeling.

Flow Matching Policy Gradients Flow Matching for Generative Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.513926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.513926Z digest=sha256:297d0d9471e4c74728180f2fdd17a01304db897bad364cbae94fd5e06cb16e63

Observation 91c46414-c837-4d29-8744-b3125bb80fa6 · outbound

This paper cites MuJoCo Playground.

Flow Matching Policy Gradients MuJoCo Playground

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.516529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.516529Z digest=sha256:ed1a29fa05d1506f2bb641462d22f21443245a802d6303b668c18b04a27105a6

Observation 21c31ee7-1365-45fa-aa20-5e285033cb70 · outbound

This paper cites Sutton, David McAllester, Satinder P.

Flow Matching Policy Gradients Sutton, David McAllester, Satinder P

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.647015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.519371Z digest=sha256:563466aa049c7467c1e1c2f2d6f9393c3f3079b04bd97291a7cf6b14744512fd

Observation 48973bbe-0726-4fad-833d-1556a987a731 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.

Flow Matching Policy Gradients Simple statistical gradient-following algorithms for connectionist reinforce- ment learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.636777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.522696Z digest=sha256:4bcbf99227734b1bf52c74a427d1fe453f76d774e37f93e6df519d3c3310a735

Observation 0c314530-806c-4eb5-9392-de4e099117ec · outbound

This paper cites an unresolved cited work.

Flow Matching Policy Gradients Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:07:10.627228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.526084Z digest=sha256:36486d97241c612816053db5cbbea5fac25010f4eff0e8795caab5a2188ae28f

Observation 495be691-4082-452c-aa65-6fa5c7ccfafb · outbound

This paper cites Natural actor–critic.

Flow Matching Policy Gradients Natural actor–critic

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.615614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.529159Z digest=sha256:b785132e971224b34940fe9b8a19ea97425c27ebf7cdfdf579ca36d192f12f47

Observation c942879b-aa84-40dc-8ebe-48948c74ef51 · outbound

This paper cites Trust region policy optimization.

Flow Matching Policy Gradients Trust region policy optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.604985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.532677Z digest=sha256:22120cd9f19840af37801b48c0b5ab3bc514b63313764a065ac7cf9ce35c9f1b

Observation 186eca33-80ff-48ef-bf90-69f4bacada81 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Flow Matching Policy Gradients Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.535762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.535762Z digest=sha256:b0d851ea9638121d82085b455826dfa8eeee99ea5923a5bc6c0f2f8259e21fac

Observation 3ba1e354-460b-4978-8393-5be06c232d93 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Flow Matching Policy Gradients Asynchronous methods for deep reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.592842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.538837Z digest=sha256:3a71c4bd9801da37c9c9e9f6e683b9c21ee85387dbc1d3c070422a961b34aea2

Observation ed193931-a53c-4b6b-a351-c52365a42aea · outbound

This paper cites Sample efficient actor–critic with experience replay.

Flow Matching Policy Gradients Sample efficient actor–critic with experience replay

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.581050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.542629Z digest=sha256:d4cf22ffad3fea13b6b6228894fa1f4a5ef0f11f0c16b6ef8c898e277a5dd5f8

Observation c982a532-ee28-4f22-a50f-5a19931f6c5e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Flow Matching Policy Gradients DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.545715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.545715Z digest=sha256:4cc6c3a6e8f29c8472be5d97dcdbd990ca325e87bc04df2a6ef225535963c067

Observation 779f8cd8-dee5-4ecc-a2e7-4e9df2984d21 · outbound

This paper cites Benchmarking deep reinforcement learning for continuous control.

Flow Matching Policy Gradients Benchmarking deep reinforcement learning for continuous control

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.570353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.548913Z digest=sha256:8108aeb5aa894e8a151ff45be5458b657ed9eb16ae8651fc82244df7a889543f

Observation 86d4becd-bdb6-4f38-a85a-ec1a24722036 · outbound

This paper cites Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning.

Flow Matching Policy Gradients Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.551956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.551956Z digest=sha256:9345284c763ff8dc213443b0d01b683ac9dc845356ea3a6530c04190551bd30e

Observation 7708fd5f-4d27-4671-9520-997b8c14c45d · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

Flow Matching Policy Gradients Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.561389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.554808Z digest=sha256:8bae7ede4220e5534bab5f50145109133c9e429d59fbe168866bec4fe476c080

Observation e1edfbab-7162-44f1-97f3-925e696e1a8f · outbound

This paper cites Curiosity-driven learning of joint locomotion and manipulation tasks.

Flow Matching Policy Gradients Curiosity-driven learning of joint locomotion and manipulation tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.550166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.557355Z digest=sha256:f29f0a2e6ab1447bd55dc0fb19229c14b8346be4e7376b97beec5b37291c51ad

Observation c475e2b2-e639-485c-9083-183f583662c2 · outbound

This paper cites Sym- metry considerations for learning task symmetric robot policies.

Flow Matching Policy Gradients Sym- metry considerations for learning task symmetric robot policies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.560136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.560136Z digest=sha256:cfee8173a5e4394970e4a63382295aa348b01c433af6fd2a0acf6fa5ac4805c8

Observation 3e435712-f0cc-4ace-a2c4-196104c75814 · outbound

This paper cites Visual Imitation Enables Contextual Humanoid Control.

Flow Matching Policy Gradients Visual Imitation Enables Contextual Humanoid Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.563148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.563148Z digest=sha256:cfd35c670f3b5791fdab018250750c85dd80205baba0b3abdad246c4aacaa26b

Observation 4d984d50-21bc-48ce-b02e-5150c11eec4a · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Flow Matching Policy Gradients Solving Rubik's Cube with a Robot Hand

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.566118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.566118Z digest=sha256:1fcff9cd39a1a4feeff7daa59322b5bea976a81ea9102d47f44af9fbe5649976

Observation cb9b6c92-0a5f-437c-94b8-78b23ddef145 · outbound

This paper cites A system for general in-hand object re-orientation.

Flow Matching Policy Gradients A system for general in-hand object re-orientation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.540286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.569282Z digest=sha256:ec9370193e1b07e49ba8639e6294dae12ccf570f02c40516d83a838efeec4431

Observation 8dd4c96f-4605-421d-b162-cc0972a884a1 · outbound

This paper cites General in-hand object rotation with vision and touch.

Flow Matching Policy Gradients General in-hand object rotation with vision and touch

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.572868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.572868Z digest=sha256:8e49b58bbef76dfcb7ffc65fd44c1aed2956348e2dc0a2213ec24e7920547a03

Observation 19f2bce2-c743-464c-a611-e38be38635cb · outbound

This paper cites From Simple to Complex Skills: The Case of In-Hand Object Reorientation.

Flow Matching Policy Gradients From Simple to Complex Skills: The Case of In-Hand Object Reorientation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.575849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.575849Z digest=sha256:e3c554dfd4be8c59703f127b263c5bc180319e0546ae007dd35fbd066c242343

Observation 3a61accb-3afc-4990-b998-80ffce8087cf · outbound

This paper cites Training language models to follow instructions with human feedback.

Flow Matching Policy Gradients Training language models to follow instructions with human feedback

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.522784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.579022Z digest=sha256:773eb709c384969a9ba2bbb89c08ed5a8089fd44f8c62cd491dccf85afd5b382

Observation b146a6b7-6b72-4676-aab7-f1e09d02b197 · outbound

This paper cites Deep reinforcement learning from human preferences.

Flow Matching Policy Gradients Deep reinforcement learning from human preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.582589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.582589Z digest=sha256:37b4249d0d7b4d6c10de1aecd7e9fd9748000e4d3f3057326204fb0b80d73350

Observation cfcbdff0-fef5-443f-91e4-606972a019ad · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Flow Matching Policy Gradients DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.585305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.585305Z digest=sha256:1e469143d282cc0edd9ba6e42b899b3bec2b9cd02270d3b6b0f1010ec77c7a06

Observation b779c2bb-1bff-4a9e-b9c9-72297ed03027 · outbound

This paper cites Magistral.

Flow Matching Policy Gradients Magistral

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.588485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.588485Z digest=sha256:8b70a146a3f68cb454b72b322a05a0f38b954506287e1348edc3ec22d5cfa801

Observation b06b7f54-55a3-4150-b4f2-20e4844256e6 · outbound

This paper cites Denoising diffusion probabilistic models.

Flow Matching Policy Gradients Denoising diffusion probabilistic models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.512063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.592235Z digest=sha256:da41d48ce2ed16e0d56652b61a47758ca16d94ed10fb9d3f2b65a2d19aaa93aa

Observation 01d1919f-4302-4b1d-a44d-c47cee24c712 · outbound

This paper cites Denoising Diffusion Implicit Models.

Flow Matching Policy Gradients Denoising Diffusion Implicit Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.595323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.595323Z digest=sha256:81eeb36be40e32bbe28883a25f04778b83f1e74925403e0f79f59b10f3463e9d

Observation 0a6b6828-98b8-461b-9a8a-6c3a8c510549 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Flow Matching Policy Gradients High-Resolution Image Synthesis with Latent Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.598137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.598137Z digest=sha256:8196940f09c0b6692c5a5a75d305e90bbdce8339dcbc830a6dbbfbdbc2f047d2

Observation 7fb39a58-50c2-4bed-a7a2-d05067b99914 · outbound

This paper cites Generative Modeling by Estimating Gradients of the Data Distribution.

Flow Matching Policy Gradients Generative Modeling by Estimating Gradients of the Data Distribution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.601546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.601546Z digest=sha256:590482177fcc5fb2af93979168331c36e66487347d03764591cca46e6aee6eb4

Observation c3cf78de-b476-405f-97bd-ea912929706c · outbound

This paper cites Video Diffusion Models.

Flow Matching Policy Gradients Video Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.604328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.604328Z digest=sha256:1c0076ce9b5372cf779af5a0a4a2389238a2b92caf80d0b83758431648579554

Observation 7e4e6e22-3ebc-4df0-9ec2-109cfb748b9a · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Flow Matching Policy Gradients Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.607246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.607246Z digest=sha256:357b002450142bff20ada0847530870a024e6f98ff95004746876ea393f438b7

Observation 39b0b94b-6eab-4796-b043-c23cc67f78b6 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Flow Matching Policy Gradients Imagen Video: High Definition Video Generation with Diffusion Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.610022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.610022Z digest=sha256:bbe662a4dfbc76d0e4779d72efefbb5dccc04c43f95f35e87d203038d604b9ed

Observation 90ec9298-fd08-447f-b4cc-4535480966b8 · outbound

This paper cites Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech.

Flow Matching Policy Gradients Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.612846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.612846Z digest=sha256:c1e198259ac92a0f18266c55fcbef6c63a90eabe4388c71930e820ce9cb2ba98

Observation 29b9bb6b-7f80-42a5-9a3c-dfe1a2e48497 · outbound

This paper cites WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis.

Flow Matching Policy Gradients WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T13:07:10.021511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.615625Z digest=sha256:545ad3cdb2575a53584ab70cf5a099afb0472c1fba424135e387734079d95b73

Observation 0ce2233f-4dad-46fa-b9aa-660cf115acc1 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Flow Matching Policy Gradients $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.618360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.618360Z digest=sha256:fa1148cd6df5219ef8a49da462b7447cd6d99e562af2dac55c87f7547fa84dcd

Observation 2c11213d-6d19-4863-9189-f602f75545ad · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Flow Matching Policy Gradients GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.621887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.621887Z digest=sha256:412ae4507513ae51343f427084a008717c59d5b3571d1714450d98ba0f7d819b

Observation f0cf4abc-8f27-431c-87b8-b8669653bb02 · outbound

This paper cites The Superposition of Diffusion Models Using the It\^o Density Estimator.

Flow Matching Policy Gradients The Superposition of Diffusion Models Using the It\^o Density Estimator

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.624787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.624787Z digest=sha256:ce1abe8e0e5a56b013806fdb6f49b61430759afc9568b4238af733540e6f519c

Observation c8264315-8621-46be-883f-c88f7f5f9ec6 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Flow Matching Policy Gradients Diffusion policy: Visuomotor policy learning via action diffusion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.627522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.627522Z digest=sha256:2e793cd0250325841a3c7f7a55f81c69fe42419740607767d99bf37b4e597af6

Observation 7b275053-1472-4143-81b3-889ac7bbb763 · outbound

This paper cites Tenenbaum, Tommi S.

Flow Matching Policy Gradients Tenenbaum, Tommi S

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.502803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.630454Z digest=sha256:7dcbd875f7b17b3a3ec9b443c23f4b0bbfaa738182be34dbc8a1758afe837217

Observation 609b7561-f021-4622-8c15-a979737e5a3b · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Flow Matching Policy Gradients Planning with Diffusion for Flexible Behavior Synthesis

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.633971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.633971Z digest=sha256:c337b823ca6844d0fbffdad935c3137b24dc93dc0885736cd44ca230dfa155cb

Observation 72de1d12-188f-43e4-89f4-c072db53ad94 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Flow Matching Policy Gradients Aligning Text-to-Image Models using Human Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.636844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.636844Z digest=sha256:4a1146d81db0a010624c675ab4c0302b61c28b8a43ab9171b14082421c59dfd4

Observation 608e9c56-f9b2-4354-8679-a2ae634a744a · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Flow Matching Policy Gradients Training Diffusion Models with Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.639595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.639595Z digest=sha256:fecf11669397da96f748f035a55155e921ce6c76b8b8bd502509330bab78a132

Observation 1f94d5fb-f315-478a-91bd-91b0b2591b2b · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Flow Matching Policy Gradients Flow-GRPO: Training Flow Matching Models via Online RL

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.642447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.642447Z digest=sha256:7f3384119736b180f490d5963d618857fa58ee73ff4e032efcc8a775d2a11cb2

Observation 307a410a-5ad7-4481-83ff-e91af1ade30c · outbound

This paper cites Learning a Diffusion Model Policy from Rewards via Q-Score Matching.

Flow Matching Policy Gradients Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.645809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.645809Z digest=sha256:0a553c275aa3bf4473287c3217b4e61b81c0f0ed2335562a23c75fbd2268870f

Observation fbdc4bdc-a63f-468e-84ca-2fba47c89ba3 · outbound

This paper cites FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control.

Flow Matching Policy Gradients FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.649092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.649092Z digest=sha256:29ff75b8a093fde14140f93633c29a25495b6a74b2b749748960d404faa9b0ff

Observation ca47697d-98b3-4c1b-bc21-a3ad36d553dc · outbound

This paper cites Addressing Function Approximation Error in Actor-Critic Methods.

Flow Matching Policy Gradients Addressing Function Approximation Error in Actor-Critic Methods

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.652056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.652056Z digest=sha256:1f1994b2d8524e66c50a06c7b11fbc0ab49633af80433a98ba299c5feb1c82a3

Observation 1e583592-7a90-4ece-af9d-ac27cb444553 · outbound

This paper cites Diffusion Policy Policy Optimization.

Flow Matching Policy Gradients Diffusion Policy Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.655123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.655123Z digest=sha256:f37f71e6c864739b66e4a833c6c266998d80571c2946f885794bc39c6f5de948

Observation c4a1389d-374d-4b3d-9fda-68b9f397afd1 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Flow Matching Policy Gradients High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.658297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.658297Z digest=sha256:aaaeab47846b2b44ceb0c8591868401544c11d221a620b8d15551d84fe711069

Observation d3c463ad-121d-4b9b-a546-b3314b3a03b5 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

Flow Matching Policy Gradients Elucidating the design space of diffusion-based generative models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.491496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.661292Z digest=sha256:6c25b41a776a74ef9c96826f023f064393c9d443e0440c5a809ee2a6b9895764

Observation 4125040b-071b-4473-86bb-994b4fd49c8d · outbound

This paper cites Kingma, Tim Salimans, Ben Poole, and Jonathan Ho.

Flow Matching Policy Gradients Kingma, Tim Salimans, Ben Poole, and Jonathan Ho

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.664803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.664803Z digest=sha256:4045f79a1fb73b50ef501dfe0453166d24d32cdb9d55deaacd213b43b0534249

Observation f87efa6e-56a4-4dcd-9361-095d50f425f2 · outbound

This paper cites Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation.

Flow Matching Policy Gradients Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.671327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.671327Z digest=sha256:7986b6c6182bbd5225734516302c30d4c073aa149a2a11ae00871194454daabc

Observation ea5d47b3-29c3-43b5-aa7a-34dd658b3c87 · outbound

This paper cites Openai gym, 2016.

Flow Matching Policy Gradients Openai gym, 2016

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.674640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.674640Z digest=sha256:00aacc8f9156d4ac6ce8d35e900f83bb71f85c39fd0509aebf900de061f33bb0

Observation 781d152e-9e66-45d5-ae39-c03650a189a4 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Flow Matching Policy Gradients Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.677389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.677389Z digest=sha256:d77c2dfc6854a576318759f3d598a8be17047cce965a9b15f4eb8d0725820f30

Observation 50e52d53-fe43-4d92-8b60-f8f8f4a83b6c · outbound

This paper cites Mujoco: A physics engine for model-based control.

Flow Matching Policy Gradients Mujoco: A physics engine for model-based control

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.469608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.680213Z digest=sha256:ac309e9c2bdb3cf7d4aa63d26c6ab3a8a40c1556dde80b9b5e4bbd138cdc56f7

Observation 4a7887dd-6be2-40f4-9483-172d05172d79 · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

Flow Matching Policy Gradients Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.682883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.682883Z digest=sha256:71110fb0de386ed08a2a1f0f8ec2c9bc8d7358642a62a32fb493431c9ea94f9d

Observation 24209886-9031-4947-9111-e6fc18d5ca08 · outbound

This paper cites Ppo-for-beginners: A simple, well-styled ppo implementation in pytorch.

Flow Matching Policy Gradients Ppo-for-beginners: A simple, well-styled ppo implementation in pytorch

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.458927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.685709Z digest=sha256:7111ddeee582c951ef0c4ee1543d5797ca2b805b102d842c7383db59f99b49f3

Observation f4537ca1-3755-46b8-9780-389e03224349 · outbound

This paper cites DeepMind Control Suite.

Flow Matching Policy Gradients DeepMind Control Suite

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.689388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.689388Z digest=sha256:eca412782ce3beaaafb3e0a3ce09fb98b710464b5729d024ea0f92fb736e6ef1

Observation ecec4bba-2bd5-4fe4-87e7-277aaef169b2 · outbound

This paper cites dm_control: Software and tasks for continuous control.

Flow Matching Policy Gradients dm_control: Software and tasks for continuous control

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.447525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.692647Z digest=sha256:5cc7c6ef906941a0767ecfb24491330bfdf55a7b07f670a92f5b2d28ebd422ee

Observation f16deb1d-05d3-4830-add6-e7f2cfc7833f · outbound

This paper cites Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation.

Flow Matching Policy Gradients Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.695342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.695342Z digest=sha256:cf217b20cb282028e069f3a14529ec63618edb1d9b673afc8282951a3fa2840e

Observation eb7e53cf-ead2-4328-83ba-89932559dc55 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Flow Matching Policy Gradients Adam: A Method for Stochastic Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.698292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.698292Z digest=sha256:f0a9f0663c0c15541d4d26f212c458c0147a254a176d2696cd9ff2098e73a609

Observation 2d1d3ee5-2ff8-4e33-b5b2-6831349903e6 · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Flow Matching Policy Gradients Perpetual humanoid control for real-time simulated avatars

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.701629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.701629Z digest=sha256:0b514c3c9a0e82b46ac82a2924f2e9e585b62d576ff206e8246b49f21af9c00d

Observation 71d4ce65-67a4-48df-a9cd-22f56efb8c39 · outbound

This paper cites Deepmimic: Example- guided deep reinforcement learning of physics-based character skills.

Flow Matching Policy Gradients Deepmimic: Example- guided deep reinforcement learning of physics-based character skills

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.431225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.704789Z digest=sha256:c4fee7999eb63dd4dbd34500fdc729da8484473456484dcf21b8f81a7914e838

Observation 5ec79a79-0638-48c3-9eef-cc92b07fefdc · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Flow Matching Policy Gradients Amass: Archive of motion capture as surface shapes

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.707404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.707404Z digest=sha256:b87e6911e97768eaf29a856372d4e08e8df60ca2aa9913709955cdbf63bb5bee

Observation 3f6620e2-326b-4dba-8f1b-6df9e88abedc · outbound

This paper cites Maskedmimic: Unified physics-based character control through masked motion inpainting.

Flow Matching Policy Gradients Maskedmimic: Unified physics-based character control through masked motion inpainting

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.414824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.709957Z digest=sha256:8b6e77caf5e44f0edce94e85ec14e26a1ae8cb478af68fec9bb65cc59c8e37b6

Observation da5426c2-df28-4a30-8564-0ee7fa8871f7 · outbound

This paper cites CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks.

Flow Matching Policy Gradients CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.712831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.712831Z digest=sha256:32efb832efc38f29b7e5eb4195cff81089fac3835ba16102fe7ca6a654199f69

Observation fb531a57-4844-48df-be70-a7f9efc57875 · outbound

This paper cites Universal Humanoid Motion Representations for Physics-Based Control.

Flow Matching Policy Gradients Universal Humanoid Motion Representations for Physics-Based Control

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.715858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.715858Z digest=sha256:c789046e790f03c803ba9f64c87598cb12900c6be15f2af91f60e78390484c65

Observation f15fcd56-9845-4b78-875e-14fbbc5ef57a · outbound

This paper cites Omnigrasp: Grasping diverse objects with simulated humanoids.

Flow Matching Policy Gradients Omnigrasp: Grasping diverse objects with simulated humanoids

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.404263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.718744Z digest=sha256:2c8b55e5b7ff6a98112dfd8198dff4aa1ce0beede472bc1e05eca876000634a9

Observation fb6389f7-a526-425c-9733-ba6a960702ba · outbound

This paper cites Ai models collapse when trained on recursively generated data.

Flow Matching Policy Gradients Ai models collapse when trained on recursively generated data

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.394437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.721355Z digest=sha256:e6b8d03eeab3a430f4c7c0dd790db831b9c3b895bf897c09145db3ed54ae0e14

Observation 5fcc6f94-7a74-4d85-9d2b-613b187b927d · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Flow Matching Policy Gradients The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.724594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.724594Z digest=sha256:15381c6d09300a7201828441923865a0b5f23654cfb3d20b9d432dbbef24283d

Observation 9609017e-463f-4320-a755-009de32ace2f · outbound

This paper cites Self-consuming generative models go mad.

Flow Matching Policy Gradients Self-consuming generative models go mad

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.382354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.727485Z digest=sha256:3642d2b8430806813219e3f74c1746720faae880153e9aefc60e9701ca1c5c3d

Observation 982b49df-d100-47d8-b962-11d9eee2036e · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Flow Matching Policy Gradients Progressive Distillation for Fast Sampling of Diffusion Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.730024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.730024Z digest=sha256:96d90410dfbda180f5131fab8851a64aec459bab50326d56c630e37b36ae6669

Observation 6d70b079-8e59-4e96-bb42-e693d40ef204 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Flow Matching Policy Gradients Classifier-Free Diffusion Guidance

Reference 84

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:07:09.734008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.734008Z digest=sha256:0319853d1eb710f032ae8440853a74b789189b97e375ccd5f305e5e926c331e0

Observation cbd8537c-7cdd-4b41-b843-fb363b4b1b1e · outbound

This paper cites Low CFG scales tend to encourage bluriness while high CFG scales encourage saturation and sharp geometric artifacts.

Flow Matching Policy Gradients Low CFG scales tend to encourage bluriness while high CFG scales encourage saturation and sharp geometric artifacts

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:07:10.371016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T13:07:09.738381Z digest=sha256:88560855616c9f6586aae0d88cdd27a4d1344a9a08eb5de870964eab435a2a23

Observation ff534587-8822-4c30-9401-e11a3b7bb46c · outbound

This paper cites Variational Diffusion Models.

Flow Matching Policy Gradients Variational Diffusion Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.668237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.668237Z digest=sha256:cd589acc3f638cd29acdfdbfc706f775b14500b2333839924e698d59b2f6d030

Pith citing papers

Observation 05e84f68-da6f-4416-a7fe-69d441ee2b57 · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Flow Matching Policy Gradients

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:09.497998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:09.497998Z digest=sha256:76bcae1b0f5ce59a64c4f8f829c91eb01d23069299f3b2b2cd8f0e79d0966eb1

Observation 42306e54-3ee8-4118-a18e-aabd58d51bbb · inbound

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models cites this paper.

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models Flow Matching Policy Gradients

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:24:57.490427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:24:57.490427Z digest=sha256:518fc71ca8fc4b16ebf09f8688e0160e46908224fb887f00a2af74625e60b2d3

Observation 813d2a66-f479-4ab2-9210-41d0a7f218b0 · inbound

Training Diffusion Policies via Prior-Mapping Co-Evolution cites this paper.

Training Diffusion Policies via Prior-Mapping Co-Evolution Flow Matching Policy Gradients

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T19:03:03.053275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:03:03.053275Z digest=sha256:cff4b25ce6394e815a007ad6ef900934341aae40c9f81880d3e874e96de8b7b6

Observation 02cc83af-386b-48db-be1b-38a7ae9bbf7a · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Flow Matching Policy Gradients

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T02:55:50.559508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:55:50.559508Z digest=sha256:0f3b973887ad8b51227ba231f37eeea92920f1ba1c50c7e9a7608b131ed99288

Observation 6110491d-3f52-421b-9f2f-7a52c6b18727 · inbound

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation cites this paper.

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation Flow Matching Policy Gradients

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T23:59:16.138832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:59:16.138832Z digest=sha256:cd8343f7d2ced20503c2acdca7edf1159f380c2f5dc12aa1b8162049959d81df

Observation e72ba1c7-d96b-452b-88d1-1d6dfddec020 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Flow Matching Policy Gradients

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:2c4aa785e683109674525dad167a1daf4cc8119006778c638b037d2f8f50cac9

Observation 4dbeea93-a4ba-4039-aa29-189b66e09596 · inbound

Genuine pair density wave order on the kagome lattice cites this paper.

Genuine pair density wave order on the kagome lattice Flow Matching Policy Gradients

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T13:17:06.175932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:17:06.175932Z digest=sha256:e218a81e51eda79f482293230600811f422c851e3ca98f9c4108f1053bbf54bb

Observation 81cc340c-653a-4b55-88ec-34322423e827 · inbound

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling cites this paper.

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling Flow Matching Policy Gradients

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:08.590825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:10:06.994557Z digest=sha256:25f24e47517c460fba1a49cbfde02dbf8d35ef204222c6e1f310c56cc54bfa49

Observation e789f086-ded3-4079-b850-39e5af030598 · inbound

Positive-Only Drifting Policy Optimization cites this paper.

Positive-Only Drifting Policy Optimization Flow Matching Policy Gradients

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:25:26.373715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:22:55.408035Z digest=sha256:a91d1d353c286ee0d9a7d3c9f8f3e35bc6a5541c9cf32d04c227c75d211e6e7b

Observation 4cd40595-a073-464a-baad-2aefc12f74a9 · inbound

V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think cites this paper.

V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think Flow Matching Policy Gradients

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.462900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T08:27:17.629238Z digest=sha256:048459ec0184b1fca4928e8af69c93caa16f1e7967f4e77efc35e7fee1dc23e5

Observation 99dd1daf-654c-4afd-9632-0e9e5074a3c9 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Flow Matching Policy Gradients

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.364053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:17f70f5212468d9418770742d0520b3247f729efb7bdd293fdf8dd869e286250

Observation d1f51504-4b4f-4ca0-9bee-40f1946ef096 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Flow Matching Policy Gradients

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.701041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:a081bbe7d8228e9ba6cba1b670335ba195fd278619eef04772508ffa60fa506d

Observation 7b1ba904-5f14-4f41-a195-a8a8dd17e79b · inbound

Generative Actor-Critic with Soft Bridge Policies cites this paper.

Generative Actor-Critic with Soft Bridge Policies Flow Matching Policy Gradients

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:28.323897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:01:08.896356Z digest=sha256:d426443a91025788efe17805fd0ea8c5519b4be4b49a0aae38d286e27bd6560e

Observation 9c0034ed-4c0b-4de1-abd3-a6314507a6c2 · inbound

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT cites this paper.

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT Flow Matching Policy Gradients

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:27.215226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:31:33.829939Z digest=sha256:b552527f3404ac0481cf27b18233f64bf526a7d85766b6b90de16ce94740f4f6

Observation 525ed746-329f-4c0b-8e25-a99392d0698d · inbound

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT cites this paper.

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT Flow Matching Policy Gradients

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:12.123481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T23:08:30.205397Z digest=sha256:9a68f116eedf0e3c7e81963b95b54b7623bd3953fb993b7803365cfa4a0092df

Observation 9925d1a7-0a26-4370-9f8d-a360209d5528 · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation Flow Matching Policy Gradients

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:25.470250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:08:43.222818Z digest=sha256:5a94b8ea4e62c5167d2b7c79a2dc8cce27245608b333745e14b1b4ec736df117

Observation 26ebde54-747a-4775-b632-bda0ab5f1f03 · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation Flow Matching Policy Gradients

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T14:26:02.424496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:26:02.424496Z digest=sha256:00ab2bf2f566df1c51dbf190b7754a8a30494bb4c839b6a49a29105ca1fae739

Observation 011018f0-67a2-453d-8fcf-545ddeb72532 · inbound

Discrete Flow Matching for Offline-to-Online Reinforcement Learning cites this paper.

Discrete Flow Matching for Offline-to-Online Reinforcement Learning Flow Matching Policy Gradients

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:52:22.645847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T05:48:40.468890Z digest=sha256:6528c2171a32fca82439a1e60fbc6099820338da97fdb410b87cb5dfbd7e122a

Observation 9f39d2f9-dc00-49a1-b84e-4f54eec339d3 · inbound

Driving Intents Amplify Planning-Oriented Reinforcement Learning cites this paper.

Driving Intents Amplify Planning-Oriented Reinforcement Learning Flow Matching Policy Gradients

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:52:58.816678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T20:50:14.602822Z digest=sha256:c0ba1672748b9c3d50e5efa94cb56b30b1c15d15554e98fca39c93ef5a966807

Observation 49869b54-f9c1-4f88-acf4-a3fed8f6ad0f · inbound

Driving Intents Amplify Planning-Oriented Reinforcement Learning cites this paper.

Driving Intents Amplify Planning-Oriented Reinforcement Learning Flow Matching Policy Gradients

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:45.635819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T04:58:07.943615Z digest=sha256:0f516c6a26e2068bb3203eacbd747468bd21d185b60b0382ed1e29aebb13b64b

Observation bb8ed032-0dba-404a-86d1-d4871142448b · inbound

Video Models Can Reason with Verifiable Rewards cites this paper.

Video Models Can Reason with Verifiable Rewards Flow Matching Policy Gradients

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:07:37.118524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T15:03:14.894952Z digest=sha256:46ff6b2179ea7a980616b251b898239583dc84f728a30698967a06683ab3223a

Observation 6a075063-c23c-48ff-97fe-63a7d6b107a0 · inbound

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL cites this paper.

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL Flow Matching Policy Gradients

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:13:24.894272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T15:11:43.235574Z digest=sha256:d55bf578cd4b62ef05a8e6684bbab4f167f24d27145cc9e9c8f31d67b692df46

Observation 23dc9362-165e-4fa4-94e4-ee0bd0bbce84 · inbound

DEFLECT: Delay-Robust Execution via Flow-matching Likelihood-Estimated Counterfactual Tuning for VLA Policies cites this paper.

DEFLECT: Delay-Robust Execution via Flow-matching Likelihood-Estimated Counterfactual Tuning for VLA Policies Flow Matching Policy Gradients

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:08:05.161093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T06:03:24.963347Z digest=sha256:3bc0dde5401eb6074a40d8d940e190e2255c5dd0679746d56deeb83907fd3978

Observation c6dda188-44b6-40bd-b964-c8790ddc3744 · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Flow Matching Policy Gradients

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:03:51.548209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:2c77faec1059bf1d66a2e61543ef00893c5f60bc9347e70a0f53eb6223df6021

Observation ba7332b4-aa86-4a3a-93fe-e8b92efd1677 · inbound

Explicit Critic Guidance for Aligning Diffusion Models cites this paper.

Explicit Critic Guidance for Aligning Diffusion Models Flow Matching Policy Gradients

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.779099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T18:18:46.456767Z digest=sha256:edd446e7fd4cff2fe767d845b883b299cb108e87ebc7e6a9f14c4faaaa858da5

Observation 6f5d792f-788b-4217-9a3c-f27b0434a31c · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Flow Matching Policy Gradients

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.380720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:f4254b126de4a6a837e702359f58ad10c5f7d492da6aa12325f0f0d4cd66dd95

Observation dfe8937c-42c2-4882-a823-a84fba206635 · inbound

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance cites this paper.

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Flow Matching Policy Gradients

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:12.574209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:21:56.382343Z digest=sha256:8d23590ac78d7519a6a90f696ec8b3f426da46168abf7c71321b5f28747561ad

Observation 0fefc763-7a47-4966-a122-eddd48ef2301 · inbound

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios cites this paper.

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios Flow Matching Policy Gradients

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:08.566580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T22:47:10.062975Z digest=sha256:1bc4f0d00530dcf030d9e2dd5bec97fc8bd2b4d2af4a2b7875250741f1255b99

Observation 54dff922-423a-4c63-9fd8-c2dc04662816 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Flow Matching Policy Gradients

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.910495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:1d6fa31d1f54577b873eae05eb79c4b2ee519aa2796bb85ec77e9d94a90407ca

Observation 644ab85c-d687-4f57-9b97-ce1248b4bbbe · inbound

DiPOD: Diffusion Policy Optimization without Drifting Apart cites this paper.

DiPOD: Diffusion Policy Optimization without Drifting Apart Flow Matching Policy Gradients

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.932973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:24bb4892ac27869e18e87cce5c773711aa02689fe62bac1c7c196ffec5559e62

Observation 3084d04a-393b-4e9c-802a-4c0bc3245ad7 · inbound

Transferring Contact, Not Just Motion: Compliant Grasping Across Dexterous Hands cites this paper.

Transferring Contact, Not Just Motion: Compliant Grasping Across Dexterous Hands Flow Matching Policy Gradients

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.312220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T04:43:54.248289Z digest=sha256:aad403975a466696c9e132c842ac498fbf7c78b5f26972b65e5fbfae0b67cb70

Observation 7ec13d94-630e-4ff1-9b03-e26a0bba8df9 · inbound

ReFPO: Reflow Regularization for Flow Matching Policy Gradients cites this paper.

ReFPO: Reflow Regularization for Flow Matching Policy Gradients Flow Matching Policy Gradients

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:37.773920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:38:54.754049Z digest=sha256:9b71b0bd2d056d48285fb68ff78dca7c681d45d7b373e6a67935ebdd108ce1f3

Observation ac6772b6-e3e0-4d8e-909f-058b20597375 · inbound

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience cites this paper.

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Flow Matching Policy Gradients

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:58.769831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T01:57:04.293058Z digest=sha256:faaa657482a462f213603707e0cbfe98e7b6a78e18af1e4105452c5e7678ff25

Observation 43635d58-6a0a-4d67-88f2-cebaaff81699 · inbound

NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation cites this paper.

NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation Flow Matching Policy Gradients

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:33:13.452711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:33:13.452711Z digest=sha256:039f9ee005283ab999bde57e2e2008f7868345cea8325b17f24698688b72b578

Observation fe9cea05-9176-457a-a327-ee2baa00f1af · inbound

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration cites this paper.

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration Flow Matching Policy Gradients

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T20:30:43.078545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:30:43.078545Z digest=sha256:a2d6c1e389b27da986c30660e12ad44fad84aeb5fe372360a8d23776eb337d37

Observation dda17a4b-e2c0-45c6-b9ca-5f0cc6d45e4f · inbound

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning cites this paper.

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning Flow Matching Policy Gradients

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:26:40.667700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:26:40.667700Z digest=sha256:42643f57b771df100f1a62c3b42a5d7dde01bbf9ee0afea9c8039a6bd3f31495

Observation 5869d352-9738-41ae-8acc-8e07dbfb20cf · inbound

GASP: GPU-Accelerated Safe Planner for Real-Time Collision-Aware Motion Generation with Latent Trajectory Sampling cites this paper.

GASP: GPU-Accelerated Safe Planner for Real-Time Collision-Aware Motion Generation with Latent Trajectory Sampling Flow Matching Policy Gradients

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:50:11.728877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:50:11.728877Z digest=sha256:92ac141bca7aceddd5f82108420d0f5b397e1345173d79dd6c464647d763ed3a