Pith. sign in

Paper Citation Record · LEDGER

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

As of 17 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2601.08136.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.08136 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:58:48.872832Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T23:47:45.866615Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T03:49:31.218132Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 833d0adb-47a0-4ea0-924f-c6d48cd7f6a8 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Deep unsupervised learning using nonequilibrium thermodynamics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:45.251938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:45.251938Z digest=sha256:a90dc46a1ce4b379d4ba2b9472efa6e190d6e113d7d934ee96d0e1e30e1cc974

Observation 19a4c57e-f037-422a-abe4-36526df3b4e2 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:45.350199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:45.350199Z digest=sha256:3d9c4ed84350609a80b07ccbfa6039a956dcf4557d1e78dbdef702268102e380

Observation 4e80d303-751b-41e2-b78b-62fafacfd02b · outbound

This paper cites an unresolved cited work.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:45.464332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:45.464332Z digest=sha256:4c0dd105e772ff61b0e8651b0a0225cecc3523a7e2d50c78d93071d80f72ac79

Observation e1fb49eb-d8af-498e-a4c3-2fb9c53c1899 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies High- resolution image synthesis with latent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:45.603358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:45.603358Z digest=sha256:2a0fd1541d30861f5f6256ce5ce49d57a47944ea39c2a88edb5039917e1337ca

Observation 045b5e92-9cf2-4784-9fbe-387c1a6bfcfd · outbound

This paper cites Scaling rectified flow transformers for high- resolution image synthesis.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Scaling rectified flow transformers for high- resolution image synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:45.695502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:45.695502Z digest=sha256:db885d2a2acef59962cd01c76e1fa7e1099ebfe84c8c7a22fa95bf7924951d63

Observation 93e151d7-96e4-4bbd-b3dd-d75a6ba8f06b · outbound

This paper cites Video diffusion models.Advances in neural information processing systems, 35:8633–8646, 2022.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Video diffusion models.Advances in neural information processing systems, 35:8633–8646, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:45.767045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:45.767045Z digest=sha256:ebe2fb643c71089819622426a1c0afe981d9f183e50d797b9a19b7b8d5d1b8a6

Observation b9143424-a43b-469f-b09d-a2d026402b97 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Pyramidal flow matching for efficient video generative modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:45.946573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:45.946573Z digest=sha256:af8f12a908ab635dde96f4f7a8fef426b10370519487e37a479c101273a23a80

Observation 639113db-7e08-4421-a2d7-a8fd16c804d4 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.095599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.095599Z digest=sha256:6e30b8a855246910a01e065379862b6c6a3120c1e291c3fad12f6ccf1110a4a3

Observation e94d59a0-348d-4f41-adc4-8f0ea1afcd56 · outbound

This paper cites Fast and robust visuomotor riemannian flow matching policy.IEEE Transactions on robotics, 2025.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Fast and robust visuomotor riemannian flow matching policy.IEEE Transactions on robotics, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.179713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.179713Z digest=sha256:93a60254f049bf07f7de7976bae3dcf9523646f0cf71c9767ed7d59dd255561e

Observation 68eb1f44-d696-4974-836a-e71e889d719e · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.231464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.231464Z digest=sha256:efd729b627e5a7b0391dd0e37cb813175f29018205870a141bbac544ca4b3fe2

Observation feb4117f-b30c-4a73-afba-2f79c66a7dee · outbound

This paper cites Flow q-learning.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Flow q-learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.389886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.389886Z digest=sha256:4fc43c5faf229ef71659cc61e56a7058b8d48c658f593d589cf8528b0933faf0

Observation d226ec83-7d80-4821-880a-5e848c1e63a7 · outbound

This paper cites Soft actor-critic: Off-policy max- imum entropy deep reinforcement learning with a stochastic actor.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Soft actor-critic: Off-policy max- imum entropy deep reinforcement learning with a stochastic actor

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.535345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.535345Z digest=sha256:51b28e2b8e7c5cdcf73b620c4a96fc6357e2277ec39d444e3de79189f98d541f

Observation 3b654983-344a-4ce0-93f1-370d431825f8 · outbound

This paper cites Efficient online reinforcement learning for diffusion policy.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Efficient online reinforcement learning for diffusion policy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.670434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.670434Z digest=sha256:144b774d0e6a976a0d8ade99020bb1e4a4f503388a112e7e5e274bedac04483f

Observation dc75b755-4b91-401d-a5e7-f4889c33db41 · outbound

This paper cites Maximum entropy reinforcement learning with diffusion policy.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Maximum entropy reinforcement learning with diffusion policy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.759163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.759163Z digest=sha256:ddc8d384daf00525a782fd515a226bd5107b266bc56e5e2140d647e1a57bd7f8

Observation 456f27fe-8874-4bef-b35f-d4162c0e4045 · outbound

This paper cites Iterated denois- ing energy matching for sampling from boltzmann densities.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Iterated denois- ing energy matching for sampling from boltzmann densities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:46.937229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:46.937229Z digest=sha256:da3f71a25400f31f1e0e3a9ea2216e599cbaa2119fe264c84342167808a1b101

Observation 626e9e36-6abf-4181-b04f-328101049240 · outbound

This paper cites Sampling from energy-based policies using diffusion.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Sampling from energy-based policies using diffusion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:47.093186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:47.093186Z digest=sha256:b094b1e3f882b7a99be3c1718643540bc01969bd6f35754f9684f0748561a40d

Observation 64495332-ef8c-4b65-a517-7c9423aebd2a · outbound

This paper cites Diffusion actor-critic with entropy regulator.Advances in Neural Information Processing Systems, 37:54183–54204, 2024.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Diffusion actor-critic with entropy regulator.Advances in Neural Information Processing Systems, 37:54183–54204, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:47.233661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:47.233661Z digest=sha256:4c2780d9d579513183d2c6790b974c23f0ef3e4dcacdb40238fd1bf4e1d664b5

Observation ef793eed-2799-4051-a587-48d3570bd33f · outbound

This paper cites Flow-based policy for online reinforcement learning.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Flow-based policy for online reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:47.312614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:47.312614Z digest=sha256:310714be35cb86e3e2c4e3f7a15299a31ef64b3f07b63fc7a24cb64e4a3aac42

Observation 91e5720d-a937-45ef-ac06-4cd058a0da86 · outbound

This paper cites Learning a diffusion model policy from rewards via q-score matching.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Learning a diffusion model policy from rewards via q-score matching

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:47.438409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:47.438409Z digest=sha256:f2e5f829f63016608a0a68ecf8e9229443c61fcb3d9ce692cf0e29d1ee1a8049

Observation 30f9cc1d-de0a-4856-91b4-f44db276e851 · outbound

This paper cites Langevin soft actor-critic: Effi- cient exploration through uncertainty-driven critic learning.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Langevin soft actor-critic: Effi- cient exploration through uncertainty-driven critic learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:47.553359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:47.553359Z digest=sha256:0ecf7e4b6c21402dcfab957faf741b3abaaa6635ea9c2bd6464d52e69391091d

Observation e275e09e-ad33-4679-9fde-a5f33f4f7c5e · outbound

This paper cites Diffusion-based reinforcement learning via q-weighted variational policy optimization.Advances in Neural Information Processing Systems, 37:53945–53968, 2024.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Diffusion-based reinforcement learning via q-weighted variational policy optimization.Advances in Neural Information Processing Systems, 37:53945–53968, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:47.725924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:47.725924Z digest=sha256:776fdb40036f8a2478584e6e571bc89b25d156185265dd125d03341ffba532a3

Observation 3cec7a6d-f65b-436b-8d96-691b3fb1cba3 · outbound

This paper cites Online reward- weighted fine-tuning of flow matching with wasserstein regularization.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Online reward- weighted fine-tuning of flow matching with wasserstein regularization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:47.874754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:47.874754Z digest=sha256:31d44292130a873a76a5ea2094f92ea5975d34a7069a8d61fbfa04ab8e3d3dad

Observation d534975d-0181-4ed3-9c0d-a36e8abd9cbb · outbound

This paper cites Control functionals for monte carlo integration.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Control functionals for monte carlo integration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:48.038292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:48.038292Z digest=sha256:f2ec0026d535ce33d237ec1dd3e4ffbe85fa1e29680a63bb766c8781c17dc863

Observation 08f3acb7-afed-46a5-8bd2-ed26d3c9b98e · outbound

This paper cites Measuringsamplequalitywithkernels.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Measuringsamplequalitywithkernels

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:48.137246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:48.137246Z digest=sha256:149054ea3d1173f33dc7618b634fca1983be415ccc9df0443408f4c7dcabba2c

Observation 90a25998-d82a-45f7-90cc-0976f53cd262 · outbound

This paper cites Importance sampling: a review.Wiley Interdisciplinary Reviews: Computational Statistics, 2(1):54–60, 2010.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Importance sampling: a review.Wiley Interdisciplinary Reviews: Computational Statistics, 2(1):54–60, 2010

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:48.308114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:48.308114Z digest=sha256:29d5798fd791c27991e83c1092cbfd9f213014c09a585d71787506695e7748be

Observation 4fc221bb-0438-461a-bc41-e29dbabe7eda · outbound

This paper cites Model-based diffusion for trajectory optimization.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Model-based diffusion for trajectory optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:48.440171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:48.440171Z digest=sha256:4748564d632d05658dadbaa043f32c6f2271b6fa33c015481d3e8afc1689b5f5

Observation 33272ca3-4f45-4821-bffb-ca03d67b300f · outbound

This paper cites Reverse diffusion monte carlo.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Reverse diffusion monte carlo

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:48.632743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:48.632743Z digest=sha256:36d96366f546f485f5eaf3c18fde3910f291c958b5016afcddf678c5ee898ae6

Observation 55fc4954-db14-4fb9-a5f1-10e968e96066 · outbound

This paper cites Stochastic localization via iterative posterior sampling.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies Stochastic localization via iterative posterior sampling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:48.771946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:48.771946Z digest=sha256:025e6e047e8f5808958955e89a20ae60229b788aab7841415e0fb87ffc7bf2cb

Observation 93896640-bd77-4dff-b79b-bed03ad9d493 · outbound

This paper cites DeepMind Control Suite.

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies DeepMind Control Suite

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T10:58:48.872832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:58:48.872832Z digest=sha256:d7a348338889ce3d07b51030fbda162c21f558710ecf9e679f0370994eb2fc0c

Pith citing papers

Observation bc753560-b09f-462f-b993-e381b25043a7 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:75111d9feb2912b1cffb75e30466221044874d07cc4d9c473377ae588b7d5078

Observation 241656fa-c536-466f-ac90-5aee52c75cf7 · inbound

Mean Flow Policy Optimization cites this paper.

Mean Flow Policy Optimization Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T20:04:29.261622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:04:29.261622Z digest=sha256:4eb0b1886431105e4b802247ee9645db737cacb0cbd54e810a22c78791af8121

Observation 5b1319b1-dc1b-40a1-94fd-5966b4c23068 · inbound

Start Right, Arrive Right: Asynchronous Execution via Initial Noise Selection cites this paper.

Start Right, Arrive Right: Asynchronous Execution via Initial Noise Selection Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:49:31.220478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T17:32:02.560340Z digest=sha256:8942ac272cfebdb574e80a9b4fab422d783ad28a21febd1da80ec2d344d76c50