Pith. sign in

Paper Citation Record · LEDGER

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

As of 9 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 3 inbound Pith citation observations for arXiv:2502.06061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06061 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:57:42.654116Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:42:22.214222Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T19:03:51.552172Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved7
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation add414f3-fad3-488d-a656-799906ed1732 · outbound

This paper cites left of”, “on top of.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization left of”, “on top of

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.064886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.556146Z digest=sha256:ff3978d241c07196dd761bb17afc3906b11c5d0ee55da7ca8730fee8c7de8d15

Observation 4047a014-4084-48fa-bfc8-b790c9635c23 · outbound

This paper cites To address this challenge, we introduce W2 regularization, which effectively prevents over-optimization and policy collapse (Lemma 1).

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization To address this challenge, we introduce W2 regularization, which effectively prevents over-optimization and policy collapse (Lemma 1)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.048103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.562201Z digest=sha256:983ad1bc4d1f463d507933dccd9c4254c16a6ee51e2634feb0dd0deb122ae8a4

Observation c6adba18-3860-4df4-88a8-87d193498cc8 · outbound

This paper cites left of”, “on top of.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization left of”, “on top of

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.081852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.550430Z digest=sha256:a433d0421fec2fbe1c0793e4711974135c10665a342a6a73ee88f86d17c5353c

Observation 8848c0cf-3223-4696-b1c9-b9e4d71afab0 · outbound

This paper cites a cat in the sky.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization a cat in the sky

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.029134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.567605Z digest=sha256:e59fa3d5cc7f1c8f2658b3ee073953b81e627ab2b91d22c930e13932177bf8f0

Observation f7cdfa02-8a77-4434-8a4c-3dc8e23792bd · outbound

This paper cites a train on top of a surfboard,.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization a train on top of a surfboard,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.006766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.573522Z digest=sha256:d5abc348de861fd3bde8b713b6b049c7efce9954c66680acf52b3502be1af5bc

Observation b9e49203-b038-4437-af6c-c5399e991a44 · outbound

This paper cites 3) Bottom row: Alpha Clip reward showcases consis- tent performance even with different text-image alignment rewards, validating the reward-agnostic nature/property of our approach.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization 3) Bottom row: Alpha Clip reward showcases consis- tent performance even with different text-image alignment rewards, validating the reward-agnostic nature/property of our approach

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.984986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.578823Z digest=sha256:98be3389c072cb1ff49254f0c4fdb823e528108b21d9af2fa73c61a923d5f9a8

Observation 7e8aa6a2-b43a-4b18-8778-2e6f50948a7e · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.968182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.584564Z digest=sha256:c81cd7ed6940bf4fc1722e31682e23094d51011b659aa05ab6dab705f0d586ee

Observation 9f8a3694-bc36-4e72-90d6-3ead5fe9bcbd · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.950919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.589732Z digest=sha256:448f4b4f6e966509633039e91375b428f6bf29f1bf300a9ed55aaa2e748e8e10

Observation d28eda71-a0da-4321-847d-cbecdb6d56fc · outbound

This paper cites on top",.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization on top",

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.933892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.595160Z digest=sha256:9db00f4f79300e45f07122d7111c70d265884f2d1991fa2b3e86ba469417be13

Observation 4066ae4b-8442-4f8a-aa12-ffbb3ade7660 · outbound

This paper cites Given that w (x1) > 0 for all x1 ∈ Xand attains its maximum at x∗ 1, we define: ϵ (x1) = w (x1) w (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Given that w (x1) > 0 for all x1 ∈ Xand attains its maximum at x∗ 1, we define: ϵ (x1) = w (x1) w (x∗

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.879540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.612598Z digest=sha256:f2d88e4c0781b14701049e39fac9dc5f7f70aee1b891c2bc81edb711d42f2395

Observation a396bb3a-3861-4a00-84dc-28a8157f6d16 · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 15

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T16:57:42.861835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.617699Z digest=sha256:c3c2a8e1dc36a8c46169e692e8f25269e73e005fc7a69f619d981c9114215b1c

Observation 440d2089-c959-4adc-bfa9-168ca71fe817 · outbound

This paper cites Then, we can rewrite qN θ (x1) using ϵ(x1) as: qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) ZN (73) Then for x1 ̸= x∗ 1, we can have: ϵ(x1)N → 0 as N → ∞since ϵ(x1) < 1.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Then, we can rewrite qN θ (x1) using ϵ(x1) as: qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) ZN (73) Then for x1 ̸= x∗ 1, we can have: ϵ(x1)N → 0 as N → ∞since ϵ(x1) < 1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.844221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.623385Z digest=sha256:50649e168e26d6dbe0053e8777a9d6ac0eb22598ea7d4fcd7e3951cd74a39d45

Observation 12896a55-1a87-4433-b125-2cd1f4707661 · outbound

This paper cites And we can have the normalization constant as follows: ZN = Z X w (x1)N q (x1) dx1 = [w (x∗ 1)]N Z X ϵ (x1)N q (x1) dx1.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization And we can have the normalization constant as follows: ZN = Z X w (x1)N q (x1) dx1 = [w (x∗ 1)]N Z X ϵ (x1)N q (x1) dx1

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.826271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.628335Z digest=sha256:e8a7b04001d9487af8433f87f452b3c807cbfb2cd348c019ca4af3c21b5c4621

Observation 630b8d02-6064-4967-aa5d-b271f77edd9c · outbound

This paper cites Then, we can have the limit behavior: For x1 ̸= x∗ 1 : qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) [w (x∗ 1)]N q (x∗ 1) = ϵ (x1)N q (x1) q (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Then, we can have the limit behavior: For x1 ̸= x∗ 1 : qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) [w (x∗ 1)]N q (x∗ 1) = ϵ (x1)N q (x1) q (x∗

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.807617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.633170Z digest=sha256:4dbcf919fa6c98dbea2350f36d054e6cb241e7d0dd0e201069b180c1b67711d1

Observation bfea99bd-4bc6-456c-aebf-059103fa33c8 · outbound

This paper cites (75) For x1 = x∗ 1 : qN θ (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization (75) For x1 = x∗ 1 : qN θ (x∗

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.789504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.637958Z digest=sha256:93c16e3b0e5ecdb3d050cee93fb4c462cbeff98ff81b5b050ddaf9791cb414aa

Observation 68dbfe68-f857-4469-8984-b3681be2010d · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.772687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.642634Z digest=sha256:7f84e65e86c650b4fc52b83d7d32d2d3d6f78d1c34bc73a26a2432aff5ebfe60

Observation 96598b70-5c5e-4947-b81d-82bee4ab30e4 · outbound

This paper cites cat") = pclip (x,.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization cat") = pclip (x,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.756262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.647660Z digest=sha256:831c1e52aef3dc83d4db29f2d18f54609689450204badb3fc6527c70a182c33a

Observation e95af9ff-2d57-4371-b120-ad8d9cf1bf3e · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.738323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.654116Z digest=sha256:2b1790e041ac02fe41b59c2f90e5f2b87d74b43b34efa8a673edc4ce1b0f764d

Observation e5a1e71e-039e-4ec2-a381-5cf4f4b03a79 · outbound

This paper cites In this paper, we introduce two methods to handle the Overoptimization and ease the mode collapse risk in online RW-CFM algorithms.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization In this paper, we introduce two methods to handle the Overoptimization and ease the mode collapse risk in online RW-CFM algorithms

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.897400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.606891Z digest=sha256:cecc1100869c4d3937dc07e665299919e3def35f1b3c45662479e34029806629

Observation b87d7f19-3b44-4ef1-834d-3d443c4d9783 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Overcoming catastrophic forgetting in neural networks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:42.537531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:42.537531Z digest=sha256:074bb0d77689e297c444a25f67d9d7ad5f1236e06fefb47cb280e7860cb5db66

Observation c132e1f4-b409-48e7-a311-447bfe713566 · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.916713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:42.600862Z digest=sha256:fdd4cd3cd36ea6011a99cf99f9023e3932e0be617c4665467e4d60028a1b2dff

Observation 1435c790-34aa-44ff-b440-baa7ed828949 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Understanding the performance gap between online and offline alignment algorithms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:42.544295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:42.544295Z digest=sha256:ecce7f1923f7e22e5543a36680523a4d194d2ac0374d390d71b306e9684c27bf

Pith citing papers

Observation 5b73787a-0011-4fb3-8fcf-a88a0e1b9001 · inbound

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance cites this paper.

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:22.214222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:22.214222Z digest=sha256:cbb1026c9aaab7c99ffb7051b6129d0bbc3038884b4061ddebb8c6ad2689bb90

Observation b909be6c-2528-444c-96c2-60f7217c6854 · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:08.654870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:08.654870Z digest=sha256:ad33260d599c34576aa83069c6d252a940fcc64e31c431711d93c5f89cc49fbb

Observation 5a44649f-859a-42bf-97b9-d208d985208c · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.553716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:a40185bfd9d8d02bb8cdb85875b8c8af307aae2dfcda4f7d120c1d8c19007062