Pith. sign in

Paper Citation Record · LEDGER

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 3 inbound Pith citation observations for arXiv:2502.06061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06061 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:57:42.654116Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:42:22.214222Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T19:03:51.552172Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved7
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation add414f3-fad3-488d-a656-799906ed1732 · outbound

This paper cites left of”, “on top of.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization left of”, “on top of

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.064886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.556146Z digest=sha256:e237aa3f7951b29b96fa7cf4d12e236bb04b71a0fea0dd3548d0873394352a9f

Observation 4047a014-4084-48fa-bfc8-b790c9635c23 · outbound

This paper cites To address this challenge, we introduce W2 regularization, which effectively prevents over-optimization and policy collapse (Lemma 1).

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization To address this challenge, we introduce W2 regularization, which effectively prevents over-optimization and policy collapse (Lemma 1)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.048103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.562201Z digest=sha256:2d4919690d9eaf275983fd8c4cfc8cb041568363763ca86940449a0737fdbbe7

Observation c6adba18-3860-4df4-88a8-87d193498cc8 · outbound

This paper cites left of”, “on top of.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization left of”, “on top of

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.081852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.550430Z digest=sha256:3d5dfac8c9455515d6d5f8d94363224e26c05212bd6c6395bf3a97911e55bc91

Observation 8848c0cf-3223-4696-b1c9-b9e4d71afab0 · outbound

This paper cites a cat in the sky.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization a cat in the sky

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.029134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.567605Z digest=sha256:c94bcc5c496a9434592a52c342aed357388518ad7741c24a1808b51721f55135

Observation f7cdfa02-8a77-4434-8a4c-3dc8e23792bd · outbound

This paper cites a train on top of a surfboard,.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization a train on top of a surfboard,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.006766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.573522Z digest=sha256:7a6e57f5d10495a67cf75e5c9d2e28b04b58490ae03a2477dd1589afc19ba672

Observation b9e49203-b038-4437-af6c-c5399e991a44 · outbound

This paper cites 3) Bottom row: Alpha Clip reward showcases consis- tent performance even with different text-image alignment rewards, validating the reward-agnostic nature/property of our approach.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization 3) Bottom row: Alpha Clip reward showcases consis- tent performance even with different text-image alignment rewards, validating the reward-agnostic nature/property of our approach

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.984986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.578823Z digest=sha256:eb967c105f256cf4c057029ae46a42e0cee898adf4bfcd4c2c7cb44b1e445194

Observation 7e8aa6a2-b43a-4b18-8778-2e6f50948a7e · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.968182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.584564Z digest=sha256:a7b4cccbe36bcd53037fdc8ec3e8d00220c389ff3aaa66bcd484a2258e3812bc

Observation 9f8a3694-bc36-4e72-90d6-3ead5fe9bcbd · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.950919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.589732Z digest=sha256:8b9359e03f73b5846dba11f9e832d1395a0329af2a91e7095acda7b8ad81ee8b

Observation d28eda71-a0da-4321-847d-cbecdb6d56fc · outbound

This paper cites on top",.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization on top",

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.933892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.595160Z digest=sha256:f970841f19b2114a075f7b885c5679060689eeb83ba4e90c3c583985ca4c4f04

Observation 4066ae4b-8442-4f8a-aa12-ffbb3ade7660 · outbound

This paper cites Given that w (x1) > 0 for all x1 ∈ Xand attains its maximum at x∗ 1, we define: ϵ (x1) = w (x1) w (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Given that w (x1) > 0 for all x1 ∈ Xand attains its maximum at x∗ 1, we define: ϵ (x1) = w (x1) w (x∗

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.879540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.612598Z digest=sha256:b6680edbac5a6d90da0d9a39dbe99f817deb93ca332b5bd1ab45d7ac3b1eef97

Observation a396bb3a-3861-4a00-84dc-28a8157f6d16 · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 15

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T16:57:42.861835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.617699Z digest=sha256:87392c3c5d4f94de426774c4bdb0713787aad12322e65cc3e96f1b8c9494d760

Observation 440d2089-c959-4adc-bfa9-168ca71fe817 · outbound

This paper cites Then, we can rewrite qN θ (x1) using ϵ(x1) as: qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) ZN (73) Then for x1 ̸= x∗ 1, we can have: ϵ(x1)N → 0 as N → ∞since ϵ(x1) < 1.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Then, we can rewrite qN θ (x1) using ϵ(x1) as: qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) ZN (73) Then for x1 ̸= x∗ 1, we can have: ϵ(x1)N → 0 as N → ∞since ϵ(x1) < 1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.844221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.623385Z digest=sha256:ecb8f6c53579a6d3f9f98278e55fa90350861dc9ff776973d27d0a71cac1986a

Observation 12896a55-1a87-4433-b125-2cd1f4707661 · outbound

This paper cites And we can have the normalization constant as follows: ZN = Z X w (x1)N q (x1) dx1 = [w (x∗ 1)]N Z X ϵ (x1)N q (x1) dx1.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization And we can have the normalization constant as follows: ZN = Z X w (x1)N q (x1) dx1 = [w (x∗ 1)]N Z X ϵ (x1)N q (x1) dx1

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.826271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.628335Z digest=sha256:15aa4e5d02396e8afcda7eed95ebacc6293fecc725b129e1f41780d7a411c636

Observation 630b8d02-6064-4967-aa5d-b271f77edd9c · outbound

This paper cites Then, we can have the limit behavior: For x1 ̸= x∗ 1 : qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) [w (x∗ 1)]N q (x∗ 1) = ϵ (x1)N q (x1) q (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Then, we can have the limit behavior: For x1 ̸= x∗ 1 : qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) [w (x∗ 1)]N q (x∗ 1) = ϵ (x1)N q (x1) q (x∗

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.807617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.633170Z digest=sha256:98da438a3a68cea157e9e6db12c148de2675480b6dc24655f8b66fdc238edef5

Observation bfea99bd-4bc6-456c-aebf-059103fa33c8 · outbound

This paper cites (75) For x1 = x∗ 1 : qN θ (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization (75) For x1 = x∗ 1 : qN θ (x∗

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.789504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.637958Z digest=sha256:d5345ba886bb81ee9cd163c78664f220700e9e74f9f4fbfc54cbd4d8b7367876

Observation 68dbfe68-f857-4469-8984-b3681be2010d · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.772687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.642634Z digest=sha256:7865268da42cbbb79e236106127033bfbda0afdd586e7f7d74a12e6db180caa5

Observation 96598b70-5c5e-4947-b81d-82bee4ab30e4 · outbound

This paper cites cat") = pclip (x,.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization cat") = pclip (x,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.756262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.647660Z digest=sha256:efe843a79b94236573d06a16ea57e0a90076bc4663fc63df67d70062e136fbbe

Observation e95af9ff-2d57-4371-b120-ad8d9cf1bf3e · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.738323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.654116Z digest=sha256:c2c26d241de5b1b3a30bb4362ccaad965020a07cdb2a5bfaab99ed651236aad6

Observation e5a1e71e-039e-4ec2-a381-5cf4f4b03a79 · outbound

This paper cites In this paper, we introduce two methods to handle the Overoptimization and ease the mode collapse risk in online RW-CFM algorithms.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization In this paper, we introduce two methods to handle the Overoptimization and ease the mode collapse risk in online RW-CFM algorithms

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.897400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.606891Z digest=sha256:066f82ef972cb05f7a8891703073bcf81d88d314ed56df9f9187cf4f88f98471

Observation b87d7f19-3b44-4ef1-834d-3d443c4d9783 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Overcoming catastrophic forgetting in neural networks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:42.537531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:42.537531Z digest=sha256:e21618cccf0cb552b3f72e45f0600cffbf6d16d47e25b006072ba3919f9cb1bb

Observation c132e1f4-b409-48e7-a311-447bfe713566 · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.916713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:57:42.600862Z digest=sha256:a6f21aa1da5077412e4b45fc27716497353de7d4c4d6290e8796ee2a48e2a01f

Observation 1435c790-34aa-44ff-b440-baa7ed828949 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Understanding the performance gap between online and offline alignment algorithms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:42.544295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:42.544295Z digest=sha256:e772f10c7fa4422312f7764538dc40938f54b9fb907f4554ea2937cf12e631da

Pith citing papers

Observation 5b73787a-0011-4fb3-8fcf-a88a0e1b9001 · inbound

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance cites this paper.

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:22.214222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:22.214222Z digest=sha256:af907e6ff1e4f24eb38bb5a09ee799ecf47fe6f2be1f66c750d012a8b7997da5

Observation b909be6c-2528-444c-96c2-60f7217c6854 · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:08.654870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:08.654870Z digest=sha256:5296fc5ab8a22797c362237cef25882fbf10e84e04f306972fc32ba9efa0366f

Observation 5a44649f-859a-42bf-97b9-d208d985208c · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.553716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:6107b0cbd5e30a0d3711c8512e5294f67c813035a8ea0ac52d1a31187818bc44