Pith. sign in

Paper Citation Record · LEDGER

Flow-Based Policy for Online Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2506.12811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12811 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:16:11.130271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:19:37.801846Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3f2938f-aa5c-4a56-b8e3-6a47e1da3c77 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Flow-Based Policy for Online Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.961170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:efdf3530749671715e7e3650a527e520ea560a6df9d3c03969a63c4f8640e5b9

Observation 7a5816e2-7622-49d0-9370-8686d0084056 · inbound

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data cites this paper.

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data Flow-Based Policy for Online Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.772811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T14:24:58.919333Z digest=sha256:2404ea21b05aacc945ea9a336969ba37462f2b8969e827ce9c7fb0490bb90140

Observation 1274de69-c92b-4002-9014-8ce57310826e · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Flow-Based Policy for Online Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.879087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:ed68e4940226a89d393199b661957aabc8b1e6f129079f9835f8f781478da1f8

Observation 53d83936-0ddb-46c2-85f7-48147fa6f40b · inbound

ReFPO: Reflow Regularization for Flow Matching Policy Gradients cites this paper.

ReFPO: Reflow Regularization for Flow Matching Policy Gradients Flow-Based Policy for Online Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:37.803584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:38:54.754049Z digest=sha256:9a0ff7d2c96ceb8444e31b81f1a49c7a143576b8e78381e44bca2eed08925fc8

Observation 450a06e7-e743-494f-bce1-df4265c25f7d · inbound

Dual-Flow Reinforcement Learning with State-Aware Exploration cites this paper.

Dual-Flow Reinforcement Learning with State-Aware Exploration Flow-Based Policy for Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.792831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T07:16:11.130271Z digest=sha256:efb000e15cc9bab3c0bf9c1f4dda01a6410990db5fc161f8d086715ce6456c4c