Pith. sign in

Paper Citation Record · LEDGER

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization

As of 6 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2606.05468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.05468 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T05:38:11.089753Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:31:34.620402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact15
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e4a2a7b-9acb-4ea5-9ba6-216eb7efa12a · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.347422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:ebd11c9084e40f6ae99052937ffd89722f01e4839e2658ddc64fe5794e73f031

Observation a84d4502-3a0b-483e-a35b-b828431ec7ec · outbound

This paper cites Zitkovich, T.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Zitkovich, T

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:3fd00fed89b8e2846e884f776c7d6ab0d6473530d8680af9b24324a5638261fc

Observation 9d4b9225-f0bd-4788-b614-29b90f28b27f · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization OpenVLA: An Open-Source Vision-Language-Action Model

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.369607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:aeb1b5496dc74d2c2e631b93f39678cd7eb7134e3ab5d31cbb297b0eab7df28f

Observation da506971-fa08-47eb-94c5-e908cac25bdb · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Octo: An Open-Source Generalist Robot Policy

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.355588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:8181e8b884980a0a3d5e8768a60a7f59e901fc6d174887d92e11e14a73e0881b

Observation 8015cbd1-3dec-4918-8d9e-bc5ffa58d7fd · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.346699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:a805ba00eb417e2c4fdf7a28089165f5af89378828f0ed2df69a2d0cca80aee3

Observation bcd60db1-2b4b-409d-bda8-91382679e21b · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:a915cb140280b1817124fde02e95f9c48355c072928a2b7d8e8d003f3f3aa3f6

Observation 84f5b215-e300-4979-817b-fb53cbe1ee15 · outbound

This paper cites Flow Matching for Generative Modeling.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Flow Matching for Generative Modeling

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.381284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:dd3b8a238b4f216e80d4d5494a67d0e2834d143f65e1bed4154fa4ff17efc912

Observation 88195603-b028-4393-a0a0-5c3de5f9c7ca · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.383928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:4537c18788475a2f7770961fff7b55bcfecfc51c883ebb5166cb2a136d78e4b2

Observation 9aedf7db-f547-48ec-a8f4-d04a5e68147d · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.392944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:da7dfe8ae9c67273e575eb2a838520ed6063eeb59d1a248bef0cdcb691b89f17

Observation bf6dabe3-9446-46be-9a50-3740297c0a36 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:053d35733f734e27206a209a12bcdd9e261eb8aae547145cfe7807a15f7c4c97

Observation 366f44bc-a214-4235-9c59-d04498b368cf · outbound

This paper cites Ouyang, J.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Ouyang, J

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:29932459df569418cae19e7c9daff14e5a62228c3606ab00f99e8469bcea542a

Observation 1d71d986-d49f-4fc2-bb8a-0621822811f8 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:0444db353b2b67e53e8bb4c145e41c08d97e56ad18e5eeb414912f3b6b56170d

Observation f5f1f064-0950-4c8a-807e-fa9479fcca7f · outbound

This paper cites Schulman, F.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Schulman, F

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:6e866452f2e5730e84bb571dca51cecb971f4411461552c85f883b1d3ccda1e9

Observation 66141af6-a326-4bb8-94bc-a52f2edd3006 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.369951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:74df80be930931f9ebceff81d9bf682d76abfe6706f0b672a76a0718b036ef47

Observation 62ec178a-1942-4964-aefe-d46acd212ed5 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:053a35e06df3ae787f3c877a734e05afe9ff57cf4635902d9fb2b3f84ff85dbc

Observation b1351cff-46a7-4dc9-a200-303dd7a210c5 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:b84f5d08f00db1f28d407c8fc8d7e87c058a561072e0b3cc0d5f645bf6e5a7de

Observation 1b030e9c-c707-4ca2-9d50-7611f9179f14 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.372377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:f96a237608859a9b9a1dc25d24a9c33f3f4c8f6076ad79d555562cd5a8a33513

Observation 50bdbd37-f54a-4d62-9397-1792e9d6d04f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.367329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:d3a8de0a1c6e57934e1088f688d87025d1a73cf189c1b27df4265adbe9b7bdc1

Observation 5b939387-b20c-4efc-b60f-0a503c1c955f · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.384131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:7a2f0104ba7e1b18286e6df076b37da2361687653e9ef3c2547bdb0ace76c4f4

Observation 02a91f91-65be-4ab1-a6f7-ca0aaa2d4ece · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:06:49.389175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:2597cc771fc8211eff7f0a9c165d3644685fa828042401ac2222b91cc36616d3

Observation c801d4a2-490e-4708-84d3-ae93e50b8d97 · outbound

This paper cites Rafailov, A.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Rafailov, A

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:d1e11cc8743a1c856e7e95489dc5a74dde9e50830f535c5ee85cce75e675d2e7

Observation bbb28a0e-da5e-4591-8c1a-15b065f0726b · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:0932f6049207815312e1961187f3411a96736f58556649b77cf367ef9c5608ca

Observation 6ef050e2-f988-407d-9b35-1216cdef32f2 · outbound

This paper cites Kelly, C.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Kelly, C

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:be486d2b30068b5b7f182f76699b990e05efef8815a16d0c3e341fa41b771a24

Observation 98a38f4f-e9cb-4295-9fe7-61d3107ab9cb · outbound

This paper cites Kostrikov, A.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Kostrikov, A

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:79d00988823f8999f6b6349387b59b702ac537226e928cfa138727d63ce389df

Observation 3a2c723f-5941-4219-b31b-ef33615b6131 · outbound

This paper cites Nakamoto, S.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Nakamoto, S

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:494df25b8da875768ed9187f2c9050b1edf970d05780a2d242c3eafb6353b6f0

Observation a51271d3-4c59-4733-b050-58bffa008673 · outbound

This paper cites Black, M.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Black, M

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:b68f46829e509f8ff49f552c068cd3877ff6f9417490245a8ffe69039b344ffe

Observation 6d466824-7d3e-46fd-a6c4-3d9b9c9e6d93 · outbound

This paper cites Zhang, Y.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Zhang, Y

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:aabae07b32d01280339ec83fb2207c306cdad58b99a5b0583e4d172a07ff8cb1

Observation 83df395c-3f9b-475d-a3e2-6049c4da235c · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.380510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:1289a46c222c44cdfddf48c308df7b14d40fd5ae4da86b98981d328c10ad6992

Observation 75c04942-f7a4-49e3-81b1-6bda7cfbccb3 · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:06:49.361855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:ef8ee07bd0a23b4222bcb0044879c4e4f4a99c15ef6a455d818c252720a3ae80

Observation b3e48ca0-1ce2-4742-bb8f-0fdabe1d4422 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:f047ed43570af16e07ba1bdb85d153663f748d65884590ed9c3feeab90a28d8c

Observation 34de7e37-25f0-4224-92fc-8cf88aed9166 · outbound

This paper cites Xiong, H.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Xiong, H

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:f01c5d89cb99eb597104908b66d9cb4daca6dba60ec411ae06b1450192e6d1b0

Observation 50dd785a-3d7a-4927-b24e-109c417bd9e0 · outbound

This paper cites Hejna, R.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Hejna, R

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:c8984452080f4d33bd66525b4c4c3dc12752eab28bb973325dfc62b823f96dba

Observation 52c89370-f519-4b01-8e73-5cf3b319c8f9 · outbound

This paper cites Wallace, M.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Wallace, M

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:ab99e2afdc056c98182ff6d55fbd4e1be8d4c34882434bdd5cb06a65385b8976

Observation 23c65083-dab4-4214-b78c-3511fa8d7f6c · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:a330f371aa3c2b5cf7ffb09f302b13813f36709676184d965b938630d5b0e89e

Observation 3706f97f-1fbf-4b2c-a064-d5da69ac0fc3 · outbound

This paper cites Improving Video Generation with Human Feedback.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Improving Video Generation with Human Feedback

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-02T09:06:49.358503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:ebb9645ca4439da2c4b3a8e5b5a7b9bdfa4140b81dca774b29558aa599897a0f

Observation c5c31fad-de0e-4ae0-8776-5f4865e16516 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:d3b98c8b65d2c0b77e59cee1fd1de6112451a8b945d6051e414ebae4ca257e0e

Observation d36f2619-0a99-40c7-8f4a-a36d8a92d57d · outbound

This paper cites Robins, S.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Robins, S

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:523e04fade5bc44268768fdd0412e37496f85b977a6910f5f78eb11b37f2d4b1

Observation 3eadd73a-5309-4ada-bdc8-70d20e3c1ada · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:f364ce7e961f1ed8b9880a76f53b8054626e8bca4a9dd8db516c07376e8d6b66

Observation 85320e61-e839-4b03-be3e-649330e1be89 · outbound

This paper cites Mantel and W.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Mantel and W

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:b41c069cb3f224e2d744402efff6cb4fcb0fbdea2caa5618d69faea4cb2b973c

Observation 7d353e49-49c0-4a2d-be97-c822563a6b65 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:af6c0e2ccab03ee0d87605d304b1258b4c4ae3dcf8bcd75c7dc3931b446dbeec

Observation 072df478-282c-42ec-864e-7c7561d2172e · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:a1fc34f9510ac6b1cf490f5c7bfcabeebb14cbe5a04a47070bfe5b8f4935ae0b

Observation 4695fdcd-8dc7-4f30-80fc-93934b4ed4af · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:4cbd80e10773412069180f80496f77acb3f8f03a6a5fc4f54e73b02045570b2a

Observation bfe98040-db98-4073-8a96-fab6e0ca3119 · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:9467bab53e9e1f30eb26b656471c9363de885e7e4f5885a354c0f418c834dfa3

Observation 6e4a9f5b-cf17-4a89-9562-193885a123da · outbound

This paper cites an unresolved cited work.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:f131ae2c61221aa97f412af5ecf5fa245043c9a66ee171f4f033528b2605f956

Observation 88880598-d42e-4036-88d5-ef4840487b44 · outbound

This paper cites correct-then-retrain.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization correct-then-retrain

Reference 45

Resolution
malformed identifier
no resolver link, observed 2026-06-28T05:38:11.089753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:bba110e4f68ab57c00ad8f4d42144f515fa69e57ae2441497fb4f6a8ce30d541

Pith citing papers

Observation 3329dd95-3ae2-4c66-811f-3187e50d2f02 · inbound

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack cites this paper.

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T11:31:34.620402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:31:34.620402Z digest=sha256:b97243339df6d0b9f9a8efe32cfc413033b6496d647a78f561d1891e6d591f8a