Pith. sign in

Paper Citation Record · LEDGER

Decision Flow Policy Optimization

As of 8 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2505.20350.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20350 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:20:42.023257Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact7
  • verified fuzzy17
  • unresolved71
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bb54068-3566-43d6-9ef6-5368890198c5 · outbound

This paper cites Diffusion policies for out-of-distribution generalization in offline reinforcement learning.

Decision Flow Policy Optimization Diffusion policies for out-of-distribution generalization in offline reinforcement learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.794468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.794468Z digest=sha256:7f537332e3543df5a0b9e2b8e4acb82b420cc740488baf409af21e215c88e7d6

Observation f119d52d-6872-42d2-8739-dad7a8a1e273 · outbound

This paper cites Is Conditional Generative Modeling all you need for Decision-Making?.

Decision Flow Policy Optimization Is Conditional Generative Modeling all you need for Decision-Making?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.841381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.841381Z digest=sha256:efe6bb58aaa37bba0ac446c5d65c1273802789aa1612c4118292122b3dd5c9e6

Observation de253d64-9fad-452f-b44c-a5243b4d6ea9 · outbound

This paper cites Let Offline RL Flow: Training Conservative Agents in the Latent Space of Normalizing Flows.

Decision Flow Policy Optimization Let Offline RL Flow: Training Conservative Agents in the Latent Space of Normalizing Flows

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.925222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.925222Z digest=sha256:9837282cb9cbbae2652edd71afd48c747df9141033bac3acf07e2f7d7201ba51

Observation 0f129e7f-1e58-4fca-a9d3-2fb211336e0d · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Decision Flow Policy Optimization Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.996839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.996839Z digest=sha256:b16b4411925e96688fc69052081c2c0c4d270bdf29c1d244e2680d54d83bc67a

Observation b4df2d69-3fe9-4ec9-81b5-adf5c3b19d50 · outbound

This paper cites Model-Based Offline Planning.

Decision Flow Policy Optimization Model-Based Offline Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.027635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.027635Z digest=sha256:b26c8991de875aa12b6f3d9ca5489b671a7be38fbddc129c68e96a34b66eeb4b

Observation 0efb84d6-7b3e-45e6-ad7b-bd796887085b · outbound

This paper cites The peano-baker series.

Decision Flow Policy Optimization The peano-baker series

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.087639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.087639Z digest=sha256:55f03cdb1b9d49b2e064cbeb499431c212369319f6ccd90bf4a73df18e27d0ec

Observation f669e893-c37a-4c18-aedb-6dc3c9e72cc4 · outbound

This paper cites Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning.

Decision Flow Policy Optimization Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.145203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.145203Z digest=sha256:c4510582d2422bd00db055ea17db438f33bec367bea53d12022a78977056e387

Observation 84d6bdb3-f2f8-4b43-b4c0-e1357b662ae1 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Decision Flow Policy Optimization Training Diffusion Models with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.208602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.208602Z digest=sha256:fc1fef05ffac991a1cab053511484c4e16d9ce709a2614c8d89fb6e24e7ca289

Observation 348dc640-46d7-4a64-9952-8c7840faa9e3 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Decision Flow Policy Optimization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.260626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.260626Z digest=sha256:3be8ad23a383b0e40dd65a5fa539a2db91032bb243ec7589f4c538e8c2e421e0

Observation 7bcaa2fe-154a-4da1-8370-bcc01f7a57c5 · outbound

This paper cites Reinforcement Learning for Generative AI: A Survey.

Decision Flow Policy Optimization Reinforcement Learning for Generative AI: A Survey

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:43.903572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:36.325840Z digest=sha256:91c4f2a96a850a5821a2dc18182a2ef438e788336689d3c38aa99add9b55a8ef

Observation 820a9767-3a44-4d6b-9503-ab3a35904761 · outbound

This paper cites Simple Hierarchical Planning with Diffusion.

Decision Flow Policy Optimization Simple Hierarchical Planning with Diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.384955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.384955Z digest=sha256:fba5b0f6c69bd1f9f693f5a010b55f456930b8a67aa717a2658cf846af0c8f67

Observation ad21c487-7e76-4c69-aab3-0e6db50e2a36 · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Decision Flow Policy Optimization Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.514447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.514447Z digest=sha256:7ecd77a54b8277c904e25634e58f4313ac68b248b3901d373e60457a71486b44

Observation 7bbc88ad-efbd-4cf0-9dda-be2de6e039af · outbound

This paper cites Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions.

Decision Flow Policy Optimization Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.669857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.669857Z digest=sha256:828c6cf98c8ba5b9664f645120f348b83b6497e1cdef4bd5efbaa1b037443289

Observation 3c723759-4267-4c73-ac31-05085568bd52 · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Decision Flow Policy Optimization Decision transformer: Reinforcement learning via sequence modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.788091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.788091Z digest=sha256:67782012400a80242d3d0d5ff7f845cc6e0dca11dcd08deb3a1419deaa175149

Observation 2278a0db-9356-4ad5-bd8a-7920cd691edf · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Decision Flow Policy Optimization Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.878337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.878337Z digest=sha256:5a6fb308d3ef56f1fdcf4c739ff0aeff554093d9885e26ac349ff8b6748b6058

Observation db435bce-f951-41d5-9933-609c4375bfae · outbound

This paper cites Flow Matching in Latent Space.

Decision Flow Policy Optimization Flow Matching in Latent Space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.912260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.912260Z digest=sha256:87a6cb0f50c5542d3cb52b21af2032171eaeddaefdd562ff6edac88f78a124b9

Observation db110ff5-eb32-42fb-b950-82196006e03a · outbound

This paper cites Fisher flow matching for generative modeling over discrete data.

Decision Flow Policy Optimization Fisher flow matching for generative modeling over discrete data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.964361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.964361Z digest=sha256:d7dfa508616559ee049206088d4e7e13ac8b50447dd47a9cfd69a49bc776f02e

Observation de1150a3-cfae-485b-a4eb-b33f5c0b678b · outbound

This paper cites Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization.

Decision Flow Policy Optimization Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.020291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.020291Z digest=sha256:0832e8b36248cba5669b4763252c9339da1c815ee6ddca30c9a370c5b47d6cb3

Observation 738948f2-c8df-4a37-bca2-b4a02c58691d · outbound

This paper cites DiffuserLite: Towards Real-time Diffusion Planning.

Decision Flow Policy Optimization DiffuserLite: Towards Real-time Diffusion Planning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.065822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.065822Z digest=sha256:957db9ea2f651da107fe5c9215bcbc8c7a75963d1d1ca989d820f2f337001e74

Observation 96595d25-2083-4c04-a978-81d105de798d · outbound

This paper cites Probabilistic number theory I: Mean-value theorems, volume 239.

Decision Flow Policy Optimization Probabilistic number theory I: Mean-value theorems, volume 239

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.130184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.130184Z digest=sha256:ab9b8339c4ae3f21d83df7ec4bc5ee2bc9ddeaf18301a65414f52f484bf2ee11

Observation 25e213c1-9081-4b08-b16a-346397d1a105 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Decision Flow Policy Optimization Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.190416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.190416Z digest=sha256:818323c30507397942ae4be0d3bd4152ba53b3fff491d1481e3a4d917a0c0d99

Observation e0ada8b7-b884-46c2-a7b3-cd3f3d362923 · outbound

This paper cites A reinforcement learning diffusion decision model for value-based decisions.

Decision Flow Policy Optimization A reinforcement learning diffusion decision model for value-based decisions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.230498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.230498Z digest=sha256:a4d718159689271322e19ee47634df89f3df5fd9909435b5b8d94b0b47bf3caa

Observation 4e7cf567-5aba-4e08-8562-5dae2e2047e3 · outbound

This paper cites Reinforcement learning for generative ai: State of the art, opportunities and open research challenges.

Decision Flow Policy Optimization Reinforcement learning for generative ai: State of the art, opportunities and open research challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.281105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.281105Z digest=sha256:fd03820ce88c806a0e6c8f7155fa4e80d18570d5bb1f6410cc7283eb30a38346

Observation 9af8a4e9-a85c-4332-b706-7a7bec547408 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Decision Flow Policy Optimization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.372971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.372971Z digest=sha256:079c2ae0d27273c61d533a715e9f1200456c1aaf6f8b7d4ae2a7449d9920d735

Observation 9a03472b-99bd-44ef-9143-fee30503e1db · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Decision Flow Policy Optimization A minimalist approach to offline reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.438497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.438497Z digest=sha256:3036670383ef6de3f9d6c486ff01da2421d29923f18a0e840e3a70753ee14f3a

Observation e125cbdf-408d-4e94-88ba-db8bc3189dc0 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Decision Flow Policy Optimization Off-policy deep reinforcement learning without exploration

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.510693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.510693Z digest=sha256:6c738a32da4c1a1730349d6d65ac35b2b9a63d9be6ffc740d90a8b8a7df2f12a

Observation d953dd6a-2031-45a3-9913-27c10a576f89 · outbound

This paper cites Generalized Decision Transformer for Offline Hindsight Information Matching.

Decision Flow Policy Optimization Generalized Decision Transformer for Offline Hindsight Information Matching

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.550897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.550897Z digest=sha256:c1fd2b963a6d335e6d478d5eea9c8efa5379721885489a1518039860c10452f5

Observation 2dff5418-ebe0-492d-be11-b3819b4cd23c · outbound

This paper cites Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning.

Decision Flow Policy Optimization Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.619351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.619351Z digest=sha256:aa640b712477c670d27856d161f16e516c6a55dd8c42225ba3ff458777f54e78

Observation 2206e9e9-7177-4d91-8879-450d1184d261 · outbound

This paper cites Discrete flow matching.

Decision Flow Policy Optimization Discrete flow matching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.684967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.684967Z digest=sha256:dfbac07e6c6029d73e9756b07ffa4059c06f7ab8ecc64a9c4e73536bcb36528e

Observation d631f224-5059-4175-9157-3e54022fb3d9 · outbound

This paper cites Offline rl policies should be trained to be adaptive.

Decision Flow Policy Optimization Offline rl policies should be trained to be adaptive

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.741192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.741192Z digest=sha256:6a8f319bda03eada6d25cf2f30c35ab214ea709a617e9b2d755faa3865ab72cb

Observation a5047400-14cb-44ea-968e-41a56053beaf · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Decision Flow Policy Optimization Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.793878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.793878Z digest=sha256:a294f7bfea49b05a0b5ebb0fdfc0ffa82b3d7830a789f397ae17eabf8af1d461

Observation 51eff6bf-9a58-4084-b442-0edb90f51bf6 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Decision Flow Policy Optimization IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.857764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.857764Z digest=sha256:3f6ba3c6540b3457930db777444fa1ba159052ca538eb72e7f9bc363fb0ffd89

Observation 56f6ab40-9ae6-4d94-a4f1-86474d6ada37 · outbound

This paper cites Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning.

Decision Flow Policy Optimization Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.931785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.931785Z digest=sha256:16536f0e7c845c52847cbd3fdb21c31a351e35e3d4c2a4e30232171432ceb005

Observation 45995f37-337d-4382-bcb8-5e9c76826337 · outbound

This paper cites Lectures on Lipschitz analysis.

Decision Flow Policy Optimization Lectures on Lipschitz analysis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.998904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.998904Z digest=sha256:5b88277f4176ecb5e40d1048b0895ac01433612eea5374b2470ab113d01041bc

Observation a942017a-8f99-480c-bb0a-77b6a401fe28 · outbound

This paper cites Flow++: Improving flow-based generative models with variational dequantization and architecture design.

Decision Flow Policy Optimization Flow++: Improving flow-based generative models with variational dequantization and architecture design

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:46.278627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:38.069389Z digest=sha256:0eebb0da32b9de3f56730447e3f8611b158f45c3fd6e7bdab7be00a651381677

Observation 20f550fd-9574-465a-818a-2a7c68710f0b · outbound

This paper cites Instructed Diffuser with Temporal Condition Guidance for Offline Reinforcement Learning.

Decision Flow Policy Optimization Instructed Diffuser with Temporal Condition Guidance for Offline Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.121882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.121882Z digest=sha256:bf5825d793b8014a24359d432c81b62d9363b8996ef8e85f8e342a6d1c04a348

Observation a05039dd-c937-4793-acdb-53693627ce24 · outbound

This paper cites On Transforming Reinforcement Learning by Transformer: The Development Trajectory.

Decision Flow Policy Optimization On Transforming Reinforcement Learning by Transformer: The Development Trajectory

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:43.445599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:38.165967Z digest=sha256:d73d25a3dbb0983918963363874d7fd12d9c87b7066580e9271befd1af328251

Observation b280403c-910d-4944-adf1-5da6f0f07d93 · outbound

This paper cites Graph Decision Transformer.

Decision Flow Policy Optimization Graph Decision Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.225712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.225712Z digest=sha256:ba1ca3ba305f1621aad8227dec106700e1e8973b4c406f15291f6e4e303262a0

Observation f3e6866e-bb55-41c4-947e-cd9b8066ff5e · outbound

This paper cites Adaflow: Imitation learning with variance- adaptive flow-based policies.

Decision Flow Policy Optimization Adaflow: Imitation learning with variance- adaptive flow-based policies

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:46.112420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:38.289200Z digest=sha256:cc247ef71f970fc60fed84fb4531f02308d9150708e04eefe6f88438bff7f31f

Observation ede41e49-49ce-4c59-8bfd-0440c2f19fb2 · outbound

This paper cites Diffusion models as optimizers for efficient planning in offline rl.

Decision Flow Policy Optimization Diffusion models as optimizers for efficient planning in offline rl

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:46.003051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:38.338249Z digest=sha256:6023f195bee30794742ac5742567517d63ccc83f183cf7769b0b7a594852c09e

Observation f402d366-30cd-4120-8488-c0978b239947 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.

Decision Flow Policy Optimization Offline reinforcement learning as one big sequence modeling problem

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.883231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:38.382943Z digest=sha256:9ff307fc33b5ef84287e93cfe757c2071d6ae6dd4301952e07f4098a49f64792

Observation cecf9b24-9199-4a7a-a826-8c0511ff1539 · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Decision Flow Policy Optimization Planning with Diffusion for Flexible Behavior Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.445055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.445055Z digest=sha256:0cadf411f565b33c5e7b1c6cde48618957d5a8aa552af93f915d7835923088c8

Observation 4861505a-1e2d-487c-a132-c445086446ce · outbound

This paper cites Efficient Planning in a Compact Latent Action Space.

Decision Flow Policy Optimization Efficient Planning in a Compact Latent Action Space

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.497294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.497294Z digest=sha256:33c9b397fbf8517851a3f706bc6f1963bcfa9677422cba691222a4040b18272d

Observation 2d8098ab-c4e2-459f-9614-4cfac6062659 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

Decision Flow Policy Optimization Pyramidal flow matching for efficient video generative modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.547205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.547205Z digest=sha256:67ad4df21165f01ff3b2a160364901afe6b7d62f882eec1d31a0febf8e777ef4

Observation eb7e562b-9b99-433e-b89a-5d65b1f14354 · outbound

This paper cites Efficient diffusion policies for offline reinforcement learning.

Decision Flow Policy Optimization Efficient diffusion policies for offline reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.709405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:38.624432Z digest=sha256:b79f2ac6284761a5393d56207f7bb1cb8cda413934b532c20352379d2be4b2d5

Observation 147e7268-7bad-41a7-84cd-ca58fb1e5aa2 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Decision Flow Policy Optimization Morel: Model-based offline reinforcement learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.689398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.689398Z digest=sha256:591eb07ee9fe4fcd36c634eab6ea9c515681c5b3d89200462bfb899959972730

Observation aade2317-0a92-436b-a2d0-33ca8966987f · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Decision Flow Policy Optimization Offline Reinforcement Learning with Implicit Q-Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.742099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.742099Z digest=sha256:126bdbf8643eda6f4449962e109e0719f21154ac2627f7b10a8175f1caa8fd49

Observation bb3f2554-6bce-47e0-a4e2-e4fbd730a15e · outbound

This paper cites Stabilizing off- policy q-learning via bootstrapping error reduction.

Decision Flow Policy Optimization Stabilizing off- policy q-learning via bootstrapping error reduction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.799102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.799102Z digest=sha256:0fe532c85b0767677b637271d62055ddc69725a54b50782508c5ae4b8fa2afde

Observation 913dcc55-4788-431b-bf37-de8794841bb8 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Decision Flow Policy Optimization Conservative q-learning for offline reinforcement learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.842872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.842872Z digest=sha256:a218e1f0cb7cc405271952a9ec99f3bb7743667c544a83c1e29bfc4cbf92cf84

Observation d507ba1f-1aec-4ef2-b44b-8ca0ebbcad97 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Decision Flow Policy Optimization Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.891334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.891334Z digest=sha256:81efd98c1f08f9ca5adcee83a74c55cfb4c322d6a4440a72c60b3c13f434f0e1

Observation 4c22e361-d8f6-477d-9dbc-b58ce2e550ac · outbound

This paper cites DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching.

Decision Flow Policy Optimization DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.939358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.939358Z digest=sha256:9bd162d787728f39785a37d079417a1882ef4a766b35a1694d5c147ee50c4b07

Observation 1a19bf54-4dc7-4172-b70c-84e0a8517511 · outbound

This paper cites Learning multimodal behaviors from scratch with diffusion policy gradient.

Decision Flow Policy Optimization Learning multimodal behaviors from scratch with diffusion policy gradient

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.570026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:38.995078Z digest=sha256:23cb6f5addcc8c9bc848bc4bd3e44cf64cdfc15a4682ab6ed9e9230db3630e30

Observation 65e40278-984b-4041-8211-9b2dc7e9e0a3 · outbound

This paper cites Efficient Planning with Latent Diffusion.

Decision Flow Policy Optimization Efficient Planning with Latent Diffusion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.045014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.045014Z digest=sha256:9c61cbf25e4d11b2f182e139db2cec8eba4cbe9705df1f74f6fd046279ea8f8c

Observation 38d867f4-4693-4e1b-81ca-6b0a6d249a75 · outbound

This paper cites Hierarchical diffusion for offline decision making.

Decision Flow Policy Optimization Hierarchical diffusion for offline decision making

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.411100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:39.111910Z digest=sha256:f89385edcdde91c454c739a135de1e0664a208714658ca2c5ce86923ac92ceac

Observation cfe9aa23-86c1-407b-90fc-015a29e29dcd · outbound

This paper cites Generative models in decision making: A survey.

Decision Flow Policy Optimization Generative models in decision making: A survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.162569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.162569Z digest=sha256:5362985a9d8e07990575a9330300de571e182aa986941b64099756e6d7ccb174

Observation 32e09caf-2a45-4896-9ecb-97a749b894da · outbound

This paper cites Flowvid: Taming imperfect optical flows for consistent video-to-video synthesis.

Decision Flow Policy Optimization Flowvid: Taming imperfect optical flows for consistent video-to-video synthesis

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.229974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:39.252580Z digest=sha256:a99c40564fbc0b59e6e3f9219b3afa5f3b1226c70f15904604e781738922f808

Observation f3f5e4cf-2988-4a3b-8c56-6046e29e676e · outbound

This paper cites AdaptDiffuser: Diffusion Models as Adaptive Self-evolving Planners.

Decision Flow Policy Optimization AdaptDiffuser: Diffusion Models as Adaptive Self-evolving Planners

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.290915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.290915Z digest=sha256:3e001093cf622d7e911985e8cb054f2f462a9782cfc5d34b6154fdd61fb8461a

Observation 59e99e51-59f7-4256-8a97-fd6895b8c622 · outbound

This paper cites Dataset distillation for offline reinforcement learning.

Decision Flow Policy Optimization Dataset distillation for offline reinforcement learning

Reference 58

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:20:43.045019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:39.370378Z digest=sha256:535a2df98e12410fa954ec00cf4132c27d1b6acba7457e4024737159e04d0727

Observation ab098d77-bed5-4008-b1fa-e314f1024513 · outbound

This paper cites Flow Matching for Generative Modeling.

Decision Flow Policy Optimization Flow Matching for Generative Modeling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.454108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.454108Z digest=sha256:8c1d12dbea985b186f34e65aa0776f0025745c29c02779ea7ac013ade38ac28e

Observation b8d5be13-e190-4420-89b9-8fb618ada7a5 · outbound

This paper cites Flow Matching Guide and Code.

Decision Flow Policy Optimization Flow Matching Guide and Code

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.518713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.518713Z digest=sha256:f07abaccb3e67897aae20c41930959540779b101a6965eeda71787f1d5b53f44

Observation 9233db61-bb90-43e3-804e-88b4e89d4d0b · outbound

This paper cites Generative Pre-training for Speech with Flow Matching.

Decision Flow Policy Optimization Generative Pre-training for Speech with Flow Matching

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.553836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.553836Z digest=sha256:3a2551a5d1e8dd81ebe7a137ece53a4d74bdb61eae90ea22f9305139a83bcb19

Observation 4f762be6-8ef3-4a64-8971-1ee7b7732092 · outbound

This paper cites SelfBC: Self Behavior Cloning for Offline Reinforcement Learning.

Decision Flow Policy Optimization SelfBC: Self Behavior Cloning for Offline Reinforcement Learning

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.862741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:39.589213Z digest=sha256:0957a3db9df2428146b1238f317ee2733c6014b351d16fbd811fb2db4ab69e9e

Observation bfb96da5-a6e1-454f-9a76-5ed920a99fc4 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Decision Flow Policy Optimization Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.668066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.668066Z digest=sha256:0e867087bcd503281c73793904b1af66b44befbab0536bf030e9efd8432d3268

Observation 5d8b44d6-4635-4d16-b8a7-05a935b8b799 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Decision Flow Policy Optimization Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.718510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.718510Z digest=sha256:625537ae5edfd48c027a9f3dca9f42702a2f0cee8a05d8f8cbbcb0163bb01ec9

Observation 8ad00c81-8a7f-4c6f-aa43-aa704d9760cd · outbound

This paper cites Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning.

Decision Flow Policy Optimization Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.813467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.813467Z digest=sha256:3184c957015336804bc98150868839ed15b63209cc310d2d54118fc675bb4a60

Observation 38f3ad79-c04a-4fb3-9f11-329072bf7a95 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Decision Flow Policy Optimization AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.855798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.855798Z digest=sha256:0e0ab1d51c3e6d527dfbbd1b561977013e76f20cc02ca542d9259ac0427dabec

Observation 85234d8d-2dcc-48bf-9bb7-ecccdb2ee88d · outbound

This paper cites Normalizing flows for probabilistic modeling and inference.

Decision Flow Policy Optimization Normalizing flows for probabilistic modeling and inference

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.960749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.960749Z digest=sha256:5a28859f2d5d7525d0c96c360eecf5d710bd2a36284faf92c11940e61fb41b5d

Observation c9adb103-e45c-4af9-8025-061ce3b280dd · outbound

This paper cites Flow Q-Learning.

Decision Flow Policy Optimization Flow Q-Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.029877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.029877Z digest=sha256:bbff6eb4fd9d5456eccdcaf6b8e58f1f246729bf5d189cef5851931e26b3e25d

Observation 7573d0f4-41ed-4f64-90d9-6fbf94a9db52 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Decision Flow Policy Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.119405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.119405Z digest=sha256:f7a32c6c494c5d4a869b058e0e465681a8dcbb088df2d0de6ba20dbe45e2c445

Observation b6a752ef-5d03-40be-9253-bb2ee4a3ab12 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Decision Flow Policy Optimization Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.163607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.163607Z digest=sha256:7edf4d812a36b319af5982150e72df3cc0d4c8f61cf3a36aa72885349e51069e

Observation 294b33b5-01f2-46d9-8400-56e849367032 · outbound

This paper cites FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching.

Decision Flow Policy Optimization FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.249756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.249756Z digest=sha256:09fccef82c8e3021e39d1f01622278bf4939723c8fd0ba5827073b03039b26d1

Observation e4601625-e387-4bc5-850e-c7d99470c718 · outbound

This paper cites Offline reinforcement learning as anti-exploration.

Decision Flow Policy Optimization Offline reinforcement learning as anti-exploration

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.103384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:40.346131Z digest=sha256:a812ee1afb18cba99ec3060215771bcd33858aca727c9873fa659f269908646d

Observation 1359d8ca-070a-490e-b627-436869ba5b16 · outbound

This paper cites Flow matching imitation learning for multi-support manipulation.

Decision Flow Policy Optimization Flow matching imitation learning for multi-support manipulation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.033689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:40.424534Z digest=sha256:e509ee753029031524156df9d557da54fbb7a6e81a7e56c99be54c6dad2dab5d

Observation 3dd87819-0a97-400f-b2ca-18147394877c · outbound

This paper cites Universal Value Density Estimation for Imitation Learning and Goal-Conditioned Reinforcement Learning.

Decision Flow Policy Optimization Universal Value Density Estimation for Imitation Learning and Goal-Conditioned Reinforcement Learning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.624081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:40.480698Z digest=sha256:8b7b186818e89f04cd7dab99683c0ba06f5189b68ef1904a0b9a7cc8d0fd498f

Observation 95c2ff49-a3ee-44e9-b473-67f0eae07cd3 · outbound

This paper cites Video prediction by modeling videos as continu- ous multi-dimensional processes.

Decision Flow Policy Optimization Video prediction by modeling videos as continu- ous multi-dimensional processes

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.881671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:40.557491Z digest=sha256:a659af2cd550bd5621df9e8bdd7f48d9f8b3b6422c79ce9043bf8467d14133f8

Observation a9556fd5-b4ff-43c3-ac7f-117dbc136841 · outbound

This paper cites Ensemble reinforcement learning: A survey.

Decision Flow Policy Optimization Ensemble reinforcement learning: A survey

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.775098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:40.620346Z digest=sha256:559609a3efa02225cec810e53629695a816788ce25e7207e5e2bc65b8b0da532

Observation 21c57fae-dcc9-4d7d-b384-60b51ad6ebd2 · outbound

This paper cites Flowllm: Flow matching for material generation with large language models as base distributions.

Decision Flow Policy Optimization Flowllm: Flow matching for material generation with large language models as base distributions

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.668710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:40.692841Z digest=sha256:d1c4e1647c70389af7791acd7db93e33aef42cea3ab0a09ebf73ef901f650e53

Observation c137d2e6-7826-482f-a691-963b72e1096a · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Decision Flow Policy Optimization Reinforcement learning: An introduction, volume 1

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.739672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.739672Z digest=sha256:40fc42b215601c420fd1bb82248a6160be19410774d9fab8b3867102eb5969d5

Observation f4f26b48-b58e-421b-a12c-173f4c8ed829 · outbound

This paper cites Imitationflow: Learning deep stable stochastic dynamic systems by normalizing flows.

Decision Flow Policy Optimization Imitationflow: Learning deep stable stochastic dynamic systems by normalizing flows

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.526585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:40.800133Z digest=sha256:3583fe3a95086e1673212715d4a287c881747c4fda6b2f5a64a31bd1dc86d994

Observation 791942d1-952a-4327-8c45-833414bec49b · outbound

This paper cites Deep reinforcement learning with double q-learning.

Decision Flow Policy Optimization Deep reinforcement learning with double q-learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.893058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.893058Z digest=sha256:76dfcc71c5f4e9c0231787c4fef0747ff15f44d19c7320f8dee72d99ce8f3b52

Observation 4141acbe-6d66-4c89-9103-23b54f1a66c8 · outbound

This paper cites Matrix calculus operations and taylor expansions.

Decision Flow Policy Optimization Matrix calculus operations and taylor expansions

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.377627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:41.012416Z digest=sha256:50bc088bc6c082c471694bacabdaecdb073aadd0d8aacae68d475f1076274208

Observation 4abbb693-68da-4591-a9cf-10b94f6997cf · outbound

This paper cites Bootstrapped Transformer for Offline Reinforcement Learning.

Decision Flow Policy Optimization Bootstrapped Transformer for Offline Reinforcement Learning

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.534167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:41.057275Z digest=sha256:99dbc2a5f0e178680598efd929a1dcc6e0e8a6169778979aecaf78b1eb62e581

Observation e161bd7c-6c5e-48ef-a647-89c905ddf6f5 · outbound

This paper cites Prioritized Generative Replay.

Decision Flow Policy Optimization Prioritized Generative Replay

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.144029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.144029Z digest=sha256:a2463e87c4cfa7cf312d3d2ebbe60a17ea4c6ccb7e93ad3bc4327b4fb790d69b

Observation d52a9e49-a01c-4f4b-9059-477a16ecc566 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Decision Flow Policy Optimization Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.247019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.247019Z digest=sha256:f88ad024dd543cbb9d2b984865f83df63366cee9d544e6587ec81a774fbb628d

Observation b99e1af5-4281-4452-af04-637c5620b629 · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Decision Flow Policy Optimization Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.291215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:41.312573Z digest=sha256:fec63781d00c5ee85cfab28e9401af177fcdb1b010f184b2e9a8cc076eebd4c7

Observation 432d8a0e-306f-423c-b0bf-dd14ab8a6b5e · outbound

This paper cites A Behavior Regularized Implicit Policy for Offline Reinforcement Learning.

Decision Flow Policy Optimization A Behavior Regularized Implicit Policy for Offline Reinforcement Learning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.361128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.361128Z digest=sha256:ecb97fbdf43c5aafca04e2a5ac44ce48a157d28cd4fdf12635e92aa3263322f8

Observation 9bc740fc-e567-476a-b66e-cd6ed73ada50 · outbound

This paper cites Policy-to-language: Train llms to explain decisions with flow-matching generated rewards.

Decision Flow Policy Optimization Policy-to-language: Train llms to explain decisions with flow-matching generated rewards

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.429401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.429401Z digest=sha256:8d023319fbb6cbdbcee5fa5718ccf77119f182111ee2350f3d384e71727bb01a

Observation fb0554b1-a476-4222-a6be-78aec2d74647 · outbound

This paper cites Flow to control: Offline reinforcement learning with lossless primitive discovery.

Decision Flow Policy Optimization Flow to control: Offline reinforcement learning with lossless primitive discovery

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.222293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:41.499645Z digest=sha256:ada5734e935e024836f03a9929b06c7a28a1c4e588611a028fe957563b828b16

Observation f99fefd8-8556-41f7-a0d0-c7de39ef5974 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Decision Flow Policy Optimization Mopo: Model-based offline policy optimization

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.575848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.575848Z digest=sha256:dd6a90ad8244e1cac0f780be9943b19da18fdef9cfbb99b68fefa79ab00881d7

Observation 1598d88e-dd30-43c9-beef-1edfb68bf0ce · outbound

This paper cites Combo: Conservative offline model-based policy optimization.

Decision Flow Policy Optimization Combo: Conservative offline model-based policy optimization

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.631335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.631335Z digest=sha256:30cf919acf4ab0979e70ab6d64edb3b1cc7f816da78229a95927c64e73e42577

Observation 57bd9d08-6527-423c-9fec-719af8a1f20a · outbound

This paper cites Affordance-based robot manipulation with flow matching.

Decision Flow Policy Optimization Affordance-based robot manipulation with flow matching

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.678599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.678599Z digest=sha256:a65477053ebf1ee88c3b00e9df9b1d2f95ffe2ce34e72e7418c1c2dc09320386

Observation 0840203d-2792-4814-b7b9-6059b49ff997 · outbound

This paper cites SaFormer: A Conditional Sequence Modeling Approach to Offline Safe Reinforcement Learning.

Decision Flow Policy Optimization SaFormer: A Conditional Sequence Modeling Approach to Offline Safe Reinforcement Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.739960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.739960Z digest=sha256:c486ac8786a4cb6bb7f2ef408862032e8c2543f5b471922eb48513db3d323dfd

Observation 863b8cb8-2f9d-4b82-9b87-409e89188567 · outbound

This paper cites Energy-Weighted Flow Matching for Offline Reinforcement Learning.

Decision Flow Policy Optimization Energy-Weighted Flow Matching for Offline Reinforcement Learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.812673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.812673Z digest=sha256:2c5e30fda9f2bb2060db3a0d33695ac518b729d16808352bde766149c5bf018e

Observation 1319d61f-646e-43b1-acc6-e301f0bad3a3 · outbound

This paper cites Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning.

Decision Flow Policy Optimization Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.215919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:20:41.895384Z digest=sha256:6453ca9a0cac68cc51f7c8ee1c8bd57f3de4ef5955f00b40fe8ab5f7681b521b

Observation edd50545-d00f-4770-bc6e-9e6cce26abe9 · outbound

This paper cites Guided Flows for Generative Modeling and Decision Making.

Decision Flow Policy Optimization Guided Flows for Generative Modeling and Decision Making

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.964655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.964655Z digest=sha256:2627641d7173cfdb5a819bc5d420b903f96857c14e3417924d8ce759ba2828b1

Observation f4db1893-30e5-461d-acef-c364e040e909 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Decision Flow Policy Optimization Open-Sora: Democratizing Efficient Video Production for All

Reference 96

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:20:42.023257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:42.023257Z digest=sha256:1d6de9dcb21e868cc367d41415326dc78f66a0947f4a00ca0620d41ad4db90bd

Pith citing papers

No inbound Pith citation observations are available.