Pith. sign in

Paper Citation Record · LEDGER

Score-Based One-step MeanFlow Policy Optimization

As of 21 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2605.23365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23365 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:21:52.748214Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:36:10.359049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T14:37:15.953758Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact9
  • verified fuzzy23
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e16b1662-262e-4316-8611-47a11a2f0a76 · outbound

This paper cites Is Conditional Generative Modeling all you need for Decision-Making?.

Score-Based One-step MeanFlow Policy Optimization Is Conditional Generative Modeling all you need for Decision-Making?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.060227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8bb67d073c664a5fc1b2cb935cd2f303fe10d6c9a2fa8b19884de4d6793c401c

Observation fe2a9ef8-634b-4dca-a98e-0ff79b082d4e · outbound

This paper cites Iterated denoising energy matching for sampling from boltzmann densities.

Score-Based One-step MeanFlow Policy Optimization Iterated denoising energy matching for sampling from boltzmann densities

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.169906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8060a91f41cdaeab7bb23f311b2cc8ce051c3aa122922ee2acdcbace3cd77ec2

Observation e835e503-70fd-45ff-b5ee-bae68fe6b10f · outbound

This paper cites Score regularized policy optimization through diffusion behavior.

Score-Based One-step MeanFlow Policy Optimization Score regularized policy optimization through diffusion behavior

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.150860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:880bdea129579429323c8bcf7fa8ba027c63de59876af89b2eb7b41030b00699

Observation dac36487-d07e-49c8-9cf7-6fcfb7e5e17d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Score-Based One-step MeanFlow Policy Optimization Diffusion policy: Visuomotor policy learning via action diffusion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.154017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:70da6ca9258a0a616caea0bebbe87b190d0d9254fa3a401e0c2b41e6392244e6

Observation c6566601-a61d-4164-a854-fefbedd48cc9 · outbound

This paper cites Diffusion-based reinforcement learning via q-weighted variational policy optimization.

Score-Based One-step MeanFlow Policy Optimization Diffusion-based reinforcement learning via q-weighted variational policy optimization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.137898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:4f0642690dfec293adc3bea189e6e92d191e9b842403dc94586b634a5e1259bf

Observation c4fac339-7b83-440c-9ee2-ed955f3510f0 · outbound

This paper cites One step diffusion via shortcut models.

Score-Based One-step MeanFlow Policy Optimization One step diffusion via shortcut models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.156901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:c029495186557bd9098556cca94789b37faffb1f0199b247e2ecc133414e86eb

Observation 4aebde4d-8f20-4261-b539-b27e5d06e10a · outbound

This paper cites Mean flows for one-step generative modeling.

Score-Based One-step MeanFlow Policy Optimization Mean flows for one-step generative modeling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.160118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:2b58c5d847e811350b99fb75d48843fb0b437c79e7cf93c18c90ae37aab6e509

Observation 80ca5bd5-9953-4cc2-b317-dd9e1246883c · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Score-Based One-step MeanFlow Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.166758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:6ed4f20584c73378b9ba4abaf2c0100dc892509bf460dca443b4fa8b90252ba6

Observation e145d41d-209f-4da1-aa6a-88860d1d8024 · outbound

This paper cites Denoising diffusion probabilistic models.

Score-Based One-step MeanFlow Policy Optimization Denoising diffusion probabilistic models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.173743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8b3f84eba1b874ebb1a81c4a820541bb291ad7741d742c013ea59ce3f6cc93b3

Observation fb889358-7f67-45a3-a65c-8679d722f27d · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Score-Based One-step MeanFlow Policy Optimization Planning with Diffusion for Flexible Behavior Synthesis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.049945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:05b80fb0084a4c84c6f532ccf7ba32a01d0b10639e2d467af6d8c5af7b082d15

Observation b6ac2a15-732d-442c-b150-1d29aa4051a7 · outbound

This paper cites Prior-guided diffusion planning for offline reinforcement learning.

Score-Based One-step MeanFlow Policy Optimization Prior-guided diffusion planning for offline reinforcement learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:26:39.035451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:22abdfecbed96e4fba243cc5db59b458c054a5263dd6569f0fe75e52a7d0fd9e

Observation 492b20f6-9de9-4291-bda5-490666b1ca5e · outbound

This paper cites an unresolved cited work.

Score-Based One-step MeanFlow Policy Optimization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-25T11:56:57.122735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:d16e99a53163c646301c1bf6159f4d573fdd37330489657efe0c4b768d72af98

Observation 07029e77-d08d-4d2a-bac3-0bbb973597dc · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Score-Based One-step MeanFlow Policy Optimization Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.108656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:ed2f21f7b2912d38df4f958f58ada6db75cb08fe9f9630944260d8863e414368

Observation d8c9407d-4691-446e-8b3d-fdb98f9f369a · outbound

This paper cites Simplifying, stabilizing and scaling continuous-time consistency models.

Score-Based One-step MeanFlow Policy Optimization Simplifying, stabilizing and scaling continuous-time consistency models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.180228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:b5a236f9423e3a0d97330f461e65b0fe51e3eaeed44bd74a7ae96d7890de0d52

Observation e791643d-1a99-4c84-a9fc-2adffdbb1539 · outbound

This paper cites Efficient online reinforcement learning for diffusion policy.

Score-Based One-step MeanFlow Policy Optimization Efficient online reinforcement learning for diffusion policy

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.186599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:bc59c09e734e57368bfe88f2cf187ad2e25c235c676cf54f6590dfce833e84e1

Observation 0b95ac95-a5d2-49c9-ab48-91457004b53e · outbound

This paper cites Learning a diffusion model policy from rewards via q-score matching.

Score-Based One-step MeanFlow Policy Optimization Learning a diffusion model policy from rewards via q-score matching

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.141016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:3e7fa5c5d01844e31d0c98d077d9789eeb8d7c0d59d3a8bd3a7b587a3395b17f

Observation dfd762db-b4d8-4092-912e-22a5e0bae85f · outbound

This paper cites Diffusion Policy Policy Optimization.

Score-Based One-step MeanFlow Policy Optimization Diffusion Policy Policy Optimization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.040336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:e0add1974658b998e5c0146526dbcebf008aff220fcae95e0091917cf2e5ba6f

Observation d6c1805b-c865-41ad-a931-fde69f2eb2e1 · outbound

This paper cites Progressive distillation for fast sampling of diffusion models.

Score-Based One-step MeanFlow Policy Optimization Progressive distillation for fast sampling of diffusion models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.119449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:d3b76ad831110bb21dc024db81b1374e31398100f2a6d52113a176142cd4eef4

Observation b8694883-1729-4fa4-a207-7653430550c3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Score-Based One-step MeanFlow Policy Optimization Proximal Policy Optimization Algorithms

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.045159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:314db3c68ef19dff0028489a7913961f56a62c4d326887b103ab96435e9bfa56

Observation f15f3bab-bf9b-4d07-bee2-f661700fecd3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Score-Based One-step MeanFlow Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.054865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:cc392a61b88d4c490a371a052cdb4beccc24bee143d7c3563f47d7f4d0a11247

Observation 7333592f-794c-4f02-bfc1-5aca43672304 · outbound

This paper cites Denoising Diffusion Implicit Models.

Score-Based One-step MeanFlow Policy Optimization Denoising Diffusion Implicit Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:25:23.791074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8feb42a06ce9578ca0233bb2b67a551612205b436f8dd2448da74074814366c6

Observation 763f04c7-b259-44b1-b986-4d16d979675a · outbound

This paper cites Consistency models.

Score-Based One-step MeanFlow Policy Optimization Consistency models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.126168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:43655ecb3a5b200bb0016423656f94b465bd3e758a11a48d058e3c408d39c6b0

Observation 8b9826e1-7c37-4249-807e-6aacddd5eddb · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.Advances in Neural Information Processing Systems, 32.

Score-Based One-step MeanFlow Policy Optimization Generative modeling by estimating gradients of the data distribution.Advances in Neural Information Processing Systems, 32

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.177176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:e0a0210c09f5be818423103d7bb54188fb607dfa5648003f3ee3ab64a9de272f

Observation 724ac68d-ab5e-4319-a08a-376234bc98d7 · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Score-Based One-step MeanFlow Policy Optimization Score-based generative modeling through stochastic differential equations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.147342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:ed99aa7fc2d98bf38c3c19c2a095153c795fbb55a5d1ee56a194acc8d85b9e33

Observation 589c7b6b-05c1-41fd-a090-a312414a711e · outbound

This paper cites MIT press Cambridge.

Score-Based One-step MeanFlow Policy Optimization MIT press Cambridge

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.115682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:cd4c80ad637499950f5ddfb2cddb27a0c8cb317a968d9333951e68610381ebe7

Observation 0d1e467d-bd2a-4a05-968b-fa0232021a69 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Score-Based One-step MeanFlow Policy Optimization Mujoco: A physics engine for model-based control

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.183032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8622afc68abd8d80bda1f5bd8d938208200213bda1e9eb82c9c95244b090751d

Observation 2fb51365-d222-4e43-98f5-7ba26eb10f64 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Score-Based One-step MeanFlow Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.029853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:af63aa0c6287ce556e3bcaf7a55a884e4aaaf3e719d8916454d0fbe71a637663

Observation 8ccdbf9e-525a-4d85-b8e7-04ee27ca024f · outbound

This paper cites Diffusion actor-critic with entropy regulator.

Score-Based One-step MeanFlow Policy Optimization Diffusion actor-critic with entropy regulator

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.163332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:6e923df288fecc154d78cfc2549f16467215d78b13f0f8f2f6a25548a5df409b

Observation 6fcb2bcd-c38f-441e-b0a7-ca7b2faff9cb · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Score-Based One-step MeanFlow Policy Optimization Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.130845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:01c0e552ebd7f6cdb179cd428d67637f91c0c128aca0abf79608fc8cefb1887c

Observation c49ab6de-e540-40d5-97ee-f5641be5a5e2 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Score-Based One-step MeanFlow Policy Optimization Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:26:39.024863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:280466aafd8b2af8a42efdfc1b35223d952a966e998d036b6013dceaa2647599

Observation e1970c2d-ee3b-4ff8-a082-f1acb63f6d56 · outbound

This paper cites One-step diffusion with distribution matching distillation.

Score-Based One-step MeanFlow Policy Optimization One-step diffusion with distribution matching distillation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.144048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:83a09a9e618aee9c1d28b53bcc6dc6cd72b9b6ff66efe33e405cf61353e5e184

Observation 1ffc8f5f-c9b4-4a12-86ba-e4dc119b479b · outbound

This paper cites Mean flow policy with instantaneous velocity constraint for one-step action generation.

Score-Based One-step MeanFlow Policy Optimization Mean flow policy with instantaneous velocity constraint for one-step action generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.134704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:a01e8f939be650d713ea91e024ee29377bd41256a023ccb39f5ed3fbb93c45fc

Observation 23a9e2cc-d14d-47f3-95c9-589d4565155f · outbound

This paper cites The final reward is given by the normalized mixture density, producing a smooth multimodal reward landscape with values in[0,1].

Score-Based One-step MeanFlow Policy Optimization The final reward is given by the normalized mixture density, producing a smooth multimodal reward landscape with values in[0,1]

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.112390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:b7465150529abf5f6bea3e2b35e1a22d829d5f05956b1a8afa317bb2c7f5e7c1

Pith citing papers

Observation c8575dd3-97fd-4132-989d-5617a15fd9d5 · inbound

Expressivity and Statistical Trade-offs in Diffusion Policy Learning cites this paper.

Expressivity and Statistical Trade-offs in Diffusion Policy Learning Score-Based One-step MeanFlow Policy Optimization

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:37:15.956480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:36:10.359049Z digest=sha256:fe23e67da59784cd8379264d02a8df9e4fa72839d581596ac5779f8829291f6a