Pith. sign in

Paper Citation Record · LEDGER

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies

As of 6 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2604.24622.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24622 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:33:32.677434Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact41
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4db0091-02fb-4af6-9b52-4c4e6797b2ca · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.511304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:9c68ceeabb53586955491fd7a70b46d05ba670d7c579a1f3b11a62966b955e29

Observation cc8a1bdc-262e-4d06-b904-113e05806cab · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:41:17.595310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:67519b742e19831056a0591a2aec5f6f16a79410c6fdde3b165759cc4f497b8f

Observation fa00c6a2-7811-4707-849c-0c81c10ad620 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:17.603111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:6322da81148363b3f9f6962786ce58eb5495fef8d2f5db3475f2f24f9c15464d

Observation d044066d-e0eb-4774-bf9e-e2cac14d96e5 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:28:07.265740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:87054bcd6188f1157f47b02da3f92e1276984e6b9b46aaf911ea2020782885ee

Observation 95565bb6-4325-4a47-aaa2-d5dc7ae27ad5 · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.514523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:687f125fdf1cdc7271927f10b74b97d46112556884d98149fdbd5e46232b6465

Observation a44099be-7509-4a0c-a065-0062ec0113d5 · outbound

This paper cites Neural Ordinary Differential Equations.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Neural Ordinary Differential Equations

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T13:00:58.588645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:9e65d1c004af6c0fbc43d2ec1f6a4864e1b72cdf6bdb4618b4e8e8d26a5b1130

Observation 8e4f56d7-0f47-4011-8e16-bc6c8fea8bdf · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T00:20:30.996755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:973a444e3cf49610dbe8ca34b3a87478c25d7ee5ae386909a952896b20785d9c

Observation 461bc3a3-8e83-4df4-9c90-c073b557e7c2 · outbound

This paper cites Everydayvla: A vision-language-action model for affordable robotic manipulation.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Everydayvla: A vision-language-action model for affordable robotic manipulation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.378626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:da1a3c52e97d759a2561fd8d3cf339833cc710ad71c0887bf5a232c33847cffb

Observation ac936cdf-ca40-4529-8ad1-52c65c94325c · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T00:04:27.600547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:f08712ce3ad187419ce95bc7ba984ee5e1ebaabeb4e4ff4e2fde052cf46925e9

Observation fd72977d-f9bd-4527-96bd-9e79208691d1 · outbound

This paper cites Diffusion Transformer Policy.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Diffusion Transformer Policy

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.441942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:697712625b0917ebda334b3c988475a23a9c1c5d865053c6f1c90ddb5213140f

Observation 095c613e-c21c-443e-8a60-13fc10ea80f3 · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-09T00:04:27.587044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:6494bf43bbd02d78a585f07039dc38fa681e95f3dcb48d82934f55126f6d68ea

Observation 7a01f14a-cc69-404d-9562-ea6970dc1cb3 · outbound

This paper cites LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.468112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:03e185be06f1e199ce64bc240c5d32fe36fd1df8701a7c728c852afa1feecd58

Observation 19d4402e-75a9-4e94-a305-cc1192d31fdd · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.497620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:8b5b744723b0146896a1f3680f67acaaa9ffdcd74082a2aa14b96aa2d17aa0c5

Observation 0777bc97-bac7-4a89-884d-d64394c78e23 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:41:17.336064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:90ec03a0facd84359acc1d764a629c9952a3c541ff094bfbada92ecfaad2c657

Observation 8e820b19-1c1c-45eb-a04d-77ac4cddb05e · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 15

Resolution
verified exact
doi, observed 2026-05-09T00:04:27.607003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:bbf14b9c9cb1abeb01feed1dbd2b5acbcc4c4ea1cb2fd6559c6be557b61b1825

Observation ec6b7c93-a076-41e3-85e4-11f7cedf296b · outbound

This paper cites Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.362544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:a9754fcaf432d6d9267cd52c4d37460a5543e440df6ebb978062ad2c6673b0c4

Observation 985a483f-720c-47d0-bfee-64795c6762e6 · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.494424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:82947ab50348fd057959354c161e46425eb33c33183e5a2c7167a416472b6a7d

Observation da674afa-b932-40e8-ad81-fa030b02b84f · outbound

This paper cites IEEE Access8, 199523–199538 (2020) https://doi.org/10.1109/ACCESS.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies IEEE Access8, 199523–199538 (2020) https://doi.org/10.1109/ACCESS

Reference 18

Resolution
malformed identifier
doi, observed 2026-05-09T00:04:27.591736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:3bc918172a47e446729fcf5f21b34bcccfd7f44a10d0c4d10d5fbeade0db4f9b

Observation 6a480ba9-be12-4ea3-acf4-34753c1b6b6e · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:17.372080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:55fd4d0e6aba50631f699100cc553e8cc73e4d1d35c5261695d0faaa7e3556b6

Observation 27dfbcfc-33d2-40f5-9794-3827c4c897c4 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:13.054714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:14708dceacb4ba6ed035e088ab1a53de6bec669341f21f4f41792b8cb0866f6a

Observation 849cac78-29b2-4cee-a8c2-88c98a2135f0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies OpenVLA: An Open-Source Vision-Language-Action Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:17.129070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:d237d808179a553515133d28234c9b8cb05805bc8fdcaa46ec16f80dcee386bd

Observation 27f610aa-a74b-46e3-8728-570d184afe15 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies MolmoAct: Action Reasoning Models that can Reason in Space

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:35:23.409560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:6c7387d473bcb1a366aa12312675e27d58a43ef716cd25a308542e293be02359

Observation 79d9053d-eb92-4e87-938a-fed39f189647 · outbound

This paper cites Motion Manifold Flow Primitives for Task-Conditioned Trajectory Generation under Complex Task-Motion Dependencies.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Motion Manifold Flow Primitives for Task-Conditioned Trajectory Generation under Complex Task-Motion Dependencies

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.247334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:15ae073a0d969e8137a4456641d281989b44581ccb5a79a2430a7de8daa3a892

Observation 35f991a9-c143-45d4-8d1c-0f96dc745542 · outbound

This paper cites CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.137638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:6b1ba7e02f49f59d834f6b175cc1d32e24302cdc22b2de7643eac5345504f40b

Observation 5de76cfd-694b-4aa1-8eb5-614911f7d18b · outbound

This paper cites Flow Matching for Generative Modeling.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Flow Matching for Generative Modeling

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:17.220370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:969d53b23a4fe5753adde0e79d4aa437bdd94da11ea4053535cfcfec72a93245

Observation 031a9715-2d42-4456-bbb6-a70638520db9 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:04:42.196761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:4fd3510296790cee0d6cef098eaa64056d8244258d45c78a0fa3430d0dffbbe0

Observation 1cfd74bd-31e6-4b2c-a57a-0cc27c0ca102 · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.488117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:d38114cfa379cc71ea0dceae570b1a200007c00ec6ac159891a8a370b1ad87f7

Observation 17bb1c14-c5c0-484e-bfaa-89b3807347e0 · outbound

This paper cites RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:17.483455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:d0e54e3450287daed62468f8601ecfacdab0da1fdfe3dbec36deac4d6f96e4bb

Observation f9998e60-4259-450d-9ad7-0cbd19e30a84 · outbound

This paper cites From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:17.316875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:60e0ac47bc963b5f62afd0e8e205e1a510552086c7df75204c01d5d7843864e4

Observation 267efd11-25d1-447a-9d35-80655bf77735 · outbound

This paper cites DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.310361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:c01b242a26a03b7af701b3519e26354df47318ca1bd0c8fa7126d92720836056

Observation d3c496a3-870e-44d0-940f-25dfc6f85ba2 · outbound

This paper cites Grounding Language with Visual Affordances over Unstructured Data.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Grounding Language with Visual Affordances over Unstructured Data

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.237636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:76ff52fa5f1c4126e6738de409bed389d19adee2f7c030eb85e2451a55aee078

Observation 3d8d5d28-d80b-4154-80df-c2812abfc048 · outbound

This paper cites What Matters in Language Conditioned Robotic Imitation Learning over Unstructured Data.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies What Matters in Language Conditioned Robotic Imitation Learning over Unstructured Data

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.275614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:0484af2fa76a23b85b4f1e7cd0caafe64fa37c6c430a54498f229fdfd8028fbc

Observation 9c884ffd-5d0e-405c-90ba-168aecccae11 · outbound

This paper cites CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.301807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:d1463590803e04328aa6bcb41512e6d1c9d8d0f1828ea6a16950eab9de6d2478

Observation 82265426-da11-4e7f-8d4c-6d791e5b6363 · outbound

This paper cites Much ado about noising: Dispelling the myths of gener- ative robotic control.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Much ado about noising: Dispelling the myths of gener- ative robotic control

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.356378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:cd7896c21d789c9c89425abf16f6bf90b1867a1134540bd357a089ea577936e0

Observation 213be866-abb3-4160-b188-6d28d09b6f19 · outbound

This paper cites FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.347821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:1227f49dafbc70952e222db6c2581fafecdcb5d862ea19048547acceda360872

Observation 79d3b403-90e9-4519-925a-b5dbc414507e · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.491387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:c85e50f13fc3f887f337f0df19f58bb95d53e9f2d1e7c0317ad1578a3666f172

Observation 931e186a-74ee-431a-a788-864391bcccf0 · outbound

This paper cites Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:17.564383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:2b377f2bc3a6e2cd0ecafb2af04c43a0587e8341cbe2539a3f04cd244fc5eb5d

Observation 8bf41518-4f59-4004-becd-6e4b7b41af3f · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Progressive Distillation for Fast Sampling of Diffusion Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:17.323334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:f461bec02bda1ff4e3f3299cb7328309331b7e212e1909acad10fc43ecdc1637

Observation 21c40d15-e509-44f4-8d61-8348ce8861c8 · outbound

This paper cites Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.822386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:7bee5ec1cdca6c835d52f3149d3be1aefcd1096d56cbc3e74a6d7fd3c68c4929

Observation 18ed64d7-1c73-4de1-933b-a949d8e30391 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:43:24.547911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:738a132a92f539dfa08e53a681792a431606f10e7a7642ce49d60ad35e14930c

Observation 4ef7912e-fb32-432b-a927-c6692122d21c · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.507966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:a4284908c5de426efa3b7460b2b71a08920646028684cfa346ecd7672aed1993

Observation daf749a8-fbe9-494a-a599-14d7e314287b · outbound

This paper cites Nina: Normalizing flows in action.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Nina: Normalizing flows in action

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.421654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:cf77b09cf005f9197284cde8308cad98edc2eb5ae53b92a6ff1f6d396a77c3c2

Observation c7bcba76-6aee-4ecf-aa5d-2dfe5acc0bf6 · outbound

This paper cites Unified Vision-Language-Action Model.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unified Vision-Language-Action Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.572350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:0a1ea2d15bb572a3e70739e4f63dbeb703a5e4f2834a1132dcf4e37ccee048c9

Observation e3371cd5-dfda-4a41-8872-d76a8e449280 · outbound

This paper cites Diffusion Models for Robotic Manipulation: A Survey.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Diffusion Models for Robotic Manipulation: A Survey

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.198471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:6f689018246edca27022e8ddededf9dde2c67dcdb4b1cf7c4732fb233c530caf

Observation a2e36438-1fc2-412a-85f8-9d385d47290d · outbound

This paper cites On-Device Diffusion Transformer Policy for Efficient Robot Manipulation.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies On-Device Diffusion Transformer Policy for Efficient Robot Manipulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.268531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:dbb83be575ef63fd9258991dd84d5da1f6b68e7b6e08d0409d382735dff595aa

Observation ee2f2627-f23f-440b-8582-7bceab4bb8de · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.501678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:b1d05a257b64726e67c0bd59e2c54ec1a9e2f975845b449bf66895e45b9d4dd7

Observation b827fd55-e022-45e9-9adb-ecf942e9ee43 · outbound

This paper cites arXiv preprint arXiv:2502.02175 (2025).

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies arXiv preprint arXiv:2502.02175 (2025)

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.253696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:4c28d7c49907831454a6a1c8d5ed438ffc208640f57b4d9544db62ac47572cd7

Observation a857de87-181b-40e2-b008-7fb6990aeebc · outbound

This paper cites RoboTron-Mani: All-in-one multimodal large model for robotic manipula- tion.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies RoboTron-Mani: All-in-one multimodal large model for robotic manipula- tion

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.543056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:f212eead8fdf579413276e3fd038bef8980eaac10c15690f8f9fe30fdc0e8695

Observation c1ac4bea-1027-461d-aa7a-388bbc8d7f6f · outbound

This paper cites Instructvla: Vision-language-action instruction tuning from understanding to manipulation.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Instructvla: Vision-language-action instruction tuning from understanding to manipulation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.228657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:e8d67f7f155efdf4250cd78524b5c536886d02a89caa0cfb4c80f1f9611f0081

Observation c171cee7-8765-4db5-ace2-30fc0e93c3c8 · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:17.552448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:95fceae9b94da2b044d0eb43cee108c1ecc3bc101ea344eb0631460ef6fb8af3

Observation 3bfa6642-5f83-4cc1-8284-2ff754ee907a · outbound

This paper cites Dysl-vla: Efficientvision-language-action model inference via dynamic-static layer-skipping for robot manipulation.arXiv preprint arXiv:2602.22896.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Dysl-vla: Efficientvision-language-action model inference via dynamic-static layer-skipping for robot manipulation.arXiv preprint arXiv:2602.22896

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.513281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:5b82d20c04a2e924f804c1f38de66eb591fe90f40f8b62c13e39aa8c52df86a6

Observation 408fe971-2d02-4c9d-8dbd-edd5e82aa782 · outbound

This paper cites arXiv preprint arXiv:2510.24795 (2025).

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies arXiv preprint arXiv:2510.24795 (2025)

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.523744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:ea1361b0daa3b0b6976e31f047864babc0c3985522a4f29d692455aaf7540f47

Observation 7b9b3274-5c39-4ca7-89eb-1146134812dc · outbound

This paper cites DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.536187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:1e7af216fcc0f698e9007fbdd55487e58bdd2776d4d7252318ed3ac571fd1a98

Observation eabca663-b967-4b56-a2d6-a3eaed0216fe · outbound

This paper cites Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.578376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:5bc7d1723ed9cf5564b799ff582e5447cce0435e7e08e57e8af0c8a20823f0c2

Observation 3f77fe12-352e-4ab7-8edf-5b23af1a12b1 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.804405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:b22b63186b183ac0bd94f4a434f15a5f8b738fb6c63f07e22d6fe019c02ac74f

Observation 06e6e170-3532-48a5-8649-73579ce0f612 · outbound

This paper cites an unresolved cited work.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-26T20:48:21.504681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:b3b2ca5fe440aff31a3a225a3ce9edb0fdd58fe5387557b5ff5ac3908d0f5687

Observation 47295cb1-c337-4dff-aff8-44590b897454 · outbound

This paper cites Universal Actions for Enhanced Embodied Foundation Models.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Universal Actions for Enhanced Embodied Foundation Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:17.452556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:c35d6ecea2730c0754860be0a2a6b7813b6b3c82363d3932114f36494ad40692

Observation 1f5889d6-edf2-415f-a17f-a7265e30ec36 · outbound

This paper cites X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Reference 58

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T14:57:47.957311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:5af30884d1cf67a8e420ea37e50b0c0d78c93beb7dd57670a392e2d85c329df0

Pith citing papers

No inbound Pith citation observations are available.