Pith. sign in

Paper Citation Record · LEDGER

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

As of 21 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 10 inbound Pith citation observations for arXiv:2412.06782.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06782 v3

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:22:28.412011Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:13:47.283364Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 55f7dcc9-fa0c-4ada-bd86-3894b7ab388b · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.147601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.147601Z digest=sha256:a0301d49cab17d76defbffd4a59ac132026aece6e42571a174617d3ed55342a5

Observation b8f4d257-2fa2-4a97-81e7-5c34804c2660 · outbound

This paper cites Is Conditional Generative Modeling all you need for Decision-Making?.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Is Conditional Generative Modeling all you need for Decision-Making?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.152321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.152321Z digest=sha256:c43f17080046663b051db871f747585f5373f478f7d727106ede257f07ad6dfe

Observation c427b20e-0ef0-4a9f-af84-8e07907cf130 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.156930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.156930Z digest=sha256:c2d16e9e220ed867564a132c7e700dfc92df74b9f5eb8224f8b350e6ca894a5b

Observation 001896e7-0b5f-4dcc-b911-bbc4069bf5e2 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.160877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.160877Z digest=sha256:0fcd93d83e04a7f8737ad58bb449f3208bfb9fd648ab15b1e6e92f42332bb7bd

Observation d2fa5dfb-c75e-4fc0-a0a4-438f2d67d263 · outbound

This paper cites From Imitation to Refinement -- Residual RL for Precise Assembly.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction From Imitation to Refinement -- Residual RL for Precise Assembly

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.164738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.164738Z digest=sha256:89b34819c5df5a97af52a813e5c2e187cf9a74ae069fdf96d56ab2d56ffe58be

Observation 1c416812-8b1b-4f4f-8d08-e7dbeaa16d04 · outbound

This paper cites End to End Learning for Self-Driving Cars.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction End to End Learning for Self-Driving Cars

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.168612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.168612Z digest=sha256:7e47377cf95dcee01b83c9e540fe427fdc513627d2877801162367fa73f8f090

Observation b519844a-7607-41f4-8b74-1cd4eb869f43 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.172355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.172355Z digest=sha256:7ca38f30f36b37f20c662a1512d5ab380821ffd54f8618e9f518909291048212

Observation 78a86c1d-e4eb-48c0-8851-651fd6113430 · outbound

This paper cites Language Models are Few-Shot Learners.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.176569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.176569Z digest=sha256:7c88ad9772b9d01f5833a7f35eadbcd839ebff8fc28c75873a0ccc79cfc2aa16

Observation 2ae3f363-1b6e-4b95-957d-39af23e01c8c · outbound

This paper cites Polydoros, So- nia Chernova, and Aude Billard.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Polydoros, So- nia Chernova, and Aude Billard

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.238101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.180558Z digest=sha256:95cce446cfc69baa9d93151b58f36238388061ee71a9dd832cbbc225db7c4394

Observation e8dd733d-cb3a-43a0-93c6-cfa197012d6d · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.185153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.185153Z digest=sha256:d51c3e6877c1e14ef7863ebc5bc039d71601fbd4a1d22add6d309d595f6295d3

Observation 38ee46f8-dd74-4065-a3b6-d1134d7c630f · outbound

This paper cites Decision transformer: Reinforcement learn- ing via sequence modeling.Advances in neural information processing systems, 34:15084–15097, 2021.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Decision transformer: Reinforcement learn- ing via sequence modeling.Advances in neural information processing systems, 34:15084–15097, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.227770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.189573Z digest=sha256:243643a7e47aa93bdfc5c7a3b3af921f255d2a634783f3cff3ef2e9fc095d41c

Observation 914c00b8-9438-4960-b2b8-d8909561d4c1 · outbound

This paper cites Diffu- sion policy: Visuomotor policy learning via action diffusion.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Diffu- sion policy: Visuomotor policy learning via action diffusion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.216464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.193225Z digest=sha256:40f108e62c5c55dcb7ded0fdedeebea29c5ac2ff0554344b79a552a32af45f4b

Observation de4ad8de-9ee2-4755-89f0-18ebc2e4b292 · outbound

This paper cites From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.196622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.196622Z digest=sha256:189bf2c92099d39476c58a300bc98de2e9fc194517633a368f92d15035acf11f

Observation a63c7b83-8310-4f3b-b318-0a943d440e62 · outbound

This paper cites Quar-vla: Vision-language-action model for quadruped robots.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Quar-vla: Vision-language-action model for quadruped robots

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.205419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.200114Z digest=sha256:48db15d9d7a79800de0f028c108ef2fa093af4c7de487a3452d1c44dca83a30b

Observation 7aa1a872-acc1-4e60-9c14-3710ef472aaa · outbound

This paper cites Taming transformers for high-resolution image synthesis.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Taming transformers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.193772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.204676Z digest=sha256:a799353374f703f8fa7978a1cf80c8a4959808c14994b72e405b1f5dc8e84de7

Observation 1fc3a0e9-c209-4741-ae06-55ee5b259b66 · outbound

This paper cites Implicit behavioral cloning.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Implicit behavioral cloning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.181708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.208588Z digest=sha256:8fd76f37e9ceb399bd716880e2c8f1af12c8ab4abaa6e6ca9b830fb400de2b56

Observation 8a50c6ac-fe68-46ce-8fd7-72f9543a142c · outbound

This paper cites DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.212045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.212045Z digest=sha256:d84d84d2eeca2e968a4f9bbd49c82744a11d2182a71cea353e96899acfd50546

Observation c2634ca2-8f26-4c60-b588-d76f96c92c49 · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.216681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.216681Z digest=sha256:04d198017a8bd43e3971a9e00c772238e765cf3111d537cc2ccde45a51dc29fe

Observation 3aca2982-0e05-4e34-b917-711ad41fbb36 · outbound

This paper cites An exponential moving average algorithm.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction An exponential moving average algorithm

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.171023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.220712Z digest=sha256:c05a42f48fc4c3e1eb6e1ae5be774349a5872e1f335275c293452354f94c3dbc

Observation 3b602760-c557-401a-89f4-09219ff2a71f · outbound

This paper cites Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.The International Journal of Robotics Research, page 02783649241304789,.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.The International Journal of Robotics Research, page 02783649241304789,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.160030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.224507Z digest=sha256:538a983d67960cad28a4dfb436dd72632f5bb6b0938bebf086e4df3cfe948a05

Observation 9a4d993f-bfae-4c43-9523-72dce98a7eca · outbound

This paper cites Denoising Diffusion Probabilistic Models.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Denoising Diffusion Probabilistic Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.228203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.228203Z digest=sha256:76d273de844c520fa80902aeadf1961a0e7760d108920933235286deb96adca3

Observation 77482673-08e7-4aad-9327-e7729069d9d2 · outbound

This paper cites Towards accurate image coding: Improved au- toregressive image generation with dynamic vector quantiza- tion.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Towards accurate image coding: Improved au- toregressive image generation with dynamic vector quantiza- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.148588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.232414Z digest=sha256:b6e3a12d36487a1e774142eb1c67df7ad9a638c54ea9c63af6876899cfac5ae3

Observation 762e4144-809b-4629-8d8b-f0918c4de2ea · outbound

This paper cites Offline rein- forcement learning as one big sequence modeling problem.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Offline rein- forcement learning as one big sequence modeling problem

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.137061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.236497Z digest=sha256:2f97526c7059b592e6e26b154bc7e5c214881b719d28a83740fd473bb68940cc

Observation e98018b5-2429-49a9-a44f-5b426cd8d4ea · outbound

This paper cites Planning with diffusion for flexible behavior synthe- sis.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Planning with diffusion for flexible behavior synthe- sis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.127061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.239975Z digest=sha256:bb41d79f5f437b1c22a3fee16567f8a4d5cce6ef7200f2df928a37c3f7da4cc2

Observation 9bc55d83-81f6-41a5-a220-408c5a2fa07c · outbound

This paper cites Strictly batch imitation learning by energy-based distribu- tion matching.Advances in Neural Information Processing Systems, 33:7354–7365, 2020.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Strictly batch imitation learning by energy-based distribu- tion matching.Advances in Neural Information Processing Systems, 33:7354–7365, 2020

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.114397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.242769Z digest=sha256:4e61bf30983993b714cb01ee331ac4006246be4dc6f8b43de14b6d950e6f7658

Observation eb34e3c4-eb03-479a-9028-d697fd8d2fc9 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction VIMA: General Robot Manipulation with Multimodal Prompts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.245813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.245813Z digest=sha256:fee6c29ba4916c80bbdce5f0246356a59c2846fa0568021b6c7d0a9515545cbb

Observation 252e95b5-6ac4-4316-94f9-f3996aafbf52 · outbound

This paper cites Scaling Laws for Neural Language Models.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Scaling Laws for Neural Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.249548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.249548Z digest=sha256:de081204dd5a2da8bb06ca1a6dd49d7d54ee426ef74c6f24b0fac4228bfe22c1

Observation d5222d84-d23c-407f-832a-c5d985a0820c · outbound

This paper cites Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.252888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.252888Z digest=sha256:f3800de4f896cc8feefd63641c0e58e4b5056a02810fa424e686be1c683e7247

Observation e8fd5b9d-5af4-4db1-bd40-d89d19eb1701 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction OpenVLA: An Open-Source Vision-Language-Action Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.256724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.256724Z digest=sha256:d7059b83c18a92443231ef62bd6d6c88e3bbe149fecad976b7698aef34363514

Observation 53596f3c-0799-47e7-8088-f69ed4f6114c · outbound

This paper cites Ac- tion chunking as policy compression.PsyArXiv, 2022.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Ac- tion chunking as policy compression.PsyArXiv, 2022

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.100496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.259976Z digest=sha256:14913711ce3beffc28fc5c654185e7675b7e03e6cf36e1017f00e89b6ed1034f

Observation 7e965d40-4f03-4329-90c0-431db59500d1 · outbound

This paper cites Autoregressive image generation using resid- ual quantization.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Autoregressive image generation using resid- ual quantization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.081194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.263462Z digest=sha256:2b99e7682165bbef0868c912decbfe121a380ddd71265f669cd18b8333ea0a9d

Observation b8ed92d7-533b-4f2d-8521-7175cc068eca · outbound

This paper cites Behavior Generation with Latent Actions.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Behavior Generation with Latent Actions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.266343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.266343Z digest=sha256:c509c10bda2a8d420933717a94fe042e01ad38520126dc141bdd473d5066b27a

Observation 7486c90b-204f-4f2b-8b9e-e26d128809bb · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Vision-Language Foundation Models as Effective Robot Imitators

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.270135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.270135Z digest=sha256:d8fbcad617273edf26d2c8f67ee2b169ea429f4d66840aacc6480537a21107f2

Observation df8efca0-fc8c-4a60-8b5f-7d84a0996983 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.274515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.274515Z digest=sha256:ed85382007b0e6987a9704b7e42aa5c7ec11d27042b4e19db00d9d084289999e

Observation 77995876-d201-4f9d-a96a-d0c30c576aa8 · outbound

This paper cites ControlVAR: Exploring Controllable Visual Autoregressive Modeling.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction ControlVAR: Exploring Controllable Visual Autoregressive Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.278269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.278269Z digest=sha256:82a34dc6bbdeafdc30aa5f94489eb1bd55c26e7d66c5acfe1afb9cf16a3e731b

Observation 62aba07f-80e3-4a93-8961-56202fd9e4ef · outbound

This paper cites Skilldiffuser: Interpretable hierarchical planning via skill abstractions in diffusion-based task execution.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Skilldiffuser: Interpretable hierarchical planning via skill abstractions in diffusion-based task execution

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.066514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.282066Z digest=sha256:4d1abb3ff82255b3678a518a92da254e4fe5a0d5d0df26afd17d56fecb7b9138

Observation c10e4b1d-b693-4682-b66c-6675b28f2fa9 · outbound

This paper cites Pite: Pixel-temporal alignment for large video-language model.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Pite: Pixel-temporal alignment for large video-language model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.051001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.285661Z digest=sha256:6dfda5c5758f1cac536a2185598c71a0ad1d15086c0a8ee4fa15ee8599b9cf38

Observation a88f7a37-4cb9-469a-955b-9c1c99cda92e · outbound

This paper cites ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.289653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.289653Z digest=sha256:72caa33e5b3daca85ee26b1544d49ca1d0ebeb22d7d36efcf37188b2e4ddbe02

Observation 4cbf180a-f2ca-4f2b-b571-6f43657ea68e · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.293295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.293295Z digest=sha256:2cab1928a6bf192059ef27348f11570964b71e06a00a2eb3b3085877aa19ed4f

Observation 0d4fd0d6-9f95-4456-89ce-650711ff5f39 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.297075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.297075Z digest=sha256:89552b1c09085cb21663268882b96d468d400773b959604d40c763da14da017b

Observation cf1667db-e74f-4f6d-953f-ed85392cb3c2 · outbound

This paper cites Mimicgen: A data generation system for scalable robot learning using human demonstrations.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Mimicgen: A data generation system for scalable robot learning using human demonstrations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.038084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.300633Z digest=sha256:b5290a20153f3f46f10166b844fcf08f4b7a695ddb65e38dc57aff497c52662d

Observation 27dfdf9f-d1f0-47ce-ae0c-a0f145047505 · outbound

This paper cites Semantic image synthesis with spatially-adaptive nor- malization.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Semantic image synthesis with spatially-adaptive nor- malization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.304227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.304227Z digest=sha256:b2f4d9985b9cc54b3214dfa65a084ea5c5b8c016592f773aa8950863621a9964

Observation bf6f9798-af8f-400e-9f17-a321027b3c2b · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:29.011647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.307708Z digest=sha256:be14db33da33b6d3f1f3771c4b0c033d6e9f6b2a36d9759a2480f5e999008230

Observation 8a009dd7-5ee9-479d-ad69-3522f8e44a32 · outbound

This paper cites Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.311290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.311290Z digest=sha256:c7a372f922fb089a06e6bbeae00ec9c3363c312092b99a36843e4e9da816c000

Observation 7051cfd0-4d89-4887-a15a-778824c6e2a4 · outbound

This paper cites Efficient Autoregressive Audio Modeling via Next-Scale Prediction.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Efficient Autoregressive Audio Modeling via Next-Scale Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.314982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.314982Z digest=sha256:9dd82ff9104c01846360b34872dec8868317358ccbb81a294e7fa1deb013f651

Observation 4aa16135-56c6-4a3e-9d0b-c75ebd5f734b · outbound

This paper cites Language models are unsuper- vised multitask learners.OpenAI blog, 1(8):9, 2019.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Language models are unsuper- vised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.996780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.317965Z digest=sha256:f1fa5da68f2f65bb122f6fad301e7d285b40300fecf162445349efe60df51511

Observation c1f3afd6-ecd6-4d61-9767-88bb6012c4a8 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Learning transferable visual models from natural language supervi- sion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.320989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.320989Z digest=sha256:a8da184ad613c02d0fa7c3e93ae2a7d89e4116be272a17596d4926e72404e14e

Observation e5a155de-70e5-4c92-b896-a327479be701 · outbound

This paper cites Robot Learning with Sensorimotor Pre-training.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Robot Learning with Sensorimotor Pre-training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.323827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.323827Z digest=sha256:08b6506763ef67f6305e7f567b343d11407f494e098f67dd38496de6a9a1e3ef

Observation 05918570-0700-4f25-af1b-1e78a3790e96 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.973031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.326790Z digest=sha256:a28a40d8ece90f40e6d4dc86e92be85b20ec5aa5ba30ec3613dc53367588109c

Observation 01158c68-f1af-4b5e-bae8-c94732cd9b81 · outbound

This paper cites Generat- ing diverse high-fidelity images with vq-vae-2.Advances in neural information processing systems, 32, 2019.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Generat- ing diverse high-fidelity images with vq-vae-2.Advances in neural information processing systems, 32, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.961873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.330058Z digest=sha256:9215e5c9195c675f65e751345c4d91ca765270edd47040b3fb425fa378c98146

Observation 0f85a2dc-aa93-4e11-92ff-89991c07a101 · outbound

This paper cites A Generalist Agent.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction A Generalist Agent

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.332675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.332675Z digest=sha256:f9473cc8b86e66abebfd937783d863720a4c9f7fc88b60c8ce671f67c6e77e61

Observation b8d4309b-a497-4cbc-88b7-befd4d283595 · outbound

This paper cites Goal-Conditioned Imitation Learning using Score-based Diffusion Policies.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.335779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.335779Z digest=sha256:58b1aefec8f98f0286ee96940c34e92a0388605f5da75177d56a363f99a95008

Observation 14a9b48f-a295-4372-9279-d9dcdf8547ea · outbound

This paper cites Multimodal diffusion transformer: Learn- ing versatile behavior from multimodal goals.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Multimodal diffusion transformer: Learn- ing versatile behavior from multimodal goals

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.951802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.338744Z digest=sha256:658ebb9fa3944e68f6b736d6fd0d660ebe0ec5039a16f3ff7e45951c73c407af

Observation d39de70a-237b-42bc-b8cb-ecd03fe71db7 · outbound

This paper cites Behavior transformers: Cloning k modes with one stone.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Behavior transformers: Cloning k modes with one stone

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.941992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.341472Z digest=sha256:c1d096aa36681871375904324870895f21fcb064609bae007c5f51e674cb0ff2

Observation bf34ffde-07a6-46bb-8f60-fc04621c9d6e · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.344118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.344118Z digest=sha256:12540ce53a8d496dc3a61bd2f5771e58844b9ff88f4ea7b7dff81311bcdefe7c

Observation b788ffa5-bb49-42c1-a629-0b04ef00baf7 · outbound

This paper cites GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:22:28.554827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.347973Z digest=sha256:f5c063c8502eb0ff2b76fc0c14edf0ab1d68733cdf88f2a3599bd1facc21108d

Observation 0d750f89-31d2-46c0-bcec-09b0200bb6e6 · outbound

This paper cites Consistency Models.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Consistency Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.351313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.351313Z digest=sha256:2ee373e11d5bdff52bc3bc34cdbe119d0b71e977c5a05fc710ee68a58618bc99

Observation 0f74ff47-3b4a-4485-ac9a-d3d92ce3a10e · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.355116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.355116Z digest=sha256:41e57a381c80b70e1186b67661d4f6ddd6cbe3e5caeaedd051c34f15c2a6ad46

Observation e03dc90f-18dc-4529-ba1b-bc641c8986de · outbound

This paper cites Plex: Making the most of the available data for robotic manipulation pretraining.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Plex: Making the most of the available data for robotic manipulation pretraining

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.931120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.358602Z digest=sha256:5d7f840c446c5595cf779b95629773fcedb061eef6e0b3938ef9b007bb840894

Observation 30b47bc2-4b7f-4b97-b100-47b90c6b854a · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.361854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.361854Z digest=sha256:2b69ded4f1f0f4a6d6dd716020e9c16ff667fb2a45b85932d495e156decbc917

Observation 491f0713-0e5c-4b19-873d-b0d9dc2de843 · outbound

This paper cites Behavioral Cloning from Observation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Behavioral Cloning from Observation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.365101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.365101Z digest=sha256:decc627f44013febaec8330b5ab69627256ec55e5f1d891908b09d78cbc28606

Observation d364c54c-f4d1-4008-9a7c-934c48825366 · outbound

This paper cites Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.919405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.368932Z digest=sha256:3b9ce13d1d58dbf3e79527cbbbd7794d7626f9bc4044262f9bc5eb97fd64a595

Observation b4922ef0-a369-49aa-af7a-c623ceede681 · outbound

This paper cites Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.372333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.372333Z digest=sha256:1545e92300445f9bc93a3dfc94d8abc46df95c8e76e3831b731f6a46f09278b3

Observation f7dce8d4-e50d-4d92-a28e-cc81386bcd5e · outbound

This paper cites EquiBot: SIM(3)-Equivariant Diffusion Policy for Generalizable and Data Efficient Learning.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction EquiBot: SIM(3)-Equivariant Diffusion Policy for Generalizable and Data Efficient Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.376130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.376130Z digest=sha256:093b50ca134fed7a5cf61817d450c3b572e9695b2f8ea53128373d853432e351

Observation 4dd89a9f-db35-4717-a184-9829602d4a1b · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Vector-quantized Image Modeling with Improved VQGAN

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.379707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.379707Z digest=sha256:1a832aadd55122e8177dfa70e315c5ce7fb717b053a01634ecceaaa2e6ff0eb4

Observation d32554af-0201-4da8-9c7a-2de2fa4f684c · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.905795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.385008Z digest=sha256:d5769e31e29af450e263b49cc7d264423cdfe5eddbaeb330490fe6ed303df87b

Observation d1bda2db-d591-4b42-8d1e-863e5f488065 · outbound

This paper cites Transporter networks: Rearranging the visual world for robotic manip- ulation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Transporter networks: Rearranging the visual world for robotic manip- ulation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.892670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.388708Z digest=sha256:19d90c8fb4da730c50867ba406855fac6be5158ee8949e6d3c30ba8d3a3bdb7c

Observation 7194d88f-b819-4f6c-ae10-bc017b05377f · outbound

This paper cites G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.392246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.392246Z digest=sha256:cb100ba820d734df66ceace2f48b6e4775fc1e5ae01328fd7a61dc153e76caea

Observation ff3b3d24-35ba-4631-9f8d-e3251be183ef · outbound

This paper cites VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.396044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.396044Z digest=sha256:75889f9160415b46807ed0a1986122cb65e9d1ca23afa416f889e0b5410a0aea

Observation b394ca45-d1aa-4f08-8dff-452d83f96a42 · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:28.878979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T19:22:28.399192Z digest=sha256:a2f43d441ed6fbfb3c77fcf29425e78a673568eb463425e85d8cc692c9bdf84e

Observation 45b55927-226e-4144-af31-bcdfe6413a28 · outbound

This paper cites Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.402112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.402112Z digest=sha256:3bf12a1318d5829aeccc2bf778305cdd1394e44e0fb9e987c7d21f0c78cb9c11

Observation d25bf4eb-98d2-4d8f-82b5-49f5fbe2705f · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.405358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.405358Z digest=sha256:d7fe51fd4955c9d251904e61190dd7c97ca7679946837108edf4798b7ec2e6b9

Observation caab3125-931d-4add-8e04-9f12c436a9c7 · outbound

This paper cites On the continuity of rotation representations in neural networks.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction On the continuity of rotation representations in neural networks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.408569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.408569Z digest=sha256:268e82087b0f70e365672f3cad35856e749fad36cac1dc4af11e797bfc3c1db2

Observation f682314d-c34f-4b78-9e06-1c11a350ab1b · outbound

This paper cites robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction robosuite: A Modular Simulation Framework and Benchmark for Robot Learning

Reference 74

Resolution
malformed identifier
no resolver link, observed 2026-08-11T19:22:28.412011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.412011Z digest=sha256:91c0383b8e5e2011cc6a2f7294813aee61c41674382f7dc233b3daa3cc8ddccd

Pith citing papers

Observation d2ca76bf-d49e-46ba-b45a-5f38693d2bc5 · inbound

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation cites this paper.

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T10:56:57.626304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:56:57.626304Z digest=sha256:46678c7803fa9fa725b737a790e7b7dba161f1365c815641bfe027e90b4efb6e

Observation 97da9e2f-65cb-4d5d-acf4-6413c2ef0867 · inbound

STDArm: Transferring Visuomotor Policies From Static Data Training to Dynamic Robot Manipulation cites this paper.

STDArm: Transferring Visuomotor Policies From Static Data Training to Dynamic Robot Manipulation CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:47.283364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:47.283364Z digest=sha256:a643d392a37cf644b2b400fcc0b5048487794833ea3f23cf7a2820ee9a55cddc

Observation b97efb9a-6cdd-40d2-aa35-e58f0a5d033a · inbound

OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation cites this paper.

OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:46:51.110229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:46:51.110229Z digest=sha256:a545be9277f37c9882aea20f1918c4c4198d4fb0944df7cc175b06cb046cf24a

Observation 5405284b-1eb7-4a8a-9e6f-744a851ee246 · inbound

H$^3$DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning cites this paper.

H$^3$DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:12:23.057880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:12:23.057880Z digest=sha256:5d958f8388d3c256651b271bc57cfb37abc100d4235f1cf91927033ae2da0dce

Observation 4c838aa9-c32d-4012-9f29-c71ba02061b4 · inbound

CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion cites this paper.

CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:35.630082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:35.630082Z digest=sha256:0e4f657c8a691495ef3a576d5821c63d4f7c4ffedeea49d69d49d8b5f0bdc644

Observation 9ab494ec-5051-41ee-8bf8-e93bc122867c · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.179410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.179410Z digest=sha256:ef5a3c4a72e93b03b1f7292fb11d753f59bdb8875a8de5cca35f70f90a61e986

Observation cbd469a7-e88f-4f30-be27-ac30c1732ccc · inbound

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation cites this paper.

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:26:40.608665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:26:40.608665Z digest=sha256:7138b5d93f5eb99c37e18ce3b0307bd31dcc3b5ad8e3213f7484789262d9155a

Observation 459d07c5-153d-40d0-b6f9-28e6a5b918cc · inbound

Referring-Aware Visuomotor Policy Learning for Closed-Loop Manipulation cites this paper.

Referring-Aware Visuomotor Policy Learning for Closed-Loop Manipulation CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:48.198410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T19:59:02.039358Z digest=sha256:9bbd5c943d35307874cf01d8d9fd1602c02a278c461b789272dd3e0ac7eda991

Observation df958cfe-5cd8-4d25-9ca6-671dfe47e1a2 · inbound

HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning cites this paper.

HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:54.973221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:27:54.107915Z digest=sha256:20798ae3b7b8cefb65bd5937148a155436de26607ceb6d54157423eed3321427

Observation a756e857-d266-474f-9365-5d22f5e8c62b · inbound

Hierarchical Policy Learning via Spectral Decomposition cites this paper.

Hierarchical Policy Learning via Spectral Decomposition CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:04:20.723275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T06:54:29.971971Z digest=sha256:b6ad44a47fb601109a91fb0a5b49cc9603c2332e7a11436419fe24bfb0822bf4