Pith. sign in

REVIEW 3 major objections 4 minor 68 references

Feeding raw patch tokens, not compressed image summaries, to robot policies sharply boosts manipulation success — a 40% relative gain over global-feature policies and a win over a 7B-parameter VLA using roughly 0.7% of its parameters.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:35 UTC pith:GQK53BPK

load-bearing objection Useful empirical result, but abstract numbers don't match the tables and the 0.7% parameter claim conflates configs. the 3 major comments →

arxiv 2607.18236 v1 pith:GQK53BPK submitted 2026-07-20 cs.RO cs.LG

Patch Policy: Efficient Embodied Control via Dense Visual Representations

classification cs.RO cs.LG
keywords dense visual representationspatch tokensvision transformer policiesblock-causal attentionbehavior cloningmanipulationimitation learningfrozen visual backbone
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Patch Policy claims that the limiting factor for vision-based robot manipulation is not model scale or language grounding, but access to spatially dense visual features. By feeding frozen, pretrained ViT patch tokens directly into standard transformer policies through a block-causal attention mask, it preserves fine-grained spatial detail that global pooling or CLS tokens discard. Across four simulated and three real-world manipulation suites, this simple change yields a 40% relative improvement over state-of-the-art global-feature representations and even outperforms a fine-tuned 7B-parameter vision-language-action model while using a fraction of the compute. The paper's message is that the robotics community can immediately benefit from advances in visual representation learning without inheriting the cost of billion-parameter generative models.

Core claim

The central claim is that dense pretrained patch features from ViTs are a powerful, underutilized representation for robot policies, and that a minimal architectural change—accepting all patch tokens per observation with a block-causal attention mask—lets standard transformer policies exploit them. This yields consistent gains over global-pooled or CLS-token features, and competitive or superior performance relative to heavy VLA baselines. The paper shows that spatial compression of features, whether via pooling or learned downsampling, degrades performance, while the choice of pretrained encoder (WebSSL or DINOv2 over SigLIP2) materially affects results. The authors frame the contribution a

What carries the argument

The block-causal attention mask: patch tokens within a single observation frame attend fully to each other (intra-frame bidirectional), while attention across frames remains causally masked to preserve temporal order. This lets a transformer policy reason over T×P tokens per context window—T frames times P patches—without violating the causality required for sequential decision-making. The mask is what allows the policy to consume dense patch features directly instead of a single global token, and the paper shows it outperforms both full attention (which leaks future info) and token-causal masking (which fragments each frame).

Load-bearing premise

The paper assumes that the four simulated and three real-robot suites are representative of 'embodied control' more broadly, so that the observed benefit of dense features over global features generalizes beyond these precise, spatially demanding manipulation tasks; if the intended scope includes semantic, language-conditioned, or navigation-heavy tasks, the claim is unsupported.

What would settle it

A direct falsifier would be a manipulation task with strong semantic or goal-conditioning demands (e.g., LIBERO's language-based tasks or a navigation benchmark) where a CLS-token policy matches or beats the patch-token policy, or a real-robot comparison with more than 100 trials showing the patch advantage shrinking to statistical noise.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Robot policies can achieve large gains on precise, multi-object manipulation without fine-tuning the visual encoder, by simply switching from global-pooled features to frozen patch tokens.
  • Heavy, billion-parameter VLAs are not necessary for in-domain manipulation learning; a lightweight transformer with frozen dense features matches or exceeds them while running at 11 ms inference latency.
  • The quality of the pretrained visual representation remains a primary bottleneck for policy learning, suggesting that continued progress in self-supervised visual representation learning will directly transfer to robot control.
  • Spatial compression of visual features, whether by pooling or learned downsampling, degrades control performance, so preserving patch resolution is recommended whenever compute permits.
  • The approach is policy-architecture-agnostic, working with both VQ-BeT and Diffusion Policy heads, making it a drop-in replacement for existing transformer-based policies.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The 40% relative improvement is task-dependent: on LIBERO Goal, the CLS-token policy actually matches the patch policy (0.95 vs 0.94 for VQ-BeT), and the largest gains appear on spatially challenging multi-object tasks (BlockPush, Cube). The paper's own data suggest the claim is strongest for tasks demanding fine-grained spatial reasoning, not for all embodied control.
  • The real-robot results rest on 20 trials per task with no error bars; the reported gap between patch and CLS features (e.g., 0.70 vs 0.60 on cable insertion) is a difference of two successes. A reader should weight these numbers accordingly, while the simulation results with multiple seeds provide a more reliable basis.
  • The paper explicitly notes that SigLIP2, a semantically-oriented encoder, underperforms everywhere, while V-JEPA2 does not consistently beat WebSSL or DINOv2. This hints that the spatial-density hypothesis is about the kind of feature geometry, not just the patch format: encoders trained with dense geometric objectives outperform language-aligned ones for manipulation.
  • A natural testable extension is whether the same block-causal dense-feature recipe transfers to other control paradigms, such as reinforcement learning or multi-task generalist policies, as the paper's Limitations section identifies RL as an unexplored direction.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes Patch Policy, a minimal modification to transformer-based behavior-cloning policies: instead of compressing each visual observation into a single global token, the policy consumes all dense patch tokens from a frozen pretrained ViT, using a block-causal attention mask over the flattened spatio-temporal token sequence. The method is evaluated with two policy heads (VQ-BeT and Diffusion Policy) and five frozen visual encoders across four simulated and three real-robot manipulation suites, plus additional zero-shot CAP experiments. The central claims are that dense patch features give a 40% relative improvement over global-pooled representations, that Patch Policy surpasses fine-tuned OpenVLA-OFT by 18% while using roughly 0.7% of the parameters, and that preserving spatial token density—rather than model scale—is the key ingredient.

Significance. If the quantitative claims survive scrutiny, the paper makes a useful contribution: it isolates representation density from model scale, shows that frozen off-the-shelf encoders can be used directly for control, and provides a drop-in architectural pattern compatible with existing transformer policy heads. The within-policy comparison is clean—same encoder, same policy head, only the tokenization changes—and the paper includes helpful ablations of the attention mask, spatial compression, and model size, plus a broad encoder sweep and real-robot evidence. The headline numbers, however, are not currently derivable from the reported tables, so the strength of the contribution is conditional on a revised, fully specified quantitative presentation.

major comments (3)
  1. [Abstract; Tables 1–3] The headline numbers do not follow from the tables under any stated aggregation. In Table 1, WebSSL Patch VQ-BeT versus WebSSL CLS gives per-task relative changes of +15% (Push-T), -1% (LIBERO Goal), +118% (BlockPush), and +630% (Cube); the ratio of summed raw scores is about +96%. WebSSL Patch DP versus AvgPool gives about +55% by summed score. No common pooling of these heterogeneous metrics (coverage fractions, success rates, mean block counts) yields 40%. For the OpenVLA-OFT comparison, the closest figure to 18% is a DP-only average of per-task relative gains from Table 1 (about 17.4%), while the VQ-BeT average is about 11% and the Table 2 final-stage relative gains are +133%, +42%, and +38%. The 0.7% parameter claim also refers to the DINOv2 ViT-S configuration used in the real-robot experiments, whereas the simulation results use WebSSL (334M total parameters, about 4.4% of OpenVLA
  2. [Section 3.3; Table 1; Figure 1 caption] The text states that Patch Policy 'consistently outperforms global representations' and 'outperforms the fine-tuned OpenVLA-OFT baseline ... on all four environments.' Table 1 contradicts this: WebSSL Patch VQ-BeT is below OpenVLA-OFT on LIBERO Goal (0.94 vs 0.95), below WebSSL AvgPool VQ-BeT on LIBERO Goal (0.94 vs 0.97), and WebSSL Patch DP is equal to WebSSL CLS DP on LIBERO Goal (0.98 vs 0.99). The effect is clearly task-dependent: gains are large on BlockPush and Cube, while LIBERO Goal shows parity or slight losses. Please replace 'consistently' and 'all four' with a precise statement such as 'matches or exceeds on spatial/multi-object tasks and is comparable on goal-conditioned tasks.'
  3. [Section 3.3; Table 1 footnote] The OpenVLA-OFT comparison on LIBERO Goal is not evidently controlled. The paper states that all policies in the study receive only visual inputs, but for LIBERO Goal the OpenVLA-OFT number is 'reported directly as provided in the original manuscript [25].' If that baseline used language instructions or a different evaluation protocol, the comparison is not apples-to-apples. Please either specify the exact conditioning and evaluation protocol of the OpenVLA-OFT baseline in this work or rerun it under the same input conditions as the other rows.
minor comments (4)
  1. [Table 2; Section 3.4] Real-robot results are based on 20 trials per task with no confidence intervals. The claim that the improvement is 'most significant' on Cable Insertion is not statistically supported. Please report binomial confidence intervals or raw trial counts, and avoid the word 'significant' unless a test is performed.
  2. [Section 3.5; Tables 7–8] The text calls DINOv2 'top-performing' in Section 3.4, but Tables 7 and 8 show WebSSL substantially ahead on BlockPush and Cube. Please qualify the claim as 'top-performing overall' or 'best on average,' and note the task-dependent ranking.
  3. [Section 3.6; Table 3] The 'as little as 0.7%' parameter claim should be paired with the WebSSL number (4.4%) where the simulation results are described, to avoid the impression that the headline 18% VLA comparison used the 0.7% configuration.
  4. [Section 3.7; Table 4] The compression ablation is run on Push-T only, with no standard errors or multiple seeds. The text says compression causes a 'significant decrease'; please soften to 'a large decrease in this environment' or add variance and additional tasks.

Circularity Check

0 steps flagged

No circular derivation found: the central claim is an empirical comparison, and the self-citations are not load-bearing reductions.

full rationale

The paper's central claim is an empirical measurement: policies consuming frozen, dense pretrained patch tokens are compared with policies using global-pooled or CLS features under matched policy heads (VQ-BeT, Diffusion Policy) in Tables 1, 5, 6, 7, and 8, and against OpenVLA-OFT in Tables 1 and 2. No parameter is fitted to the reported outcome and then renamed a prediction; the gains are observed rollouts, not constructed identities. The block-causal attention mask is borrowed from the authors' own DINO-WM paper [20], and CAP [27] is also self-cited, but neither citation forces the empirical conclusion. The paper applies an existing masking scheme to a new setting and measures its effect; it does not invoke a self-cited uniqueness theorem or smuggle in an ansatz that already contains the result. The result is also not definitional: dense patch tokens are not defined in terms of downstream success, and the global-feature baselines are independent architectures. The limitations section (frozen backbones, behavior-cloning-only evaluation) and the task-dependent results (e.g., LIBERO Goal CLS 0.95 vs Patch 0.94) weaken the generality of the claim but are not circularity. The abstract's headline numbers (40%, 18%, 0.7%) are not obviously reproducible from the tables under a stated aggregation, and the parameter-count comparison conflates the WebSSL/ViT-L simulation configuration with the DINOv2/ViT-S real-robot configuration; this is a reporting and reproducibility concern, not a circular reduction of the kind this pass flags. Overall, the derivation chain is self-contained as an empirical study, with no load-bearing step reducing to its own inputs.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The central contribution is an architecture/application, not a derivation. The only 'free' choices are standard hyperparameters and the attention-mask design; no constants are fitted to the headline result. The main burden is the domain assumption that frozen dense patch features transfer across tasks, and the specific encoding choices (1D positional embedding) are not independently validated.

free parameters (2)
  • Policy context window T (block size) = per-env: 5 (Push-T), 10 (LIBERO), 3 (BlockPush), 5 (Cube); 2 (real-world)
    Hand-picked per environment (Tables 9, 11). The attention-mask ablation varies masking, but T itself is not swept, so performance depends on this unsearched choice.
  • Policy backbone size (N, n_heads, d_emb) = e.g., 8/8/512 for many tasks; 6/6/120 for LIBERO VQ-BeT
    Selected per task; the model-size ablation (Table 13) demonstrates performance scales with capacity, so the reported numbers assume a generous, task-specific setting rather than a fixed architecture.
axioms (4)
  • domain assumption Frozen internet-pretrained ViT patch features retain sufficient spatial and semantic detail for closed-loop control; no encoder fine-tuning is needed.
    This is the central hypothesis; the paper only tests it on the chosen benchmarks, so any transfer beyond them assumes this.
  • domain assumption A 1D learned positional embedding over the flattened patch sequence is sufficient to convey the 2D spatial layout of each frame.
    Section 2.2: 'add a learned 1D positional embedding indexed by the token's position in the flattened sequence'. The paper does not compare with 2D positional encodings or coordinate-aware attention, so adequacy of the 1D layout is assumed.
  • standard math Standard transformer attention and the VQ-BeT/Diffusion Policy losses are taken as given and correctly implemented.
    Sequence-modeling properties of transformers (Vaswani et al.) and the two action heads are used without proof.
  • domain assumption The chosen benchmark tasks (Push-T, LIBERO Goal, BlockPush, Cube, plus three real-robot tasks) are representative of embodied manipulation, and the observed ordering of representations generalizes.
    The paper draws general conclusions about dense vs global features from these tasks; the claim of generality is an assumption beyond the measured set.

pith-pipeline@v1.3.0-alltime-deepseek · 19031 in / 13016 out tokens · 104738 ms · 2026-08-01T15:35:26.208884+00:00 · methodology

0 comments
read the original abstract

Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sacrificing both fine-grained spatial detail and the benefits of large-scale visual pre-training. While there exist policies that do operate on dense patch features like large vision-language-action models (VLAs), they tend to be heavy and slow, inheriting the full cost of a billion-parameter vision-language model (VLM) backbone. We close this gap with Patch Policy, a minimal architectural extension that enables transformer-based policies to consume dense pre-trained patch tokens directly without the computational overhead of a full VLM. At its core is a block-causal attention mask that preserves the temporal causality of standard policies while letting the model attend over many patch tokens per observation, alongside other state information. Patch Policy is lightweight, fast, and highly effective. Across four simulated and three real-world environment suites, our method achieves a 40% relative improvement over policies using state-of-the-art global-pooled representations. Furthermore, it surpasses fine-tuned OpenVLA-OFT by 18% while using roughly 0.7% of the parameters. We believe Patch Policy provides a pipeline for the robotics community to readily leverage continuing progress in visual representation learning, without sacrificing the training efficiency or inference speed required for high-frequency, reactive control. Videos can be viewed at https://patch-policy.github.io

Figures

Figures reproduced from arXiv: 2607.18236 by Ada Langford, Bowen Tan, Gaoyue Zhou, Lerrel Pinto, Yann LeCun, Zichen Jeff Cui.

Figure 1
Figure 1. Figure 1: We introduce PATCH POLICY, an efficient policy architecture that harnesses the power of pre-trained dense visual features. PATCH POLICY demonstrates superior performance while remain￾ing computationally lean in both parameter count and inference latency. Through a rigorous analysis across five state-of-the-art visual representations, we show that our dense patch-based approach pro￾vides consistent performa… view at source ↗
Figure 2
Figure 2. Figure 2: Architecture. PATCH POLICY consists of an observation trunk (left) with a policy head (right). We encode multiview observations into patch features and optionally concatenate goal em￾beddings (images/states) into the current timestep sequence. PATCH POLICY is compatible with any transformer-based policy head. A block-wise causal mask is applied to enforce conditioning on past observations, allowing the pol… view at source ↗
Figure 3
Figure 3. Figure 3: We evaluate PATCH POLICY on four simulated and three real-world environments. havior Transformer (VQ-BeT) [21], which uses a hybrid classification-regression loss, and Diffusion Policy (DP) [22], which uses a denoising objective. During training, we forward a sequence of patch tokens through the policy transformer trunk and action head, and compute a loss between the predicted and ground-truth actions for … view at source ↗
Figure 4
Figure 4. Figure 4: Real-robot rollout examples for the three evaluation tasks: cable insertion, pen collection, [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Comparison of PATCH POLICY across various pretrained visual representations. We report the mean and standard deviation across runs with three random seeds. Our results suggest that DINOv2 and WebSSL are the most effective vision backbones for robot learning tasks. We normalize BlockPush and Cube results to [0, 1] to facilitate comparison across the entire task suite. For all experiments, we freeze the imag… view at source ↗
Figure 6
Figure 6. Figure 6: We evaluate on 10 unseen objects for Real Franka Pickup, following the evaluation criteria [PITH_FULL_IMAGE:figures/full_fig_p018_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Successful rollouts of PATCH POLICY for Cable Insertion, Pen Collection, and Tool Hang￾ing [PITH_FULL_IMAGE:figures/full_fig_p022_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Failure rollouts of PATCH POLICY for Cable Insertion, Pen Collection, and Tool Hanging. 22 [PITH_FULL_IMAGE:figures/full_fig_p022_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: CAP real Franka object pickup successes. [PITH_FULL_IMAGE:figures/full_fig_p023_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: CAP real Franka object pickup failure modes. [PITH_FULL_IMAGE:figures/full_fig_p024_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Push-T, Cube, and LIBERO Goal environment Ours - VQ-BeT evaluation rollouts. [PITH_FULL_IMAGE:figures/full_fig_p025_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: EgoGym pick, open, and close PATCH POLICY evaluation rollouts. 26 [PITH_FULL_IMAGE:figures/full_fig_p026_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: CAP dataset pick, open, and close trajectory samples. [PITH_FULL_IMAGE:figures/full_fig_p027_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 28 linked inside Pith

  1. [1]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale.ArXiv, abs/2010.11929, 2020. URL https://api.semanticscholar.org/CorpusID:225039882

  2. [2]

    K. He, X. Chen, S. Xie, Y . Li, P. Doll’ar, and R. B. Girshick. Masked autoencoders are scalable vision learners.2022 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 15979–15988, 2021. URLhttps://api.semanticscholar.org/ CorpusID:243985980

  3. [3]

    Caron, H

    M. Caron, H. Touvron, I. Misra, H. J’egou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9630–9640, 2021. URLhttps://api.semanticscholar. org/CorpusID:233444273

  4. [4]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. InInternational Conference on Machine Learning, 2021. URL https://api.semanticscholar.org/CorpusID:231591445

  5. [5]

    X. Chen, S. Xie, and K. He. An empirical study of training self-supervised vision transform- ers.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9620–9629,

  6. [6]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . B. Huang, S.-W. Li, I. Misra, M. G. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. J´egou, J. Mairal, 10 P. Labatut, A. Joulin, and P. Bojanowski. Dinov2: Learning robust visual features...

  7. [7]

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierar- chical vision transformer using shifted windows.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9992–10002, 2021. URLhttps://api.semanticscholar. org/CorpusID:232352874

  8. [8]

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European conference on computer vision, pages 38–55. Springer, 2024

  9. [9]

    Tschannen, A

    M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. M. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y . Xia, B. Mustafa, O. H’enaff, J. Harmsen, A. Steiner, and X.-Q. Zhai. Siglip 2: Multilingual vision-language encoders with improved semantic understand- ing, localization, and dense features.ArXiv, abs/2502.14786, 2025. URLhttps://api. semantics...

  10. [10]

    Carion, L

    N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Sur ´ıs, C. K. Ryali, K. V . Al- wala, H. Khedr, A. Huang, J. Lei, T. Ma, B. Guo, A. Kalla, M. Marks, J. Greer, M. Wang, P. Sun, R. Radle, T. Afouras, E. Mavroudi, K. Xu, T.-H. Wu, Y . Zhou, L. Momeni, R. Hazra, S. Ding, S. Vaze, F. Porcher, F. Li, S. Li, A. Kamath, H. K. Cheng, P. Doll’ar, N. Ravi...

  11. [11]

    Sim’eoni, H

    O. Sim’eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Cou- prie, J. Mairal, H. J’egou, P. Labatut, and P. Bojanowski. Dinov3. 2025. URLhttps: //api.semantic...

  12. [12]

    D. Fan, S. Tong, J. Zhu, K. Sinha, Z. Liu, X. Chen, M. Rabbat, N. Ballas, Y . LeCun, A. Bar, and S. Xie. Scaling language-free visual representation learning.ArXiv, abs/2504.01017, 2025. URLhttps://api.semanticscholar.org/CorpusID:277467353

  13. [13]

    Bardes, Q

    A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y . LeCun, M. Assran, and N. Bal- las. Revisiting feature prediction for learning visual representations from video, 2024. URL https://arxiv.org/abs/2404.08471

  14. [14]

    Assran, A

    M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, et al. V-jepa 2: Self-supervised video models enable understanding, prediction and planning.arXiv preprint arXiv:2506.09985, 2025

  15. [15]

    K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. InPro- ceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  16. [16]

    Perez, F

    E. Perez, F. Strub, H. De Vries, V . Dumoulin, and A. Courville. Film: Visual reasoning with a general conditioning layer. InProceedings of the AAAI conference on artificial intelligence, volume 32, 2018

  17. [17]

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022

  18. [18]

    Dasari, M

    S. Dasari, M. K. Srirama, U. Jain, and A. Gupta. An unbiased look at datasets for visuo-motor pre-training. InConference on Robot Learning. PMLR, 2023. 11

  19. [19]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polo- sukhin. Attention is all you need. InNeural Information Processing Systems, 2017. URL https://api.semanticscholar.org/CorpusID:13756489

  20. [20]

    G. Zhou, H. Pan, Y . LeCun, and L. Pinto. DINO-WM: world models on pre-trained vi- sual features enable zero-shot planning. In A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu, editors,Forty-second International Con- ference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267...

  21. [21]

    S. Lee, Y . Wang, H. Etukuru, H. J. Kim, N. Muhammad, M. Shafiullah, and L. Pinto. Be- havior generation with latent actions.ArXiv, abs/2403.03181, 2024. URLhttps://api. semanticscholar.org/CorpusID:268248763

  22. [22]

    C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. doi:10.15607/RSS.2023.XIX.026. URLhttps://doi.org/10.15607/RSS. 2023.XIX.026

  23. [23]

    Z. J. Cui, H. Pan, A. Iyer, S. Haldar, and L. Pinto. Dynamo: In-domain dynamics pre- training for visuo-motor control.ArXiv, abs/2409.12192, 2024. URLhttps://api. semanticscholar.org/CorpusID:272709175

  24. [24]

    T. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware.ArXiv, abs/2304.13705, 2023. URLhttps://api.semanticscholar. org/CorpusID:258331658

  25. [25]

    M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success.ArXiv, abs/2502.19645, 2025. URLhttps://api.semanticscholar. org/CorpusID:276647709

  26. [26]

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre- training. InProceedings of the IEEE/CVF international conference on computer vision, pages 11975–11986, 2023

  27. [27]

    Z. J. Cui, O. Rayyan, H. Etukuru, B. Tan, Z. Andrianarivo, Z. Teng, Y . Zhou, K. Mehta, N. Wojno, K. Y . Wu, M. H. Anjaria, Z. Wu, M. Mao, G. Zhang, B. Shah, Y . Kim, S. Chintala, L. Pinto, and N. M. M. Shafiullah. Contact-anchored policies: Contact conditioning creates strong robot utility models.arXiv preprint arXiv:2602.09017, 2026

  28. [28]

    R. G. Goswami, A. Bar, D. Fan, T.-Y . Yang, G. Zhou, P. Krishnamurthy, M. Rabbat, F. Khor- rami, and Y . LeCun. World models can leverage human videos for dexterous manipulation. ArXiv, abs/2512.13644, 2025. URLhttps://api.semanticscholar.org/CorpusID: 283896258

  29. [29]

    D. A. Pomerleau. Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988

  30. [30]

    N. M. M. Shafiullah, Z. J. Cui, A. Altanzaya, and L. Pinto. Behavior transformers: Cloning k modes with one stone.ArXiv, abs/2206.11251, 2022. URLhttps://api. semanticscholar.org/CorpusID:249926747

  31. [31]

    Z. J. Cui, Y . Wang, N. M. M. Shafiullah, and L. Pinto. From play to policy: Conditional behavior generation from uncurated robot data.ArXiv, abs/2210.10047, 2022. URLhttps: //api.semanticscholar.org/CorpusID:252968170. 12

  32. [32]

    Florence, C

    P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mor- datch, and J. Tompson. Implicit behavioral cloning. InConference on robot learning, pages 158–168. PMLR, 2022

  33. [33]

    K. Rana, R. Lee, D. Pershouse, and N. Suenderhauf. Imle policy: Fast and sample effi- cient visuomotor policy learning via implicit maximum likelihood estimation.arXiv preprint arXiv:2502.12371, 2025

  34. [34]

    O’Neill, A

    A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024

  35. [35]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024

  36. [36]

    K. He, H. Fan, Y . Wu, S. Xie, and R. B. Girshick. Momentum contrast for unsupervised visual representation learning.2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9726–9735, 2019. URLhttps://api.semanticscholar.org/ CorpusID:207930212

  37. [37]

    Grill, F

    J.-B. Grill, F. Strub, F. Altch’e, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B.´A. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko. Bootstrap your own latent: A new approach to self-supervised learning.ArXiv, abs/2006.07733, 2020. URLhttps://api.semanticscholar.org/CorpusID:219687798

  38. [38]

    Caron, H

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF interna- tional conference on computer vision, pages 9650–9660, 2021

  39. [39]

    S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta. R3M: A universal visual rep- resentation for robot manipulation. In K. Liu, D. Kulic, and J. Ichnowski, editors,Confer- ence on Robot Learning, CoRL 2022, 14-18 December 2022, Auckland, New Zealand, vol- ume 205 ofProceedings of Machine Learning Research, pages 892–909. PMLR, 2022. URL https://proc...

  40. [40]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, K. Choromanski, T. Ding, D. Driess, K. A. Dubey, C. Finn, P. R. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y . Kuang, I. Leal, S. Levine, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, ...

  41. [41]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InNorth American Chapter of the Association for Computational Linguistics, 2019. URLhttps://api.semanticscholar.org/CorpusID: 52967399

  42. [42]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

  43. [43]

    Jaegle, F

    A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira. Perceiver: General perception with iterative attention.ArXiv, abs/2103.03206, 2021. URLhttps: //api.semanticscholar.org/CorpusID:232110866. 13

  44. [44]

    L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch. Decision transformer: Reinforcement learning via sequence modeling. InNeu- ral Information Processing Systems, 2021. URLhttps://api.semanticscholar.org/ CorpusID:235294299

  45. [45]

    Janner, Q

    M. Janner, Q. Li, and S. Levine. Offline reinforcement learning as one big sequence mod- eling problem. InNeural Information Processing Systems, 2021. URLhttps://api. semanticscholar.org/CorpusID:235313679

  46. [46]

    Y . Su, X. Zhan, H. Fang, H. Xue, H.-S. Fang, Y .-L. Li, C. Lu, and L. Yang. Dense policy: Bidirectional autoregressive learning of actions. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 14486–14495, 2025

  47. [47]

    Haldar, Z

    S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learn- ing.ArXiv, abs/2406.07539, 2024. URLhttps://api.semanticscholar.org/CorpusID: 270379931

  48. [48]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Perts...

  49. [49]

    S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gim´enez, Y . Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. D. Edwards, N. M. O. Heess, Y . Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas. A generalist agent. ArXiv, abs/2205.06175, 2022. URLhttps://api.semanticscholar.org/CorpusID: 248722148

  50. [50]

    Z. Hou, T. Zhang, Y . Xiong, H. Pu, C. Zhao, R. Tong, Y . Qiao, J. Dai, and Y . Chen. Diffusion transformer policy.arXiv preprint arXiv:2410.15959, 2024

  51. [51]

    O. M. Team, D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, P. R. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy.ArXiv, abs/2405.12213, 2024. URL https://api.semanticscholar.org/CorpusID:266379116

  52. [52]

    Goyal, J

    A. Goyal, J. Xu, Y . Guo, V . Blukis, Y .-W. Chao, and D. Fox. Rvt: Robotic view transformer for 3d object manipulation. InConference on Robot Learning, pages 694–710. PMLR, 2023

  53. [53]

    Shridhar, L

    M. Shridhar, L. Manuelli, and D. Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. InConference on Robot Learning, pages 785–799. PMLR, 2023

  54. [54]

    Driess, F

    D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tomp- son, Q. H. Vuong, T. Yu, W. Huang, Y . Chebotar, P. Sermanet, D. Duckworth, S. Levine, V . Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. R. Florence. Palm-e: An embodied multimodal language model. InInternational Conference on Machine L...

  55. [55]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky.π0: A vision-language-action flow model for general robot control.ArXiv, abs/2410.24164, 2024...

  56. [56]

    Intelligence, K

    P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. V...

  57. [57]

    Intelligence, A

    P. Intelligence, A. A. Amin, R. J. Aniceto, A. Balakrishna, K. Black, K. Conley, G. Connors, J. Darpinian, K. Dhabalia, J. DiCarlo, D. Driess, M. Equi, A. Esmail, Y . Fang, C. Finn, C. Glos- sop, T. Godden, I. Goryachev, L. Groom, H. Hancock, K. Hausman, G. Hussein, B. Ichter, S. Jakubczak, R. Jen, T. Jones, B. Katz, L. Ke, C. Kuchi, M. Lamb, D. LeBlanc, ...

  58. [58]

    T. Dao, D. Fu, S. Ermon, A. Rudra, and C. R ´e. Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35: 16344–16359, 2022

  59. [59]

    B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

  60. [60]

    S. Park, K. Frans, B. Eysenbach, and S. Levine. Ogbench: Benchmarking offline goal- conditioned rl.ArXiv, abs/2410.20092, 2024. URLhttps://api.semanticscholar.org/ CorpusID:273654871

  61. [61]

    X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kir- mani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao. Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024. 15 A Appendix A.1 Implementation and Baselines Our implementation and baselines are built upon the ...

  62. [64]

    DynaMo:https://github.com/jeffacce/dynamo_ssl

  63. [65]

    VQ-BeT:https://github.com/jayLEE0301/vq_bet_official

  64. [66]

    Diffusion Policy:https://github.com/real-stanford/diffusion_policy

  65. [67]

    OpenVLA-OFT:https://github.com/moojink/openvla-oft

  66. [68]

    The real-world tasks comprise: (1)Tool Hanging, (2)Pen Collection (long-horizon, small objects), (3)Cable Insertion(low-tolerance insertion)

    ACT:https://github.com/tonyzhaozh/act A.2 Environments and Tasks We evaluate PATCHPOLICYacross four simulated environments (Push-T, LIBERO Goal, Block- Push, Cube) with 2D-to-7D action spaces, and three real-world tasks using a 7-DoF Franka arm with a parallel-jaw gripper. The real-world tasks comprise: (1)Tool Hanging, (2)Pen Collection (long-horizon, sm...

  67. [2021]

    URLhttps://api.semanticscholar.org/CorpusID:233024948

  68. [2023]

    URLhttps://api.semanticscholar.org/CorpusID:260293142