REVIEW 3 major objections 4 minor 68 references
Feeding raw patch tokens, not compressed image summaries, to robot policies sharply boosts manipulation success — a 40% relative gain over global-feature policies and a win over a 7B-parameter VLA using roughly 0.7% of its parameters.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 15:35 UTC pith:GQK53BPK
load-bearing objection Useful empirical result, but abstract numbers don't match the tables and the 0.7% parameter claim conflates configs. the 3 major comments →
Patch Policy: Efficient Embodied Control via Dense Visual Representations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that dense pretrained patch features from ViTs are a powerful, underutilized representation for robot policies, and that a minimal architectural change—accepting all patch tokens per observation with a block-causal attention mask—lets standard transformer policies exploit them. This yields consistent gains over global-pooled or CLS-token features, and competitive or superior performance relative to heavy VLA baselines. The paper shows that spatial compression of features, whether via pooling or learned downsampling, degrades performance, while the choice of pretrained encoder (WebSSL or DINOv2 over SigLIP2) materially affects results. The authors frame the contribution a
What carries the argument
The block-causal attention mask: patch tokens within a single observation frame attend fully to each other (intra-frame bidirectional), while attention across frames remains causally masked to preserve temporal order. This lets a transformer policy reason over T×P tokens per context window—T frames times P patches—without violating the causality required for sequential decision-making. The mask is what allows the policy to consume dense patch features directly instead of a single global token, and the paper shows it outperforms both full attention (which leaks future info) and token-causal masking (which fragments each frame).
Load-bearing premise
The paper assumes that the four simulated and three real-robot suites are representative of 'embodied control' more broadly, so that the observed benefit of dense features over global features generalizes beyond these precise, spatially demanding manipulation tasks; if the intended scope includes semantic, language-conditioned, or navigation-heavy tasks, the claim is unsupported.
What would settle it
A direct falsifier would be a manipulation task with strong semantic or goal-conditioning demands (e.g., LIBERO's language-based tasks or a navigation benchmark) where a CLS-token policy matches or beats the patch-token policy, or a real-robot comparison with more than 100 trials showing the patch advantage shrinking to statistical noise.
If this is right
- Robot policies can achieve large gains on precise, multi-object manipulation without fine-tuning the visual encoder, by simply switching from global-pooled features to frozen patch tokens.
- Heavy, billion-parameter VLAs are not necessary for in-domain manipulation learning; a lightweight transformer with frozen dense features matches or exceeds them while running at 11 ms inference latency.
- The quality of the pretrained visual representation remains a primary bottleneck for policy learning, suggesting that continued progress in self-supervised visual representation learning will directly transfer to robot control.
- Spatial compression of visual features, whether by pooling or learned downsampling, degrades control performance, so preserving patch resolution is recommended whenever compute permits.
- The approach is policy-architecture-agnostic, working with both VQ-BeT and Diffusion Policy heads, making it a drop-in replacement for existing transformer-based policies.
Where Pith is reading between the lines
- The 40% relative improvement is task-dependent: on LIBERO Goal, the CLS-token policy actually matches the patch policy (0.95 vs 0.94 for VQ-BeT), and the largest gains appear on spatially challenging multi-object tasks (BlockPush, Cube). The paper's own data suggest the claim is strongest for tasks demanding fine-grained spatial reasoning, not for all embodied control.
- The real-robot results rest on 20 trials per task with no error bars; the reported gap between patch and CLS features (e.g., 0.70 vs 0.60 on cable insertion) is a difference of two successes. A reader should weight these numbers accordingly, while the simulation results with multiple seeds provide a more reliable basis.
- The paper explicitly notes that SigLIP2, a semantically-oriented encoder, underperforms everywhere, while V-JEPA2 does not consistently beat WebSSL or DINOv2. This hints that the spatial-density hypothesis is about the kind of feature geometry, not just the patch format: encoders trained with dense geometric objectives outperform language-aligned ones for manipulation.
- A natural testable extension is whether the same block-causal dense-feature recipe transfers to other control paradigms, such as reinforcement learning or multi-task generalist policies, as the paper's Limitations section identifies RL as an unexplored direction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Patch Policy, a minimal modification to transformer-based behavior-cloning policies: instead of compressing each visual observation into a single global token, the policy consumes all dense patch tokens from a frozen pretrained ViT, using a block-causal attention mask over the flattened spatio-temporal token sequence. The method is evaluated with two policy heads (VQ-BeT and Diffusion Policy) and five frozen visual encoders across four simulated and three real-robot manipulation suites, plus additional zero-shot CAP experiments. The central claims are that dense patch features give a 40% relative improvement over global-pooled representations, that Patch Policy surpasses fine-tuned OpenVLA-OFT by 18% while using roughly 0.7% of the parameters, and that preserving spatial token density—rather than model scale—is the key ingredient.
Significance. If the quantitative claims survive scrutiny, the paper makes a useful contribution: it isolates representation density from model scale, shows that frozen off-the-shelf encoders can be used directly for control, and provides a drop-in architectural pattern compatible with existing transformer policy heads. The within-policy comparison is clean—same encoder, same policy head, only the tokenization changes—and the paper includes helpful ablations of the attention mask, spatial compression, and model size, plus a broad encoder sweep and real-robot evidence. The headline numbers, however, are not currently derivable from the reported tables, so the strength of the contribution is conditional on a revised, fully specified quantitative presentation.
major comments (3)
- [Abstract; Tables 1–3] The headline numbers do not follow from the tables under any stated aggregation. In Table 1, WebSSL Patch VQ-BeT versus WebSSL CLS gives per-task relative changes of +15% (Push-T), -1% (LIBERO Goal), +118% (BlockPush), and +630% (Cube); the ratio of summed raw scores is about +96%. WebSSL Patch DP versus AvgPool gives about +55% by summed score. No common pooling of these heterogeneous metrics (coverage fractions, success rates, mean block counts) yields 40%. For the OpenVLA-OFT comparison, the closest figure to 18% is a DP-only average of per-task relative gains from Table 1 (about 17.4%), while the VQ-BeT average is about 11% and the Table 2 final-stage relative gains are +133%, +42%, and +38%. The 0.7% parameter claim also refers to the DINOv2 ViT-S configuration used in the real-robot experiments, whereas the simulation results use WebSSL (334M total parameters, about 4.4% of OpenVLA
- [Section 3.3; Table 1; Figure 1 caption] The text states that Patch Policy 'consistently outperforms global representations' and 'outperforms the fine-tuned OpenVLA-OFT baseline ... on all four environments.' Table 1 contradicts this: WebSSL Patch VQ-BeT is below OpenVLA-OFT on LIBERO Goal (0.94 vs 0.95), below WebSSL AvgPool VQ-BeT on LIBERO Goal (0.94 vs 0.97), and WebSSL Patch DP is equal to WebSSL CLS DP on LIBERO Goal (0.98 vs 0.99). The effect is clearly task-dependent: gains are large on BlockPush and Cube, while LIBERO Goal shows parity or slight losses. Please replace 'consistently' and 'all four' with a precise statement such as 'matches or exceeds on spatial/multi-object tasks and is comparable on goal-conditioned tasks.'
- [Section 3.3; Table 1 footnote] The OpenVLA-OFT comparison on LIBERO Goal is not evidently controlled. The paper states that all policies in the study receive only visual inputs, but for LIBERO Goal the OpenVLA-OFT number is 'reported directly as provided in the original manuscript [25].' If that baseline used language instructions or a different evaluation protocol, the comparison is not apples-to-apples. Please either specify the exact conditioning and evaluation protocol of the OpenVLA-OFT baseline in this work or rerun it under the same input conditions as the other rows.
minor comments (4)
- [Table 2; Section 3.4] Real-robot results are based on 20 trials per task with no confidence intervals. The claim that the improvement is 'most significant' on Cable Insertion is not statistically supported. Please report binomial confidence intervals or raw trial counts, and avoid the word 'significant' unless a test is performed.
- [Section 3.5; Tables 7–8] The text calls DINOv2 'top-performing' in Section 3.4, but Tables 7 and 8 show WebSSL substantially ahead on BlockPush and Cube. Please qualify the claim as 'top-performing overall' or 'best on average,' and note the task-dependent ranking.
- [Section 3.6; Table 3] The 'as little as 0.7%' parameter claim should be paired with the WebSSL number (4.4%) where the simulation results are described, to avoid the impression that the headline 18% VLA comparison used the 0.7% configuration.
- [Section 3.7; Table 4] The compression ablation is run on Push-T only, with no standard errors or multiple seeds. The text says compression causes a 'significant decrease'; please soften to 'a large decrease in this environment' or add variance and additional tasks.
Circularity Check
No circular derivation found: the central claim is an empirical comparison, and the self-citations are not load-bearing reductions.
full rationale
The paper's central claim is an empirical measurement: policies consuming frozen, dense pretrained patch tokens are compared with policies using global-pooled or CLS features under matched policy heads (VQ-BeT, Diffusion Policy) in Tables 1, 5, 6, 7, and 8, and against OpenVLA-OFT in Tables 1 and 2. No parameter is fitted to the reported outcome and then renamed a prediction; the gains are observed rollouts, not constructed identities. The block-causal attention mask is borrowed from the authors' own DINO-WM paper [20], and CAP [27] is also self-cited, but neither citation forces the empirical conclusion. The paper applies an existing masking scheme to a new setting and measures its effect; it does not invoke a self-cited uniqueness theorem or smuggle in an ansatz that already contains the result. The result is also not definitional: dense patch tokens are not defined in terms of downstream success, and the global-feature baselines are independent architectures. The limitations section (frozen backbones, behavior-cloning-only evaluation) and the task-dependent results (e.g., LIBERO Goal CLS 0.95 vs Patch 0.94) weaken the generality of the claim but are not circularity. The abstract's headline numbers (40%, 18%, 0.7%) are not obviously reproducible from the tables under a stated aggregation, and the parameter-count comparison conflates the WebSSL/ViT-L simulation configuration with the DINOv2/ViT-S real-robot configuration; this is a reporting and reproducibility concern, not a circular reduction of the kind this pass flags. Overall, the derivation chain is self-contained as an empirical study, with no load-bearing step reducing to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (2)
- Policy context window T (block size) =
per-env: 5 (Push-T), 10 (LIBERO), 3 (BlockPush), 5 (Cube); 2 (real-world)
- Policy backbone size (N, n_heads, d_emb) =
e.g., 8/8/512 for many tasks; 6/6/120 for LIBERO VQ-BeT
axioms (4)
- domain assumption Frozen internet-pretrained ViT patch features retain sufficient spatial and semantic detail for closed-loop control; no encoder fine-tuning is needed.
- domain assumption A 1D learned positional embedding over the flattened patch sequence is sufficient to convey the 2D spatial layout of each frame.
- standard math Standard transformer attention and the VQ-BeT/Diffusion Policy losses are taken as given and correctly implemented.
- domain assumption The chosen benchmark tasks (Push-T, LIBERO Goal, BlockPush, Cube, plus three real-robot tasks) are representative of embodied manipulation, and the observed ordering of representations generalizes.
read the original abstract
Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sacrificing both fine-grained spatial detail and the benefits of large-scale visual pre-training. While there exist policies that do operate on dense patch features like large vision-language-action models (VLAs), they tend to be heavy and slow, inheriting the full cost of a billion-parameter vision-language model (VLM) backbone. We close this gap with Patch Policy, a minimal architectural extension that enables transformer-based policies to consume dense pre-trained patch tokens directly without the computational overhead of a full VLM. At its core is a block-causal attention mask that preserves the temporal causality of standard policies while letting the model attend over many patch tokens per observation, alongside other state information. Patch Policy is lightweight, fast, and highly effective. Across four simulated and three real-world environment suites, our method achieves a 40% relative improvement over policies using state-of-the-art global-pooled representations. Furthermore, it surpasses fine-tuned OpenVLA-OFT by 18% while using roughly 0.7% of the parameters. We believe Patch Policy provides a pipeline for the robotics community to readily leverage continuing progress in visual representation learning, without sacrificing the training efficiency or inference speed required for high-frequency, reactive control. Videos can be viewed at https://patch-policy.github.io
Figures
Reference graph
Works this paper leans on
-
[1]
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale.ArXiv, abs/2010.11929, 2020. URL https://api.semanticscholar.org/CorpusID:225039882
Pith/arXiv arXiv 2010
-
[2]
K. He, X. Chen, S. Xie, Y . Li, P. Doll’ar, and R. B. Girshick. Masked autoencoders are scalable vision learners.2022 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 15979–15988, 2021. URLhttps://api.semanticscholar.org/ CorpusID:243985980
2022
-
[3]
Caron, H
M. Caron, H. Touvron, I. Misra, H. J’egou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9630–9640, 2021. URLhttps://api.semanticscholar. org/CorpusID:233444273
2021
-
[4]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. InInternational Conference on Machine Learning, 2021. URL https://api.semanticscholar.org/CorpusID:231591445
2021
-
[5]
X. Chen, S. Xie, and K. He. An empirical study of training self-supervised vision transform- ers.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9620–9629,
2021
-
[6]
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . B. Huang, S.-W. Li, I. Misra, M. G. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. J´egou, J. Mairal, 10 P. Labatut, A. Joulin, and P. Bojanowski. Dinov2: Learning robust visual features...
Pith/arXiv arXiv 2023
-
[7]
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierar- chical vision transformer using shifted windows.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9992–10002, 2021. URLhttps://api.semanticscholar. org/CorpusID:232352874
2021
-
[8]
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European conference on computer vision, pages 38–55. Springer, 2024
2024
-
[9]
M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. M. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y . Xia, B. Mustafa, O. H’enaff, J. Harmsen, A. Steiner, and X.-Q. Zhai. Siglip 2: Multilingual vision-language encoders with improved semantic understand- ing, localization, and dense features.ArXiv, abs/2502.14786, 2025. URLhttps://api. semantics...
Pith/arXiv arXiv 2025
-
[10]
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Sur ´ıs, C. K. Ryali, K. V . Al- wala, H. Khedr, A. Huang, J. Lei, T. Ma, B. Guo, A. Kalla, M. Marks, J. Greer, M. Wang, P. Sun, R. Radle, T. Afouras, E. Mavroudi, K. Xu, T.-H. Wu, Y . Zhou, L. Momeni, R. Hazra, S. Ding, S. Vaze, F. Porcher, F. Li, S. Li, A. Kamath, H. K. Cheng, P. Doll’ar, N. Ravi...
Pith/arXiv arXiv 2025
-
[11]
Sim’eoni, H
O. Sim’eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Cou- prie, J. Mairal, H. J’egou, P. Labatut, and P. Bojanowski. Dinov3. 2025. URLhttps: //api.semantic...
2025
-
[12]
D. Fan, S. Tong, J. Zhu, K. Sinha, Z. Liu, X. Chen, M. Rabbat, N. Ballas, Y . LeCun, A. Bar, and S. Xie. Scaling language-free visual representation learning.ArXiv, abs/2504.01017, 2025. URLhttps://api.semanticscholar.org/CorpusID:277467353
Pith/arXiv arXiv 2025
-
[13]
A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y . LeCun, M. Assran, and N. Bal- las. Revisiting feature prediction for learning visual representations from video, 2024. URL https://arxiv.org/abs/2404.08471
Pith/arXiv arXiv 2024
-
[14]
M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, et al. V-jepa 2: Self-supervised video models enable understanding, prediction and planning.arXiv preprint arXiv:2506.09985, 2025
Pith/arXiv arXiv 2025
-
[15]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. InPro- ceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016
2016
-
[16]
Perez, F
E. Perez, F. Strub, H. De Vries, V . Dumoulin, and A. Courville. Film: Visual reasoning with a general conditioning layer. InProceedings of the AAAI conference on artificial intelligence, volume 32, 2018
2018
-
[17]
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022
2022
-
[18]
Dasari, M
S. Dasari, M. K. Srirama, U. Jain, and A. Gupta. An unbiased look at datasets for visuo-motor pre-training. InConference on Robot Learning. PMLR, 2023. 11
2023
-
[19]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polo- sukhin. Attention is all you need. InNeural Information Processing Systems, 2017. URL https://api.semanticscholar.org/CorpusID:13756489
2017
-
[20]
G. Zhou, H. Pan, Y . LeCun, and L. Pinto. DINO-WM: world models on pre-trained vi- sual features enable zero-shot planning. In A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu, editors,Forty-second International Con- ference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267...
2025
-
[21]
S. Lee, Y . Wang, H. Etukuru, H. J. Kim, N. Muhammad, M. Shafiullah, and L. Pinto. Be- havior generation with latent actions.ArXiv, abs/2403.03181, 2024. URLhttps://api. semanticscholar.org/CorpusID:268248763
Pith/arXiv arXiv 2024
-
[22]
C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. doi:10.15607/RSS.2023.XIX.026. URLhttps://doi.org/10.15607/RSS. 2023.XIX.026
-
[23]
Z. J. Cui, H. Pan, A. Iyer, S. Haldar, and L. Pinto. Dynamo: In-domain dynamics pre- training for visuo-motor control.ArXiv, abs/2409.12192, 2024. URLhttps://api. semanticscholar.org/CorpusID:272709175
Pith/arXiv arXiv 2024
-
[24]
T. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware.ArXiv, abs/2304.13705, 2023. URLhttps://api.semanticscholar. org/CorpusID:258331658
Pith/arXiv arXiv 2023
-
[25]
M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success.ArXiv, abs/2502.19645, 2025. URLhttps://api.semanticscholar. org/CorpusID:276647709
Pith/arXiv arXiv 2025
-
[26]
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre- training. InProceedings of the IEEE/CVF international conference on computer vision, pages 11975–11986, 2023
2023
-
[27]
Z. J. Cui, O. Rayyan, H. Etukuru, B. Tan, Z. Andrianarivo, Z. Teng, Y . Zhou, K. Mehta, N. Wojno, K. Y . Wu, M. H. Anjaria, Z. Wu, M. Mao, G. Zhang, B. Shah, Y . Kim, S. Chintala, L. Pinto, and N. M. M. Shafiullah. Contact-anchored policies: Contact conditioning creates strong robot utility models.arXiv preprint arXiv:2602.09017, 2026
arXiv 2026
-
[28]
R. G. Goswami, A. Bar, D. Fan, T.-Y . Yang, G. Zhou, P. Krishnamurthy, M. Rabbat, F. Khor- rami, and Y . LeCun. World models can leverage human videos for dexterous manipulation. ArXiv, abs/2512.13644, 2025. URLhttps://api.semanticscholar.org/CorpusID: 283896258
arXiv 2025
-
[29]
D. A. Pomerleau. Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988
1988
-
[30]
N. M. M. Shafiullah, Z. J. Cui, A. Altanzaya, and L. Pinto. Behavior transformers: Cloning k modes with one stone.ArXiv, abs/2206.11251, 2022. URLhttps://api. semanticscholar.org/CorpusID:249926747
Pith/arXiv arXiv 2022
-
[31]
Z. J. Cui, Y . Wang, N. M. M. Shafiullah, and L. Pinto. From play to policy: Conditional behavior generation from uncurated robot data.ArXiv, abs/2210.10047, 2022. URLhttps: //api.semanticscholar.org/CorpusID:252968170. 12
Pith/arXiv arXiv 2022
-
[32]
Florence, C
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mor- datch, and J. Tompson. Implicit behavioral cloning. InConference on robot learning, pages 158–168. PMLR, 2022
2022
-
[33]
K. Rana, R. Lee, D. Pershouse, and N. Suenderhauf. Imle policy: Fast and sample effi- cient visuomotor policy learning via implicit maximum likelihood estimation.arXiv preprint arXiv:2502.12371, 2025
Pith/arXiv arXiv 2025
-
[34]
O’Neill, A
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024
2024
-
[35]
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024
Pith/arXiv arXiv 2024
-
[36]
K. He, H. Fan, Y . Wu, S. Xie, and R. B. Girshick. Momentum contrast for unsupervised visual representation learning.2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9726–9735, 2019. URLhttps://api.semanticscholar.org/ CorpusID:207930212
2020
-
[37]
J.-B. Grill, F. Strub, F. Altch’e, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B.´A. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko. Bootstrap your own latent: A new approach to self-supervised learning.ArXiv, abs/2006.07733, 2020. URLhttps://api.semanticscholar.org/CorpusID:219687798
Pith/arXiv arXiv 2006
-
[38]
Caron, H
M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF interna- tional conference on computer vision, pages 9650–9660, 2021
2021
-
[39]
S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta. R3M: A universal visual rep- resentation for robot manipulation. In K. Liu, D. Kulic, and J. Ichnowski, editors,Confer- ence on Robot Learning, CoRL 2022, 14-18 December 2022, Auckland, New Zealand, vol- ume 205 ofProceedings of Machine Learning Research, pages 892–909. PMLR, 2022. URL https://proc...
2022
-
[40]
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, K. Choromanski, T. Ding, D. Driess, K. A. Dubey, C. Finn, P. R. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y . Kuang, I. Leal, S. Levine, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, ...
-
[41]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InNorth American Chapter of the Association for Computational Linguistics, 2019. URLhttps://api.semanticscholar.org/CorpusID: 52967399
2019
-
[42]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
1901
-
[43]
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira. Perceiver: General perception with iterative attention.ArXiv, abs/2103.03206, 2021. URLhttps: //api.semanticscholar.org/CorpusID:232110866. 13
Pith/arXiv arXiv 2021
-
[44]
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch. Decision transformer: Reinforcement learning via sequence modeling. InNeu- ral Information Processing Systems, 2021. URLhttps://api.semanticscholar.org/ CorpusID:235294299
2021
-
[45]
Janner, Q
M. Janner, Q. Li, and S. Levine. Offline reinforcement learning as one big sequence mod- eling problem. InNeural Information Processing Systems, 2021. URLhttps://api. semanticscholar.org/CorpusID:235313679
2021
-
[46]
Y . Su, X. Zhan, H. Fang, H. Xue, H.-S. Fang, Y .-L. Li, C. Lu, and L. Yang. Dense policy: Bidirectional autoregressive learning of actions. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 14486–14495, 2025
2025
-
[47]
S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learn- ing.ArXiv, abs/2406.07539, 2024. URLhttps://api.semanticscholar.org/CorpusID: 270379931
Pith/arXiv arXiv 2024
-
[48]
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Perts...
Pith/arXiv arXiv 2022
-
[49]
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gim´enez, Y . Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. D. Edwards, N. M. O. Heess, Y . Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas. A generalist agent. ArXiv, abs/2205.06175, 2022. URLhttps://api.semanticscholar.org/CorpusID: 248722148
Pith/arXiv arXiv 2022
-
[50]
Z. Hou, T. Zhang, Y . Xiong, H. Pu, C. Zhao, R. Tong, Y . Qiao, J. Dai, and Y . Chen. Diffusion transformer policy.arXiv preprint arXiv:2410.15959, 2024
Pith/arXiv arXiv 2024
-
[51]
O. M. Team, D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, P. R. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy.ArXiv, abs/2405.12213, 2024. URL https://api.semanticscholar.org/CorpusID:266379116
Pith/arXiv arXiv 2024
-
[52]
Goyal, J
A. Goyal, J. Xu, Y . Guo, V . Blukis, Y .-W. Chao, and D. Fox. Rvt: Robotic view transformer for 3d object manipulation. InConference on Robot Learning, pages 694–710. PMLR, 2023
2023
-
[53]
Shridhar, L
M. Shridhar, L. Manuelli, and D. Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. InConference on Robot Learning, pages 785–799. PMLR, 2023
2023
-
[54]
Driess, F
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tomp- son, Q. H. Vuong, T. Yu, W. Huang, Y . Chebotar, P. Sermanet, D. Duckworth, S. Levine, V . Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. R. Florence. Palm-e: An embodied multimodal language model. InInternational Conference on Machine L...
2023
-
[55]
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky.π0: A vision-language-action flow model for general robot control.ArXiv, abs/2410.24164, 2024...
Pith/arXiv arXiv 2024
-
[56]
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. V...
Pith/arXiv arXiv 2025
-
[57]
P. Intelligence, A. A. Amin, R. J. Aniceto, A. Balakrishna, K. Black, K. Conley, G. Connors, J. Darpinian, K. Dhabalia, J. DiCarlo, D. Driess, M. Equi, A. Esmail, Y . Fang, C. Finn, C. Glos- sop, T. Godden, I. Goryachev, L. Groom, H. Hancock, K. Hausman, G. Hussein, B. Ichter, S. Jakubczak, R. Jen, T. Jones, B. Katz, L. Ke, C. Kuchi, M. Lamb, D. LeBlanc, ...
Pith/arXiv arXiv 2025
-
[58]
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. R ´e. Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35: 16344–16359, 2022
2022
-
[59]
B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023
2023
-
[60]
S. Park, K. Frans, B. Eysenbach, and S. Levine. Ogbench: Benchmarking offline goal- conditioned rl.ArXiv, abs/2410.20092, 2024. URLhttps://api.semanticscholar.org/ CorpusID:273654871
Pith/arXiv arXiv 2024
-
[61]
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kir- mani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao. Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024. 15 A Appendix A.1 Implementation and Baselines Our implementation and baselines are built upon the ...
Pith/arXiv arXiv 2024
-
[64]
DynaMo:https://github.com/jeffacce/dynamo_ssl
-
[65]
VQ-BeT:https://github.com/jayLEE0301/vq_bet_official
-
[66]
Diffusion Policy:https://github.com/real-stanford/diffusion_policy
-
[67]
OpenVLA-OFT:https://github.com/moojink/openvla-oft
-
[68]
The real-world tasks comprise: (1)Tool Hanging, (2)Pen Collection (long-horizon, small objects), (3)Cable Insertion(low-tolerance insertion)
ACT:https://github.com/tonyzhaozh/act A.2 Environments and Tasks We evaluate PATCHPOLICYacross four simulated environments (Push-T, LIBERO Goal, Block- Push, Cube) with 2D-to-7D action spaces, and three real-world tasks using a 7-DoF Franka arm with a parallel-jaw gripper. The real-world tasks comprise: (1)Tool Hanging, (2)Pen Collection (long-horizon, sm...
-
[2021]
URLhttps://api.semanticscholar.org/CorpusID:233024948
-
[2023]
URLhttps://api.semanticscholar.org/CorpusID:260293142
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.