REVIEW 3 major objections 6 minor 67 references
Rethinking the Intermediate Features in Adversarial Attacks: Misleading Robotic Models via Adversarial Distillation
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A fixed adversarial text prefix, optimized against continuous action features and intermediate self-attention features, makes a language-conditioned robot policy fail at the commanded task; across 13 VIMA-Bench tasks the average attack…
desk verdict A plausible new attack idea for language-conditioned robot policies, but the ablation table contradicts the headline claim at the 25-token setting and the evaluation lacks clean baselines; worth peer review after the authors fix the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the adversarial-distillation objective, a composite loss over two intermediate feature representations. The first term, $L_{\mathrm{continuous}} = -\lVert D_c(p_a\oplus p,h) - D_c(p,h)\rVert_2$, maximizes the distance between continuous action vectors produced by the controller decoder with and without the adversarial prefix, bypassing the robust discretized action decoder. The second term, $L_{\mathrm{self\text{-}attn}} = -\lVert F_s(p_a\oplus p,h) - F_s(p,h)\rVert_2$, maximizes the distance between self-attention feature maps across intermediate layers, which the paper argues are particularly susceptible to perturbation because attention weights propagate small input changes through the whole sequence. The full loss is $L = \alpha L_{\mathrm{continuous}} + \beta L_{\mathrm{self\text{-}attn}}$, and the prefix tokens are optimized with the Greedy Coordinate Gradient algorithm.
What would settle it
Re-run the evaluation on the same 13 VIMA tasks and report the clean (no-prefix) task success rate alongside the random-prefix and adversarial-prefix attack success rates; if the clean success rate is close to the adversarial ASR on several tasks, the central claim of superior attack performance is not supported.
Extended reading notes
Core claim
The paper's central discovery is that attacking the right intermediate representations, rather than the final output distribution, is what makes adversarial text work against robotic policies. Concretely, the optimization loss is the negative alignment between (i) the controller decoder's continuous action vector with and without the adversarial prefix and (ii) the self-attention feature maps of all intermediate layers; optimizing this composite loss with a coordinate-gradient search yields a universal prefix. On the VIMA model across 13 Level 1 tasks, the method reports an average attack success rate of 47.08%, compared with 39.59% for M GCG, 35.26% for GCG, 20.79% for random prefixes, and 18.96% for gradient descent. The authors further report that a prefix optimized on a single Visual Manipulation demonstration transfers to the other 12 task types and, in a gray-box setting, attacks a smaller 92M-parameter VIMA variant.
Load-bearing premise
The evaluation assumes that a failed task under an adversarial prompt is caused by the prompt; because no attack-free success rate is reported, tasks with high natural failure rates could inflate the reported attack success rate.
Editorial extensions
If this is right
- A fixed 25-token prefix trained on one Visual Manipulation demonstration transfers to 12 other task types, so an attacker needs no per-task optimization to disrupt a deployed robot.
- Attack success rises with prefix length: the method reports 33.83% average ASR with 10 tokens and 68.75% with 48 tokens, while baselines plateau or falter as token count grows.
- Attacking continuous action features before discretization is the key improvement: the ablation shows ASR rises from 34% to 51% at 25 tokens when the loss moves from discrete outputs to continuous action features.
- Adding self-attention feature misalignment on top of continuous features gives a 16% ASR gain at 48 tokens, larger than the 3% gain from adding cross-attention features, supporting the paper's emphasis on self-attention as the vulnerable intermediate representation.
- Prefixes optimized on a 200M-parameter model also attack a 92M-parameter variant, with 52.2% ASR at only 10 tokens, indicating potential transferability across model sizes.
Reading between the lines
- The paper does not report the clean, attack-free task success rate, so part of the reported ASR on high-random-failure tasks such as Twist (84.67% under a random prefix) may reflect task difficulty rather than the adversarial prefix; a fair comparison would subtract natural failure rates.
- The same adversarial-distillation recipe may apply to other multimodal policies with discretized outputs, such as diffusion-based visuomotor policies or action-quantized reinforcement learning agents, wherever a continuous intermediate action representation exists before discretization.
- The attack objective suggests a corresponding defense: training a policy to keep self-attention features aligned under input perturbation may directly counteract the vulnerability the paper exploits, turning adversarial distillation into a robustness regularizer.
- A testable extension is whether feature-based prefix optimization yields systematically better cross-model transfer than output-based optimization; the paper's gray-box result hints at this, but the mechanism is not isolated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a white-box adversarial prompt attack against language-conditioned robotic manipulation models, specifically VIMA. The attack optimizes a universal adversarial prefix with a GCG-style discrete search, using a weighted loss over continuous action features from the controller decoder and intermediate self-attention features, with the stated goal of bypassing the robustness conferred by the action discretization module. Experiments on 13 VIMA-Bench Level 1 tasks report an average attack success rate of 47.08% with 25-token prefixes, compared with 39.59% for M GCG and 35.26% for GCG, plus transfer experiments to a 92M VIMA variant. The paper also reports ablations over token count and loss components, and concludes that intermediate feature misalignment further enhances attack efficacy.
Significance. Adversarial attacks on language-conditioned robot policies are understudied, and the paper identifies a plausible and practically relevant mechanism: optimizing against continuous action representations instead of the final discrete output distribution improves attack success. If the reported numbers are not confounded, the universal-prefix attack and its cross-model transferability would be a useful contribution to the robotics-security literature. The paper is also honest about its simulation-only evaluation and the one-time offline cost of optimization. However, the missing attack-free task-success baseline, the internal inconsistency in the ablation table, and the absence of variance reporting currently prevent the empirical claims from being accepted as stated.
major comments (3)
- [IV-A2, Table I] Attack success is defined as the model failing to complete the task, yet no clean (attack-free) task success rate is reported for any of the 13 tasks. Because Random already achieves ASRs as high as 84.67% on Twist and 65.11% on Rearrange then Restore, a high ASR does not by itself demonstrate that the adversarial prefix caused the failure. Please report clean success rates per task and, preferably, attack-induced degradation relative to clean performance; otherwise the headline 47.08% average and the claimed superiority over GCG/M GCG are confounded by natural task failures.
- [IV-D, Table II vs IV-B] The central claim that self-attention feature misalignment further enhances attack efficacy is contradicted at the operating point used for the headline result. Main results use 25 adversarial tokens (Section IV-B), where Table II reports L_continuous = 0.51 and L_continuous + L_self-attn = 0.47, i.e., adding the self-attention loss lowers ASR. Self-attention only helps at 10 tokens (0.29 to 0.34) and at 48 tokens (0.53 to 0.69), neither of which is the setting for the main comparison. Please clarify the table or the operating point; as printed, the conclusion in Section V that intermediate feature misalignment further enhances attack efficacy is not supported by the reported ablation.
- [IV-A2, Table I] The average ASR comparison (47.08% vs 39.59%) is reported without error bars, confidence intervals, or significance tests, even though the manuscript states that results come from three random seeds. Given that some task-level differences are very small (e.g., Twist: Ours 85.56 vs M GCG 85.33; Pick in Order then Restore: Ours 46.44 vs M GCG 47.33), the 7.49-point average gap may be within seed-to-seed noise. Please report per-seed results or standard deviations and test the significance of the average difference.
minor comments (6)
- [III-C and IV-A4] Equations (4) and (5) define losses as negative L2 norms, but Section IV-A4 states that cosine similarity is used to calculate the loss between features; please align the mathematical definitions with the actual implementation.
- [Equation (1)] There is a typo in Equation (1): "grounth-truth" should be "ground-truth".
- [References] The reference list contains uncited template references [1] through [10] and a duplicate of the Vaswani et al. paper ([11] and [24]); please clean up the bibliography.
- [IV-B and IV-D] When reporting improvements such as "a substantial 28% improvement" and "a 17% improvement," please specify whether the numbers are absolute percentage points or relative improvements, since the baseline values make the two interpretations differ substantially.
- [IV-E, Figure 5] The gray-box transferability claim would also benefit from clean task-success rates on the 92M model, and from variance reporting, since the comparison to the white-box number appears to be a single point estimate.
- [Figures 4 and 5] Please describe in the captions whether the plotted curves are averages over the three seeds and whether error bars or shading are shown; currently no uncertainty information is visible.
Circularity Check
No circularity: the attack is an optimization objective evaluated against its own target, not a derivation of a prediction.
full rationale
The paper's claimed contribution is an adversarial-prefix optimization. There is no predictive or first-principles derivation whose output is equivalent to an input. The losses Lcontinuous (Eq. 4) and Lself-attn (Eq. 5) are optimization objectives defined from the target model's own continuous features and self-attention features; optimizing them and measuring attack success rate on VIMA tasks is the intended attack construction, not a fitted parameter renamed as a prediction. The GCG optimizer is taken from an external prior work ([16]); no load-bearing result is justified only by a self-citation, and the paper contains no self-citations of the authors. Two non-circular weaknesses exist: (1) no clean task success rate is reported, so high ASR on tasks like Twist (random prefix 84.67%, Table I) may partly reflect natural failures; and (2) at the 25-token setting used for the headline comparison, the ablation in Table II reports Lcontinuous alone at 0.51 and Lcontinuous + Lself-attn at 0.47, so the claimed benefit of the self-attention term is not supported at the operating point of Table I. These are internal-consistency and evaluation-baseline concerns, not circularity: they do not make any derivation equivalent to its own inputs. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- alpha (weight of L_continuous) =
1
- beta (weight of L_self-attn) =
20
assumptions (5)
- domain assumption White-box access to model gradients and intermediate features
- domain assumption Action discretization via Da is a non-differentiable bottleneck that reduces attack transfer
- domain assumption Self-attention features are more manipulable than cross-attention features
- standard math GCG discrete optimization converges to effective adversarial prefixes
- domain assumption Cosine similarity is an appropriate metric for feature misalignment
Cite this review
Pith. "Pith review of Rethinking the Intermediate Features in Adversarial Attacks: Misleading Robotic Models via Adversarial Distillation." pith.science (2026). https://pith.science/paper/3WX4WDRE
@misc{pith2026241115222,
author = {Pith},
title = {Pith review of: Rethinking the Intermediate Features in Adversarial Attacks: Misleading Robotic Models via Adversarial Distillation},
year = {2026},
howpublished = {\url{https://pith.science/paper/3WX4WDRE}},
note = {Machine review of arXiv:2411.15222}
}
read the original abstract
Language-conditioned robotic learning has significantly enhanced robot adaptability by enabling a single model to execute diverse tasks in response to verbal commands. Despite these advancements, security vulnerabilities within this domain remain largely unexplored. This paper addresses this gap by proposing a novel adversarial prompt attack tailored to language-conditioned robotic models. Our approach involves crafting a universal adversarial prefix that induces the model to perform unintended actions when added to any original prompt. We demonstrate that existing adversarial techniques exhibit limited effectiveness when directly transferred to the robotic domain due to the inherent robustness of discretized robotic action spaces. To overcome this challenge, we propose to optimize adversarial prefixes based on continuous action representations, circumventing the discretization process. Additionally, we identify the beneficial impact of intermediate features on adversarial attacks and leverage the negative gradient of intermediate self-attention features to further enhance attack efficacy. Extensive experiments on VIMA models across 13 robot manipulation tasks validate the superiority of our method over existing approaches and demonstrate its transferability across different model variants.
Figures
Reference graph
Works this paper leans on
-
[1]
R. Engelmore and A. Morgan, Eds., Blackboard Systems . Reading, Mass.: Addison-Wesley, 1986
work page 1986
-
[12]
Pluto: The ’other’ red planet,
NASA, “Pluto: The ’other’ red planet,” https://www.nasa.gov/nh/ pluto-the-other-red-planet, 2015, accessed: 2018-12-06. IEEE TRANSACTIONS ON ROBOTICS 8
work page 2015
-
[2]
W. J. Clancey, “Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education,” in Proceedings of the Eighth International Joint Conference on Artificial Intelligence (IJCAI-83). Menlo Park, Calif: IJCAI Organization, 1983, pp. 556–560
work page 1983
-
[3]
Classification Problem Solving,
W. J. Clancey, “Classification Problem Solving,” in Proceedings of the Fourth National Conference on Artificial Intelligence. Menlo Park, Calif.: AAAI Press, 1984, pp. 45–54
work page 1984
-
[4]
New ways to make microcircuits smaller,
A. L. Robinson, “New ways to make microcircuits smaller,” Science, vol. 208, no. 4447, pp. 1019–1022, 1980. [Online]. Available: https://science. sciencemag.org/content/208/4447/1019
work page 1980
-
[5]
New Ways to Make Microcircuits Smaller—Duplicate Entry,
A. L. Robinson, “New Ways to Make Microcircuits Smaller—Duplicate Entry,” Science, vol. 208, pp. 1019–1026, 1980
work page 1980
-
[6]
Strategic explanations for a diagnostic consultation system,
D. W. Hasling, W. J. Clancey, and G. Rennels, “Strategic explanations for a diagnostic consultation system,” International Journal of Man-Machine Studies, vol. 20, no. 1, pp. 3–19, 1984. [Online]. Available: https://www. sciencedirect.com/science/article/pii/S0020737384800036
work page 1984
-
[7]
Strategic Explanations in Consultation—Duplicate,
D. W. Hasling, W. J. Clancey, G. R. Rennels, and T. Test, “Strategic Explanations in Consultation—Duplicate,” The International Journal of Man-Machine Studies, vol. 20, no. 1, pp. 3–19, 1983
work page 1983
Show all 67 references
-
[8]
Poligon: A System for Parallel Problem Solving,
J. Rice, “Poligon: A System for Parallel Problem Solving,” Dept. of Computer Science, Stanford Univ., Technical Report KSL-86-19, 1986
1986
-
[9]
Transfer of Rule-Based Expertise through a Tutorial Dialogue,
W. J. Clancey, “Transfer of Rule-Based Expertise through a Tutorial Dialogue,” Ph.D. diss., Dept. of Computer Science, Stanford Univ., Stanford, Calif., 1979
1979
-
[10]
The Engineering of Qualitative Models,
W. J. Clancey, “The Engineering of Qualitative Models,” 2021, forth- coming
2021
-
[11]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” 2017
2017
-
[13]
Vima: General robot manipula- tion with multimodal prompts,
Y . Jiang, A. Gupta, Z. Zhang, G. Wang, Y . Dou, Y . Chen, L. Fei-Fei, A. Anandkumar, Y . Zhu, and L. Fan, “Vima: General robot manipula- tion with multimodal prompts,” in Fortieth International Conference on Machine Learning, 2023
2023
-
[14]
Diffusion policy attacker: Craft- ing adversarial attacks for diffusion-based policies,
Y . Chen, H. Xue, and Y . Chen, “Diffusion policy attacker: Craft- ing adversarial attacks for diffusion-based policies,” arXiv preprint arXiv:2405.19424, 2024
2024 arXiv
-
[15]
Diffusion policy: Visuomotor policy learning via action diffusion,
C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” arXiv preprint arXiv:2303.04137, 2023
2023 arXiv
-
[16]
Universal and trans- ferable adversarial attacks on aligned language models,
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and trans- ferable adversarial attacks on aligned language models,” arXiv preprint arXiv:2307.15043, 2023
2023 arXiv
-
[17]
Explaining and harnessing adversarial examples,
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” arXiv preprint arXiv:1412.6572 , 2014
2014 arXiv
-
[18]
Towards deep learning models resistant to adversarial attacks,
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in International Conference on Learning Representations , 2018
2018
-
[19]
Adversarial examples are not easily detected: Bypassing ten detection methods,
N. Carlini and D. Wagner, “Adversarial examples are not easily detected: Bypassing ten detection methods,” in Proceedings of the 10th ACM workshop on artificial intelligence and security , 2017, pp. 3–14
2017
-
[20]
Jailbreaking black box large language models in twenty queries,
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, “Jailbreaking black box large language models in twenty queries,” arXiv preprint arXiv:2310.08419, 2023
2023 arXiv
-
[21]
Jailbreaking lead- ing safety-aligned llms with simple adaptive attacks,
M. Andriushchenko, F. Croce, and N. Flammarion, “Jailbreaking lead- ing safety-aligned llms with simple adaptive attacks,” arXiv preprint arXiv:2404.02151, 2024
2024 arXiv
-
[22]
Guiding multi-step rearrangement tasks with natural language instructions,
E. Stengel-Eskin, A. Hundt, Z. He, A. Murali, N. Gopalan, M. Gombo- lay, and G. Hager, “Guiding multi-step rearrangement tasks with natural language instructions,” in Conference on Robot Learning. PMLR, 2022, pp. 1486–1501
2022
-
[23]
Language conditioned imitation learning over unstructured data,
C. Lynch and P. Sermanet, “Language conditioned imitation learning over unstructured data,” Robotics: Science and Systems , 2021
2021
-
[24]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[25]
Safe learning in robotics: From learning-based control to safe reinforcement learning,
L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig, “Safe learning in robotics: From learning-based control to safe reinforcement learning,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 5, no. 1, pp. 411–444, 2022
2022
-
[26]
Rt-1: Robotics transformer for real-world control at scale,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu et al. , “Rt-1: Robotics transformer for real-world control at scale,” arXiv preprint arXiv:2212.06817, 2022
2022 arXiv
-
[27]
Adaptive dis- cretization for model-based reinforcement learning,
S. Sinclair, T. Wang, G. Jain, S. Banerjee, and C. Yu, “Adaptive dis- cretization for model-based reinforcement learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 3858–3871, 2020
2020
-
[28]
Action- quantized offline reinforcement learning for robotic skill learning,
J. Luo, P. Dong, J. Wu, A. Kumar, X. Geng, and S. Levine, “Action- quantized offline reinforcement learning for robotic skill learning,” in Conference on Robot Learning . PMLR, 2023, pp. 1348–1361
2023
-
[29]
Dota 2 with large scale deep reinforcement learning,
C. Berner, G. Brockman, B. Chan, V . Cheung, P. D ´ebiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse et al., “Dota 2 with large scale deep reinforcement learning,” arXiv preprint arXiv:1912.06680 , 2019
1912 arXiv
-
[30]
Bc-z: Zero-shot task generalization with robotic imitation learning,
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” in Conference on Robot Learning. PMLR, 2022, pp. 991–1002
2022
-
[31]
Language-conditioned imitation learning for robot ma- nipulation tasks,
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor, “Language-conditioned imitation learning for robot ma- nipulation tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 13 139–13 150, 2020
2020
-
[32]
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation,
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel, “Deep imitation learning for complex manipulation tasks from virtual reality teleoperation,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 5628–5635
2018
-
[33]
Energy-based imitation learning,
M. Liu, T. He, M. Xu, and W. Zhang, “Energy-based imitation learning,” arXiv preprint arXiv:2004.09395 , 2020
2004 arXiv
-
[34]
Intriguing properties of neural networks,
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. J. Goodfellow, and R. Fergus, “Intriguing properties of neural networks,” in 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings , Y...
2014
-
[35]
Attacking large language models with projected gradient descent,
S. Geisler, T. Wollschl ¨ager, M. Abdalla, J. Gasteiger, and S. G¨unnemann, “Attacking large language models with projected gradient descent,” arXiv preprint arXiv:2402.09154, 2024
2024 arXiv
-
[36]
Amplegcg: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed llms,
Z. Liao and H. Sun, “Amplegcg: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed llms,” arXiv preprint arXiv:2404.07921 , 2024
2024 arXiv
-
[37]
Adversarial example does good: Preventing painting imi- tation from diffusion models via adversarial examples,
C. Liang, X. Wu, Y . Hua, J. Zhang, Y . Xue, T. Song, Z. Xue, R. Ma, and H. Guan, “Adversarial example does good: Preventing painting imi- tation from diffusion models via adversarial examples,” in International Conference on Machine Learning . PMLR, 2023, pp. 20 763–20 786
2023
-
[38]
On the adversarial robustness of multi- modal foundation models,
C. Schlarmann and M. Hein, “On the adversarial robustness of multi- modal foundation models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3677–3685
2023
-
[39]
Attacking deep reinforcement learning with decoupled adversarial policy,
K. Mo, W. Tang, J. Li, and X. Yuan, “Attacking deep reinforcement learning with decoupled adversarial policy,” IEEE Transactions on De- pendable and Secure Computing , vol. 20, no. 1, pp. 758–768, 2022
2022
-
[40]
Revisiting the adversarial robustness-accuracy tradeoff in robot learning,
M. Lechner, A. Amini, D. Rus, and T. A. Henzinger, “Revisiting the adversarial robustness-accuracy tradeoff in robot learning,”IEEE Robotics and Automation Letters , vol. 8, no. 3, pp. 1595–1602, 2023
2023
-
[41]
Studying adversarial attacks on behavioral cloning dynamics,
G. Hall, A. Das, J. Quarles, and P. Rad, “Studying adversarial attacks on behavioral cloning dynamics,” in 2020 IEEE 32nd International Conference on Tools with Artificial Intelligence (ICTAI) . IEEE, 2020, pp. 452–459
2020
-
[42]
Is deep learning safe for robot vision? adversarial examples against the icub humanoid,
M. Melis, A. Demontis, B. Biggio, G. Brown, G. Fumera, and F. Roli, “Is deep learning safe for robot vision? adversarial examples against the icub humanoid,” in Proceedings of the IEEE international conference on computer vision workshops , 2017, pp. 751–759
2017
-
[43]
Analyzing adversarial attacks against deep learning for robot navigation
M. I. Khedher and M. Rezzoug, “Analyzing adversarial attacks against deep learning for robot navigation.” in ICAART (2), 2021, pp. 1114–1121
2021
-
[44]
Video pretraining (vpt): Learning to act by watching unlabeled online videos,
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune, “Video pretraining (vpt): Learning to act by watching unlabeled online videos,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 639–24 654, 2022
2022
-
[45]
Is bert really robust? a strong baseline for natural language attack on text classification and entailment,
D. Jin, Z. Jin, J. T. Zhou, and P. Szolovits, “Is bert really robust? a strong baseline for natural language attack on text classification and entailment,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 05, 2020, pp. 8018–8025
2020
-
[46]
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery,
Y . Wen, N. Jain, J. Kirchenbauer, M. Goldblum, J. Geiping, and T. Gold- stein, “Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery,”Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[47]
Automatically auditing large language models via discrete optimization,
E. Jones, A. Dragan, A. Raghunathan, and J. Steinhardt, “Automatically auditing large language models via discrete optimization,” inInternational Conference on Machine Learning . PMLR, 2023, pp. 15 307–15 329
2023
-
[48]
Character-level white-box adversarial attacks against transformers via attachable subwords substitution,
A. Liu, H. Yu, X. Hu, L. Lin, F. Ma, Y . Yang, L. Wen et al. , “Character-level white-box adversarial attacks against transformers via attachable subwords substitution,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 7664– 7676
2022
-
[49]
Generating natural language adversarial examples,
M. Alzantot, Y . Sharma, A. Elgohary, B.-J. Ho, M. Srivastava, and K.-W. Chang, “Generating natural language adversarial examples,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2018
2018
-
[50]
Fitnets: Hints for thin deep nets,
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y . Ben- gio, “Fitnets: Hints for thin deep nets,” arXiv preprint arXiv:1412.6550 , 2014
2014 arXiv
-
[51]
Knowledge transfer via distilla- tion of activation boundaries formed by hidden neurons,
B. Heo, M. Lee, S. Yun, and J. Y . Choi, “Knowledge transfer via distilla- tion of activation boundaries formed by hidden neurons,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 3779–3787
2019
-
[52]
Au- toprompt: Eliciting knowledge from language models with automatically generated prompts,
T. Shin, Y . Razeghi, R. L. Logan IV , E. Wallace, and S. Singh, “Au- toprompt: Eliciting knowledge from language models with automatically generated prompts,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for C...
2020
-
[53]
Hotflip: White-box adversarial examples for text classification,
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou, “Hotflip: White-box adversarial examples for text classification,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2018, pp. 31–36
2018
-
[54]
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters , vol. 7, no. 3, pp. 7327–7334, 2022
2022
-
[55]
Modularity through attention: Efficient training and transfer of language- conditioned policies for robot manipulation,
Y . Zhou, S. Sonawani, M. Phielipp, S. Stepputtis, and H. Amor, “Modularity through attention: Efficient training and transfer of language- conditioned policies for robot manipulation,” in Conference on Robot Learning. PMLR, 2023, pp. 1684–1695
2023
-
[56]
Rearrangement: A challenge for embodied ai,
D. Batra, A. X. Chang, S. Chernova, A. J. Davison, J. Deng, V . Koltun, S. Levine, J. Malik, I. Mordatch, R. Mottaghi et al., “Rearrangement: A challenge for embodied ai,” arXiv preprint arXiv:2011.01975 , 2020. IEEE TRANSACTIONS ON ROBOTICS 9
2011 arXiv
-
[57]
Ocrtoc: A cloud-based competition and benchmark for robotic grasping and manipulation,
Z. Liu, W. Liu, Y . Qin, F. Xiang, M. Gou, S. Xin, M. A. Roa, B. Calli, H. Su, Y . Sunet al., “Ocrtoc: A cloud-based competition and benchmark for robotic grasping and manipulation,” IEEE Robotics and Automation Letters, vol. 7, no. 1, pp. 486–493, 2021
2021
-
[58]
Transporter networks: Rearranging the visual world for robotic manipulation,
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V . Sindhwani et al. , “Transporter networks: Rearranging the visual world for robotic manipulation,” in Conference on Robot Learning . PMLR, 2021, pp. 726–747
2021
-
[59]
Cliport: What and where pathways for robotic manipulation,
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in Conference on robot learning . PMLR, 2022, pp. 894–906
2022
-
[60]
Decision transformer: Reinforcement learn- ing via sequence modeling,
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learn- ing via sequence modeling,” Advances in neural information processing systems, vol. 34, pp. 15 084–15 097, 2021
2021
-
[61]
Relay pol- icy learning: Solving long horizon tasks via imitation and reinforcement learning,
A. Gupta, V . Kumar, C. Lynch, S. Levine, and K. Hausman, “Relay pol- icy learning: Solving long horizon tasks via imitation and reinforcement learning,” Conference on Robot Learning (CoRL) , 2019
2019
-
[62]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2020
2020
-
[63]
Mask r-cnn,
K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969
2017
-
[64]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research, vol. 21, no. 140, pp. 1–67, 2020
2020
-
[65]
Cheating suffix: Targeted attack to text-to-image diffusion models with multi-modal priors,
D. Yang, Y . Bai, X. Jia, Y . Liu, X. Cao, and W. Yu, “Cheating suffix: Targeted attack to text-to-image diffusion models with multi-modal priors,” arXiv preprint arXiv:2402.01369 , 2024
2024 arXiv
-
[66]
Automatic and universal prompt injection attacks against large language models,
X. Liu, Z. Yu, Y . Zhang, N. Zhang, and C. Xiao, “Automatic and universal prompt injection attacks against large language models,” arXiv preprint arXiv:2403.04957, 2024
2024 arXiv
-
[67]
Enhancing adversarial example transferability with an intermediate level attack,
Q. Huang, I. Katsman, H. He, Z. Gu, S. Belongie, and S.-N. Lim, “Enhancing adversarial example transferability with an intermediate level attack,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4733–4742
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.