Pith. sign in

REVIEW 7 cited by

On the Vulnerability of LLM/VLM-Controlled Robotics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10340 v5 pith:R6CB7RKX submitted 2024-02-15 cs.RO cs.AI

On the Vulnerability of LLM/VLM-Controlled Robotics

classification cs.RO cs.AI
keywords inputvlm-controlledroboticsystemsmodelsvulnerabilitiesexecutionmodality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this work, we highlight vulnerabilities in robotic systems integrating large language models (LLMs) and vision-language models (VLMs) due to input modality sensitivities. While LLM/VLM-controlled robots show impressive performance across various tasks, their reliability under slight input variations remains underexplored yet critical. These models are highly sensitive to instruction or perceptual input changes, which can trigger misalignment issues, leading to execution failures with severe real-world consequences. To study this issue, we analyze the misalignment-induced vulnerabilities within LLM/VLM-controlled robotic systems and present a mathematical formulation for failure modes arising from variations in input modalities. We propose empirical perturbation strategies to expose these vulnerabilities and validate their effectiveness through experiments on multiple robot manipulation tasks. Our results show that simple input perturbations reduce task execution success rates by 22.2% and 14.6% in two representative LLM/VLM-controlled robotic systems. These findings underscore the importance of input modality robustness and motivate further research to ensure the safe and reliable deployment of advanced LLM/VLM-controlled robotic systems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space

    cs.CR 2026-01 conditional novelty 6.0

    A backdoor attack on vision-language-action robot policies uses the arm's initial joint configuration as the trigger, achieving >90% triggered failure with only small clean-task degradation.

  2. High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models

    cs.CV 2025-12 unverdicted novelty 6.0

    High-entropy tokens act as concentrated multimodal failure points in VLMs, enabling sparse Entropy-Guided Attacks that achieve 93-95% success and 30-38% harmful rates with cross-model transfer.

  3. High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models

    cs.CV 2025-12 conditional novelty 6.0

    Attacking only the top 20% high-entropy token positions in vision-language models causes comparable semantic damage and more harmful outputs than global attacks, and these vulnerable tokens transfer across model archi...

  4. AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models

    cs.RO 2025-11 conditional novelty 6.0

    Structured prompting plus a discrete skill library lets a frozen VLM direct aerial manipulation, reaching 87.5% simulated and 80% hardware success in pick-and-place tasks.

  5. When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective

    cs.SE 2025-09 conditional novelty 6.0

    The first empirical taxonomy of LLM tasks in UAVs, with an academia-industry comparison and survey, shows LLMs are used mainly for planning and interaction, not direct control.

  6. A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics

    cs.RO 2025-11 reject novelty 4.0

    Storing operator-validated diagnoses in a vector database and retrieving them during anomaly characterization cuts diagnostic dialog turns by 71% and raises characterization specificity from 2.7 to 4.8 in a BlueROV2 t...

  7. When control meets large language models: From words to dynamics

    eess.SY 2026-02 unverdicted novelty 3.0

    The paper proposes a bidirectional continuum between LLMs and control systems, covering LLM-assisted controller design, control-based LLM steering, and state-space modeling of LLMs.