Pith. sign in

REVIEW 3 major objections 6 minor 188 references

World models give robots a predictive core—and a new way digital compromise becomes physical harm.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 14:01 UTC pith:7GIWQTBF

load-bearing objection Useful lifecycle survey that reframes known attacks around predictive world models and safety-checker illusions; organizational, not empirical, and the taxonomy is the bet. the 3 major comments →

arxiv 2607.28226 v1 pith:7GIWQTBF submitted 2026-07-30 cs.CR cs.AI

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

classification cs.CR cs.AI
keywords embodied AI securityworld modelsvision-language-action modelspredictive safety illusionadversarial attacksbackdoor attacksruntime safetycyber-physical systems
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This survey argues that when embodied AI agents use world models to compress observations into states, imagine action-conditioned futures, and plan before acting, they open a security boundary that ordinary perception or policy defenses do not cover. Compromise in data, sensors, prompts, generators, memory, or feedback can corrupt the agent’s internal picture of the world so that a dangerous future looks safe, and that error can compound from imagination into real motion. Familiar attacks—poisoning, backdoors, spoofing, prompt injection, trajectory manipulation, supply-chain compromise—take on distinct meanings once they hit world states, learned dynamics, affordances, safety costs, or trajectory ranking. The paper also stresses a duality: the same predictive models can act as runtime safety shields, yet when compromised or over-trusted they produce predictive safety illusions—false certificates that an unsafe action is cleared. It reorganizes scattered attack literatures into a lifecycle taxonomy, maps them onto world-model security properties, and outlines how to evaluate and defend across provenance, grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.

Core claim

World-model-based embodied AI is not just another attack surface on robots; it is a distinct security boundary in which compromise propagates through predicted futures into physical action, familiar attack families gain new meanings by corrupting world-model security objects, and world models used as safety checkers can themselves generate predictive safety illusions when compromised or over-trusted.

What carries the argument

A lifecycle taxonomy of world-model-mediated decision making—from data construction and representation learning through state grounding, imagination, trajectory evaluation, execution, and agentic memory/tools—tied to seven security objects (state integrity, dynamics fidelity, affordance correctness, constraint compliance, trajectory-ranking integrity, uncertainty calibration, feedback provenance) and five recurring insights, especially predictive safety illusion.

Load-bearing premise

That cutting threats into these lifecycle stages and seven security properties is stable and complete enough to reorganize scattered attack work without leaving major risks uncounted or double-counted.

What would settle it

Build paired logs of imagined rollouts, safety-monitor decisions, and real outcomes under adaptive attacks on a world-model safety checker; if the predicted-safe-but-actually-unsafe rate does not separate cleanly from ordinary policy failure, or if the lifecycle map fails to place major published attacks without forced duplication, the central framing does not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Security evaluation must report predicted-safe-but-actually-unsafe executions, not only task success or textual refusal.
  • Defenses that only filter prompts or inspect generated video are incomplete if state estimates, dynamics, ranking, or feedback stay untrusted.
  • Generative world models used for synthetic demos become persistent poisoning sources and need provenance and physical-consistency audits.
  • Runtime shields that depend on a learned world model need independent monitors so policy and checker do not share the same corrupted state or dynamics.
  • Benchmarks must cover imagination-under-attack and feedback-stage contamination, which current suites largely miss.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Regulators and standards bodies writing robot or AV assurance rules will eventually need explicit requirements for logging assumed world state, predicted futures, and blocked actions—not only final commands.
  • The dual-use pattern (predictor as shield vs predictor as attack amplifier) likely generalizes to any agent that plans over a learned simulator, including non-embodied tool-using agents.
  • If trajectory-ranking integrity is the real decision surface, red teams should prioritize ranking and cost-model attacks over single-frame perception tricks.
  • Human operators shown confident generated futures may over-trust them; interface design becomes part of the security boundary, not only a usability issue.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This survey argues that world-model-based embodied AI creates a distinct security boundary: familiar attack families (poisoning, backdoors, spoofing, prompt injection, trajectory manipulation, supply-chain compromise) acquire new meaning when they corrupt world states, dynamics, affordances, safety costs, trajectory ranking, uncertainty, or feedback provenance. The authors organize threats along a lifecycle from data construction through grounding, imagination, trajectory evaluation, execution, and agentic extension; introduce five recurring insights (SP, SD, PA, NC, PI); and emphasize a duality in which world models can act as runtime safety shields yet, when compromised or over-trusted, produce predictive safety illusions. The manuscript maps attacks and defenses to seven security objects, catalogs benchmarks and coverage gaps, and proposes a paired prediction/execution evaluation protocol plus lifecycle-structured defenses.

Significance. If the organizational claim holds, the paper supplies a usable vocabulary and evaluation checklist for a rapidly fragmenting literature spanning model-based RL, VLA security, generative video world models, CPS sensor attacks, and agentic memory. The dual-use treatment of world models as safety checkers (Section V-B), the honest coverage-gap analysis (Table VII), and the concrete four-step evaluation protocol (Section VI-D) are particularly valuable: they turn a diffuse set of adjacent results into falsifiable measurement targets (especially predicted-safe-but-actually-unsafe rate). As a survey it does not ship new theorems or experiments, but the lifecycle cut and PI framing are concrete enough to guide benchmark design and defense comparison in subsequent work.

major comments (3)
  1. [Section I] Section I asserts two gaps relative to prior embodied-AI security surveys [15]–[19] and world-model safety notes [20], [21], but the manuscript never systematically contrasts stage/object coverage. A compact comparison (even one table row per prior survey: stages covered, whether WM-as-checker/PI is treated, whether trajectory ranking is separated from single-step prediction) is load-bearing for the claim that this lifecycle is not merely a relabeling. Without it, readers cannot judge whether the seven objects and SP/SD/PA/NC/PI cut reduce missed or double-counted risks relative to origin- or capability-centric taxonomies.
  2. [Table II; Sections IV–V] Table II’s evidence column marks most stages as adjacent (G#) rather than direct WM/WAM evidence, while the abstract and contributions state that familiar attacks “take on distinct meanings” for world-model objects. The distinction is acknowledged in places (e.g., PhysCond-WMA, TRAP, BadWAM, SafeDreamer) but not enforced in the prose of Sections IV–V, where VLA, CPS, and video-generation results are often narrated as if they already establish WM-object corruption. Please separate, per stage, (i) direct attacks on explicit world-model prediction/ranking/safety rollouts from (ii) extrapolated risks from implicit VLA/CPS work, so the central “distinct security boundary” claim is evidence-graded rather than uniformly asserted.
  3. [Section II-D; Table V] Section II-D’s seven security objects are used as the organizing vocabulary, yet several pairs are not operationally distinguished in the attack tables: constraint compliance vs. trajectory-ranking integrity (both appear under evaluation-stage attacks), and state integrity vs. feedback provenance (both absorb sensor/memory poisoning). For a taxonomy survey this boundary work is load-bearing. Add brief decision rules or examples showing when an attack is labeled as one object rather than another (e.g., in Table V rows for TRAP, BadWAM, AgentPoison, CHAI), or merge overlapping objects if the cut is not stable.
minor comments (6)
  1. [Figure 1] Figure 1 is central to the lifecycle claim but is only described in text in the submission copy; ensure the published figure clearly separates pre- vs. post-action-selection stages and marks cross-stage backdoor/injection paths so it matches Table II.
  2. [Section I] The abbreviations SP, SD, PA, NC, PI are introduced in Section I and reused heavily; a single reminder box or table listing one-sentence definitions would help readers who enter at Sections V–VII.
  3. [Table I] Table I mixes 2018–2026 systems and is useful, but the “WM Security Object” column sometimes lists one object where several apply (e.g., Cosmos 3: State; constraints). Consider allowing multi-labels consistently with Table II.
  4. [Section VI-D] Section VI-D’s protocol is strong; briefly note which existing resources in Table VI can already log paired imagined rollout vs. executed trajectory without new simulators, to make the protocol immediately actionable.
  5. [Throughout] Minor consistency: “W AM” / “WAM” spacing and “V oyager” / “Voyager” appear with stray spaces in several places; clean typography in production.
  6. [Section I] Related-work positioning of Baraldi et al. [20] and Parmar [21] could be one paragraph sharper: state explicitly what safety-risk reviews cover that this security-lifecycle survey adds (attacker model, supply chain, PI subversion metrics).

Circularity Check

0 steps flagged

No significant circularity: survey taxonomy reorganizes external attack literatures without self-derived predictions or load-bearing self-citation chains.

full rationale

This is a lifecycle survey of world-model-based embodied-AI security. It does not fit parameters, derive quantitative predictions, or invoke uniqueness theorems. The central contributions—lifecycle stages, seven security objects, five framing insights (SP/SD/PA/NC/PI), attack-to-object maps (Tables II/V), benchmark coverage gaps (Table VII), and the paired prediction/execution evaluation protocol—are organizational taxonomies built from cited external work (Dreamer, VLA attacks, PhysCond-WMA, TRAP, SafeDreamer, AgentPoison, sensor spoofing, etc.). Author-chosen labels do not reduce any claimed result to its own inputs by construction. No self-citation is load-bearing for a uniqueness or derivation claim. Honest non-finding: score 0; steps empty.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 4 invented entities

As a survey, the load-bearing commitments are definitional and organizational rather than fitted constants. The central claim depends on a broadened definition of world-model-based embodied AI, a chosen lifecycle decomposition, a seven-object security vocabulary, and the dual-use premise that predictive monitors can certify unsafe actions.

axioms (4)
  • ad hoc to paper A system is world-model-based when it maintains predictive structure used to simulate, score, or execute embodied behavior, including implicit VLA/physical priors.
    Section II-C explicitly adopts this broad scope rather than restricting to RSSM/Dreamer-style latent dynamics; the taxonomy's coverage depends on it.
  • domain assumption Embodied security failures are safety-critical because corrupted observations, plans, or controllers can become physical harm.
    Stated in Sections I and II-A and standard in CPS/robotics security; grounds why predictive corruption matters beyond digital error.
  • domain assumption Familiar attack families can be meaningfully remapped onto world-model security objects across lifecycle stages without losing essential distinctions.
    Implicit throughout Section III and Table II; the survey's value proposition depends on this remapping being clarifying rather than forced.
  • domain assumption When used as runtime safety checkers, world models can produce false safety certificates if state, dynamics, constraints, or uncertainty channels are subverted.
    Core of Section V-B and insight PI; supported by cited safe RL/shielding literature plus attack surfaces, but treated as a general principle.
invented entities (4)
  • Lifecycle taxonomy of world-model-mediated embodied decision making no independent evidence
    purpose: Organize threats, defenses, and evaluation from data construction through agentic extension.
    Introduced in Section III and Fig. 1 as the paper's primary organizing contribution.
  • Five named insights SP, SD, PA, NC, PI no independent evidence
    purpose: Compress recurring failure modes into reusable security insights that later drive defense priorities.
    Defined in Section I and reused in Tables V and IX; labels are paper-specific even if underlying phenomena are cited.
  • Seven world-model security objects no independent evidence
    purpose: Provide vocabulary for what attacks corrupt beyond CIA triad.
    Defined in Section II-D and used as the mapping target in Tables I-II and V.
  • Predictive safety illusion (PI) no independent evidence
    purpose: Name the failure mode where a compromised or over-trusted safety world model falsely certifies unsafe action.
    Highlighted in abstract, Section I, and Section V-B as the dual-use centerpiece.

pith-pipeline@v1.2.0-daily-grok45 · 36357 in / 3185 out tokens · 58462 ms · 2026-07-31T14:01:28.287303+00:00 · methodology

0 comments
read the original abstract

World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reactive control. This predictive layer, however, opens a new security boundary-compromise can propagate from data, sensors, prompts, or feedback into physical action. Rather than treating world models as an isolated component, this survey traces threats across their entire lifecycle-from data construction and representation learning, through state grounding and imagination, to trajectory evaluation, execution, and long-term adaptation via memory and tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states, learned dynamics, affordance estimates, or safety costs. We also highlight a duality: world models can serve as runtime safety shields, yet when compromised or over-trusted they generate predictive safety illusions. The survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.

Figures

Figures reproduced from arXiv: 2607.28226 by Fazhong Liu, Guoxing Chen, Haojin Zhu, Haozhen Tan, Yan Meng, Zhuoyan Chen.

Figure 1
Figure 1. Figure 1: Lifecycle taxonomy of security threats in world-model-based embodied AI. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

188 extracted references · 45 linked inside Pith

  1. [1]

    World models,

    D. Ha and J. Schmidhuber, “World models,”arXiv preprint arXiv:1803.10122, 2018

  2. [2]

    Dream to control: Learning behaviors by latent imagination,

    D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” inInternational Conference on Learning Representations, 2020

  3. [3]

    Mastering diverse domains through world models,

    D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap, “Mastering diverse domains through world models,”arXiv preprint arXiv:2301.04104, 2023

  4. [4]

    SafeDreamer: Safe reinforcement learning with world models,

    W. Huang, J. Ji, C. Xia, B. Zhang, and Y . Yang, “SafeDreamer: Safe reinforcement learning with world models,” inInternational Conference on Learning Representations (ICLR), 2024

  5. [5]

    Vlm-safe: Vision-language model-guided safety-aware reinforcement learning with world models for autonomous driving,

    Y . Qu, Z. Huang, Z. Sheng, J. Chen, Y . Leng, S. Labi, and S. Chen, “Vlm-safe: Vision-language model-guided safety-aware reinforcement learning with world models for autonomous driving,”arXiv preprint arXiv:2505.16377, 2025

  6. [6]

    Genie: Generative interactive environments,

    J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y . Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rockt ¨aschel, “Genie: Generative interactive environments,”arXiv preprint ...

  7. [7]

    Palm-e: An embodied multimodal language model,

    D. Driess, F. Xia, M. S. M. Sajjadiet al., “Palm-e: An embodied multimodal language model,” inProceedings of the 40th International Conference on Machine Learning, 2023

  8. [8]

    Rt-1: Robotics transformer for real-world control at scale,

    A. Brohan, N. Brown, J. Carbajalet al., “Rt-1: Robotics transformer for real-world control at scale,”arXiv preprint arXiv:2212.06817, 2022

  9. [9]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control,

    B. Zitkovich, T. Xu, T. Xiaoet al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” inConference on Robot Learning, 2023

  10. [10]

    Open- vla: An open-source vision-language-action model,

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketiet al., “Open- vla: An open-source vision-language-action model,”arXiv preprint arXiv:2406.09246, 2024

  11. [11]

    Octo: An open-source generalist robot policy,

    O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xuet al., “Octo: An open-source generalist robot policy,”arXiv preprint arXiv:2405.12213, 2024

  12. [12]

    Exploring the adversarial vulnerabilities of vision-language-action models in robotics,

    T. Wang, C. Han, J. Liang, W. Yang, D. Liu, L. X. Zhang, Q. Wang, J. Luo, and R. Tang, “Exploring the adversarial vulnerabilities of vision-language-action models in robotics,” inProceedings of the 15 IEEE/CVF International Conference on Computer Vision, 2025, pp. 6948–6958

  13. [13]

    Lidar spoofing meets the new-gen: Capability improvements, broken assumptions, and new attack strategies,

    T. Sato, Y . Hayakawa, R. Suzuki, Y . Shiiki, K. Yoshioka, and Q. A. Chen, “Lidar spoofing meets the new-gen: Capability improvements, broken assumptions, and new attack strategies,” inNetwork and Dis- tributed System Security Symposium (NDSS), 2024

  14. [14]

    Badrobot: Jailbreaking embodied llm agents in the physical world,

    H. Zhang, C. Zhu, X. Wang, Z. Zhou, C. Yin, M. Li, L. Xue, Y . Wang, S. Hu, A. Liuet al., “Badrobot: Jailbreaking embodied llm agents in the physical world,” inThe Thirteenth International Conference on Learning Representations, 2025

  15. [15]

    Aligning cyber space with physical world: A comprehensive survey on embodied ai,

    Y . Liu, W. Chen, Y . Bai, X. Liang, G. Li, W. Gao, and L. Lin, “Aligning cyber space with physical world: A comprehensive survey on embodied ai,”IEEE/ASME Transactions on Mechatronics, 2025

  16. [16]

    Towards robust and secure embodied ai: A survey on vulnerabilities and attacks,

    W. Xing, M. Li, M. Li, and M. Han, “Towards robust and secure embodied ai: A survey on vulnerabilities and attacks,”arXiv preprint arXiv:2502.13175, 2025

  17. [17]

    What breaks embodied ai security: Llm vulnerabilities, cps flaws, or something else?

    B. Ma, H. Guo, P. Lv, M. Xu, X. Dai, Y . Zhang, Y . Yang, and Y . Zhang, “What breaks embodied ai security: Llm vulnerabilities, cps flaws, or something else?”arXiv preprint arXiv:2602.17345, 2026

  18. [18]

    Safety of embodied navigation: A survey,

    Z. Wang, J. Hu, and R. Mu, “Safety of embodied navigation: A survey,” arXiv preprint arXiv:2508.05855, 2025

  19. [19]

    Security considerations in ai-robotics: A survey of current methods, challenges, and opportunities,

    S. Neupane, S. Mitra, I. A. Fernandez, S. Saha, S. Mittal, J. Chen, N. Pillai, and S. Rahimi, “Security considerations in ai-robotics: A survey of current methods, challenges, and opportunities,”IEEE Access, vol. 12, pp. 22 072–22 097, 2024

  20. [20]

    The safety challenge of world models for embodied ai agents: A review,

    L. Baraldi, Z. Zeng, C. Zhang, A. Nayak, H. Zhu, F. Liu, Q. Zhang, P. Wang, S. Liu, Z. Hu, and A. Cangelosi, “The safety challenge of world models for embodied ai agents: A review,”arXiv preprint arXiv:2510.05865, 2025

  21. [21]

    Safety, security, and cognitive risks in world models,

    M. Parmar, “Safety, security, and cognitive risks in world models,” arXiv preprint arXiv:2604.01346, 2026

  22. [22]

    Jailbreaking llm-controlled robots,

    A. Robey, Z. Ravichandran, V . Kumar, H. Hassani, and G. J. Pap- pas, “Jailbreaking llm-controlled robots,” in2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 11 948–11 956

  23. [23]

    Poex: Towards policy executable jailbreak attacks against the llm-based robots,

    X. Lu, Z. Huang, X. Li, C. Zhang, W. Xuet al., “Poex: Towards policy executable jailbreak attacks against the llm-based robots,”arXiv preprint arXiv:2412.16633, 2024

  24. [24]

    When world models dream wrong: Physical- conditioned adversarial attacks against world models,

    Z. Guo, S. Liang, A. Balogh, N. Lunberry, R.-C. Tu, M. Jela- sity, and D. Tao, “When world models dream wrong: Physical- conditioned adversarial attacks against world models,”arXiv preprint arXiv:2602.18739, 2026

  25. [25]

    State backdoor: Towards stealthy real-world poisoning attack on vision-language-action model in state space,

    J. Guo, W. Jiang, Y . Lin, Y . Liu, R. Zhang, G. Lu, A. Chen, X. Han, and H. Li, “State backdoor: Towards stealthy real-world poisoning attack on vision-language-action model in state space,”arXiv preprint arXiv:2601.04266, 2026

  26. [26]

    A systematic study of physical sensor attack hardness,

    H. Kim, R. Bandyopadhyay, M. O. Ozmen, Z. B. Celik, A. Bianchi, Y . Kim, and D. Xu, “A systematic study of physical sensor attack hardness,” in2024 IEEE Symposium on Security and Privacy (SP). IEEE, 2024, pp. 2328–2347

  27. [27]

    ANNIE: Be careful of your robots,

    Y . Huang, Z. Wang, Z. Wan, Y . Tian, H. Xu, Y . Han, and Y . Gan, “ANNIE: Be careful of your robots,”arXiv preprint arXiv:2509.03383, 2025

  28. [28]

    Silentdrift: Exploiting action chunking for stealthy backdoor attacks on vision-language- action models,

    B. Xu, Y . Shang, B. Wang, and E. Ferrara, “Silentdrift: Exploiting action chunking for stealthy backdoor attacks on vision-language- action models,”arXiv preprint arXiv:2601.14323, 2026

  29. [29]

    A first{Physical-World}trajectory prediction attack via{LiDAR- induced}deceptions in autonomous driving,

    Y . Lou, Y . Zhu, Q. Song, R. Tan, C. Qiao, W.-B. Lee, and J. Wang, “A first{Physical-World}trajectory prediction attack via{LiDAR- induced}deceptions in autonomous driving,” in33rd USENIX Security Symposium (USENIX Security 24), 2024, pp. 6291–6308

  30. [30]

    Trap: Tail-aware ranking attack for world-model planning,

    S. Duan, K. Zhang, and X. Luo, “Trap: Tail-aware ranking attack for world-model planning,”arXiv preprint arXiv:2605.01950, 2026

  31. [31]

    A framework for benchmarking and aligning task- planning safety in llm-based embodied agents,

    Y . Huang, L. Ding, Z. Tang, T. Wang, X. Lin, W. Zhang, M. Ma, and Y . Zhang, “A framework for benchmarking and aligning task- planning safety in llm-based embodied agents,”arXiv preprint arXiv:2504.14650, 2025

  32. [32]

    Safety guardrails for llm-enabled robots,

    Z. Ravichandran, A. Robey, V . Kumar, G. J. Pappas, and H. Hassani, “Safety guardrails for llm-enabled robots,”IEEE Robotics and Automa- tion Letters, 2026

  33. [33]

    A survey of embodied ai: From simulators to research tasks,

    J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan, “A survey of embodied ai: From simulators to research tasks,”IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 6, no. 2, pp. 230–244, 2022

  34. [34]

    A survey of robot learning from demonstration,

    B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,”Robotics and Autonomous Systems, vol. 57, no. 5, pp. 469–483, 2009

  35. [35]

    Llm-enabled cyber- physical systems: Survey, research opportunities, and challenges,

    W. Xu, M. Liu, O. Sokolsky, I. Lee, and F. Kong, “Llm-enabled cyber- physical systems: Survey, research opportunities, and challenges,” in 2024 IEEE International Workshop on Foundation Models for Cyber- Physical Systems & Internet of Things (FMSys). IEEE, 2024, pp. 50–55

  36. [36]

    V oyager: An open-ended embodied agent with large language models,

    G. Wang, Y . Xie, Y . Jianget al., “V oyager: An open-ended embodied agent with large language models,”Transactions on Machine Learning Research, 2024

  37. [37]

    Learning latent dynamics for planning from pixels,

    D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in International Conference on Machine Learning, 2019

  38. [38]

    Day- dreamer: World models for physical robot learning,

    P. Wu, A. Escontrela, D. Hafner, K. Goldberg, and P. Abbeel, “Day- dreamer: World models for physical robot learning,” inConference on Robot Learning, 2022

  39. [39]

    Safe model-based reinforcement learning with stability guarantees,

    F. Berkenkamp, M. Turchetta, A. P. Schoellig, and A. Krause, “Safe model-based reinforcement learning with stability guarantees,” inAd- vances in Neural Information Processing Systems, 2017

  40. [40]

    Nightshade: Prompt-specific poisoning attacks on text-to-image gen- erative models,

    S. Shan, W. Ding, J. Passananti, S. Wu, H. Zheng, and B. Y . Zhao, “Nightshade: Prompt-specific poisoning attacks on text-to-image gen- erative models,”arXiv preprint arXiv:2310.13828, 2023

  41. [41]

    Badvideo: Stealthy backdoor attack against text-to-video generation,

    R. Wang, M. Zhu, J. Ou, R. Chen, X. Tao, P. Wan, and B. Wu, “Badvideo: Stealthy backdoor attack against text-to-video generation,” arXiv preprint arXiv:2504.16907, 2025

  42. [42]

    T2vattack: Adversarial attack on text-to-video diffusion models,

    C. Li, Y . Min, J. Zhang, Z. Yuan, S. Shan, and X. Chen, “T2vattack: Adversarial attack on text-to-video diffusion models,”arXiv preprint arXiv:2512.23953, 2025

  43. [43]

    Badwam: When world-action models dream right but act wrong,

    Q. Li, X. Yang, and X. Wang, “Badwam: When world-action models dream right but act wrong,”arXiv preprint arXiv:2607.15207, 2026

  44. [44]

    Attackvla: Benchmarking adversarial and backdoor attacks on vision- language-action models,

    J. Li, Y . Zhao, X. Zheng, Z. Xu, Y . Li, X. Ma, and Y .-G. Jiang, “Attackvla: Benchmarking adversarial and backdoor attacks on vision- language-action models,”arXiv preprint arXiv:2511.12149, 2025

  45. [45]

    Bad- VLA: Towards backdoor attacks on vision-language-action models via objective-decoupled optimization,

    X. Zhou, G. Tie, G. Zhang, H. Wang, P. Zhou, and L. Sun, “Bad- VLA: Towards backdoor attacks on vision-language-action models via objective-decoupled optimization,” inAdvances in Neural Information Processing Systems (NeurIPS), 2025

  46. [46]

    DropVLA: An action-level backdoor attack on vision-language-action models,

    Z. Xu, J. Li, Y . Zhao, X. Zheng, X. Ma, and Y .-G. Jiang, “DropVLA: An action-level backdoor attack on vision-language-action models,” arXiv preprint arXiv:2510.10932, 2025

  47. [47]

    Do as i can, not as i say: Ground- ing language in robotic affordances,

    M. Ahn, A. Brohan, N. Brownet al., “Do as i can, not as i say: Ground- ing language in robotic affordances,”arXiv preprint arXiv:2204.01691, 2022

  48. [48]

    Inner monologue: Embodied reasoning through planning with language models,

    W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng et al., “Inner monologue: Embodied reasoning through planning with language models,” inConference on Robot Learning (CoRL), 2022

  49. [49]

    Code as policies: Language model programs for embodied control,

    J. Liang, W. Huang, F. Xiaet al., “Code as policies: Language model programs for embodied control,” in2023 IEEE International Conference on Robotics and Automation, 2023

  50. [50]

    Vipergpt: Visual inference via python execution for reasoning,

    D. Sur ´ıs, S. Menon, and C. V ondrick, “Vipergpt: Visual inference via python execution for reasoning,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023

  51. [51]

    Vista: A generalizable driving world model with high fidelity and versatile controllability,

    S. Gao, J. Yang, L. Chen, K. Chitta, Y . Qiu, A. Geiger, J. Zhang, and H. Li, “Vista: A generalizable driving world model with high fidelity and versatile controllability,”arXiv preprint arXiv:2405.17398, 2024

  52. [52]

    Urbanworld: An urban world model for 3d city generation,

    Y . Shang, Y . Lin, Y . Zheng, H. Fan, J. Ding, J. Feng, J. Chen, L. Tian, and Y . Li, “Urbanworld: An urban world model for 3d city generation,” arXiv preprint arXiv:2407.11965, 2024

  53. [53]

    Cosmos world foundation model platform for physical ai,

    NVIDIA, N. Agarwal, A. Ali, M. Bala, Y . Balajiet al., “Cosmos world foundation model platform for physical ai,”arXiv preprint arXiv:2501.03575, 2025

  54. [54]

    V-jepa 2: Self-supervised video models enable understanding, prediction and planning,

    M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Komeili, M. Muckley, A. Rizvi, C. Roberts, Y . LeCunet al., “V-jepa 2: Self-supervised video models enable understanding, prediction and planning,”arXiv preprint arXiv:2506.09985, 2025

  55. [55]

    Cosmos 3: Omnimodal world models for physical ai,

    NVIDIA, N. Agarwal, A. Ali, J. Allen, M. Antolini, A. Aubame, A. Azzolini, J. Bai, M. Bala, Y . Balajiet al., “Cosmos 3: Omnimodal world models for physical ai,”arXiv preprint arXiv:2606.02800, 2026

  56. [56]

    Causal world modeling for robot control,

    L. Li, Q. Zhang, Y . Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y . Shen, and Y . Xu, “Causal world modeling for robot control,”arXiv preprint arXiv:2601.21998, 2026

  57. [57]

    World action models are zero- shot policies,

    S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xianget al., “World action models are zero- shot policies,”arXiv preprint arXiv:2602.15922, 2026

  58. [58]

    Hy-world 2.0: A multi-modal world model for reconstructing, generating, and simulating 3d worlds,

    Team HY-World, C. Cao, X. Zuo, Z. Wang, Y . Zhang, J. Wu, Z. Liu, Y . Gong, Y . Liu, B. Yuanet al., “Hy-world 2.0: A multi-modal world model for reconstructing, generating, and simulating 3d worlds,”arXiv preprint arXiv:2604.14268, 2026. 16

  59. [59]

    Webworld: A large-scale world model for web agent training,

    Z. Xiao, J. Tu, C. Zou, Y . Zuo, Z. Li, P. Wang, B. Yu, F. Huang, J. Lin, and Z. Liu, “Webworld: A large-scale world model for web agent training,”arXiv preprint arXiv:2602.14721, 2026

  60. [60]

    A tutorial on world models and physical ai,

    I.-S. Oh, “A tutorial on world models and physical ai,”arXiv preprint arXiv:2606.12783, 2026

  61. [61]

    Poisoning attacks against support vector machines,

    B. Biggio, B. Nelson, and P. Laskov, “Poisoning attacks against support vector machines,” inProceedings of the 29th International Conference on Machine Learning, 2012

  62. [62]

    Badnets: Evaluating backdooring attacks on deep neural networks,

    T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg, “Badnets: Evaluating backdooring attacks on deep neural networks,”IEEE Access, vol. 7, pp. 47 230–47 244, 2019

  63. [63]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacyet al., “Learning transferable visual models from natural language supervision,” inProceedings of the 38th International Conference on Machine Learning, 2021

  64. [64]

    Poison frogs! targeted clean-label poisoning attacks on neural networks,

    A. Shafahi, W. R. Huang, M. Najibi, O. Suciu, C. Studer, T. Dumitras, and T. Goldstein, “Poison frogs! targeted clean-label poisoning attacks on neural networks,” inAdvances in Neural Information Processing Systems, 2018

  65. [65]

    Certified defenses for data poisoning attacks,

    J. Steinhardt, P. W. Koh, and P. Liang, “Certified defenses for data poisoning attacks,” inAdvances in Neural Information Processing Systems, 2017

  66. [66]

    Adversarial patch,

    T. B. Brown, D. Man ´e, A. Roy, M. Abadi, and J. Gilmer, “Adversarial patch,”arXiv preprint arXiv:1712.09665, 2017

  67. [67]

    Robust physical-world attacks on deep learning visual classification,

    K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno, and D. Song, “Robust physical-world attacks on deep learning visual classification,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018

  68. [68]

    Chai: Command hijacking against embodied ai,

    L. Burbano, D. Ortiz, Q. Sun, S. Yang, H. Tu, C. Xie, Y . Cao, and A. A. Cardenas, “Chai: Command hijacking against embodied ai,” arXiv preprint arXiv:2510.00181, 2025

  69. [69]

    Freezevla: Action-freezing attacks against vision-language-action models,

    X. Wang, J. Li, Z. Weng, Y . Wang, Y . Gao, T. Pang, C. Du, Y . Teng, Y . Wang, Z. Wuet al., “Freezevla: Action-freezing attacks against vision-language-action models,”arXiv preprint arXiv:2509.19870, 2025

  70. [70]

    {SA VIOR}: Securing autonomous vehicles with robust physi- cal invariants,

    R. Quinonez, J. Giraldo, L. Salazar, E. Bauman, A. Cardenas, and Z. Lin, “{SA VIOR}: Securing autonomous vehicles with robust physi- cal invariants,” in29th USENIX security symposium (USENIX Security 20), 2020, pp. 895–912

  71. [71]

    Real-time data-predictive attack-recovery for complex cyber-physical systems,

    L. Zhang, K. Sridhar, M. Liu, P. Lu, X. Chen, F. Kong, O. Sokolsky, and I. Lee, “Real-time data-predictive attack-recovery for complex cyber-physical systems,” in2023 IEEE 29th Real-Time and Embedded Technology and Applications Symposium (RTAS). IEEE, 2023, pp. 209–222

  72. [72]

    More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models,

    K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, M. Fritz, and P. Ho- henecker, “More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models,”arXiv preprint arXiv:2302.12173, 2023

  73. [73]

    Security for the robot operating system,

    B. Dieber, B. Breiling, S. Taurer, S. Kacianka, P. Schartner, and M. Hofbaur, “Security for the robot operating system,”Robotics and Autonomous Systems, vol. 98, pp. 192–203, 2017

  74. [74]

    Practical secure aggre- gation for privacy-preserving machine learning,

    K. Bonawitz, V . Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth, “Practical secure aggre- gation for privacy-preserving machine learning,” inproceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 2017, pp. 1175–1191

  75. [75]

    Manipulating machine learning: Poisoning attacks and countermea- sures for regression learning,

    M. Jagielski, A. Oprea, B. Biggio, C. Liu, C. Nita-Rotaru, and B. Li, “Manipulating machine learning: Poisoning attacks and countermea- sures for regression learning,” inProceedings of the IEEE Symposium on Security and Privacy, 2018

  76. [76]

    Adversarial back- door attack by naturalistic data poisoning on trajectory prediction in autonomous driving,

    M. Pourkeshavarz, M. Sabokrou, and A. Rasouli, “Adversarial back- door attack by naturalistic data poisoning on trajectory prediction in autonomous driving,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 14 885–14 894

  77. [77]

    Trojaning attack on neural networks,

    Y . Liu, S. Ma, Y . Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang, “Trojaning attack on neural networks,” inProceedings of the 25th Annual Network and Distributed System Security Symposium, 2018

  78. [78]

    Hidden trigger backdoor attacks,

    A. Saha, A. Subramanya, and H. Pirsiavash, “Hidden trigger backdoor attacks,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 7, pp. 11 957–11 965, 2020

  79. [79]

    Label-consistent backdoor attacks,

    A. Turner, D. Tsipras, and A. Madry, “Label-consistent backdoor attacks,”arXiv preprint arXiv:1912.02771, 2019

  80. [80]

    Goal-oriented backdoor attack against vision-language-action models via physical objects,

    Z. Zhou, Z. Xiao, H. Xu, J. Sun, D. Wang, and J. Zhang, “Goal-oriented backdoor attack against vision-language-action models via physical objects,”arXiv preprint arXiv:2510.09269, 2025

Showing first 80 references.