Pith. sign in

Safety, Security, and Cognitive Risks in World Models

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it
abstract

World models - learned internal simulators of environment dynamics - are rapidly becoming foundational to autonomous decision-making in robotics, autonomous vehicles, and agentic AI. By predicting future states in compressed latent spaces, they enable sample-efficient planning and long-horizon imagination without direct environment interaction. Yet this predictive power introduces a distinctive set of safety, security, and cognitive risks. Adversaries can corrupt training data, poison latent representations, and exploit compounding rollout errors to cause significant degradation in safety-critical deployments. At the alignment layer, world model-equipped agents are more capable of goal misgeneralisation, deceptive alignment, and reward hacking. At the human layer, authoritative world model predictions foster automation bias, miscalibrated trust, and planning hallucination. This paper surveys the world model landscape; introduces formal definitions of trajectory persistence and representational risk; presents a five-profile attacker taxonomy; and develops a unified threat model drawing on MITRE ATLAS and the OWASP LLM Top 10. We provide an empirical proof-of-concept demonstrating trajectory-persistent adversarial attacks on a GRU-based RSSM ($\mathcal{A}_1 = 2.26\times$ amplification, $-59.5\%$ reward reduction under adversarial fine-tuning), validate architecture-dependence via a stochastic RSSM proxy ($\mathcal{A}_1 = 0.65\times$), and probe a real DreamerV3 checkpoint (non-zero action drift confirmed). We propose interdisciplinary mitigations spanning adversarial hardening, alignment engineering, NIST AI RMF and EU AI Act governance, and human-factors design, arguing that world models require the same rigour as flight-control software or medical devices.

years

2026 3

representative citing papers

A Definition and Roadmap for World Models

cs.AI · 2026-07-07 · conditional · novelty 5.0

A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.

citing papers explorer

Showing 3 of 3 citing papers.

  • BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving cs.CV · 2026-06-19 · unverdicted · none · ref 46 · internal anchor

    Introduces BadDreamer, a backdoor attack that poisons the transition dynamics of video world models so that a trigger causes hallucination of obstacle-free futures, transferring to unsafe action predictions in autonomous driving.

  • ROBOSHACKLES: A Safety Dataset for Human-Injury Prevention in Embodied Foundation Models cs.RO · 2026-06-17 · unverdicted · none · ref 10 · internal anchor

    ROBOSHACKLES is a new safety dataset for embodied foundation models created through a pipeline of hazard-aware editing and video synthesis from real observations, with all six tested models generating unsafe actions at a 100% rate.

  • A Definition and Roadmap for World Models cs.AI · 2026-07-07 · conditional · none · ref 23 · internal anchor

    A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.