Target-Bench shows the best off-the-shelf video world model scores only 0.341 on semantic target-approaching and directional consistency, with fine-tuning on a small robot dataset yielding measurable gains.
Foundation models in autonomous driving: A survey on scenario generation and scenario analysis
8 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
NuRisk is a new VQA dataset for agent-level risk assessment in autonomous driving that benchmarks VLMs at 33% peak accuracy and shows a fine-tuned 7B model reaching 41% with 75% lower latency.
A constraint-satisfaction system that converts natural language traffic scenario descriptions into solvable constraints for precise closed-loop autonomous vehicle testing, outperforming baselines on diverse benchmarks.
SAVANT reformulates semantic anomaly detection as layered consistency verification, raising VLM recall by 18.5% on real driving images and enabling a fine-tuned 7B open model to reach 90.8% recall and 93.8% accuracy.
FrozenDrive enables zero-shot text-guided generation of consistent multi-view driving scenes via a parameter-free frozen diffusion backbone with spatio-temporal attention, improving autonomous driving models on adverse conditions via data augmentation.
NAN-SPOT detects unknown objects better than retraining-heavy methods by using Negative-Aware Norm from off-the-shelf detectors and introduces the expanded COCO-Open dataset.
Vision-language models can serve as zero-shot ODD sensors for autonomous driving when using definition-anchored chain-of-thought prompting with persona decomposition.
A tutorial that unifies explicit and implicit world models through shared predictive structure for applications in physical AI such as robotics.
citing papers explorer
-
Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
Target-Bench shows the best off-the-shelf video world model scores only 0.341 on semantic target-approaching and directional consistency, with fine-tuning on a small robot dataset yielding measurable gains.
-
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving
NuRisk is a new VQA dataset for agent-level risk assessment in autonomous driving that benchmarks VLMs at 33% peak accuracy and shows a fine-tuned 7B model reaching 41% with 75% lower latency.
-
Traffic Scenario Orchestration from Language via Constraint Satisfaction
A constraint-satisfaction system that converts natural language traffic scenario descriptions into solvable constraints for precise closed-loop autonomous vehicle testing, outperforming baselines on diverse benchmarks.
-
Can VLMs Unlock Semantic Anomaly Detection? A Framework for Structured Reasoning
SAVANT reformulates semantic anomaly detection as layered consistency verification, raising VLM recall by 18.5% on real driving images and enabling a fine-tuned 7B open model to reach 90.8% recall and 93.8% accuracy.
-
FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model
FrozenDrive enables zero-shot text-guided generation of consistent multi-view driving scenes via a parameter-free frozen diffusion backbone with spatio-temporal attention, improving autonomous driving models on adverse conditions via data augmentation.
-
Beyond Known Objects: A Novel Framework for Open-Set Object Detection using Negative-Aware Norm
NAN-SPOT detects unknown objects better than retraining-heavy methods by using Negative-Aware Norm from off-the-shelf detectors and introduces the expanded COCO-Open dataset.
-
Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models
Vision-language models can serve as zero-shot ODD sensors for autonomous driving when using definition-anchored chain-of-thought prompting with persona decomposition.
-
A Tutorial on World Models and Physical AI
A tutorial that unifies explicit and implicit world models through shared predictive structure for applications in physical AI such as robotics.