A three-branch model (VideoMAE, a two-layer sensor MLP, BERT, fused into a BART-based explainer) reports 92.5% action accuracy and a 0.75 BLEU-4 on nuScenes, with simulated attention maps and no released code.
Why did the AI make that decision?,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency
A three-branch model (VideoMAE, a two-layer sensor MLP, BERT, fused into a BART-based explainer) reports 92.5% action accuracy and a 0.75 BLEU-4 on nuScenes, with simulated attention maps and no released code.