Adding stage-wise temporal and spatial memory to a heatmap-prediction 3D VLA policy yields strong results on memory-dependent manipulation benchmarks while keeping data efficiency.
3D Diffuser Actor: Policy diffusion with 3D scene representations,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
baseline 1
citation-polarity summary
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Adding stage-wise temporal and spatial memory to a heatmap-prediction 3D VLA policy yields strong results on memory-dependent manipulation benchmarks while keeping data efficiency.