← back to paper
arxiv: 2608.03084 · 2 revisions
SUV: Future Scene Understanding as Video Generation for End-to-End Driving