REVIEW 2 major objections 29 references
A DVDrive Approach for doScenes Instructed Driving Challenge
T0 review · 2 major / 0 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read A divided-view perception module reduces cross-view interference to better align language instructions with local driving visuals in trajectory prediction.
desk verdict Incremental OmniDrive adaptation for the doScenes challenge with a divided-view module, but no results or ablations to support the claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The DVPE-style divided-view perception module that partitions features into local view spaces for intra-view visibility-aware cross-attention.
What would settle it
A controlled experiment on the doScenes challenge showing that an otherwise identical OmniDrive model without the divided-view module matches or exceeds the submitted model's trajectory accuracy and instruction-following metrics would falsify the claimed benefit.
Extended reading notes
Core claim
The paper claims that inserting a DVPE-style divided-view perception module into OmniDrive, by grouping query features and image tokens into divided local view spaces and performing visibility-aware cross-attention within each view, reduces irrelevant cross-view interference and helps the model better align language instructions with local driving-relevant visual evidence.
Load-bearing premise
That restricting cross-attention to local divided views improves alignment without discarding global scene context required for accurate trajectory prediction.
Editorial extensions
If this is right
- The adapted model produces 6-second ego trajectories that more closely follow given natural-language instructions.
- Multi-view visual grounding improves because each view processes only its own relevant tokens.
- Irrelevant information from other cameras is filtered before it reaches the language-alignment stage.
- The approach can be applied to other instruction-conditioned driving agents that use multiple cameras.
Reading between the lines
- The same local-view grouping might reduce compute in other multi-camera vision-language tasks outside driving.
- If the module preserves global context as claimed, it could be tested by measuring performance drop when global attention is fully removed.
- Extensions to longer prediction horizons or additional sensor types remain unexamined in the submission.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a submission to the doScenes Instructed Driving Challenge by adapting the OmniDrive vision-language-action model. It trains on instruction-annotated nuScenes scenes to predict 6-second ego trajectories (12 waypoints) and introduces a DVPE-style divided-view perception module in the perception head. This module groups query features and image tokens into local view spaces and applies visibility-aware cross-attention within each view, with the stated goal of reducing irrelevant cross-view interference and improving alignment between natural-language instructions and driving-relevant visual evidence.
Significance. If the divided-view module demonstrably improves instruction grounding without loss of necessary global context, the approach could provide a targeted architectural change for multi-camera VLA driving agents. The public release of code at https://github.com/feel12348/doscenes-omnidrive is a clear strength that supports reproducibility.
major comments (2)
- [Abstract] Abstract: the central claim that the DVPE-style module 'reduces irrelevant cross-view interference and helps the model better align language instructions with local driving-relevant visual evidence' is presented without any quantitative results, ablation studies, baseline comparisons, or error analysis, leaving the performance benefit unsupported.
- [Method] Method description (DVPE-style module): the design groups features into divided local view spaces for per-view attention, but the manuscript supplies no analysis or equations showing that critical cross-view information (e.g., relating a maneuver instruction to combined front/side views for trajectory planning) is retained rather than discarded.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our submission to the doScenes challenge. We address the major comments below and will revise the manuscript to strengthen the presentation of the DVPE-style module's benefits.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that the DVPE-style module 'reduces irrelevant cross-view interference and helps the model better align language instructions with local driving-relevant visual evidence' is presented without any quantitative results, ablation studies, baseline comparisons, or error analysis, leaving the performance benefit unsupported.
Authors: We agree that the abstract claim would benefit from direct empirical support. In the revised manuscript we will add ablation studies (with/without the divided-view module), baseline comparisons against the original OmniDrive, and error analysis on instruction grounding to quantitatively demonstrate the claimed reductions in cross-view interference and improved alignment. revision: yes
-
Referee: [Method] Method description (DVPE-style module): the design groups features into divided local view spaces for per-view attention, but the manuscript supplies no analysis or equations showing that critical cross-view information (e.g., relating a maneuver instruction to combined front/side views for trajectory planning) is retained rather than discarded.
Authors: We will expand the method section with additional analysis and equations that formalize how visibility-aware cross-attention within local views preserves necessary inter-view relations (e.g., front/side view fusion for maneuver instructions) while attenuating irrelevant tokens. This will include a brief derivation showing information flow across views is not discarded. revision: yes
Circularity Check
No circularity: architectural adaptation with no derivation chain
full rationale
The paper describes an engineering adaptation of OmniDrive plus addition of a DVPE-style divided-view module whose effect is stated directly as a design property (grouping into local view spaces and visibility-aware attention reduces interference). No equations, parameter fits, predictions, or self-citation chains are present that reduce any claimed result to its own inputs by construction. The central claim is a straightforward description of an architectural choice rather than a derived quantity, making the work self-contained against external benchmarks with no load-bearing circular steps.
Assumptions & free parameters
Cite this review
Pith. "Pith review of A DVDrive Approach for doScenes Instructed Driving Challenge." pith.science (2026). https://pith.science/paper/GEUBH7K6
@misc{pith2026260621623,
author = {Pith},
title = {Pith review of: A DVDrive Approach for doScenes Instructed Driving Challenge},
year = {2026},
howpublished = {\url{https://pith.science/paper/GEUBH7K6}},
note = {Machine review of arXiv:2606.21623}
}
read the original abstract
Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and historical motion, but also from a natural-language maneuver instruction. This paper presents our submission to the doScenes Instructed Driving Challenge, built upon OmniDrive, a vision-language-action driving agent with 3D perception, reasoning, and planning capabilities. We adapt OmniDrive to the doScenes setting by training it on instruction-annotated nuScenes scenes and generating a 6-second ego trajectory represented by 12 future waypoints. To improve multi-view visual grounding, we further introduce a DVPE-style divided-view perception module into the OmniDrive perception head. Instead of attending globally to all camera features, the proposed module groups query features and image tokens into divided local view spaces and performs visibility-aware cross-attention within each view. This design reduces irrelevant cross-view interference and helps the model better align language instructions with local driving-relevant visual evidence. The code is publicly available at: https://github.com/feel12348/doscenes-omnidrive.
Reference graph
Works this paper leans on
-
[1]
nuscenes: A multi- modal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom. nuscenes: A multi- modal dataset for autonomous driving. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11621–11631, 2020. 1
2020
-
[2]
Talk2car: Taking control of your self-driving car
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic, Luc Van Gool, and Marie-Francine Moens. Talk2car: Taking control of your self-driving car. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Pro- cessing (EMNLP-IJCNLP), pages 2088–2098, 2019. 2
2019
-
[3]
St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning
Shengchao Hu, Li Chen, Penghao Wu, Hongyang Li, Junchi Yan, and Dacheng Tao. St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning. In European Conference on Computer Vision, pages 533–549. Springer, 2022. 1
2022
-
[4]
Planning-oriented autonomous driving
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al. Planning-oriented autonomous driving. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17853–17862, 2023. 1
2023
-
[5]
Vad: Vectorized scene representa- tion for efficient autonomous driving
Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang. Vad: Vectorized scene representa- tion for efficient autonomous driving. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 8340–8350, 2023. 1
2023
-
[6]
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chong- hao Sima, Tong Lu, Yu Qiao, and Jifeng Dai. Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. InEuropean con- ference on computer vision, pages 1–18. Springer, 2022. 2
2022
-
[7]
Dacheng Liao, Mengshi Qi, Liang Liu, and Huadong Ma. Improving batch normalization with tta for robust object de- tection in self-driving.arXiv preprint arXiv:2411.18860,
-
[8]
Petr: Position embedding transformation for multi-view 3d object detection
Yingfei Liu, Tiancai Wang, Xiangyu Zhang, and Jian Sun. Petr: Position embedding transformation for multi-view 3d object detection. InEuropean Conference on Computer Vi- sion, pages 531–548. Springer, 2022. 2
2022
Show all 29 references
-
[9]
T2sg: Traffic topology scene graph for topology reasoning in autonomous driving
Changsheng Lv, Mengshi Qi, Liang Liu, and Huadong Ma. T2sg: Traffic topology scene graph for topology reasoning in autonomous driving. InProceedings of the IEEE/CVF 4 Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 2
2025
-
[10]
Gpt-driver: Learning to drive with gpt.arXiv preprint arXiv:2310.01415, 2023
Jiageng Mao, Yuxi Qian, Junjie Ye, Hang Zhao, and Yue Wang. Gpt-driver: Learning to drive with gpt.arXiv preprint arXiv:2310.01415, 2023. 2
2023 arXiv
-
[11]
Robust disentangled counterfactual learning for physical audiovi- sual commonsense reasoning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Mengshi Qi, Changsheng Lv, and Huadong Ma. Robust disentangled counterfactual learning for physical audiovi- sual commonsense reasoning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 2
2025
-
[12]
Towards balanced multi-modal learning in 3d human pose estimation.arXiv preprint arXiv:2501.05264, 2025
Mengshi Qi, Jiaxuan Peng, Xianlin Zhang, and Huadong Ma. Towards balanced multi-modal learning in 3d human pose estimation.arXiv preprint arXiv:2501.05264, 2025. 2
2025
-
[13]
Semantics-aware spatial-temporal binaries for cross- modal video retrieval.IEEE Transactions on Image Process- ing, 30:2989–3004, 2021
Mengshi Qi, Jie Qin, Yi Yang, Yunhong Wang, and Jiebo Luo. Semantics-aware spatial-temporal binaries for cross- modal video retrieval.IEEE Transactions on Image Process- ing, 30:2989–3004, 2021. 2
2021
-
[14]
Few-shot ensemble learning for video clas- sification with slowfast memory networks
Mengshi Qi, Jie Qin, Xiantong Zhen, Di Huang, Yi Yang, and Jiebo Luo. Few-shot ensemble learning for video clas- sification with slowfast memory networks. InProceedings of the 28th ACM International Conference on Multimedia,
-
[15]
Sports video captioning via attentive motion representation and group relationship modeling.IEEE Transactions on Cir- cuits and Systems for Video Technology, 30(8):2617–2633,
Mengshi Qi, Yunhong Wang, Annan Li, and Jiebo Luo. Sports video captioning via attentive motion representation and group relationship modeling.IEEE Transactions on Cir- cuits and Systems for Video Technology, 30(8):2617–2633,
-
[16]
Explainable action form assessment by exploiting multimodal chain-of-thoughts reasoning.arXiv preprint arXiv:2512.15153, 2025
Mengshi Qi, Yeteng Wu, Xianlin Zhang, and Huadong Ma. Explainable action form assessment by exploiting multimodal chain-of-thoughts reasoning.arXiv preprint arXiv:2512.15153, 2025. 2
2025 arXiv
-
[17]
Action quality assessment via hierarchical pose- guided multi-stage contrastive regression.arXiv preprint arXiv:2501.03674, 2025
Mengshi Qi, Hao Ye, Jiaxuan Peng, and Huadong Ma. Action quality assessment via hierarchical pose- guided multi-stage contrastive regression.arXiv preprint arXiv:2501.03674, 2025. 2
2025
-
[18]
Dc-sam: In-context segment anything in images and videos via dual consistency
Mengshi Qi, Pengfei Zhu, Xiangtai Li, Xiaoyang Bi, Lu Qi, Huadong Ma, and Ming-Hsuan Yang. Dc-sam: In-context segment anything in images and videos via dual consistency. arXiv preprint arXiv:2504.12080, 2025. 2
2025
-
[19]
doscenes: An autonomous driving dataset with natural language in- struction for human interaction and vision-language naviga- tion
Parthib Roy, Srinivasa Perisetla, Shashank Shriram, Harsha Krishnaswamy, Aryan Keskar, and Ross Greer. doscenes: An autonomous driving dataset with natural language in- struction for human interaction and vision-language naviga- tion. In2025 IEEE 28th International Conference ...
2025
-
[20]
Lmdrive: Closed-loop end-to- end driving with large language models
Hao Shao, Yuxuan Hu, Letian Wang, Steven L Waslander, Yu Liu, and Hongyang Li. Lmdrive: Closed-loop end-to- end driving with large language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15120–15130, 2024. 2
2024
-
[21]
Drivelm: Driving with graph visual question answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens Beißwenger, Ping Luo, Andreas Geiger, and Hongyang Li. Drivelm: Driving with graph visual question answering. InEuropean Conference on Computer Vision, pages 256–274. Springer, 2024. 2
2024
-
[22]
Dvpe: Divided view position embedding for multi- view 3d object detection.arXiv preprint arXiv:2407.16955,
Jiasen Wang, Zhenglin Li, Ke Sun, Xianyuan Liu, and Yang Zhou. Dvpe: Divided view position embedding for multi- view 3d object detection.arXiv preprint arXiv:2407.16955,
-
[23]
Exploring object-centric temporal modeling for efficient multi-view 3d object detection
Shihao Wang, Yingfei Liu, Tiancai Wang, Ying Li, and Xi- angyu Zhang. Exploring object-centric temporal modeling for efficient multi-view 3d object detection. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 3621–3631, 2023. 2
2023
-
[24]
Omnidrive: A holistic vision-language dataset for autonomous driving with counterfactual reasoning
Shihao Wang, Zhiding Yu, Xiaohui Jiang, Shiyi Lan, Min Shi, Nadine Chang, Jan Kautz, Ying Li, and Jose M Al- varez. Omnidrive: A holistic vision-language dataset for autonomous driving with counterfactual reasoning. InPro- ceedings of the computer vision and pattern recognitio...
2025
-
[25]
Detr3d: 3d object detection from multi-view images via 3d-to-2d queries
Yue Wang, Vitor Campagnolo Guizilini, Tianyuan Zhang, Yilun Wang, Hang Zhao, and Justin Solomon. Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. InConference on Robot Learning, pages 180–191. PMLR, 2022. 2
2022
-
[26]
Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee K Wong, Zhenguo Li, and Hengshuang Zhao. Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024. 2
2024
-
[27]
Safedriverag: Towards safe autonomous driv- ing with knowledge graph-based retrieval-augmented gener- ation
Hao Ye, Mengshi Qi, Zhaohong Liu, Liang Liu, and Huadong Ma. Safedriverag: Towards safe autonomous driv- ing with knowledge graph-based retrieval-augmented gener- ation. InProceedings of the 33rd ACM International Con- ference on Multimedia, 2025. 2
2025
-
[28]
Weakly-supervised temporal action localization by in- ferring salient snippet-feature
Wulian Yun, Mengshi Qi, Chuanming Wang, and Huadong Ma. Weakly-supervised temporal action localization by in- ferring salient snippet-feature. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, 2024. 2
2024
-
[29]
Unsupervised self-driving attention prediction via un- certainty mining and knowledge embedding
Pengfei Zhu, Mengshi Qi, Xia Li, Weijian Li, and Huadong Ma. Unsupervised self-driving attention prediction via un- certainty mining and knowledge embedding. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 8558–8568, 2023. 2 5
2023
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.