REVIEW 4 major objections 5 minor 1 cited by
Spatially Visual Perception for End-to-End Robotic Learning
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A policy that fuses corrupted RGB frames with monocular depth maps keeps 82-88 percent success across camera exposures from 10 to 170 ms, where standard imitation-learning baselines collapse.
desk verdict Real-robot exposure robustness, but the missing combined-data control makes the main attribution unsupported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a two-branch perception module feeding a transformer-based denoising diffusion policy. AugBlender, an extension of AugMix, applies random color-only corruptions such as hue, saturation, solarization, and gamma changes, with mixing weights drawn from a Dirichlet distribution and a logic gate $\beta=0.16$ deciding whether the output is an in-distribution or out-of-distribution image; restricting corruptions to color keeps the geometry aligned with depth. The depth branch runs the monocular estimator Depth Anything V2 on the same frames, producing spatial maps that stay consistent under exposure change. The fused features pass through separate per-view ResNet34 Feature Pyramid Network encoders and then into the action prediction module, so the policy can condition on depth when the RGB channel is corrupted.
What would settle it
Train a vanilla Diffusion Policy on the same combined dataset (62.5 percent fixed at 120 ms, 37.5 percent varied between 50 and 160 ms) without AugBlender or depth, and evaluate it at the ten test exposures. If its average success approaches the proposed model's 82-88 percent, the paper's attribution to the fused perception pipeline collapses; if it stays near the other baselines, the attribution is supported.
Extended reading notes
Core claim
The paper's central claim is that fusing corrupted RGB observations with uncorrupted monocular depth maps yields an imitation-learning perception module that stays stable under exposure variation. Concretely, the model combining AugBlender and Depth Anything V2 reaches average success rates of 88 percent on CupStack, 82 percent on PickSmall, and 83 percent on PickBig across exposures of 10 to 170 ms, whereas the vanilla Diffusion Policy baseline averages 23, 43, and 47 percent on those tasks. At the three lowest exposures, the proposed model keeps an 81 percent average on CupStack while all four baselines score zero at the lowest two levels. The paper attributes this to the depth channel remaining informative when RGB is nearly unusable, and to AugBlender teaching the policy to lean on depth when RGB departs from training conditions.
Load-bearing premise
The method-compensation claim depends on the combined training dataset being neutral; because only the full model trains on the larger mixed-exposure dataset, the extra data could explain part of the gain.
Editorial extensions
If this is right
- If the central claim holds, an RGB-only imitation learning stack can gain exposure stability by adding monocular depth and color-only augmentation, with no hardware change.
- The full model's average success is 82-88 percent across 10-170 ms exposures on all three tasks, more than double the second-best baseline on CupStack.
- At extreme low exposures (10-40 ms), the proposed model keeps an 81 percent average success on CupStack while every baseline scores zero at 10-20 ms.
- The ablation pattern suggests depth mainly rescues high-exposure performance while AugBlender widens the usable exposure range, so the two components act together rather than either alone.
- Because the depth estimator runs in real time on a single RTX 3090, the stability gain is available to low-cost two-camera robot setups.
Reading between the lines
- The paper leaves open whether dataset scale matters: the combined dataset is larger and only the full model trains on it, so part of the gain could come from data volume rather than the perception modules.
- If the mechanism is really depth reliability under lighting change, swapping Depth Anything V2 for a different monocular depth estimator should preserve most of the benefit; if not, the gain may come from DINOv2 features or from augmentation alone.
- The paper's own CupStack failure at exposure 10 suggests depth preserves geometry but not color identity, so tasks requiring color discrimination or strict sequencing will need a third cue, such as segmentation or object embeddings, in near-darkness.
- A standardized exposure-robustness benchmark for manipulation policies would let this single-robot evaluation be compared across labs and methods.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a multimodal perception pipeline for real-robot imitation learning: AugBlender, a stochastic RGB augmentation scheme, plus monocular depth maps from Depth Anything V2, fused through separate FPN encoders before a transformer-based Diffusion Policy. The central empirical claim is that this pipeline maintains 82–88% average success across camera exposures from 10 to 170 ms on three manipulation tasks, whereas baselines collapse at extreme exposures. The paper also argues that the robustness comes from method compensation rather than from the larger combined training dataset used for the proposed model.
Significance. If the attribution were supported, the result would be practically valuable: it demonstrates a low-cost (single arm, two RGB cameras, one RTX 3090) way to improve lighting robustness in visuomotor policies, using an off-the-shelf monocular depth model and a composable augmentation module. The paper also makes a useful conceptual distinction between data compensation and method compensation, and it attempts a real-robot evaluation with 10 exposure levels. The pipeline is plausible and the low-cost setup is a strength, as is the use of an externally pretrained depth model and an explicit algorithmic description of AugBlender. However, as reported, the main attribution is not established because the proposed model is the only condition trained on the enlarged combined dataset; no control with the same data but without AugBlender or depth is reported, and no uncertainty estimates are provided.
major comments (4)
- [Table 1; §4.4; §5.1] The proposed system is the only condition in Table 1 that can be trained on the combined dataset defined in §4.4 (62.5% fixed-exposure demos plus 37.5% varied-exposure demos). The paper's own §5.1 states that the model trained on the combined dataset 'demonstrated improved robustness' relative to the original and varied datasets. Consequently, the large success-rate advantage of Ours over DP+Depth and DP+AugBlender in Table 1 is confounded with training-data scale and exposure composition. A DP+CombinedDataset baseline, trained with the same data budget and no AugBlender or depth inputs, is the minimal control needed to support the attribution in the abstract and §5. Without it, the headline statement that 'our approach significantly boosts the success rate' is not established.
- [§4.3; Table 1] The central empirical comparison rests on success-rate point estimates with no error bars, confidence intervals, per-cell trial counts, or multiple training seeds. The text says each model was tested 20–50 times per exposure level and that 2–3 human evaluators had to agree unanimously, but Table 1 does not report n, and no inter-rater agreement statistic is given. Given that the claimed effect is large and precisely the kind of result that can be driven by one favorable seed or evaluation batch, the authors should report seed-level variance, per-cell trial counts, and a statistical comparison (or at least confidence intervals) before the robustness claim can be assessed.
- [§5.2] The section concludes that 'increasing training data diversity alone does not enhance model robustness to varying exposure levels,' but the evidence shown is inconsistent with this conclusion. DP+Varied Data is trained on the smaller varied-exposure dataset, not on the combined dataset that §5.1 credits with improved robustness. Thus the paper's own data-compensation result—that a larger combined dataset improves robustness—undercuts the method-compensation interpretation. The missing DP+CombinedDataset row would resolve this contradiction; in its absence the paragraph's causal conclusion should be withdrawn or substantially weakened.
- [§3.1] Depth maps for training episodes were preprocessed with the ViT-B variant of Depth Anything V2, while inference uses the ViT-S variant. This train/inference mismatch changes the input distribution of the depth channel that the policy is supposed to rely on under corrupted RGB. The paper asserts that the lighter model does not 'significantly compromis[e] depth estimation quality,' but no experiment or quantitative comparison is provided. Please either use the same depth model in both phases or report a comparison showing that the mismatch has no material effect on the reported success rates.
minor comments (5)
- [§3.2; Algorithm 1] The pseudocode has presentation issues: 'if ξ > βthen' is missing a space, and the else branch applies 'ai(xt)' without explicitly defining which augmentation ai is selected. The role of λ in line 4, where λ is set to 1 when ξ < β, should also be explained in the surrounding text.
- [§4.4] The sentence 'it is larger than prvevious dataset as we intutively mix them up' contains typos, and the three dataset sizes are not reported numerically; please give exact episode or demonstration counts for the original, varied, and combined datasets.
- [§3.2] The key AugBlender hyperparameters β, α, k, and λ are fixed (e.g., β=0.16) without any sensitivity study; since the ID/OOD balance is a central design choice, reporting at least a small sweep over β and λ would strengthen the method description.
- [§5.3] The explanation that 'present noise in the generated policy trajectories' caused CupStack failures at exposure 10 is not backed by trajectory-level evidence; either provide such evidence or phrase this passage as a hypothesis.
- [§7; References] Future Work contains spacing artifacts ('PointV oxelNet', 'V oxFormer'), and the reference list has issues: [14] and [15] are duplicates and [45] is an incomplete citation of Vaswani et al.
Circularity Check
No significant circularity: the robustness claim rests on externally benchmarked real-robot success rates, not on a derivation whose inputs already contain the outputs.
full rationale
The paper's central result is a measured success rate on physical robot tasks across exposure values 10–170 ms, which is an external benchmark not contained in the training inputs. AugBlender is a training-time augmentation procedure and Depth Anything V2 is a pretrained, internet-scale external depth model; neither is defined in terms of the reported success-rate metric. The 'Ours' condition's use of the combined dataset (62.5% fixed-exposure and 37.5% varied-exposure demos, per Section 4.4) is a genuine experimental confound because no DP+CombinedDataset control is reported, and it weakens the attribution of the gain to AugBlender-plus-depth fusion. However, that is a missing control and an attribution problem, not circularity: the paper does not fit a parameter and then report that parameter as a prediction, and no equation or construction makes the output equivalent to its input. The one self-citation, reference [47] used to justify the ResNet34 encoder choice, is not load-bearing for the headline robustness claim. Under the requirement to exhibit a specific reduction such as Eq. X = Eq. Y by construction or a fitted parameter renamed as a prediction, no circular step is present.
Assumptions & free parameters
free parameters (3)
- AugBlender logic gate threshold beta =
0.16
- AugBlender Dirichlet alpha, chain count k, and mixing parameter lambda sampling =
not reported
- Combined dataset mixing proportions =
62.5% fixed / 37.5% varied exposure
assumptions (3)
- domain assumption Depth Anything V2 depth estimates remain reliable and informative under extreme exposure and AugBlender color corruptions.
- domain assumption Color-only augmentations preserve the spatial alignment between corrupted RGB and depth maps.
- domain assumption Human unanimous success judgments over 20-50 trials per condition are precise enough to resolve the reported differences.
Cite this review
Pith. "Pith review of Spatially Visual Perception for End-to-End Robotic Learning." pith.science (2026). https://pith.science/paper/ZNPSAWM5
@misc{pith2026241117458,
author = {Pith},
title = {Pith review of: Spatially Visual Perception for End-to-End Robotic Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZNPSAWM5}},
note = {Machine review of arXiv:2411.17458}
}
read the original abstract
Recent advances in imitation learning have shown significant promise for robotic control and embodied intelligence. However, achieving robust generalization across diverse mounted camera observations remains a critical challenge. In this paper, we introduce a video-based spatial perception framework that leverages 3D spatial representations to address environmental variability, with a focus on handling lighting changes. Our approach integrates a novel image augmentation technique, AugBlender, with a state-of-the-art monocular depth estimation model trained on internet-scale data. Together, these components form a cohesive system designed to enhance robustness and adaptability in dynamic scenarios. Our results demonstrate that our approach significantly boosts the success rate across diverse camera exposures, where previous models experience performance collapse. Our findings highlight the potential of video-based spatial perception models in advancing robustness for end-to-end robotic learning, paving the way for scalable, low-cost solutions in embodied intelligence.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
Spatial RoboGrasp combines AugFusion, monocular depth, and grasp prompts in a diffusion policy, claiming large gains under exposure change, without released artifacts or error bars.
Reference graph
Works this paper leans on
-
[1]
Remembr: Building and reasoning over long- horizon spatio-temporal memory for robot navigation
Abrar Anwar, John Welsh, Joydeep Biswas, Soha Pouya, and Yan Chang. Remembr: Building and reasoning over long- horizon spatio-temporal memory for robot navigation. arXiv preprint arXiv:2409.13682, 2024. 8
arXiv 2024
-
[2]
Apollo: An open autonomous driving platform
ApolloAuto. Apollo: An open autonomous driving platform. https : / / github . com / ApolloAuto / apollo, 2023. Accessed: October 18, 2023. 2
work page 2023
-
[3]
Midas v3.1 – a model zoo for robust monocular relative depth estimation,
Reiner Birkl, Diana Wofk, and Matthias M ¨uller. Midas v3.1 – a model zoo for robust monocular relative depth estimation,
-
[4]
Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba. End to end learning for self-driving cars, 2016. 1
work page 2016
-
[5]
Emerg- ing properties in self-supervised vision transformers, 2021
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers, 2021. 8
work page 2021
-
[6]
Diffusion policy: Visuomotor policy learning via action dif- fusion, 2024
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action dif- fusion, 2024. 1, 4
work page 2024
-
[7]
Benchmarking robustness of 3d object detection to common corruptions in autonomous driving, 2023
Yinpeng Dong, Caixin Kang, Jinlai Zhang, Zijian Zhu, Yikai Wang, Xiao Yang, Hang Su, Xingxing Wei, and Jun Zhu. Benchmarking robustness of 3d object detection to common corruptions in autonomous driving, 2023. 2
work page 2023
-
[8]
Scene memory transformer for embodied agents in long-horizon tasks
Kuan Fang, Alexander Toshev, Li Fei-Fei, and Silvio Savarese. Scene memory transformer for embodied agents in long-horizon tasks. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 538–547, 2019. 8
work page 2019
Show all 54 references
-
[9]
Zhao, and Chelsea Finn
Zipeng Fu, Tony Z. Zhao, and Chelsea Finn. Mobile aloha: Learning bimanual mobile manipulation with low- cost whole-body teleoperation, 2024. 1
2024
-
[10]
Paschalidis
Vittorio Giammarino, James Queeney, and Ioannis Ch. Paschalidis. Visually robust adversarial imitation learning from videos with contrastive learning, 2024. 2
2024
-
[11]
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015. 4
2015
-
[12]
Benchmarking neu- ral network robustness to common corruptions and perturba- tions, 2019
Dan Hendrycks and Thomas Dietterich. Benchmarking neu- ral network robustness to common corruptions and perturba- tions, 2019. 2, 8
2019
-
[13]
Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan
Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. Augmix: A simple data processing method to improve robustness and uncertainty, 2020. 3
2020
-
[14]
Denoising diffu- sion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models, 2020. 2
2020
-
[15]
Denoising diffu- sion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models, 2020. 4
2020
-
[16]
Imitation with spatial-temporal heatmap: 2nd place solution for nuplan challenge, 2023
Yihan Hu, Kun Li, Pingyuan Liang, Jingyu Qian, Zhening Yang, Haichao Zhang, Wenxin Shao, Zhuangzhuang Ding, Wei Xu, and Qiang Liu. Imitation with spatial-temporal heatmap: 2nd place solution for nuplan challenge, 2023. 1
2023
-
[17]
Rekep: Spatio-temporal reasoning of rela- tional keypoint constraints for robotic manipulation, 2024
Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, and Li Fei-Fei. Rekep: Spatio-temporal reasoning of rela- tional keypoint constraints for robotic manipulation, 2024. 1, 2
2024
-
[18]
Yeying Jin, Beibei Lin, Wending Yan, Yuan Yuan, Wei Ye, and Robby T. Tan. Enhancing visibility in nighttime haze images using guided apsf and gradient adaptive convolution,
-
[19]
FORTRESS: Feature optimization and robustness techniques for 3d object detection systems
Caixin Kang, Xinning Zhou, Chengyang Ying, Wentao Shang, Xingxing Wei, Yinpeng Dong, and Hang Su. FORTRESS: Feature optimization and robustness techniques for 3d object detection systems. In ECCV 2024 Workshop on Multimodal Perception and Comprehension of Corner Cases in Auton...
2024
-
[20]
Mikhail Kennerley, Jian-Gang Wang, Bharadwaj Veeravalli, and Robby T. Tan. 2pcnet: Two-phase consistency training for day-to-night unsupervised domain adaptive object detec- tion, 2023. 1
2023
-
[21]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4), 2023. 2
2023
-
[22]
Domain adaptive imitation learning, 2020
Kuno Kim, Yihong Gu, Jiaming Song, Shengjia Zhao, and Stefano Ermon. Domain adaptive imitation learning, 2020. 2
2020
-
[23]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything, 2023. 1
2023
-
[24]
End-to-end planning of au- tonomous driving in industry and academia: 2022-2023,
Gongjin Lan and Qi Hao. End-to-end planning of au- tonomous driving in industry and academia: 2022-2023,
2022
-
[25]
Okami: Teaching hu- manoid robots manipulation skills through single video imi- tation
Jinhan Li, Yifeng Zhu, Yuqi Xie, Zhenyu Jiang, Mingyo Seo, Georgios Pavlakos, and Yuke Zhu. Okami: Teaching hu- manoid robots manipulation skills through single video imi- tation. In 8th Annual Conference on Robot Learning (CoRL),
-
[26]
Robust visual imi- tation learning with inverse dynamics representations, 2023
Siyuan Li, Xun Wang, Rongchang Zuo, Kewu Sun, Lingfei Cui, Jishiyu Ding, Peng Liu, and Zhe Ma. Robust visual imi- tation learning with inverse dynamics representations, 2023. 1
2023
-
[27]
Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar
Yiming Li, Zhiding Yu, Christopher Choy, Chaowei Xiao, Jose M. Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar. V oxformer: Sparse voxel transformer for camera-based 3d semantic scene completion, 2023. 8
2023
-
[28]
Data scaling laws in imi- tation learning for robotic manipulation
Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Ji- acheng You, and Yang Gao. Data scaling laws in imi- tation learning for robotic manipulation. arXiv preprint arXiv:2410.18647, 2024. 1
2024 arXiv
-
[29]
Feature pyramid networks for object detection, 2017
Tsung-Yi Lin, Piotr Doll ´ar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection, 2017. 4
2017
-
[30]
Occupancy prediction-guided neural planner for autonomous driving,
Haochen Liu, Zhiyu Huang, and Chen Lv. Occupancy prediction-guided neural planner for autonomous driving,
-
[31]
Robust imitation learning from corrupted demonstrations, 2022
Liu Liu, Ziyang Tang, Lanqing Li, and Dijun Luo. Robust imitation learning from corrupted demonstrations, 2022. 2
2022
-
[32]
The practice of mass produc- tion autonomous driving
Langechuan Patrick Liu. The practice of mass produc- tion autonomous driving. Presented at the CVPR 2023 E2EAD Workshop, 2023. Available at https : / / opendrivelab.com/e2ead/cvpr23. 2
2023
-
[33]
Point- voxel cnn for efficient 3d deep learning, 2019
Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. Point- voxel cnn for efficient 3d deep learning, 2019. 8
2019
-
[34]
Ecker, Matthias Bethge, and Wieland Brendel
Claudio Michaelis, Benjamin Mitzkus, Robert Geirhos, Evgenia Rusak, Oliver Bringmann, Alexander S. Ecker, Matthias Bethge, and Wieland Brendel. Benchmarking ro- bustness in object detection: Autonomous driving when win- ter is coming, 2020. 2, 8
2020
-
[35]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis, 2020. 2
2020
-
[36]
Dinov2: Learning robust visual features with- out supervision, 2024
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mah- moud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michae...
2024
-
[37]
Ro- bust multimodal vehicle detection in foggy weather using complementary lidar and radar signals
Kun Qian, Shilin Zhu, Xinyu Zhang, and Li Erran Li. Ro- bust multimodal vehicle detection in foggy weather using complementary lidar and radar signals. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 444–453, 2021. 2
2021
-
[38]
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. 1
2021
-
[39]
3d-outdet: A fast and memory efficient outlier detector for 3d lidar point clouds in adverse weather
Abu Mohammed Raisuddin, Tiago Cortinhal, Jesper Holm- blad, and Eren Erdal Aksoy. 3d-outdet: A fast and memory efficient outlier detector for 3d lidar point clouds in adverse weather. 2023. 1
2023
-
[40]
Sam 2: Segment anything in images and videos,
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Doll´ar, and Christoph Feic...
-
[41]
Raychaudhuri, Sujoy Paul, Jeroen van Baar, and Amit K
Dripta S. Raychaudhuri, Sujoy Paul, Jeroen van Baar, and Amit K. Roy-Chowdhury. Cross-domain imitation from ob- servations, 2021. 2
2021
-
[42]
Cliport: What and where pathways for robotic manipulation, 2021
Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Cliport: What and where pathways for robotic manipulation, 2021. 2
2021
-
[43]
Robust imitation learning from noisy demon- strations
V oot Tangkaratt, Nontawat Charoenphakdee, and Masashi Sugiyama. Robust imitation learning from noisy demon- strations. In Proceedings of The 24th International Confer- ence on Artificial Intelligence and Statistics, pages 298–306. PMLR, 2021. 2
2021
-
[44]
ALOHA 2 Team, Jorge Aldaco, Travis Armstrong, Robert Baruch, Jeff Bingham, Sanky Chan, Kenneth Draper, De- bidatta Dwibedi, Chelsea Finn, Pete Florence, Spencer Goodrich, Wayne Gramlich, Torr Hage, Alexander Herzog, Jonathan Hoech, Thinh Nguyen, Ian Storz, Baruch Taban- pour, ...
2024
-
[45]
Attention is all you need
A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017. 2
2017
-
[46]
4seasons: A cross-season dataset for multi-weather slam in autonomous driving, 2020
Patrick Wenzel, Rui Wang, Nan Yang, Qing Cheng, Qadeer Khan, Lukas von Stumberg, Niclas Zeller, and Daniel Cre- mers. 4seasons: A cross-season dataset for multi-weather slam in autonomous driving, 2020. 1, 2, 8
2020
-
[47]
Generalized robot learn- ing framework, 2024
Jiahuan Yan, Zhouyang Hong, Yu Zhao, Yu Tian, Yunxin Liu, Travis Davies, and Luhui Hu. Generalized robot learn- ing framework, 2024. 4
2024
-
[48]
Depth anything: Unleashing the power of large-scale unlabeled data, 2024
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data, 2024. 1, 2
2024
-
[49]
Depth any- thing v2, 2024
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2, 2024. 2, 3
2024
-
[50]
Primedepth: Efficient monocular depth estimation with a sta- ble diffusion preimage, 2024
Denis Zavadski, Damjan Kal ˇsan, and Carsten Rother. Primedepth: Efficient monocular depth estimation with a sta- ble diffusion preimage, 2024. 1, 2
2024
-
[51]
Multi-object detection at night for traffic in- vestigations based on improved ssd framework
Qiang Zhang, Xiaojian Hu, Yutao Yue, Yanbiao Gu, and Yizhou Sun. Multi-object detection at night for traffic in- vestigations based on improved ssd framework. Heliyon, 8 (11):e11570, 2022. 2
2022
-
[52]
Safe occlusion-aware au- tonomous driving via game-theoretic active perception
Zixu Zhang and Jaime Fisac. Safe occlusion-aware au- tonomous driving via game-theoretic active perception. In Robotics: Science and Systems XVII . Robotics: Science and Systems Foundation, 2021. 1
2021
-
[53]
Autofed: Heterogeneity-aware federated multimodal learning for robust autonomous driving, 2023
Tianyue Zheng, Ang Li, Zhe Chen, Hongbo Wang, and Jun Luo. Autofed: Heterogeneity-aware federated multimodal learning for robust autonomous driving, 2023. 2
2023
-
[54]
Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Anima Anandkumar, Jiashi Feng, and Jose M. Alvarez. Understand- ing the robustness in vision transformers, 2022. 8
2022
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.