Pith. sign in

REVIEW 1 cited by

Robust Bird's Eye View Segmentation by Adapting DINOv2

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.10228 v1 pith:DTYUHBGQ submitted 2024-09-16 cs.CV

classification cs.CV
keywords dinov2adaptingbirdcameracorruptionsmodelperceptionrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Extracting a Bird's Eye View (BEV) representation from multiple camera images offers a cost-effective, scalable alternative to LIDAR-based solutions in autonomous driving. However, the performance of the existing BEV methods drops significantly under various corruptions such as brightness and weather changes or camera failures. To improve the robustness of BEV perception, we propose to adapt a large vision foundational model, DINOv2, to BEV estimation using Low Rank Adaptation (LoRA). Our approach builds on the strong representation space of DINOv2 by adapting it to the BEV task in a state-of-the-art framework, SimpleBEV. Our experiments show increased robustness of BEV perception under various corruptions, with increasing gains from scaling up the model and the input resolution. We also showcase the effectiveness of the adapted representations in terms of fewer learnable parameters and faster convergence during training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Birds Eye View Perception Models with Frozen Foundation Models: DINOv2 and Metric3Dv2

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Frozen DINOv2 features and Metric3Dv2 depth improve Lift-Splat-Shoot BEV segmentation by up to 8.9 IoU and a Metric3Dv2 PseudoLiDAR cloud adds about 3 IoU to Simple-BEV camera-only.

Pith tools