REVIEW 3 cited by
Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a novel method for scene change detection that leverages the robust feature extraction capabilities of a visual foundational model, DINOv2, and integrates full-image cross-attention to address key challenges such as varying lighting, seasonal variations, and viewpoint differences. In order to effectively learn correspondences and mis-correspondences between an image pair for the change detection task, we propose to a) ``freeze'' the backbone in order to retain the generality of dense foundation features, and b) employ ``full-image'' cross-attention to better tackle the viewpoint variations between the image pair. We evaluate our approach on two benchmark datasets, VL-CMU-CD and PSCD, along with their viewpoint-varied versions. Our experiments demonstrate significant improvements in F1-score, particularly in scenarios involving geometric changes between image pairs. The results indicate our method's superior generalization capabilities over existing state-of-the-art approaches, showing robustness against photometric and geometric variations as well as better overall generalization when fine-tuned to adapt to new environments. Detailed ablation studies further validate the contributions of each component in our architecture. Our source code is available at: https://github.com/ChadLin9596/Robust-Scene-Change-Detection.
Forward citations
Cited by 3 Pith papers
-
Multi-View Pose-Agnostic Change Localization with Zero Labels
A label-free method embeds change information into a 3D Gaussian Splatting model, enabling multi-view and unseen-view change localization.
-
Environmental Change Detection: Toward a Practical Task of Scene Change Detection
Environmental Change Detection removes the aligned-reference assumption from scene change detection, and a retrieval-plus-aggregation framework outperforms a strong baseline on reconstructed benchmarks.
-
ViewDelta: Scaling Scene Change Detection through Text-Conditioning
ViewDelta uses text prompts to define relevant scene changes, enabling one model to work across multiple change-detection datasets and view angles.
Discussion (0). Continue with ORCID to comment.