Pith. sign in

Can Visual Foundation Models Achieve Long-term Point Tracking?

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Large-scale vision foundation models have demonstrated remarkable success across various tasks, underscoring their robust generalization capabilities. While their proficiency in two-view correspondence has been explored, their effectiveness in long-term correspondence within complex environments remains unexplored. To address this, we evaluate the geometric awareness of visual foundation models in the context of point tracking: (i) in zero-shot settings, without any training; (ii) by probing with low-capacity layers; (iii) by fine-tuning with Low Rank Adaptation (LoRA). Our findings indicate that features from Stable Diffusion and DINOv2 exhibit superior geometric correspondence abilities in zero-shot settings. Furthermore, DINOv2 achieves performance comparable to supervised models in adaptation settings, demonstrating its potential as a strong initialization for correspondence learning.

citation-role summary

extension 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

extension 1

polarities

extend 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Exploring Temporally-Aware Features for Point Tracking cs.CV · 2025-01-21 · conditional · none · ref 2 · internal anchor

    A DINOv2 backbone augmented with temporal adapters tracks video points accurately using only soft-argmax matching, without iterative refinement.