REVIEW 2 cited by
Context-Aware Integration of Language and Visual References for Natural Language Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Tracking by natural language specification (TNL) aims to consistently localize a target in a video sequence given a linguistic description in the initial frame. Existing methodologies perform language-based and template-based matching for target reasoning separately and merge the matching results from two sources, which suffer from tracking drift when language and visual templates miss-align with the dynamic target state and ambiguity in the later merging stage. To tackle the issues, we propose a joint multi-modal tracking framework with 1) a prompt modulation module to leverage the complementarity between temporal visual templates and language expressions, enabling precise and context-aware appearance and linguistic cues, and 2) a unified target decoding module to integrate the multi-modal reference cues and executes the integrated queries on the search image to predict the target location in an end-to-end manner directly. This design ensures spatio-temporal consistency by leveraging historical visual information and introduces an integrated solution, generating predictions in a single step. Extensive experiments conducted on TNL2K, OTB-Lang, LaSOT, and RefCOCOg validate the efficacy of our proposed approach. The results demonstrate competitive performance against state-of-the-art methods for both tracking and grounding.
Forward citations
Cited by 2 Pith papers
-
ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
A new vision-language tracking model uses LLM-annotated target words and a global target-context memory heatmap to achieve state-of-the-art precision on MGIT, TNL2K, and LaSOT benchmarks.
-
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
CTVLT converts text descriptions into spatial heatmaps via Grounding DINO and fuses them into a visual tracker, achieving reported state-of-the-art performance on MGIT, TNL2K, and LaSOT.
Discussion (0). Continue with ORCID to comment.