REVIEW 3 cited by
Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a novel Siamese Natural Language Tracker (SNLT), which brings the advancements in visual tracking to the tracking by natural language (NL) descriptions task. The proposed SNLT is applicable to a wide range of Siamese trackers, providing a new class of baselines for the tracking by NL task and promising future improvements from the advancements of Siamese trackers. The carefully designed architecture of the Siamese Natural Language Region Proposal Network (SNL-RPN), together with the Dynamic Aggregation of vision and language modalities, is introduced to perform the tracking by NL task. Empirical results over tracking benchmarks with NL annotations show that the proposed SNLT improves Siamese trackers by 3 to 7 percentage points with a slight tradeoff of speed. The proposed SNLT outperforms all NL trackers to-date and is competitive among state-of-the-art real-time trackers on LaSOT benchmarks while running at 50 frames per second on a single GPU.
Forward citations
Cited by 3 Pith papers
-
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
CTVLT converts text descriptions into spatial heatmaps via Grounding DINO and fuses them into a visual tracker, achieving reported state-of-the-art performance on MGIT, TNL2K, and LaSOT.
-
MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking
MambaVLT applies Mamba state space models to vision-language tracking with a time-evolving memory, beating several baselines on three of four benchmarks.
-
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
A fine-grained benchmark combining 10 challenge labels and 6 text types shows the value of language in vision-language tracking varies by scenario and tracker.
Discussion (0). Continue with ORCID to comment.