Pith. sign in

REVIEW 3 cited by

Exploring Enhanced Contextual Information for Video-Level Object Tracking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.11023 v1 pith:OFO3LU3K submitted 2024-12-15 cs.CV

classification cs.CV
keywords informationcontextuallayermcitrackobjecttrackingmambavisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contextual information at the video level has become increasingly crucial for visual object tracking. However, existing methods typically use only a few tokens to convey this information, which can lead to information loss and limit their ability to fully capture the context. To address this issue, we propose a new video-level visual object tracking framework called MCITrack. It leverages Mamba's hidden states to continuously record and transmit extensive contextual information throughout the video stream, resulting in more robust object tracking. The core component of MCITrack is the Contextual Information Fusion module, which consists of the mamba layer and the cross-attention layer. The mamba layer stores historical contextual information, while the cross-attention layer integrates this information into the current visual features of each backbone block. This module enhances the model's ability to capture and utilize contextual information at multiple levels through deep integration with the backbone. Experiments demonstrate that MCITrack achieves competitive performance across numerous benchmarks. For instance, it gets 76.6% AUC on LaSOT and 80.0% AO on GOT-10k, establishing a new state-of-the-art performance. Code and models are available at https://github.com/kangben258/MCITrack.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What You Have is What You Track: Adaptive and Robust Multimodal Tracking

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FlexTrack claims state-of-the-art multimodal tracking on complete and simulated missing-modality benchmarks, using heterogeneous mixture-of-experts fusion and a video-level masking training strategy.

  2. AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.

  3. Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning

    cs.CV 2026-08 conditional novelty 4.0 of 10

    A pruned-head student with dual spatial and semantic distillation reaches 54 FPS and near-teacher accuracy on RGB-T and RGB-E tracking.

Pith tools