Pith. sign in

REVIEW 21 cited by

TransTrack: Multiple Object Tracking with Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.15460 v2 pith:WWBEVTCR submitted 2020-12-31 cs.CV

classification cs.CV
keywords objecttranstrackmultipletrackingframemethodsnoveltransformer
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this work, we propose TransTrack, a simple but efficient scheme to solve the multiple object tracking problems. TransTrack leverages the transformer architecture, which is an attention-based query-key mechanism. It applies object features from the previous frame as a query of the current frame and introduces a set of learned object queries to enable detecting new-coming objects. It builds up a novel joint-detection-and-tracking paradigm by accomplishing object detection and object association in a single shot, simplifying complicated multi-step settings in tracking-by-detection methods. On MOT17 and MOT20 benchmark, TransTrack achieves 74.5\% and 64.5\% MOTA, respectively, competitive to the state-of-the-art methods. We expect TransTrack to provide a novel perspective for multiple object tracking. The code is available at: \url{https://github.com/PeizeSun/TransTrack}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The generator is the tracker: Multi-object tracking by painting persistent identity colours

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A video generator fine-tuned to emit persistent per-person colors tracks DanceTrack dancers at 40.3 HOTA with no detector or tracker.

  2. CAMELTrack: Context-Aware Multi-cue ExpLoitation for Online Multi-Object Tracking

    cs.CV 2025-05 conditional novelty 7.0 of 10

    CAMELTrack is an online tracker whose association step is learned end to end from multiple cues, reaching state-of-the-art HOTA on DanceTrack, SportsMOT, PoseTrack21 and BEE24, and competitive results on MOT17.

  3. CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Cross-domain language-guided tracking degrades sharply under weather/viewpoint/realism shifts, and a query-centric adaptation method partially restores performance.

  4. COVTrack++: Learning Open-Vocabulary Multi-Object Tracking from Continuous Videos via a Synergistic Paradigm

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Continuous TAO annotations plus multi-cue fusion, hierarchical aggregation, and temporal confidence propagation raise novel TETA to 35.4%/30.5% on TAO val/test.

  5. Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework

    cs.CV 2026-01 reject novelty 6.0 of 10

    A new LLM-generated dataset and an MLLM-based tracker claim state-of-the-art semantic multi-object tracking, but the evaluation protocol masks missed objects and ID switches.

  6. To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A taxonomy that classifies unified perception methods in autonomous driving into Early, Late, and Full Unified Perception based on task integration, tracking formulation, and representation flow.

  7. FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FusionTrack jointly optimizes single-view tracking and cross-view re-identification in a Transformer, and the new MDMOT benchmark covers arbitrary multi-drone views with overlapping and non-overlapping cameras.

  8. CoMotion: Concurrent Multi-person 3D Motion

    cs.CV 2025-04 conditional novelty 6.0 of 10

    CoMotion tracks multiple people's 3D poses online from monocular video using recurrent pose updates from image features, reaching 71.4 MOTA on PoseTrack21 and running about 12x faster than a strong prior 3D pose tracker.

  9. PD-SORT: Occlusion-Robust Multi-Object Tracking Using Pseudo-Depth Cues

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PD-SORT achieves higher HOTA than its OC-SORT baseline on DanceTrack, MOT17, and MOT20 by adding pseudo-depth states to the Kalman filter and using depth-volume IoU and quantized depth costs in data association.

  10. Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HGT-Track fuses visible and thermal drone video with a heterogeneous graph transformer and reports the best MOTA and IDF1 on the authors' new VT-Tiny-MOT benchmark.

  11. Temporally Consistent Dynamic Scene Graphs: An End-to-End Approach for Action Tracklet Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An end-to-end transformer couples detection with a temporal matching penalty and feedback queries, boosting temporal consistency of scene-graph predictions on Action Genome, OpenPVSG, and MEVA.

  12. FastTrackTr:Towards Fast Multi-Object Tracking with Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FastTrackTr reaches 166 FPS with TensorRT at 640x640 while scoring 62.4 HOTA on DanceTrack and 62.4 HOTA on MOT17 test.

  13. From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

    cs.CV 2025-09 accept novelty 5.0 of 10

    A survey organizing recent camera-based AI methods for vulnerable road user safety into four interlocking visual tasks and four open deployment challenges.

  14. CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios

    cs.CV 2025-07 conditional novelty 5.0 of 10

    CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.

  15. Depth-Aware Scoring and Hierarchical Alignment for Multiple Object Tracking

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A training-free MOT framework that adds zero-shot depth histograms and a hierarchical box/mask alignment score to association, with mixed state-of-the-art results.

  16. MapExpert: Online HD Map Construction with Simple and Efficient Sparse Map Element Expert

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MapExpert uses shape-specific sparse expert networks and a learnable temporal fusion module to improve online HD map construction by about 1.4-1.8 mAP over MapTracker on nuScenes and Argoverse2.

  17. Enhanced Multi-Object Tracking Using Pose-based Virtual Markers in 3x3 Basketball

    cs.CV 2024-12 reject novelty 5.0 of 10

    A pose-based virtual marker overlay, applied to test videos, tracks 3x3 basketball players with zero ID switches and a 72.6 HOTA on a private dataset, but the markers supply identity at test time.

  18. A Deep Dive into Generic Object Tracking: A Survey

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A survey that categorizes generic object tracking into Siamese, discriminative, and transformer-based paradigms and compares them across architecture and performance.

  19. YOLOv8-SMOT: An Efficient and Robust Framework for Real-Time Small Object Tracking via Slice-Assisted Training and Adaptive Association

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A YOLOv8 detector trained on overlapping slices plus an OC-SORT tracker with EMA motion direction and expanded IoU distance penalty achieves 55.205 SO-HOTA on the SMOT4SB public test set.

  20. Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A comprehensive survey that classifies modern MOT methods and shows, via cross-benchmark aggregation, that heuristic trackers dominate crowded linear-motion scenes while deep learning association methods dominate comp...

  21. Enhancing Thermal MOT: A Novel Box Association Method Leveraging Thermal Identity and Motion Similarity

    cs.CV 2024-11 conditional novelty 4.0 of 10

    Combining thermal histogram similarity with motion similarity improves ByteTrack and OCSORT MOT scores by about one to two MOTA points on a new RGB-thermal pedestrian dataset.

Pith tools