REVIEW 67 cited by
BoT-SORT: Robust Associations Multi-Pedestrian Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The goal of multi-object tracking (MOT) is detecting and tracking all the objects in a scene, while keeping a unique identifier for each object. In this paper, we present a new robust state-of-the-art tracker, which can combine the advantages of motion and appearance information, along with camera-motion compensation, and a more accurate Kalman filter state vector. Our new trackers BoT-SORT, and BoT-SORT-ReID rank first in the datasets of MOTChallenge [29, 11] on both MOT17 and MOT20 test sets, in terms of all the main MOT metrics: MOTA, IDF1, and HOTA. For MOT17: 80.5 MOTA, 80.2 IDF1, and 65.0 HOTA are achieved. The source code and the pre-trained models are available at https://github.com/NirAharon/BOT-SORT
Forward citations
Showing 60 of 67 Pith papers that cite this
-
CAMELTrack: Context-Aware Multi-cue ExpLoitation for Online Multi-Object Tracking
CAMELTrack is an online tracker whose association step is learned end to end from multiple cues, reaching state-of-the-art HOTA on DanceTrack, SportsMOT, PoseTrack21 and BEE24, and competitive results on MOT17.
-
CHIRLA: Comprehensive High-resolution Identification and Re-identification for Large-scale Analysis
A benchmark dataset with 963,554 bounding boxes of 22 people across seven cameras over seven months shows current Re-ID models drop to 18.81% CMC@1 on long-term indoor re-identification.
-
JitTrack: Onboard Multi-Object Tracking Against Viewpoint Jitter for Agile UAVs
JitTrack adds three motion-aware modifications to a query-based transformer tracker, improving identity-preserving multi-object tracking under UAV camera jitter on benchmarks and on a real quadrotor.
-
Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football
A junk-possession index and a video-based Space-Creation Index together classify sterile ball circulation versus space-creating possession, and the event flag adds outcome information beyond on-ball value models.
-
Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline
A two-pass coarse-to-fine pipeline using a frozen Qwen3-VL model, YOLO11x, and BoT-SORT achieves a harmonic mean of 0.504 on the ACCIDENT@CVPR zero-shot test set, beating the best published baseline by 22%.
-
Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers
A system that spatializes dialogue, adds diegetic sound, and offers tap-to-hear descriptions improved blind viewers' understanding of character position, movement, actions, and visual details in dialogue-only scenes.
-
Training-Free Off-Screen Player Imputation for Broadcast-Based Spatial Football Analytics
Role-anchored centroid voting, a training-free online imputer, roughly halves hidden-zone pitch-control error from ignoring off-screen players and cuts control-share error to 28–48% of the ignore baseline across three...
-
Unsupervised Detection of Entry and Exit Regions from Vehicle Trajectories for Camera-Agnostic Turning Movement Counts
Clustering start and end points of vehicle trajectories yields persistent entry/exit polygons that classify turning movements at ~3% median error across 25 uncalibrated cameras.
-
WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps
A new insect-tracking benchmark reveals that five standard MOT methods suffer severe identity fragmentation on 8-minute sequences even with oracle detections, with simple spatial stitching recovering significant gains.
-
Resonance-enhanced integrated acousto-optic beam steering
A TFLN ring-resonator-enhanced acousto-optic beam steerer reaches 26% efficiency and 18° FOV and supports FMCW LiDAR via electro-optic resonance locking.
-
Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods
GD3A and DVTrack, driven by optimal-transport descriptor matching with an adaptive dustbin score, set state-of-the-art results on a new moving-drone dense-crowd counting and tracking benchmark.
-
Learning Association via Track-Detection Matching for Multi-Object Tracking
TDLP uses a link-prediction head to match tracks to detections, beating heuristic and metric-learning trackers on several MOT benchmarks while underperforming on MOT17.
-
GenTrack2: An Improved Hybrid Approach for Multi-Object Tracking
GenTrack2 combines particle filtering, swarm optimization, and Hungarian data association with a velocity regression, and reports the highest tracking accuracy and fewest ID switches on one MOT17 sequence.
-
DeepSea MOT: A benchmark dataset for multi-object tracking on deep-sea video
DeepSea MOT, a public four-video benchmark with manually corrected ground truth, enables the first standardized HOTA comparisons of multi-object detectors and trackers on deep-sea ROV footage.
-
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
An end-to-end multi-view pedestrian tracker that aggregates motion and appearance costs over K past timestamps, outperforming prior methods on GMVD, Wildtrack, and MultiviewX, though with a validation-protocol concern...
-
VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation
Claims a view-invariant pose representation method for figure skating jump segmentation reaches over 92% F1@50, but the submitted text is a different manuscript, leaving the claim unverified.
-
GRASPTrack: Geometry-Reasoned Association via Segmentation and Projection for Multi-Object Tracking
A depth-aware MOT tracker using mask-guided 3D point clouds, voxelized 3D IoU association, adaptive Kalman noise, and 3D motion consistency surpasses prior TBD methods on MOT17, MOT20, and DanceTrack.
-
TrackOR: Towards Personalized Intelligent Operating Rooms Through Robust Tracking
TrackOR keeps persistent staff identities in the operating room by matching 3D geometric signatures, reporting +11% association accuracy over its strongest baseline and enabling per-person workflow trajectories.
-
Multi-tracklet Tracking for Generic Targets with Adaptive Detection Clustering
A tracklet-based multi-hypothesis tracker with adaptive detection clustering achieves competitive MOTA and IDF1 on GMOT-40 without category-specific knowledge.
-
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
A natural-language video analytics system combining bandit-based sampling, open-vocabulary detection, and trajectory linking reports higher query accuracy than closed-world baselines on a new 18-predicate traffic benchmark.
-
RoundaboutHD: High-Resolution Real-World Urban Environment Benchmark for Multi-Camera Vehicle Tracking
A high-resolution, real-world roundabout dataset for multi-camera vehicle tracking, with 512 identities, four non-overlapping 4K cameras, and baselines across four tasks.
-
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
A new USV dataset combining 4D radar, camera, GPS, and IMU data for tracking boats, ships, and vessels on inland waterways, together with a radar-camera matching method that consistently improves two-stage trackers.
-
NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
NOVA lets a drone track and follow a moving target at over 50 km/h through GPS-denied, unstructured terrain using only a stereo camera and an IMU, with no maps, motion capture, or external localization.
-
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
A new benchmark, DRAMA-X, adds multi-class directional intents, risk labels, and action suggestions for vulnerable road users to frames from the DRAMA dataset, and shows that scene-graph reasoning improves VLM risk sc...
-
Multiple Object Tracking in Video SAR: A Benchmark and Tracking Baseline
A new public video SAR multi-object tracking benchmark, plus a DETR-based tracker with line feature enhancement and motion-aware association, reports state-of-the-art results on that benchmark.
-
MOSE: A Novel Orchestration Framework for Stateful Microservice Migration at the Edge
A new edge-computing framework orchestrates stateful container migration, preserving live connections and meeting user-chosen KPI targets with up to 77 percent lower downtime.
-
Out of the Past: An AI-Enabled Pipeline for Traffic Simulation from Noisy, Multimodal Detector Data and Stakeholder Feedback
An AI pipeline (computer vision, quadratic optimization, and LLM-generated constraints) builds a Strongsville traffic simulation from noisy camera and loop detector data, but the simulation's accuracy is only checked ...
-
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
LASP adds continuous-time history fusion and trajectory-based motion compensation to 3D object detection, keeping online accuracy close to offline on edge devices.
-
EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
An LLM-based dialogue understanding module and a visual saliency model are combined with cinematic constraints and dynamic programming to automatically edit static wide-angle stage recordings into engaging multi-shot videos.
-
PD-SORT: Occlusion-Robust Multi-Object Tracking Using Pseudo-Depth Cues
PD-SORT achieves higher HOTA than its OC-SORT baseline on DanceTrack, MOT17, and MOT20 by adding pseudo-depth states to the Kalman filter and using depth-volume IoU and quantized depth costs in data association.
-
RoHan: Robust Hand Detection in Operation Room
Combining synthetic glove augmentation with iterative self-training on unlabeled videos gives OR hand detection mAP50 of about 0.91 on two surgical benchmarks.
-
TGOSPA Metric Parameters Selection and Evaluation for Visual Multi-object Tracking
The paper derives practical rules for setting the switching penalty and cut-off of the TGOSPA tracking metric and recommends three application-specific parameter setups for computer vision.
-
SEAL: Semantic Attention Learning for Long Video Representation
SEAL compresses long videos into scene, object, and action tokens, then selects a query-relevant, diverse subset, reporting improvements on LVBench, MovieChat-1K, and Ego4D.
-
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
ConsisID generates identity-preserving videos by injecting low-frequency facial features into shallow layers and high-frequency identity features into attention blocks of a DiT video model.
-
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
The paper defines the online episodic-memory query task OVQ2D and shows ESOM, a detect-track-memorize-retrieve system, achieves only ~4% success on Ego4D, rising to 81.92% with oracle detection and tracking.
-
Event-RGB Adaptive Tracking for Nighttime Highway Perception
JEAT jointly associates RGB and event detections with NIS-adapted measurement noise, raising MOTA on unlit nighttime highways from 46% (RGB) / 69% (event) to 77% on a new CARLA dataset.
-
GenTrack3: Hybrid Stochastic-Deterministic Online Multi-Object Tracking with Cluster-Aware Association
GenTrack3 presents a cluster-aware association method that partitions the track-detection cost matrix into smaller local matrices and reports competitive MOT scores on two pedestrian-tracking sequences.
-
CLIFE: Camera-LiDAR Fusion Framework for Edge-Deployable Roadside VRU Perception
An edge-deployed camera–LiDAR late-fusion system with targetless online calibration achieves real-time VRU tracking on a single Jetson, but its robustness claims are only partially supported by the experiments.
-
Efficient Perception in Automotive Detection and Tracking Using Neuromorphic Computing
Transfer-learned SpikeYOLO achieves mAP 0.937/0.771 and HOTA 0.701/0.445 on KITTI and BDD100K for two-class automotive detection and tracking, competitive with conventional deep networks.
-
A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog
Fog degrades UAV detection and tracking mainly via missed detections; fog-inclusive training is more robust than test-time dehazing, and restoration quality does not proportionally improve downstream perception.
-
GenTrack: A New Generation of Multi-Object Tracking
GenTrack combines particle filters, PSO guidance, and Hungarian data association with social interaction terms, reporting fewer ID switches than SORT, DeepSORT, ByteTrack, BoT-SORT, OC-SORT, SMILEtrack, and ConfTrack ...
-
From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
A survey organizing recent camera-based AI methods for vulnerable road user safety into four interlocking visual tasks and four open deployment challenges.
-
Motion Estimation for Multi-Object Tracking using KalmanNet with Semantic-Independent Encoding
SIKNet, a learned Kalman filter with a semantic-independent encoder, improves bounding-box motion prediction in multi-object tracking over the classic Kalman filter and prior KalmanNet variants.
-
SoccerNet 2025 Challenges Results
The SoccerNet 2025 challenges benchmarked four football video understanding tasks, and top submissions beat the organizer baselines by large margins.
-
MeMoSORT: Memory-Assisted Filtering and Motion-Adaptive Association Metric for Multi-Person Tracking
MeMoSORT reports state-of-the-art multi-person tracking HOTA of 67.9% on DanceTrack and 82.1% on SportsMOT by combining a memory-compensated Kalman filter with a motion-adaptive IoU association metric.
-
Head Anchor Enhanced Detection and Association for Crowded Pedestrian Tracking
FocusTrack, a YOLOX-based tracker using head keypoints and detector features, matches but does not beat simple baselines on MOT17 and MOT20, and its 3D trajectory completion underperforms 2D interpolation.
-
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.
-
SportMamba: Adaptive Non-Linear Multi-Object Tracking with State Space Models for Team Sports
A Mamba-plus-attention motion predictor with a height-adaptive IoU matching metric achieves state-of-the-art HOTA on SportsMOT and strong zero-shot results on VIP-HTD.
-
Depth-Aware Scoring and Hierarchical Alignment for Multiple Object Tracking
A training-free MOT framework that adds zero-shot depth histograms and a hierarchical box/mask alignment score to association, with mixed state-of-the-art results.
-
EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
EWMBench is a benchmark dataset and evaluation toolkit for scoring embodied world models on scene consistency, motion correctness, and semantic alignment, applied to seven video generation models.
-
Improving trajectory continuity in drone-based crowd monitoring using a set of minimal-cost techniques and deep discriminative correlation filters
A point-oriented SORT variant with camera-motion compensation, altitude-aware assignment, trajectory classification, and reused-feature correlation filters reduces trajectory counting error to 15% on UP-COUNT-TRACK an...
-
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
HaWoR estimates metric world-space hand trajectories from egocentric video by masking hands from SLAM bundle adjustment, aligning SLAM scale with Metric3D depth, and infilling missing hand frames with a transformer.
-
Corn Ear Detection and Orientation Estimation Using Deep Learning
A YOLOv8-based pipeline with keypoints and multi-view triangulation estimates 3D corn ear orientation from field video with 18 degree mean absolute error, close to the 12 to 15 degree spread between human measurers.
-
Enhanced Multi-Object Tracking Using Pose-based Virtual Markers in 3x3 Basketball
A pose-based virtual marker overlay, applied to test videos, tracks 3x3 basketball players with zero ID switches and a 72.6 HOTA on a private dataset, but the markers supply identity at test time.
-
Real-Time AIoT for AAV Antenna Interference Detection via Edge-Cloud Collaboration
An edge-cloud UAV inspection system with a new lightweight detector (EdgeAnt) and keyframe selection reduces end-to-end latency by 88.9% and reports 42.1% mAP on a custom antenna dataset.
-
SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory
A training-free adaptation of SAM 2 that adds Kalman-filter motion scoring and motion-aware memory selection improves zero-shot visual tracking across multiple benchmarks.
-
GTA: Global Tracklet Association for Multi-Object Tracking in Sports
GTA, a post-processing module, splits and merges player tracklets using ReID features and clustering, improving HOTA scores on SportsMOT and SoccerNet across three trackers.
-
Glance-MCMT: A General MCMT Framework with Glance Initialization and Progressive Association
Glance-MCMT combines BoT-SORT single-camera tracks, a short glance phase to seed global IDs, and progressive cross-view association, reaching 51.34 HOTA on AI City 2025 validation data.
-
VisionGuard: Synergistic Framework for Helmet Violation Detection
A two-module post-processing framework (tracking-based relabeling plus virtual bounding box injection) reports +1.6% to +3.1% mAP@50 on helmet violation detection, but only on a self-annotated test set with fitted con...
-
YASMOT: Yet another stereo image multi-object tracker
A lightweight detection-only multi-object tracker uses Gaussian distance plus Hungarian matching to track objects over time, link stereo views, and combine ensemble detector output.
Discussion (0). Continue with ORCID to comment.