Pith. sign in

REVIEW 3 major objections 5 minor 300 references

Vision Technologies with Applications in Traffic Surveillance Systems: A Holistic Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A unified survey maps traffic-surveillance vision from pixels to behavior, and argues that five persistent gaps can be closed by pairing today's methods with foundation models.

desk verdict Useful survey with a clear taxonomy and broad coverage; the five-limitation roadmap is asserted without a stated selection protocol, so treat it as informed opinion rather than a proven map. read the letter →

arxiv 2412.00348 v2 pith:Z5D6G7JZ submitted 2024-11-30 cs.CV

classification cs.CV
keywords trafficsurveillancesystemscomputervisionfoundationmodelsintelligenttransportationobjectdetectionanomalybehaviorunderstandingsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Vision-based traffic surveillance is usually reviewed task by task, leaving the field fragmented. This paper tries to establish a single analytical framework that runs from low-level perception (detecting, classifying, and tracking vehicles) up to high-level perception (estimating speeds, spotting anomalies, and understanding behavior). It argues that the field is held back by five fundamental limitations—degraded image quality under real conditions, heavy reliance on labeled data, weak semantic understanding, limited camera coverage, and high compute cost—and that five families of approaches address them. The payoff, if the framework holds, is a shared map for researchers and a concrete case for betting on foundation models as the transformative next step. The paper is a review: its contribution is the organizing framework and roadmap, not a new experiment.

What carries the argument

The organizing device is a two-tier perception framework: a low-level layer that extracts where things are and what they are, and a high-level layer that interprets what is happening and what will happen next. Around that spine the survey builds a second mapping that pairs each of the five stated limitations with a solution category, and then overlays foundation models (large language, vision, and vision-language models plus world models) as a cross-cutting lever. That three-part structure—task hierarchy, limitation-to-solution map, and foundation-model overlay—is what carries the survey's argument, giving it a way to place any individual method and to justify the claim that semantic gaps and data constraints are the bottlenecks foundation models are best suited to break.

What would settle it

Find one production-scale traffic surveillance system whose dominant failure mode falls outside all five categories—for example, a system brought down mainly by a privacy regulation, a cyberattack, or a vendor lock-in constraint. Documenting such a case would show that the five limitations are not fundamental, only common. Conversely, the claim would be supported by a systematic study in which reported field failures map cleanly onto the five categories.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that the scattered literature on vision technologies in traffic surveillance systems can be read as one coherent pipeline with two levels: low-level perception (2D/3D detection, vehicle classification and re-identification, single- and multi-object tracking) and high-level perception (camera calibration, speed estimation, vehicle counting, anomaly detection, and behavior understanding). Surveying representative methods and benchmarks for each, the authors conclude that five limitations are fundamental—perceptual data degradation, data-driven learning constraints, semantic understanding gaps, sensing coverage limitations, and computational resource demands—and map each to a solution family: perception enhancement, efficient learning paradigms, knowledge-enhanced understanding, cooperative sensing, and efficient computing. They then argue that foundation models, with zero-shot learning, open-vocabulary detection, visual question answering, and world-model scene generation, are the most promising single direction for closing the semantic and data gaps. If read sympathetically, the paper's contribution is to reorganize the field so that future work can be positioned and compared within one roadmap.

Load-bearing premise

The roadmap's completeness rests on the assumption that the papers the authors chose to review fairly represent the whole field, so the list of five fundamental limitations is genuinely exhaustive.

Editorial extensions

If this is right

  • A researcher can now locate any TSS method on a shared map and see which of the five limitations it addresses, which the paper argues was previously hard because reviews were fragmented.
  • If the five limitations are fundamental, then progress on the five solution families—especially foundation-model-driven data efficiency and reasoning—should improve all high-level tasks at once, not just one benchmark.
  • The paper's performance tables imply that detection, classification, and re-identification are near practical maturity (accuracy over 90 percent on several benchmarks), so the next gains should come from semantic understanding and deployment efficiency.
  • Foundation models are predicted to move TSS from reactive detection to anticipatory reasoning, e.g., describing safety-critical events in natural language and using world models to synthesize rare accident scenarios for training.
  • Standardized, open benchmarks for speed estimation and behavior understanding are singled out as a necessary condition for fair comparison and generalization claims.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I infer that the framework's real test is completeness: if a future system's main failure is privacy compliance or adversarial tampering—neither of which appears in the five limitations—the map would need a sixth or seventh slot.
  • The survey's own evidence suggests that benchmark scarcity, not algorithm design, is the binding constraint: several tables report high accuracies on existing datasets but the text repeatedly flags lack of standardized TSS benchmarks for speed estimation.
  • A testable extension would be to measure, via a bibliometric or empirical study, how often published TSS failures trace to each of the five categories; the framework predicts those five account for nearly all reported bottlenecks.
  • Another extension the paper leaves implicit: the limitation-to-solution pairing could be turned into a scorecard for selecting between cooperative sensing and efficient computing when deploying a TSS under budget, trading coverage against latency.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript is a survey of vision technologies for traffic surveillance systems (TSS). It organizes the field into low-level perception tasks (2D/3D detection, classification including vehicle model recognition and Re-ID, and single/multi-object tracking) and high-level perception tasks (traffic parameter estimation, anomaly detection, and behavior understanding). For each task it provides a methodological taxonomy, dataset descriptions, and tables of representative reported performances. The paper then claims to identify five fundamental limitations of current TSS—perceptual data degradation, data-driven learning constraints, semantic understanding gaps, sensing coverage limitations, and computational resource demands—and maps each to a category of current solutions and future trends. A final section argues that foundation models, including vision-language models and world models, offer a transformative path toward data-efficient, semantically grounded TSS. The paper contains no new empirical results; its contributions are the organizational framework, the literature compilation, and the proposed roadmap.

Significance. If accepted as a map of the field, the survey has real utility: the low-level/high-level distinction is sensible, the dataset tables in Section 3.4 and Section 4.4 are useful reference material, and the explicit pairing of limitations with solution categories in Section 5.2 gives practitioners a structured entry point. The authors also deserve credit for acknowledging several benchmarking weaknesses themselves, notably in speed estimation (§4.1.2) and in the synthetic-to-real gap (§5.2.2). However, the central forward-looking claim—that the five limitations of Section 5.1 are fundamental and complete—is asserted rather than demonstrated, and the performance tables compare numbers across incompatible benchmarks without adequate caveats. Because the paper's roadmap is built one-to-one on this five-item list, these issues affect the core contribution and require substantive revision rather than cosmetic correction.

major comments (3)
  1. [§5.1] The step from 'issues observed in the surveyed papers' to 'five fundamental limitations persist in TSS' is load-bearing but not methodologically justified. The manuscript never states its search strategy, inclusion/exclusion criteria, or method-selection rationale, so the completeness, independence, and granularity of the five limitations cannot be checked. For example, privacy/security constraints are mentioned only in passing under limitation (b), the evolving-normality drift in unsupervised anomaly detection is discussed in §4.2.2 but not elevated to a fundamental limitation, and sim-to-real transfer appears as a sub-issue in §5.2.2. Since Section 5.2's solution map and the foundation-model outlook in Section 5.3 mirror these five items one-to-one, an incomplete or skewed list propagates through the entire roadmap. I recommend adding a short methodology subsection that states the review protocol and inclusion criteria, and either justifying the exhaustiveness of the five-item list or softening the 'fundamental' claim to 'five challenges emphasized in the reviewed literature.'
  2. [§3.4, Table 2] Table 2 presents performance numbers from mutually incompatible benchmarks as though they support cross-method conclusions. Rows for 2D detection use SEU_PML, VisDrone-DET, BIT-Vehicle, and UA-DETRAC; 3D detection rows mostly use Rope3D but with different AP3D protocols; SOT rows use OTB2015 and VOT2018/2019; MOT rows use MOT16 and MOT17. Without a benchmark/protocol column and without uncertainty estimates, the narrative in §3.4 ('clear progression from two-stage methods toward more efficient one-stage approaches', 'direct estimation methods significantly outperform geometric approaches') is not supported by the numbers as displayed. The same issue affects Table 4, where speed-estimation rows are mostly 'Proprietary' datasets and yet the text infers historical improvement ('from early implementations (3-7 km/h errors) to recent methods (below 2-3 km/h)'). At minimum, add an explicit column for dataset/protocol in each table, state that numbers are not directly comparable across rows, and restrict comparative statements to rows evaluated on the same benchmark.
  3. [§4.3.1 and Table 4] There is a reference-identity error that undermines a specific performance entry. Reference [7] is the survey by Santhosh et al. (ACM Computing Surveys, 2020), but in §4.3.1 the text says '[7] developing a CNN-VAE architecture' and Table 4 lists 'Hybrid CNN-VAE [7] 2021 T15: ACC=99.0%; QMUL: ACC=97.3%'. A survey paper is not the method's source. The correct reference for the CNN-VAE trajectory-anomaly method appears to be Santhosh et al. 2021 (which is otherwise missing from the table). This needs correction because the anomaly-detection performance comparison currently attributes a benchmark result to the wrong publication.
minor comments (5)
  1. [§3.2.1] The opening sentence of §3.2.1 is duplicated verbatim from §3.2: 'Classification in TSS extends beyond basic categorization to fine-grained vehicle model recognition and cross-camera vehicle re-identification (Re-ID), as shown in Figure 4.' One of the two occurrences should be removed.
  2. [Figure 6 caption] The caption of Figure 6 lists items as '(c) virtual section-based speed estimation methods; (b) homography transformation-based speed estimation methods; (e) detection and tracking-based vehicle counting methods and (f) direct regression-based vehicle counting methods.' The second '(b)' should be '(d)'.
  3. [Table 2] The VeRi row for Shen et al. [55] reports 'mAP=80.3%%' with a doubled percent sign; this should be fixed.
  4. [§5.3] The foundation-model section is largely programmatic and would benefit from one or two concrete, critical case studies that show both a successful TSS adaptation and a documented failure mode (for example, hallucination in traffic QA or poor fine-grained localization), rather than citing capabilities only in the positive direction.
  5. [References] Several in-text mentions lack citations or have informal labels: 'ChatGPT 3.5' in the Introduction and Section 5.3 has no reference, and 'YOLO V11' in Table 2 does not appear to have a corresponding numbered reference in the bibliography.

Circularity Check

0 steps flagged · score 0.0 of 10

Survey contains no derivation chain; the five-limitation taxonomy is an asserted synthesis, so no circularity found.

full rationale

This is a literature survey with no fitted parameters, no quantitative predictions, and no theorems. The closest thing to a central claim is the five-item limitation list in Section 5.1 (a)-(e), but it is presented as 'Building upon the comprehensive analysis of strengths and limitations of existing methods across different tasks' — an inductive synthesis of the reviewed literature, not a formal derivation. The solution map in Section 5.2 is organized in parallel to the five limitations, but the paper never claims the limitations are defined by the solutions or that the solutions are derived from them; it is a taxonomic arrangement of existing research areas. The authors' own prior works (e.g., refs [1, 14, 16, 30, 145, 157, 231, 285]) are used as representative examples of particular methods and limitations, but no load-bearing argument reduces to those citations; the limitations are supported by the general body of reviewed work. No equation-level equivalence, no fitted-parameter-as-prediction pattern, and no uniqueness-imported-from-self-citation pattern is present. The absence of a search protocol or exhaustiveness criterion for the limitation list is a methodological limitation of the survey, not a circularity, because the list is not claimed to be the output of a formal derivation from stated premises. Therefore, no significant circularity is found.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central survey rests on domain assumptions about sensor priority, exhaustiveness of its limitation list, and transferability of foundation model capabilities. None are mathematically axiomatic; they are framing choices.

assumptions (3)
  • domain assumption Cameras are the predominant sensing choice for TSS and vision is the foundation of traffic perception.
    Section 1 states this directly; the entire survey scope excludes radar, LiDAR, and inductive loops as primary inputs, so conclusions about limitations depend on this choice.
  • domain assumption The five limitations in Section 5.1 are fundamental and comprehensive for current TSS.
    Stated as 'five fundamental limitations' without a systematic search or formal analysis proving exhaustiveness; it is a synthesis judgment.
  • domain assumption Foundation models possess zero-shot learning, strong generalization, and reasoning capabilities that will transfer to traffic surveillance.
    Section 5.3 relies on general capabilities of SAM, CLIP, and LLMs from outside TSS and projects them onto TSS tasks without benchmark validation inside the domain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Vision Technologies with Applications in Traffic Surveillance Systems: A Holistic Survey." pith.science (2026). https://pith.science/paper/Z5D6G7JZ

@misc{pith2026241200348,
  author       = {Pith},
  title        = {Pith review of: Vision Technologies with Applications in Traffic Surveillance Systems: A Holistic Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z5D6G7JZ}},
  note         = {Machine review of arXiv:2412.00348}
}
read the original abstract

Traffic Surveillance Systems (TSS) have become increasingly crucial in modern intelligent transportation systems, with vision technologies playing a central role for scene perception and understanding. While existing surveys typically focus on isolated aspects of TSS, a comprehensive analytical framework bridging low-level and high-level perception tasks, particularly considering emerging technologies, remains lacking. This paper presents a systematic review of vision technologies in TSS, examining both low-level perception tasks (object detection, classification, and tracking) and high-level perception tasks (parameter estimation, anomaly detection, and behavior understanding). Specifically, we first provide a detailed methodological categorization and comprehensive performance evaluation for each task. Our investigation reveals five fundamental limitations in current TSS: perceptual data degradation in complex scenarios, data-driven learning constraints, semantic understanding gaps, sensing coverage limitations and computational resource demands. To address these challenges, we systematically analyze five categories of current approaches and potential trends: advanced perception enhancement, efficient learning paradigms, knowledge-enhanced understanding, cooperative sensing frameworks and efficient computing frameworks, critically assessing their real-world applicability. Furthermore, we evaluate the transformative potential of foundation models in TSS, which exhibit remarkable zero-shot learning abilities, strong generalization, and sophisticated reasoning capabilities across diverse tasks. This review provides a unified analytical framework bridging low-level and high-level perception tasks, systematically analyzes current limitations and solutions, and presents a structured roadmap for integrating emerging technologies, particularly foundation models, to enhance TSS capabilities.

Figures

Figures reproduced from arXiv: 2412.00348 by the authors.

Figure 1
Figure 1. Overview of vision-based TSS: core components and future prospects [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Evolution and categorization of mainstream methods for 2D/3D detection [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Depth-based methods fall short in accurately detecting vehicles that are either moving at high speeds or are situated far from the camera. In [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Schematic diagram of (a) vehicle model recognition; (b) cross-camera vehicle re-identification; (c) categorization of vehicle Re-ID techniques. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Pipeline and timeline of methodological development for (a) Correlation filter-based SOT methods; (b) Siamese network-based SOT methods; (c) [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Schematic diagram of (a) vanishing point-based camera calibration methods; (b) vehicle keypoint-based camera calibration methods; (c) virtual [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Categorization and development timeline of current traffic anomaly detection (TAD), which includes weakly supervised and unsupervised learning [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Two-stage nature of Unsupervised Traffic Anomaly Detection (UTAD), which learns normal patterns at training stage and detects anomalies at [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Categorization and literature of Vehicle Behavior Understanding (VBU) and Vulnerable Road User Behavior Understanding (VRBU). [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Overview of challenges, solutions, and potential trends for vision technologies in traffic surveillance [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 59 canonical work pages

  1. [7]

    Kelathodi Kumaran Santhosh, Debi Prosad Dogra, and Partha Pratim Roy. 2020. Anomaly detection in road traffic using visual surveillance: A survey. ACM Computing Surveys (CSUR) 53, 6 (2020), 1–26

  2. [1]

    Wei Zhou, Yuqing Liu, Lei Zhao, Sixuan Xu, and Chen Wang. 2023. Pedestrian crossing intention prediction from surveillance videos for over-the-horizon safety warning. IEEE Transactions on Intelligent Transportation Systems (2023)

  3. [2]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770–778

  4. [3]

    Dosovitskiy Alexey. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv: 2010.11929 (2020)

  5. [4]

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. 2023. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4015–4026

  6. [5]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al . 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763

  7. [6]

    Sokemi Rene Emmanuel Datondji, Yohan Dupuis, Peggy Subirats, and Pascal Vasseur. 2016. A survey of vision-based traffic monitoring of road intersections. IEEE transactions on intelligent transportation systems 17, 10 (2016), 2681–2698. 26 Wei Zhou et al

  8. [8]

    Azzedine Boukerche and Zhijun Hou. 2021. Object detection using deep learning methods in traffic scenarios. ACM Computing Surveys (CSUR) 54, 2 (2021), 1–35

Show all 300 references
  1. [9]

    Jorge E Espinosa, Sergio A Velastín, and John W Branch. 2020. Detection of motorcycles in urban traffic using video analysis: A review. IEEE Transactions on Intelligent Transportation Systems 22, 10 (2020), 6115–6130

  2. [10]

    Chenghuan Liu, Du Q Huynh, Yuchao Sun, Mark Reynolds, and Steve Atkinson. 2020. A vision-based pipeline for vehicle counting, speed estimation, and classification. IEEE transactions on intelligent transportation systems 22, 12 (2020), 7547–7560

  3. [11]

    Xingchen Zhang, Yuxiang Feng, Panagiotis Angeloudis, and Yiannis Demiris. 2022. Monocular visual traffic surveillance: A review. IEEE Transactions on Intelligent Transportation Systems 23, 9 (2022), 14148–14165

  4. [12]

    Hadi Ghahremannezhad, Hang Shi, and Chengjun Liu. 2023. Object detection in traffic videos: A survey. IEEE Transactions on Intelligent Transportation Systems 24, 7 (2023), 6780–6799

  5. [13]

    Diego M Jiménez-Bravo, Álvaro Lozano Murciego, André Sales Mendes, Héctor Sánchez San Blás, and Javier Bajo. 2022. Multi-object tracking in traffic environments: A systematic literature review. Neurocomputing 494 (2022), 43–55

  6. [14]

    Wei Zhou, Yuqing Liu, Chen Wang, Yunfei Zhan, Yulu Dai, and Ruiyu Wang. 2022. An automated learning framework with limited and cross-domain data for traffic equipment detection from surveillance videos. IEEE Transactions on Intelligent Transportation Systems 23, 12 (2022), 24891–24903

  7. [15]

    Zhishan Li, Hongxu Chen, Battista Biggio, Yifan He, Haoran Cai, Fabio Roli, and Lei Xie. 2024. Toward effective traffic sign detection via two-stage fusion neural networks. IEEE Transactions on Intelligent Transportation Systems (2024)

  8. [16]

    Wei Zhou, Chen Wang, Jingxin Xia, Zhendong Qian, and Yuan Wu. 2023. Monitoring-based traffic participant detection in urban mixed traffic: A novel dataset and a tailored detector. IEEE Transactions on Intelligent Transportation Systems (2023)

  9. [17]

    Li Kang, Zhiwei Lu, Lingyu Meng, and Zhijian Gao. 2024. YOLO-FA: Type-1 fuzzy attention based YOLO detector for vehicle detection. Expert Systems with Applications 237 (2024), 121209

  10. [18]

    Jia-wei Liu, Da Yang, Ting-wei Feng, and Jun-jie Fu. 2024. MDFD2-DETR: A Real-Time Complex Road Object Detection Model Based on Multi-Domain Feature Decomposition and De-Redundancy. IEEE Transactions on Intelligent Vehicles (2024)

  11. [19]

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence 39, 6 (2016), 1137–1149

  12. [20]

    Zhaowei Cai and Nuno Vasconcelos. 2018. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. 6154–6162

  13. [21]

    Zhengxia Zou, Keyan Chen, Zhenwei Shi, Yuhong Guo, and Jieping Ye. 2023. Object detection in 20 years: A survey. Proc. IEEE 111, 3 (2023), 257–276

  14. [22]

    J Redmon. 2016. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition

  15. [23]

    Joseph Redmon. 2018. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 (2018)

  16. [24]

    Ao Wang, Hui Chen, Lihao Liu, Kai Chen, Zijia Lin, Jungong Han, and Guiguang Ding. 2024. Yolov10: Real-time end-to-end object detection. arXiv preprint arXiv:2405.14458 (2024)

  17. [25]

    Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. 2016. Ssd: Single shot multibox detector. InComputer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I ...

  18. [26]

    Z Tian, C Shen, H Chen, and T He. 2019. FCOS: Fully convolutional one-stage object detection. arXiv 2019. arXiv preprint arXiv:1904.01355 (2019)

  19. [27]

    Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl. 2019. Objects as points. arXiv preprint arXiv:1904.07850 (2019)

  20. [28]

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In European conference on computer vision . Springer, 213–229

  21. [29]

    Yian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei, Guanzhong Wang, Qingqing Dang, Yi Liu, and Jie Chen. 2024. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16965–16974

  22. [30]

    Wei Zhou, Chen Wang, Yiran Ge, Longhui Wen, and Yunfei Zhan. 2023. All-day vehicle detection from surveillance videos based on illumination-adjustable generative adversarial network. IEEE Transactions on Intelligent Transportation Systems 25, 5 (2023), 3326–3340

  23. [31]

    Wei Zhou, Lei Zhao, Hongpu Huang, and Chen Wang. 2025. Data-Efficient Object Detection on Construction Sites Using Reweighting Mechanism and Cross-Batch Contrastive Learning. IEEE Transactions on Industrial Informatics (2025)

  24. [32]

    Matthijs H Zwemer, D Scholte, Rob GJ Wijnhoven, and Peter HN de With. 2022. 3D Detection of Vehicles from 2D Images in Traffic Surveillance.. In VISIGRAPP (5: VISAPP). 97–106

  25. [33]

    Markéta Dubská, Adam Herout, and Jakub Sochor. 2014. Automatic camera calibration for traffic understanding.. In BMVC, Vol. 4. 8

  26. [34]

    Viktor Kocur and Milan Ftáčnik. 2020. Detection of 3D bounding boxes of vehicles using perspective transformation for accurate speed measurement. Machine Vision and Applications 31, 7 (2020), 62

  27. [35]

    Yiqiang Chen, Feng Liu, and Ke Pei. 2022. Monocular vehicle 3d bounding box estimation using homograhy and geometry in traffic scene. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1995–1999

  28. [36]

    Peixuan Li and Huaici Zhao. 2021. Monocular 3d detection with geometric constraint embedding and semi-supervised training. IEEE Robotics and Automation Letters 6, 3 (2021), 5565–5572

  29. [37]

    Xiaoqing Ye, Mao Shu, Hanyu Li, Yifeng Shi, Yingying Li, Guangjie Wang, Xiao Tan, and Errui Ding. 2022. Rope3d: The roadside perception dataset for autonomous driving and monocular 3d object detection task. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patte...

  30. [38]

    Garrick Brazil and Xiaoming Liu. 2019. M3d-rpn: Monocular 3d region proposal network for object detection. In Proceedings of the IEEE/CVF international conference on computer vision. 9287–9296

  31. [39]

    Xinzhu Ma, Yinmin Zhang, Dan Xu, Dongzhan Zhou, Shuai Yi, Haojie Li, and Wanli Ouyang. 2021. Delving into localization errors for monocular 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4721–4730

  32. [40]

    Yunpeng Zhang, Jiwen Lu, and Jie Zhou. 2021. Objects are different: Flexible monocular 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3289–3298. Vision Technologies with Applications in Traffic Surveillance Systems: A...

  33. [41]

    Lei Yang, Kaicheng Yu, Tao Tang, Jun Li, Kun Yuan, Li Wang, Xinyu Zhang, and Peng Chen. 2023. Bevheight: A robust framework for vision-based roadside 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 21611–21620

  34. [42]

    Jia Jinrang, Zhenjia Li, and Yifeng Shi. 2024. MonoUNI: A unified vehicle and infrastructure-side monocular 3d object detection network with sufficient depth clues. Advances in Neural Information Processing Systems 36 (2024)

  35. [43]

    Yuanchang Ou, Huicheng Zheng, Shuyue Chen, and Jiangtao Chen. 2014. Vehicle logo recognition based on a weighted spatial pyramid framework. In 17th International IEEE Conference on Intelligent Transportation Systems (ITSC) . IEEE, 1238–1244

  36. [44]

    Li-Chih Chen, Jun-Wei Hsieh, Yilin Yan, and Duan-Yu Chen. 2015. Vehicle make and model recognition using sparse representation and symmetrical SURFs. Pattern Recognition 48, 6 (2015), 1979–1998

  37. [45]

    Ye Yu, Jun Wang, Jingting Lu, Yang Xie, and Zhenxing Nie. 2018. Vehicle logo recognition based on overlapping enhanced patterns of oriented edge magnitudes. Computers & Electrical Engineering 71 (2018), 273–283

  38. [46]

    Yue Huang, Ruiwen Wu, Ye Sun, Wei Wang, and Xinghao Ding. 2015. Vehicle logo recognition system based on convolutional neural networks with a pretraining strategy. IEEE Transactions on Intelligent Transportation Systems 16, 4 (2015), 1951–1960

  39. [47]

    Foo Chong Soon, Hui Ying Khaw, Joon Huang Chuah, and Jeevan Kanesan. 2018. Hyper-parameters optimisation of deep CNN architecture for vehicle logo recognition. IET Intelligent Transport Systems 12, 8 (2018), 939–946

  40. [48]

    Yang Li, Doudou Zhang, and Jianli Xiao. 2024. A New Method for Vehicle Logo Recognition Based on Swin Transformer. arXiv preprint arXiv:2401.15458 (2024)

  41. [49]

    Ye Yu, Hua Li, Jun Wang, Hai Min, Wei Jia, Jun Yu, and Changwen Chen. 2020. A multilayer pyramid network based on learning for vehicle logo recognition. IEEE Transactions on Intelligent Transportation Systems 22, 5 (2020), 3123–3134

  42. [50]

    Yuqi Li, Yanghao Li, Hongfei Yan, and Jiaying Liu. 2017. Deep joint discriminative learning for vehicle re-identification and retrieval. In2017 IEEE international conference on image processing (ICIP). IEEE, 395–399

  43. [51]

    Yiheng Zhang, Dong Liu, and Zheng-Jun Zha. 2017. Improving triplet-wise training of convolutional neural network for vehicle re-identification. In 2017 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 1386–1391

  44. [52]

    Xiaobin Liu, Shiliang Zhang, Qingming Huang, and Wen Gao. 2018. Ram: a region-aware deep model for vehicle re-identification. In 2018 IEEE international conference on multimedia and expo (ICME) . IEEE, 1–6

  45. [53]

    Fuxiang Huang, Xuefeng Lv, and Lei Zhang. 2023. Coarse-to-fine sparse self-attention for vehicle re-identification. Knowledge-Based Systems 270 (2023), 110526

  46. [54]

    Jiawei Lian, Da-Han Wang, Yun Wu, and Shunzhi Zhu. 2023. Multi-Branch Enhanced Discriminative Network for Vehicle Re-Identification. IEEE Transactions on Intelligent Transportation Systems (2023)

  47. [55]

    Fei Shen, Yi Xie, Jianqing Zhu, Xiaobin Zhu, and Huanqiang Zeng. 2023. Git: Graph interactive transformer for vehicle re-identification. IEEE Transactions on Image Processing 32 (2023), 1039–1051

  48. [56]

    Ruihang Chu, Yifan Sun, Yadong Li, Zheng Liu, Chi Zhang, and Yichen Wei. 2019. Vehicle re-identification with viewpoint-aware metric learning. InProceedings of the IEEE/CVF international conference on computer vision . 8282–8291

  49. [57]

    Pirazh Khorramshahi, Amit Kumar, Neehar Peri, Sai Saketh Rambhatla, Jun-Cheng Chen, and Rama Chellappa. 2019. A dual-path model with adaptive attention for vehicle re-identification. In Proceedings of the IEEE/CVF international conference on computer vision . 6132–6141

  50. [58]

    Rodolfo Quispe, Cuiling Lan, Wenjun Zeng, and Helio Pedrini. 2021. AttributeNet: Attribute enhanced vehicle re-identification. Neurocomputing 465 (2021), 84–92

  51. [59]

    Zhi Yu, Zhiyong Huang, Jiaming Pei, Lamia Tahsin, and Daming Sun. 2023. Semantic-oriented feature coupling transformer for vehicle re-identification in intelligent transportation system. IEEE Transactions on Intelligent Transportation Systems 25, 3 (2023), 2803–2813

  52. [60]

    David S Bolme, J Ross Beveridge, Bruce A Draper, and Yui Man Lui. 2010. Visual object tracking using adaptive correlation filters. In 2010 IEEE computer society conference on computer vision and pattern recognition . IEEE, 2544–2550

  53. [61]

    João F Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista. 2014. High-speed tracking with kernelized correlation filters. IEEE transactions on pattern analysis and machine intelligence 37, 3 (2014), 583–596

  54. [62]

    Luca Bertinetto, Jack Valmadre, Stuart Golodetz, Ondrej Miksik, and Philip HS Torr. 2016. Staple: Complementary learners for real-time tracking. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1401–1409

  55. [63]

    Luca Bertinetto, Jack Valmadre, Joao F Henriques, Andrea Vedaldi, and Philip HS Torr. 2016. Fully-convolutional siamese networks for object tracking. In Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part II 14 . Spr...

  56. [64]

    Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. 2018. High performance visual tracking with siamese region proposal network. In Proceedings of the IEEE conference on computer vision and pattern recognition . 8971–8980

  57. [65]

    Yuechen Yu, Yilei Xiong, Weilin Huang, and Matthew R Scott. 2020. Deformable siamese attention networks for visual object tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 6728–6737

  58. [66]

    Hong Zhang, Wanli Xing, Yifan Yang, Yan Li, and Ding Yuan. 2023. SiamST: Siamese network with spatio-temporal awareness for object tracking. Information Sciences 634 (2023), 122–139

  59. [67]

    Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft. 2016. Simple online and realtime tracking. In 2016 IEEE international conference on image processing (ICIP). IEEE, 3464–3468

  60. [68]

    Nicolai Wojke, Alex Bewley, and Dietrich Paulus. 2017. Simple online and realtime tracking with a deep association metric. In 2017 IEEE international conference on image processing (ICIP). IEEE, 3645–3649

  61. [69]

    Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. 2022. Bytetrack: Multi-object tracking by associating every detection box. In European conference on computer vision . Springer, 1–21

  62. [70]

    Yunhao Du, Zhicheng Zhao, Yang Song, Yanyun Zhao, Fei Su, Tao Gong, and Hongying Meng. 2023. Strongsort: Make deepsort great again.IEEE Transactions on Multimedia 25 (2023), 8725–8737

  63. [71]

    Yu-Hsiang Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Hung-Hin So, and Xin Li. 2024. Smiletrack: Similarity learning for occlusion-aware multiple object tracking. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 5740–5748

  64. [72]

    Zhongdao Wang, Liang Zheng, Yixuan Liu, Yali Li, and Shengjin Wang. 2020. Towards real-time multi-object tracking. In European conference on computer vision . Springer, 107–122. 28 Wei Zhou et al

  65. [73]

    Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. 2021. Fairmot: On the fairness of detection and re-identification in multiple object tracking. International journal of computer vision 129 (2021), 3069–3087

  66. [74]

    Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. 2022. Trackformer: Multi-object tracking with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 8844–8854

  67. [75]

    Ruopeng Gao and Limin Wang. 2023. MeMOTR: Long-term memory-augmented transformer for multi-object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 9901–9910

  68. [76]

    Longyin Wen, Dawei Du, Zhaowei Cai, Zhen Lei, Ming-Ching Chang, Honggang Qi, Jongwoo Lim, Ming-Hsuan Yang, and Siwei Lyu. 2020. UA-DETRAC: A new benchmark and protocol for multi-object detection and tracking. Computer Vision and Image Understanding 193 (2020), 102907

  69. [77]

    Huansheng Song, Haoxiang Liang, Huaiyu Li, Zhe Dai, and Xu Yun. 2019. Vision-based vehicle detection and counting system using deep learning in highway scenes. European Transport Research Review 11, 1 (2019), 1–16

  70. [78]

    Zhiming Luo, Frederic Branchaud-Charron, Carl Lemaire, Janusz Konrad, Shaozi Li, Akshaya Mishra, Andrew Achkar, Justin Eichel, and Pierre-Marc Jodoin. 2018. MIO-TCD: A new benchmark dataset for vehicle classification and localization. IEEE Transactions on Image Processing 27, ...

  71. [79]

    Deng Yongqiang, Wang Dengjiang, Cao Gang, Ma Bing, Guan Xijia, Wang Yajun, Liu Jianchao, Fang Yanming, and Li Juanjuan. 2021. Baai-vanjee roadside dataset: Towards the connected automated vehicle highway technologies in challenging environments of china. arXiv preprint arXiv:2...

  72. [80]

    Huanan Wang, Xinyu Zhang, Zhiwei Li, Jun Li, Kun Wang, Zhu Lei, and Ren Haibing. 2022. Ips300+: a challenging multi-modal data sets for intersection perception system. In 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2539–2545

  73. [81]

    Christian Creß, Walter Zimmer, Leah Strand, Maximilian Fortkord, Siyi Dai, Venkatnarayanan Lakshminarasimhan, and Alois Knoll. 2022. A9-dataset: Multi-sensor infrastructure-based dataset for mobility research. In 2022 IEEE Intelligent Vehicles Symposium (IV) . IEEE, 965–970

  74. [82]

    Haibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo, Zebang Yang, Yifeng Shi, Zhenglong Guo, Hanyu Li, Xing Hu, Jirui Yuan, et al . 2022. Dair-v2x: A large-scale dataset for vehicle-infrastructure cooperative 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Visi...

  75. [83]

    Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 2013. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops . 554–561

  76. [84]

    Linjie Yang, Ping Luo, Chen Change Loy, and Xiaoou Tang. 2015. A large-scale car dataset for fine-grained categorization and verification. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3973–3981

  77. [85]

    Shuo Yang, Chunjuan Bo, Junxing Zhang, Pengxiang Gao, Yujie Li, and Seiichi Serikawa. 2021. VLD-45: A big dataset for vehicle logo recognition and detection. IEEE Transactions on Intelligent Transportation Systems 23, 12 (2021), 25567–25573

  78. [86]

    Hongye Liu, Yonghong Tian, Yaowei Yang, Lu Pang, and Tiejun Huang. 2016. Deep relative distance learning: Tell the difference between similar vehicles. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2167–2175

  79. [87]

    Xinchen Liu, Wu Liu, Tao Mei, and Huadong Ma. 2016. A deep learning-based approach to progressive vehicle re-identification for urban surveillance. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14 ...

  80. [88]

    Zheng Tang, Milind Naphade, Ming-Yu Liu, Xiaodong Yang, Stan Birchfield, Shuo Wang, Ratnesh Kumar, David Anastasiu, and Jenq-Neng Hwang. 2019. Cityflow: A city- scale benchmark for multi-target multi-camera vehicle tracking and re-identification. In Proceedings of the IEEE/CVF...

  81. [89]

    Yan Bai, Jun Liu, Yihang Lou, Ce Wang, and Ling-Yu Duan. 2021. Disentangled feature learning network and a comprehensive benchmark for vehicle re-identification. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 10 (2021), 6854–6871

  82. [90]

    UT Benchmark. 2016. A benchmark and simulator for UAV tracking. In European Conference on Computer Vision

  83. [91]

    Pengfei Zhu, Longyin Wen, Xiao Bian, Haibin Ling, and Qinghua Hu. 2018. Vision meets drones: A challenge. arXiv preprint arXiv:1804.07437 (2018)

  84. [92]

    Patrick Dendorfer, Aljosa Osep, Anton Milan, Konrad Schindler, Daniel Cremers, Ian Reid, Stefan Roth, and Laura Leal-Taixé. 2021. Motchallenge: A benchmark for single-camera multiple target tracking. International Journal of Computer Vision 129 (2021), 845–881

  85. [93]

    Ziying Song, Lin Liu, Feiyang Jia, Yadan Luo, Caiyan Jia, Guoxin Zhang, Lei Yang, and Li Wang. 2024. Robustness-aware 3d object detection in autonomous driving: A review and outlook. IEEE Transactions on Intelligent Transportation Systems (2024)

  86. [94]

    Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. 2013. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research 32, 11 (2013), 1231–1237

  87. [95]

    Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. 2020. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and patt...

  88. [96]

    Ali Amiri, Aydin Kaya, and Ali Seydi Keceli. 2024. A Comprehensive Survey on Deep-Learning-based Vehicle Re-Identification: Models, Data Sets and Challenges. arXiv preprint arXiv:2401.10643 (2024)

  89. [97]

    Wenhan Luo, Junliang Xing, Anton Milan, Xiaoqin Zhang, Wei Liu, and Tae-Kyun Kim. 2021. Multiple object tracking: A literature review. Artificial intelligence 293 (2021), 103448

  90. [98]

    Yu-Jin Zhang. 2023. Camera calibration. In 3-D Computer Vision: Principles, Algorithms and Applications . Springer, 37–65

  91. [99]

    Anup Basu and Kavita Ravi. 1997. Active camera calibration using pan, tilt and roll. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics) 27, 3 (1997), 559–566

  92. [100]

    Zhanfei Chen, Xuelong Si, Dan Wu, Fengnian Tian, Zhenxing Zheng, and Renfu Li. 2024. A novel camera calibration method based on known rotations and translations. Computer Vision and Image Understanding 243 (2024), 103996

  93. [101]

    Tuan Hue Thi, Sijun Lu, and Jian Zhang. 2008. Self-calibration of traffic surveillance camera using motion tracking. In 2008 11th International IEEE Conference on Intelligent Transportation Systems. IEEE, 304–309

  94. [102]

    Yuan Zheng and Silong Peng. 2012. Model based vehicle localization for urban traffic surveillance using image gradient based matching. In 2012 15th International IEEE Conference on Intelligent Transportation Systems . IEEE, 945–950

  95. [103]

    Radu Orghidan, Joaquim Salvi, Mihaela Gordan, and Bogdan Orza. 2012. Camera calibration using two or three vanishing points. In 2012 Federated Conference on Computer science and information systems (FedCSIS) . IEEE, 123–130

  96. [104]

    Zhaoxiang Zhang, Tieniu Tan, Kaiqi Huang, and Yunhong Wang. 2012. Practical camera calibration from moving objects for traffic scene surveillance. IEEE transactions on circuits and systems for video technology 23, 3 (2012), 518–533. Vision Technologies with Applications in Tra...

  97. [105]

    Viktor Kocur and Milan Ftáčnik. 2021. Traffic camera calibration via vehicle vanishing point detection. In Artificial Neural Networks and Machine Learning–ICANN 2021: 30th International Conference on Artificial Neural Networks, Bratislava, Slovakia, September 14–17, 2021, Proc...

  98. [106]

    Wentao Zhang, Huansheng Song, and Lichen Liu. 2023. Automatic calibration for monocular cameras in highway scenes via vehicle vanishing point detection. Journal of transportation engineering, Part A: Systems 149, 7 (2023), 04023050

  99. [107]

    Shusen Guo, Xianwen Yu, Yuejin Sha, Yifan Ju, Mingchen Zhu, and Jiafu Wang. 2024. Online camera auto-calibration appliable to road surveillance. Machine Vision and Applications 35, 4 (2024), 91

  100. [108]

    Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159 (2020)

  101. [109]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision . 10012–10022

  102. [110]

    Jun-Wei Hsieh, Li-Chih Chen, and Duan-Yu Chen. 2014. Symmetrical SURF and its applications to vehicle detection and vehicle make and model recognition. IEEE Transactions on intelligent transportation systems 15, 1 (2014), 6–20

  103. [111]

    Jakub Sochor, Jakub Špaňhel, and Adam Herout. 2018. Boxcars: Improving fine-grained recognition of vehicles using 3-d bounding boxes in traffic surveillance. IEEE transactions on intelligent transportation systems 20, 1 (2018), 97–108

  104. [112]

    Wei Sun, Guoce Zhang, Xiaorui Zhang, Xu Zhang, and Nannan Ge. 2021. Fine-grained vehicle type classification using lightweight convolutional neural network with feature optimization and joint learning strategy. Multimedia Tools and Applications 80, 20 (2021), 30803–30816

  105. [113]

    Azzedine Boukerche and Xiren Ma. 2021. A novel smart lightweight visual attention model for fine-grained vehicle recognition.IEEE Transactions on Intelligent Transportation Systems 23, 8 (2021), 13846–13862

  106. [114]

    Hongchun Lu, Min Han, Chaoqing Wang, and Junlong Cheng. 2023. AMLNet: Attention Multibranch Loss CNN Models for Fine-Grained Vehicle Recognition. IEEE Transactions on Vehicular Technology (2023)

  107. [115]

    Ruilong Chen, Matthew Hawes, Lyudmila Mihaylova, Jingjing Xiao, and Wei Liu. 2016. Vehicle logo recognition by spatial-SIFT combined with logistic regression. In 2016 19th International Conference on Information Fusion (FUSION) . IEEE, 1228–1235

  108. [116]

    Sugang Ma, Zhixian Zhao, Zhiqiang Hou, Lei Zhang, Xiaobao Yang, and Lei Pu. 2022. Correlation filters based on multi-expert and game theory for visual object tracking. IEEE Transactions on Instrumentation and Measurement 71 (2022), 1–14

  109. [117]

    Zedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang, and Rongrong Ji. 2020. Siamese box adaptive network for visual tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 6668–6677

  110. [118]

    Jing Liu, Han Wang, Chao Ma, Yuting Su, and Xiaokang Yang. 2024. Siamdmu: Siamese dual mask update network for visual object tracking. IEEE Transactions on Emerging Topics in Computational Intelligence (2024)

  111. [119]

    Romil Bhardwaj, Gopi Krishna Tummala, Ganesan Ramalingam, Ramachandran Ramjee, and Prasun Sinha. 2018. Autocalib: Automatic traffic camera calibration at scale. ACM Transactions on Sensor Networks (TOSN) 14, 3-4 (2018), 1–27

  112. [120]

    Vojtěch Bartl, Roman Juranek, Jakub Špaňhel, and Adam Herout. 2020. Planecalib: Automatic camera calibration by multiple observations of rigid objects on plane. In 2020 Digital Image Computing: Techniques and Applications (DICTA) . IEEE, 1–8

  113. [121]

    Vojtěch Bartl, Jakub Špaňhel, Petr Dobeš, Roman Juranek, and Adam Herout. 2021. Automatic camera calibration by landmarks on rigid objects. Machine Vision and Applications 32, 1 (2021), 2

  114. [122]

    Turgay Celik and Huseyin Kusetogullari. 2009. Solar-powered automated road surveillance system for speed violation detection. IEEE Transactions on Industrial Electronics 57, 9 (2009), 3216–3227

  115. [123]

    Chomtip Pornpanomchai and Kaweepap Kongkittisan. 2009. Vehicle speed detection system. In 2009 IEEE international conference on signal and image processing applications . IEEE, 135–139

  116. [124]

    Mallikarjun Anandhalli, Pavana Baligar, Santosh S Saraf, and Pooja Deepsir. 2022. Image projection method for vehicle speed estimation model in video system. Machine Vision and Applications 33, 1 (2022), 7

  117. [125]

    Muhammad Hassaan Ashraf, Farhana Jabeen, Hamed Alghamdi, M Sultan Zia, and Mubarak S Almutairi. 2023. HVD-net: a hybrid vehicle detection network for vision- based vehicle tracking and speed estimation. Journal of King Saud University-Computer and Information Sciences 35, 8 (2...

  118. [126]

    Tingting Huang. 2018. Traffic speed estimation from surveillance video data. InProceedings of the IEEE conference on computer vision and pattern recognition workshops. 161–165

  119. [127]

    D Bell, W Xiao, and P James. 2020. Accurate vehicle speed estimation from monocular camera footage. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences 2 (2020), 419–426

  120. [128]

    Ervin Yohannes, Chih-Yang Lin, Timothy K Shih, Tipajin Thaipisutikul, Avirmed Enkhbat, and Fitri Utaminingrum. 2023. An improved speed estimation using deep homography transformation regression network on monocular videos. IEEE Access 11 (2023), 5955–5965

  121. [129]

    Igor Lashkov, Runze Yuan, and Guohui Zhang. 2023. Edge-Computing-Empowered Vehicle Tracking and Speed Estimation Against Strong Image Vibrations Using Surveillance Monocular Camera. IEEE Transactions on Intelligent Transportation Systems (2023)

  122. [130]

    David Fernández Llorca, Antonio Hernández Martínez, and Iván García Daza. 2021. Vision-based vehicle speed estimation: A survey. IET Intelligent Transport Systems 15, 8 (2021), 987–1005

  123. [131]

    Zhe Dai, Huansheng Song, Xuan Wang, Yong Fang, Xu Yun, Zhaoyang Zhang, and Huaiyu Li. 2019. Video-based vehicle counting framework. IEEE Access 7 (2019), 64460–64470

  124. [132]

    Zhongji Liu, Wei Zhang, Xu Gao, Hao Meng, Xiao Tan, Xiaoxing Zhu, Zhan Xue, Xiaoqing Ye, Hongwu Zhang, Shilei Wen, et al. 2020. Robust movement-specific vehicle counting at crowded intersections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  125. [133]

    Mishuk Majumder and Chester Wilmot. 2023. Automated vehicle counting from pre-recorded video using you only look once (YOLO) object detection model. Journal of imaging 9, 7 (2023), 131

  126. [134]

    Muhammad Asif Khan, Hamid Menouar, and Ridha Hamila. 2023. Revisiting crowd counting: State-of-the-art, trends, and future perspectives. Image and Vision Computing 129 (2023), 104597

  127. [135]

    Daniel Onoro-Rubio and Roberto J López-Sastre. 2016. Towards perspective-free object counting with deep learning. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part VII 14 . Springer, 615–629. 30 Wei Zhou et al

  128. [136]

    Shanghang Zhang, Guanhang Wu, Joao P Costeira, and José MF Moura. 2017. Fcn-rlstm: Deep spatio-temporal neural networks for vehicle counting in city cameras. In Proceedings of the IEEE international conference on computer vision . 3667–3676

  129. [137]

    Henglong Yang, Youmei Zhang, Yu Zhang, Hailong Meng, Shuang Li, and Xianglin Dai. 2021. A fast vehicle counting and traffic volume estimation method based on convolutional neural network. IEEe Access 9 (2021), 150522–150531

  130. [138]

    Xiangyu Guo, Mingliang Gao, Wenzhe Zhai, Qilei Li, and Gwanggil Jeon. 2023. Scale region recognition network for object counting in intelligent transportation system. IEEE Transactions on Intelligent Transportation Systems (2023)

  131. [139]

    Yang Liu, Dingkang Yang, Yan Wang, Jing Liu, Jun Liu, Azzedine Boukerche, Peng Sun, and Liang Song. 2024. Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models. Comput. Surveys 56, 7 (2024), 1–38

  132. [140]

    Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy, and Reinhard Klette. 2017. Deep-cascade: Cascading 3d deep neural networks for fast anomaly detection and localization in crowded scenes. IEEE Transactions on Image Processing 26, 4 (2017), 1992–2004

  133. [141]

    Elizaveta Batanina, Imad Eddine Ibrahim Bekkouch, Youssef Youssry, Adil Khan, Asad Masood Khattak, and Mikhail Bortnikov. 2019. Domain adaptation for car accident detection in videos. In 2019 ninth international conference on image processing theory, tools and applications (IP...

  134. [142]

    Zhenbo Lu, Wei Zhou, Shixiang Zhang, and Chen Wang. 2020. A New Video-Based Crash Detection Method: Balancing Speed and Accuracy Using a Feature Fusion Deep Learning Framework. Journal of advanced transportation 2020, 1 (2020), 8848874

  135. [143]

    Jia-Xing Zhong, Nannan Li, Weijie Kong, Shan Liu, Thomas H Li, and Ge Li. 2019. Graph convolutional label noise cleaner: Train a plug-and-play action classifier for anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1237–1246

  136. [144]

    Jia-Chang Feng, Fa-Ting Hong, and Wei-Shi Zheng. 2021. Mist: Multiple instance self-training framework for video anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 14009–14018

  137. [145]

    Wei Zhou, Longhui Wen, Yunfei Zhan, and Chen Wang. 2023. An appearance-motion network for vision-based crash detection: Improving the accuracy in congested traffic. IEEE transactions on intelligent transportation systems (2023)

  138. [146]

    Hongyang Yu, Xinfeng Zhang, Yaowei Wang, Qingming Huang, and Baocai Yin. 2024. Fine-grained accident detection: database and algorithm. IEEE transactions on image processing (2024)

  139. [147]

    Waqas Sultani, Chen Chen, and Mubarak Shah. 2018. Real-world anomaly detection in surveillance videos. In Proceedings of the IEEE conference on computer vision and pattern recognition. 6479–6488

  140. [148]

    Yi Zhu and Shawn Newsam. 2019. Motion-aware feature for improved video anomaly detection. arXiv preprint arXiv:1907.10211 (2019)

  141. [149]

    Muhammad Zaigham Zaheer, Arif Mahmood, Hochul Shin, and Seung-Ik Lee. 2020. A self-reasoning framework for anomaly detection using video-level labels. IEEE Signal Processing Letters 27 (2020), 1705–1709

  142. [150]

    Wenhao Shao, Ruliang Xiao, Praboda Rajapaksha, Mengzhu Wang, Noel Crespi, Zhigang Luo, and Roberto Minerva. 2023. Video anomaly detection with NTCN-ML: A novel TCN for multi-instance learning. Pattern Recognition 143 (2023), 109765

  143. [151]

    Silas SL Pereira and José Everardo Bessa Maia. 2024. MC-MIL: video surveillance anomaly detection with multi-instance learning and multiple overlapped cameras. Neural Computing and Applications 36, 18 (2024), 10527–10543

  144. [152]

    Yang Liu, Jing Liu, Wei Ni, and Liang Song. 2022. Abnormal event detection with self-guiding multi-instance ranking framework. In 2022 International Joint Conference on Neural Networks (IJCNN). IEEE, 01–07

  145. [153]

    Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, and Anton van den Hengel. 2019. Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection. In Proceedings of the IEEE/CVF international...

  146. [154]

    Mahmudul Hasan, Jonghyun Choi, Jan Neumann, Amit K Roy-Chowdhury, and Larry S Davis. 2016. Learning temporal regularity in video sequences. In Proceedings of the IEEE conference on computer vision and pattern recognition . 733–742

  147. [155]

    K Deepak, S Chandrakala, and C Krishna Mohan. 2021. Residual spatiotemporal autoencoder for unsupervised video anomaly detection. Signal, Image and Video Processing 15, 1 (2021), 215–222

  148. [156]

    Kelathodi Kumaran Santhosh, Debi Prosad Dogra, Partha Pratim Roy, and Adway Mitra. 2021. Vehicular trajectory classification and traffic anomaly detection in videos using a hybrid CNN-VAE architecture. IEEE Transactions on Intelligent Transportation Systems 23, 8 (2021), 11891–11902

  149. [157]

    Wei Zhou, Yunhong Yu, Yunfei Zhan, and Chen Wang. 2022. A vision-based abnormal trajectory detection framework for online traffic incident alert on freeways. Neural Computing and Applications 34, 17 (2022), 14945–14958

  150. [158]

    Wen Liu, Weixin Luo, Dongze Lian, and Shenghua Gao. 2018. Future frame prediction for anomaly detection–a new baseline. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6536–6545

  151. [159]

    Zhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang, and Guiqing Li. 2021. A hybrid video anomaly detection framework via memory-augmented flow reconstruction and flow-guided frame prediction. In Proceedings of the IEEE/CVF international conference on computer vision . 13588–13597

  152. [160]

    Xuanzhao Wang, Zhengping Che, Bo Jiang, Ning Xiao, Ke Yang, Jian Tang, Jieping Ye, Jingyu Wang, and Qi Qi. 2021. Robust unsupervised video anomaly detection by multipath frame prediction. IEEE transactions on neural networks and learning systems 33, 6 (2021), 2301–2312

  153. [161]

    Tung Minh Tran, Doanh C Bui, Tam V Nguyen, and Khang Nguyen. 2024. Transformer-based Spatio-Temporal Unsupervised Traffic Anomaly Detection in Aerial Videos. IEEE Transactions on Circuits and Systems for Video Technology (2024)

  154. [162]

    Huan-Sheng Song, Sheng-Nan Lu, Xiang Ma, Yuan Yang, Xue-Qin Liu, and Peng Zhang. 2014. Vehicle behavior analysis using target motion trajectories. IEEE Transactions on Vehicular Technology 63, 8 (2014), 3580–3591

  155. [163]

    Aaron Christian P Uy, Rhen Anjerome Bedruz, Ana Riza Quiros, Argel Bandala, and Elmer P Dadios. 2015. Machine vision for traffic violation detection system through genetic algorithm. In 2015 International Conference on Humanoid, Nanotechnology, Information Technology, Communic...

  156. [164]

    Georges S Aoude, Vishnu R Desaraju, Lauren H Stephens, and Jonathan P How. 2012. Driver behavior classification at intersections and validation on large naturalistic data set. IEEE Transactions on Intelligent Transportation Systems 13, 2 (2012), 724–736

  157. [165]

    Hailun Zhang and Rui Fu. 2021. An ensemble learning–online semi-supervised approach for vehicle behavior recognition. IEEE Transactions on Intelligent Transportation Systems 23, 8 (2021), 10610–10626

  158. [166]

    Da Xu, Mengfei Liu, Xinpeng Yao, and Nengchao Lyu. 2023. Integrating surrounding vehicle information for vehicle trajectory representation and abnormal lane-change behavior detection. Sensors 23, 24 (2023), 9800. Vision Technologies with Applications in Traffic Surveillance Sy...

  159. [167]

    Arya Haghighat and Anuj Sharma. 2023. A computer vision-based deep learning model to detect wrong-way driving using pan–tilt–zoom traffic cameras. Computer-Aided Civil and Infrastructure Engineering 38, 1 (2023), 119–132

  160. [168]

    Guotao Xie, Hongbo Gao, Lijun Qian, Bin Huang, Keqiang Li, and Jianqiang Wang. 2017. Vehicle trajectory prediction by integrating physics-and maneuver-based approaches using interactive multiple models. IEEE Transactions on Industrial Electronics 65, 7 (2017), 5999–6008

  161. [169]

    Cyrus Anderson, Ram Vasudevan, and Matthew Johnson-Roberson. 2021. A kinematic model for trajectory prediction in general highway scenarios. IEEE Robotics and Automation Letters 6, 4 (2021), 6757–6764

  162. [170]

    Nishant Nikhil and Brendan Tran Morris. 2018. Convolutional neural network for trajectory prediction. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops. 0–0

  163. [171]

    Vinit Katariya, Mohammadreza Baharani, Nichole Morris, Omidreza Shoghli, and Hamed Tabkhi. 2022. Deeptrack: Lightweight deep learning for vehicle trajectory prediction in highways. IEEE Transactions on Intelligent Transportation Systems 23, 10 (2022), 18927–18936

  164. [172]

    Lei Lin, Weizi Li, Huikun Bi, and Lingqiao Qin. 2021. Vehicle trajectory prediction using LSTMs with spatial–temporal attention mechanisms. IEEE Intelligent Transportation Systems Magazine 14, 2 (2021), 197–208

  165. [173]

    Yuzhen Zhang, Wentong Wang, Weizhi Guo, Pei Lv, Mingliang Xu, Wei Chen, and Dinesh Manocha. 2022. D2-TPred: Discontinuous dependency for trajectory prediction under traffic lights. In European Conference on Computer Vision . Springer, 522–539

  166. [174]

    Hongyan Guo, Qingyu Meng, Dongpu Cao, Hong Chen, Jun Liu, and Bingxu Shang. 2022. Vehicle trajectory prediction method coupled with ego vehicle motion trend under dual attention mechanism. IEEE Transactions on Instrumentation and Measurement 71 (2022), 1–16

  167. [175]

    Yilong Ren, Zhengxing Lan, Lingshan Liu, and Haiyang Yu. 2024. Emsin: enhanced multi-stream interaction network for vehicle trajectory prediction. IEEE Transactions on Fuzzy Systems (2024)

  168. [176]

    Armin Danesh Pazho, Ghazal Alinezhad Noghre, Vinit Katariya, and Hamed Tabkhi. 2024. VT-Former: An Exploratory Study on Vehicle Trajectory Prediction for Highway Surveillance through Graph Isomorphism and Transformer. In Proceedings of the IEEE/CVF Conference on Computer Visio...

  169. [177]

    Renteng Yuan, Mohamed Abdel-Aty, Qiaojun Xiang, Zijin Wang, and Xin Gu. 2023. A temporal multi-gate mixture-of-experts approach for vehicle trajectory and driving intention prediction. IEEE Transactions on Intelligent Vehicles (2023)

  170. [178]

    Michael Goldhammer, Sebastian Köhler, Stefan Zernetsch, Konrad Doll, Bernhard Sick, and Klaus Dietmayer. 2019. Intentions of vulnerable road users—Detection and forecasting by means of machine learning. IEEE transactions on intelligent transportation systems 21, 7 (2019), 3035–3045

  171. [179]

    Khaled Saleh, Mohammed Hossny, and Saeid Nahavandi. 2018. Intent prediction of pedestrians via motion trajectories using stacked recurrent neural networks. IEEE Transactions on Intelligent Vehicles 3, 4 (2018), 414–424

  172. [180]

    Chenxin Xu, Weibo Mao, Wenjun Zhang, and Siheng Chen. 2022. Remember intentions: Retrospective-memory-based trajectory prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6488–6497

  173. [181]

    Shile Zhang, Mohamed Abdel-Aty, Yina Wu, and Ou Zheng. 2021. Pedestrian crossing intention prediction at red-light using pose estimation. IEEE transactions on intelligent transportation systems 23, 3 (2021), 2331–2339

  174. [182]

    Zhijie Fang and Antonio M López. 2019. Intention recognition of pedestrians and cyclists by 2d pose estimation. IEEE Transactions on Intelligent Transportation Systems 21, 11 (2019), 4773–4783

  175. [183]

    Feiyi Xu, Feng Xu, Jiucheng Xie, Chi-Man Pun, Huimin Lu, and Hao Gao. 2021. Action recognition framework in traffic scene for autonomous driving system. IEEE Transactions on Intelligent Transportation Systems 23, 11 (2021), 22301–22311

  176. [184]

    Amir Rasouli, Iuliia Kotseruba, and John K Tsotsos. 2017. Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior. In Proceedings of the IEEE International Conference on Computer Vision Workshops . 206–213

  177. [185]

    Joseph Gesnouin, Steve Pechberti, Bogdan Stanciulcscu, and Fabien Moutarde. 2021. TrouSPI-Net: Spatio-temporal attention on parallel atrous convolutions and U-GRUs for skeletal pedestrian crossing prediction. In 2021 16th IEEE International Conference on Automatic Face and Ges...

  178. [186]

    Iuliia Kotseruba, Amir Rasouli, and John K Tsotsos. 2021. Benchmark for evaluating pedestrian action prediction. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 1258–1268

  179. [187]

    Mohsen Azarmi, Mahdi Rezaei, He Wang, and Sebastien Glaser. 2024. PIP-Net: Pedestrian Intention Prediction in the Wild. arXiv preprint arXiv:2402.12810 (2024)

  180. [188]

    Chi Zhang and Christian Berger. 2023. Pedestrian behavior prediction using deep learning methods for urban scenarios: A review. IEEE Transactions on Intelligent Transportation Systems 24, 10 (2023), 10279–10301

  181. [189]

    Biao Yang, Zhiwen Wei, Hongyu Hu, Rui Wang, Changchun Yang, and Rongrong Ni. 2023. DPCIAN: A novel dual-channel pedestrian crossing intention anticipation network. IEEE Transactions on Intelligent Transportation Systems (2023)

  182. [190]

    Alexandre Alahi, Vignesh Ramanathan, and Li Fei-Fei. 2014. Socially-aware large-scale crowd forecasting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2203–2210

  183. [191]

    Cao Ningbo, Wei Wei, Qu Zhaowei, Zhao Liying, and Bai Qiaowen. 2017. Simulation of pedestrian crossing behaviors at unmarked roadways based on social force model. Discrete Dynamics in Nature and Society 2017, 1 (2017), 8741534

  184. [192]

    Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. 2016. Social lstm: Human trajectory prediction in crowded spaces. In Proceedings of the IEEE conference on computer vision and pattern recognition . 961–971

  185. [193]

    Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. 2018. Social gan: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2255–2264

  186. [194]

    Liushuai Shi, Le Wang, Sanping Zhou, and Gang Hua. 2023. Trajectory unified transformer for pedestrian trajectory prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 9675–9684

  187. [195]

    Weicheng Zhang, Hao Cheng, Fatema T Johora, and Monika Sester. 2023. ForceFormer: exploring social force and transformer for pedestrian trajectory prediction. In 2023 IEEE Intelligent Vehicles Symposium (IV) . IEEE, 1–7

  188. [196]

    Liushuai Shi, Le Wang, Chengjiang Long, Sanping Zhou, Mo Zhou, Zhenxing Niu, and Gang Hua. 2021. SGCN: Sparse graph convolution network for pedestrian trajectory prediction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 8994–9003

  189. [197]

    Jing Lian, Weiwei Ren, Linhui Li, Yafu Zhou, and Bin Zhou. 2023. Ptp-stgcn: pedestrian trajectory prediction based on a spatio-temporal graph convolutional neural network. Applied Intelligence 53, 3 (2023), 2862–2878. 32 Wei Zhou et al

  190. [198]

    Pei Lv, Wentong Wang, Yunxin Wang, Yuzhen Zhang, Mingliang Xu, and Changsheng Xu. 2023. SSAGCN: social soft attention graph convolution network for pedestrian trajectory prediction. IEEE transactions on neural networks and learning systems (2023)

  191. [199]

    Milind Naphade, Ming-Ching Chang, Anuj Sharma, David C Anastasiu, Vamsi Jagarlamudi, Pranamesh Chakraborty, Tingting Huang, Shuo Wang, Ming-Yu Liu, Rama Chellappa, et al. 2018. The 2018 nvidia ai city challenge. In Proceedings of the IEEE conference on computer vision and patt...

  192. [200]

    Diogo C Luvizon, Bogdan T Nassu, and Rodrigo Minetto. 2014. Vehicle speed estimation by license plate detection and tracking. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6563–6567

  193. [201]

    Timothy Hospedales, Shaogang Gong, and Tao Xiang. 2012. Video behaviour mining using a dynamic topic model.International journal of computer vision 98 (2012), 303–323

  194. [202]

    Ming-Ching Chang, Chen-Kuo Chiang, Chun-Ming Tsai, Yun-Kai Chang, Hsuan-Lun Chiang, Yu-An Wang, Shih-Ya Chang, Yun-Lun Li, Ming-Shuin Tsai, and Hung- Yu Tseng. 2020. Ai city challenge 2020-computer vision for smart transportation applications. In Proceedings of the IEEE/CVF co...

  195. [203]

    Ricardo Guerrero-Gómez-Olmedo, Beatriz Torre-Jiménez, Roberto López-Sastre, Saturnino Maldonado-Bascón, and Daniel Onoro-Rubio. 2015. Extremely overlapping vehicle counting. In Pattern Recognition and Image Analysis: 7th Iberian Conference, IbPRIA 2015, Santiago de Compostela,...

  196. [204]

    Meng-Ru Hsieh, Yen-Liang Lin, and Winston H Hsu. 2017. Drone-based object counting by spatially regularized regional proposal network. In Proceedings of the IEEE international conference on computer vision . 4145–4153

  197. [205]

    Weixin Li, Vijay Mahadevan, and Nuno Vasconcelos. 2013. Anomaly detection and localization in crowded scenes. IEEE transactions on pattern analysis and machine intelligence 36, 1 (2013), 18–32

  198. [206]

    Cewu Lu, Jianping Shi, and Jiaya Jia. 2013. Abnormal event detection at 150 fps in matlab. In Proceedings of the IEEE international conference on computer vision . 2720–2727

  199. [207]

    Ankit Parag Shah, Jean-Bapstite Lamare, Tuan Nguyen-Anh, and Alexander Hauptmann. 2018. CADP: A novel dataset for CCTV traffic camera based accident analysis. In 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (A VSS) . IEEE, 1–9

  200. [208]

    Tung Minh Tran, Tu N Vu, Tam V Nguyen, and Khang Nguyen. 2023. UIT-ADrone: A novel drone dataset for traffic anomaly detection. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 16 (2023), 5590–5601

  201. [209]

    Stefano Pellegrini, Andreas Ess, and Luc Van Gool. 2010. Improving data association by joint modeling of pedestrian trajectories and groupings. In Computer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceeding...

  202. [210]

    Alon Lerner, Yiorgos Chrysanthou, and Dani Lischinski. 2007. Crowds by example. In Computer graphics forum, Vol. 26. Wiley Online Library, 655–664

  203. [211]

    Amir Rasouli, Iuliia Kotseruba, Toni Kunic, and John K Tsotsos. 2019. Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 6262–6271

  204. [212]

    Robert Krajewski, Julian Bock, Laurent Kloeker, and Lutz Eckstein. 2018. The highd dataset: A drone dataset of naturalistic vehicle trajectories on german highways for validation of highly automated driving systems. In 2018 21st international conference on intelligent transpor...

  205. [213]

    Ou Zheng, Mohamed Abdel-Aty, Lishengsa Yue, Amr Abdelraouf, Zijin Wang, and Nada Mahmoud. 2024. CitySim: a drone-based vehicle trajectory dataset for safety- oriented research and digital twins. Transportation research record 2678, 4 (2024), 606–621

  206. [214]

    Xinyu Huang, Xinjing Cheng, Qichuan Geng, Binbin Cao, Dingfu Zhou, Peng Wang, Yuanqing Lin, and Ruigang Yang. 2018. The apolloscape dataset for autonomous driving. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops . 954–960

  207. [215]

    John Houston, Guido Zuidhof, Luca Bergamini, Yawei Ye, Long Chen, Ashesh Jain, Sammy Omari, Vladimir Iglovikov, and Peter Ondruska. 2021. One thousand and one hours: Self-driving motion prediction dataset. In Conference on Robot Learning . PMLR, 409–418

  208. [216]

    Haibao Yu, Wenxian Yang, Hongzhi Ruan, Zhenwei Yang, Yingjuan Tang, Xu Gao, Xin Hao, Yifeng Shi, Yifeng Pan, Ning Sun, et al. 2023. V2x-seq: A large-scale sequential dataset for vehicle-infrastructure cooperative perception and forecasting. In Proceedings of the IEEE/CVF Confe...

  209. [217]

    Shaohua Liu, Haibo Liu, Huikun Bi, and Tianlu Mao. 2020. CoL-GAN: Plausible and collision-less trajectory prediction by attention-based GAN. IEEE Access 8 (2020), 101662–101671

  210. [218]

    Parth Kothari, Sven Kreiss, and Alexandre Alahi. 2021. Human trajectory forecasting in crowds: A deep learning perspective. IEEE Transactions on Intelligent Transportation Systems 23, 7 (2021), 7386–7400

  211. [219]

    Budi Setiyono, Dwi Ratna Sulistyaningrum, Farah Fajriyah, Danang Wahyu Wicaksono, et al . 2017. Vehicle speed detection based on gaussian mixture model using sequential of images. In Journal of Physics: Conference Series , Vol. 890. IOP Publishing, 012144

  212. [220]

    Peichao Cong, Yixuan Xiao, Xianquan Wan, Murong Deng, Jiaxing Li, and Xin Zhang. 2023. DACR-AMTP: adaptive multi-modal vehicle trajectory prediction for dynamic drivable areas based on collision risk. IEEE Transactions on Intelligent Vehicles (2023)

  213. [221]

    Xiaobo Chen, Shilin Zhang, Jun Li, and Jian Yang. 2024. Pedestrian Crossing Intention Prediction Based on Cross-Modal Transformer and Uncertainty-Aware Multi-Task Learning for Autonomous Driving. IEEE Transactions on Intelligent Transportation Systems (2024)

  214. [222]

    Abduallah Mohamed, Kun Qian, Mohamed Elhoseiny, and Christian Claudel. 2020. Social-stgcnn: A social spatio-temporal graph convolutional neural network for human trajectory prediction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 14424–14432

  215. [223]

    Eleni Kamenou, Jesús Martínez Del Rincón, Paul Miller, and Patricia Devlin-Hill. 2023. A meta-learning approach for domain generalisation across visual modalities in vehicle re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...

  216. [224]

    Yifan Jiang, Xinyu Gong, Ding Liu, Yu Cheng, Chen Fang, Xiaohui Shen, Jianchao Yang, Pan Zhou, and Zhangyang Wang. 2021. Enlightengan: Deep light enhancement without paired supervision. IEEE transactions on image processing 30 (2021), 2340–2349

  217. [225]

    Mark Schutera, Mostafa Hussein, Jochen Abhau, Ralf Mikut, and Markus Reischl. 2020. Night-to-day: Online image-to-image translation for object detection within autonomous driving by night. IEEE Transactions on Intelligent Vehicles 6, 3 (2020), 480–489

  218. [226]

    Fabio Pizzati, Pietro Cerri, and Raoul De Charette. 2021. CoMoGAN: continuous model-guided image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14288–14298

  219. [227]

    Jinlong Li, Baolu Li, Zhengzhong Tu, Xinyu Liu, Qing Guo, Felix Juefei-Xu, Runsheng Xu, and Hongkai Yu. 2024. Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous Driving. In Proceedings of the IEEE/CVF Conference on Computer ...

  220. [228]

    Jinlong Li, Zhigang Xu, Lan Fu, Xuesong Zhou, and Hongkai Yu. 2021. Domain adaptation from daytime to nighttime: A situation-sensitive vehicle detection and traffic flow parameter estimation framework. Transportation Research Part C: Emerging Technologies 124 (2021), 102946

  221. [229]

    Jie Xiao, Xueyang Fu, Aiping Liu, Feng Wu, and Zheng-Jun Zha. 2022. Image de-raining transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 11 (2022), 12978–12995. Vision Technologies with Applications in Traffic Surveillance Systems: A Holistic Survey 33

  222. [230]

    Xiang Chen, Hao Li, Mingqiang Li, and Jinshan Pan. 2023. Learning a sparse transformer network for effective image deraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5896–5905

  223. [231]

    Wei Zhou, Nan Zheng, and Chen Wang. 2024. Synthesizing Realistic Traffic Events from UAV Perspectives: A Mask-guided Generative Approach Based on Style-modulated Transformer. IEEE Transactions on Intelligent Vehicles (2024)

  224. [232]

    Yuhua Chen, Wen Li, Christos Sakaridis, Dengxin Dai, and Luc Van Gool. 2018. Domain adaptive faster r-cnn for object detection in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3339–3348

  225. [233]

    Chaoqi Chen, Zebiao Zheng, Xinghao Ding, Yue Huang, and Qi Dou. 2020. Harmonizing transferability and discriminability for adapting object detectors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 8869–8878

  226. [234]

    Muhammad Akhtar Munir, Muhammad Haris Khan, M Saquib Sarfraz, and Mohsen Ali. 2023. Domain adaptive object detection via balancing between self-training and adversarial learning. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  227. [235]

    Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. 2022. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5728–5739

  228. [236]

    Xingjia Pan, Fan Tang, Weiming Dong, Yang Gu, Zhichao Song, Yiping Meng, Pengfei Xu, Oliver Deussen, and Changsheng Xu. 2020. Self-supervised feature augmentation for large image object detection. IEEE Transactions on Image Processing 29 (2020), 6745–6758

  229. [237]

    Suvramalya Basak and S Suresh. 2024. Vehicle detection and type classification in low resolution congested traffic scenes using image super resolution. Multimedia Tools and Applications 83, 8 (2024), 21825–21847

  230. [238]

    Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. 2016. Matching networks for one shot learning. Advances in neural information processing systems 29 (2016)

  231. [239]

    Gongjie Zhang, Zhipeng Luo, Kaiwen Cui, Shijian Lu, and Eric P Xing. 2022. Meta-detr: Image-level few-shot detection with inter-class correlation exploitation. IEEE transactions on pattern analysis and machine intelligence 45, 11 (2022), 12832–12843

  232. [240]

    Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728 (2018)

  233. [241]

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16000–16009

  234. [242]

    Chao Huang, Zhihao Wu, Jie Wen, Yong Xu, Qiuping Jiang, and Yaowei Wang. 2021. Abnormal event detection using deep contrastive learning for intelligent video surveillance system. IEEE Transactions on Industrial Informatics 18, 8 (2021), 5171–5179

  235. [243]

    Antonio Barbalau, Radu Tudor Ionescu, Mariana-Iuliana Georgescu, Jacob Dueholm, Bharathkumar Ramachandra, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B Moeslund, and Mubarak Shah. 2023. SSMTL++: Revisiting self-supervised multi-task learning for video anomaly detection. Compu...

  236. [244]

    Haohan Luo and Feng Wang. 2023. A simulation-based framework for urban traffic accident detection. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5

  237. [245]

    Xiangyu Yue, Yang Zhang, Sicheng Zhao, Alberto Sangiovanni-Vincentelli, Kurt Keutzer, and Boqing Gong. 2019. Domain randomization and pyramid consistency: Simulation-to-real generalization without accessing target domain data. In Proceedings of the IEEE/CVF international confe...

  238. [246]

    Xuan Li, Haibin Duan, Bingzi Liu, Xiao Wang, and Fei-Yue Wang. 2023. A novel framework to generate synthetic video for foreground detection in highway surveillance scenarios. IEEE Transactions on Intelligent Transportation Systems 24, 6 (2023), 5958–5970

  239. [247]

    Stephan R Richter, Hassan Abu AlHaija, and Vladlen Koltun. 2022. Enhancing photorealism enhancement. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 2 (2022), 1700–1715

  240. [248]

    Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Gipyo Nam, and Ig-Jae Kim. 2022. Detection of road accidents using synthetically generated multi-perspective accident videos. IEEE Transactions on Intelligent Transportation Systems 24, 2 (2022), 1926–1935

  241. [249]

    Yancheng Ling, Zhenliang Ma, Qi Zhang, Bangquan Xie, and Xiaoxiong Weng. 2024. PedAST-GCN: Fast Pedestrian Crossing Intention Prediction Using Spatial–Temporal Attention Graph Convolution Networks. IEEE Transactions on Intelligent Transportation Systems (2024)

  242. [250]

    Rajat Koner, Hang Li, Marcel Hildebrandt, Deepan Das, Volker Tresp, and Stephan Günnemann. 2021. Graphhopper: Multi-hop scene graph reasoning for visual question answering. In The Semantic Web–ISWC 2021: 20th International Semantic Web Conference, ISWC 2021, Virtual Event, Oct...

  243. [251]

    Edward Curry, Dhaval Salwala, Praneet Dhingra, Felipe Arruda Pontes, and Piyush Yadav. 2022. Multimodal event processing: A neural-symbolic paradigm for the internet of multimedia things. IEEE Internet of Things Journal 9, 15 (2022), 13705–13724

  244. [252]

    Xu Yang, Hanwang Zhang, and Jianfei Cai. 2020. Auto-encoding and distilling scene graphs for image captioning. IEEE transactions on pattern analysis and machine intelligence 44, 5 (2020), 2313–2327

  245. [253]

    Danfei Xu, Dragomir Anguelov, and Ashesh Jain. 2018. Pointfusion: Deep sensor fusion for 3d bounding box estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 244–253

  246. [254]

    Kong Li, Zhe Dai, Chen Zuo, Xuan Wang, Hua Cui, Huansheng Song, and Mengying Cui. 2024. Scene adaptation in adverse conditions: a multi-sensor fusion framework for roadside traffic perception. Journal of Intelligent Transportation Systems (2024), 1–21

  247. [255]

    Farman Ali, Amjad Ali, Muhammad Imran, Rizwan Ali Naqvi, Muhammad Hameed Siddiqi, and Kyung-Sup Kwak. 2021. Traffic accident detection and condition analysis based on social networking data. Accident Analysis & Prevention 151 (2021), 105973

  248. [256]

    Jingxiao Liu, Siyuan Yuan, Yiwen Dong, Biondo Biondi, and Hae Young Noh. 2023. TelecomTM: A fine-grained and ubiquitous traffic monitoring system using pre-existing telecommunication fiber-optic cables as sensors. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubi...

  249. [257]

    Alexander Swerdlow, Runsheng Xu, and Bolei Zhou. 2024. Street-view image generation from a bird’s-eye view layout. IEEE Robotics and Automation Letters (2024)

  250. [258]

    Qingyi Wang, Shenhao Wang, Yunhan Zheng, Hongzhou Lin, Xiaohu Zhang, Jinhua Zhao, and Joan Walker. 2024. Deep hybrid model with satellite imagery: How to combine demand modeling and computer vision for travel behavior analysis? Transportation Research Part B: Methodological 17...

  251. [259]

    Jinchao Song, Chunli Zhao, Shaopeng Zhong, Thomas Alexander Sick Nielsen, and Alexander V Prishchepov. 2019. Mapping spatio-temporal patterns and detecting the factors of traffic congestion with multi-source data fusion and mining techniques. Computers, Environment and Urban S...

  252. [260]

    Lin Zhu, Fangce Guo, John W Polak, and Rajesh Krishnan. 2018. Urban link travel time estimation using traffic states-based data fusion. IET Intelligent Transport Systems 12, 7 (2018), 651–663. 34 Wei Zhou et al

  253. [261]

    Pu Wang, Zhiren Huang, Jiyu Lai, Zhihao Zheng, Yang Liu, and Tao Lin. 2021. Traffic speed estimation based on multi-source GPS data and mixture model.IEEE Transactions on Intelligent Transportation Systems 23, 8 (2021), 10708–10720

  254. [262]

    Hao Peng, Hongfei Wang, Bowen Du, Md Zakirul Alam Bhuiyan, Hongyuan Ma, Jianwei Liu, Lihong Wang, Zeyu Yang, Linfeng Du, Senzhang Wang, et al. 2020. Spatial temporal incidence dynamic graph neural networks for traffic flow forecasting. Information Sciences 521 (2020), 277–290

  255. [263]

    Medhavi Mishra, Sumit Mishra, and Dongsoo Har. 2024. Integrating Multi-sourced Sensor Data for Enhanced Traffic State Estimation. IEEE Sensors Journal (2024)

  256. [264]

    Eduardo Arnold, Mehrdad Dianati, Robert de Temple, and Saber Fallah. 2020. Cooperative perception for 3D object detection in driving scenarios using infrastructure sensors. IEEE Transactions on Intelligent Transportation Systems 23, 3 (2020), 1852–1864

  257. [265]

    Li Chen, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger, and Hongyang Li. 2024. End-to-end autonomous driving: Challenges and frontiers. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  258. [266]

    Sanbao Su, Yiming Li, Sihong He, Songyang Han, Chen Feng, Caiwen Ding, and Fei Miao. 2023. Uncertainty quantification of collaborative detection for self-driving. In 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 5588–5594

  259. [267]

    Runsheng Xu, Hao Xiang, Xin Xia, Xu Han, Jinlong Li, and Jiaqi Ma. 2022. Opv2v: An open benchmark dataset and fusion pipeline for perception with vehicle-to-vehicle communication. In 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2583–2589

  260. [268]

    Xin Gao, Xinyu Zhang, Yiguo Lu, Yuning Huang, Lei Yang, Yijin Xiong, and Peng Liu. 2024. A Survey of Collaborative Perception in Intelligent Vehicles at Intersections. IEEE Transactions on Intelligent Vehicles (2024)

  261. [269]

    Saquib Mazhar, Nadeem Atif, MK Bhuyan, and Shaik Rafi Ahamed. 2023. Rethinking DABNet: Light-weight Network for Real-time Semantic Segmentation of Road Scenes. IEEE Transactions on Artificial Intelligence (2023)

  262. [270]

    Jiahao Zheng, Longqi Yang, Yiying Li, Ke Yang, Zhiyuan Wang, and Jun Zhou. 2023. Lightweight Vision Transformer with Spatial and Channel Enhanced Self-Attention. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1492–1496

  263. [271]

    Yuqiao Liu, Yanan Sun, Bing Xue, Mengjie Zhang, Gary G Yen, and Kay Chen Tan. 2021. A survey on evolutionary neural architecture search. IEEE transactions on neural networks and learning systems 34, 2 (2021), 550–570

  264. [272]

    Sachin Mehta and Mohammad Rastegari. 2021. Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer. arXiv preprint arXiv:2110.02178 (2021)

  265. [273]

    Yanyu Li, Geng Yuan, Yang Wen, Ju Hu, Georgios Evangelidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. 2022. Efficientformer: Vision transformers at mobilenet speed. Advances in Neural Information Processing Systems 35 (2022), 12934–12949

  266. [274]

    Lie Guo, Pingshu Ge, Yibing Zhao, Dongxing Wang, and Liang Huang. 2023. LightMOT: a lightweight convolution neural network for real-time multi-object tracking. International Journal of Bio-Inspired Computation 22, 3 (2023), 152–161

  267. [275]

    Kang Yang, Tianzhang Xing, Yang Liu, Zhenjiang Li, Xiaoqing Gong, Xiaojiang Chen, and Dingyi Fang. 2019. cDeepArch: A compact deep neural network architecture for mobile sensing. IEEE/ACM Transactions on Networking 27, 5 (2019), 2043–2055

  268. [276]

    Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019. Importance estimation for neural network pruning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition . 11264–11272

  269. [277]

    Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2020. Improved knowledge distillation via teacher assistant. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 5191–5198

  270. [278]

    Jinqi Xiao, Chengming Zhang, Yu Gong, Miao Yin, Yang Sui, Lizhi Xiang, Dingwen Tao, and Bo Yuan. 2023. HALOC: hardware-aware automatic low-rank compression for compact neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 10464–10472

  271. [279]

    Jie Hu, Peng Lin, Huajun Zhang, Zining Lan, Wenxin Chen, Kailiang Xie, Siyun Chen, Hao Wang, and Sheng Chang. 2023. A dynamic pruning method on multiple sparse structures in deep neural networks. IEEE Access 11 (2023), 38448–38457

  272. [280]

    Yingchao Wang, Chen Yang, Shulin Lan, Liehuang Zhu, and Yan Zhang. 2024. End-edge-cloud collaborative computing for deep learning: A comprehensive survey. IEEE Communications Surveys & Tutorials (2024)

  273. [281]

    Qi Chen, Sihai Tang, Qing Yang, and Song Fu. 2019. Cooper: Cooperative perception for connected autonomous vehicles based on 3d point clouds. In 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS) . IEEE, 514–524

  274. [282]

    Ehzaz Mustafa, Junaid Shuja, Faisal Rehman, Ahsan Riaz, Mohammed Maray, Muhammad Bilal, and Muhammad Khurram Khan. 2024. Deep Neural Networks meet computation offloading in mobile edge networks: Applications, taxonomy, and open issues. Journal of Network and Computer Applicati...

  275. [283]

    Pian Qi, Diletta Chiaro, Antonella Guzzo, Michele Ianni, Giancarlo Fortino, and Francesco Piccialli. 2024. Model aggregation techniques in federated learning: A comprehensive survey. Future Generation Computer Systems 150 (2024), 272–293

  276. [284]

    Danesh Shokri, Christian Larouche, and Saeid Homayouni. 2024. Proposing an Efficient Deep Learning Algorithm Based on Segment Anything Model for Detection and Tracking of Vehicles through Uncalibrated Urban Traffic Surveillance Cameras. Electronics 13, 14 (2024), 2883

  277. [285]

    Wei Zhou, Hongpu Huang, Hancheng Zhang, and Chen Wang. 2024. Teaching Segment-Anything-Model Domain-Specific Knowledge for Road Crack Segmentation From On-Board Cameras. IEEE Transactions on Intelligent Transportation Systems (2024)

  278. [286]

    Guoyang Zhao, Fulong Ma, Weiqing Qi, Chenguang Zhang, Yuxuan Liu, Ming Liu, and Jun Ma. 2024. TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition. arXiv preprint arXiv:2409.15077 (2024)

  279. [287]

    Aaron Lohner, Francesco Compagno, Jonathan Francis, and Alessandro Oltramari. 2024. Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding. arXiv preprint arXiv:2407.05910 (2024)

  280. [288]

    Jianzong Wu, Xiangtai Li, Shilin Xu, Haobo Yuan, Henghui Ding, Yibo Yang, Xia Li, Jiangning Zhang, Yunhai Tong, Xudong Jiang, et al. 2024. Towards open vocabulary learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  281. [289]

    George Tom, Minesh Mathew, Sergi Garcia-Bordils, Dimosthenis Karatzas, and CV Jawahar. 2023. Reading Between the Lanes: Text VideoQA on the Road. In International Conference on Document Analysis and Recognition . Springer, 137–154

  282. [290]

    Mohammad Abu Tami, Huthaifa I Ashqar, Mohammed Elhenawy, Sebastien Glaser, and Andry Rakotonirainy. 2024. Using Multimodal Large Language Models (MLLMs) for Automated Detection of Traffic Safety-Critical Events. Vehicles 6, 3 (2024), 1571–1590

  283. [291]

    Kan Guo, Daxin Tian, Yongli Hu, Chunmian Lin, Zhen Qian, Yanfeng Sun, Jianshan Zhou, Xuting Duan, Junbin Gao, and Baocai Yin. 2024. CFMMC-Align: Coarse-Fine Multi-Modal Contrastive Alignment Network for Traffic Event Video Question Answering. IEEE Transactions on Circuits and ...

  284. [292]

    Jiarui Zhang, Filip Ilievski, Kaixin Ma, Aravinda Kollaa, Jonathan Francis, and Alessandro Oltramari. 2023. A study of situational reasoning for traffic understanding. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3262–3272. Vision T...

  285. [293]

    Xu Cao, Tong Zhou, Yunsheng Ma, Wenqian Ye, Can Cui, Kun Tang, Zhipeng Cao, Kaizhao Liang, Ziran Wang, James M Rehg, et al. 2024. MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding. In Proceedings of the IEEE/CVF Conference on Com...

  286. [294]

    Lening Wang, Yilong Ren, Han Jiang, Pinlong Cai, Daocheng Fu, Tianqi Wang, Zhiyong Cui, Haiyang Yu, Xuesong Wang, Hanchu Zhou, et al. 2023. Accidentgpt: Accident analysis and prevention from v2x environmental perception with multi-modal large model. arXiv preprint arXiv:2312.1...

  287. [295]

    Joseph Cho, Fachrina Dewi Puspitasari, Sheng Zheng, Jingyao Zheng, Lik-Hang Lee, Tae-Ho Kim, Choong Seon Hong, and Chaoning Zhang. 2024. Sora as an agi world model? a complete survey on text-to-video generation. arXiv preprint arXiv:2403.05131 (2024)

  288. [296]

    Lening Wang, Wenzhao Zheng, Yilong Ren, Han Jiang, Zhiyong Cui, Haiyang Yu, and Jiwen Lu. 2024. OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving. arXiv preprint arXiv:2405.20337 (2024)

  289. [297]

    Indu Joshi, Marcel Grimmer, Christian Rathgeb, Christoph Busch, Francois Bremond, and Antitza Dantcheva. 2024. Synthetic data in human analysis: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  290. [298]

    Runnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen, Xinge Zhu, Yuexin Ma, Tongliang Liu, and Wenping Wang. 2024. Towards label-free scene understanding by vision foundation models. Advances in Neural Information Processing Systems 36 (2024)

  291. [299]

    Brian HW Guo, Yang Zou, Yihai Fang, Yang Miang Goh, and Patrick XW Zou. 2021. Computer vision technologies for safety science and management in construction: A critical review and future research directions. Safety science 135 (2021), 105130

  292. [300]

    Xiao Wen, Yuanchang Xie, Lingtao Wu, and Liming Jiang. 2021. Quantifying and comparing the effects of key risk factors on various types of roadway segment crashes with LightGBM and SHAP. Accident Analysis & Prevention 159 (2021), 106261

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.