REVIEW 5 major objections 3 minor 247 references
What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies
T0 review · 5 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This survey claims that all road-safety-relevant scene-understanding research can be organized by a two-group, eleven-category taxonomy of attention-worthy elements, and uses it to analyze 40 vision-driven tasks and 78 datasets.
desk verdict A genuinely useful dataset survey with a sensible taxonomy, but the authors need to clean up their own scoring: headline numbers disagree with the abstract, and a few taxonomy assignments contradict their own definitions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the taxonomy itself, organized around the question 'Which elements are most critical in the current traffic scenario and require attention to road safety?' It is a two-level hierarchy: the reason the element demands attention (anomaly versus pertinent entity) and the specific category of that reason (spatial location, semantic category, event, kinematic pattern, status or condition, appearance, or multiple reasons for anomalies; spatial location, semantic category, status or condition, or kinematic pattern for pertinent entities). Each leaf node in the taxonomy's graphical diagram maps to the corresponding tasks and representative datasets, with gray task nodes marking the absence of publicly available datasets. The taxonomy does the work of the paper: it is the standard against which tasks are assigned, datasets are reclassified, and research gaps are identified.
What would settle it
Re-examine a sample of the reclassified datasets at the annotation level: if Pothole-600's pixel masks are not genuinely segmentation-grade, if GLARE frequently annotates more than one traffic sign per image, or if a newly published vision-driven road-safety dataset cannot be assigned to any of the eleven categories without stretching their definitions, the paper's reclassification claims and its claim of comprehensive coverage would be called into question.
Extended reading notes
Core claim
The paper's central claim is that 'critical traffic elements that demand attention' can be classified by their functional role in road safety, and that this classification fully organizes the field. The resulting taxonomy has two groups: anomalies, which demand attention because they are abnormal (by spatial location, semantic category, event, kinematic pattern, status or condition, appearance, or a hybrid of these), and pertinent entities, which are normal but critical to the current driving maneuver (by spatial location, semantic category, status or condition, or kinematic pattern). The paper reports that this yields eleven categories and twenty-three common research topics, and it uses the taxonomy to examine 40 vision-driven tasks and 78 datasets, reclassifying each dataset according to what its data and annotations actually support rather than its stated task. It further claims that cross-domain investigation reveals substantial variations in benchmark quality, with recurring limitations such as uneven task coverage, imbalanced distributions, inconsistent or insufficient annotations, and limited multimodal and cross-task support.
Load-bearing premise
The survey's conclusions about coverage and benchmark quality rest on the authors' unpublished qualitative reclassification of each dataset and on their manual visual inspection of annotation files, so if those judgments are inconsistent or incomplete, the taxonomy's assignments and the research-gap findings would change.
Editorial extensions
If this is right
- Researchers can use the taxonomy to position new tasks and datasets relative to the full field, making it easier to transfer methods between related but historically isolated areas such as obstacle segmentation, anomaly segmentation, and accident anticipation.
- The functional classification exposes concrete gaps: road construction detection, traffic salient object detection, and appearance-based anomaly tasks (criminal recognition, suspicious vehicle recognition, damaged vehicle detection) currently lack public datasets.
- Datasets with mismatched claimed tasks are reclassified by what their annotations actually support, so benchmarks such as Pothole-600 become segmentation resources rather than detection resources, and GLARE is treated as single-sign traffic sign detection.
- The distinction between anomaly and pertinent entity changes how accidents are categorized: ego-involved collisions are treated as anomalous kinematic patterns, while accidents witnessed by the ego vehicle are treated as anomalous events.
- The survey's cross-domain analysis supports a push toward unified annotation standards and dataset reuse, because many datasets can serve multiple categories once their functional format is recognized.
Reading between the lines
- If the taxonomy is adopted by the community, a natural next step would be a unified benchmark suite that samples one or more datasets from each of the eleven categories, allowing holistic road-safety perception to be measured instead of task-isolated accuracy.
- The authors' practice of reclassifying datasets by annotation format implies that published task names are not a reliable guide to dataset utility; a testable extension would be requiring dataset authors to report the functional annotation level (classification, detection, segmentation, graph, temporal) as metadata.
- The taxonomy's functional-role principle could be carried beyond vision into multimodal and planning-oriented settings, since the same attention-worthy elements appear in LiDAR point clouds and in prediction or decision-support pipelines that the survey deliberately excludes.
- A direct editorial check would be to compute inter-annotator agreement on assigning a sample of datasets to the eleven categories; low agreement would signal that the taxonomy needs sharper definitions, while high agreement would support its claim to be a stable organizing framework.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey proposes a taxonomy of attention-worthy traffic elements for road safety, organizing them into two main groups (anomalies and pertinent entities), eleven categories, and twenty-three research topics, and uses it to review 40 vision-driven tasks and 78 datasets. The paper provides detailed descriptions, annotation visualizations, and per-dataset quality critiques, and it draws cross-domain conclusions about research gaps and benchmark limitations.
Significance. If the taxonomy is consistent and the dataset assignments are sound, this survey would provide a valuable unified framework for a fragmented field, with practical utility for dataset selection and gap identification. The paper's concrete annotation visualizations and specific, checkable claims about dataset flaws (e.g., dummy boxes in CADP, label mismatches in CST-S3D, broken links in A3D) are strengths that go beyond typical survey-level description. The central claim, however, depends on the correctness and transparency of the qualitative reclassification that underlies the taxonomy assignments, and that process is neither fully specified nor reproducible, which is the main risk to the paper's contribution.
major comments (5)
- [Section 1 and Abstract] The inventory counts are inconsistent between the abstract (10 categories, 20 subclasses, 35 tasks, 73 datasets) and the full text (11 categories, 23 subclasses, 40 tasks, 78 datasets). Since the survey's contribution is framed as a complete organization of the field, these mismatches must be reconciled and a single set of counts used consistently throughout the manuscript and all tables.
- [Section 3.9.7] The text states that GLARE actually localizes only one traffic sign per image, yet GLARE is placed under Traffic Sign Detection. By the paper's own definitions in Sections 2.2 and 2.3, single-instance localization is object localization, not detection. This is an internal inconsistency in applying the stated functional-format rule, and it affects the taxonomy's credibility for this category.
- [Section 3.7.1] RiskBench is the sole dataset in the category 'Anomaly Due to Multiple Reasons' / Risk Identification, but the text admits that it 'exclusively labels one predefined type of risks per sample.' This placement contradicts the category's own justification, which requires multi-factor anomalies. The assignment should be reconsidered, or the category definition and the dataset description should be aligned.
- [Section 3.4.3] Driver Attention Prediction is filed under 'Anomaly Due to Kinematic Pattern', yet the attention maps in BDD-A and DADA-2000 are derived from gaze and fixation patterns, not from kinematic patterns of traffic elements. This is a categorical mismatch with the taxonomy's stated reason for the category, and it undermines the claim that category membership is determined by the functional-format criterion.
- [Section 3 and Table 2] The dataset-to-taxonomy assignments rest on an unpublished, qualitative reclassification of each dataset (e.g., Pothole-600 as segmentation, GLARE as detection, UCF-Crime as partially traffic-related), with no protocol and no released assignment table. Because the 'research gap' conclusions depend directly on these assignments, the paper should specify the reclassification procedure and provide a complete, citable appendix of assignments so that the claimed coverage can be independently verified.
minor comments (3)
- [Abstract] There is a missing space between 'Comparedto' and 'existingsurveys' in the abstract; these and other spacing errors should be corrected during copyediting.
- [Section 3.3.4] The paper notes that all 204 YouTube links in A3D are inaccessible, which prevents data retrieval and verification. This should be flagged more prominently in Table 2 (e.g., with a data-availability marker) so that readers are not misled about the dataset's current usability.
- [Table 2] The table's caption defines 'Level' abbreviations (C, L, D, S, I, P), but the table itself does not indicate which granularity of label is used for 'D' in datasets where detection is only claimed by the original authors; adding a footnote that distinguishes claimed vs. reclassified labels would improve clarity.
Circularity Check
No significant circularity: the taxonomy is a definitional construct; insider datasets (STU, M2S-RoAD) and inconsistent assignments (GLARE, RiskBench) are reproducibility concerns, not reductions to inputs.
full rationale
The paper's load-bearing claim is that its taxonomy, namely two groups, eleven categories, twenty-three topics, 40 tasks, and 78 datasets, organizes the field of vision-driven road-safety perception. That claim is a definitional construct, not a derived result: the paper states the taxonomy 'was initially derived from thirteen distinct perspectives, integrating insights from traffic scene understanding literature and practical road-safety considerations' and then grouped by whether the element is abnormal or normal-but-critical. There is no equation, no fitted parameter, and no predicted quantity that reduces to an input, so no classic circularity pattern (self-definitional derivation, fitted input called prediction, or uniqueness imported from prior work) is present. The survey is self-contained against external benchmarks: nearly all 78 datasets are external resources, and the taxonomy structure would stand even if every dataset were replaced. Two entries appear to come from the authors' own research group (STU [47], M2S-RoAD [64]), but neither is load-bearing: the Anomaly Segmentation category contains five independent 2D benchmarks alongside STU, and Road Damage Segmentation contains PotholeMix alongside M2S-RoAD, so the central organizing claim does not depend on these entries. The concerns worth flagging are reproducibility and consistency, not circularity. The stated rule that each dataset is classified 'according to its actual functional format' is applied inconsistently: GLARE is kept under Traffic Sign Detection although the paper says it 'actually only localizes one traffic sign per image,' which the paper's own Section 2.2/2.3 single-instance vs. multi-instance distinction places in object localization; RiskBench is the sole occupant of 'Anomaly Due to Multiple Reasons' although the paper admits it 'exclusively labels one predefined type of risks per sample.' The Road Construction Detection gap partly results from the paper's own scoping exclusion of construction-site datasets ('exhibit a clear domain shift from our focus on urban street scenes and are therefore not further analyzed'), and the abstract (10 categories, 35 tasks, 73 datasets) contradicts the full text (11 categories, 40 tasks, 78 datasets). These are correctness and verifiability risks for the survey's empirical claims, to be weighed in referee review, not evidence that the taxonomy reduces to its inputs.
Assumptions & free parameters
assumptions (2)
- domain assumption The 78 surveyed datasets and 40 tasks are representative of the field of vision-driven road-safety scene understanding.
- domain assumption The authors' manual reclassification of each dataset to its 'actual functional format' is faithful and consistent.
Cite this review
Pith. "Pith review of What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies." pith.science (2026). https://pith.science/paper/C75BYUHC
@misc{pith2026250706513,
author = {Pith},
title = {Pith review of: What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies},
year = {2026},
howpublished = {\url{https://pith.science/paper/C75BYUHC}},
note = {Machine review of arXiv:2507.06513}
}
read the original abstract
Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To facilitate the use of these improvements for road safety, this survey systematically categorizes the critical elements that demand attention in traffic scenarios and comprehensively analyzes available vision-driven tasks and datasets. Compared to existing surveys that focus on isolated domains, our taxonomy categorizes attention-worthy traffic entities into two main groups that are anomalies and normal but critical entities, integrating ten categories and twenty subclasses. It establishes connections between inherently related fields and provides a unified analytical framework. Our survey highlights the analysis of 35 vision-driven tasks and comprehensive examinations and visualizations of 73 available datasets based on the proposed taxonomy. The cross-domain investigation covers the pros and cons of each benchmark with the aim of providing information on standards unification and resource optimization. Our article concludes with a systematic discussion of the existing weaknesses, underlining the potential effects and promising solutions from various perspectives. The integrated taxonomy, comprehensive analysis, and recapitulatory tables serve as valuable contributions to this rapidly evolving field by providing researchers with a holistic overview, guiding strategic resource selection, and highlighting critical research gaps.
Figures
Figures from the paper (57 more)
Reference graph
Works this paper leans on
-
[1]
Road traffic injuries,https://www.who.int/news-room/fact-sheets/detail/road-traffic-injuries (2023)
2023
-
[2]
B. Tian, B. T. Morris, M. Tang, Y. Liu, Y. Yao, C. Gou, D. Shen, S. Tang, Hierarchical and networked vehicle surveillance in its: A survey, IEEE Transactions on Intelligent Transportation Systems 18 (1) (2017) 25–48.doi:10.1109/TITS.2016.2552778
arXiv 2017
-
[3]
9396–9405.doi:10.1109/CVPR.2019.00963
A.Kirillov,K.He,R.Girshick,C.Rother,P.Dollár,Panopticsegmentation,in:2019IEEE/CVFConferenceonComputerVisionandPattern Recognition (CVPR), 2019, pp. 9396–9405.doi:10.1109/CVPR.2019.00963
arXiv 2019
-
[4]
Y.Zhou,Y.Zhang,Z.Zhao,K.Zhang,C.Gou,Towarddrivingsceneunderstanding:Aparadigmandbenchmarkdatasetforego-centrictraffic scene graph representation, IEEE Journal of Radio Frequency Identification 6 (2022) 962–967.doi:10.1109/JRFID.2022.3207017
arXiv 2022
- [5]
-
[6]
L. Qin, Y. Shi, Y. He, J. Zhang, X. Zhang, Y. Li, T. Deng, H. Yan, Id-yolo: Real-time salient object detection based on the driver’s fixation region, IEEE Transactions on Intelligent Transportation Systems 23 (9) (2022) 15898–15908.doi:10.1109/TITS.2022.3146271
arXiv 2022
-
[7]
N. Jia, Y. Sun, X. Liu, Tfgnet: Traffic salient object detection using a feature deep interaction and guidance fusion, IEEE Transactions on Intelligent Transportation Systems 25 (3) (2024) 3020–3030.doi:10.1109/TITS.2023.3293822
arXiv 2024
-
[8]
Y. Xia, D. Zhang, J. Kim, K. Nakayama, K. Zipser, D. Whitney, Predicting driver attention in critical situations, in: Asian Conference on Computer Vision, 2017. URL https://api.semanticscholar.org/CorpusID:52019549
2017
Show all 247 references
-
[9]
J. Fang, D. Yan, J. Qiao, J. Xue, H. Yu, Dada: Driver attention prediction in driving accident scenarios, IEEE Transactions on Intelligent Transportation Systems 23 (6) (2022) 4959–4971.doi:10.1109/TITS.2020.3044678
2022
-
[10]
K. K. Santhosh, D. P. Dogra, P. P. Roy, Anomaly detection in road traffic using visual surveillance: A survey, ACM Comput. Surv. 53 (6) (Dec. 2020). doi:10.1145/3417989. URL https://doi.org/10.1145/3417989
2020 doi
-
[11]
Y. Yao, X. Wang, M. Xu, Z. Pu, Y. Wang, E. Atkins, D. J. Crandall, DoTA: Unsupervised Detection of Traffic Anomaly in Driving Videos , IEEE Transactions on Pattern Analysis & Machine Intelligence 45 (01) (2023) 444–459.doi:10.1109/TPAMI.2022.3150763. URL https://doi.ieeecomput...
2023
-
[12]
J. Yu, J. Jiang, S. Fichera, P. Paoletti, L. Layzell, D. Mehta, S. Luo, Road surface defect detection—from image-based to non-image-based: A survey, IEEE Transactions on Intelligent Transportation Systems PP (2024) 1–23.doi:10.1109/TITS.2024.3382837
2024
-
[13]
Chan, Y.-T
F.-H. Chan, Y.-T. Chen, Y. Xiang, M. Sun, Anticipating accidents in dashcam videos, in: Asian Conference on Computer Vision, 2016. URL https://api.semanticscholar.org/CorpusID:45520437
2016
-
[14]
H. Kim, K. Lee, G. Hwang, C. Suh, Crash to not crash: Learn to identify dangerous vehicles using a simulator, Proceedings of the AAAI Conference on Artificial Intelligence 33 (2019) 978–985.doi:10.1609/aaai.v33i01.3301978
2019 doi
-
[15]
206–213.doi:10.1109/ICCVW.2017.33
A.Rasouli,I.Kotseruba,J.K.Tsotsos,Aretheygoingtocross?abenchmarkdatasetandbaselineforpedestriancrosswalkbehavior,in:2017 IEEE International Conference on Computer Vision Workshops (ICCVW), 2017, pp. 206–213.doi:10.1109/ICCVW.2017.33
2017 doi
-
[16]
Pinggera, S
P. Pinggera, S. Ramos, S. Gehrig, U. Franke, C. Rother, R. Mester, Lost and found: detecting small road hazards for self-driving vehicles, in:2016IEEE/RSJInternationalConferenceonIntelligentRobotsandSystems(IROS),2016,pp.1099–1106. doi:10.1109/IROS.2016. 7759186
2016 doi
-
[17]
R. Chan, K. Lis, S. Uhlemeyer, H. Blum, S. Honari, R. Siegwart, P. Fua, M. Salzmann, M. Rottmann, Segmentmeifyoucan: A benchmark for anomaly segmentation, in: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021
2021
-
[18]
K. Lis, K. Nakka, P. Fua, M. Salzmann, Detecting the unexpected via image resynthesis, in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE Computer Society, Los Alamitos, CA, USA, 2019, pp. 2152–2161.doi:10.1109/ICCV.2019.00224. URL https://doi.ieeecompu...
2019
-
[19]
2403–2412.doi:10.1109/ICCVW.2019
H.Blum,P.-E.Sarlin,J.Nieto,R.Siegwart,C.Cadena,Fishyscapes:Abenchmarkforsafesemanticsegmentationinautonomousdriving,in: 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), 2019, pp. 2403–2412.doi:10.1109/ICCVW.2019. 00294
2019 doi
-
[20]
A. P. Shah, J.-B. Lamare, T. Nguyen-Anh, A. Hauptmann, Cadp: A novel dataset for cctv traffic camera based accident analysis, in: 2018 15thIEEEInternationalConferenceonAdvancedVideoandSignalBasedSurveillance(AVSS),2018,pp.1–9. doi:10.1109/AVSS.2018. 8639160
2018 doi
-
[21]
doi:10.1145/3394171.3413827
W.Bao,Q.Yu,Y.Kong,Uncertainty-basedtrafficaccidentanticipationwithspatio-temporalrelationallearning,in:Proceedingsofthe28th ACMInternationalConferenceonMultimedia,MM’20,AssociationforComputingMachinery,NewYork,NY,USA,2020,p.2682–2690. doi:10.1145/3394171.3413827. URL https://d...
2020
-
[22]
Y.Yao,M.Xu,Y.Wang,D.J.Crandall,E.M.Atkins,Unsupervisedtrafficaccidentdetectioninfirst-personvideos,in:IEEE/RSJInternational Conference on Intelligent Robots and Systems (IROS), 2019
2019
-
[23]
: Preprint submitted to Elsevier Page 75 of 85
J.Fang,J.Qiao,J.Xue,Z.Li,Vision-basedtrafficaccidentdetectionandanticipation:Asurvey,IEEETransactionsonCircuitsandSystems for Video Technology 34 (4) (2024) 1983–1999.doi:10.1109/TCSVT.2023.3307655. : Preprint submitted to Elsevier Page 75 of 85
2024
-
[24]
S.-Y. Yu, A. V. Malawade, D. Muthirayan, P. P. Khargonekar, M. A. A. Faruque, Scene-graph augmented data-driven risk assessment of autonomous vehicle decisions, IEEE Transactions on Intelligent Transportation Systems 23 (7) (2022) 7941–7951.doi:10.1109/TITS. 2021.3074854
2022
-
[25]
N. Bao, D. Yang, A. Carballo, Ü. Özgüner, K. Takeda, Personalized safety-focused control by minimizing subjective risk, in: 2019 IEEE Intelligent Transportation Systems Conference (ITSC), IEEE, 2019, pp. 3853–3858
2019
-
[26]
Kopuklu, J
O. Kopuklu, J. Zheng, H. Xu, G. Rigoll, Driver anomaly detection: A dataset and contrastive learning approach, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 91–100
2021
-
[27]
Zhang, S
Y. Zhang, S. Wang, Z. Shi, W. Yang, A survey on anomaly segmentation in urban scene understanding with image data, Knowledge-Based Systems 338 (2026) 115521.doi:https://doi.org/10.1016/j.knosys.2026.115521. URL https://www.sciencedirect.com/science/article/pii/S0950705126002637
2026
-
[28]
R.Jiao,Y.Wan,F.Poiesi,Y.Wang,Surveyonvideoanomalydetectionindynamicsceneswithmovingcameras,Artif.Intell.Rev.56(Suppl
-
[29]
URL https://doi.org/10.1007/s10462-023-10609-x
(2023) 3515–3570.doi:10.1007/s10462-023-10609-x. URL https://doi.org/10.1007/s10462-023-10609-x
2023 doi
-
[30]
L. Zhu, L. Wang, A. Raj, T. Gedeon, C. Chen, Advancing video anomaly detection: a concise review and a new dataset, in: Proceedings of the38thInternationalConferenceonNeuralInformationProcessingSystems,NIPS’24,CurranAssociatesInc.,RedHook,NY,USA,2024
2024
-
[31]
doi:10.1145/1541880.1541882
V.Chandola,A.Banerjee,V.Kumar,Anomalydetection:Asurvey,ACMComput.Surv.41(072009). doi:10.1145/1541880.1541882
-
[32]
arXiv:https://doi.org/10.1080/14680629.2023.2237601, doi:10.1080/14680629.2023.2237601
C.Abdollahi,M.Mollajafari,A.Golroo,S.Moridpour,H.Wang,Areviewonpavementdataacquisitionandanalyticstoolsusingautonomous vehicles,RoadMaterialsandPavementDesign25(5)(2024)914–940. arXiv:https://doi.org/10.1080/14680629.2023.2237601, doi:10.1080/14680629.2023.2237601. URL https:/...
2024
-
[33]
Ghari, A
B. Ghari, A. Tourani, A. Shahbahrami, G. Gaydadjiev, Pedestrian detection in low-light conditions: A comprehensive survey, Image Vision Comput. 148 (C) (Aug. 2024).doi:10.1016/j.imavis.2024.105106. URL https://doi.org/10.1016/j.imavis.2024.105106
2024
-
[34]
J. Cao, Y. Pang, J. Xie, F. S. Khan, L. Shao, From handcrafted to deep features for pedestrian detection: A survey, IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (9) (2022) 4913–4934.doi:10.1109/TPAMI.2021.3076733
2022
-
[35]
J. Chen, Q. Wang, H. H. Cheng, W. Peng, W. Xu, A review of vision-based traffic semantic understanding in itss, IEEE Transactions on Intelligent Transportation Systems 23 (11) (2022) 19954–19979.doi:10.1109/TITS.2022.3182410
2022
-
[36]
M.Liu,E.Yurtsever,J.Fossaert,X.Zhou,W.Zimmer,Y.Cui,B.L.Zagar,A.C.Knoll,Asurveyonautonomousdrivingdatasets:Statistics, annotation quality, and a future outlook, IEEE Transactions on Intelligent Vehicles (2024) 1–29doi:10.1109/TIV.2024.3394735
2024
-
[37]
Paneru, I
S. Paneru, I. Jeelani, Computer vision applications in construction: Current state, opportunities & challenges, Automation in Construction 132 (2021) 103940.doi:https://doi.org/10.1016/j.autcon.2021.103940. URL https://www.sciencedirect.com/science/article/pii/S0926580521003915
2021
-
[38]
Simonyan, A
K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition, in: International Conference on Learning Representations, 2015
2015
-
[39]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778.doi:10.1109/CVPR.2016.90
2016 doi
-
[40]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Repr...
2021
-
[41]
doi:10.48550/arXiv.2602.23120
A.Sabaghi,J.OramasM,Trilite:Efficientweaklysupervisedobjectlocalizationwithuniversalvisualfeaturesandtri-regiondisentanglement (02 2026). doi:10.48550/arXiv.2602.23120
2026 doi
-
[42]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255.doi:10.1109/CVPR.2009.5206848
2009
-
[43]
F.Yu,H.Chen,X.Wang,W.Xian,Y.Chen,F.Liu,V.Madhavan,T.Darrell,Bdd100k:Adiversedrivingdatasetforheterogeneousmultitask learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[44]
Everingham, L
M. Everingham, L. van Gool, C. Williams, J. Winn, A. Zisserman, The pascal visual object classes (voc) challenge, International Journal of Computer Vision 88 (2) (2010) 303–338.doi:10.1007/s11263-009-0275-4
2010 doi
-
[45]
URL https://api.semanticscholar.org/CorpusID:347907
P.Henderson,V.Ferrari,End-to-endtrainingofobjectclassdetectorsformeanaverageprecision,in:AsianConferenceonComputerVision, 2016. URL https://api.semanticscholar.org/CorpusID:347907
2016
-
[46]
A. D. Rhodes, M. H. Quinn, M. Mitchell, Fast on-line kernel density estimation for active object localization, in: 2017 International Joint Conference on Neural Networks (IJCNN), 2017, pp. 454–462.doi:10.1109/IJCNN.2017.7965889
2017
-
[47]
U. A. Computer Vision Lab, Lecture 12: Computer vision, Course Materials for COMPSCI 682, available at:https://cvl-umass. github.io/compsci682-fall-2024/docs/lecture12fall2024.pdf (2024)
2024
-
[48]
Conference on Computer Vision and Pattern Recognition (CVPR)
A. Nekrasov, M. Burdorf, S. Worrall, B. Leibe, J. S. B. Perez, Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving, in: "Conference on Computer Vision and Pattern Recognition (CVPR)", 2025
2025
-
[49]
O. O. Xiying Li, Chinese city traffic image database(cctrib) (2022). URL http://www.openits.cn/openData4/824.jhtml
2022
-
[50]
L. Wen, D. Du, Z. Cai, Z. Lei, M.-C. Chang, H. Qi, J. Lim, M.-H. Yang, S. Lyu, Ua-detrac: A new benchmark and protocol for multi-object detectionandtracking,ComputerVisionandImageUnderstanding193(2020)102907. doi:https://doi.org/10.1016/j.cviu.2020. 102907. URL https://www.sci...
2020 doi
-
[51]
URL https://dx.doi.org/10.21227/tjtg-nz28
V.Adewopo,N.Elsayed,Z.ElSayed,M.Ozer,C.Zekios,A.Abdelgawad,M.Bayoumi,Trafficaccidentdetectionvideodatasetforai-driven computer vision systems in smart city transportation (2023).doi:10.21227/tjtg-nz28. URL https://dx.doi.org/10.21227/tjtg-nz28
2023 doi
-
[52]
T. You, B. Han, Traffic Accident Benchmark for Causality Recognition, in: ECCV, 2020
2020
-
[53]
Fang, L.-l
J. Fang, L.-l. Li, J. Zhou, J. Xiao, H. Yu, C. Lv, J. Xue, T.-S. Chua, Abductive ego-view accident video understanding for safe driving perception, in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 22030–22040. doi: 10.1109/CVPR52733.2024.02080
2024
-
[54]
Sultani, C
W. Sultani, C. Chen, M. Shah, Real-world anomaly detection in surveillance videos, in: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[55]
Chan, Y.-T
F.-H. Chan, Y.-T. Chen, Y. Xiang, M. Sun, Anticipating accidents in dashcam videos, in: S.-H. Lai, V. Lepetit, K. Nishino, Y. Sato (Eds.), Computer Vision – ACCV 2016, Springer International Publishing, Cham, 2017, pp. 136–153
2016
-
[56]
Pradana, M.-S
H. Pradana, M.-S. Dao, K. Zettsu, Augmenting ego-vehicle for traffic near-miss and accident classification dataset using manipulating conditional style translation, in: 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA), 2022, pp. 1–8...
2022
-
[57]
Q. Zou, Y. Cao, Q. Li, Q. Mao, S. Wang, Cracktree: Automatic crack detection from pavement images, Pattern Recognition Letters 33 (3) (2012) 227–238. doi:https://doi.org/10.1016/j.patrec.2011.11.004. URL https://www.sciencedirect.com/science/article/pii/S0167865511003795
2012 doi
-
[58]
Y. Shi, L. Cui, Z. Qi, F. Meng, Z. Chen, Automatic road crack detection using random structured forests, IEEE Transactions on Intelligent Transportation Systems 17 (12) (2016) 3434–3445.doi:10.1109/TITS.2016.2552248
2016
-
[59]
F. Yang, L. Zhang, S. Yu, D. Prokhorov, X. Mei, H. Ling, Feature pyramid and hierarchical boosting network for pavement crack detection, IEEE Transactions on Intelligent Transportation Systems 21 (4) (2020) 1525–1535.doi:10.1109/TITS.2019.2910595
2020
-
[60]
URL https://www.sciencedirect.com/science/article/pii/S0950061820314021
Q.Mei,M.Gül,Acosteffectivesolutionforpavementcrackinspectionusingcamerasanddeepneuralnetworks,ConstructionandBuilding Materials 256 (2020) 119397.doi:https://doi.org/10.1016/j.conbuildmat.2020.119397. URL https://www.sciencedirect.com/science/article/pii/S0950061820314021
2020
-
[61]
Z.Huang,W.Chen,A.Al-Tabbaa,I.Brilakis,Nha12d:Anewpavementcrackdatasetandacomparisonstudyofcrackdetectionalgorithms, 2022 European Conference on Computing in Construction (2022)
2022
-
[62]
R. Fan, U. Ozgunalp, B. Hosking, M. Liu, I. Pitas, Pothole detection based on disparity transformation and road surface modeling, IEEE Transactions on Image Processing 29 (2020) 897–908.doi:10.1109/TIP.2019.2933750
2020
-
[63]
W. Tang, Q. Zhao, S. Huang, R. Li, L. Huangfu, An iteratively optimized patch label inference network for automatic pavement distress detection, IEEE Transactions on Intelligent Transportation Systems 23 (2020) 8652–8661. URL https://api.semanticscholar.org/CorpusID:218900648
2020
-
[64]
Moscoso Thompson, A
E. Moscoso Thompson, A. Ranieri, S. Biasotti, M. Chicchon, I. Sipiran, M.-K. Pham, T.-L. Nguyen-Ho, H.-D. Nguyen, M.-T. Tran, Shrec 2022: Pothole and crack detection in the road pavement using images and rgb-d data, Comput. Graph. 107 (C) (2022) 161–171. doi:10.1016/j.cag.2022...
2022 doi
-
[65]
T.-Y.Tseng,H.Lyu,J.Li,J.S.Berrio,M.Shan,S.Worrall,M2s-road:Multi-modalsemanticsegmentationforroaddamageusingcameraand lidar data, in: Proceedings of the 2024 Australasian Conference on Robotics and Automation (ACRA), Australian Robotics and Automation Association, 2024, in press
2024
-
[66]
Stricker, M
R. Stricker, M. Eisenbach, M. Sesselmann, K. Debes, H.-M. Gross, Improving visual road condition assessment by extensive experiments on the extended gaps dataset, in: 2019 International Joint Conference on Neural Networks (IJCNN), 2019, pp. 1–8.doi:10.1109/IJCNN. 2019.8852257
2019
-
[67]
arXiv:https://rmets.onlinelibrary.wiley.com/doi/pdf/10.1002/gdj3.260, doi:https://doi.org/10.1002/gdj3.260
D.Arya,H.Maeda,S.K.Ghosh,D.Toshniwal,Y.Sekimoto,Rdd2022:Amulti-nationalimagedatasetforautomaticroaddamagedetection, GeoscienceDataJournal11(4)(2024)846–862. arXiv:https://rmets.onlinelibrary.wiley.com/doi/pdf/10.1002/gdj3.260, doi:https://doi.org/10.1002/gdj3.260. URL https://...
2024 doi
-
[68]
Sakaridis, D
C. Sakaridis, D. Dai, L. Van Gool, ACDC: The adverse conditions dataset with correspondences for semantic driving scene understanding, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021
2021
-
[69]
X.Hu,C.-W.Fu,L.Zhu,P.-A.Heng,Depth-attentionalfeaturesforsingle-imagerainremoval,in:ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[70]
Sakaridis, D
C. Sakaridis, D. Dai, S. Hecker, L. Van Gool, Model adaptation with synthetic and real data for semantic dense foggy scene understanding, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018
2018
-
[71]
URL https://doi.org/10.1007/s11263-018-1072-8
C.Sakaridis,D.Dai,L.VanGool,Semanticfoggysceneunderstandingwithsyntheticdata,InternationalJournalofComputerVision126(9) (2018) 973–992. URL https://doi.org/10.1007/s11263-018-1072-8
2018 doi
-
[72]
3819–3824.doi:10.1109/ITSC.2018.8569387
D.Dai,L.V.Gool,Darkmodeladaptation:Semanticimagesegmentationfromdaytimetonighttime,in:201821stInternationalConference on Intelligent Transportation Systems (ITSC), 2018, pp. 3819–3824.doi:10.1109/ITSC.2018.8569387
2018
-
[73]
X. Tan, K. Xu, Y. Cao, Y. Zhang, L. Ma, R. W. H. Lau, Night-time scene parsing with a large real dataset, IEEE Transactions on Image Processing 30 (2021) 9085–9098.doi:10.1109/TIP.2021.3122004
2021
-
[74]
C.Sakaridis,D.Dai,L.VanGool,Map-guidedcurriculumdomainadaptationanduncertainty-awareevaluationforsemanticnighttimeimage segmentation, IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (6) (2022) 3139–3153.doi:10.1109/TPAMI.2020. 3045882
2022 doi
-
[75]
Neumann, M
L. Neumann, M. Karg, S. Zhang, C. Scharfenberger, E. Piegert, S. Mistr, O. Prokofyeva, R. Thiel, A. Vedaldi, A. Zisserman, B. Schiele, Nightowls: A pedestrians at night dataset, in: Asian Conference on Computer Vision, 2018. : Preprint submitted to Elsevier Page 77 of 85 URL h...
2018
-
[76]
N. Gray, M. Moraes, J. Bian, A. Wang, A. Tian, K. Wilson, Y. Huang, H. Xiong, Z. Guo, Glare: A dataset for traffic sign detection in sun glare, IEEE Transactions on Intelligent Transportation Systems 24 (11) (2023) 12323–12330.doi:10.1109/TITS.2023.3294411
2023
-
[78]
Rasouli, I
A. Rasouli, I. Kotseruba, T. Kunic, J. K. Tsotsos, Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction, in: International Conference on Computer Vision (ICCV), 2019
2019
-
[79]
B.Liu,E.Adeli,Z.Cao,K.-H.Lee,A.Shenoi,A.Gaidon,J.C.Niebles,Spatiotemporalrelationshipreasoningforpedestrianintentprediction, in: IEEE Robotics and Automation Letters (IEEE RA-L) and International Conference on Robotics and Automation (ICRA), IEEE, 2020
2020
-
[80]
W.Kim,M.S.Ramanagopal,C.Barto,M.-Y.Yu,K.Rosaen,N.Goumas,R.Vasudevan,M.Johnson-Roberson,Pedx:Benchmarkdatasetfor metric 3-d pose estimation of pedestrians in complex urban intersections, IEEE Robotics and Automation Letters 4 (2) (2019) 1940–1947
2019
-
[82]
Braun, F
M. Braun, F. B. Flohr, S. Krebs, U. Kreße, D. M. Gavrila, Simple pair pose - pairwise human pose estimation in dense urban traffic scenes, in: 2021 IEEE Intelligent Vehicles Symposium (IV), 2021, pp. 1545–1552.doi:10.1109/IV48863.2021.9575435
2021
-
[83]
A. V. Malawade, S.-Y. Yu, B. Hsu, D. Muthirayan, P. P. Khargonekar, M. A. A. Faruque, Spatiotemporal scene-graph embedding for autonomousvehiclecollisionprediction,IEEEInternetofThingsJournal9(12)(2022)9379–9388. doi:10.1109/JIOT.2022.3141044
2022
-
[84]
Yurtsever, Y
E. Yurtsever, Y. Liu, J. Lambert, C. Miyajima, E. Takeuchi, K. Takeda, J. H. L. Hansen, Risky action recognition in lane change video clips usingdeepspatiotemporalnetworkswithsegmentationmasktransfer,in:2019IEEEIntelligentTransportationSystemsConference(ITSC), 2019, pp. 3100–3...
2019
-
[85]
Dollar, C
P. Dollar, C. Wojek, B. Schiele, P. Perona, Pedestrian detection: A benchmark, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 304–311.doi:10.1109/CVPR.2009.5206631
2009
-
[86]
Zhang, R
S. Zhang, R. Benenson, B. Schiele, Citypersons: A diverse dataset for pedestrian detection, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4457–4465.doi:10.1109/CVPR.2017.474
2017 doi
-
[87]
M.Braun,S.Krebs,F.B.Flohr,D.M.Gavrila,Eurocitypersons:Anovelbenchmarkforpersondetectionintrafficscenes,IEEETransactions on Pattern Analysis and Machine Intelligence (2019) 1–1doi:10.1109/TPAMI.2019.2897684
2019
-
[88]
X. Jia, C. Zhu, M. Li, W. Tang, W. Zhou, LLVIP: A Visible-infrared Paired Dataset for Low-light Vision , in: 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), IEEE Computer Society, Los Alamitos, CA, USA, 2021, pp. 3489–3497.doi: 10.1109/ICCVW54120.2...
2021
-
[89]
X. Li, F. Flohr, Y. Yang, H. Xiong, M. Braun, S. Pan, K. Li, D. M. Gavrila, A new benchmark for vision-based cyclist detection, in: 2016 IEEE Intelligent Vehicles Symposium (IV), 2016, pp. 1028–1033.doi:10.1109/IVS.2016.7535515
2016
-
[90]
K. Zhou, C. Li, M. Wang, L. Tomato, Tusimple competitions for cvpr2017,https://github.com/TuSimple/tusimple-benchmark (2017)
2017
-
[91]
X. Pan, J. Shi, P. Luo, X. Wang, X. Tang, Spatial as deep: Spatial cnn for traffic scene understanding, in: AAAI Conference on Artificial Intelligence (AAAI), 2018
2018
-
[92]
Zhang, L
Y. Zhang, L. Zhu, W. Feng, H. Fu, M. Wang, Q. Li, C. Li, S. Wang, Vil-100: A new dataset and a baseline model for video instance lane detection, 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021) 15661–15670. URL https://api.semanticscholar.org/CorpusID:237213300
2021
-
[93]
L. Chen, C. Sima, Y. Li, Z. Zheng, J. Xu, X. Geng, H. Li, C. He, J. Shi, Y. Qiao, J. Yan, Persformer: 3d lane detection via perspective transformer and the openlane benchmark, in: European Conference on Computer Vision (ECCV), 2022
2022
-
[94]
Huang, P
X. Huang, P. Wang, X. Cheng, D. Zhou, Q. Geng, R. Yang, The apolloscape open dataset for autonomous driving and its application, IEEE Transactions on Pattern Analysis and Machine Intelligence 42 (10) (2020) 2702–2719.doi:10.1109/TPAMI.2019.2926463
2020
-
[97]
B.Wilson,W.Qi,T.Agarwal,J.Lambert,J.Singh,S.Khandelwal,B.Pan,R.Kumar,A.Hartnett,J.K.Pontes,D.Ramanan,P.Carr,J.Hays, Argoverse 2: Next generation datasets for self-driving perception and forecasting, in: Proceedings of the Neural Information Processing Systems Track on Datasets...
2021
-
[98]
D.Tabernik,D.Skočaj,Deeplearningforlarge-scaletraffic-signdetectionandrecognition,IEEETransactionsonIntelligentTransportation Systems 21 (4) (2020) 1427–1440.doi:10.1109/TITS.2019.2913588
2020
-
[99]
Z.Zhu,D.Liang,S.Zhang,X.Huang,B.Li,S.Hu,Traffic-signdetectionandclassificationinthewild,in:TheIEEEConferenceonComputer Vision and Pattern Recognition (CVPR), 2016
2016
-
[100]
Gámez Serna, Y
C. Gámez Serna, Y. Ruichek, Traffic signs detection and classification for european urban environments, IEEE Transactions on Intelligent Transportation Systems 21 (10) (2020) 4388–4399.doi:10.1109/TITS.2019.2941081
2020
-
[101]
Y. Guo, W. Feng, F. Yin, T. Xue, S. Mei, C.-L. Liu, Learning to understand traffic signs, in: Proceedings of the 29th ACM International Conference on Multimedia, MM ’21, Association for Computing Machinery, New York, NY, USA, 2021, p. 2076–2084.doi:10.1145/ : Preprint submitte...
2021
-
[103]
M. B. Jensen, M. P. Philipsen, A. Møgelmose, T. B. Moeslund, M. M. Trivedi, Vision for looking at traffic lights: Issues, survey, and perspectives, IEEE Transactions on Intelligent Transportation Systems 17 (7) (2016) 1800–1815.doi:10.1109/TITS.2015.2509509
2016
-
[104]
Behrendt, L
K. Behrendt, L. Novak, R. Botros, A deep learning approach to traffic lights: Detection, tracking, and classification, in: 2017 IEEE International Conference on Robotics and Automation (ICRA), 2017, pp. 1370–1377.doi:10.1109/ICRA.2017.7989163
2017
-
[105]
Fregin, J
A. Fregin, J. Muller, U. Krebel, K. Dietmayer, The driveu traffic light dataset: Introduction and comparison with existing datasets, in: 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, pp. 3376–3383.doi:10.1109/ICRA.2018.8460737
2018
-
[106]
X. Yang, J. Yan, W. Liao, X. Yang, J. Tang, T. He, SCRDet++: Detecting Small, Cluttered and Rotated Objects via Instance-Level Feature Denoising and Rotation Loss Smoothing , IEEE Transactions on Pattern Analysis & Machine Intelligence 45 (02) (2023) 2384–2399. doi:10.1109/TPA...
2023
-
[107]
J. He, C. Zhang, X. He, R. Dong, Visual recognition of traffic police gestures with convolutional pose machine and handcrafted features, Neurocomputing 390 (2020) 248–259.doi:https://doi.org/10.1016/j.neucom.2019.07.103. URL https://www.sciencedirect.com/science/article/pii/S0...
2020 doi
-
[108]
Izquierdo, A
R. Izquierdo, A. Quintanar, I. Parra, D. Fernández-Llorca, M. A. Sotelo, The prevention dataset: a novel benchmark for prediction of vehicles intentions, in: 2019 IEEE Intelligent Transportation Systems Conference (ITSC), 2019, pp. 3114–3121.doi:10.1109/ITSC. 2019.8917433
2019
-
[109]
International Conference on Robotics and Automation, in press, 2019
J.Xue,J.Fang,T.Li,B.Zhang,P.Zhang,Z.Ye,J.Dou,BLVD:Buildingalarge-scale5dsemanticsbenchmarkforautonomousdriving,in: Proc. International Conference on Robotics and Automation, in press, 2019
2019
-
[110]
Ettinger, S
S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V. Vasudevan, A. McCauley, J. Shlens, D. Anguelov, Large scale interactive motion forecasting for autonomous driving: The waymo open motion...
2021
-
[111]
J. Choe, S. J. Oh, S. Chun, S. Lee, Z. Akata, H. Shim, Evaluation for weakly supervised object localization: Protocol, metrics, and datasets, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2) (2023) 1732–1748.doi:10.1109/TPAMI.2022.3169881
2023
-
[112]
Gupta, S
S. Gupta, S. Lakhotia, A. Rawat, R. Tallamraju, Vitol: Vision transformer for weakly supervised object localization, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4101–4110
2022
-
[113]
O.Russakovsky,J.Deng,H.Su,J.Krause,S.Satheesh,S.Ma,Z.Huang,A.Karpathy,A.Khosla,M.Bernstein,A.Berg,L.Fei-Fei,Imagenet large scale visual recognition challenge, International Journal of Computer Vision 115 (09 2014).doi:10.1007/s11263-015-0816-y
2014 doi
-
[114]
K. He, G. Gkioxari, P. Dollár, R. B. Girshick, Mask r-cnn, 2017 IEEE International Conference on Computer Vision (ICCV) (2017) 2980– 2988. URL https://api.semanticscholar.org/CorpusID:206771194
2017
-
[115]
S.Ren,K.He,R.Girshick,J.Sun,Fasterr-cnn:Towardsreal-timeobjectdetectionwithregionproposalnetworks,in:C.Cortes,N.Lawrence, D. Lee, M. Sugiyama, R. Garnett (Eds.), Advances in Neural Information Processing Systems, Vol. 28, Curran Associates, Inc., 2015. URL https://proceedings....
2015
-
[116]
Redmon, S
J. Redmon, S. K. Divvala, R. B. Girshick, A. Farhadi, You only look once: Unified, real-time object detection, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016) 779–788. URL https://api.semanticscholar.org/CorpusID:206594738
2016
-
[117]
Welling (Eds.), Computer Vision – ECCV 2016, Springer International Publishing, Cham, 2016, pp
W.Liu,D.Anguelov,D.Erhan,C.Szegedy,S.Reed,C.-Y.Fu,A.C.Berg,Ssd:Singleshotmultiboxdetector,in:B.Leibe,J.Matas,N.Sebe, M. Welling (Eds.), Computer Vision – ECCV 2016, Springer International Publishing, Cham, 2016, pp. 21–37
2016
-
[118]
Z. Tian, C. Shen, H. Chen, T. He, Fcos: Fully convolutional one-stage object detection, in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 9626–9635.doi:10.1109/ICCV.2019.00972
2019
-
[119]
N.Carion,F.Massa,G.Synnaeve,N.Usunier,A.Kirillov,S.Zagoruyko,End-to-endobjectdetectionwithtransformers,in:ComputerVision –ECCV2020:16thEuropeanConference,Glasgow,UK,August23–28,2020,Proceedings,PartI,Springer-Verlag,Berlin,Heidelberg,2020, p. 213–229. doi:10.1007/978-3-030-584...
2020 doi
-
[120]
Robinson, P
I. Robinson, P. Robicheaux, M. Popov, D. Ramanan, N. Peri, Rf-detr: Neural architecture search for real-time detection transformers (2025). arXiv:2511.09554. URL https://arxiv.org/abs/2511.09554
2025
-
[121]
Everingham, S
M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, A. Zisserman, The pascal visual object classes challenge: A retrospective, International Journal of Computer Vision 111 (1) (2015) 98–136
2015
-
[122]
URL https://api.semanticscholar.org/CorpusID:14113767
T.-Y.Lin,M.Maire,S.J.Belongie,J.Hays,P.Perona,D.Ramanan,P.Dollár,C.L.Zitnick,Microsoftcoco:Commonobjectsincontext,in: European Conference on Computer Vision, 2014. URL https://api.semanticscholar.org/CorpusID:14113767
2014
-
[123]
of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
M.Cordts,M.Omran,S.Ramos,T.Rehfeld,M.Enzweiler,R.Benenson,U.Franke,S.Roth,B.Schiele,Thecityscapesdatasetforsemantic urban scene understanding, in: Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[124]
Girshick, Fast r-cnn, in: 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp
R. Girshick, Fast r-cnn, in: 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1440–1448.doi:10.1109/ICCV. 2015.169
2015 doi
-
[126]
Ronneberger, P
O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: N. Navab, J. Hornegger, W. M. Wells, A. F. Frangi (Eds.), Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, Springer International Publishing, Cham...
2015
-
[127]
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, A. L. Yuille, Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs, IEEE Transactions on Pattern Analysis and Machine Intelligence 40 (4) (2018) 834–848. doi:10.1109/T...
2018
-
[128]
Zheng, J
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr, L. Zhang, Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers, in: CVPR, 2021
2021
-
[129]
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, P. Luo, Segformer: Simple and efficient design for semantic segmentation with transformers, in: Neural Information Processing Systems (NeurIPS), 2021
2021
-
[130]
5000–5009.doi:10.1109/ICCV.2017.534
G.Neuhold,T.Ollmann,S.R.Bulò,P.Kontschieder,Themapillaryvistasdatasetforsemanticunderstandingofstreetscenes,in:2017IEEE International Conference on Computer Vision (ICCV), 2017, pp. 5000–5009.doi:10.1109/ICCV.2017.534
2017 doi
-
[131]
Hariharan, P
B. Hariharan, P. Arbeláez, R. Girshick, J. Malik, Simultaneous detection and segmentation, in: European Conference on Computer Vision (ECCV), 2014
2014
-
[132]
X. Wang, T. Kong, C. Shen, Y. Jiang, L. Li, Solo: Segmenting objects by locations, in: Proc. Eur. Conf. Computer Vision (ECCV), 2020
2020
-
[133]
Kirillov, Y
A. Kirillov, Y. Wu, K. He, R. B. Girshick, Pointrend: Image segmentation as rendering, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019) 9796–9805. URL https://api.semanticscholar.org/CorpusID:209386851
2019
-
[134]
B. Dong, F. Zeng, T. Wang, X. Zhang, Y. Wei, Solq: segmenting objects by learning queries, in: Proceedings of the 35th International Conference on Neural Information Processing Systems, NIPS ’21, Curran Associates Inc., Red Hook, NY, USA, 2021
2021
-
[135]
Cheng, I
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, R. Girdhar, Masked-attention mask transformer for universal image segmentation, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 1280–1289.doi:10.1109/CVPR52688.2022. 00135
2022
-
[138]
Kirillov, R
A. Kirillov, R. Girshick, K. He, P. Dollár, Panoptic feature pyramid networks, in: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 6392–6401.doi:10.1109/CVPR.2019.00656
2019
-
[139]
12472–12482
B.Cheng,M.D.Collins,Y.Zhu,T.Liu,T.S.Huang,H.Adam,L.-C.Chen,Panoptic-deeplab:Asimple,strong,andfastbaselineforbottom- up panoptic segmentation, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12472–12482. doi:10.1109/CVPR42600.2020.01249
2020
-
[140]
Chang, S.-E
C.-Y. Chang, S.-E. Chang, P. Hsiao, L. Fu, Epsnet: Efficient panoptic segmentation network with cross-layer attention fusion, ArXiv abs/2003.10142 (2020)
2020 arXiv
-
[141]
J. Jain, J. Li, M. Chiu, A. Hassani, N. Orlov, H. Shi, OneFormer: One Transformer to Rule Universal Image Segmentation, 2023
2023
-
[142]
T.-J. Yang, M. D. Collins, Y. Zhu, J.-J. Hwang, T. Liu, X. Zhang, V. Sze, G. Papandreou, L.-C. Chen, Deeperlab: Single-shot image parser, ArXiv abs/1902.05093 (2019). URL https://api.semanticscholar.org/CorpusID:61153476
2019 arXiv
-
[143]
D. Xu, Y. Zhu, C. Choy, L. Fei-Fei, Scene graph generation by iterative message passing, in: Computer Vision and Pattern Recognition (CVPR), 2017
2017
-
[144]
Johnson, R
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, L. Fei-Fei, Image retrieval using scene graphs, in: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 3668–3678.doi:10.1109/CVPR.2015.7298990
2015
-
[145]
Zareian, S
A. Zareian, S. Karaman, S.-F. Chang, Bridging knowledge graphs to generate scene graphs, in: Proceedings of the European conference on computer vision (ECCV), 2020
2020
-
[146]
R. Li, S. Zhang, B. Wan, X. He, Bipartite graph network with adaptive message passing for unbiased scene graph generation, in: 2021 IEEE/CVFConferenceonComputerVisionandPatternRecognition(CVPR),2021,pp.11104–11114. doi:10.1109/CVPR46437.2021. 01096
2021
-
[147]
Y. Cong, M. Y. Yang, B. Rosenhahn, Reltr: Relation transformer for scene graph generation, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
2023
-
[148]
24229–24238
J.Im,J.Nam,N.Park,H.Lee,S.Park,Egtr:Extractinggraphfromtransformerforscenegraphgeneration,in:ProceedingsoftheIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 24229–24238
2024
-
[149]
Zhang, S
C. Zhang, S. Stepputtis, J. Campbell, K. Sycara, Y. Xie, Hiker-sgg: Hierarchical knowledge enhanced robust scene graph generation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 28233–28243
2024
-
[150]
L. Li, C. ZHANG, D. Zhang, C. Sun, C. Li, L. Chen, Interaction-centric knowledge infusion and transfer for open vocabulary scene graph generation, in: The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=h2kwURAFkJ
2025
-
[151]
Krishna, Y
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. S. Bernstein, L. Fei-Fei, Visualgenome:Connectinglanguageandvisionusingcrowdsourceddenseimageannotations,Int.J.Comput.Vision123(1)(2017)32–73. doi:10.1007/s11263-0...
2017 doi
-
[152]
Ferrari, The open images dataset v4, International Journal of Computer Vision 128 (2018) 1956 – 1981
A.Kuznetsova,H.Rom,N.G.Alldrin,J.R.R.Uijlings,I.Krasin,J.Pont-Tuset,S.Kamali,S.Popov,M.Malloci,A.Kolesnikov,T.Duerig, V. Ferrari, The open images dataset v4, International Journal of Computer Vision 128 (2018) 1956 – 1981. URL https://api.semanticscholar.org/CorpusID:53296866
2018
-
[153]
D. A. Hudson, C. D. Manning, GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering , in: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA, 2019, pp. 6693–6702. doi:10.1109...
2019
-
[154]
C.Lu,R.Krishna,M.Bernstein,L.Fei-Fei,Visualrelationshipdetectionwithlanguagepriors,in:EuropeanConferenceonComputerVision, 2016
2016
-
[155]
T. Chen, W. Yu, R. Chen, L. Lin, Knowledge-embedded routing network for scene graph generation, in: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 6156–6164.doi:10.1109/CVPR.2019.00632
2019
-
[156]
T. Yin, X. Zhou, P. Krähenbühl, Center-based 3d object detection and tracking, CVPR (2021)
2021
-
[157]
R. Qian, X. Lai, X. Li, 3d object detection for autonomous driving: A survey, Pattern Recognition 130 (2022) 108796.doi:https: //doi.org/10.1016/j.patcog.2022.108796. URL https://www.sciencedirect.com/science/article/pii/S0031320322002771
2022
-
[158]
C. Nie, Z. Ju, Z. Sun, H. Zhang, 3d object detection and tracking based on lidar-camera fusion and imm-ukf algorithm towards highway driving, IEEE Transactions on Emerging Topics in Computational Intelligence 7 (4) (2023) 1242–1252.doi:10.1109/TETCI.2023. 3259441
2023 doi
-
[159]
6520–6530
F.Pu,Y.Wang,J.Deng,W.Yang,Monodgp:Monocular3dobjectdetectionwithdecoupled-queryandgeometry-errorpriors,in:Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 6520–6530
2025
-
[160]
Y. Liu, J. Yan, F. Jia, S. Li, Q. Gao, T. Wang, X. Zhang, J. Sun, Petrv2: A unified framework for 3d perception from multi-camera images, arXiv preprint arXiv:2206.01256 (2022)
2022 arXiv
-
[161]
Y.Chen,J.Liu,X.Zhang,X.Qi,J.Jia,Voxelnext:Fullysparsevoxelnetfor3dobjectdetectionandtracking,in:ProceedingsoftheIEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023
2023
-
[162]
Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. Rus, S. Han, Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation, in: IEEE International Conference on Robotics and Automation (ICRA), 2023
2023
-
[163]
J. Yin, J. Shen, R. Chen, W. Li, R. Yang, P. Frossard, W. Wang, Is-fusion: Instance-scene collaborative fusion for multimodal 3d object detection, in: CVPR, 2024
2024
-
[164]
Geiger, P
A. Geiger, P. Lenz, R. Urtasun, Are we ready for autonomous driving? the kitti vision benchmark suite, in: Conference on Computer Vision and Pattern Recognition (CVPR), 2012
2012
-
[167]
URL https://www.sciencedirect.com/science/article/pii/S1566253524004494
H.Xu,J.Chen,S.Meng,Y.Wang,L.-P.Chau,Asurveyonoccupancyperceptionforautonomousdriving:Theinformationfusionperspective, Information Fusion 114 (2025) 102671.doi:https://doi.org/10.1016/j.inffus.2024.102671. URL https://www.sciencedirect.com/science/article/pii/S1566253524004494
2025
-
[168]
A.-Q. Cao, R. de Charette, Monoscene: Monocular 3d semantic scene completion, in: CVPR, 2022
2022
-
[170]
Wolters, J
P. Wolters, J. Gilg, T. Teepe, F. Herzog, A. Laouichi, M. Hofmann, G. Rigoll, Unleashing hydra: Hybrid fusion, depth consistency and radar for unified 3d perception, in: 2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 7467–7474. doi:10.1109/ICRA...
2025
- [171]
-
[172]
Z. Ming, J. Stephany Berrio, M. Shan, S. Worrall, Occfusion: Multi-sensor fusion framework for 3d semantic occupancy prediction, IEEE Transactions on Intelligent Vehicles 10 (5) (2025) 3421–3433.doi:10.1109/TIV.2024.3453293
2025
-
[173]
X. Tian, T. Jiang, L. Yun, Y. Mao, H. Yang, Y. Wang, Y. Wang, H. Zhao, Occ3d: a large-scale 3d occupancy prediction benchmark for autonomous driving, in: Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Curran Associates Inc....
2023
-
[174]
Y. Wei, L. Zhao, W. Zheng, Z. Zhu, J. Zhou, J. Lu, Surroundocc: Multi-camera 3d occupancy prediction for autonomous driving, 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (2023) 21672–21683. URL https://api.semanticscholar.org/CorpusID:257557568
2023
-
[175]
17804–17813.doi:10.1109/ICCV51070.2023.01636
X.Wang,Z.Zhu,W.Xu,Y.Zhang,Y.Wei,X.Chi,Y.Ye,D.Du,J.Lu,X.Wang, OpenOccupancy:ALargeScaleBenchmarkforSurrounding Semantic Occupancy Perception , in: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE Computer Society, Los Alamitos, CA, USA, 2023, pp. 17804–178...
2023
-
[176]
Everingham, L
M. Everingham, L. V. Gool, C. K. I. Williams, J. M. Winn, A. Zisserman, The pascal visual object classes (voc) challenge, International Journal of Computer Vision 88 (2010) 303–338. URL https://api.semanticscholar.org/CorpusID:4246903 : Preprint submitted to Elsevier Page 81 of 85
2010
-
[177]
4303–4309.doi:10.1109/ITSC.2019.8917218
J.Fang,D.Yan,J.Qiao,J.Xue,H.Wang,S.Li,Dada-2000:Candrivingaccidentbepredictedbydriverattention ƒanalyzedbyabenchmark, in: 2019 IEEE Intelligent Transportation Systems Conference (ITSC), 2019, pp. 4303–4309.doi:10.1109/ITSC.2019.8917218
2000
-
[178]
2039–2047
M.Eisenbach,R.Stricker,D.Seichter,K.Amende,K.Debes,M.Sesselmann,D.Ebersbach,U.Stoeckert,H.-M.Gross,Howtogetpavement distressdetectionreadyfordeeplearning?asystematicapproach,in:2017InternationalJointConferenceonNeuralNetworks(IJCNN),2017, pp. 2039–2047. doi:10.1109/IJCNN.2017.7966101
2017
-
[179]
URL https://www.sciencedirect.com/science/article/pii/S0952197621002542
P.S.Perumal,M.Sujasree,S.Chavhan,D.Gupta,V.Mukthineni,S.R.Shimgekar,A.Khanna,G.Fortino,Aninsightintocrashavoidanceand overtaking advice systems for autonomous vehicles: A review, challenges and solutions, Engineering Applications of Artificial Intelligence 104 (2021) 104406.do...
2021
-
[180]
Bogdoll, S
D. Bogdoll, S. Uhlemeyer, K. Kowol, J. M. Zöllner, Perception Datasets for Anomaly Detection in Autonomous Driving: A Survey, in: Intelligent Vehicles Symposium (IV), 2023
2023
-
[181]
K. Lis, S. Honari, P. Fua, M. Salzmann, Detecting road obstacles by erasing them, IEEE Trans. Pattern Anal. Mach. Intell. 46 (4) (2023) 2450–2460. doi:10.1109/TPAMI.2023.3335152. URL https://doi.org/10.1109/TPAMI.2023.3335152
2023
-
[182]
Nekrasov, A
A. Nekrasov, A. Hermans, L. Kuhnert, B. Leibe, UGainS: Uncertainty Guided Anomaly Instance Segmentation, in: GCPR, 2023
2023
-
[183]
M. Park, D. Q. Tran, J. Bak, S. Park, Small and overlapping worker detection at construction sites, Automation in Construction 151 (2023) 104856. doi:https://doi.org/10.1016/j.autcon.2023.104856. URL https://www.sciencedirect.com/science/article/pii/S0926580523001164
2023
-
[184]
A.Xuehui,Z.Li,L.Zuguang,W.Chengzhi,L.Pengfei,L.Zhiwei,Datasetandbenchmarkfordetectingmovingobjectsinconstructionsites, Automation in Construction 122 (2021) 103482
2021
-
[185]
R. Feng, Y. Miao, J. Zheng, A yolo-based intelligent detection algorithm for risk assessment of construction sites, Journal of Intelligent Construction (2024). doi:10.26599/JIC.2024.9180037. URL https://www.sciopen.com/article/10.26599/JIC.2024.9180037
2024
-
[186]
R. Duan, H. Deng, M. Tian, Y. Deng, J. Lin, Soda: A large-scale open site object detection dataset for deep learning in construction, Automation in Construction 142 (2022) 104499
2022
-
[187]
X. Ke, L. Shi, W. Guo, D. Chen, Multi-dimensional traffic congestion detection based on fusion of visual features and convolutional neural network, IEEE Transactions on Intelligent Transportation Systems 20 (6) (2019) 2157–2170.doi:10.1109/TITS.2018.2864612
2019
-
[188]
S. Jiang,Y. Feng, W.Zhang, X. Liao, X.Dai, B. O.Onasanya, A new multi-branchconvolutional neural networkand feature mapextraction method for traffic congestion detection, Sensors 24 (13) (2024).doi:10.3390/s24134272. URL https://www.mdpi.com/1424-8220/24/13/4272
2024 doi
-
[189]
X. Li, T. Hao, X. Jin, B. Huang, J. Liang, Fine traffic congestion detection with hierarchical description, IEEE Transactions on Intelligent Transportation Systems 23 (12) (2022) 24439–24453.doi:10.1109/TITS.2022.3206583
2022
-
[190]
A.Shah,J.B.Lamare,N.-A.Tuan,A.Hauptmann,Accidentforecastingincctvtrafficcameravideos,arXivpreprintarXiv:1809.05782First three authors share the first authorship. (2018)
2018 arXiv
-
[191]
A.P.Shah,J.-B.Lamare,T.Nguyen-Anh,A.Hauptmann,Caraccidentsdetectionandpredictiondatasetv1, https://docs.google.com/ document/d/12F7l4yxNzzUAISZufEd9WFhQKSefVVo_QsPdTsWxZh8/edit?tab=t.0 (2018)
2018
-
[192]
W. Bao, Q. Yu, Y. Kong, DRIVE: Deep Reinforced Accident Anticipation with Visual Explanation , in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE Computer Society, Los Alamitos, CA, USA, 2021, pp. 7599–7608.doi:10.1109/ ICCV48922.2021.00752. URL https:/...
2021
-
[193]
Y. Yao, M. Xu, C. Choi, D. J. Crandall, E. M. Atkins, B. Dariush, Egocentric vision-based future vehicle localization for intelligent driving assistance systems, in: International Conference on Robotics and Automation, 2019
2019
-
[194]
G. Tang, H. Zhao, B. Yu, Low-cost and high-performance abnormal trajectory detection based on the gru model with deep spatiotemporal sequence analysis in cloud computing, J. Cloud Comput. 13 (1) (Mar. 2024).doi:10.1186/s13677-024-00611-1. URL https://doi.org/10.1186/s13677-024-00611-1
2024 doi
-
[195]
A. R. Alozi, M. Hussein, How do active road users act around autonomous vehicles? an inverse reinforcement learning approach, Transportation Research Part C: Emerging Technologies 161 (2024) 104572.doi:https://doi.org/10.1016/j.trc.2024.104572. URL https://www.sciencedirect.co...
2024
-
[196]
Palazzi, D
A. Palazzi, D. Abati, F. Solera, R. Cucchiara, Predicting the driver’s focus of attention: the dr (eye) ve project, IEEE transactions on pattern analysis and machine intelligence 41 (7) (2018) 1720–1733
2018
-
[197]
Y. Liu, M. Wang, P. Lasang, Q. Sun, Importance biased traffic scene segmentation in diverse weather conditions, IEEE Transactions on Intelligent Vehicles 9 (1) (2024) 2753–2765.doi:10.1109/TIV.2023.3272922
2024
-
[198]
Q. Zou, Z. Zhang, Q. Li, X. Qi, Q. Wang, S. Wang, Deepcrack: Learning hierarchical convolutional features for crack detection, IEEE Transactions on Image Processing 28 (3) (2019) 1498–1512
2019
-
[199]
Q. Mei, M. Gül, M. R. Azim, Densely connected deep neural network considering connectivity of pixels for automatic crack detection, Automation in Construction 110 (2019) 103018.doi:10.1016/j.autcon.2019.103018
2019
-
[200]
doi:10.1109/TITS.2019.2931297
A.Dhiman,R.Klette,Potholedetectionusingcomputervisionandlearning,IEEETransactionsonIntelligentTransportationSystems21(8) (2020) 3536–3550. doi:10.1109/TITS.2019.2931297
2020
-
[201]
B. T. Passos, M. J. Cassaniga, A. M. R. Fernandes, K. B. Medeiros, E. Comunello, Cracks and potholes in road images,https://data. mendeley.com/datasets/t576ydh9v8/4 (2020)
2020
-
[202]
doi:10.1109/BigData55660.2022
D.Arya,H.Maeda,S.K.Ghosh,D.Toshniwal,H.Omata,T.Kashiyama,Y.Sekimoto,Crowdsensing-basedroaddamagedetectionchallenge (crddc’2022),in:2022IEEEInternationalConferenceonBigData(BigData),2022,pp.6378–6386. doi:10.1109/BigData55660.2022. 10021040. : Preprint submitted to Elsevier Pag...
2022
-
[203]
T. Yu, Q. Kuang, J. Hu, J. Zheng, X. Li, Global-similarity local-salience network for traffic weather recognition, IEEE Access 9 (2021) 4607–4615. doi:10.1109/ACCESS.2020.3048116
2021
-
[204]
Zendel, K
O. Zendel, K. Honauer, M. Murschitz, D. Steininger, G. F. Dominguez, Wilddash - creating hazard-aware benchmarks, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018
2018
-
[205]
C. Qian, Y. Guo, Y. Mo, W. Li, Weatherdg: Llm-assisted procedural weather generation for domain-generalized semantic segmentation, https://arxiv.org/abs/2410.12075 (2024)
2024 arXiv
-
[206]
Toney, C
G. Toney, C. Bhargava, Adaptive headlamps in automobile: A review on the models, detection techniques, and mathematical models, IEEE Access 9 (2021) 87462–87474.doi:10.1109/ACCESS.2021.3088036
2021
-
[207]
Sakaridis, D
C. Sakaridis, D. Dai, L. Van Gool, Guided curriculum model adaptation and uncertainty-aware evaluation for semantic nighttime image segmentation, in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 7373–7382.doi:10.1109/ICCV. 2019.00747
2019
-
[209]
L. A. Zebrowitz, J. M. Montepare, Social psychological face perception: Why appearance matters., Social and personality psychology compass 2 3 (2008) 1497. URL https://api.semanticscholar.org/CorpusID:11974402
2008
-
[210]
Verma, S
H. Verma, S. Lotia, A. Singh, Convolutional neural network based criminal detection, in: 2020 IEEE REGION 10 CONFERENCE (TENCON), 2020, pp. 1124–1129.doi:10.1109/TENCON50793.2020.9293926
2020
-
[211]
Singh, Facial-recognition-for-crime-detection, https://github.com/Navu4/Facial-Recognition-for-Crime-Detection (2021)
N. Singh, Facial-recognition-for-crime-detection, https://github.com/Navu4/Facial-Recognition-for-Crime-Detection (2021)
2021
-
[212]
Anoop, H
A. Anoop, H. G, K. Nair, S. B, V. Praseedalekshmi, T. S. H, Traffic surveillance system and criminal detection using image processing and deep learning, in: 2022 International Conference on Innovations in Science and Technology for Sustainable Development (ICISTSD), 2022, pp. ...
2022
-
[213]
1–5.doi:10.1109/ICoNSIP49665.2022.10007501
V.Jagtap,D.Rajmane,V.N.More,R.C.Mahajan,Suspiciousvehiclerecognitionusingnumberplate,in:2022InternationalConferenceon Signal and Information Processing (IConSIP), 2022, pp. 1–5.doi:10.1109/ICoNSIP49665.2022.10007501
2022
-
[214]
contributors, Automatic number-plate recognition, https://en.wikipedia.org/w/index.php?title=Automatic_ number-plate_recognition&oldid=1255542574 (2024)
W. contributors, Automatic number-plate recognition, https://en.wikipedia.org/w/index.php?title=Automatic_ number-plate_recognition&oldid=1255542574 (2024)
2024
-
[215]
Ayman, Automated-car-damage-detection,https://github.com/basel-ay/Automated-Car-Damage-Detection (2023)
B. Ayman, Automated-car-damage-detection,https://github.com/basel-ay/Automated-Car-Damage-Detection (2023)
2023
-
[216]
Lim, Mask r-cnn for car damage detection and segmentation,https://github.com/louisyuzhe/car-damage-detector (2021)
Y. Lim, Mask r-cnn for car damage detection and segmentation,https://github.com/louisyuzhe/car-damage-detector (2021)
2021
-
[217]
Gupta, Automated car damage assessment,https://github.com/shubhi/car-damage-assessment (2020)
S. Gupta, Automated car damage assessment,https://github.com/shubhi/car-damage-assessment (2020)
2020
-
[218]
Dosovitskiy, G
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, V. Koltun, CARLA: An open urban driving simulator, in: Proceedings of the 1st Annual Conference on Robot Learning, 2017, pp. 1–16
2017
-
[219]
Kotseruba, A
I. Kotseruba, A. Rasouli, J. K. Tsotsos, Benchmark for evaluating pedestrian action prediction, in: 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 2021, pp. 1257–1267.doi:10.1109/WACV48630.2021.00130
2021
-
[220]
Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, Y. Sheikh, Openpose: Realtime multi-person 2d pose estimation using part affinity fields, IEEE Transactions on Pattern Analysis and Machine Intelligence 43 (1) (2021) 172–186.doi:10.1109/TPAMI.2019.2929257
2021
-
[221]
R. A. Guler, N. Neverova, I. Kokkinos, Densepose: Dense human pose estimation in the wild, 2018
2018
-
[222]
G. Zhu, S. Fan, H. Dai, E. S. L. Ho, Waymo-3dskelmo: A multi-agent 3d skeletal motion dataset for pedestrian interaction modeling in autonomous driving, in: Proceedings of the 33rd ACM International Conference on Multimedia, MM ’25, Association for Computing Machinery, New Yor...
2025
-
[223]
Zheng, X
J. Zheng, X. Shi, A. Gorban, J. Mao, Y. Song, C. R. Qi, T. Liu, V. Chari, A. Cornman, Y. Zhou, C. Li, D. Anguelov, Multi-modal 3d human pose estimation with 2d weak supervision in autonomous driving, arXiv (2021)
2021
-
[224]
Bauer, A
P. Bauer, A. Bouazizi, U. Kressel, F. B. Flohr, Weakly supervised multi-modal 3d human body pose estimation for autonomous driving, in: 2023 IEEE Intelligent Vehicles Symposium (IV), 2023, pp. 1–7.doi:10.1109/IV55152.2023.10186575
2023
-
[225]
S.-Y. Yu, A. V. Malawade, D. Muthirayan, P. P. Khargonekar, M. A. A. Faruque, Scene-graph augmented data-driven risk assessment of autonomous vehicle decisions, arXiv preprint arXiv:2009.06435 (2020)
2020 arXiv
-
[226]
University, Meaning of pedestrian in english,https://dictionary.cambridge.org/dictionary/english/pedestrian (2025)
C. University, Meaning of pedestrian in english,https://dictionary.cambridge.org/dictionary/english/pedestrian (2025)
2025
-
[227]
URL https://openreview.net/forum?id=OMOOO3ls6g
H.Wang,T.Li,Y.Li,L.Chen,C.Sima,Z.Liu,B.Wang,P.Jia,Y.Wang,S.Jiang,F.Wen,H.Xu,P.Luo,J.Yan,W.Zhang,H.Li,Openlane- v2: A topology reasoning benchmark for unified 3d HD mapping, in: Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 20...
2023
-
[228]
URL https://datasetninja.com/bdd100k
D.Ninja,Visualizationtoolsforbdd100k:Images100kdataset, https://datasetninja.com/bdd100k,visitedon2025-02-01(feb2025). URL https://datasetninja.com/bdd100k
-
[229]
F.Yu,H.Chen,X.Wang,W.Xian,Y.Chen,F.Liu,V.Madhavan,T.Darrell,Bdd100k, https://github.com/bdd100k/bdd100k(2020)
2020
-
[230]
Zhang, P
S. Zhang, P. Wang, Toolkit for apolloscape dataset,https://github.com/ApolloScapeAuto/dataset-api (2022)
2022
-
[231]
Garnett, R
N. Garnett, R. Cohen, T. Pe’er, R. Lahav, D. Levi, 3d-lanenet: End-to-end 3d multiple lane detection, in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 2921–2930.doi:10.1109/ICCV.2019.00301
2019
-
[232]
Y.Cai,W.Dong,Z.Liu,H.Wang,L.Chen,Homap:End-to-endvectorizedhdmapconstructionwithhigh-ordermodeling,IEEETransactions on Intelligent Vehicles 10 (5) (2025) 2987–2997.doi:10.1109/TIV.2024.3445374
2025
-
[233]
K. Wu, C. Yang, Z. Li, Interactionmap: Improving online vectorized hdmap construction with interaction, 2025, pp. 17176–17186.doi: 10.1109/CVPR52734.2025.01601. : Preprint submitted to Elsevier Page 83 of 85
2025
-
[234]
Y.Liu,T.Yuan,Y.Wang,Y.Wang,H.Zhao,Vectormapnet:end-to-endvectorizedhdmaplearning,in:Proceedingsofthe40thInternational Conference on Machine Learning, ICML’23, JMLR.org, 2023
2023
-
[235]
nuscenes ™ devkit, https://github.com/nutonomy/nuscenes-devkit (2020)
2020
-
[236]
S. Choi, J. Kim, H. Shin, J. W. Choi, Mask2map: Vectorized hd map construction using bird’s eye view segmentation masks, in: European Conference on Computer Vision, 2024
2024
-
[237]
Zhang, Y
Z. Zhang, Y. Zhang, X. Ding, F. Jin, X. Yue, Online vectorized hd map construction using geometry, in: Computer Vision – ECCV 2024: 18thEuropeanConference,Milan,Italy,September29–October4,2024,Proceedings,PartXLIX,Springer-Verlag,Berlin,Heidelberg,2024, p. 73–90. doi:10.1007/9...
2024 doi
-
[238]
Mogelmose, M
A. Mogelmose, M. M. Trivedi, T. B. Moeslund, Vision-based traffic sign detection and analysis for intelligent driver assistance systems: Perspectives and survey, IEEE Transactions on Intelligent Transportation Systems 13 (4) (2012) 1484–1497.doi:10.1109/TITS.2012. 2209421
2012 doi
-
[239]
Z. Zhu, D. Liang, S. Zhang, X. Huang, B. Li, S. Hu, Visualization tools for tsinghua tencent 2021 dataset,https://datasetninja.com/ tt100k-2021 (2025)
2025
-
[240]
citlag, Traffic-sign-recognition,https://github.com/citlag/Traffic-Sign-Recognition (2019)
2019
-
[241]
Vitas, M
D. Vitas, M. Tomic, M. Burul, Traffic light detection in autonomous driving systems, IEEE Consumer Electronics Magazine 9 (4) (2020) 90–96. doi:10.1109/MCE.2020.2969156
2020
-
[243]
Lateef, M
F. Lateef, M. Kas, Y. Ruichek, Saliency heat-map as visual attention for autonomous driving using generative adversarial network (gan), IEEE Transactions on Intelligent Transportation Systems 23 (6) (2022) 5360–5373.doi:10.1109/TITS.2021.3053178
2022
-
[244]
T. Deng, H. Yan, L. Qin, T. Ngo, B. S. Manjunath, How do drivers allocate their potential attention? driving fixation prediction via convolutional neural networks, IEEE Transactions on Intelligent Transportation Systems 21 (5) (2020) 2146–2154.doi:10.1109/TITS. 2019.2915540
2020
-
[245]
S. Gan, X. Pei, Y. Ge, Q. Wang, S. Shang, S. E. Li, B. Nie, Multisource adaption for driver attention prediction in arbitrary driving scenes, IEEE Transactions on Intelligent Transportation Systems 23 (11) (2022) 20912–20925.doi:10.1109/TITS.2022.3177640
2022
-
[246]
H.Kuang,K.-F.Yang,L.Chen,Y.-J.Li,L.L.H.Chan,H.Yan,Bayessaliency-basedobjectproposalgeneratorfornighttimetrafficimages, IEEE Transactions on Intelligent Transportation Systems 19 (3) (2018) 814–825.doi:10.1109/TITS.2017.2702665
2018
-
[247]
3225–3232.doi:10.1109/ITSC.2018.8569438
A.Tawari,P.Mallela,S.Martin,Learningtoattendtosalienttargetsindrivingvideosusingfullyconvolutionalrnn,in:201821stInternational Conference on Intelligent Transportation Systems (ITSC), 2018, pp. 3225–3232.doi:10.1109/ITSC.2018.8569438
2018
-
[248]
Biparva, D
M. Biparva, D. Fernández-Llorca, R. I. Gonzalo, J. K. Tsotsos, Video action recognition for lane-change classification and prediction of surrounding vehicles, IEEE Transactions on Intelligent Vehicles 7 (3) (2022) 569–578.doi:10.1109/TIV.2022.3164507
2022
-
[249]
Liang, J
K. Liang, J. Wang, A. Bhalerao, Lane change classification and prediction with action recognition networks, in: L. Karlinsky, T. Michaeli, K. Nishino (Eds.), Computer Vision – ECCV 2022 Workshops, Springer Nature Switzerland, Cham, 2023, pp. 617–632
2022
-
[251]
M. S. Kristoffersen, J. V. Dueholm, R. K. Satzoda, M. M. Trivedi, A. Møgelmose, T. B. Moeslund, Towards semantic understanding of surrounding vehicular maneuvers: A panoramic vision-based framework for real-world highway studies, in: 2016 IEEE Conference on Computer Vision and...
2016 doi
-
[252]
P. Shen, J. Fang, H. Yu, J. Xue, Vehicle behavior prediction by episodic-memory implanted ndt, in: 2024 IEEE International Conference on Robotics and Automation (ICRA), 2024, pp. 14177–14183.doi:10.1109/ICRA57147.2024.10610995
2024
-
[253]
J. Shi, J. Chen, Y. Wang, L. Sun, C. Liu, W. Xiong, T. Wo, Motion forecasting for autonomous vehicles: a survey, International Journal of Machine Learning and Cybernetics 17 (01 2026).doi:10.1007/s13042-025-02859-8
2026 doi
-
[254]
Teeti, S
I. Teeti, S. Khan, A. Shahbaz, A. Bradley, F. Cuzzolin, Vision-based intention and trajectory prediction in autonomous vehicles: A survey, in: L. D. Raedt (Ed.), Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22, International ...
2022 doi
-
[255]
Singh, Trajectory-prediction with vision: A survey, in: 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2023, pp
A. Singh, Trajectory-prediction with vision: A survey, in: 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2023, pp. 3310–3315.doi:10.1109/ICCVW60793.2023.00356
2023
-
[256]
W. Zhan, L. Sun, D. Wang, H. Shi, A. Clausse, M. Naumann, J. Kümmerle, H. Königshof, C. Stiller, A. de La Fortelle, M. Tomizuka, INTERACTIONDataset:AnINTERnational,AdversarialandCooperativemoTIONDatasetinInteractiveDrivingScenarioswithSemantic Maps, arXiv:1910.03088 [cs, eess]...
1910 arXiv
-
[257]
K. Chen, R. Ge, H. Qiu, R. Ai-Rfou, C. R. Qi, X. Zhou, Z. Yang, S. Ettinger, P. Sun, Z. Leng, M. Mustafa, I. Bogun, W. Wang, M. Tan, D. Anguelov, Womd-lidar: Raw sensor dataset benchmark for motion forecasting, in: Proceedings of the IEEE International Conference on Robotics a...
2024
-
[258]
J. Gu, C. Hu, T. Zhang, X. Chen, Y. Wang, Y. Wang, H. Zhao, ViP3D: End-to-End Visual Trajectory Prediction via 3D Agent Queries , in: 2023IEEE/CVFConferenceonComputerVisionandPatternRecognition(CVPR),IEEEComputerSociety,LosAlamitos,CA,USA,2023, pp. 5496–5506. doi:10.1109/CVPR5...
2023
-
[259]
J. Bock, R. Krajewski, T. Moers, S. Runde, L. Vater, L. Eckstein, The ind dataset: A drone dataset of naturalistic road user trajectories at germanintersections,in:2020IEEEIntelligentVehiclesSymposium(IV),2020,pp.1929–1934. doi:10.1109/IV47402.2020.9304839
2020
-
[260]
Z. Zhou, J. Wang, Y.-H. Li, Y.-K. Huang, Query-centric trajectory prediction, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 17863–17873. : Preprint submitted to Elsevier Page 84 of 85
2023
-
[261]
Z. Zhou, Z. Wen, J. Wang, Y.-H. Li, Y.-K. Huang, Qcnext: A next-generation framework for joint multi-agent trajectory prediction, arXiv preprint arXiv:2306.10508 (2023). : Preprint submitted to Elsevier Page 85 of 85
2023 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.