REVIEW 4 major objections 6 minor 74 references
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read DAVE, a 13-million-box Indian road dataset, claims that current perception methods degrade sharply on dense unstructured traffic where vulnerable road users make up 41.13% of instances.
desk verdict DAVE is a potentially valuable dataset for dense Asian traffic, but the paper's main 'more challenging' evidence has a label-space flaw and the dataset isn't released or quality-checked. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is DAVE itself, a manually annotated corpus of 1,231 one-minute dashcam videos from urban and semi-urban India in which every visible actor is treated as an atomic visual element. Its operational mechanism is the dual annotation: each of over 13 million bounding boxes gets an actor identity from 16 classes, and more than 1.6 million boxes additionally receive an action label from 16 classes such as cut-in, overtaking, zigzag movement, and U-turn. That combination lets one corpus feed five different video benchmarks, while the dense, unstructured scenes push models beyond the clean settings of Western traffic datasets.
What would settle it
Independently re-annotate a random sample of DAVE frames and measure inter-annotator agreement on bounding boxes and action labels; if agreement falls below standard thresholds (for example, mean IoU below 0.5 or action-label agreement below 80%), the ground truth and every reported benchmark number lose their foundation.
Extended reading notes
Core claim
The paper's central claim is that DAVE, the Diverse Atomic Visual Elements dataset, is a large-scale, manually annotated traffic-video benchmark from urban and semi-urban India that is substantially harder and more representative of vulnerable road users than existing datasets such as Waymo. DAVE contains 1,231 one-minute dashcam videos with over 13 million bounding boxes, more than 1.6 million of which also carry one of 16 action labels; actors span 16 categories including animals, motorized tricycles, scooters, and pedestrians. The paper reports that vulnerable road users account for 41.13% of instances, versus 23.14% in Waymo. Across five video tasks, current methods score far lower on DAVE than on their original benchmarks: ARTrack's SR0.75 drops 23.7% relative to GOT-10k, Swin-T reaches 32.5 mAP versus 50.5 on COCO, ACAR-Net gets 6.3% mAP versus 33.3% on AVA v2.2, CG-DETR attains 5.1 R1@0.5 versus 58.4 on Charades-STA, and SlowFast gets 41.0 mAP versus 45.2 on Charades. The conclusion drawn is that DAVE exposes a real perception gap for dense, unpredictable, rule-bending traffic and offers a testbed for building models that protect vulnerable road users.
Load-bearing premise
The load-bearing premise is that the manual annotations are accurate and consistent enough to serve as ground truth for all five benchmark tasks, and that the reported statistics (13,012,635 boxes, 41.13% vulnerable-road-user share, action-label counts) were computed correctly.
Editorial extensions
If this is right
- Training on DAVE's vulnerable-road-user instances instead of Waymo's lifts a YOLOv8 detector's mAP50 from 0.00266 to 0.235 on DAVE validation, and combining both datasets raises it further to 0.267.
- ARTrack's success rate at 0.75 IoU drops 23.7% on DAVE relative to GOT-10k, indicating that current trackers lose precise localization in dense, cluttered scenes.
- Methods on four other video tasks—Swin-T for detection, ACAR-Net for spatiotemporal action localization, CG-DETR for moment retrieval, and SlowFast for multi-label action recognition—all score far lower on DAVE than on their original benchmarks.
- Adding DAVE to an existing dataset like Waymo produces better vulnerable-road-user detection than either dataset alone, pointing toward merged, more globally representative training corpora.
Reading between the lines
- Editorial inference: if the annotations are released, DAVE could serve as a geographic domain-shift probe, quantifying how much each model's performance drops when moving from Western to Indian traffic.
- Editorial inference: because the paper reports no inter-annotator agreement, an independent re-annotation of a random sample would be the first validating experiment a user should run before trusting the benchmark numbers.
- Editorial inference: the action taxonomy covers vehicle maneuvers almost exclusively; extending it to pedestrian and animal behaviors would strengthen the safety argument implied by the high vulnerable-road-user share.
- Editorial inference: the recorded GPS and camera intrinsics could support trajectory forecasting and monocular 3D localization, tasks not benchmarked in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces DAVE, a manually annotated dashcam video dataset collected in India, with 16 actor categories, 16 action types, over 13 million bounding boxes, and over 1.6 million boxes that also carry action labels. The paper reports that VRUs constitute 41.13% of instances, compares this with Waymo, and evaluates existing models across five video tasks (tracking, detection, video moment retrieval, spatiotemporal action localization, multi-label video action recognition), finding large performance drops on DAVE. The authors conclude that DAVE is more challenging and more representative of unstructured traffic than existing datasets.
Significance. If the annotation statistics and benchmark protocols are accurate, DAVE addresses a real gap: most traffic datasets are Western, structured, and underrepresent VRUs, while DAVE provides dense, action-labeled footage of heterogeneous Indian road scenes. The breadth of actor categories, the weather/density variety, and the coverage of five tasks make it a potentially useful benchmark. However, the contribution is currently not verifiable: the manuscript contains no data release, no annotation-quality metrics, and several inconsistencies in core numbers, and the key Waymo comparison has a label-space confound. These are fixable with revision, after which the dataset could be a solid resource for the community.
major comments (4)
- [Section 3.2.1, Table 5] The quantitative claim that DAVE is more challenging than Waymo rests on the YOLOv8 VRU detection comparison, but the paper never specifies how the Waymo label space (pedestrian, cyclist) is mapped to DAVE's VRU classes (Animal, Bicycle, MotorBike, MotorizedTricycle, MultiWheeler, Scooter, TriCycle). If the Waymo-trained model is evaluated against all DAVE VRU classes, the 0.00266 mAP50 is largely attributable to classes never seen during training rather than to scene difficulty. The authors should either evaluate on the intersection of class labels (e.g., pedestrian/cyclist subsets) or report the class-wise mapping and per-class results; otherwise the 'more challenging' conclusion is not supported.
- [Section 2.2] The dataset description relies on manual annotation with CVAT but reports no inter-annotator agreement, no quality-control protocol, and no spot-check statistics. Since every benchmark result in Sections 3.1-3.5 is computed against these annotations, the absence of any annotation-quality evidence leaves the ground-truth validity of all reported numbers unestablished. Provide IAA on a random sample (e.g., bbox IoU agreement and label agreement per class/action), and state how ambiguous cases (occlusion, small objects) were handled.
- [Abstract, Section 2.1, Section 3.5, Tables 3, 6, 8] Core statistics are internally inconsistent. The abstract reports 23.71% VRU share in Waymo while the Introduction reports 23.14%; Section 2.1 states 1920x1080 capture resolution while Table 6 lists 1920x1280; Section 3.4/Table 8 reports 1,600k action-annotated boxes while the Introduction says 1.6 million; Table 3 gives no total count of action instances; and Section 2.1 says the dataset contains 1231 clips while Section 3.5 says the multi-label split has 10,083 clips (8,166 + 1,917). These discrepancies need to be reconciled and the exact computation of the 41.13% VRU share and the 13,012,635 total boxes stated before the quantitative claims can be trusted.
- [Data availability] The manuscript provides no download link, hosting plan, or license for DAVE, even though the contribution is a dataset. Without access, no external researcher can verify the 13M/1.6M box statistics, the split sizes, or any benchmark result. Provide a clear availability statement (including any privacy restrictions, since faces and plates are said to be blurred) and a data-release mechanism.
minor comments (6)
- [Table 3 and Table 9] 'Breaking' should be 'Braking', and 'Charedes' in Table 9's caption should be 'Charades'.
- [Section 2.1] The dashcam 'resolution of 2.3 megapixels' does not match the stated video resolution of 1920x1080 (2.07 MP); clarify whether the sensor resolution and output resolution differ.
- [Table 2] 'NuScense' should be 'nuScenes'.
- [Section 3.1, Table 4] Clarify the relationship between the 44.8k frame sequences and the 5,227 filtered validation sequences, since Table 4 lists 'Sequence number' as 44.8k while the text reports only the validation filtering.
- [Figure 3] Add axis labels and legends to the actor/action distribution panels; as rendered in the text, the quantitative distributions cannot be read from the figure.
- [Abstract/Keywords] The keyword 'Big data-driven models' is vague and could be removed or replaced with more specific terms such as 'action recognition' and 'object tracking'.
Circularity Check
No significant circularity: the dataset's statistics and benchmark degradation numbers are empirically measured, not derived from the paper's own assumptions or fitted parameters.
full rationale
DAVE is a dataset and benchmarking paper, so the derivation chain under review is not a theory-derived prediction. The central claims—over 13 million annotated boxes, 41.13% VRUs, and that existing models degrade on DAVE—are empirical measurements. The degradation numbers are produced by running externally released models (ARTrack-256, YOLOv8, Swin-T, CG-DETR, ACAR-Net, SlowFast) under stated training and evaluation protocols, so they do not reduce by construction to the dataset statistics. Several references cite the authors' prior work on dense heterogeneous traffic (e.g., [6], [9], [14], [35], [36]), but these are contextual and motivational; none is invoked as a uniqueness theorem or as the source of an ansatz that defines DAVE's annotations. The claim that DAVE is 'more challenging' than Waymo rests on Table 5, which may suffer from label-space mismatch between DAVE's eight VRU categories and Waymo's pedestrian/cyclist classes; if so, that is a validity or reproducibility concern, not circularity, because the performance gap is not logically forced by how the dataset was defined. No fitted parameter is renamed as a prediction, and no result is defined in terms of the conclusion it is used to support. Therefore no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption The dashcam videos from two vehicles in India are representative of complex and unpredictable Asian traffic.
- domain assumption Manual annotations are accurate enough to serve as ground truth for benchmarking.
Cite this review
Pith. "Pith review of DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments." pith.science (2026). https://pith.science/paper/3KKHXXGW
@misc{pith2026241220042,
author = {Pith},
title = {Pith review of: DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments},
year = {2026},
howpublished = {\url{https://pith.science/paper/3KKHXXGW}},
note = {Machine review of arXiv:2412.20042}
}
read the original abstract
Most existing traffic video datasets including Waymo are structured, focusing predominantly on Western traffic, which hinders global applicability. Specifically, most Asian scenarios are far more complex, involving numerous objects with distinct motions and behaviors. Addressing this gap, we present a new dataset, DAVE, designed for evaluating perception methods with high representation of Vulnerable Road Users (VRUs: e.g. pedestrians, animals, motorbikes, and bicycles) in complex and unpredictable environments. DAVE is a manually annotated dataset encompassing 16 diverse actor categories (spanning animals, humans, vehicles, etc.) and 16 action types (complex and rare cases like cut-ins, zigzag movement, U-turn, etc.), which require high reasoning ability. DAVE densely annotates over 13 million bounding boxes (bboxes) actors with identification, and more than 1.6 million boxes are annotated with both actor identification and action/behavior details. The videos within DAVE are collected based on a broad spectrum of factors, such as weather conditions, the time of day, road scenarios, and traffic density. DAVE can benchmark video tasks like Tracking, Detection, Spatiotemporal Action Localization, Language-Visual Moment retrieval, and Multi-label Video Action Recognition. Given the critical importance of accurately identifying VRUs to prevent accidents and ensure road safety, in DAVE, vulnerable road users constitute 41.13% of instances, compared to 23.71% in Waymo. DAVE provides an invaluable resource for the development of more sensitive and accurate visual perception algorithms in the complex real world. Our experiments show that existing methods suffer degradation in performance when evaluated on DAVE, highlighting its benefit for future video recognition research.
Figures
Reference graph
Works this paper leans on
-
[1]
In: Proceedings of the IEEE/CVF international conference on computer vision
Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., Gall, J.: Semantickitti: A dataset for semantic scene understanding of lidar sequences. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9297–9307 (2019)
work page 2019
-
[2]
Caesar, H., Bankiti, V ., Lang, A.H., V ora, S., Liong, V .E., Xu, Q., Krishnan, A., Pan, Y ., Baldan, G., Beijbom, O.: nuscenes: A multimodal dataset for autonomous driving. CVPR (2020)
work page 2020
-
[3]
In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Cai, Z., Vasconcelos, N.: Cascade r-cnn: Delving into high quality object detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 6154–6162 (2018)
work page 2018
-
[4]
In: European Conference on Computer Vision (ECCV)
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: European Conference on Computer Vision (ECCV). pp. 213–229. Springer (2020)
work page 2020
-
[5]
arXiv preprint arXiv:1907.06987 (2019)
Carreira, J., Noland, E., Hillier, C., Zisserman, A.: A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987 (2019)
arXiv 2019
-
[6]
Chandra, R.: Towards autonomous driving in dense, heterogeneous, and unstructured traffic. Ph.D. thesis, University of Maryland, College Park (2022)
work page 2022
-
[7]
In: 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Chandra, R., Bhattacharya, U., Bera, A., Manocha, D.: Densepeds: Pedestrian tracking in dense crowds using front-rvo and sparse features. In: 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 468–475. IEEE (2019)
work page 2019
-
[8]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Chandra, R., Bhattacharya, U., Bera, A., Manocha, D.: Traphic: Trajectory prediction in dense and heterogeneous traffic using weighted interactions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8483–8492 (2019)
work page 2019
Show all 74 references
-
[9]
In: 2020 IEEE International Conference on Robotics and Automation (ICRA)
Chandra, R., Bhattacharya, U., Randhavane, T., Bera, A., Manocha, D.: Roadtrack: Realtime tracking of road agents in dense and heterogeneous environments. In: 2020 IEEE International Conference on Robotics and Automation (ICRA). pp. 1270–1277. IEEE (2020)
2020
-
[10]
In: Proceedings of the 3rd ACM Computer Science in Cars Symposium
Chandra, R., Bhattacharya, U., Roncal, C., Bera, A., Manocha, D.: Robusttp: End-to-end trajectory prediction for heterogeneous road-agents in dense traffic with noisy sensor inputs. In: Proceedings of the 3rd ACM Computer Science in Cars Symposium. pp. 1–9 (2019)
2019
-
[11]
IEEE Robotics and Automation Letters 5(3), 4882–4890 (2020)
Chandra, R., Guan, T., Panuganti, S., Mittal, T., Bhattacharya, U., Bera, A., Manocha, D.: Forecast- ing trajectory and behavior of road-agents using spectral clustering in graph-lstms. IEEE Robotics and Automation Letters 5(3), 4882–4890 (2020)
2020
-
[12]
IEEE Robotics and Automation Letters 7(2), 2676–2683 (2022)
Chandra, R., Manocha, D.: Gameplan: Game-theoretic multi-agent planning with human drivers at intersections, roundabouts, and merging. IEEE Robotics and Automation Letters 7(2), 2676–2683 (2022)
2022
-
[13]
In: 2022 International Conference on Robotics and Automation (ICRA)
Chandra, R., Wang, M., Schwager, M., Manocha, D.: Game-theoretic planning for autonomous driving among risk-aware human drivers. In: 2022 International Conference on Robotics and Automation (ICRA). pp. 2876–2883 (2022)
2022
-
[14]
In: 2023 IEEE International Conference on Robotics and Automation (ICRA)
Chandra, R., Wang, X., Mahajan, M., Kala, R., Palugulla, R., Naidu, C., Jain, A., Manocha, D.: Meteor: A dense, heterogeneous, and unstructured traffic dataset with rare behaviors. In: 2023 IEEE International Conference on Robotics and Automation (ICRA). pp. 9169–9175. IEEE (2023)
2023
-
[15]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Chang, M.F., Lambert, J., Sangkloy, P., Singh, J., Bak, S., Hartnett, A., Wang, D., Carr, P., Lucey, S., Ramanan, D., et al.: Argoverse: 3d tracking and forecasting with rich maps. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 8748–...
2019
-
[16]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: The cityscapes dataset for semantic urban scene understanding. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3213–3223 (2016)
2016
-
[17]
In: Proceedings of the IEEE/CVF winter conference on applications of computer vision
Corona, K., Osterdahl, K., Collins, R., Hoogs, A.: Meva: A large-scale multiview, multimodal video dataset for activity detection. In: Proceedings of the IEEE/CVF winter conference on applications of computer vision. pp. 1060–1068 (2021)
2021
-
[18]
CV AT.ai Corporation: Computer Vision Annotation Tool (CV AT) (Nov 2023),https://github.com/ opencv/cvat
2023
-
[19]
Dave, A., Khurana, T., Tokmakov, P., Schmid, C., Ramanan, D.: Tao: A large-scale benchmark for tracking any object (2020)
2020
-
[20]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Deng, J., Guo, J., Ververas, E., Kotsia, I., Zafeiriou, S.: Retinaface: Single-shot multi-level face localisation in the wild. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 5203–5212 (2020)
2020
-
[21]
Ding, L., Terwilliger, J., Sherony, R., Reimer, B., Fridman, L.: Mit driveseg (manual) dataset for dynamic driving scene segmentation. Tech. Rep., Technical report, Massachusetts Institute of Technology (2020)
2020
-
[22]
International journal of computer vision 88(2), 303–338 (2010)
Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. International journal of computer vision 88(2), 303–338 (2010)
2010
-
[23]
International Journal of Computer Vision 129, 439–461 (2021)
Fan, H., Bai, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Harshit, Huang, M., Liu, J., et al.: Lasot: A high-quality large-scale single object tracking benchmark. International Journal of Computer Vision 129, 439–461 (2021)
2021
-
[24]
In: IEEE Interna- tional Conference on Computer Vision (ICCV)
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition. In: IEEE Interna- tional Conference on Computer Vision (ICCV). pp. 6202–6211 (2019)
2019
-
[25]
Geyer, J., Kassahun, Y ., Mahmudi, M., Ricou, X., Durgesh, R., Chung, A.S., Hauswald, L., Pham, V .H., Mühlegg, M., Dorn, S., Fernandez, T., Jänicke, M., Mirashi, S., Savani, C., Sturm, M., V orobiov, O., Oelker, M., Garreis, S., Schuberth, P.: A2d2: Audi autonomous driving da...
2020 arXiv
-
[26]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Gu, C., Sun, C., Ross, D.A., V ondrick, C., Pantofaru, C., Li, Y ., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., et al.: Ava: A video dataset of spatio-temporally localized atomic visual actions. In: Proceedings of the IEEE Conference on Computer Vision and P...
2018
-
[27]
In: Proceedings of the IEEE/CVF winter conference on applications of computer vision
Guan, T., Wang, J., Lan, S., Chandra, R., Wu, Z., Davis, L., Manocha, D.: M3detr: Multi-representation, multi-scale, mutual-relation 3d object detection with transformers. In: Proceedings of the IEEE/CVF winter conference on applications of computer vision. pp. 772–782 (2022)
2022
-
[28]
In: Proceedings of the IEEE International Conference on Computer Vision
Hendricks, L.A., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.: Localizing moments in video with natural language. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 5803–5812 (2017)
2017
-
[29]
IEEE transactions on pattern analysis and machine intelligence 43(5), 1562–1577 (2019)
Huang, L., Zhao, X., Huang, K.: Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE transactions on pattern analysis and machine intelligence 43(5), 1562–1577 (2019)
2019
-
[30]
IEEE Trans- actions on Pattern Analysis and Machine Intelligence 42(10), 2702 (2020)
HuangXinyu, W., et al.: Theapolloscape opendatasetforautonomousdrivinganditsapplication. IEEE Trans- actions on Pattern Analysis and Machine Intelligence 42(10), 2702 (2020)
2020
-
[31]
in the wild
Idrees, H., Zamir, A.R., Jiang, Y .G., Gorban, A., Laptev, I., Sukthankar, R., Shah, M.: The thumos challenge on action recognition for videos “in the wild”. Computer Vision and Image Understanding 155, 1–23 (2017)
2017
-
[32]
In: Proceedings of the IEEE International Conference on Computer Vision
Jhuang, H., Gall, J., Zuffi, S., Schmid, C., Black, M.J.: Towards understanding action recognition. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 3192–3199 (2013)
2013
-
[33]
Jocher, G., Chaurasia, A., Qiu, J.: Ultralytics yolov8 (2023), https://github.com/ultralytics/ ultralytics
2023
-
[34]
arXiv preprint arXiv:1705.06950 (2017)
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al.: The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017)
2017 arXiv
-
[35]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Kothandaraman, D., Chandra, R., Manocha, D.: Bomudanet: unsupervised adaptation for visual scene understanding in unstructured driving environments. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 3966–3975 (2021) 15
2021
-
[36]
In: Proceedings of the IEEE/CVF international conference on computer vision
Kothandaraman, D., Chandra, R., Manocha, D.: Ss-sfda: Self-supervised source-free domain adaptation for road segmentation in hazardous environments. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 3049–3059 (2021)
2021
-
[37]
In: Proceedings of the IEEE International Conference on Computer Vision (2017)
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Carlos Niebles, J.: Dense-captioning events in videos. In: Proceedings of the IEEE International Conference on Computer Vision (2017)
2017
-
[38]
In: Proceedings of the IEEE international conference on computer vision workshops
Kristan, M., Matas, J., Leonardis, A., Felsberg, M., Cehovin, L., Fernandez, G., V ojir, T., Hager, G., Nebehay, G., Pflugfelder, R.: The visual object tracking vot2015 challenge results. In: Proceedings of the IEEE international conference on computer vision workshops. pp. 1–...
2015
-
[39]
In: European Conference on Computer Vision
Lei, J., Li, L., Zhou, L., Gan, Z., Berg, T.L., Bansal, M., Luo, J.: Tvr: A large-scale dataset for video-subtitle moment retrieval. In: European Conference on Computer Vision. Springer (2020)
2020
-
[40]
arXiv preprint arXiv:2103.07514 (2021)
Li, Y ., Xu, J., Qiu, Z., Tian, Y ., Hu, T., Wang, J., Li, J., Xiong, H., Hauptmann, A.G.: Multisports: A multi-person video dataset for action spotting, localization, and detection. arXiv preprint arXiv:2103.07514 (2021)
2021 arXiv
-
[41]
IEEE Transactions on Pattern Analysis and Machine Intelligence 45(3), 3292–3310 (2022)
Liao, Y ., Xie, J., Geiger, A.: Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(3), 3292–3310 (2022)
2022
-
[42]
In: European conference on computer vision
Lin, T.Y ., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European conference on computer vision. pp. 740–755. Springer (2014)
2014
-
[43]
European conference on computer vision (2014)
Lin, T.Y ., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. European conference on computer vision (2014)
2014
-
[44]
arXiv preprint arXiv:2302.07919 (2023)
Liu, F., Yacoob, Y ., Shrivastava, A.: Covid-vts: Fact extraction and verification on short video platforms. arXiv preprint arXiv:2302.07919 (2023)
2023 arXiv
-
[45]
In: IEEE International Conference on Computer Vision (ICCV)
Liu, Z., Lin, Y ., Cao, Y ., Hu, H., Wei, Y ., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: IEEE International Conference on Computer Vision (ICCV). pp. 10012–10022 (2021)
2021
-
[46]
arXiv preprint arXiv:2402.1413 (2024)
Liu, Z., Yang, Z., Xu, X., Jiang, C., Li, X., He, X.: Gdtm: An indoor geospatial tracking dataset with distributed multimodal sensors. arXiv preprint arXiv:2402.1413 (2024)
2024
-
[47]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Malla, S., Dariush, B., Choi, C.: Titan: Future forecast using action priors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11186–11196 (2020)
2020
-
[48]
arXiv preprint arXiv:1603.00831 (2016)
Milan, A., Leal-Taixé, L., Reid, I., Roth, S., Schindler, K.: Mot17: A benchmark for multi-object tracking. arXiv preprint arXiv:1603.00831 (2016)
2016 arXiv
-
[49]
Moon, W., Hyun, S., Lee, S., Heo, J.P.: Correlation-guided query-dependency calibration in video representation learning for temporal grounding (2023)
2023
-
[50]
In: Proceedings of the European conference on computer vision (ECCV)
Muller, M., Bibi, A., Giancola, S., Alsubaihi, S., Ghanem, B.: Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In: Proceedings of the European conference on computer vision (ECCV). pp. 300–317 (2018)
2018
-
[51]
In: CVPR
Oh, S., Hoogs, A., Perera, A., Cuntoor, N., Chen, C.C., Lee, J.T., Mukherjee, S., Aggarwal, J., Lee, H., Davis, L., et al.: A large-scale benchmark dataset for event recognition in surveillance video. In: CVPR
-
[52]
In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Pan, J., Chen, S., Shou, M.Z., Liu, Y ., Shao, J., Li, H.: Actor-context-actor relation network for spatio- temporal action localization. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 464–474 (2021)
2021
-
[53]
arXiv preprint arXiv:2405.05354 (2024)
Parikh, C., Mishra, R.S., Chandra, R., Sarvadevabhatla, R.K.: Transfer-lmr: Heavy-tail driving behavior recognition in diverse traffic scenarios. arXiv preprint arXiv:2405.05354 (2024)
2024 arXiv
-
[54]
In: 2019 International Conference on Robotics and Automation (ICRA)
Patil, A., Malla, S., Gang, H., Chen, Y .T.: The h3d dataset for full-surround 3d multi-object detection and tracking in crowded urban scenes. In: 2019 International Conference on Robotics and Automation (ICRA). pp. 9552–9557. IEEE (2019)
2019
-
[55]
In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2024), https://openreview.net/forum?id=5L05sLRIlQ 16
Peng, L., Gao, J., Liu, X., Li, W., Dong, S., Zhang, Z., Fan, H., Zhang, L.: Vasttrack: Vast category visual object tracking. In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2024), https://openreview.net/forum?id=5L05sLRIlQ 16
2024
-
[56]
In: 2020 IEEE International conference on Robotics and Automation (ICRA)
Pham, Q.H., Sevestre, P., Pahwa, R.S., Zhan, H., Pang, C.H., Chen, Y ., Mustafa, A., Chandrasekhar, V ., Lin, J.: A* 3d dataset: Towards autonomous driving in challenging environments. In: 2020 IEEE International conference on Robotics and Automation (ICRA). pp. 2267–2273. IEEE (2020)
2020
-
[57]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Rasouli, A., Kotseruba, I., Kunic, T., Tsotsos, J.K.: Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 6262–6271 (2019)
2019
-
[58]
Transactions of the Association for Computational Linguistics 1, 25–36 (2013)
Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., Pinkal, M.: Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics 1, 25–36 (2013)
2013
-
[59]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Ros, G., Sellart, L., Materzynska, J., Vazquez, D., Lopez, A.M.: The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3234–3243 (2016)
2016
-
[60]
European Conference on Computer Vision (2016)
Sigurdsson, G.A., Varol, G.V ., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: Hollywood in homes: Crowdsourcing data collection for activity understanding. European Conference on Computer Vision (2016)
2016
-
[61]
In: European Conference on Computer Vision
Sigurdsson, G.A., Varol, G.V ., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: Hollywood in homes: Crowdsourcing data collection for activity understanding. In: European Conference on Computer Vision. pp. 510–526. Springer (2016)
2016
-
[62]
IEEE transactions on pattern analysis and machine intelligence 45(1), 1036–1054 (2022)
Singh, G., Akrigg, S., Di Maio, M., Fontana, V ., Alitappeh, R.J., Khan, S., Saha, S., Jeddisaravi, K., Yousefi, F., Culley, J., et al.: Road: The road event awareness dataset for autonomous driving. IEEE transactions on pattern analysis and machine intelligence 45(1), 1036–10...
2022
-
[63]
arXiv preprint arXiv:1212.0402 (2012)
Soomro, K., Zamir, A.R., Shah, M.: Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)
2012 arXiv
-
[64]
CVPR (2020)
Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V ., Tsui, P., Guo, J., Zhou, Y ., Chai, Y ., Caine, B., et al.: Scalability in perception for autonomous driving: Waymo open dataset. CVPR (2020)
2020
-
[65]
Sun, S., Akhtar, N., Song, H., Mian, A., Shah, M.: Deep affinity network for multiple object tracking (2019)
2019
-
[66]
Wang, J., Jiang, W., Ma, L., Liu, W., Xu, Y .: Bidirectional attentive fusion with context gating for dense video captioning (2018)
2018
-
[67]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Wei, X., Bai, Y ., Zheng, Y ., Shi, D., Gong, Y .: Autoregressive visual tracking. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9697–9706 (2023)
2023
-
[68]
In: IEEE transactions on pattern analysis and machine intelligence
Wu, Y ., Lim, J., Yang, M.H.: Object tracking benchmark. In: IEEE transactions on pattern analysis and machine intelligence. vol. 37, pp. 1834–1848 (2013)
2013
-
[69]
In: https://github.com/szad670401/HyperLPR (2023)
Yan, j., Yu, J., Xiaoxiao: Hyperlpr3 - high performance license plate recognition framework. In: https://github.com/szad670401/HyperLPR (2023)
2023
-
[70]
arXiv preprint arXiv:2004.03044 (2020)
Yao, Y ., Wang, X., Xu, M., Pu, Z., Atkins, E., Crandall, D.: When, where, and what? a new dataset for anomaly detection in driving videos. arXiv preprint arXiv:2004.03044 (2020)
2020 arXiv
-
[71]
In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Zhan, X., Wu, Q., Wang, M., Sun, J., He, T.: Lasot: A large-scale high-resolution benchmark for visual object tracking. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 4641–4650 (2019)
2019
-
[72]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Zhang, X., Yu, F., He, B., Yang, M., Li, X., Zhu, J., Wang, J.: Posetrack: A dataset for person search, multi-object tracking and multi-person pose tracking. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 14077–14086 (2021)
2021
-
[73]
In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Zheng, Z., Tang, S., Guo, Y ., Liu, Y ., Zhou, S., Li, H.: Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5216–5225 (2023) 17
2023
-
[2011]
3153–3160
pp. 3153–3160. IEEE (2011)
2011
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.