{"id":"dca10d72-92bc-4745-a798-cf819d136ca5","arxiv_id":"2412.20042","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DAVE is a manually annotated Indian traffic video dataset with 16 actor types and 16 action types, meant to benchmark perception in complex, vulnerable-road-user-heavy environments.","lead":"DAVE is a new video dataset capturing chaotic Indian traffic, with over 13 million annotated objects and a high share of pedestrians, animals, and two-wheelers. It is designed to test whether perception models that work well on tidy Western roads can handle dense, unpredictable Asian driving.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest evidence for 'more challenging than Waymo' uses disjoint label sets; a category-matched control is needed before the difficulty claim is accepted.","rationale":"The reader's conditional verdict is appropriate: the dataset may be valuable, but the manuscript currently does not provide enough evidence to verify its central quantitative claims. My stress-test identifies a more specific load-bearing flaw than the reader's stated weakest assumption. Table 5 is the only controlled comparison against Waymo, and its validity depends on label-space alignment that the paper does not document. If the Waymo-trained model is evaluated on DAVE VRU categories it never saw, the 20.8% gap cannot be interpreted as DAVE being more challenging. This is fixable by a class-matched evaluation, and it does not require rejecting the underlying dataset. I therefore keep the verdict unchanged. I mark agreement as partial because the reader flagged missing annotation quality metrics and data release, which are related verification gaps, but did not identify the concrete class-mismatch problem in the head-to-head detection experiment.","tokens_in":16869,"tokens_out":10965,"duration_ms":119359,"concrete_test":"Reproduce the Section 3.2.1 experiment with a matched label space: map DAVE's Bicycle/MotorBike/Scooter/TriCycle/MotorizedTricycle/MultiWheeler onto a single 'two-wheeler/VRU' class, or restrict evaluation to the pedestrian/cyclist classes shared with Waymo, and retrain and evaluate YOLOv8 under identical settings. If the Waymo-trained model's mAP50 remains near 0.00266 on the shared classes, the difficulty claim survives; if mAP50 rises substantially, the reported 20.8% gap is an artifact of disjoint categories and the 'more challenging' claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main controlled evidence that DAVE is harder than existing traffic data is the VRU detection experiment in Section 3.2.1 and Table 5, where a YOLOv8 trained on Waymo reaches mAP50=0.00266 on the DAVE validation set versus 0.235 for a model trained on DAVE. This comparison is only meaningful if the two training sets share a label space. Waymo's VRU labels are pedestrian and cyclist, while DAVE's VRU categories include Animal, Bicycle, MotorBike, MotorizedTricycle, MultiWheeler, Scooter, and TriCycle (Tables 1 and 3). The paper never states how class IDs were aligned, whether DAVE-only VRU classes were merged into Waymo labels, or whether the evaluation map excludes classes unseen in the Waymo training set. If the Waymo-trained model is evaluated against all DAVE VRU classes, the near-zero mAP is at least partly a label-space mismatch rather than evidence of scene difficulty. Since this table is the only head-to-head Waymo-versus-DAVE comparison, the quantitative 'more challenging' conclusion is not yet established. Companion concerns, such as unreported inter-annotator agreement, no data release, and conflicting action/frame statistics in Table 3 versus the 1.6M action boxes used in Section 3.4, reinforce that the dataset's central quantitative claims are not independently checkable from the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces DAVE, a manually annotated dashcam video dataset collected in India, with 16 actor categories, 16 action types, over 13 million bounding boxes, and over 1.6 million boxes that also carry action labels. The paper reports that VRUs constitute 41.13% of instances, compares this with Waymo, and evaluates existing models across five video tasks (tracking, detection, video moment retrieval, spatiotemporal action localization, multi-label video action recognition), finding large performance drops on DAVE. The authors conclude that DAVE is more challenging and more representative of unstructured traffic than existing datasets.","tokens_in":17115,"tokens_out":5782,"duration_ms":52995,"significance":"If the annotation statistics and benchmark protocols are accurate, DAVE addresses a real gap: most traffic datasets are Western, structured, and underrepresent VRUs, while DAVE provides dense, action-labeled footage of heterogeneous Indian road scenes. The breadth of actor categories, the weather/density variety, and the coverage of five tasks make it a potentially useful benchmark. However, the contribution is currently not verifiable: the manuscript contains no data release, no annotation-quality metrics, and several inconsistencies in core numbers, and the key Waymo comparison has a label-space confound. These are fixable with revision, after which the dataset could be a solid resource for the community.","major_comments":[{"comment":"The quantitative claim that DAVE is more challenging than Waymo rests on the YOLOv8 VRU detection comparison, but the paper never specifies how the Waymo label space (pedestrian, cyclist) is mapped to DAVE's VRU classes (Animal, Bicycle, MotorBike, MotorizedTricycle, MultiWheeler, Scooter, TriCycle). If the Waymo-trained model is evaluated against all DAVE VRU classes, the 0.00266 mAP50 is largely attributable to classes never seen during training rather than to scene difficulty. The authors should either evaluate on the intersection of class labels (e.g., pedestrian/cyclist subsets) or report the class-wise mapping and per-class results; otherwise the 'more challenging' conclusion is not supported.","section":"Section 3.2.1, Table 5"},{"comment":"The dataset description relies on manual annotation with CVAT but reports no inter-annotator agreement, no quality-control protocol, and no spot-check statistics. Since every benchmark result in Sections 3.1-3.5 is computed against these annotations, the absence of any annotation-quality evidence leaves the ground-truth validity of all reported numbers unestablished. Provide IAA on a random sample (e.g., bbox IoU agreement and label agreement per class/action), and state how ambiguous cases (occlusion, small objects) were handled.","section":"Section 2.2"},{"comment":"Core statistics are internally inconsistent. The abstract reports 23.71% VRU share in Waymo while the Introduction reports 23.14%; Section 2.1 states 1920x1080 capture resolution while Table 6 lists 1920x1280; Section 3.4/Table 8 reports 1,600k action-annotated boxes while the Introduction says 1.6 million; Table 3 gives no total count of action instances; and Section 2.1 says the dataset contains 1231 clips while Section 3.5 says the multi-label split has 10,083 clips (8,166 + 1,917). These discrepancies need to be reconciled and the exact computation of the 41.13% VRU share and the 13,012,635 total boxes stated before the quantitative claims can be trusted.","section":"Abstract, Section 2.1, Section 3.5, Tables 3, 6, 8"},{"comment":"The manuscript provides no download link, hosting plan, or license for DAVE, even though the contribution is a dataset. Without access, no external researcher can verify the 13M/1.6M box statistics, the split sizes, or any benchmark result. Provide a clear availability statement (including any privacy restrictions, since faces and plates are said to be blurred) and a data-release mechanism.","section":"Data availability"}],"minor_comments":[{"comment":"'Breaking' should be 'Braking', and 'Charedes' in Table 9's caption should be 'Charades'.","section":"Table 3 and Table 9"},{"comment":"The dashcam 'resolution of 2.3 megapixels' does not match the stated video resolution of 1920x1080 (2.07 MP); clarify whether the sensor resolution and output resolution differ.","section":"Section 2.1"},{"comment":"'NuScense' should be 'nuScenes'.","section":"Table 2"},{"comment":"Clarify the relationship between the 44.8k frame sequences and the 5,227 filtered validation sequences, since Table 4 lists 'Sequence number' as 44.8k while the text reports only the validation filtering.","section":"Section 3.1, Table 4"},{"comment":"Add axis labels and legends to the actor/action distribution panels; as rendered in the text, the quantitative distributions cannot be read from the figure.","section":"Figure 3"},{"comment":"The keyword 'Big data-driven models' is vague and could be removed or replaced with more specific terms such as 'action recognition' and 'object tracking'.","section":"Abstract/Keywords"}],"recommendation":"major_revision","confidential_remarks":"The absence of a data release is unusual for a dataset paper and may be a dealbreaker at some venues; however, since the benchmark protocols are standard and the limitations are fixable in revision, I recommend major revision rather than rejection. The authors should be asked to provide the data or a clear public release plan, and to resolve the label-mapping issue in the Waymo comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [Colleague],\n\nThe short version: DAVE looks like a genuinely useful dataset for perception in dense Indian traffic—16 actor categories including animals and motorized tricycles, 16 action types, 13M+ boxes—but the paper as written doesn't yet establish the central claim that it's more challenging than existing datasets. The main controlled comparison is the VRU detection experiment in Table 5, and the stress-test note is right: Waymo's VRU labels are pedestrian and cyclist, while DAVE's VRU set includes animal, scooter, motorized tricycle, etc. If the Waymo-trained YOLO is evaluated on all DAVE VRU classes, the near-zero mAP is at least partly a label-space mismatch, not evidence of scene difficulty. The paper never explains how class IDs were aligned. So that specific claim is not supported.\n\nWhat's new and good: the combination of dense annotations, high VRU fraction, and rare maneuvers in a non-Western setting is genuinely missing from public benchmarks. The five-task coverage is broad, and the experimental setups mostly follow standard protocols. I can believe this becomes a reference benchmark if the data is released and validated.\n\nSoft spots, in order of importance. (1) No data link, no inter-annotator agreement, no QC metrics. For a dataset paper that's a serious gap. (2) METEOR, the same group's earlier dense Indian traffic dataset, is cited in the references but never discussed or compared. That's a major omission, especially since it looks like the direct predecessor. (3) Minor numeric inconsistencies: Waymo's VRU share appears as both 23.71% and 23.14%, and the resolution is 1920×1080 in one place and 1920×1280 in another. Easy fixes, but noticeable. (4) The cross-dataset comparisons for the other four tasks (tracking vs GOT-10k, detection vs COCO, etc.) are indicative but not controlled, so calling DAVE 'harder' based on them is overclaiming.\n\nMy recommendation: send it to peer review, but with the expectation of major revision. Require a data release statement, annotation quality metrics, a proper class-mapping explanation for the cross-dataset detection experiment, and a discussion of METEOR. If the data is real and released, this has real value; without those changes, it's an unverifiable promise.\n\nFor the reading group, I'd say maybe—worth a look once the revisions clarify the numbers.\n\nBest,","headline":"DAVE is a potentially valuable dataset for dense Asian traffic, but the paper's main 'more challenging' evidence has a label-space flaw and the dataset isn't released or quality-checked.","tokens_in":17678,"tokens_out":4076,"would_cite":true,"duration_ms":38662,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DAVE, a 13-million-box Indian road dataset, claims that current perception methods degrade sharply on dense unstructured traffic where vulnerable road users make up 41.13% of instances.","keywords":["traffic video dataset","vulnerable road users","unstructured environments","dense bounding-box annotations","video action recognition","autonomous driving perception","domain shift","India road scenes"],"falsifier":"Independently re-annotate a random sample of DAVE frames and measure inter-annotator agreement on bounding boxes and action labels; if agreement falls below standard thresholds (for example, mean IoU below 0.5 or action-label agreement below 80%), the ground truth and every reported benchmark number lose their foundation.","tokens_in":16679,"feed_emoji":"🚸","tokens_out":11801,"duration_ms":101936,"temperature":0.7,"pith_summary":"DAVE is a new traffic-video dataset built from Indian road scenes, where traffic is denser, less structured, and more mixed than the Western highways in Waymo. The paper contends that state-of-the-art perception models trained on existing datasets degrade substantially when evaluated on DAVE, and that the dataset's 41.13% share of vulnerable road users makes it a more safety-relevant test than Waymo's 23.14%. If that is right, DAVE gives the field a way to measure how well detection, tracking, and video-understanding systems generalize to the kind of heterogeneous, rule-bending traffic that dominates much of the world. The practical stake is that perception which fails on these scenes is not ready for global deployment.","feed_headline":"New India road dataset stumps today's traffic-vision models","feed_subtitle":"Waymo-trained models lose accuracy on this dense Indian road set, where 41% of boxes are pedestrians, bikes, and animals.","key_machinery":"The central object is DAVE itself, a manually annotated corpus of 1,231 one-minute dashcam videos from urban and semi-urban India in which every visible actor is treated as an atomic visual element. Its operational mechanism is the dual annotation: each of over 13 million bounding boxes gets an actor identity from 16 classes, and more than 1.6 million boxes additionally receive an action label from 16 classes such as cut-in, overtaking, zigzag movement, and U-turn. That combination lets one corpus feed five different video benchmarks, while the dense, unstructured scenes push models beyond the clean settings of Western traffic datasets.","core_discovery":"The paper's central claim is that DAVE, the Diverse Atomic Visual Elements dataset, is a large-scale, manually annotated traffic-video benchmark from urban and semi-urban India that is substantially harder and more representative of vulnerable road users than existing datasets such as Waymo. DAVE contains 1,231 one-minute dashcam videos with over 13 million bounding boxes, more than 1.6 million of which also carry one of 16 action labels; actors span 16 categories including animals, motorized tricycles, scooters, and pedestrians. The paper reports that vulnerable road users account for 41.13% of instances, versus 23.14% in Waymo. Across five video tasks, current methods score far lower on DAVE than on their original benchmarks: ARTrack's SR0.75 drops 23.7% relative to GOT-10k, Swin-T reaches 32.5 mAP versus 50.5 on COCO, ACAR-Net gets 6.3% mAP versus 33.3% on AVA v2.2, CG-DETR attains 5.1 R1@0.5 versus 58.4 on Charades-STA, and SlowFast gets 41.0 mAP versus 45.2 on Charades. The conclusion drawn is that DAVE exposes a real perception gap for dense, unpredictable, rule-bending traffic and offers a testbed for building models that protect vulnerable road users.","pith_inferences":["Editorial inference: if the annotations are released, DAVE could serve as a geographic domain-shift probe, quantifying how much each model's performance drops when moving from Western to Indian traffic.","Editorial inference: because the paper reports no inter-annotator agreement, an independent re-annotation of a random sample would be the first validating experiment a user should run before trusting the benchmark numbers.","Editorial inference: the action taxonomy covers vehicle maneuvers almost exclusively; extending it to pedestrian and animal behaviors would strengthen the safety argument implied by the high vulnerable-road-user share.","Editorial inference: the recorded GPS and camera intrinsics could support trajectory forecasting and monocular 3D localization, tasks not benchmarked in the paper."],"forward_implications":["Training on DAVE's vulnerable-road-user instances instead of Waymo's lifts a YOLOv8 detector's mAP50 from 0.00266 to 0.235 on DAVE validation, and combining both datasets raises it further to 0.267.","ARTrack's success rate at 0.75 IoU drops 23.7% on DAVE relative to GOT-10k, indicating that current trackers lose precise localization in dense, cluttered scenes.","Methods on four other video tasks—Swin-T for detection, ACAR-Net for spatiotemporal action localization, CG-DETR for moment retrieval, and SlowFast for multi-label action recognition—all score far lower on DAVE than on their original benchmarks.","Adding DAVE to an existing dataset like Waymo produces better vulnerable-road-user detection than either dataset alone, pointing toward merged, more globally representative training corpora."],"supporting_citations":[{"why":"Provides the Waymo dataset, which supplies the main comparison for vulnerable-road-user share (23.14%) and the Western-trained baseline in detection transfer experiments.","marker":"[64]"},{"why":"Supplies GOT-10k, the tracking benchmark whose success-rate numbers are compared against DAVE's for ARTrack.","marker":"[29]"},{"why":"Supplies ARTrack, the tracker whose performance on DAVE quantifies the dataset's tracking difficulty.","marker":"[67]"},{"why":"Supplies COCO, the detection benchmark and pretraining source whose Swin-T mAP of 50.5 is compared with DAVE's 32.5.","marker":"[42]"},{"why":"Supplies Swin-T, the detector backbone used to measure DAVE's detection challenge.","marker":"[45]"},{"why":"Supplies AVA v2.2, the spatiotemporal action localization benchmark whose 33.3% mAP for ACAR-Net is compared with DAVE's 6.3%.","marker":"[26]"},{"why":"Supplies ACAR-Net, the action localization method run on DAVE to produce the spatiotemporal action localization difficulty result.","marker":"[52]"},{"why":"Supplies CG-DETR, the moment retrieval model whose R1@0.5 of 5.1 on DAVE versus 58.4 on Charades-STA is the video moment retrieval challenge result.","marker":"[49]"},{"why":"Supplies Charades-STA, the video moment retrieval comparison dataset used in the CG-DETR experiments.","marker":"[60]"},{"why":"Supplies Charades, the multi-label action recognition comparison dataset against which SlowFast's 41.0 mAP on DAVE is measured.","marker":"[61]"}],"fun_headline_variants":["AI models stumble on dense Indian traffic benchmark","Vulnerable road users reveal AI's big blind spot","India's chaotic roads crush top detection models","New DAVE dataset: Where advanced AI breaks down"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the manual annotations are accurate and consistent enough to serve as ground truth for all five benchmark tasks, and that the reported statistics (13,012,635 boxes, 41.13% vulnerable-road-user share, action-label counts) were computed correctly.","fun_headline_variants_meta":{"raw":{"variants":["AI models stumble on dense Indian traffic benchmark","Vulnerable road users reveal AI's big blind spot","India's chaotic roads crush top detection models","New DAVE dataset: Where advanced AI breaks down"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1604,"prompt_tokens":1170,"completion_tokens":434,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":786,"completion_tokens_details":{"reasoning_tokens":374}},"tokens_in":786,"tokens_out":434,"duration_ms":4856,"temperature":1.0,"reasoning_tokens":374,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:36:22.045896+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-annotate a random sample of DAVE frames and measure inter-annotator agreement on bounding boxes and action labels; if agreement falls below standard thresholds (for example, mean IoU below 0.5 or action-label agreement below 80%), the ground truth and every reported benchmark number lose their foundation.","supporting_citations":[{"cited_title":"CVPR (2020)","cited_arxiv_id":null,"evidence_quote":"Provides the Waymo dataset, which supplies the main comparison for vulnerable-road-user share (23.14%) and the Western-trained baseline in detection transfer experiments."},{"cited_title":"IEEE transactions on pattern analysis and machine intelligence 43(5), 1562–1577 (2019)","cited_arxiv_id":null,"evidence_quote":"Supplies GOT-10k, the tracking benchmark whose success-rate numbers are compared against DAVE's for ARTrack."},{"cited_title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition","cited_arxiv_id":null,"evidence_quote":"Supplies ARTrack, the tracker whose performance on DAVE quantifies the dataset's tracking difficulty."},{"cited_title":"In: European conference on computer vision","cited_arxiv_id":null,"evidence_quote":"Supplies COCO, the detection benchmark and pretraining source whose Swin-T mAP of 50.5 is compared with DAVE's 32.5."},{"cited_title":"In: IEEE International Conference on Computer Vision (ICCV)","cited_arxiv_id":null,"evidence_quote":"Supplies Swin-T, the detector backbone used to measure DAVE's detection challenge."},{"cited_title":"In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","cited_arxiv_id":null,"evidence_quote":"Supplies AVA v2.2, the spatiotemporal action localization benchmark whose 33.3% mAP for ACAR-Net is compared with DAVE's 6.3%."},{"cited_title":"In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","cited_arxiv_id":null,"evidence_quote":"Supplies ACAR-Net, the action localization method run on DAVE to produce the spatiotemporal action localization difficulty result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies CG-DETR, the moment retrieval model whose R1@0.5 of 5.1 on DAVE versus 58.4 on Charades-STA is the video moment retrieval challenge result."},{"cited_title":"European Conference on Computer Vision (2016)","cited_arxiv_id":null,"evidence_quote":"Supplies Charades-STA, the video moment retrieval comparison dataset used in the CG-DETR experiments."},{"cited_title":"In: European Conference on Computer Vision","cited_arxiv_id":null,"evidence_quote":"Supplies Charades, the multi-label action recognition comparison dataset against which SlowFast's 41.0 mAP on DAVE is measured."}],"review_version":1}