{"id":"65838101-7d88-42b6-b55d-60d5d08b00fc","arxiv_id":"2504.12696","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey that catalogs and compares collaborative perception datasets for autonomous driving across cooperation paradigms, sensors, scenarios, and tasks, with a living online repository.","lead":"This paper reviews and compares dozens of datasets used to train autonomous vehicles to share sensor information with each other and with roadside infrastructure. It organizes the field by cooperation type, sensor type, scenario, and task, and provides a continuously updated online index.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table II/Figure 2/Table V contain verifiable inconsistencies (V2X-Sim/V2XSet year, SCOPE venue); the catalog is the load-bearing part of the central 'reliable map' claim, so these errors must be audited before the review can be used as a reference.","rationale":"The reader's weakest assumption—that Table II/GitHub metadata is accurate and current—is exactly the load-bearing condition for the paper's central claim of being a reliable map. I agree with that assessment. My pass adds more concrete instances of the same failure mode (SCOPE venue, F-Cooper venue, V2X-Sim 2.0 naming, V2XSet-Noise citation), which strengthens the conditional verdict rather than changing it. I do not treat the 'first comprehensive review' wording as the primary attack: although it sits in tension with the paper's own refs [72] and [73], which are CP-dataset-focused surveys, the review's usefulness survives a softened novelty claim, whereas the catalog errors do not. The correct fix is an erratum plus versioned audit for Table II/Fig. 2/Table V and the GitHub repository; after that the survey can serve its intended reference role. Hence verdict remains CONDITIONAL (unchanged).","tokens_in":29920,"tokens_out":9128,"duration_ms":91371,"concrete_test":"Run an independent primary-source audit of every row in Table II and every row in Table V: for each dataset, fetch the official project page/repository and the cited paper's arXiv/DOI metadata, and compare year, venue, sensors, data source, size, task list, and access method. Flag all mismatches; pay specific attention to V2X-Sim and V2XSet release years, SCOPE/Table V venue, F-Cooper's venue, V2X-Sim 2.0 version naming, and the V2XSet-Noise citation. If more than, say, 5% of rows disagree with primary sources, or if the three specific contradictions above cannot be resolved by erratum, the review's 'reliable map' claim is not met and a revised/versioned table is required before endorsement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central value is being a reliable, systematic catalog of CP datasets, and that value depends entirely on the metadata in Table II, Figure 2, Table V, and the linked repository being accurate. That assumption is demonstrably unsafe as printed. Figure 2 places V2X-Sim and V2XSet in 2021, while Table II and the cited papers (refs [40], [33]) give 2022. Table V lists SCOPE as '2023 ICCV', but Table II and ref [93] say '2024 arXiv'; F-Cooper's venue appears as 'SEC' instead of ICDCS, and the row labeled 'V2VNetrobust' is not a distinct 2021 CoRL paper. In Sections VI-A and V-C the text repeatedly cites 'V2X-Sim 2.0' to ref [40], whose title is just 'V2X-Sim', and the Fig. 7 caption calls V2XSet 'V2VSet'. V2XSet-Noise is cited to ref [58], an AgentAlign method paper rather than a dataset paper. None of these is fatal individually, but they are visible symptoms that the catalog has not been verified against primary sources. The GitHub repository inherits the same risk, and the 'continuously updated' claim supplies no audit trail. Since the review's utility is precisely as a map, this is the load-bearing weak point. Separately, the 'first comprehensive review' claim is in tension with refs [72] and [73], which are CP-dataset-focused surveys, but that is a novelty matter and secondary to the accuracy issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of collaborative perception (CP) datasets for autonomous driving. It organizes datasets by collaboration paradigm (roadside, V2V, V2I, V2X, and I2I), then cross-cuts the material by data source (simulated versus real), scenario (intersections, urban streets, highways), sensor modality (camera, LiDAR, radar), and supported perception tasks (detection, segmentation, tracking, prediction), including associated metrics. The paper also provides qualitative quality comparisons, discusses challenges and future directions, and points to an online repository of CP datasets that the authors state is continuously updated. The central claim is that this is the first comprehensive review devoted specifically to CP datasets and that its multi-dimensional categorization provides a reliable map of the field.","tokens_in":30256,"tokens_out":6941,"duration_ms":70087,"significance":"If the catalog is accurate, the paper would be a useful reference resource: the paradigm-based taxonomy is sensible, the coverage of recent datasets (through early 2025) is broad, and the cross-cutting organization by scenario, modality, and task is convenient for researchers selecting datasets or benchmarks. The qualitative radar charts and the discussion of availability, standardization, and Sim2Real gaps are also potentially helpful. However, the manuscript is a literature review with no new quantitative results, so its value rests almost entirely on the completeness and correctness of its dataset metadata and its comparison tables. The paper currently contains multiple verifiable internal inconsistencies and citation mismatches in exactly those load-bearing parts, so the claimed reliability of the map is not yet established.","major_comments":[{"comment":"The catalog metadata is not internally consistent and several entries contradict the cited primary sources. For example, Figure 2 places V2X-Sim and V2XSet in 2021, while Table II and refs [40] and [33] both give 2022; Table V lists SCOPE as \"2023 ICCV\", while Table II and ref [93] state \"2024 arXiv\", and SCOPE is a dataset rather than a method; the F-Cooper row gives venue \"SEC\", although ref [45] is an ICDCS paper; and the row labeled \"V2VNetrobust\" does not correspond to a distinct 2021 CoRL paper. Because the stated contribution is a reliable map of the field, these discrepancies are load-bearing and require a full audit of Table II, Figure 2, Table V, and the linked repository against the original dataset papers.","section":"Table II, Figure 2, Table V"},{"comment":"The reference column in Table V does not support the reported performance numbers. For instance, the F-Cooper and V2VNet rows on DAIR-V2X cite refs [96] and [98], which are roadside-3D-detection papers (SGV3D and MonoGAE), not the papers reporting the cooperative-benchmark results; several other rows cite refs [97], [99], [103], etc., which are likewise unrelated method papers. The table therefore cannot be used to compare methods as claimed. Each entry must be re-sourced to the original benchmark paper or to a verifiable public leaderboard, and the numbers must be checked against those sources.","section":"Table V"},{"comment":"The text repeatedly attributes material to the wrong references or uses incorrect dataset names. Section II-D and Section VII-B cite \"V2X-Sim 2.0\" to ref [40], whose title is \"V2X-Sim\"; Figure 7 and Section II-D caption the V2XSet samples as \"V2VSet\"; Table II lists V2XSet-Noise as a dataset citing ref [58], which is the AgentAlign method paper; and Section IV-C attributes \"TUMTraf A9\" to ref [78], which is the TUMTraf Intersection Dataset. These are not stylistic issues; each is a concrete factual mismatch that must be fixed and cross-checked against primary sources.","section":"Sections II-D, V-C, VI-A, VII-B; Figure 7"},{"comment":"The \"first comprehensive review\" claim is not supportable as stated, given that refs [72] and [73] are both 2024 IEEE IV conference surveys devoted specifically to collective/collaborative perception datasets. The authors distinguish their work in Section I-B by arguing those surveys are less complete, but the categorical claim of being \"the first dedicated review\" should be softened to a claim of greater breadth, currency, or depth, or the relation to [72] and [73] must be documented in detail.","section":"Abstract and Section I-C"},{"comment":"The review does not state its dataset inclusion criteria, the literature search cutoff date, or any procedure for verifying dataset metadata. Since the contribution is a catalog and the paper points to a \"continuously updated\" repository, the absence of a documented methodology makes completeness unassessable and provides no audit trail for the repository. A short methodology subsection (databases searched, selection criteria, extraction and verification procedure, update policy for the GitHub repository) should be added.","section":"Section I (overall methodology)"}],"minor_comments":[{"comment":"The V2V row labeled \"CP-UA V\" should be written consistently (e.g., CoPerception-UAV) and checked against its source; the abbreviation as printed is confusing.","section":"Table II"},{"comment":"The GitHub URL is broken across a line break in the manuscript text (\"Collaborative-Perceptio n-Datasets\"); ensure the final PDF typesets the hyperlink correctly.","section":"Abstract"},{"comment":"Capitalization is inconsistent for RCooper (\"Rcooper\" in Figure 12 and Table II); please standardize dataset names throughout.","section":"Figure 12 and Table II"},{"comment":"The list of LiDAR-camera datasets includes IRV2V [100] and OPV2V-N [99], which appear to be method-specific variants rather than standalone datasets; clarify their status in the catalog.","section":"Section V-C"}],"recommendation":"major_revision","confidential_remarks":"The topic fits the journal and the survey concept is sound, but the accuracy audit is the primary obstacle: the catalog, comparison tables, and reference attributions must be corrected before the paper can serve as the reliable reference resource it claims to be. I found no evidence of author misconduct; the issues are consistent with insufficient verification of assembled metadata."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a real contribution: the most up-to-date structured catalog of collaborative perception datasets I've seen, with a sensible taxonomy (roadside/V2V/V2I/V2X/I2I), sensor and task breakdowns, and an accompanying GitHub repo. Second, the catalog has enough factual errors that you should not yet trust it as a reference without checking primary sources.\n\nWhat's new: prior surveys (refs 72 and 73, both IV 2024) overlap heavily, but this one is broader and more recent, and treating I2I as its own category is a genuinely useful addition. The organization is coherent, and the comparative tables (sensor modalities, tasks, sim-vs-real) are helpful for dataset selection. The qualitative radar charts work fine as intuition.\n\nWhere it falls down: the table of datasets is the load-bearing part, and it has verifiable inconsistencies. Figure 2 dates V2X-Sim and V2XSet to 2021; Table II and the original papers say 2022. Table V lists SCOPE as 2023 ICCV when it is a 2024 arXiv paper, F-Cooper's venue appears as 'SEC' instead of ICDCS, and 'V2VNetrobust' is listed as a distinct 2021 CoRL paper when it is not. The Fig. 7 caption calls V2XSet 'V2VSet', and V2XSet-Noise is cited to an AgentAlign method paper rather than a dataset paper. Individually these are minor; collectively they show the metadata was not checked against primary sources, and the linked GitHub repo inherits the same risk. The 'continuously updated' claim also has no audit trail.\n\nThe 'first comprehensive review' claim is also overreach—refs 72 and 73 are precisely CP-dataset-focused surveys—but that is secondary to the accuracy problem.\n\nBottom line: the paper is a useful map with several wrong street names. It deserves a serious referee and a major revision with a full fact-check of Table II, Figure 2, Table V, and the repo against the original dataset papers. Once corrected, I would be glad to cite it as the standard resource. As printed, I would not rely on it.\n\nWho it is for: researchers choosing CP datasets or looking for a quick landscape overview. Take the numbers with a grain of salt until the revision.","headline":"A genuinely useful catalog of collaborative perception datasets, but the metadata is not yet reliable enough to serve as the reference map it claims to be.","tokens_in":30779,"tokens_out":2015,"would_cite":false,"duration_ms":20362,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims to be the first review devoted entirely to collaborative perception datasets, organizing roughly fifty resources by cooperation paradigm, data source, sensor modality, scenario, and task, and pairing the survey with a…","keywords":["collaborative perception","autonomous driving","V2X","V2V","V2I","I2I","dataset survey","sim-to-real gap"],"falsifier":"Compare each row of Table II against the cited original dataset paper: if a meaningful fraction of entries—for example the publication year of V2X-Sim or V2XSet, or the venue labels in Table V—disagrees with the primary sources, the review's authority as a dataset map fails. A reader could also re-check the online repository against newly released datasets six months after publication to test its currency.","tokens_in":29738,"feed_emoji":"🚗","tokens_out":6737,"duration_ms":59036,"temperature":0.7,"pith_summary":"This review tries to give collaborative perception for autonomous driving a reliable map of its data resources. It claims to be the first survey dedicated exclusively to collaborative perception datasets, rather than to methods or fusion techniques, and it organizes the datasets by cooperation paradigm (roadside, V2V, V2I, V2X, I2I), by simulated versus real origin, by sensor modality, by scenario, and by supported perception task. If the map is right, researchers can pick a dataset for a given task and immediately see what is missing—real-world scale, adverse weather, geographic diversity, communication constraints—which in turn exposes where benchmarks need standardization. The review also points to a continuously updated online list, a useful feature in a field where new datasets appear every year.","feed_headline":"Fifty-plus collaborative-driving datasets, sorted into one map","feed_subtitle":"The first survey devoted solely to collaborative perception datasets, ordered by paradigm, sensor, scenario, and task.","key_machinery":"The central object is a multi-axis taxonomy of collaborative perception datasets: rows are datasets, and the columns are cooperation paradigm (roadside, V2V, V2I, V2X, I2I), source (simulated or real), sensor modality (camera, LiDAR, radar), scenario (intersection, urban street, highway), and supported tasks (detection, segmentation, tracking, prediction). This taxonomy does the work of turning a scattered collection of heterogeneous resources into a single comparative grid, with qualitative chart profiles and quantitative tables, so that gaps and appropriate dataset choices become visible at a glance.","core_discovery":"The central claim is that a systematic, multi-dimensional comparison of collaborative perception datasets is both possible and needed, and that this review supplies it. The review catalogs roughly fifty public datasets, sorts them into five cooperation paradigms, and argues that the field has moved from static roadside perception through simulated V2V and V2I benchmarks to real-world, multimodal, task-specialized resources, including infrastructure-to-infrastructure and large-language-model-driven datasets. It further argues that the binding constraints on progress are dataset scale and diversity, accessibility, inconsistent annotation and evaluation standards, the simulation-to-reality gap, privacy, and the still-early use of large language models for annotation, generation, and reasoning.","pith_inferences":["If the online catalogue remains current, it could become the field's entry point for choosing benchmarks, but its long-term value depends on maintenance; a machine-readable, versioned metadata registry would be a natural and testable extension of the paper's contribution.","The taxonomy implies a testable prediction: datasets that combine real-world collection with explicit communication and latency modeling, such as OTVIC or V2X-Radar, will be the ones adopted for deployment-oriented benchmarking, because neither pure simulation nor clean real data captures the deployment bottleneck.","The inconsistencies visible within the paper itself, such as differing publication years for the same dataset in different figures and tables, suggest that a curated metadata file would serve the community more reliably than static tables."],"forward_implications":["A researcher can use the taxonomy to select a dataset by paradigm and task; for 3D detection, for instance, the review points to DAIR-V2X, OPV2V, V2X-Sim, V2V4Real, and RCooper as complementary real and simulated choices.","The documented simulation-to-reality gap implies that models trained on simulated benchmarks such as OPV2V or V2XSet should be expected to lose substantial average precision on real-world data unless domain adaptation is applied.","The review's challenge analysis implies that future datasets should prioritize adverse weather, geographic diversity, heterogeneous agents, realistic communication constraints, and standardized annotation formats.","The inclusion of LLM-based resources such as V2V-QA suggests that collaborative perception benchmarks will expand from detection and tracking toward reasoning and planning-style question answering."],"supporting_citations":[{"why":"An open benchmark V2V dataset used throughout the review as the reference for simulated cooperative 3D detection.","marker":"[39]"},{"why":"A large-scale real-world V2I dataset that anchors the review's V2I section and many benchmark comparisons.","marker":"[35]"},{"why":"A simulated multi-agent dataset used for cooperative detection, segmentation, and prediction benchmarks; also an example of the paper's metadata inconsistencies.","marker":"[40]"},{"why":"Introduces the V2XSet simulation benchmark and a widely used cooperative detection method.","marker":"[33]"},{"why":"Real-world V2V dataset used to quantify the simulation-to-reality gap and benchmark robust fusion.","marker":"[52]"},{"why":"Claims to be the first large-scale real-world roadside cooperative dataset and anchors the I2I section.","marker":"[60]"},{"why":"A prior review with partially overlapping dataset coverage, cited to show what the present review adds in breadth and depth.","marker":"[72]"},{"why":"A prior collaborative-perception-dataset survey that the paper distinguishes from its own task-centric, multi-dimensional review.","marker":"[73]"}],"fun_headline_variants":["First survey maps 50+ collaborative perception datasets by paradigm","First collaborative-perception dataset survey: 50+ datasets, 5 paradigms","Five paradigms, 50+ datasets: the first survey of collaborative perception","First collaborative perception dataset review: 50+ resources, 5 cooperation paradigms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the metadata compiled in the tables and the online list accurately reflect the cited dataset papers and are current; if entries are stale or wrong, the review's map misleads the researchers it aims to serve.","fun_headline_variants_meta":{"raw":{"variants":["First survey maps 50+ collaborative perception datasets by paradigm","First collaborative-perception dataset survey: 50+ datasets, 5 paradigms","Five paradigms, 50+ datasets: the first survey of collaborative perception","First collaborative perception dataset review: 50+ resources, 5 cooperation paradigms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001709,"raw_usage":{"total_tokens":6728,"prompt_tokens":870,"completion_tokens":5858,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":5780}},"tokens_in":486,"tokens_out":5858,"duration_ms":40468,"temperature":1.0,"reasoning_tokens":5780,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:23:55.085479+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare each row of Table II against the cited original dataset paper: if a meaningful fraction of entries—for example the publication year of V2X-Sim or V2XSet, or the venue labels in Table V—disagrees with the primary sources, the review's authority as a dataset map fails. A reader could also re-check the online repository against newly released datasets six months after publication to test its currency.","supporting_citations":[{"cited_title":"Collective perception datasets for autonomous driving: A compre- hensive review,","cited_arxiv_id":null,"evidence_quote":"A prior review with partially overlapping dataset coverage, cited to show what the present review adds in breadth and depth."},{"cited_title":"Collaborative perception datasets in autonomous driving: A survey,","cited_arxiv_id":null,"evidence_quote":"A prior collaborative-perception-dataset survey that the paper distinguishes from its own task-centric, multi-dimensional review."}],"review_version":1}