{"id":"4d89831d-5809-40cd-a973-38b06bf3fbab","arxiv_id":"2505.00747","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that frames V2X communication as an information sensor and categorizes cooperative perception research into representation, fusion, and scalability.","lead":"This survey reviews cooperative perception for autonomous vehicles, treating V2X wireless communication as an information sensor alongside cameras and LiDAR. It organizes the field into three challenges: information representation, fusion, and large-scale deployment, and compares recent methods in each area.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Representation taxonomy omits BEV/map-level methods, undercutting the survey's central claim of a comprehensive organizing perspective.","rationale":"The reader's weakest assumption—that the three-level taxonomy is exhaustive and standard—is close to the concern I identify, but I refine it with internal evidence: the survey cites [15] as an occupancy-grid task that object-level representation cannot support, yet gives it no place in the taxonomy. This is not a disagreement with field consensus; it is a completeness gap relative to the paper's own references. A survey's utility depends on its organizing categories covering the relevant state of the art, and here a prominent representation family (dense BEV/occupancy maps) is left unclassified. The paper has real strengths: it is readable, cites relevant work, and its ideal/non-ideal fusion split is useful. The concern is fixable by revision—either adding a map-level category or explicitly defending why such methods are included under feature-level—so I recommend CONDITIONAL rather than REJECT. I partially agree with the reader because they flagged taxonomy exhaustiveness generally, while I identify a specific missing category and a concrete test to verify it.","tokens_in":11384,"tokens_out":7700,"duration_ms":88492,"concrete_test":"Run a systematic literature search for 2019-2025 combining 'cooperative perception' or 'collaborative perception' with 'BEV fusion', 'occupancy grid', 'semantic occupancy', or 'map-level fusion', on OPV2V, V2X-Sim, DAIR-V2X, and Rope3D. For each top-cited method, attempt to file its transmitted representation under data-level, feature-level, or object-level exactly as Section II-A defines them. If any method requires redefinition or a new category (e.g., shared occupancy grids or BEV semantic maps), the taxonomy is incomplete; the revision should either add a map-level/occupancy category or explicitly justify subsuming it under feature-level. This check settles whether the three-level schema is exhaustive or only a convenient subset.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The survey's central value claim rests on the three-level representation taxonomy in Section II-A (data-level, feature-level, object-level) being both standard and exhaustive. That assumption is not established, and the paper itself supplies evidence against it: Section II-A cites collaborative semantic occupancy prediction [15] and end-to-end driving on cooperative features [16], [17] as tasks that object-level information cannot support, yet it never assigns these methods a place in the taxonomy. Dense, semantically structured BEV or occupancy-grid representations are neither raw sensor data, in the sense of [7], [8], nor model-specific intermediate features in the sense of [10], [11], nor object lists. Such map-level representations are increasingly common in V2X cooperative perception. As written, a reader cannot tell whether Section II-A is meant to exclude them, fold them silently into 'feature-level', or treat them as a fourth category. Because comprehensiveness is the survey's main promise, this ambiguity is load-bearing: the organizing schema does not demonstrably cover a major and growing body of work. The numerical slip in Section II-B ('10 LiDAR points per second') is a separate editing issue; the taxonomic gap is the substantive concern for the paper's central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews cooperative perception for autonomous driving from an information-centric viewpoint, treating V2X communication as a dynamic 'information sensor' with four characteristics: mobility, heterogeneity, communication dependence, and scalability. It organizes recent work along three dimensions—information representation (data-level, feature-level, object-level), information fusion under ideal and non-ideal conditions (heterogeneity, latency, packet loss, pose errors), and large-scale deployment (system architectures and information scheduling). The paper identifies open challenges such as task-specific information selection, reliance on joint training, lack of standardized benchmarks, and suggests future directions including explicit representations and universal feature spaces. The survey contributes no new technical results; its value rests on the usefulness and completeness of its organizing taxonomy and coverage.","tokens_in":11572,"tokens_out":5097,"duration_ms":52492,"significance":"If its organizing taxonomy is accepted, this survey offers a useful complement to fusion-centric surveys by foregrounding representation choices and deployment scalability. It assembles a broad set of recent methods, including several from 2023-2025 venues, and gives balanced treatment to compression, heterogeneity, latency, packet loss, and pose calibration. The 'information sensor' framing is a plausible pedagogical contribution. However, the survey's significance is conditional on the three-level representation taxonomy being both complete and clearly defined; as discussed below, that condition is not currently met, and one quantitative motivation contains an apparent error.","major_comments":[{"comment":"The three-level taxonomy (data-, feature-, object-level) is presented as exhaustive ('cooperative perception can be categorized into three approaches'), but the paper itself cites collaboration methods built on Bird's-Eye-View (BEV) or occupancy representations: collaborative semantic occupancy prediction [15] and end-to-end cooperative driving [16], [17] are invoked as tasks that object-level information cannot support, yet no fourth category is defined to cover them. Dense map-level or occupancy-grid representations are not raw sensor data in the sense of [7], [8], nor model-specific intermediate features in the sense of [10], [11], nor object lists. Please add an explicit fourth representation category (e.g., map/BEV/occupancy-level) or justify subsuming these works under feature-level; as written, the taxonomy does not demonstrably cover a major body of cooperative perception work, which undercuts the survey's claim of a comprehensive organizing perspective.","section":"II-A"},{"comment":"The claim that 'less than 10 Mbps' translates to 'about 4.16 million pixels, 10 LiDAR points, or 4,800 64-channel depth features per second' is not derived and is numerically implausible: at 8 bits per pixel, 4.16 million pixels would require about 33 Mbps, and '10 LiDAR points' is several orders of magnitude too low for any reasonable point encoding. Moreover, the sentence attributes the 10 Mbps C-V2X bound to reference [16], which is the Coopernaut paper on end-to-end driving, not a V2X throughput measurement study. Please correct the derivation, replace the numbers, and cite an appropriate source for the throughput claim.","section":"II-B"}],"minor_comments":[{"comment":"The acronym C-V2X is used without defining 'Cellular Vehicle-to-Everything' at first use; please expand it for readers who are not specialists in vehicular communications.","section":"Section II-B"},{"comment":"References [8] and [34] are duplicate entries for the same Cooper paper, and references [7] and [48] are duplicate entries for the same multivehicle cooperative driving paper; please merge these duplicate entries.","section":"References"},{"comment":"The phrase 'deep generative models such as autoencoders and their variations' is imprecise in relation to V2VNet, which uses a CNN-based compression module; suggest distinguishing learned compression from generative-model-based compression to avoid conflating the two.","section":"II-C"},{"comment":"A summary table comparing the surveyed large-scale systems (e.g., EMP, AutoCast, Harbor) in terms of agent count, architecture type, communication assumptions, and reported performance would make the comparison easier to follow and would strengthen the survey's usefulness.","section":"IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid survey but needs to resolve the taxonomy gap in Section II-A and correct the quantitative and citation error in Section II-B before it can serve as a reliable reference. The duplicate reference entries and other presentation issues are straightforward to fix. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. This is a useful survey, not a breakthrough. The information-sensor framing is a sensible way to organize known V2X constraints—mobility, heterogeneity, communication dependence, scalability—but it doesn't add new theory. What the survey does well is collection: it pulls together recent work through early 2025, including the scalability/system-level line (EMP, AutoCast, Harbor, Select2Coll), which earlier surveys treat thinly. The references are real, and the self-citations are used as examples, not to prop up the framing. If I were new to cooperative perception, this would give me a quicker map than the 2023 surveys. The biggest issue is the representation taxonomy in Section II-A. The paper defines three levels—data, feature, object—but then cites occupancy prediction [15] and end-to-end cooperative driving [16,17] as tasks object-level can't support, without saying where those methods live. Dense BEV or occupancy-grid representations are common now, and the survey itself uses BEV features in Section III-B (HM-ViT). That leaves the reader unable to tell whether BEV/map-level is supposed to fold into feature-level, form a fourth category, or be excluded. Since comprehensiveness is the survey's main pitch, this is load-bearing, not cosmetic. The fix is easy in a revision: state the relationship explicitly. The other concrete problem is numbers. Section II-B says 10 Mbps supports about 10 LiDAR points per second. That is off by orders of magnitude for any reasonable point encoding, and no derivation or source is given for 4.16 million pixels either. A reader who spots this will start doubting the qualitative claims even where they are sound. It's a small editing fix, but it should not survive review. Novelty is modest. The three-level taxonomy is not new; prior surveys use it. The information-sensor lens is a reframing. That's fine for a survey, but the prose should be careful not to oversell it. This paper is for graduate students and researchers entering the area, or for people in the V2X comms community who want the perception side summarized. Specialists won't learn much they don't already know. I probably won't cite it in my own work, but I'd point students to it. It deserves peer review rather than desk rejection: it's coherent, current, and citable once the taxonomy and numbers are fixed. Send it to a venue with constructive referees; ask for revision, not rejection.","headline":"A current, useful survey with a real organizing gap around BEV/occupancy representations and a bad bandwidth number; worth peer review after revision.","tokens_in":12108,"tokens_out":3390,"would_cite":false,"duration_ms":37480,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that Vehicle-to-Everything (V2X) wireless communication is best understood as an \"information sensor\" for autonomous vehicles, defined by mobility, heterogeneity, communication dependence, and scalability, and it…","keywords":["cooperative perception","V2X communication","information sensor","multi-agent perception","information fusion","autonomous driving","communication-efficient perception","large-scale deployment"],"falsifier":"A concrete check on the central claim: run a representative cooperative perception stack with and without communication-aware representation and compression under a measured V2X link of less than 10 Mbps; if the communication-agnostic version matches its accuracy and latency, the case for treating the wireless link as a first-class perception sensor weakens. A meta-analytic falsifier would be a sizable cluster of published cooperative perception methods whose core contribution fits none of the survey's three axes.","tokens_in":11174,"feed_emoji":"📡","tokens_out":5807,"duration_ms":60191,"temperature":0.7,"pith_summary":"This paper argues that Vehicle-to-Everything (V2X) wireless communication should be treated as an \"information sensor\" for autonomous vehicles, not merely as a data link. On this view, the wireless channel behaves like a perception sensor whose readings are other agents' observations, with four defining traits: mobility, heterogeneity, communication dependence, and scalability. The survey organizes the field around three questions—how shared information should be represented, how it should be fused, and how the system should scale—and shows that each question is governed by limited bandwidth. A sympathetic reader would take away that representation choice and communication-aware fusion are core design decisions, on par with sensor choice, rather than implementation details.","feed_headline":"Treat vehicle-to-everything radio as a sensor for self-driving cars","feed_subtitle":"A survey reorders cooperative perception around representation, fusion, and scale, with bandwidth as the load-bearing constraint.","key_machinery":"The load-bearing object of the survey is the \"information sensor\" concept: a way of treating the V2X wireless link as a virtual perception sensor whose inputs are measurements made by other agents and whose output is constrained by bandwidth, signal stability, and mobility. This concept carries the argument by converting communication constraints from an engineering nuisance into a first-class property of the perception system, on equal footing with a camera's field of view or a LiDAR's range. The second piece of machinery is the three-level taxonomy of information representation—data-level, feature-level, object-level—which the survey uses to index both compression methods and fusion strategies, and which lets it identify the open problem of a universal, task-agnostic intermediate representation.","core_discovery":"On the paper's own terms, the central discovery is organizational: cooperative perception research has matured enough to be viewed through an information-centric lens, and doing so reveals a consistent structure. Raw sensor data can be shared at the data level, the feature level, or the object level; data level preserves detail but swamps the network, feature level compresses but suffers standardization and heterogeneity problems, and object level is bandwidth-friendly but loses information needed for prediction and end-to-end driving. Fusion methods that work under ideal, homogeneous conditions degrade under real-world latency, packet loss, and localization error, in some cases falling below single-vehicle perception. The paper further claims that large-scale deployment requires explicit system-level choices—edge-assisted, fully decentralized, or hybrid architectures, plus communication scheduling—because the number of cooperating agents varies from a few to hundreds. The conclusion is that the field's next step is not better detectors alone but generalizable, communication-aware representations and standardized, realistic benchmarks.","pith_inferences":["The information-sensor lens likely extends beyond vehicles to any multi-robot or edge-AI system where perception data cross a wireless bottleneck, such as drones or warehouse robots; the same representation-fusion-deployment triad would apply.","The paper's suggested directions—3D Gaussian ellipsoids as explicit representations and a universal feature space—point toward a task-agnostic compressed world model; a testable extension is whether such a shared representation lets agents collaborate on the fly without any joint training.","The observed sharp accuracy collapse under compression resembles a rate-distortion-perception tradeoff; quantifying that tradeoff with an information-theoretic bound could give codec designers a target to optimize.","A concrete missing piece implied by the survey is a common test harness that injects measured packet loss, latency, and localization noise into standard datasets; building one would test whether robustness methods actually generalize."],"forward_implications":["Bandwidth becomes a perception budget: choosing between data-, feature-, and object-level sharing is a perception design decision with direct accuracy and latency consequences.","Fusion algorithms must assume imperfect communication and pose error, because under latency, packet loss, or misalignment cooperation can perform worse than a single vehicle.","Compression is not free: once compression exceeds a threshold, cooperative perception accuracy drops sharply, so codecs for this setting need to preserve semantic content, not just geometry.","Scalability requires system-level planning of who talks to whom and when, using architectures that range from edge servers to fully decentralized schemes.","Progress in real-world deployment depends on standardized benchmarks and realistic large-scale datasets that include heterogeneity, localization noise, and genuine communication limits."],"supporting_citations":[{"why":"Supplies the early GNN-based cross-agent feature fusion method that later feature-level approaches build on.","marker":"[1]"},{"why":"Documents that high compression drives cooperative perception accuracy below single-vehicle perception, underpinning the communication-dependence claim.","marker":"[2]"},{"why":"A prior survey that focused mainly on information fusion, providing the contrast for this paper's new perspective.","marker":"[5]"},{"why":"A prior survey categorizing intermediate fusion methods by real-world challenges, used as a baseline for the proposed organization.","marker":"[6]"},{"why":"Introduces a widely used V2V benchmark dataset and an autoencoder-based feature compression pipeline.","marker":"[11]"},{"why":"Provides a feature-level information selection method using spatial confidence maps, a key exemplar of reducing information volume.","marker":"[21]"},{"why":"Demonstrates a fully decentralized scalable architecture tested in dense scenarios of up to 40 vehicles.","marker":"[27]"},{"why":"Offers attention-based fusion robust to spatial misalignment, illustrating fusion under non-ideal conditions.","marker":"[36]"},{"why":"Presents an edge-assisted architecture with real-time operation for up to six vehicles, grounding the scalability discussion.","marker":"[56]"}],"fun_headline_variants":["Why V2X radio is the next sensor for self-driving cars","Bandwidth is the real bottleneck in cooperative perception","Three levels of data sharing for self-driving cooperation","Communication-aware perception: the next step for AVs","Treating V2X as a sensor to scale cooperative perception"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness rests on accepting that cooperative perception research can be cleanly divided into the three dimensions of representation, fusion, and scalability, and that the papers it reviews are representative of the field.","fun_headline_variants_meta":{"raw":{"variants":["Why V2X radio is the next sensor for self-driving cars","Bandwidth is the real bottleneck in cooperative perception","Three levels of data sharing for self-driving cooperation","Communication-aware perception: the next step for AVs","Treating V2X as a sensor to scale cooperative perception"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3072,"prompt_tokens":915,"completion_tokens":2157,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":2078}},"tokens_in":531,"tokens_out":2157,"duration_ms":17442,"temperature":1.0,"reasoning_tokens":2078,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:58:53.354956+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check on the central claim: run a representative cooperative perception stack with and without communication-aware representation and compression under a measured V2X link of less than 10 Mbps; if the communication-agnostic version matches its accuracy and latency, the case for treating the wireless link as a first-class perception sensor weakens. A meta-analytic falsifier would be a sizable cluster of published cooperative perception methods whose core contribution fits none of the survey's three axes.","supporting_citations":[{"cited_title":"V2VNet: Vehicle-to-vehicle communication for joint perception and prediction,","cited_arxiv_id":null,"evidence_quote":"Supplies the early GNN-based cross-agent feature fusion method that later feature-level approaches build on."},{"cited_title":"A survey on intermediate fusion methods for collaborative perception categorized by real world challenges,","cited_arxiv_id":null,"evidence_quote":"A prior survey categorizing intermediate fusion methods by real-world challenges, used as a baseline for the proposed organization."},{"cited_title":"OPV2V: An open benchmark dataset and fusion pipeline for perception with Vehicle-to- Vehicle communication,","cited_arxiv_id":null,"evidence_quote":"Introduces a widely used V2V benchmark dataset and an autoencoder-based feature compression pipeline."},{"cited_title":"Where2comm: Communication-efficient collaborative perception via spatial confidence maps,","cited_arxiv_id":null,"evidence_quote":"Provides a feature-level information selection method using spatial confidence maps, a key exemplar of reducing information volume."},{"cited_title":"Autocast: scalable infrastructure-less cooperative perception for distributed collaborative driving,","cited_arxiv_id":null,"evidence_quote":"Demonstrates a fully decentralized scalable architecture tested in dense scenarios of up to 40 vehicles."},{"cited_title":"V2X-ViT: Vehicle-to-everything cooperative perception with vision transformer,","cited_arxiv_id":null,"evidence_quote":"Offers attention-based fusion robust to spatial misalignment, illustrating fusion under non-ideal conditions."},{"cited_title":"Emp: Edge-assisted multi-vehicle perception,","cited_arxiv_id":null,"evidence_quote":"Presents an edge-assisted architecture with real-time operation for up to six vehicles, grounding the scalability discussion."}],"review_version":1}