{"id":"ec6cfa97-d3b2-4bf8-ba7f-934093b6432b","arxiv_id":"2507.01059","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A vision paper recommending natural language as the universal communication medium for connected and automated vehicles.","lead":"This paper argues that self-driving cars should talk to each other in natural human language rather than raw sensor data or numerical detection outputs. It makes the case that language handles bandwidth limits, bridges different vehicle brands, and keeps human drivers in the loop, though it offers no experimental proof.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper asserts model-agnostic interoperability without specifying a shared semantics for natural-language V2X messages; this is more fundamental than the acknowledged latency gap and remains unaddressed.","rationale":"The reader's conditional verdict is appropriate. The paper is a readable position piece with a coherent limitation section, but its central normative claim rests on natural language delivering universal interoperability. I read the strongest_claim as 'natural language should be the foundational V2X protocol.' For that to hold, natural-language messages must have stable semantics across heterogeneous agents. The paper asserts this rather than supporting it. My concern is not that natural language is 'ambiguous' in the everyday sense; it is that the paper gives no mechanism by which different LVLMs or human drivers converge on the same decision-relevant interpretation, and no evaluation. This is not refuted by its own rebuttal in §5.1, which cites consistency within a single model or prompting regime. The reader's latency concern is genuine and explicitly acknowledged in §5.2, but I see the semantic-interoperability gap as more load-bearing because it threatens the core advantage even under ideal computational conditions. The proposed test—cross-model agreement on the paper's own example messages—would settle it. Given the paper's genre (vision/position) and its explicit hybrid-approach caveat in §5.4, the verdict should remain conditional rather than reject or accept.","tokens_in":14262,"tokens_out":6533,"duration_ms":78581,"concrete_test":"Select the example messages in Figures 1/3 and §4.5 (for example, 'Entering now; intersection clear in three seconds' and 'I'm yielding because you arrived first'). Using two publicly available driving LVLMs with different training pipelines (for instance, DriveLM and V2V-LLM, or two differently fine-tuned checkpoints of the same backbone), generate a forced-choice action decision (accelerate, brake, yield, hold, other) and any referenced spatial/temporal predicate for each message in 200 simulated V2X scenarios drawn from the paper's negotiation cases. Compute agreement rate against each other and against human-expert labels. If cross-model agreement on safety-critical actions is below, say, 95%, the 'model-agnostic interoperability' claim fails; if it is high, the semantic-standardization gap reduces to a specification task.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central argument (Introduction, §4.3) is that natural language is a universal protocol because 'language carries semantics in the words themselves' and any machine with a shared ontology can interpret it. That assumption is the load-bearing element: if two agents do not map the same sentence to the same driving-relevant meaning, the claimed interoperability—the paper's principal advantage over raw data, features, and perception results—collapses. The paper never defines the 'shared ontology' it invokes, and it provides no evidence that independently developed LVLMs will agree on interpretations. Its own §3.2 warns that models trained on slightly different data exhibit silent divergences; LVLMs trained on different driving corpora will likewise ground phrases such as 'Entering now; intersection clear in three seconds' differently. §5.1 only says 'consistent terminology and contextual grounding' and 'appropriate training and prompting' can help, citing work on single models; it does not show cross-model consistency or provide a protocol-level grammar/ontology. A communication protocol requires an agreed encoding and decoding relation; 'natural language' as used here is a proposal to build such a protocol, not the protocol itself. The latency issue noted by the reader (§5.2) is real but secondary: if semantic agreement is not established, faster LVLMs do not make the protocol universal.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that natural language, processed by large vision-language models (LVLMs), should become the foundational communication protocol for multi-agent collaborative driving. It reviews existing communication media (raw sensor data, neural-network features, perception results), identifies challenges (bandwidth, heterogeneity, decision-level fusion, scenario variability, transparency), and then makes a case for natural language based on semantic richness, bandwidth efficiency, adaptability, interoperability, human compatibility, and support for intent communication. The paper acknowledges and responds to several counterarguments, including precision, latency, security, and hybrid approaches, and concludes with a call to prioritize research on natural-language V2X frameworks.","tokens_in":14470,"tokens_out":2596,"duration_ms":27049,"significance":"If the central thesis is correct, it would motivate a significant redesign of V2X architectures and a new research agenda for protocol-level natural-language semantics. The paper is clearly written and does a service by articulating a coherent alternative to data-oriented collaboration, gathering relevant recent literature, and explicitly engaging with counterarguments. It also makes honest concessions: Section 5.2 admits that LVLM latency is currently unproven, and Section 5.4 concedes that precise numeric data remains useful. However, the paper is argumentative rather than empirical, and its key load-bearing claim of universal interoperability rests on an unexamined assumption about shared semantics across independently trained LVLMs. The absence of quantitative support for the bandwidth comparisons and the lack of a concrete semantic-interoperability mechanism are the main weaknesses.","major_comments":[{"comment":"The claim of model-agnostic interoperability is load-bearing but not established. Section 4.3 asserts that 'natural language carries semantics in the words themselves' and that any agent with a shared ontology can interpret it, yet no shared ontology, message grammar, or grounding mechanism is defined anywhere in the paper. The paper's own Section 3.2 warns that models with identical architectures trained on slightly different datasets exhibit 'silent divergences'; this applies directly to LVLMs trained on different driving corpora, which could ground phrases such as 'Entering now; intersection clear in three seconds' differently. Section 5.1's response appeals to 'consistent terminology and contextual grounding' and cites works on single models, but does not demonstrate cross-model agreement or propose a protocol-level specification of the encoding/decoding relation. Without this, natural language is a proposal for a protocol rather than the universal protocol the conclusions claim.","section":"§4.3, §5.1"},{"comment":"The quantitative evidence for the bandwidth argument is unsubstantiated. Figure 2 presents a per-agent data budget comparison across communication protocols and data modalities, but the manuscript provides no data source, no calculation methodology, and no assumptions about message sizes or protocol overhead; the caption even says 'Language and bounding box messages remain efficient' without definition. Table 1 lists bandwidths and latencies for DSRC, LTE-V2X, and 5G-V2X, but only DSRC and 5G-V2X rows are backed by the cited references [56,57] in the text; the LTE-V2X figures lack a source. These figures are central to the paper's 'bandwidth efficiency' advantage, so they should either be properly sourced and computed or explicitly labeled as illustrative.","section":"Figure 2, Table 1"},{"comment":"The feasibility of real-time LVLM-based communication is acknowledged as an open problem but then effectively assumed away. Section 5.2 states that 'specialized models optimized for driving can run with less resource requirements' and that 'it is promising' that LVLMs will become efficient enough, citing [103], but provides no measured latency or throughput figures. The final sentence of the section says the use of natural language 'should be encouraged' rather than showing that it can meet the real-time constraints of safety-critical maneuvers. As a position paper, an explicit research agenda is acceptable, but the strength of the Conclusions—which describe natural language as 'the foundational communication protocol'—should be tempered to match the level of evidence, or the paper should include concrete numbers from compressed models.","section":"§5.2"}],"minor_comments":[{"comment":"There are multiple typos: 'their is a tendency' should be 'there is a tendency', 'communincation' should be 'communication', 'drive-specicalized' should be 'driving-specialized', and 'A Vs' should be 'AVs'.","section":"§5.2"},{"comment":"The sentence 'An natural language text like' should be 'A natural language text like', and 'converting such accidental alert into structural numerical data is either trivial nor necessary' should be 'is neither trivial nor necessary'.","section":"§5.1"},{"comment":"The figure axis label reads 'Per-agent data budget at 30 Hz (KB)', but the caption says 'Data budgets are shown in kilobytes'; clarify what message size is assumed for a 'language' message and how the curves are computed.","section":"Figure 2"},{"comment":"The boxed 'Heterogeneity Metaphor' and 'Perception is a Tool, Driving is the Goal' passages are stylistic and interrupt the technical narrative; consider moving them to a footnote or removing them.","section":"General"},{"comment":"Several references are unpublished arXiv preprints, including several by the authors themselves; while this is common in fast-moving areas, the paper would benefit from a note that key claims rely on these preprints.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-positioned as an opinion/outlook piece, but the central interoperability claim needs a concrete protocol-level definition of shared semantics before it can support the strong conclusion. The heavy reliance on the authors' own preprints (LangCoop, STAMP, DRAMA-X, V2X-LLM) as feasibility evidence is worth checking for novelty disclosure. Also, the unsubstantiated Figure 2 should be fixed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'll be direct: this is a position paper, not a research paper. The core idea — natural language as a foundational V2X communication protocol — is already present in the authors' own references (Talking Vehicles, LangCoop, V2V-LLM) and their earlier work. What this submission adds is a systematic argument: a structured critique of raw data, features, and perception results; a clear list of core challenges; a useful breakdown of SAE J2735 message types; and a genuinely honest limitations section that even concedes hybrid approaches. The writing is clear, and the reasoning is coherent. That is not nothing.\n\nThe soft spot is the load-bearing assumption in Section 4.3: natural language is model-agnostic because any machine with a shared ontology can interpret it. The paper never defines that shared ontology and provides no evidence that independently trained LVLMs will map the same sentence to the same driving-relevant meaning. In fact, Section 3.2 warns that models trained on slightly different datasets exhibit silent divergences; the same reasoning applies to LVLMs grounded on different driving corpora. Natural language is not a protocol. It is a promising medium that still requires an agreed encoding and decoding relation. This is more fundamental than the latency concern in Section 5.2, which the paper itself admits is speculative. The reader's stress-test note on missing semantics is on target.\n\nOther issues are minor for a position piece: Table 1 lacks sources, Figure 2's bandwidth numbers are unsubstantiated, and the efficiency claims for driving-specialized LVLMs are explicitly hopeful rather than measured. None of these undercut the paper's value as a vision document, but they do keep it from being a technical contribution.\n\nThe paper is for researchers working on LLM-based driving or V2X who want a compact statement of this research agenda, and for referees who need to see both sides of the language-vs-structured-data debate. It deserves serious peer review, not a desk reject, because the position is timely, the authors admit the open problems, and the field can use explicit agenda-setting. My recommendation: engage with it. If I were reviewing, I'd press the authors to either produce cross-model consistency evidence or substantially narrow the interoperability claim.","headline":"A clear, honest position paper arguing for natural language as the V2X layer; the case holds as a research agenda, but the interoperability claim rests on an unproven shared-semantics assumption.","tokens_in":15028,"tokens_out":2990,"would_cite":true,"duration_ms":29289,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that natural language should become the foundational communication protocol for multi-agent collaborative driving, replacing raw sensor data, neural network features, and perception results as the primary exchange medium.","keywords":["multi-agent collaborative driving","natural language communication","vehicle-to-everything (V2X)","intent and reasoning communication","large vision-language models","connected autonomous vehicles","decision-level fusion","interoperability"],"falsifier":"Measure the end-to-end latency and misinterpretation rate of a language-based V2X alert in a time-critical scenario, such as a pedestrian dart-out at an unsignalized intersection, using representative vision-language models on automotive-grade hardware. If message round-trip consistently exceeds the emergency-response budget on the order of 100 milliseconds, or produces a material fraction of wrong or ambiguous intents, the universal-language protocol fails at its core.","tokens_in":14045,"feed_emoji":"💬","tokens_out":5305,"duration_ms":55189,"temperature":0.7,"pith_summary":"This paper argues that natural language, rather than raw sensor data, neural network features, or perception outputs, should be the foundational communication protocol for collaborative autonomous driving. The reason is that language packs perception, intent, rationale, and negotiation into a few bytes, works across heterogeneous vehicles and infrastructure, and remains readable by human drivers. The paper walks through the bandwidth, interoperability, decision-level, scenario-variability, and transparency failures of current V2X media, then argues language fixes each. If the argument is right, V2X design should shift from perception-data exchange to intent-and-reasoning messages, with structured numeric data kept as a complement.","feed_headline":"Connected cars should talk to each other in plain language","feed_subtitle":"Intent-based, human-readable messages beat raw sensor streams on bandwidth, interoperability, and safety coordination.","key_machinery":"The load-bearing object is the natural-language V2X message as a universal exchange unit, for example: \"I am slowing down because there is a cyclist on the right shoulder who appears unsteady.\" One short sentence simultaneously carries perception, assessment, intent, and rationale. The argument rests on the claim that this message type outperforms raw sensor data, neural features, and bounding-box or occupancy outputs on bandwidth, interoperability, adaptability, human compatibility, and decision-level coordination, and that large vision-language models make generation and interpretation feasible.","core_discovery":"The central claim is a design thesis: multi-agent collaborative driving has hit limits set by its communication media, and the medium that removes those limits is human natural language, generated and interpreted by large vision-language models. Existing media trade off bandwidth, completeness, and interoperability: raw sensor data exceeds available budgets, learned features break across heterogeneous agents, and perception results lose context and skip decision-level coordination. Natural language carries perception, state assessment, intended action, and causal reasoning in a single short message, scales its detail to channel conditions without protocol renegotiation, and provides a common channel for vehicles, roadside units, drones, pedestrians, and humans. The authors explicitly position language as the primary, universal protocol with structured numeric data as a complement, not as the exclusive medium.","pith_inferences":["Beyond roads, the same logic likely extends to any setting where heterogeneous machines and people share space, such as warehouse robots, delivery drones, and pedestrian devices.","The paper leaves implicit that language could serve as the semantic layer that labels and explains structured data, so hybrid protocols may converge on language-plus-numbers even if pure language falls short.","A fair test of the latency premise would be whether distilled driving-specialized vision-language models can sustain a 10 Hz message exchange on automotive-grade hardware, since the paper concedes this is not yet demonstrated.","The paper's hybrid caveat implies an open design problem: specifying a shared driving vocabulary and fallback rules for when language messages are ambiguous or missing."],"forward_implications":["V2X systems should be redesigned around intent and reasoning messages, with structured numeric data used only where precision is genuinely required.","Heterogeneous agents with different sensor suites, models, and manufacturers could interoperate without feature alignment or shared neural architectures.","Decision-level collaboration, such as negotiation at intersections, merging, and emergency yielding, becomes direct and explicit rather than inferred from perception data.","Human drivers and pedestrians can participate in the same communication channel, improving mixed-traffic safety and transparency.","Messages can compress to short high-priority alerts during bandwidth limits or emergencies without protocol renegotiation."],"supporting_citations":[{"why":"Demonstrates a cooperative driving system that uses natural language for vehicle-to-vehicle intent communication, serving as the main positive evidence for the thesis.","marker":"[25]"},{"why":"Provides a language-based collaborative driving framework that the paper repeatedly cites to show language messages can convey perception and intent.","marker":"[26]"},{"why":"Shows that LLM-based negotiation improves cooperative autonomous driving, supporting the decision-level collaboration argument.","marker":"[24]"},{"why":"Documents the heterogeneous V2X message formats used in practice, motivating the need for a unified representation such as natural language.","marker":"[59]"},{"why":"Supplies the DSRC bandwidth and latency figures used to argue raw data transmission is impractical at scale.","marker":"[56]"},{"why":"Supplies the 5G-V2X bandwidth and latency figures used in the per-agent data budget comparison.","marker":"[57]"},{"why":"Supports the claim that vehicle-to-vehicle communication with multimodal language models is feasible and can be made resource-efficient.","marker":"[73]"},{"why":"Grounds large vision-language models in driving context, supporting the claim that language messages can be visually grounded and interpreted consistently.","marker":"[27]"}],"fun_headline_variants":["Cars should speak plain language to coordinate","Why self-driving cars should chat in natural language","Natural language is the missing link for connected cars","For safer fleets, try cars that talk like humans","Beyond sensor data: cars that explain their moves"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument depends on large vision-language models being able to generate and interpret safety-critical driving messages quickly and reliably enough for real-time use; the paper concedes the latency issue and only says specialized driving models are promising.","fun_headline_variants_meta":{"raw":{"variants":["Cars should speak plain language to coordinate","Why self-driving cars should chat in natural language","Natural language is the missing link for connected cars","For safer fleets, try cars that talk like humans","Beyond sensor data: cars that explain their moves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1352,"prompt_tokens":814,"completion_tokens":538,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":430,"tokens_out":538,"duration_ms":5536,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:45:06.409210+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the end-to-end latency and misinterpretation rate of a language-based V2X alert in a time-critical scenario, such as a pedestrian dart-out at an unsignalized intersection, using representative vision-language models on automotive-grade hardware. If message round-trip consistently exceeds the emergency-response budget on the order of 100 milliseconds, or produces a material fraction of wrong or ambiguous intents, the universal-language protocol fails at its core.","supporting_citations":[{"cited_title":"V2X Communications Message Set Dictionary","cited_arxiv_id":null,"evidence_quote":"Documents the heterogeneous V2X message formats used in practice, motivating the need for a unified representation such as natural language."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DSRC bandwidth and latency figures used to argue raw data transmission is impractical at scale."},{"cited_title":"C-v2x use cases, methodology, and service level requirements, 2020","cited_arxiv_id":null,"evidence_quote":"Supplies the 5G-V2X bandwidth and latency figures used in the per-agent data budget comparison."}],"review_version":1}