{"id":"cfbcdeec-2056-4d06-aaf8-108286b22c60","arxiv_id":"2606.29115","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Analysis of AV stopping incidents categorizes failures in perception, planning, and control, arguing for human-interactive autonomy over passive minimal risk conditions.","lead":"Autonomous vehicles stop when uncertain to avoid crashes, but this can block traffic, emergency vehicles, and create accessibility issues. The paper analyzes real-world incidents to argue that AV safety needs human interaction capabilities instead of relying only on stopping.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Paper attributes interaction failures primarily to AV perception/planning/control limits without evidence these dominate over regulatory/infrastructure factors","rationale":"Reader's weakest assumption directly identifies the unsupported causal priority. Because the manuscript is an incident taxonomy plus position statement with no quantitative disambiguation of causes, the same concern remains load-bearing even after full-text review.","tokens_in":1714,"tokens_out":287,"duration_ms":13199,"concrete_test":"For each publicly cited incident in the paper's taxonomy, extract the original report text and code whether an external factor (regulation, infrastructure, or human driver behavior) is mentioned as contributory; if >30% of incidents list such factors as primary or co-equal, recompute the gap analysis excluding purely internal attributions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that documented stopping incidents stem chiefly from internal AV architecture gaps (so that adding human-interactive perception, language-grounded planning, and teleoperation would enable safe urban deployment). The taxonomy in the analysis section maps incidents to those three categories but supplies no comparative count, counterfactual, or exclusion of external causes (e.g., traffic rules that treat stopped AVs as obstacles, lack of V2I infrastructure, or municipal response protocols). Without that, the inference that passive MRCs must be replaced by cooperative autonomy rather than complemented by policy or infrastructure changes remains unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that AV safety frameworks centered on minimal risk conditions (MRCs) such as stopping are insufficient for urban environments, as public incidents show these behaviors obstruct traffic, interfere with emergency response, and create accessibility issues. It presents an analysis of documented incidents categorized by limitations in perception, planning, and control; reviews research on human-interactive perception, language-grounded planning, and teleoperation; and concludes that reliable deployment requires augmenting current paradigms with cooperative human-interactive autonomy rather than relying on passive fallbacks.","tokens_in":1826,"tokens_out":385,"duration_ms":26897,"significance":"If the analysis and taxonomy are substantiated, the work could usefully direct attention to gaps in passive MRC strategies and synthesize directions for interactive autonomy research, potentially informing safety standards for urban AV integration. The review of emerging capabilities in multimodal interaction and remote guidance provides a constructive bridge between incident observations and technical research agendas.","major_comments":[{"comment":"Analysis section: the taxonomy maps incidents to perception, planning, and control limitations but supplies no incident counts, explicit selection criteria for the public records, or comparative evaluation against external factors (e.g., traffic regulations treating stopped AVs as obstacles or lack of V2I infrastructure). This is load-bearing for the central inference that internal AV architecture gaps are the primary driver necessitating replacement of passive MRCs with human-interactive autonomy rather than complementary policy or infrastructure measures.","section":"Analysis of publicly documented incidents"}],"minor_comments":[{"comment":"Abstract: states that an analysis of incidents is presented but provides no overview of methodology, data sources, or quantitative scope, reducing immediate evaluability.","section":"Abstract"},{"comment":"The review of research directions could benefit from more explicit linkage back to the specific incident categories identified in the taxonomy.","section":"Review of emerging research directions"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback. We address the major comment below and propose revisions to improve clarity and transparency.","responses":[{"response":"The analysis is qualitative, using publicly documented incidents to illustrate recurring failure patterns rather than providing a statistical survey. We agree that greater transparency is warranted. In revision, we will add a 'Data Sources and Methodology' subsection specifying the public records reviewed (municipal reports, news archives from cities with AV deployments), inclusion criteria (incidents where stopping behavior caused documented traffic obstruction or interaction failures), and approximate counts of incidents examined. On external factors, we will expand the discussion to acknowledge that regulations treating stopped vehicles as obstacles and the absence of V2I infrastructure contribute to the observed problems; however, the incidents still demonstrate that current AV architectures lack mechanisms for interpreting human authority or adapting to socially regulated conditions. This supports our position that interactive capabilities are a necessary complement to policy and infrastructure measures, not a sole replacement. These additions will clarify the scope of the inference without overstating the evidence.","revision_made":"yes","referee_comment":"[Analysis of publicly documented incidents] Analysis section: the taxonomy maps incidents to perception, planning, and control limitations but supplies no incident counts, explicit selection criteria for the public records, or comparative evaluation against external factors (e.g., traffic regulations treating stopped AVs as obstacles or lack of V2I infrastructure). This is load-bearing for the central inference that internal AV architecture gaps are the primary driver necessitating replacement of passive MRCs with human-interactive autonomy rather than complementary policy or infrastructure measures."}],"tokens_in":1329,"tokens_out":343,"duration_ms":32771,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The key point is that AV minimal risk conditions like stopping can block traffic, slow emergency vehicles, and strand passengers, so the authors want safety systems that handle human instructions and remote help instead.\n\nWhat the paper does is pull public incident reports into three buckets—perception limits, planning shortfalls, and control problems—then link those to missing human authority handling. It also rounds up work on language-based planning and teleoperation. That framing is useful for anyone tracking deployment friction.\n\nThe soft spot is the evidence base. The abstract and description mention an analysis of incidents and a taxonomy, but give no numbers, selection rules, or breakdown of how many cases fall into each category. Without that, it's hard to tell whether AV architecture really drives most failures or if traffic laws, missing V2I, and city response protocols matter more. The stress-test concern lands: the call to replace passive MRCs with cooperative autonomy assumes internal limits are primary, yet the paper does not test that against external factors.\n\nThis is a position paper aimed at AV safety researchers and urban planners who already follow HRI work. It synthesizes gaps rather than delivering new data or proofs.\n\nIt deserves peer review because the incidents it cites are documented and the research directions it flags are active. Referees can push for clearer methodology and a direct comparison of causes.","headline":"The paper flags real problems with AV stopping in cities but rests on a taxonomy without counts or proof that AV tech gaps are the main cause over rules and infrastructure.","tokens_in":2302,"tokens_out":352,"would_cite":false,"duration_ms":15085,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Autonomous vehicle safety must shift from passive stopping to interactive responses with humans for reliable urban use.","keywords":["autonomous vehicles","minimal risk conditions","human-AV interaction","safety frameworks","urban deployment","traffic incidents","interactive autonomy"],"falsifier":"A set of incidents where AV stopping behaviors occur without causing obstructions or interferences, or a demonstration that external factors alone explain all failures without architectural changes.","tokens_in":2620,"feed_emoji":"🚗","tokens_out":555,"duration_ms":22415,"temperature":0.7,"pith_summary":"The paper examines real-world cases where autonomous vehicles stopping as a safety measure instead cause traffic blocks, emergency interference, and accessibility problems. It breaks down these failures into issues with how AVs perceive, plan, and control in uncertain situations. The analysis reveals missing abilities to understand human directions or adapt to social traffic rules. It argues that safe city deployment needs AVs that can cooperate with people rather than just halt. This matters because current designs leave AVs unable to handle the human-governed nature of roads.","feed_headline":"Stopping alone fails to keep AVs safe in human traffic","feed_subtitle":"Review of real incidents shows AVs need abilities to interpret human authority and respond to instructions.","key_machinery":"Taxonomy of incidents categorized by limitations in perception, planning, and control within AV architectures, which exposes the absence of mechanisms for human authority interpretation and multimodal response in existing minimal risk conditions.","core_discovery":"Publicly documented incidents demonstrate that minimal risk conditions relying on stopping obstruct traffic flow, disrupt emergency services, and create barriers for users, stemming from gaps in perception, planning, and control that prevent interpreting human authority or responding to dynamic instructions; therefore, safety frameworks require augmentation with human-interactive capabilities for cooperative operation.","pith_inferences":["Urban AV testing protocols may need to include scenarios involving human interactions beyond stopping.","Policy for AV deployment could require interactive features to prevent obstruction issues.","Integration with smart infrastructure might enable better human-AV coordination."],"forward_implications":["AV architectures need added support for interpreting human authority in traffic situations.","Planning systems must handle multimodal instructions from humans and infrastructure.","Control mechanisms should adapt to socially regulated and dynamic traffic conditions.","Research into human-interactive perception and teleoperation can address these gaps."],"fun_headline_variants":["AV fallback stops block traffic and emergency response","Incidents reveal AV stops ignore human authority signals","AV safety paradigms must incorporate human interactive planning","Rethinking minimal risk conditions for dynamic human traffic"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That the documented failures stem primarily from internal AV limitations in perception, planning, and control rather than from regulatory, infrastructure, or other external factors.","fun_headline_variants_meta":{"raw":{"variants":["AV fallback stops block traffic and emergency response","Incidents reveal AV stops ignore human authority signals","AV safety paradigms must incorporate human interactive planning","Rethinking minimal risk conditions for dynamic human traffic"]},"model":"grok-4.3","cost_usd":0.005474,"raw_usage":{"total_tokens":2629,"prompt_tokens":664,"num_sources_used":0,"completion_tokens":56,"cost_in_usd_ticks":54737000,"prompt_tokens_details":{"text_tokens":664,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1909,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":664,"tokens_out":56,"duration_ms":19398,"temperature":1.0,"reasoning_tokens":1909,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T09:04:57.486296+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A set of incidents where AV stopping behaviors occur without causing obstructions or interferences, or a demonstration that external factors alone explain all failures without architectural changes.","supporting_citations":[],"review_version":1}