{"id":"2a56e061-b9e5-40df-8c06-ca04ede2968f","arxiv_id":"2603.16013","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"RAISE supplies reusable Reject-Instruction and Accept-Adequate-Instructions patterns plus an extended HARA that includes safe events, enabling structured safety cases for VLA driving systems, shown on SimLingo.","lead":"The paper proposes RAISE, a safety-case method with two new GSN patterns and an extended HARA process for Vision-Language-Action driving systems that take open-ended natural-language instructions. It demonstrates the method by building a partial safety case for the SimLingo system in the CARLA simulator.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The central claim that RAISE yields rigorous, evidence-based safety cases rests on unshown evidence linkage and untested pattern transfer beyond one CARLA system.","rationale":"The Reader correctly isolates the single-case, CARLA-centric threat to validity (§4.2) as the weakest assumption. My stress test sharpens the same point: even for the one system that was studied, the manuscript stops at goal decomposition and never closes the argument with evidence. That gap is load-bearing for the adjective “rigorous, evidence-based.” Because the methodological contribution (Safe-Events HARA extension + two reusable GSN patterns + constructive algorithm) is still novel and useful, and the authors themselves flag the limitation, the appropriate verdict remains CONDITIONAL rather than REJECT. The concrete test above would settle whether the published SimLingo case already meets the claim or whether additional work is required before the claim can be accepted. No stronger internal inconsistency or formal error was found; the concern is incompleteness of demonstration, not contradiction.","tokens_in":9973,"tokens_out":657,"duration_ms":6500,"concrete_test":"Take the published SimLingo GSN fragments (G3.2/G3.3 and their children in Figs. 6–7). For every leaf goal, attach the concrete evidence artifact that the algorithm claims exists (e.g., CARLA rejection-rate logs for dangerous reverse/overtake instructions under OS1–OS9). If more than half the leaves remain unsupported by measurable evidence, or if the same RI/AAI skeleton cannot be instantiated for a second VLA system (e.g., LMDrive) without redesigning the hot-spots, the “rigorous, evidence-based” claim does not hold for the published work.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim asserts that RAISE (extended HARA + RI/AAI patterns + algorithm) enables systematic construction of rigorous, evidence-based safety claims, illustrated by SimLingo. The paper shows partial GSN trees (Figs. 5–7) whose leaf goals are intended to be supported by concrete evidence, and the algorithm description states that decomposition continues “until they can be supported by concrete evidence from the VLA-based driving system under analysis.” Yet the manuscript never exhibits any actual evidence nodes (test logs, rejection rates, scenario coverage metrics, or CARLA Leaderboard outcomes) attached to those leaves, nor any quantitative demonstration that SimLingo’s instruction-handling behavior satisfies the safety goals derived from HARA. The patterns themselves are derived almost exclusively from SimLingo’s system functions (Table 1) and a short list of operational scenarios (Table 2), with only a secondary nod to LMDrive. Consequently the claim that the same decomposition structure produces rigorous, transferable safety cases for other VLA architectures or real-world ODDs is an untested extrapolation. Without either (a) completed evidence attachment for the published SimLingo case or (b) a second independent instantiation, the “rigorous, evidence-based” part of the central claim remains aspirational rather than demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes RAISE, a safety-case design approach for Vision-Language-Action (VLA) driving systems that accept open-ended natural-language instructions. RAISE comprises (i) an extension of classic HARA that adds a “Safe Events” step recording expected outcomes of non-hazardous instructions, (ii) two novel GSN patterns—Reject Instruction (RI) and Accept Adequate Instructions (AAI)—that decompose the dual objectives of rejecting dangerous instructions and accepting safe ones under common operational scenarios, and (iii) a constructive algorithm that iteratively instantiates those patterns from HARA outcomes down to evidence leaves. Applicability is illustrated by a partial GSN safety case for the CARLA-based SimLingo system (system functions in Table 1, operational scenarios in Table 2, top-level argument in Fig. 5, and pattern instantiations in Figs. 6–7), with supporting HARA tables and algorithm code released on GitHub.","tokens_in":10302,"tokens_out":976,"duration_ms":9861,"significance":"Instruction-conditioned VLA controllers introduce a genuinely new hazard class that existing ML-assurance frameworks (AMLAS, BIG-Argument, etc.) do not address. Supplying reusable GSN patterns, an explicit Safe-Events extension of HARA, and a constructive algorithm is a useful first step for the assurance community. The public GitHub artifacts (HARA tables + algorithm) and the concrete SimLingo GSN fragments give practitioners a concrete starting point. If the patterns prove transferable, the work would lower the cost of producing instruction-aware safety cases for an emerging class of autonomous systems.","major_comments":[{"comment":"The abstract and §1 claim that RAISE constructs “rigorous, evidence-based safety claims,” yet the published SimLingo case (Figs. 5–7) stops at intermediate goals. No evidence nodes (test logs, rejection-rate metrics, scenario-coverage results, or CARLA Leaderboard outcomes) are attached to the leaves, nor is any quantitative demonstration given that SimLingo’s instruction-handling satisfies the HARA-derived safety goals. Without at least one completed evidence chain for the published case, the “evidence-based” part of the central claim remains aspirational.","section":null},{"comment":"§3.2 and §4.2 acknowledge that the RI/AAI patterns and the operational-scenario catalogue (Table 2) were extracted primarily from SimLingo (with only a secondary nod to LMDrive). The claim of systematic applicability to other VLA architectures, training regimes, or real-world ODDs therefore rests on an untested transfer assumption. A second, independent instantiation—or an explicit argument why the same decomposition structure is architecture- and ODD-invariant—is needed before the general-applicability claim can be regarded as demonstrated.","section":null}],"minor_comments":[{"comment":"Table 2 header reads “System Function” while the column lists operational scenarios; the header should be corrected to “Operational Scenario”.","section":null},{"comment":"§3.2 refers to “our algorithm is available on GitHub” but does not reproduce even a high-level pseudocode listing in the manuscript; a short algorithmic sketch would improve self-containment.","section":null},{"comment":"Figures 3 and 4 are described as containing “hot spots,” yet the published figures do not visually distinguish placeholders from fixed text; a typographic convention would help readers instantiate the patterns.","section":null},{"comment":"Several sentences contain missing spaces or concatenated words (e.g., “addressingVLA-baseddrivingsystems,” “Thecreationofsafetycases”); a careful copy-edit pass is needed.","section":null},{"comment":"Related-work discussion of AMLAS, BIG-Argument and STAMP/STPA is present but brief; a short comparison table clarifying which hazards each method can (and cannot) address would strengthen positioning.","section":null}],"recommendation":"major_revision","confidential_remarks":"The contribution is timely and the dual-objective framing (reject dangerous / accept adequate) is sound. The main risk is over-claiming: the manuscript currently demonstrates a reusable argument skeleton, not a finished evidence-based safety case. If the authors can attach concrete evidence leaves for SimLingo and either add a second case or carefully scope the transfer claim, the paper would be a solid contribution for an assurance or robotics venue. Scope fit for a pure robotics journal is acceptable given the VLA focus; a safety/assurance venue might be an even better home."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that this paper is the first clear attempt I have seen to treat open-ended natural-language instructions as first-class hazards that a VLA driver must accept or reject. That gap is real; AMLAS, BIG-Argument and the usual STPA applications do not give you reusable structure for it.\n\nWhat is actually new: the RI and AAI GSN patterns, the “Safe Events” extension of HARA that records non-hazardous instruction outcomes, the RAISE construction algorithm, and the concrete (if partial) SimLingo fragments. The GitHub release of the algorithm and HARA tables is a genuine plus. The authors correctly derive the two top-level goals from system functions, keep the threats-to-validity section honest, and do not oversell transfer beyond CARLA. Citation pattern is clean; they know the prior assurance literature and the VLA driving papers.\n\nSoft spots, in proportion. The stress-test note is right on the evidence gap: Figures 5–7 stop at leaf goals; no test logs, rejection rates, scenario coverage numbers or CARLA Leaderboard outcomes are attached. So the claim that RAISE produces “rigorous, evidence-based” safety cases is still a method claim, not a demonstrated result. The patterns and the short operational-scenario list are extracted almost entirely from SimLingo (with a nod to LMDrive), then reapplied to SimLingo; that is mild self-application, not circularity, but it leaves transfer untested. Expert judgment is acknowledged but unquantified. None of these are fatal for a methods paper; they simply mean the central claim is not yet fully cashed.\n\nWho it is for: automotive assurance engineers and regulators who need a starting template for instruction-following AVs. A serious referee should see it; the gap is genuine and the artifacts are usable. I would bring it to reading group as a methods discussion, cite the patterns when I next write about instruction hazards, and expect revision that either attaches evidence to the SimLingo leaves or adds a second system.","headline":"Useful first patterns for instruction-driven VLA safety cases, but the SimLingo case stops at partial GSN trees without attached evidence, so the “rigorous, evidence-based” claim is still aspirational.","tokens_in":10903,"tokens_out":515,"would_cite":true,"duration_ms":5991,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"RAISE shows how to build safety cases for cars that take open-ended language instructions, by arguing when to reject dangerous ones and when to accept safe ones.","keywords":["safety case","safety case patterns","VLA-based driving system","HARA","Goal Structuring Notation","instruction-based autonomy","SimLingo","assurance cases"],"falsifier":"Apply RAISE unchanged to a second VLA driving system trained on different data and operating outside the CARLA simulator; if the same patterns cannot produce a coherent, evidence-linked safety case without inventing substantially new structure, the transfer claim fails.","tokens_in":10832,"feed_emoji":"🚗","tokens_out":809,"duration_ms":8722,"temperature":0.7,"pith_summary":"Vision-language-action driving systems can take free-form user or navigation instructions and turn them into vehicle actions. That flexibility creates hazards that older autonomous-driving safety methods never covered: a vehicle might reverse, accelerate, or turn on command without checking whether the scene is safe. This paper claims those risks can be managed systematically with RAISE. RAISE extends classic hazard analysis so it records both dangerous and safe instruction outcomes, supplies two reusable argument patterns (Reject Instruction and Accept Adequate Instructions), and gives an algorithm that turns those patterns plus the analysis results into a Goal Structuring Notation safety case. The authors walk through the method on the SimLingo system, producing concrete claims that the vehicle can refuse unsafe commands and follow safe ones under the operational scenarios they examined. If the method holds, safety engineers finally have a structured way to justify instruction-following autonomy rather than treating language inputs as an afterthought.","feed_headline":"Safety cases for cars that obey spoken instructions","feed_subtitle":"RAISE patterns show when a VLA driver must refuse a command and when it may obey","key_machinery":"RAISE: an extended HARA that adds Safe Events, two GSN patterns (Reject Instruction and Accept Adequate Instructions), and an algorithm that decomposes top-level safety goals into evidence-backed claims using those patterns and the HARA results.","core_discovery":"The paper establishes that safety cases for VLA-based driving systems can be constructed systematically by combining an extended HARA that captures both hazardous and safe instruction outcomes, two novel GSN patterns focused on rejecting dangerous instructions and accepting adequate ones, and a constructive algorithm that instantiates those patterns against operational scenarios. The SimLingo case study shows the resulting argument structure.","pith_inferences":["The same reject/accept pattern pair may apply to any embodied agent that takes open-ended natural-language commands, not only road vehicles.","Once Safe Events are routine in HARA, argument libraries can be auto-populated from scenario catalogues, reducing manual safety-case cost.","Regulators may eventually require evidence that a VLA system both refuses unsafe instructions and still reaches its destination under safe ones."],"forward_implications":["Safety engineers can treat instruction acceptance and rejection as first-class safety goals rather than informal add-ons.","Hazard analysis for language-driven vehicles must document safe instruction outcomes, not only hazardous ones.","Reusable GSN patterns for reject/accept decisions become available for other instruction-based autonomous systems.","Regulatory and corporate assurance teams gain an explicit checklist for decomposing VLA safety claims down to concrete evidence."],"fun_headline_variants":["RAISE patterns build safety cases for VLA drivers rejecting unsafe spoken orders","Safety arguments for cars that obey language via extended HARA and GSN patterns","SimLingo shows how to claim VLA systems only accept adequate instructions","Patterns that prove instruction-following drivers refuse hazardous commands","From HARA to GSN: constructing safety cases for VLA spoken-command control"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The two patterns and the short list of operational scenarios drawn mainly from SimLingo (and secondarily from another system) are representative enough to transfer to other VLA architectures and real-world driving domains without major redesign.","fun_headline_variants_meta":{"raw":{"variants":["RAISE patterns build safety cases for VLA drivers rejecting unsafe spoken orders","Safety arguments for cars that obey language via extended HARA and GSN patterns","SimLingo shows how to claim VLA systems only accept adequate instructions","Patterns that prove instruction-following drivers refuse hazardous commands","From HARA to GSN: constructing safety cases for VLA spoken-command control"]},"model":"grok-4.5","effort":"low","cost_usd":0.003796,"raw_usage":{"total_tokens":1210,"prompt_tokens":770,"num_sources_used":0,"completion_tokens":98,"cost_in_usd_ticks":37960000,"prompt_tokens_details":{"text_tokens":770,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":342,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":770,"tokens_out":98,"duration_ms":4192,"temperature":1.0,"reasoning_tokens":342,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T00:00:48.864828+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Apply RAISE unchanged to a second VLA driving system trained on different data and operating outside the CARLA simulator; if the same patterns cannot produce a coherent, evidence-linked safety case without inventing substantially new structure, the transfer claim fails.","supporting_citations":[],"review_version":1}