{"id":"baa45a39-9a4a-4a45-a38a-20aef46be7f4","arxiv_id":"2607.09045","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Cloud driving for AVs is communication-bound first, then VLA-compute-bound under 100 ms, then cost-bound; feature-level offloading concentrates the VLA economic crossover once admissible.","lead":"Cloud offloading for frontier autonomous-driving models is feasible only under three nested constraints—uplink capacity, reactive latency, then cost—with feature-level splits becoming economical for VLA models once memory bandwidth improves. The analysis maps when shared cloud GPUs can replace expensive idle onboard hardware in dense cities like New York.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The VLA compute wall is stack-specific; non-AR or FP8 decoders would erase the year-by-year reactive-admissibility claims the paper treats as structural.","rationale":"The reader correctly isolates the load-bearing premise: the Alpamayo/TensorRT-LLM autoregressive-FP16 roofline that produces the ~114 ms decode wall. That premise is what makes compute bind after communication and before cost, and what generates the concrete year numbers (floor clears ~2027; 6G admits feature-level VLA ~2028). The paper is unusually honest about the stack-specificity in Sec. 4.2 and the Limitations, so the concern is not a hidden error but an untested sensitivity that the main claims still treat as structural. Communication-regime results (S1 eliminated early; 5G-Advanced as S2 threshold) and the qualitative cost-sharing argument for VLA are more robust and do not depend on the decoder stack. Because the paper already frames the answer as conditional on model class, GPU year, and latency budget, the appropriate verdict remains CONDITIONAL; the concrete decoder-sensitivity check is the single experiment that would decide whether the year-by-year compute map survives. No stronger internal inconsistency is present; the equations and tables are complete enough that the test is well-defined.","tokens_in":29730,"tokens_out":786,"duration_ms":8748,"concrete_test":"Re-evaluate Eq. (9) and the jointly-feasible branch counts of Sec. 4.2 under two alternate stacks: (i) FP8 weights (halve W_reasoning and W_traj) and (ii) a non-autoregressive action head with D=0 and n_traj replaced by a single parallel decode of ~20–40 ms. Recompute L_det(t) and the 2026/2028/2030 VLA feasible-branch fractions under the 100 ms budget. If either stack admits VLA-S2 under 5G-Advanced at the dense corridor (10% pen., u=0.45) before 2027, the compute-bound regime and its year claims do not hold as stated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The nested-regime claim that compute binds next—and that near-term VLA is latency-infeasible regardless of bandwidth—rests on Eq. (9)’s roofline service time remaining dominated by fully autoregressive FP16 decode. On B300 the paper reports a 153 ms floor with ~114 ms HBM-bound decode (D=16, W_reasoning=15.6 GB, n_traj=17, W_traj=22.9 GB). That floor first clears 100 ms only around 2027, which is what creates the compute-bound regime and all subsequent year/generation admissibility numbers (0/432 VLA branches in 2026 under the reactive budget; 6G admits VLA-S2 ~2028; 5G-Advanced only at light loading). The paper itself notes (Sec. 4.2 and Limitations) that FP8 roughly halves the per-token read (~57 ms) and that parallel diffusion/flow action decoders remove the autoregressive re-reads entirely, moving reactive admissibility several years earlier. If either becomes the serving default before ~2027, the compute-bound regime collapses into the communication-bound regime and the year-by-year map ceases to be structural. The nesting of regimes is therefore conditional on a conservative decoder stack that the paper flags but does not sensitivity-test in the main results.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper asks whether cloud/edge inference can economically drive autonomous vehicles under closed-loop latency constraints. It couples three systems—an offloading-strategy spectrum (S1 raw-sensor / S2 feature-level / S3 query-level) for E2E, VLM, and VLA models; an interference-aware 5G/5G-Advanced/6G uplink admission model; and a roofline GPU service model with stochastic tail latency and utilization-aware TCO—into a joint feasibility gate and cost-crossover analysis applied to a 1,296-branch New York City matrix. Under a reactive 100 ms budget and a deliberative 300 ms tier (the latter only behind an onboard reactive fallback), the authors report three nested binding regimes: communication binds first in dense cells (5G fails early; 5G-Advanced is the practical threshold for S2), compute binds next for near-term VLA because autoregressive FP16 decode is memory-bandwidth-bound (~114 ms on 2025 hardware; floor clears ~2027), and cost binds last once a branch is admissible (utilization-pooled cloud GPUs undercut expensive idle VLA onboard hardware, with the crossover concentrating at S2). Latency decides which model is admissible in which year; cost decides whether it is economical.","tokens_in":30173,"tokens_out":1418,"duration_ms":11600,"significance":"If the nested-regime picture holds, the paper supplies a concrete, operator-facing deployment sequence that the three literatures (driving models, vehicular communications, vehicular edge computing) have not previously integrated. The framework is internally consistent, parameterizes independently sourced inputs (3GPP/ITU link budgets, NVIDIA datasheets, DOE parking statistics, published model FLOPs), and produces falsifiable year/generation admissibility maps and cost crossovers rather than a binary feasibility claim. The explicit separation of reactive and deliberative budgets, the residual-TOPS accounting for S1–S3, and the utilization-driven TCO comparison are useful contributions for spectrum planning, shared-edge co-investment, and safety certification of cloud-dependent AV stacks. The main result is conditional on a conservative decoder stack, but the paper itself flags that condition and the framework can absorb alternative architectures by updating Table 2 and Eq. (9).","major_comments":[{"comment":"§4.2 and Eq. (9): The compute-bound regime—and all subsequent year/generation VLA admissibility numbers (0/432 reactive branches in 2026; floor clears ~2027; 6G admits VLA-S2 ~2028; 5G-Advanced only at light loading)—rests on the premise that service time remains dominated by fully autoregressive FP16 decode (~114 ms HBM-bound on B300). The paper correctly notes that FP8 roughly halves the per-token read and that parallel diffusion/flow action decoders remove the autoregressive re-reads, moving reactive admissibility years earlier. Because this is load-bearing for the nesting claim, the main results need at least a one- or two-scenario sensitivity (e.g., FP8 and a non-AR decoder) so readers can see which regime boundaries survive when the conservative stack is relaxed. Without that, the year-by-year map is presented as more structural than the Limitations section itself allows.","section":null},{"comment":"§3.8 / Eq. (28) and §4.3 / Fig. 10: Deliberative-tier cost ratios charge only each strategy’s residual onboard hardware and explicitly exclude the onboard reactive-fallback controller that the 300 ms tier requires by construction (§3.5.3, §5). The manuscript labels the resulting VLA cost-attractive region an “optimistic bound,” but the bound is not quantified. Because the paper’s economic takeaway is that S2 concentrates the VLA crossover, a simple additive fallback cost (or a sensitivity band) is needed so the deliberative TCO comparison is not systematically biased toward cloud.","section":null},{"comment":"§3.4.3 / Table 4 and Eqs. (10)–(11): GPU and HBM evolution are calibrated to four A100→B300 points with a decelerating-rate model (r0_ϕ=64%/yr → r∞_ϕ=10%/yr, λ=0.15). The 2027 floor-clearing year and the 2028 6G-admission claim inherit this extrapolation. A short sensitivity on the long-run floors (or on λ) would show how much the compute-regime timeline moves under slower or faster memory-bandwidth growth; without it the year labels are more precise than the calibration supports.","section":null}],"minor_comments":[{"comment":"Table 3 / residual H_s: The residual INT8 TOPS for VLA under S2 (550) and S3 (2900) versus the 3000 TOPS full baseline should be cross-checked against the phase FLOPs in Table 2 and the 2:1 INT8/FP16 conversion; a one-sentence derivation would help reproducibility.","section":null},{"comment":"Fig. 3 and Fig. 5: Axis labels and the capacity-cliff annotation are dense; a short caption note defining N_max_c and Δ would improve readability for non-communications readers.","section":null},{"comment":"§3.5.2: The processor-sharing approximation for L_nq and the Chernoff/MGF tail for Eq. (19) are standard but cited lightly; a pointer to the exact MGF forms used would aid replication.","section":null},{"comment":"References: Alpamayo is cited as a CES 2026 / Hugging Face release; ensure the archival citation is stable before camera-ready. A few 6G V2X surveys already in the bibliography could be tied more explicitly to the uplink-heavy framing in §5.2.","section":null},{"comment":"Notation: μ_eff(s,m,t) is used both as service rate and (inverted) as service time in Eq. (9); a single convention would reduce friction.","section":null}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a strong systems-integration piece for eess.SY / vehicular networking venues. The main risk is over-precision of the year-by-year VLA map under a stack the authors themselves flag as conservative; requiring the sensitivity and the deliberative-fallback cost band should be sufficient. No novelty or citation-pattern concerns. Fit is good for a journal that values infrastructure feasibility over pure algorithmic novelty."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is a single pipeline that forces the three constraints into order—uplink admission, then reactive latency, then utilization cost—and runs them on one dense-city matrix instead of treating them as separate literatures. That nesting is the real contribution, not any single equation.\n\nWhat they do well: the communication gate is interference-aware (not just bandwidth/n), the dual 100/300 ms budgets are honest about onboard fallback, residual TOPS vs cloud FLOPs are tracked per split, and the cost side correctly prices shared GPUs against parked vehicles. The NYC branch counts (487 reactive vs 903 deliberative in 2026) and the S2-centric VLA crossover are concrete and previously unpublished. Citations look right—3GPP/ITU, NVIDIA datasheets, UniAD/DriveLM/Alpamayo FLOPs, DOE parking stats—and the limitations section flags the soft spots instead of burying them.\n\nThe stress-test is fair but overstated as a collapse. The ~114 ms HBM-bound decode on B300 is stack-specific (AR FP16, Alpamayo/TensorRT-LLM). FP8 or parallel action decoders would move the reactive wall earlier, as the paper itself says. That weakens the precise year map (2027 floor, 2028 6G S2), not the qualitative order: communication still binds first in dense cells, cost still binds last once latency is cleared, and S2 remains the robust middle. They should have put a one-page sensitivity on decoder variants in the main results; without code that is the main re-implementation friction. 6G is aspirational and utilization is static—both minor relative to the framework.\n\nThis is for people who size edge fleets, spectrum, or AV safety stacks, not for pure ML or pure radio theorists. Math and data are solid enough for a serious referee. I would engage, cite the regime map and S2 economics, and ask for the decoder sensitivity plus code. Send it to review.","headline":"Solid systems map of when cloud AV inference is feasible; the year-by-year VLA wall is stack-specific, but the nested-regime framing and S2 economics still hold.","tokens_in":30792,"tokens_out":515,"would_cite":true,"duration_ms":7903,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Cloud can drive only after uplink, then memory-bound VLA latency, then cost each clear in sequence.","keywords":["autonomous driving","vehicular edge computing","task offloading","5G/6G communications","vision-language-action models","infrastructure feasibility","roofline GPU model","utilization-aware cost"],"falsifier":"Measure end-to-end reactive-loop latency of a production VLA stack on 2025-class hardware that uses FP8 weights or a parallel (diffusion/flow) action decoder; if the decode floor falls well below 100 ms before 2027, the compute-bound regime and all year-by-year admissibility claims collapse.","tokens_in":30608,"feed_emoji":"🚗","tokens_out":849,"duration_ms":7071,"temperature":0.7,"pith_summary":"Frontier autonomous-driving models, especially vision-language-action systems that cost tens of TFLOPs per decision, make full onboard hardware wasteful: peak chips sit idle most of the day. Cloud inference can share GPUs across active vehicles, but only if the vehicle can upload through a capacity-limited cell, reach a free GPU, and return a decision inside a closed-loop deadline. The paper couples communication limits, a roofline GPU service model, stochastic latency, and utilization-aware cost across three model classes, three offloading splits, and three radio generations, then applies the whole stack to New York City under a tight 100 ms reactive budget and a looser 300 ms deliberative tier. Three nested constraints bind in order: dense cells first make the uplink the bottleneck (5G fails early; 5G-Advanced is the practical threshold for feature-level offloading); under the reactive budget, near-term VLA remains latency-infeasible regardless of bandwidth because autoregressive decoding is memory-bandwidth-bound; only after both gates clear does utilization-pooled cloud cost undercut expensive idle onboard VLA hardware. Latency therefore decides which model is admissible in which year; cost decides whether it is economical.","feed_headline":"Cloud can drive only after three nested gates clear","feed_subtitle":"Uplink first, then memory-bound VLA latency, then cost; feature-level split is the sweet spot","key_machinery":"Three nested binding regimes produced by a joint analytical pipeline: interference-aware uplink admission, a roofline GPU service model that separates compute-bound encoder/prefill from HBM-bandwidth-bound autoregressive decode, stochastic tail-latency feasibility under 100 ms reactive and 300 ms deliberative budgets, and utilization-aware total-cost-of-ownership crossover.","core_discovery":"In a single dense-city setting the feasibility of cloud-driven autonomy is governed by three nested binding regimes. Communication binds first: 5G cannot sustain feature-level offloading under realistic loading, 5G-Advanced is the practical threshold, and 6G supplies headroom. Compute binds next under the 100 ms reactive budget: near-term VLA is latency-infeasible regardless of bandwidth because autoregressive FP16 decode is memory-bandwidth-bound (about 114 ms of a 153 ms floor on 2025 hardware); the floor clears 100 ms around 2027, after which 6G admits feature-level VLA by about 2028 while 5G-Advanced does so only at light loading. Cost binds last: once admissible, shared cloud GPUs under","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Three nested gates: uplink then VLA latency then cost","Cloud drives only after comms, compute, cost regimes clear","Feature-level split is the VLA cloud sweet spot","5G fails early; 5G-Advanced opens feature-level offload","Memory-bound VLA decode blocks reactive cloud until ~2027"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim that near-term VLA stays latency-infeasible under a 100 ms reactive loop rests on the premise that service time remains dominated by fully autoregressive FP16 decode whose memory-bandwidth floor is accurately given by the paper's roofline calibration.","fun_headline_variants_meta":{"raw":{"variants":["Three nested gates: uplink then VLA latency then cost","Cloud drives only after comms, compute, cost regimes clear","Feature-level split is the VLA cloud sweet spot","5G fails early; 5G-Advanced opens feature-level offload","Memory-bound VLA decode blocks reactive cloud until ~2027"]},"model":"grok-4.5","effort":"low","cost_usd":0.004902,"raw_usage":{"total_tokens":1521,"prompt_tokens":1008,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":49020000,"prompt_tokens_details":{"text_tokens":1008,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":441,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":1008,"tokens_out":72,"duration_ms":4922,"temperature":1.0,"reasoning_tokens":441,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T00:44:20.799122+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Measure end-to-end reactive-loop latency of a production VLA stack on 2025-class hardware that uses FP8 weights or a parallel (diffusion/flow) action decoder; if the decode floor falls well below 100 ms before 2027, the compute-bound regime and all year-by-year admissibility claims collapse.","supporting_citations":[],"review_version":1}