{"id":"8bb29b7f-533a-41e7-a119-2e509ce9048d","arxiv_id":"2505.11474","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"REACT uses a kinetic-energy risk field and grid-based directional warnings to trigger collision avoidance, reporting 100% success in four vehicle trials.","lead":"This paper introduces REACT, a driving-risk scoring system that turns surrounding vehicles into a risk field and triggers warnings or evasive actions in real time. The authors report perfect avoidance and zero false alarms in four small on-road tests, but the evidence and calibration are weaker than the claims.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Zero false alarms/misses claim rests on 8 trials with hand-set thresholds; no held-out validation or quantitative baseline comparison.","rationale":"The reader's weakest assumption was that the hand-set thresholds and field coefficients may have been tuned to the same test scenarios, making the evaluation non-independent. My concern is broader but related: even if the thresholds were fixed a priori, the empirical basis is statistically insufficient (8 trials, 0 events) and lacks any quantitative baseline comparison. Thus the central claim of 100% safe avoidance with zero false alarms or missed detections is not supported. I agree with the reader's REJECT verdict, so no change to the verdict is needed. The concrete test I propose would settle whether the concern lands: a pre-registered held-out evaluation with thresholds frozen before data collection and a larger sample, plus a quantitative highD metric. This is a fair and feasible check, not an accusation of dishonesty. The paper could become a valid incremental proposal after such validation, but as written the evidence does not back the headline claims.","tokens_in":17082,"tokens_out":3330,"duration_ms":33562,"concrete_test":"Run a pre-registered held-out evaluation: freeze T1=0.3, T2=0.7 and all field parameters (a, b, β, k_lane, λ_j) before any data collection, execute at least 30 trials across the four scenarios, and compute exact binomial 95% confidence intervals for the false-alarm and miss rates. If the intervals are not entirely below 5% (which they will not be for any finite sample), the 'zero false alarms/misses' claim must be weakened to a rate bound. Additionally, report a quantitative highD comparison (e.g., warning-time error or precision-recall) against TTC, THW, RSS, and DRF.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—100% safe avoidance with zero false alarms or missed detections—rests on Table 4, which reports only 8 REACT trials in total (and 8 human-driver trials). With zero observed events, the exact 95% binomial confidence interval for the event rate is approximately [0, 0.37] (rule of three: 3/8), so the data cannot support a literal zero rate. Moreover, the thresholds T1=0.3 and T2=0.7 (Section 3.2.1, Eq. 11) and the field parameters a=0.2·||v_j||, b=5 m, λ_j, β, k_lane (Eqs. 5–7) are presented as fixed constants with no documented calibration procedure or held-out split. If these values were chosen or adjusted using the same eight trials whose outcomes are reported as 0% false alarms and 0% misses, the evaluation is not an independent test. The highD comparison (Figs. 3–4) is also only qualitative: no quantitative metric such as warning-time error, AUC, or precision-recall is computed, and no significance testing is performed against the five baselines, so the 'state-of-the-art accuracy' claim is unsupported. Finally, Theorem 1 is tautological: it assumes the existence of a true trajectory ℓ* and that the system actively responds to that trajectory, which is exactly the property the experiments are supposed to establish. The combination of a tiny sample, potentially tuned parameters, and no independent baseline comparison means the headline claims are not supported by the evidence presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes REACT, a runtime risk-assessment and active collision-avoidance framework for autonomous driving. The method builds a kinetic-energy-based risk field with directional and road-constraint terms, discretizes it on a grid around the ego vehicle, computes global and directional risk values, and maps them through dynamically adjusted thresholds (T1=0.3, T2=0.7) to three warning levels. The evaluation consists of a qualitative comparison on two highD highway scenarios against TTC, THW, RSS, and DRF, and on-vehicle experiments in four scenarios (car-following braking, cut-in, rear-approaching, intersection conflict). The paper claims 100% safe avoidance with zero false alarms or missed detections, warning lead time under 0.4 s, latency below 50 ms, and state-of-the-art accuracy.","tokens_in":17482,"tokens_out":3342,"duration_ms":34082,"significance":"If properly supported, REACT would be a useful engineering contribution: it combines a lightweight, interpretable risk field with a deployable warning system and demonstrates real-vehicle operation. The authors deserve credit for building a working prototype, measuring latency on embedded hardware, and comparing qualitative risk curves against human driver responses in realistic scenarios. However, the headline claims—zero false alarms/misses, warning times consistent with human cognition, and state-of-the-art accuracy—are not supported by the evidence as presented. The empirical base is eight REACT trials, the thresholds and field coefficients are hand-set with no documented calibration, the highD comparison is purely qualitative, and Theorem 1 is tautological. These limitations directly affect the central contribution, so the manuscript is not ready for publication in its current form.","major_comments":[{"comment":"The abstract and conclusion claim 'zero false alarms or missed detections,' but Table 4 reports only eight REACT trials total. With zero observed events across eight trials, the exact 95% binomial confidence interval for the event rate extends to roughly 0.31–0.37 (rule of three: 3/8), so the data cannot support a literal zero rate. Please report per-scenario trial counts, define the operational criteria for a false alarm and a miss, and provide confidence intervals or Bayesian posterior intervals for the miss and false-alarm rates.","section":"§4.3, Table 4"},{"comment":"The thresholds T1=0.3 and T2=0.7, together with the field coefficients λj, β, k_lane, a=0.2·||v_j||, and b=5 m in Eqs. (5)–(7), are presented as fixed constants with no documented calibration procedure or held-out validation. If these values were chosen or adjusted with the same eight trials whose outcomes are reported as 0% false alarms and 0% misses, the evaluation is conditional on the tuned parameters and is not an independent test. Please document the calibration data, the fitting procedure, and either a held-out split or a sensitivity analysis over plausible parameter ranges.","section":"§3.2.1, Eq. (11) and §4.2"},{"comment":"Theorem 1 is stated as a foundational result but is tautological: it assumes the existence of a true trajectory ℓ* and that the system actively responds to ℓ*, which is precisely the property that the experiments are meant to establish. No proof or formal definition of the ADS policy set is provided. Either remove the theorem or replace it with a substantive, provable safety statement (e.g., a condition under which the risk field guarantees a minimum separation distance).","section":"§3.1, Theorem 1"},{"comment":"The highD comparison against TTC, THW, RSS, and DRF is qualitative only. The text asserts that REACT 'accurately captures interaction dynamics' and that DRF 'shows delayed response,' but no quantitative metrics—such as warning-time error, ROC/AUC, precision-recall, or statistical significance tests—are computed for any baseline. The claim of 'state-of-the-art accuracy' in the abstract is therefore unsupported. Please add quantitative comparisons on a defined set of conflict events with appropriate metrics and error bars.","section":"§4.1, Figs. 3–4"},{"comment":"Several mathematical inconsistencies compromise reproducibility. Eq. (2) is missing the squares under the square root; Eq. (3) is numbered twice with different content; Eq. (5) uses exp(−r̃_ij) while Algorithm 1 line 6 uses exp(−r̃_ij^2), and the exponent notation in Eq. (5) is garbled. In addition, v_ij is defined as a scalar magnitude in Eq. (1) but used as a vector in Eqs. (3) and (5). These must be corrected, and the definitions must be consistent between the main text and the algorithm pseudocode.","section":"§3.1.1, Eqs. (1)–(5) and Algorithm 1"}],"minor_comments":[{"comment":"The relative position r_ij should be defined with squared coordinate differences: r_ij = sqrt((x_i − x_j)^2 + (y_i − y_j)^2). The current expression is dimensionally incorrect.","section":"§3.1.1, Eq. (2)"},{"comment":"Table 3 references Fig. 5(a)–(d) for the scenario illustrations, but the text and Fig. 6 indicate that the scenarios are shown in Fig. 6. Please make the cross-references consistent.","section":"Table 3 and §4.2.2"},{"comment":"Two of the risk-curve panels contain untranslated Chinese text ('风险曲线' and '风 险 曲 线'). Since the manuscript is in English, these labels should be translated.","section":"Figs. 7–8"},{"comment":"The text says 'In the RV scenario, false alarms and missed detections occurred in driver behavior,' and attributes the human driver's 25% false alarm and miss rates to premature or absent responses. This is interesting, but it would be clearer to define how a false alarm or miss is scored for a human driver versus the REACT system.","section":"§4.3"},{"comment":"The data availability statement says 'Data will be made available on request.' Given the paper's reproducibility claims, please consider releasing the scenario trajectories, parameter settings, and evaluation scripts alongside the paper.","section":"Data availability"}],"recommendation":"reject","confidential_remarks":"The paper has a genuine engineering contribution in the form of a deployed lightweight risk-field warning system, but the current evaluation does not support the central claims of zero false alarms/misses and state-of-the-art accuracy. The issues are not merely editorial: they require a substantially redesigned empirical study (larger trial counts, independent calibration, quantitative baselines) and a rework of the theoretical framing. This is better addressed in a new submission than in a revision of the current manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take on arXiv:2505.11474. The paper is a reasonable incremental engineering contribution: REACT combines a kinetic-energy-based risk field, elliptical directional distance, grid-based sector attribution, and hierarchical warning thresholds into a lightweight closed-loop framework. The method is clearly described, the modular architecture is sensible, and they did real on-vehicle trials in four high-risk scenarios. That's more than many papers in this space do. The highD comparison, though qualitative, shows that TTC/THW miss lateral cut-in risk while REACT flags it. So there is a genuine within-subfield contribution here.\n\nThe soft spots are in the evidence, not the concept. The central claim—0% false alarms, 0% misses, 100% safe avoidance—rests on eight REACT trials. With zero observed events, the 95% CI for the event rate is roughly [0, 0.37]. That cannot support a literal zero rate. The thresholds T1=0.3, T2=0.7 and the other field coefficients are presented as fixed, but there is no calibration procedure or held-out split. If they were tuned with these trials in view, the evaluation is not independent. The highD figures are illustrative only; no quantitative metric (AUC, warning-time error, precision-recall) is reported, so 'state-of-the-art accuracy' is unsupported. Theorem 1 is a tautology: assuming a true trajectory and that the system responds to it is exactly what the experiments are supposed to show. There are also technical slips—duplicate equation numbers, a missing square in the elliptical distance, and Table 4 does not cleanly match the text's claim that warnings are earlier than human reaction in all scenarios (the rear-approaching row shows REACT at 2.7 s vs the driver at 0.8 s, later, which the text then explains as a driver false alarm). These are fixable, but they are exactly the kind of thing a referee should catch.\n\nSo my sense: the architecture is worth taking seriously, but the paper as written overclaims. If they calibrate thresholds on held-out data, report proper confidence intervals, add quantitative baseline comparisons, and rewrite Theorem 1 or drop it, I'd be happy to see it in the literature. As is, it should not be published with those claims intact. That said, it deserves a real referee—this is not a desk-reject candidate; it's a revise-and-resubmit at a good workshop or a major-revision at a solid journal. I would not cite it in my own work until the empirical claims are cleaned up. Reading group? Maybe, if you want to discuss what a real safety evaluation requires. Recommendation: send it out for review, with an expectation of substantial revision.","headline":"Incremental risk-field architecture with real hardware trials, but the zero-fault safety claims are statistically unsupported and the evaluation needs a proper calibration split.","tokens_in":17975,"tokens_out":2670,"would_cite":false,"duration_ms":24802,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a kinetic-energy risk field plus adaptive thresholds lets an autonomous vehicle avoid all collisions in four dynamic traffic scenarios, with zero false alarms and zero missed detections.","keywords":["autonomous driving","runtime risk assessment","collision avoidance","risk field modeling","active avoidance","hierarchical warning strategy","kinetic energy","real-time safety"],"falsifier":"Re-run the four scenarios with the published constants fixed ($T_1=0.3$, $T_2=0.7$, $a=0.2\\|v_j\\|$, $b=5$ m, $\\beta$, $\\lambda_j$, $k_{\\text{lane}}$) on a new set of trials that include a two-ton truck approaching from the rear at 35 km/h and a cyclist cutting in at 15 km/h; if any trial produces a warning while the gap is still large enough for normal driving (false alarm) or fails to warn before a collision path becomes unavoidable (miss), the claimed 0% / 0% performance is falsified.","tokens_in":16910,"feed_emoji":"🚗","tokens_out":9740,"duration_ms":84003,"temperature":0.7,"pith_summary":"The paper proposes REACT, a closed-loop collision-avoidance system that builds a continuous risk field around the ego vehicle from the kinetic energies, relative velocities, and positions of nearby traffic, plus a road-boundary penalty. The field is sampled on a small grid to yield a single normalized risk value and a dominant risk direction, which a pair of adaptive thresholds converts into three warning levels. On four real-road high-risk scenarios (car-following braking, cut-in, rear-approaching, intersection conflict), the authors report warnings issued earlier than the driver's natural response (lead time under 0.4 s), 100% safe avoidance with zero false alarms and zero misses, and under 50 ms latency on embedded hardware. If these results hold, a light, interpretable field model could handle real-time proactive safety without heavyweight trajectory prediction.","feed_headline":"REACT avoids every crash in four traffic scenarios","feed_subtitle":"A runtime risk field and adaptive thresholds warn in under 0.4 s, with zero false alarms.","key_machinery":"The carrying object is the interaction risk field $U_{ij}(x_i,y_i) = \\tfrac12 \\lambda_j m_j \\|v_j\\|^2 \\left(1 + \\beta \\cos(\\theta_{ij}) \\frac{\\|v_{ij}\\|}{\\|v_j\\|+\\epsilon}\\right)\\exp(-\\tilde{r}_{ij})$, where the elliptical distance $\\tilde{r}_{ij}$ concentrates risk along the threat's velocity direction (because $a=0.2\\|v_j\\|$ grows with the threat's speed) and $\\cos(\\theta_{ij})$ amplifies oncoming threats while attenuating receding ones. The field is superimposed with the road-constrained potential $U_E^a$ (a symmetric dual-spring lane-boundary penalty) and sampled on an $8$-direction, $m\\times n$ grid inside a reachable region around the ego vehicle. The normalized mean of all grid cells gives the global runtime risk, while per-sector means define the dominant danger direction $d^* = \\arg\\max_d \\bar{\\mathcal{R}}_d$, and the dynamic thresholds $T_1', T_2'$ map the global risk onto three warning levels. This machine reduces a dense multi-agent interaction to a handful of control-relevant scalars that an embedded controller can evaluate in under 50 ms.","core_discovery":"The central claim is that runtime driving risk can be captured by an anisotropic scalar field whose source strength is the kinetic energy $\\tfrac12 \\lambda_j m_j \\|v_j\\|^2$ of each surrounding participant, stretched along the motion direction through an elliptical distance $\\tilde{r}_{ij} = (x_i-x_j)^2/a + (y_i-y_j)^2/b$ with $a = 0.2\\|v_j\\|$ and $b = 5$ m, and modulated by the relative-velocity direction through $\\cos(\\theta_{ij})$. Superposing these fields with a road-boundary spring potential and integrating over an $m\\times n$ grid centered on the ego vehicle yields a normalized global risk $\\bar{\\mathcal{R}}_t$ and per-sector directional risks. Two adaptive thresholds, $T_1' = 0.3(1+\\Delta v/30)$ and $T_2' = 0.7(1-S_{\\text{brake}})$, turn that scalar into Level 0/1/2 warnings with a named dominant risk direction. The authors argue that this construction captures front, rear, and lateral multi-source interactions that TTC, THW, and RSS miss, and they support it with highD highway comparisons against four baselines and with on-vehicle trials in four scenarios, reporting zero false alarms, zero misses, and warning times within 0.4 s of human perception.","pith_inferences":["Because the field strength scales with the threat vehicle's kinetic energy, the same parameters would rate heavy trucks as systematically riskier than cars at equal speed and distance; a natural test is whether this matches human risk perception in mixed traffic.","The eight-sector grid and scalar risk value could serve as a common interface between perception and motion planning; the paper leaves trajectory generation to the driver or controller, so a promising extension is to replace the advisory action with gradient-descent steering along $F_a = -\\nabla U_a$.","The 0.4 s alignment with human cognition is measured on a small number of trials; a broader study on naturalistic near-crash recordings would reveal whether the threshold values or decay constants need to be re-tuned per site or weather condition.","If the thresholds are truly parameter-free, the same field should transfer to intersections of different geometries; one could check the authors' assertion of automatic adaptation by feeding the same constants into a simulator with an orthogonal crossing at varied approach angles."],"forward_implications":["If the claimed generalization holds, REACT can be deployed as a drop-in warning and advisory module on existing drive-by-wire vehicles without high-compute prediction models.","The directional risk output provides semantic warnings ('vehicle approaching from rear-right') that can be piped directly to human-machine interfaces or downstream planners.","The dynamic threshold logic implies earlier warnings when the ego is slower than surrounding traffic and earlier emergency escalation when the driver is already braking, a safety-prioritizing behavior.","The highD comparisons suggest the field model degrades gracefully where longitudinal-only metrics and RSS are blind to cut-ins, offering a path toward unified longitudinal-lateral risk assessment.","If the zero-miss, zero-false-alarm performance holds across more scenarios, the framework could reduce reliance on Monte-Carlo trajectory sampling for real-time safety."],"supporting_citations":[{"why":"Supplies the electrostatic/energy-transfer field analogy and behavioral risk model that the REACT field formula extends.","marker":"Zheng et al., 2021"},{"why":"Supplies the driving safety field theory and the road-boundary spring potential used in the total field.","marker":"Wang et al., 2016"},{"why":"Supplies the Driver Risk Field baseline whose delayed rear-threat response REACT claims to improve.","marker":"Kolekar et al., 2020"},{"why":"Defines the RSS baseline model used in the comparative table and highD experiments.","marker":"Hasuo, 2022"},{"why":"Motivates the intention-aware directional modulation of the risk field in dynamic interactions.","marker":"Huang et al., 2020"}],"fun_headline_variants":["REACT runtime risk field: zero false alarms, zero misses","Energy-based risk field powers REACT's 100% avoidance","REACT: runtime risk field, adaptive thresholds, zero misses","REACT's energy field warns in 0.4s, avoids all crashes","Runtime risk field with adaptive thresholds: REACT beats baselines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire zero-false-alarm, zero-miss result rests on the assumption that the hand-picked risk thresholds ($T_1 = 0.3$, $T_2 = 0.7$) and the field coefficients ($\\lambda_j$, $\\beta$, $k_{\\text{lane}}$, $a = 0.2\\|v_j\\|$, $b = 5$ m) are valid across all four test scenarios without being tuned on the same trials that produced the reported 100% success.","fun_headline_variants_meta":{"raw":{"variants":["REACT runtime risk field: zero false alarms, zero misses","Energy-based risk field powers REACT's 100% avoidance","REACT: runtime risk field, adaptive thresholds, zero misses","REACT's energy field warns in 0.4s, avoids all crashes","Runtime risk field with adaptive thresholds: REACT beats baselines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001275,"raw_usage":{"total_tokens":5267,"prompt_tokens":1053,"completion_tokens":4214,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":4124}},"tokens_in":669,"tokens_out":4214,"duration_ms":29462,"temperature":1.0,"reasoning_tokens":4124,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:52:41.530050+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the four scenarios with the published constants fixed ($T_1=0.3$, $T_2=0.7$, $a=0.2\\|v_j\\|$, $b=5$ m, $\\beta$, $\\lambda_j$, $k_{\\text{lane}}$) on a new set of trials that include a two-ton truck approaching from the rear at 35 km/h and a cyclist cutting in at 15 km/h; if any trial produces a warning while the gap is still large enough for normal driving (false alarm) or fails to warn before a collision path becomes unavoidable (miss), the claimed 0% / 0% performance is falsified.","supporting_citations":[],"review_version":1}