REVIEW 3 major objections 6 minor 26 references
Spatial Association Between Near-Misses and Accident Blackspots in Sydney, Australia: A Getis-Ord $G_i^*$ Analysis
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Near-miss telematics data exposes 2,681 high-risk Sydney cells that crash records miss.
desk verdict Standard spatial stats on a new dataset yields an interesting Sydney mismatch, but two load-bearing errors make the results non-reproducible as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a two-stage local spatial statistics pipeline on a uniform 400m grid with Queen-contiguity spatial weights. First, the Getis-Ord Gi* statistic is applied to crash counts alone to identify single-variable hotspots and coldspots. Then Bivariate Local Moran's I is applied to the paired cell counts (crashes, High-G near-misses) to classify each cell into four profiles—HH, HL, LH, LL—depending on whether the local association is concordant high, discordant in either direction, or concordant low. High-G events are defined as individual peak G-force readings above 0.47g, a threshold the paper takes from automated emergency braking and harsh-cornering systems. The POI characterization step adds Mann-Whitney U tests to see which land-use features differ between profile classes.
What would settle it
Take a sample of the raw accelerometer traces for a straight, uncongested motorway in the dataset and compute the paper's G-force formula; if the median value is above 0.47g, the vertical axis was not gravity-compensated and the near-miss counts used in every analysis are an artifact of the units.
Extended reading notes
Core claim
The paper's claim, on its own terms, is that in the Sydney study area during 2022, high-G near-miss events and reported crashes are only partially aligned spatially. Bivariate LISA at p<0.05 classifies 825 grid cells as concordant high-high, 91 as high-crash/low-near-miss, and 2,681 as low-crash/high-near-miss, with zero significant low-low cells. Because the low-crash/high-near-miss class is more than three times larger than the concordant hotspot class, the discovery is that near-miss severity data carries information about risk that historical crash data does not encode. These LH areas are places where risky maneuvers are frequent but crashes have not yet accumulated, so they are the paper's candidate 'pre-blackspots' for proactive safety management.
Load-bearing premise
The whole analysis depends on the assumption that the vertical acceleration measurement in the G-force formula has already had gravity removed; if it has not, the computed G-force would be at least 1g in ordinary driving and the 0.47g threshold would label nearly every trip as a near-miss.
Editorial extensions
If this is right
- Crash-only blackspot programs miss the largest risk class: the 2,681 low-crash/high-near-miss cells would not appear in a hotspot map built from historical crashes alone.
- The 825 HH cells are the priority for dual interventions, since both realized crashes and risky maneuvers concentrate there.
- The LH cells are the natural target for monitoring and low-cost proactive treatments; if near-misses are leading indicators, these are where crash risk builds before it shows up in records.
- The absence of significant LL cells implies that the absence of crashes is not a reliable safety signal, because near-miss activity can still be high.
- Transport agencies can apply the same grid-and-LISA workflow to any city with connected-vehicle telematics to produce a leading-indicator risk map without waiting for crashes to accumulate.
Reading between the lines
- If the vertical acceleration channel in the G-force calculation is not gravity-compensated, the 0.47g threshold would be meaningless; a quick sensitivity check on smooth motorway segments could settle this before the framework is used operationally.
- The one-year cross-section cannot by itself prove temporal precedence; tracking whether LH cells later develop crash hotspots would test the 'leading indicator' interpretation directly.
- Near-miss counts depend on fleet penetration: cells with more connected-vehicle equipped trips will mechanically record more events, so normalizing by vehicle-kilometres travelled or exposure per cell could shrink or redistribute the LH class.
- The LH pattern might also indicate places where good road design successfully absorbs frequent risky maneuvers, rather than latent danger; an outcome-based follow-up checking for severe crashes or complaints would separate risk from resilience.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a spatial-statistical workflow for comparing crash hotspots with high-G near-miss events in Sydney, Australia. It applies Getis-Ord Gi* to both reported crashes and high-G event counts on a 400m grid, then uses bivariate Local Moran's I (LISA) to classify cells into concordant and discordant risk profiles (HH, HL, LH, LL). The central empirical claim is that the large number of Low-Crash/High-High-G (LH) cells (2,681) relative to High-High (HH) cells (825) indicates that telematics-derived high-G data reveal additional high-risk areas not captured by crash-only hotspot analysis. The paper then characterizes the HH and LL areas using Mann-Whitney U tests on point-of-interest (POI) counts.
Significance. The motivating question is practically important: if severity-filtered near-miss data are spatially informative beyond historical crash records, they could support proactive road-safety management. The paper is among the first to apply bivariate LISA to paired crash and telematics near-miss data in an Australian city, and the methodological workflow is clearly described. However, the central quantitative results rest on a physically undefined G-force threshold and are undermined by an internal contradiction between the reported LISA cell counts and the POI comparison. If the G-force definition is corrected and the inconsistency resolved, the paper would provide a useful case study, but as submitted the main numerical claims are not reliable.
major comments (3)
- [IV, Eq. (1)] Equation (1) defines G-force(t) as the Euclidean norm of the three-axis acceleration vector divided by g, where a(t) is said to be 'measured by inertial sensors'. If a_z(t) includes the static gravitational component, as is standard for a raw MEMS accelerometer in the vehicle frame, then ||a(t)||/g is at least 1g whenever the vehicle is on level ground, and the 0.47g threshold would classify essentially all driving as High-G, not 24,137 rare near-miss events. The manuscript does not state that a_z is gravity-compensated or that the norm is computed on detrended or high-passed data. The cited Bendix and Geotab thresholds are applied to dedicated braking or cornering channels, not to the full three-axis Euclidean norm. The counts feeding Table I and the LH-versus-HH argument cannot be reproduced from the text as written.
- [V, Tables I and II] Table I reports zero Low-Low (LL) cells (p<0.05), yet Section V and Table II present Mann-Whitney U tests comparing 'Mean Outliers (LL)' with 'Mean Clusters (HH)' and report significant differences (e.g., traffic signals, p=0.017). With n=0 in the LL group, no U statistic, mean, or p-value can be computed; the two tables are mutually inconsistent. This contradiction invalidates the POI-based characterization of HH versus LL areas, including the conclusion that HH areas have a lower prevalence of traffic signals.
- [IV, V] The manuscript states in Section IV that the 400 m grid size was chosen after a sensitivity analysis, but no such analysis is reported anywhere, and no sensitivity analysis is reported for the 0.47g threshold. The central empirical result is a single numerical comparison of cell counts (2681 LH vs 825 HH) derived from these choices. Without evidence that the qualitative pattern of concordance/discordance is robust to reasonable variations in grid size and threshold, the claim that High-G data identify 'additional' high-risk areas is not established.
minor comments (6)
- [Throughout] Terminology is inconsistent: 'near-miss', 'nearmiss', 'NM+G', and 'High-G' are used interchangeably; please define each term at first use and use one convention throughout.
- [IV] There is a typo in the sentence 'The trajectory of both the the steering (lateral deviations) and braking (longitudinal deceleration) maneuvers...'.
- [I] The phrase 'For This study we use data provided by Compass IoT' should be 'For this study'.
- [Fig. 1] The y-axis label 'acceleration ing' is ambiguous; clarify whether the units are dimensionless g-force or m/s².
- [V] The permutation tests on 38,824 grid cells use a nominal p<0.05 threshold without multiple-testing correction; please report FDR-adjusted q-values or justify the nominal level.
- [Conclusion] The text contains 'p¡0.05' (encoding error); it should read 'p<0.05'.
Circularity Check
No significant circularity: the spatial-statistical analysis is empirical, uses an externally sourced threshold, and the central conclusion is a descriptive interpretation of the LISA output rather than a prediction derived from fitted inputs.
full rationale
The paper's derivation chain is: crash counts and High-G event counts are aggregated onto a 400 m grid; Getis-Ord Gi* and bivariate Local Moran's I are applied with permutation-based significance testing; cells are classified as HH, HL, LH, or LL. The 0.47g threshold defining a High-G event is adopted from cited external sources (Bendix automatic-emergency-braking literature and Geotab harsh-driving documentation), not fitted to maximize the number of LH cells or to force the paper's conclusion. No parameter is estimated from the target data to produce the LH finding; the 2,681 LH cells are the direct empirical output of a standard LISA procedure on the two count variables. The statement that High-G data 'identifies additional locations requiring monitoring' is an interpretive gloss on the LH category, not a formal derivation that reduces to its own definition: the paper does not claim that these LH cells were predicted from a model fitted to crash data. There is no load-bearing self-citation: the surrogate-safety literature cited (e.g., Guo et al.) is external, and the authors' connection to Compass IoT appears as data provenance and an acknowledgement, not as an analytical premise. The potential concern that Eq. (1) may not account for the gravitational component of az is a data-validity or correctness issue, not a circularity. Accordingly, the paper is not circular; its main limitations are interpretive and data-quality related, not logical.
Assumptions & free parameters
free parameters (2)
- 0.47g threshold =
0.47 g
- Grid cell size =
400 m
assumptions (3)
- ad hoc to paper The inertial sensor measurements in Eq. (1) are gravity-compensated, so the Euclidean norm can fall below 1g and the 0.47g threshold selects braking/cornering events.
- domain assumption The 'causal continuum' hypothesis: near-misses and crashes share causal factors, so High-G event frequency is a proxy for crash risk.
- domain assumption Compass IoT connected-vehicle coverage is spatially representative of Sydney traffic exposure; otherwise the High-G count distribution is biased by sampling.
Cite this review
Pith. "Pith review of Spatial Association Between Near-Misses and Accident Blackspots in Sydney, Australia: A Getis-Ord $G_i^*$ Analysis." pith.science (2026). https://pith.science/paper/B22QKQPD
@misc{pith2026250603356,
author = {Pith},
title = {Pith review of: Spatial Association Between Near-Misses and Accident Blackspots in Sydney, Australia: A Getis-Ord $G_i^*$ Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/B22QKQPD}},
note = {Machine review of arXiv:2506.03356}
}
read the original abstract
Conventional road safety management is inherently reactive, relying on analysis of sparse and lagged historical crash data to identify hazardous locations, or crash blackspots. The proliferation of vehicle telematics presents an opportunity for a paradigm shift towards proactive safety, using high-frequency, high-resolution near-miss data as a leading indicator of crash risk. This paper presents a spatial-statistical framework to systematically analyze the concordance and discordance between official crash records and near-miss events within urban environment. A Getis-Ord statistic is first applied to both reported crashes and near-miss events to identify statistically significant local clusters of each type. Subsequently, Bivariate Local Moran's I assesses spatial relationships between crash counts and High-G event counts, classifying grid cells into distinct profiles: High-High (coincident risk), High-Low and Low-High. Our analysis reveals significant amount of Low-Crash, High-Near-Miss clusters representing high-risk areas that remain unobservable when relying solely on historical crash data. Feature importance analysis is performed using contextual Point of Interest data to identify the different infrastructure factors that characterize difference between spatial clusters. The results provide a data-driven methodology for transport authorities to transition from a reactive to a proactive safety management strategy, allowing targeted interventions before severe crashes occur.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Mehdizadeh, M. Cai, Q. Hu, M. A. Alamdar Yazdi, N. Mohabbati- Kalejahi, A. Vinel, S. E. Rigdon, K. C. Davis, and F. M. Megahed, “A review of data analytic applications in road traffic safety. part 1: Descriptive and predictive modeling,”Sensors, vol. 20, no. 4, 2020
work page 2020
-
[2]
S. Hallmark, G. Basulto-Elias, N. Oneyear, O. Smadi, S. Chrysler, and G. Ullman, “Evaluation of work zone safety using the SHRP2 natural- istic driving study data – volume 2 description of research,” Minnesota Department of Transportation, Office of Research & Innovation and Iowa State University, Institute for Transportation, St. Paul, MN and Ames, IA, F...
work page 2022
-
[4]
A review of surrogate safety measures uses in historical crash investigations,
D. Nikolaou, A. Ziakopoulos, and G. Yannis, “A review of surrogate safety measures uses in historical crash investigations,”Sustainability, vol. 15, p. 7580, 05 2023
work page 2023
-
[5]
Comparison of traffic conflict indicators for crash estimation using peak over threshold approach,
L. Zheng and T. Sayed, “Comparison of traffic conflict indicators for crash estimation using peak over threshold approach,”Transportation Research Record: Journal of the Transportation Research Board, vol. 2673, p. 036119811984155, 04 2019
work page 2019
-
[6]
M. Balawi and G. Tenekeci, “Time series traffic collision analysis of london hotspots: Patterns, predictions and prevention strategies,” Heliyon, vol. 10, no. 4, p. e25710, 2024
work page 2024
-
[7]
S.-H. Park and M.-K. Bae, “Effects influencing pedestrian–vehicle crash frequency by severity level: A case study of seoul metropolitan city, south korea,”Safety, vol. 6, no. 2, 2020. [Online]. Available: https://www.mdpi.com/2313-576X/6/2/25
work page 2020
-
[8]
T. Alsahfi, “Spatial and temporal analysis of road traffic accidents in major californian cities using a geographic information system,” ISPRS International Journal of Geo-Information, vol. 13, no. 5, 2024. [Online]. Available: https://www.mdpi.com/2220-9964/13/5/157
work page 2024
-
[9]
Comparative analysis of the spatial analysis methods for hotspot identification,
H. Yu, P. Liu, J. Chen, and H. Wang, “Comparative analysis of the spatial analysis methods for hotspot identification,”Accident Analysis & Prevention, vol. 66, pp. 80–88, 2014
work page 2014
Show all 26 references
-
[10]
Local indicators of spatial association—lisa,
L. Anselin, “Local indicators of spatial association—lisa,”Geographical Analysis, vol. 27, no. 2, pp. 93–115, 1995
1995
-
[11]
Gis-based spatial analysis of accident hotspots: A nigerian case study,
A. Afolayan, S. M. Easa, O. S. Abiola, F. M. Alayaki, and O. Folorunso, “Gis-based spatial analysis of accident hotspots: A nigerian case study,”Infrastructures, vol. 7, no. 8, 2022. [Online]. Available: https://www.mdpi.com/2412-3811/7/8/103
2022
-
[12]
Surrogate safety assessment in heterogeneous traffic environment prevailing in developing countries: a systematic literature review,
A. Kumar and A. M. and, “Surrogate safety assessment in heterogeneous traffic environment prevailing in developing countries: a systematic literature review,”International Journal of Injury Control and Safety Promotion, vol. 0, no. 0, pp. 1–19, 2025, pMID: 40279179. [Online]. ...
2025
-
[13]
Investigating surrogate safety measures’ threshold consistency on different types of curves,
K. Sukhanya, R. N. Shilpa, and B. K. B. and, “Investigating surrogate safety measures’ threshold consistency on different types of curves,”Journal of Transportation Safety & Security, vol. 0, no. 0, pp. 1–13, 2025. [Online]. Available: https://doi.org/10.1080/19439962.2024.2447985
2025
-
[14]
Modeling traffic conflicts for use in road safety analysis: A review of analytic methods and future directions,
L. Zheng, T. Sayed, and F. Mannering, “Modeling traffic conflicts for use in road safety analysis: A review of analytic methods and future directions,”Analytic Methods in Accident Research, vol. 29, p. 100142, 2021
2021
-
[15]
Conflict-based safety evaluations at unsignalized intersections using surrogate safety measures,
D. Singh, P. Das, and I. Ghosh, “Conflict-based safety evaluations at unsignalized intersections using surrogate safety measures,”Heliyon, vol. 10, no. 5, p. e27665, 2024
2024
-
[16]
When intelligent transportation systems sensing meets edge computing: Vision and challenges,
X. Zhou, R. Ke, H. Yang, and C. Liu, “When intelligent transportation systems sensing meets edge computing: Vision and challenges,” Applied Sciences, vol. 11, no. 20, 2021. [Online]. Available: https://www.mdpi.com/2076-3417/11/20/9680
2021
-
[17]
Using video analytics to improve traffic intersection safety and performance,
A. Mishra, K. Chen, S. Poddar, E. Posadas, A. Rangarajan, and S. Ranka, “Using video analytics to improve traffic intersection safety and performance,”Vehicles, vol. 4, no. 4, pp. 1288–1313, 2022. [Online]. Available: https://www.mdpi.com/2624-8921/4/4/68
2022
-
[18]
(2018) Strategic highway research program (SHRP2)
Federal Highway Administration (FHW A). (2018) Strategic highway research program (SHRP2). U.S. Department of Transportation, Federal Highway Administration
2018
-
[19]
Supporting large scale connected vehicle data analysis using hive,
W. Xu, N. Ru ´ız-Juri, A. Gupta, A. Deering, C. Bhat, J. Kuhr, and J. Archer, “Supporting large scale connected vehicle data analysis using hive,” 12 2016, pp. 2296–2304
2016
-
[20]
Causation analysis of crashes and near crashes using naturalistic driving data,
X. Wang, Q. Liu, F. Guo, S. Fang, X. Xu, and X. Chen, “Causation analysis of crashes and near crashes using naturalistic driving data,” Accident Analysis & Prevention, vol. 177, p. 106821, 2022
2022
-
[21]
Traffic conflict techniques for road safety analysis: Open questions and some insights,
L. Zheng, K. Ismail, and X. Meng, “Traffic conflict techniques for road safety analysis: Open questions and some insights,”Canadian Journal of Civil Engineering, vol. 41, 07 2014
2014
-
[22]
Near crashes as crash surrogate for naturalistic driving studies,
F. Guo, S. G. Klauer, J. M. Hankey, and T. A. Dingus, “Near crashes as crash surrogate for naturalistic driving studies,”Transportation Research Record, vol. 2147, no. 1, pp. 66–74, 2010
2010
-
[23]
Freeway safety estimation using extreme value theory approaches: A comparative study,
L. Zheng, K. Ismail, and X. Meng, “Freeway safety estimation using extreme value theory approaches: A comparative study,”Accident; analysis and prevention, vol. 62C, pp. 32–41, 09 2013
2013
-
[24]
Analysis of trajectory transection rate on four-lane divided rural highway curves,
V . K. Sharma and G. S. and, “Analysis of trajectory transection rate on four-lane divided rural highway curves,”Traffic Injury Prevention, vol. 0, no. 0, pp. 1–9, 2025, pMID: 39879567. [Online]. Available: https://doi.org/10.1080/15389588.2025.2450710
2025
-
[25]
Analysis of stopping sight distance (ssd) parameters: A review study,
C. J. Samson, Q. Hussain, and W. K. Alhajyaseen, “Analysis of stopping sight distance (ssd) parameters: A review study,” Procedia Computer Science, vol. 201, pp. 126–133, 2022, the 13th International Conference on Ambient Systems, Networks and Technologies (ANT) / The 5th Inte...
2022
-
[26]
A target population for automatic emergency braking in heavy vehicles,
D. Glassbrenner, A. Morgan, R. Kreeb, A. Svenson, H. Liddell, and F. Barickman, “A target population for automatic emergency braking in heavy vehicles,” National Highway Traffic Safety Administration, Washington, DC, Tech. Rep. DOT HS 812 390, July 2017, mathematical Analysis ...
2017
-
[27]
What is g-force and how is it related to harsh driving?
M. Broughall, “What is g-force and how is it related to harsh driving?” Geotab Blog. [Online]. Available: https://www.geotab.com/blog/what- is-g-force/
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.