Pith. sign in

REVIEW 3 major objections 6 minor 26 references

Spatial Association Between Near-Misses and Accident Blackspots in Sydney, Australia: A Getis-Ord $G_i^*$ Analysis

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Near-miss telematics data exposes 2,681 high-risk Sydney cells that crash records miss.

desk verdict Standard spatial stats on a new dataset yields an interesting Sydney mismatch, but two load-bearing errors make the results non-reproducible as written. read the letter →

arxiv 2506.03356 v2 pith:B22QKQPD submitted 2025-06-03 eess.SY cs.CYcs.SYstat.AP

classification eess.SYcs.CYcs.SYstat.AP
keywords connectedvehiclesnear-misseventscrashblackspotsGetis-OrdGi*BivariateLocalMoran'sIspatialclusteringproactiveroadsafetytelematics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to decide whether high-severity near-miss events recorded by vehicle telematics are located where crashes historically happen, or whether they reveal risk that crash records hide. On a 400-metre grid covering Sydney in 2022, it applies Getis-Ord Gi* to crash counts and then Bivariate Local Moran's I to crash counts paired with high-G near-miss counts. The central result is an asymmetry: 825 grid cells are high-crash/high-near-miss, but 2,681 cells are low-crash/high-near-miss, and no cells form significant low-low clusters. The paper reads this as evidence that near-miss severity data identifies many additional high-risk locations beyond traditional crash blackspots, and that the absence of crashes is not the same as the absence of risk. That is why the result matters: it offers transport authorities a data-driven way to shift from reactive crash counting toward proactive monitoring.

What carries the argument

The machinery is a two-stage local spatial statistics pipeline on a uniform 400m grid with Queen-contiguity spatial weights. First, the Getis-Ord Gi* statistic is applied to crash counts alone to identify single-variable hotspots and coldspots. Then Bivariate Local Moran's I is applied to the paired cell counts (crashes, High-G near-misses) to classify each cell into four profiles—HH, HL, LH, LL—depending on whether the local association is concordant high, discordant in either direction, or concordant low. High-G events are defined as individual peak G-force readings above 0.47g, a threshold the paper takes from automated emergency braking and harsh-cornering systems. The POI characterization step adds Mann-Whitney U tests to see which land-use features differ between profile classes.

What would settle it

Take a sample of the raw accelerometer traces for a straight, uncongested motorway in the dataset and compute the paper's G-force formula; if the median value is above 0.47g, the vertical axis was not gravity-compensated and the near-miss counts used in every analysis are an artifact of the units.

Watch

Extended reading notes

Core claim

The paper's claim, on its own terms, is that in the Sydney study area during 2022, high-G near-miss events and reported crashes are only partially aligned spatially. Bivariate LISA at p<0.05 classifies 825 grid cells as concordant high-high, 91 as high-crash/low-near-miss, and 2,681 as low-crash/high-near-miss, with zero significant low-low cells. Because the low-crash/high-near-miss class is more than three times larger than the concordant hotspot class, the discovery is that near-miss severity data carries information about risk that historical crash data does not encode. These LH areas are places where risky maneuvers are frequent but crashes have not yet accumulated, so they are the paper's candidate 'pre-blackspots' for proactive safety management.

Load-bearing premise

The whole analysis depends on the assumption that the vertical acceleration measurement in the G-force formula has already had gravity removed; if it has not, the computed G-force would be at least 1g in ordinary driving and the 0.47g threshold would label nearly every trip as a near-miss.

Editorial extensions

If this is right

  • Crash-only blackspot programs miss the largest risk class: the 2,681 low-crash/high-near-miss cells would not appear in a hotspot map built from historical crashes alone.
  • The 825 HH cells are the priority for dual interventions, since both realized crashes and risky maneuvers concentrate there.
  • The LH cells are the natural target for monitoring and low-cost proactive treatments; if near-misses are leading indicators, these are where crash risk builds before it shows up in records.
  • The absence of significant LL cells implies that the absence of crashes is not a reliable safety signal, because near-miss activity can still be high.
  • Transport agencies can apply the same grid-and-LISA workflow to any city with connected-vehicle telematics to produce a leading-indicator risk map without waiting for crashes to accumulate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the vertical acceleration channel in the G-force calculation is not gravity-compensated, the 0.47g threshold would be meaningless; a quick sensitivity check on smooth motorway segments could settle this before the framework is used operationally.
  • The one-year cross-section cannot by itself prove temporal precedence; tracking whether LH cells later develop crash hotspots would test the 'leading indicator' interpretation directly.
  • Near-miss counts depend on fleet penetration: cells with more connected-vehicle equipped trips will mechanically record more events, so normalizing by vehicle-kilometres travelled or exposure per cell could shrink or redistribute the LH class.
  • The LH pattern might also indicate places where good road design successfully absorbs frequent risky maneuvers, rather than latent danger; an outcome-based follow-up checking for severe crashes or complaints would separate risk from resilience.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents a spatial-statistical workflow for comparing crash hotspots with high-G near-miss events in Sydney, Australia. It applies Getis-Ord Gi* to both reported crashes and high-G event counts on a 400m grid, then uses bivariate Local Moran's I (LISA) to classify cells into concordant and discordant risk profiles (HH, HL, LH, LL). The central empirical claim is that the large number of Low-Crash/High-High-G (LH) cells (2,681) relative to High-High (HH) cells (825) indicates that telematics-derived high-G data reveal additional high-risk areas not captured by crash-only hotspot analysis. The paper then characterizes the HH and LL areas using Mann-Whitney U tests on point-of-interest (POI) counts.

Significance. The motivating question is practically important: if severity-filtered near-miss data are spatially informative beyond historical crash records, they could support proactive road-safety management. The paper is among the first to apply bivariate LISA to paired crash and telematics near-miss data in an Australian city, and the methodological workflow is clearly described. However, the central quantitative results rest on a physically undefined G-force threshold and are undermined by an internal contradiction between the reported LISA cell counts and the POI comparison. If the G-force definition is corrected and the inconsistency resolved, the paper would provide a useful case study, but as submitted the main numerical claims are not reliable.

major comments (3)
  1. [IV, Eq. (1)] Equation (1) defines G-force(t) as the Euclidean norm of the three-axis acceleration vector divided by g, where a(t) is said to be 'measured by inertial sensors'. If a_z(t) includes the static gravitational component, as is standard for a raw MEMS accelerometer in the vehicle frame, then ||a(t)||/g is at least 1g whenever the vehicle is on level ground, and the 0.47g threshold would classify essentially all driving as High-G, not 24,137 rare near-miss events. The manuscript does not state that a_z is gravity-compensated or that the norm is computed on detrended or high-passed data. The cited Bendix and Geotab thresholds are applied to dedicated braking or cornering channels, not to the full three-axis Euclidean norm. The counts feeding Table I and the LH-versus-HH argument cannot be reproduced from the text as written.
  2. [V, Tables I and II] Table I reports zero Low-Low (LL) cells (p<0.05), yet Section V and Table II present Mann-Whitney U tests comparing 'Mean Outliers (LL)' with 'Mean Clusters (HH)' and report significant differences (e.g., traffic signals, p=0.017). With n=0 in the LL group, no U statistic, mean, or p-value can be computed; the two tables are mutually inconsistent. This contradiction invalidates the POI-based characterization of HH versus LL areas, including the conclusion that HH areas have a lower prevalence of traffic signals.
  3. [IV, V] The manuscript states in Section IV that the 400 m grid size was chosen after a sensitivity analysis, but no such analysis is reported anywhere, and no sensitivity analysis is reported for the 0.47g threshold. The central empirical result is a single numerical comparison of cell counts (2681 LH vs 825 HH) derived from these choices. Without evidence that the qualitative pattern of concordance/discordance is robust to reasonable variations in grid size and threshold, the claim that High-G data identify 'additional' high-risk areas is not established.
minor comments (6)
  1. [Throughout] Terminology is inconsistent: 'near-miss', 'nearmiss', 'NM+G', and 'High-G' are used interchangeably; please define each term at first use and use one convention throughout.
  2. [IV] There is a typo in the sentence 'The trajectory of both the the steering (lateral deviations) and braking (longitudinal deceleration) maneuvers...'.
  3. [I] The phrase 'For This study we use data provided by Compass IoT' should be 'For this study'.
  4. [Fig. 1] The y-axis label 'acceleration ing' is ambiguous; clarify whether the units are dimensionless g-force or m/s².
  5. [V] The permutation tests on 38,824 grid cells use a nominal p<0.05 threshold without multiple-testing correction; please report FDR-adjusted q-values or justify the nominal level.
  6. [Conclusion] The text contains 'p¡0.05' (encoding error); it should read 'p<0.05'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the spatial-statistical analysis is empirical, uses an externally sourced threshold, and the central conclusion is a descriptive interpretation of the LISA output rather than a prediction derived from fitted inputs.

full rationale

The paper's derivation chain is: crash counts and High-G event counts are aggregated onto a 400 m grid; Getis-Ord Gi* and bivariate Local Moran's I are applied with permutation-based significance testing; cells are classified as HH, HL, LH, or LL. The 0.47g threshold defining a High-G event is adopted from cited external sources (Bendix automatic-emergency-braking literature and Geotab harsh-driving documentation), not fitted to maximize the number of LH cells or to force the paper's conclusion. No parameter is estimated from the target data to produce the LH finding; the 2,681 LH cells are the direct empirical output of a standard LISA procedure on the two count variables. The statement that High-G data 'identifies additional locations requiring monitoring' is an interpretive gloss on the LH category, not a formal derivation that reduces to its own definition: the paper does not claim that these LH cells were predicted from a model fitted to crash data. There is no load-bearing self-citation: the surrogate-safety literature cited (e.g., Guo et al.) is external, and the authors' connection to Compass IoT appears as data provenance and an acknowledgement, not as an analytical premise. The potential concern that Eq. (1) may not account for the gravitational component of az is a data-validity or correctness issue, not a circularity. Accordingly, the paper is not circular; its main limitations are interpretive and data-quality related, not logical.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central analysis is empirical and relies on standard spatial statistics; no parameters are fitted to the target result. The most important ledger item is the implicit gravity-compensation assumption behind the 0.47g threshold, and the unstated coverage assumptions for the Compass IoT fleet. The grid size is a hand-chosen modeling parameter.

free parameters (2)
  • 0.47g threshold = 0.47 g
    Chosen from Bendix and Geotab literature to define High-G near-miss events; the central event definition depends on this hand-selected cutoff, and the paper does not test sensitivity to it.
  • Grid cell size = 400 m
    Selected after an unreported sensitivity analysis; the LISA and Gi* results depend on the aggregation scale.
assumptions (3)
  • ad hoc to paper The inertial sensor measurements in Eq. (1) are gravity-compensated, so the Euclidean norm can fall below 1g and the 0.47g threshold selects braking/cornering events.
    Required for the G-force definition to be consistent with the threshold; the paper never states this assumption.
  • domain assumption The 'causal continuum' hypothesis: near-misses and crashes share causal factors, so High-G event frequency is a proxy for crash risk.
    Invoked in Related Works as the motivation for using near-misses as a leading indicator; this is a substantive domain assumption not established by the data.
  • domain assumption Compass IoT connected-vehicle coverage is spatially representative of Sydney traffic exposure; otherwise the High-G count distribution is biased by sampling.
    The telematics fleet covers a small fraction of Sydney traffic and the paper does not normalize High-G counts by exposure; low coverage in some areas could create spurious LH cells.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spatial Association Between Near-Misses and Accident Blackspots in Sydney, Australia: A Getis-Ord $G_i^*$ Analysis." pith.science (2026). https://pith.science/paper/B22QKQPD

@misc{pith2026250603356,
  author       = {Pith},
  title        = {Pith review of: Spatial Association Between Near-Misses and Accident Blackspots in Sydney, Australia: A Getis-Ord $G_i^*$ Analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B22QKQPD}},
  note         = {Machine review of arXiv:2506.03356}
}
read the original abstract

Conventional road safety management is inherently reactive, relying on analysis of sparse and lagged historical crash data to identify hazardous locations, or crash blackspots. The proliferation of vehicle telematics presents an opportunity for a paradigm shift towards proactive safety, using high-frequency, high-resolution near-miss data as a leading indicator of crash risk. This paper presents a spatial-statistical framework to systematically analyze the concordance and discordance between official crash records and near-miss events within urban environment. A Getis-Ord statistic is first applied to both reported crashes and near-miss events to identify statistically significant local clusters of each type. Subsequently, Bivariate Local Moran's I assesses spatial relationships between crash counts and High-G event counts, classifying grid cells into distinct profiles: High-High (coincident risk), High-Low and Low-High. Our analysis reveals significant amount of Low-Crash, High-Near-Miss clusters representing high-risk areas that remain unobservable when relying solely on historical crash data. Feature importance analysis is performed using contextual Point of Interest data to identify the different infrastructure factors that characterize difference between spatial clusters. The results provide a data-driven methodology for transport authorities to transition from a reactive to a proactive safety management strategy, allowing targeted interventions before severe crashes occur.

Figures

Figures reproduced from arXiv: 2506.03356 by the authors.

Figure 1
Figure 1. Vehicle trajectory segments represented with G-force profiles (x-axis: [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Spatial distribution of statistically significant [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Map of LISA results on the 400m grid, showing the spatial correlation between crash counts and High-G event counts per cell. Colors represent the type of statistically significant (p < 0.05) local correlation: High-High (Red: high crashes, high High-G), Low-Low (Blue: low crashes, low High-G), High-Low (Pink: high crashes, low High-G), Low-High (Light Blue: low crashes, high High-G). Grey areas indicate no significa… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

26 extracted references · 25 canonical work pages

  1. [1]

    A review of data analytic applications in road traffic safety. part 1: Descriptive and predictive modeling,

    A. Mehdizadeh, M. Cai, Q. Hu, M. A. Alamdar Yazdi, N. Mohabbati- Kalejahi, A. Vinel, S. E. Rigdon, K. C. Davis, and F. M. Megahed, “A review of data analytic applications in road traffic safety. part 1: Descriptive and predictive modeling,”Sensors, vol. 20, no. 4, 2020

  2. [2]

    Evaluation of work zone safety using the SHRP2 natural- istic driving study data – volume 2 description of research,

    S. Hallmark, G. Basulto-Elias, N. Oneyear, O. Smadi, S. Chrysler, and G. Ullman, “Evaluation of work zone safety using the SHRP2 natural- istic driving study data – volume 2 description of research,” Minnesota Department of Transportation, Office of Research & Innovation and Iowa State University, Institute for Transportation, St. Paul, MN and Ames, IA, F...

  3. [4]

    A review of surrogate safety measures uses in historical crash investigations,

    D. Nikolaou, A. Ziakopoulos, and G. Yannis, “A review of surrogate safety measures uses in historical crash investigations,”Sustainability, vol. 15, p. 7580, 05 2023

  4. [5]

    Comparison of traffic conflict indicators for crash estimation using peak over threshold approach,

    L. Zheng and T. Sayed, “Comparison of traffic conflict indicators for crash estimation using peak over threshold approach,”Transportation Research Record: Journal of the Transportation Research Board, vol. 2673, p. 036119811984155, 04 2019

  5. [6]

    Time series traffic collision analysis of london hotspots: Patterns, predictions and prevention strategies,

    M. Balawi and G. Tenekeci, “Time series traffic collision analysis of london hotspots: Patterns, predictions and prevention strategies,” Heliyon, vol. 10, no. 4, p. e25710, 2024

  6. [7]

    Effects influencing pedestrian–vehicle crash frequency by severity level: A case study of seoul metropolitan city, south korea,

    S.-H. Park and M.-K. Bae, “Effects influencing pedestrian–vehicle crash frequency by severity level: A case study of seoul metropolitan city, south korea,”Safety, vol. 6, no. 2, 2020. [Online]. Available: https://www.mdpi.com/2313-576X/6/2/25

  7. [8]

    Spatial and temporal analysis of road traffic accidents in major californian cities using a geographic information system,

    T. Alsahfi, “Spatial and temporal analysis of road traffic accidents in major californian cities using a geographic information system,” ISPRS International Journal of Geo-Information, vol. 13, no. 5, 2024. [Online]. Available: https://www.mdpi.com/2220-9964/13/5/157

  8. [9]

    Comparative analysis of the spatial analysis methods for hotspot identification,

    H. Yu, P. Liu, J. Chen, and H. Wang, “Comparative analysis of the spatial analysis methods for hotspot identification,”Accident Analysis & Prevention, vol. 66, pp. 80–88, 2014

Show all 26 references
  1. [10]

    Local indicators of spatial association—lisa,

    L. Anselin, “Local indicators of spatial association—lisa,”Geographical Analysis, vol. 27, no. 2, pp. 93–115, 1995

  2. [11]

    Gis-based spatial analysis of accident hotspots: A nigerian case study,

    A. Afolayan, S. M. Easa, O. S. Abiola, F. M. Alayaki, and O. Folorunso, “Gis-based spatial analysis of accident hotspots: A nigerian case study,”Infrastructures, vol. 7, no. 8, 2022. [Online]. Available: https://www.mdpi.com/2412-3811/7/8/103

  3. [12]

    Surrogate safety assessment in heterogeneous traffic environment prevailing in developing countries: a systematic literature review,

    A. Kumar and A. M. and, “Surrogate safety assessment in heterogeneous traffic environment prevailing in developing countries: a systematic literature review,”International Journal of Injury Control and Safety Promotion, vol. 0, no. 0, pp. 1–19, 2025, pMID: 40279179. [Online]. ...

  4. [13]

    Investigating surrogate safety measures’ threshold consistency on different types of curves,

    K. Sukhanya, R. N. Shilpa, and B. K. B. and, “Investigating surrogate safety measures’ threshold consistency on different types of curves,”Journal of Transportation Safety & Security, vol. 0, no. 0, pp. 1–13, 2025. [Online]. Available: https://doi.org/10.1080/19439962.2024.2447985

  5. [14]

    Modeling traffic conflicts for use in road safety analysis: A review of analytic methods and future directions,

    L. Zheng, T. Sayed, and F. Mannering, “Modeling traffic conflicts for use in road safety analysis: A review of analytic methods and future directions,”Analytic Methods in Accident Research, vol. 29, p. 100142, 2021

  6. [15]

    Conflict-based safety evaluations at unsignalized intersections using surrogate safety measures,

    D. Singh, P. Das, and I. Ghosh, “Conflict-based safety evaluations at unsignalized intersections using surrogate safety measures,”Heliyon, vol. 10, no. 5, p. e27665, 2024

  7. [16]

    When intelligent transportation systems sensing meets edge computing: Vision and challenges,

    X. Zhou, R. Ke, H. Yang, and C. Liu, “When intelligent transportation systems sensing meets edge computing: Vision and challenges,” Applied Sciences, vol. 11, no. 20, 2021. [Online]. Available: https://www.mdpi.com/2076-3417/11/20/9680

  8. [17]

    Using video analytics to improve traffic intersection safety and performance,

    A. Mishra, K. Chen, S. Poddar, E. Posadas, A. Rangarajan, and S. Ranka, “Using video analytics to improve traffic intersection safety and performance,”Vehicles, vol. 4, no. 4, pp. 1288–1313, 2022. [Online]. Available: https://www.mdpi.com/2624-8921/4/4/68

  9. [18]

    (2018) Strategic highway research program (SHRP2)

    Federal Highway Administration (FHW A). (2018) Strategic highway research program (SHRP2). U.S. Department of Transportation, Federal Highway Administration

  10. [19]

    Supporting large scale connected vehicle data analysis using hive,

    W. Xu, N. Ru ´ız-Juri, A. Gupta, A. Deering, C. Bhat, J. Kuhr, and J. Archer, “Supporting large scale connected vehicle data analysis using hive,” 12 2016, pp. 2296–2304

  11. [20]

    Causation analysis of crashes and near crashes using naturalistic driving data,

    X. Wang, Q. Liu, F. Guo, S. Fang, X. Xu, and X. Chen, “Causation analysis of crashes and near crashes using naturalistic driving data,” Accident Analysis & Prevention, vol. 177, p. 106821, 2022

  12. [21]

    Traffic conflict techniques for road safety analysis: Open questions and some insights,

    L. Zheng, K. Ismail, and X. Meng, “Traffic conflict techniques for road safety analysis: Open questions and some insights,”Canadian Journal of Civil Engineering, vol. 41, 07 2014

  13. [22]

    Near crashes as crash surrogate for naturalistic driving studies,

    F. Guo, S. G. Klauer, J. M. Hankey, and T. A. Dingus, “Near crashes as crash surrogate for naturalistic driving studies,”Transportation Research Record, vol. 2147, no. 1, pp. 66–74, 2010

  14. [23]

    Freeway safety estimation using extreme value theory approaches: A comparative study,

    L. Zheng, K. Ismail, and X. Meng, “Freeway safety estimation using extreme value theory approaches: A comparative study,”Accident; analysis and prevention, vol. 62C, pp. 32–41, 09 2013

  15. [24]

    Analysis of trajectory transection rate on four-lane divided rural highway curves,

    V . K. Sharma and G. S. and, “Analysis of trajectory transection rate on four-lane divided rural highway curves,”Traffic Injury Prevention, vol. 0, no. 0, pp. 1–9, 2025, pMID: 39879567. [Online]. Available: https://doi.org/10.1080/15389588.2025.2450710

  16. [25]

    Analysis of stopping sight distance (ssd) parameters: A review study,

    C. J. Samson, Q. Hussain, and W. K. Alhajyaseen, “Analysis of stopping sight distance (ssd) parameters: A review study,” Procedia Computer Science, vol. 201, pp. 126–133, 2022, the 13th International Conference on Ambient Systems, Networks and Technologies (ANT) / The 5th Inte...

  17. [26]

    A target population for automatic emergency braking in heavy vehicles,

    D. Glassbrenner, A. Morgan, R. Kreeb, A. Svenson, H. Liddell, and F. Barickman, “A target population for automatic emergency braking in heavy vehicles,” National Highway Traffic Safety Administration, Washington, DC, Tech. Rep. DOT HS 812 390, July 2017, mathematical Analysis ...

  18. [27]

    What is g-force and how is it related to harsh driving?

    M. Broughall, “What is g-force and how is it related to harsh driving?” Geotab Blog. [Online]. Available: https://www.geotab.com/blog/what- is-g-force/

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.