{"id":"8b41ac77-2cc6-4bba-8bff-278bf677c5ed","arxiv_id":"2507.19173","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new way to compare ray tracing results by treating each propagation path as a point in a 6D space and using Hausdorff and Chamfer distances.","lead":"This paper proposes two metrics, Hausdorff ray tracing (HRT) and Chamfer ray tracing (CRT), to measure how much two ray tracing simulations differ when the 3D environment changes. The metrics are tested on a digital twin of Milan with parked vehicles or segmented building windows, highlighting where radio propagation changes most.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported dB/ns/degree values do not follow from Eq. (3): per-set standardization yields dimensionless distances, and no conversion is stated; numerical results in Section 4 are therefore unsupported as defined.","rationale":"The central claim of the paper is that HRT and CRT metrics consistently compare temporal, angular, and power features between two ray tracing simulations. For that claim to hold, the metric definitions must support the quantitative results used to validate them. The mismatch between Eq. (3) and Section 4 is direct and load-bearing: Eq. (3) operates on per-set standardized, unitless quantities, while Section 4 presents physical units (dB, nsec, degrees) with no stated conversion. This makes the reported average distances (e.g., ≈16 dB, ≈277 ns, ≈46 deg) non-derivable from the proposed definitions, so the numerical evidence for the metrics' usefulness is unsubstantiated. The per-set standardization also means uniform offsets in power or delay between scenarios are invisible to the metric, further undermining the claim of comparing power features in an absolute sense. This is not a stylistic issue: it affects the validity of every quantitative statement in the results. The concern is fixable—by either defining a common normalization, reporting standardized units, or explicitly converting with the relevant σ—but as submitted the paper cannot be fully verified. I agree with the reader's identification of this as the weakest assumption. I did not find a more severe internal inconsistency; the qualitative spatial patterns (differences concentrated near the BS) are plausible and consistent with the simulation setup. The concern therefore warrants a conditional acceptance rather than rejection, matching the reader's verdict.","tokens_in":15230,"tokens_out":4058,"duration_ms":39370,"concrete_test":"For one representative grid point near the BS, take the raw Sionna path parameters used for Fig. 7 and compute dτ and dP exactly as in Eq. (3) with per-set standardization, plus dDoD/dDoA. Then attempt to reproduce the quoted ≈16 dB, ≈277 ns, and ≈46 deg values. If the reported values are not equal to the computed standardized distances scaled by a stated constant (e.g., σP and στ from one scenario), or to arccos(1−d) for angles, then the results in Section 4 are inconsistent with the metric definition. This check only needs one point and the authors' raw data; it settles whether the unit mismatch is a presentation error or a definitional gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (3) defines dτ and dP as absolute differences of per-set standardized variables \\barτ=(τ−μτ)/στ and \\barP=(P−μP)/σP, and dDoD/dDoA as cosine distances. All four components are dimensionless and unit-free. Section 4, however, reports and color-scales 'Power distance (dB)', 'Delay distance (nsec)', and 'Direction (deg)', with quantitative averages such as ≈16 dB, ≈277 ns, ≈46 deg. The paper never states a transformation from the standardized distances to physical units. Because standardization is per-set (μ,σ computed on X or Y separately), a constant power offset between scenarios cancels, so dP cannot represent an absolute dB difference; a constant delay offset similarly cancels. To recover dB or ns one would need to multiply by σP or στ, but which σ is used is unspecified, and the reported values are not derivable from Eq. (3) alone. Similarly, dDoD=1−u_v·u_w is in [0,2], not degrees; reporting 46 deg would require arccos(1−d), which is not defined in the metric. Thus the specific quantitative findings in Section 4 and the validation of HRT/CRT as physical-comparison metrics do not follow from the proposed definitions. The central claim—that HRT/CRT 'consistently compare temporal, angular and power features'—is not supported by the numerical evidence as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two dissimilarity measures, the Hausdorff ray tracing (HRT) and Chamfer ray tracing (CRT) distances, to compare the multipath parameters (power, delay, departure and arrival angles) produced by ray tracing simulations of the same transmitter-receiver pair under different environmental-modeling assumptions. The method represents each set of propagation paths as a point cloud in a six-dimensional space, standardizes the power and delay dimensions using per-set means and standard deviations, and computes a composite tuple distance dR (Eq. 4) that combines standardized power and delay absolute differences with cosine distances for DoD and DoA. HRT and CRT are then defined as bidirectional Hausdorff and Chamfer statistics on these distances. The authors evaluate the metrics on a digital twin of an area of Milan, Italy, comparing a baseline scenario with (i) a scenario enriched with 505 parked vehicle meshes and (ii) a scenario with segmented building windows assigned a different radio material, using Sionna RT and SUMO at 28 GHz. The results are presented as maps and trajectory plots reporting the distances in dB, nanoseconds, and degrees.","tokens_in":15502,"tokens_out":5449,"duration_ms":57927,"significance":"If the proposed metrics were correctly defined and the reported numbers were derivable from them, the approach would be a useful tool for assessing the fidelity of environmental models in digital twins: it offers a quantitative and spatially localizable comparison of how changes in 3D scene geometry and materials alter simulated radio channels, and it handles sets of paths with different cardinality. The experimental setup is substantial: a high-fidelity urban model, realistic vehicle placement, manual facade segmentation, and integration of two open-source simulators. The paper also contributes two concrete case studies with a reproducible simulation pipeline using publicly documented tools (Sionna RT, SUMO, Blender), although no code or dataset is released. However, the central quantitative claims are undermined by inconsistencies between the formal definitions and the reported physical units, so the current version does not support the stated objective of 'consistently comparing temporal, angular and power features' in physical terms.","major_comments":[{"comment":"","section":"Section 3.2 (Eq. 3) and Section 4.2"},{"comment":"","section":"Section 3.2 (standardization) and Section 4.3"},{"comment":"","section":"Section 3.3 and Section 4.2"}],"minor_comments":[{"comment":"","section":"Section 4.2, paragraph on grid-based simulations"},{"comment":"","section":"Table 1"},{"comment":"","section":"Section 3.2 and Section 4.4"},{"comment":"","section":"Section 4.2"},{"comment":"","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core idea of using Hausdorff and Chamfer distances on multipath parameter point clouds is sound and the experimental scenario is substantial. However, the quantitative results as presented are not derivable from the formal definitions: the standardization in Eq. (3) produces dimensionless distances, yet the paper reports and color-scales physical units without a stated conversion. The per-set standardization also undermines cross-comparison claims. These are load-bearing issues that can be fixed either by revising the definitions to use raw values or by reframing all reported numbers as relative standardized dissimilarities and removing the dB/ns/degree labels. I recommend major revision rather than rejection because the concept is salvageable and the experimental data appear carefully produced. The authors should also decide whether they wish to call the quantities 'metrics' given that the angular component is not a true metric."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is worth a look: representing ray tracing paths as points in R6 and comparing two simulations with bidirectional Hausdorff and Chamfer distances is a clean, useful construction for digital twin work. The composite distance dR, with standardized power/delay and cosine angular terms, is reasonably thought-out, and the two case studies (parked vehicles, window segmentation) are concrete and credible. The integration of Sionna RT with SUMO is a good practical contribution, and the qualitative finding that differences concentrate near the base station is plausible. But the stress-test note is right, and it lands on a load-bearing part of the paper. Eq. (3) defines d_tau and d_P on per-set standardized values, and d_DoD/d_DoA as cosine distances; all are dimensionless. Section 4 reports averages in dB, nanoseconds, and degrees, with no stated conversion. You cannot get about 16 dB or 277 ns from those definitions as written. The per-set standardization also means any constant power or delay offset between scenarios cancels, so the reported numbers cannot represent absolute physical differences. The figures and color bars inherit the same problem. This is not a minor presentational slip; it separates the main quantitative findings from the mathematics. There are two smaller soft spots. First, HRT and CRT are just the max and mean of the same nearest-neighbor distances, so their agreement in flagging the same areas is partly a consequence of construction, not an independent check. The paper does note this relation, but uses it as validation without much caution. Second, no code, data, or baseline comparison is provided, so the actual metric values have no external reference point. The cited prior work is mostly fine, including some self-citations that are relevant. Who is this for? People building 6G digital twins and comparing ray tracing outputs across environmental model choices. The metric framework is useful despite the flaws. A serious referee should engage with it, but the revision must reconcile Eq. (3) with Section 4, clarify the standardization and its implications (or provide the re-scaling), and ideally release artifacts. As it stands, I would not rely on the numerical results, but the conceptual contribution is solid enough to merit careful peer review rather than a desk rejection.","headline":"The metric idea is sensible and the case studies are real, but the reported dB/ns/degree numbers do not follow from the equations, so the quantitative claims need a major fix.","tokens_in":682,"tokens_out":948,"would_cite":false,"duration_ms":29950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes Hausdorff and chamfer point-cloud distances that quantify how environmental modeling changes—adding parked vehicles or segmenting building windows—alter the delay, power, and angle features of simulated radio channels…","keywords":["Digital Twin","Ray tracing","Environmental modeling","Hausdorff distance","Chamfer distance","Point cloud comparison","28 GHz propagation","Vehicular simulation"],"falsifier":"Take two simulations of the same scene and add a constant power offset, say +6 dB, to every path in one of them; under the stated per-scenario standardization both HRT and CRT report a power distance of zero, which would demonstrate that the metric measures only relative shape and cannot detect absolute level changes. Conversely, recomputing the reported 16 dB average power distance directly from equations (3) and (7) would require the missing transformation from standardized units to decibels, and its absence would settle whether the reported numbers are well-defined.","tokens_in":15009,"feed_emoji":"📡","tokens_out":12169,"duration_ms":106401,"temperature":0.7,"pith_summary":"This paper is trying to establish a quantitative, localizable way to ask when a digital twin of an electromagnetic environment is good enough: given two ray tracing simulations for the same transmitter and receiver over 3D scenes whose modeling details differ, the proposed Hausdorff and chamfer ray-tracing distances (HRT and CRT) convert the difference into numbers in the delay, power, and angular domains. The motivation is practical, because 6G network digital twins depend on environmental fidelity and builders need to know which mesh details actually change propagation before investing in them. Tested on a 550 by 670 meter digital twin of a Milan urban area at 28 GHz, with parked vehicles added and building facades split into glass window components, both metrics flag the same geographic regions as most affected, concentrated near the base station, and the vehicular simulations trace how those differences evolve along realistic routes. If the approach is right, it gives network designers an interpretable tool for deciding where modeling effort matters and for checking that a simplified environment still reproduces the radio behavior of a detailed one.","feed_headline":"Two metrics find where 3D model edits alter 6G radio paths","feed_subtitle":"Treating simulated ray paths as point clouds, Hausdorff and Chamfer distances flag the zones model changes perturb most.","key_machinery":"The central object is the ray-path point cloud: each simulated path is a 6-tuple $(P,\\tau,\\theta_{\\mathrm{DoD}},\\varphi_{\\mathrm{DoD}},\\theta_{\\mathrm{DoA}},\\varphi_{\\mathrm{DoA}})$ in $\\mathbb{R}^6$, with power and delay standardized to zero mean and unit variance within each scenario and angular differences measured by the cosine distance $1-\\hat{u}(\\theta,\\varphi)\\cdot\\hat{u}(\\theta',\\varphi')$ between free-space unit vectors. On these clouds the aggregate distance $d_{\\mathrm{R}}=d_\\tau+d_P+d_{\\mathrm{DoD}}+d_{\\mathrm{DoA}}$ defines the nearest-neighbor assignment, and two bidirectional set distances summarize the result: the Hausdorff distance, a worst-case (maximum) statistic, and the Chamfer distance, an average statistic, each symmetrized over the two sets. This machinery does the work of comparing simulations that produce different numbers of paths, since it never requires a one-to-one correspondence between rays.","core_discovery":"The paper's central claim is that a propagation path can be represented as a point in a six-dimensional space—received power, delay, departure azimuth and elevation, and arrival azimuth and elevation—and that the bidirectional Hausdorff and Chamfer distances between two such point clouds form a consistent and interpretable measure of how an environmental change alters the simulated channel. Distances in power and delay are computed on per-scenario standardized values, the angular parts are cosine distances between unit vectors, and the four components are summed into one aggregate distance that drives the nearest-neighbor search while each feature's contribution is separately recorded. Applied to the Milan digital twin, the metrics show that adding 505 parked vehicles changes simulated propagation mostly near the base station, with average HRT power differences around 16 dB on the grid and up to 35 dB along a vehicular trajectory, while switching building windows from concrete to glass changes only the power feature, up to about 2.4 dB in HRT and 0.75 dB in CRT. The fact that HRT and CRT flag the same areas, with the maximum-type statistic consistently above the average-type one, is what the authors take as evidence that the comparison is trustworthy.","pith_inferences":["Because each scenario's power and delay are standardized by their own mean and standard deviation, a uniform offset between the two scenarios—say a global shadowing event or a wrong absorption constant for a material—produces zero distance, so the metric is a relative rather than an absolute fidelity check.","The concentration of differences near the base station may be partly an artifact of the shooting-and-bouncing-ray sampling, which launches a fixed number of candidate rays and therefore rarely reaches distant vehicles, so the claimed geography of modeling impact should be re-checked with an exhaustive ray tracer before being treated as physical.","A natural extension the paper does not develop is to use HRT and CRT as a loss function for calibrating material parameters: adjust the simulated materials until the point-cloud distance to a measured set of paths is minimized.","The windows case, where only the power feature moves and reachable geometry is identical, suggests the metrics can be used to separate material-calibration errors from geometric modeling errors in mixed scenes."],"forward_implications":["Because the distances operate on sets of unequal cardinality through nearest-neighbor matching under the joint metric, any two simulation runs can be compared without one-to-one path correspondence.","The per-feature tracking means one computation yields separate delay, power, departure-angle, and arrival-angle distances, so an engineer can see which physical feature a modeling change actually perturbs.","In the two Milan case studies, both metrics locate the main differences near the base station, suggesting that for dense urban deployments the fidelity of close-range modeling dominates the simulated channel.","The vehicular runs show how HRT and CRT evolve along a trajectory, identifying the stretches of a route where the environmental model has the largest effect on the simulated link.","The authors report analogous patterns at 7 GHz, indicating the comparison procedure is not specific to the 28 GHz band."],"supporting_citations":[{"why":"Supplies the ray tracing engine whose output paths (power, delay, departure and arrival angles) form the point clouds that the proposed metrics compare.","marker":"[33]"},{"why":"Supplies the microscopic traffic simulation that provides realistic receiver positions along vehicular routes.","marker":"[32]"},{"why":"Supplies the material parameters (concrete, glass, perfect conductor) assigned to the buildings, windows, and vehicle meshes at 28 GHz.","marker":"[30]"},{"why":"Supplies the 3D vehicle meshes used to model the 505 parked cars added to the base scenario.","marker":"[28]"},{"why":"Supplies the road topology for the Milan area, which is converted into the traffic network used by the vehicular simulator.","marker":"[27]"},{"why":"Supplies the satellite imagery used to place and orient the parked vehicle meshes at realistic positions.","marker":"[31]"}],"fun_headline_variants":["Hausdorff and Chamfer distances expose 3D edits on 6G paths","New metrics flag zones where model changes alter 6G radio","Ray tracing metrics measure fidelity of 6G digital twins","Chamfer and Hausdorff ray tracing distances compare 6G scenarios","Metrics reveal how 3D mesh edits shift simulated 6G channels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the power and delay distances, which are computed on values standardized separately for each scenario, can be reported as physical differences in decibels and nanoseconds as done in the results section, even though the paper never states the mapping between the standardized distances and those physical units, and the per-scenario standardization makes any uniform offset between the two scenarios invisible to the metrics.","fun_headline_variants_meta":{"raw":{"variants":["Hausdorff and Chamfer distances expose 3D edits on 6G paths","New metrics flag zones where model changes alter 6G radio","Ray tracing metrics measure fidelity of 6G digital twins","Chamfer and Hausdorff ray tracing distances compare 6G scenarios","Metrics reveal how 3D mesh edits shift simulated 6G channels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1495,"prompt_tokens":1024,"completion_tokens":471,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":375}},"tokens_in":640,"tokens_out":471,"duration_ms":4991,"temperature":1.0,"reasoning_tokens":375,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:59:24.450845+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take two simulations of the same scene and add a constant power offset, say +6 dB, to every path in one of them; under the stated per-scenario standardization both HRT and CRT report a power distance of zero, which would demonstrate that the metric measures only relative shape and cannot detect absolute level changes. Conversely, recomputing the reported 16 dB average power distance directly from equations (3) and (7) would require the missing transformation from standardized units to decibels, and its absence would settle whether the reported numbers are well-defined.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the microscopic traffic simulation that provides realistic receiver positions along vehicular routes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the material parameters (concrete, glass, perfect conductor) assigned to the buildings, windows, and vehicle meshes at 28 GHz."},{"cited_title":"Dosovitskiy, G","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D vehicle meshes used to model the 505 parked cars added to the base scenario."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the road topology for the Milan area, which is converted into the traffic network used by the vehicular simulator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the satellite imagery used to place and orient the parked vehicle meshes at realistic positions."}],"review_version":2}