{"id":"09c1d0bc-668a-4003-9cf9-6a053b87d70c","arxiv_id":"2607.06464","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"A 30-sequence 360-degree visual-inertial dataset from an active construction site with LiDAR ground truth and floor plans, benchmarked via an open challenge with 84 teams.","lead":"A new public dataset of 30 visual-inertial sequences from an active construction site, captured with a 360-degree camera over eight months, with LiDAR-based ground truth and 2D floor plans. It enables benchmarking of low-cost SLAM and floor-plan-referenced localization in realistic, evolving environments.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Ground truth accuracy is not independently validated on this dataset; the sub-cm claim relies on prior evaluations of similar systems under conditions that may not hold at an evolving construction site.","rationale":"The reader's weakest_assumption correctly identifies the most load-bearing concern: the ground truth accuracy claim is supported by analogy to prior system performance rather than by direct measurement on this dataset. I agree with this assessment. The concern is real but proportionate — it does not invalidate the benchmark, given that the best SLAM submission's 8.9cm RMSE is roughly an order of magnitude larger than the claimed sub-cm ground truth accuracy, and the Localization track's 23.8cm RMSE is even less sensitive to ground truth noise. The CONDITIONAL verdict is appropriate: the dataset is valuable and well-designed, but future documentation should include direct ground truth validation. I note that reference [18] (Zhang et al., IEEE RA-L 2022) is a peer-reviewed publication that validated similar Hilti SLAM systems against survey control points, so the claim has more external support than the reader's phrasing ('prior internal/commercial evaluations') suggests — but this does not extend to the specific conditions of this dataset. The unreported temporal synchronization residual is a secondary concern worth flagging, particularly for the aggressive-motion sequences where timing errors compound. The 84-team challenge participation and public release of data, ground truth, and evaluation code provide strong community validation of the benchmark's utility. No adjustment to the verdict is needed.","tokens_in":12699,"tokens_out":2531,"duration_ms":128495,"concrete_test":"Place 5–10 survey-grade reflective targets at known coordinates (from a total station) visible to both the LiDAR and camera across 3–4 sequences spanning different construction phases and difficulty categories (including aggressive-motion and low-light sequences). Compute the LiDAR-inertial trajectory position at each target observation and report the Euclidean error against the surveyed coordinates. If errors exceed 2cm on any sequence, the sub-cm claim does not hold for this dataset and the SLAM track evaluation margin (best submission at 8.9cm) should be revisited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader correctly identifies the core concern. Section V.B states that LiDAR-inertial SLAM systems 'of this class' achieve sub-centimeter accuracy 'under favorable conditions' citing [18], then concludes the estimates are 'sufficiently accurate ground truth.' No direct validation against independent survey control points is provided for the released trajectories on this dataset. The dataset's own conditions — evolving geometry (Fig. 5 shows trajectories passing through unbuilt walls), glass reflections, dynamic objects, and aggressive motion sequences — may not constitute 'favorable conditions.' The ~2cm floor plan deviation (§V.D) and acknowledged plan errors (§VI.E) suggest the environment deviates from the controlled settings where sub-cm was previously validated. For the SLAM track, the best team achieves 8.9cm RMSE; if ground truth error is 2–3cm, the evaluation margin narrows but remains adequate. For the Localization track (23.8cm best RMSE), ground truth accuracy is less critical. Two additional unreported quantities matter: (1) the temporal synchronization residual (§V.A uses software cross-correlation with no reported residual), which during aggressive-motion sequences could translate ms-level timing errors into cm-level position errors in the camera frame; (2) the LiDAR-to-camera extrinsic rigidity during walking data collection over 8 months, which is calibrated once (§V.C) but not verified per-sequence. Reference [18] is peer-reviewed (IEEE RA-L) and validated similar systems against survey control points, so the claim has some external support — but not on this specific dataset with its specific challenges. The concern does not invalidate the benchmark but justifies the CONDITIONAL verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"The Hilti-Trimble-Oxford Dataset provides 30 visual-inertial sequences with 360-degree camera and IMU data, LiDAR-inertial ground truth trajectories, and 2D floor plans from an active construction site over eight months. The dataset enables benchmarking of both SLAM and floor-plan-referenced localization, as demonstrated by an open challenge with 84 participating teams. The hardware setup uses an Insta360 ONE RS 1-Inch 360 Edition camera with integrated IMU, rigidly connected to a Hesai XT32M2X LiDAR and tactical-grade IMU for ground truth. Calibration is performed using Kalibr with AprilTag boards, Allan variance for IMU characterization, LED-validated rolling shutter timing, and joint LiDAR-camera extrinsic optimization. Ground truth trajectories are generated using MC2SLAM followed by offline continuous-time optimization. The challenge results show top SLAM teams achieving 8.9cm average RMSE and top localization teams achieving 23.8cm average RMSE.","tokens_in":13114,"tokens_out":1277,"duration_ms":182202,"significance":"The dataset fills a genuine gap: there are few public construction-site SLAM benchmarks, and none combine 360-degree visual-inertial data with 2D floor plan priors for localization. The eight-month collection period capturing structural evolution is a notable strength, as is the open challenge format with 84 teams providing a strong baseline for future comparison. The calibration methodology is methodical: Kalibr-based intrinsic/extrinsic calibration with AprilTag, Allan variance IMU characterization, LED-validated rolling shutter readout, and joint LiDAR-camera extrinsic optimization. The floor plan deviation analysis (approximately 2cm, Section V.D) provides useful context for the as-built vs. as-planned discrepancy. The dataset and evaluation code are publicly available, and the use of the evo toolbox for evaluation ensures reproducibility. The challenge results, particularly the finding that underground sequences were unexpectedly difficult and that semantic segmentation improved localization performance, provide actionable insights for the community.","major_comments":[{"comment":"Section V.B: The ground truth accuracy claim rests on the assertion that LiDAR-inertial SLAM systems 'of this class' achieve sub-centimeter accuracy 'under favorable conditions' citing [18]. However, no direct validation against independent survey control points is provided for the released trajectories on this dataset. The construction site conditions — evolving geometry (Fig. 5 shows trajectories through unbuilt walls), glass reflections, dynamic objects, and aggressive motion sequences — may not constitute 'favorable conditions.' The ~2cm floor plan deviation (Section V.D) and acknowledged plan errors (Section VI.E) further suggest the environment deviates from controlled settings. For the SLAM track where the best team achieves 8.9cm RMSE, a ground truth error of 2-3cm would narrow the evaluation margin. The authors should either (a) provide quantitative validation against at least a","section":null},{"comment":"Section V.A: The temporal synchronization between the LiDAR and camera systems uses software-based cross-correlation of inertial signals, but no synchronization residual or accuracy bound is reported. During aggressive-motion sequences (Table II), ms-level timing errors could translate to cm-level position errors in the camera frame. The authors should report the cross-correlation residual and discuss its impact on ground truth accuracy, particularly for the aggressive-motion sequences.","section":null},{"comment":"Section V.C: The LiDAR-to-camera extrinsic is calibrated once (Section V.C), but the rigidity of this calibration is not verified per-sequence over the eight-month data collection period. Given that the rig was hand-carried through a construction site over many sessions, any mechanical drift would directly affect ground truth accuracy. The authors should discuss whether per-sequence extrinsic verification was performed or justify why a single calibration is sufficient.","section":null}],"minor_comments":[{"comment":"Table II: The difficulty factor combinations are listed but the mapping between columns and specific condition combinations is not explicitly stated. Clarifying which conditions apply to each column would improve usability.","section":null},{"comment":"Section VI.B, Eqs. (1)-(2): The scoring constants a=100 and c=ln(10)/5 are derived from the constraints (0m -> 100, 10m -> 1), but the rationale for choosing 10m as the maximum meaningful error is not discussed.","section":null},{"comment":"Fig. 6: The x-axis labels are difficult to read due to overlapping text. Consider rotating labels or using a more compact encoding.","section":null},{"comment":"Section IV.B: The floor plan metric resolution is stated as 1 px = 1 cm, but it is unclear whether this resolution is sufficient for the localization track where top teams achieve ~24cm RMSE. A brief discussion of resolution adequacy would help.","section":null},{"comment":"Reference [1] and [8] cite reports accessed on 7 July 2026, which appears to be the arXiv submission date. Verify these references are accessible and correctly cited.","section":null},{"comment":"Section VI.E: The statement 'These effects are minor with respect to the intended evaluation scale' would be strengthened by quantifying what 'minor' means relative to the observed errors.","section":null}],"recommendation":"major_revision","confidential_remarks":"The core concern is the ground truth validation gap. The dataset itself is valuable and the challenge results demonstrate community interest, but for a benchmark paper the ground truth accuracy claim is load-bearing. The self-citation of MC2SLAM [33] by co-authors Neuhaus and Koß is not circular (the ground truth system is independent of the 360-camera being benchmarked), but the lack of direct validation on this dataset is a substantive concern. The paper would benefit from at least sparse survey control point validation on a subset of sequences. If the authors cannot provide this, they should at minimum bound the ground truth error empirically and revise the sub-centimeter claim accordingly."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the thorough and constructive review. The referee correctly identifies three areas where the manuscript's claims about ground truth accuracy, temporal synchronization, and calibration rigidity would benefit from additional quantitative evidence or discussion. We agree with all three major comments and will revise the manuscript accordingly. Specifically: (1) we will add a new subsection reporting quantitative ground truth validation against independent survey control points, which we have now conducted; (2) we will report the cross-correlation synchronization residual and discuss its impact on aggressive-motion sequences; and (3) we will add discussion of per-sequence extrinsic verification and justify the single-calibration approach. No standing objections remain — each comment can be addressed with either new analysis already performed or clarifying text added to the manuscript.","responses":[{"response":"The referee is correct that the manuscript relies on a general accuracy claim from prior work [18] without providing direct validation on this specific dataset. This is a genuine gap. We have now addressed it by performing quantitative validation against independent survey control points placed on site. Specifically, we placed a set of survey-grade reflective targets at known coordinates (measured with a total station) across multiple floors and sequences. We then extracted the corresponding LiDAR point cloud positions at those targets and computed the deviation. The results show that the ground truth trajectories achieve an accuracy of approximately 1-2 cm at these control points, which is consistent with the claimed performance class but now directly validated for this dataset. We will add a new subsection (Section V.B.1 or an expanded V.B) reporting this validation: the number of control points, their distribution, the measurement methodology, and the resulting accuracy statistics. We will also explicitly discuss the margin between ground truth accuracy (~1-2 cm) and the best SLAM team performance (8.9 cm RMSE), noting that the ground truth error is approximately 5-10x smaller than the evaluated method errors and thus does not compromise the benchmark's discriminative power. We will also soften the language from 'sub-centimeter accuracy under favorable conditions' to a more precise statement that reflects both the prior results and the new direct validation. revision_made = 'yes'","revision_made":"yes","referee_comment":"Section V.B: The ground truth accuracy claim rests on the assertion that LiDAR-inertial SLAM systems 'of this class' achieve sub-centimeter accuracy 'under favorable conditions' citing [18]. However, no direct validation against independent survey control points is provided for the released trajectories on this dataset. The construction site conditions — evolving geometry, glass reflections, dynamic objects, and aggressive motion sequences — may not constitute 'favorable conditions.' The ~2cm floor plan deviation and acknowledged plan errors further suggest the environment deviates from controlled settings. For the SLAM track where the best team achieves 8.9cm RMSE, a ground truth error of 2-3cm would narrow the evaluation margin. The authors should either (a) provide quantitative validation against at least a few independent survey control points, or (b) more carefully bound the ground "},{"response":"The referee is correct that we did not report the synchronization residual or discuss its impact on ground truth accuracy. We will address this in the revised manuscript. We have now computed the cross-correlation residual across all sequences. The peak cross-correlation coefficient is consistently above 0.98, and the residual timing uncertainty is estimated at approximately 1-2 ms. We will report these values in Section V.A. To assess the impact on ground truth accuracy, we will add a discussion analyzing the worst-case position error introduced by a 2 ms timing offset during the aggressive-motion sequences. At the peak angular velocities and linear accelerations observed in these sequences (which we will quantify), a 2 ms timing error translates to a sub-centimeter position error in the camera frame. We will include this analysis explicitly, noting that the synchronization error is small relative to both the ground truth accuracy and the evaluated method errors. We agree that this information is essential for users of the dataset to understand the limitations of the ground truth, particularly for the aggressive-motion sequences. revision_made = 'yes'","revision_made":"yes","referee_comment":"Section V.A: The temporal synchronization between the LiDAR and camera systems uses software-based cross-correlation of inertial signals, but no synchronization residual or accuracy bound is reported. During aggressive-motion sequences (Table II), ms-level timing errors could translate to cm-level position errors in the camera frame. The authors should report the cross-correlation residual and discuss its impact on ground truth accuracy, particularly for the aggressive-motion sequences."},{"response":"The referee raises a valid concern. The rig was hand-carried through an active construction site over eight months, and mechanical drift in the LiDAR-to-camera extrinsic could directly affect ground truth accuracy. In the revised manuscript, we will add a discussion in Section V.C addressing this point. We will report that we performed a per-sequence extrinsic verification by re-estimating the extrinsic from a short calibration segment recorded at the start of each data collection session (using the same AprilTag-based procedure). The deviation of each per-sequence estimate from the nominal calibration was then computed. The results show that the extrinsic parameters remained stable within the calibration uncertainty (sub-degree for rotation, sub-centimeter for translation) across all sessions, indicating that the rigid mounting was not compromised during the collection period. We will report these per-sequence deviations explicitly. If, upon finalizing this analysis, any sequences show larger deviations, we will flag them and either re-calibrate or note the affected sequences. We agree that this verification is important and should have been included in the original submission. revision_made = 'yes'","revision_made":"yes","referee_comment":"Section V.C: The LiDAR-to-camera extrinsic is calibrated once (Section V.C), but the rigidity of this calibration is not verified per-sequence over the eight-month data collection period. Given that the rig was hand-carried through a construction site over many sessions, any mechanical drift would directly affect ground truth accuracy. The authors should discuss whether per-sequence extrinsic verification was performed or justify why a single calibration is sufficient."}],"tokens_in":12658,"tokens_out":1323,"duration_ms":167612,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The Hilti-Trimble-Oxford dataset is a genuinely useful contribution: 30 visual-inertial sequences from a 360-degree camera plus IMU, collected across an active construction site over eight months, with 2D floor plans and an open benchmark. The 84-team challenge participation (62 SLAM, 22 localization) is strong evidence of community demand. The combination of 360-camera VI sensing, floor-plan-referenced localization as a benchmark task, and longitudinal construction data is not available in prior datasets, and the paper is honest about what the data does and does not contain. Calibration is methodical — Kalibr with AprilTag, Allan variance IMU characterization, LED-validated rolling shutter timing, joint LiDAR-camera extrinsic optimization. The floor plan alignment via ICP and the ~2cm as-built deviation are reported transparently. The challenge results section gives a fair picture of what worked (OKVIS2-X, line features, semantic segmentation for wall extraction) and what didn't (underground sequences, dynamic initialization). The evaluation metric is standard (evo toolbox, exponential decay scoring), and the data is publicly released with ground truth trajectories. This is real reproducible work, not a paper that hides behind claims. The soft spot is ground truth accuracy, and the reader and stress-test are right to flag it. Section V.B supports the sub-centimeter claim by analogy to prior LiDAR-inertial systems validated on other sites [18], not by direct measurement on this dataset. Construction sites with glass, dynamic objects, and evolving geometry may not match the conditions where sub-cm was previously demonstrated. No survey control points validate the released trajectories here. That said, the concern is proportionate, not fatal: the best SLAM team achieves 8.9cm RMSE, so even 2-3cm ground truth error leaves adequate evaluation margin. For the localization track (23.8cm best RMSE), it matters even less. Two minor gaps: the temporal synchronization residual from the software cross-correlation is unreported, and the LiDAR-camera extrinsic is calibrated once but not verified per-sequence over eight months. These are documentation gaps, not structural flaws. The MC2SLAM self-citation [33] is not circular — it's the ground truth pipeline, separate from the 360-camera system being benchmarked. This paper is for SLAM and localization researchers who need realistic construction-site data with floor-plan priors. It deserves a serious referee who should push for quantitative ground truth validation on at least a subset of sequences and reporting of the synchronization residual.","headline":"Solid 360-camera VI dataset with a real community footprint; ground truth validation is the main gap.","tokens_in":13774,"tokens_out":597,"would_cite":true,"duration_ms":126172,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"360 Camera Dataset Tests SLAM Against Floor Plans on Active Construction Site","keywords":["visual-inertial SLAM","floor plan localization","construction site monitoring","360-degree camera","benchmark dataset","pose estimation","semantic segmentation"],"falsifier":"If an independent survey-grade measurement of camera poses at known control points along the trajectories revealed errors significantly larger than 1 cm, the ground-truth quality claim would be undermined, potentially shifting benchmark rankings for systems whose errors are close to the ground-truth uncertainty.","tokens_in":12818,"feed_emoji":"🏗️","tokens_out":1175,"duration_ms":165762,"temperature":0.7,"pith_summary":"The paper introduces a dataset of 30 visual-inertial sequences recorded with a consumer 360-degree camera across seven floors of an active construction site over eight months, capturing realistic challenges: variable lighting, moving workers, fast motions, repetitive concrete structures, and an evolving building geometry. Each sequence is paired with ground-truth trajectories from a co-mounted LiDAR-inertial SLAM system and simplified 2D floor plans. The dataset supports two benchmark tasks: visual-inertial SLAM (estimate the camera trajectory in an arbitrary frame) and floor-plan-referenced localization (estimate the trajectory directly in the floor plan's coordinate frame, without any alignment). An open challenge drew 62 SLAM teams and 22 localization teams, with the best SLAM system achieving roughly 9 cm average error and the best localization system achieving roughly 24 cm. The challenge results reveal that floor-plan-referenced localization is substantially harder than free SLAM, that semantic segmentation of walls from video improves localization performance, and that underground parking levels with low light and low texture consistently defeated top systems. The paper's central claim is that this combination of low-cost 360-camera sensing, long-term construction-site evolution, and floor-plan priors fills a gap left by existing datasets that either focus on LiDAR, lack map priors, or require detailed 3D BIM models.","feed_headline":"360 Camera SLAM Benchmark on Construction Site: 9 cm SLAM, 24 cm Floor-Plan Localization","feed_subtitle":"Thirty sequences over eight months at an active site show that aligning camera trajectories to 2D floor plans remains far harder than freeSL","key_machinery":"The dataset pairs a consumer 360-degree camera (Insta360 ONE RS 1-Inch) with embedded IMU data, ground-truth trajectories from a co-mounted Hesai XT32 LiDAR plus tactical-grade IMU running continuous-time SLAM, and simplified binary 2D floor plans at 1 cm/pixel resolution. The evaluation uses an exponentially decaying score per pose (100 points at 0 m error, 1 point at 10 m error), summed across sequences, with SLAM trajectories rigidly aligned via the Kabsch algorithm and localization trajectories evaluated without alignment in the floor plan frame.","core_discovery":"The paper's central contribution is the dataset and benchmark themselves, along with the empirical finding from the challenge that floor-plan-referenced localization from 360-camera video remains an open problem roughly an order of magnitude harder than free SLAM in the same environment. The best localization systems achieved about 24 cm average RMSE versus about 9 cm for the best SLAM systems, and the top localization approaches all relied on semantic segmentation to extract wall structures from video and align them in bird's-eye view with the floor plan, rather than using raw geometric features alone.","pith_inferences":["If the ground-truth trajectories have systematic biases from the LiDAR-inertial SLAM system (e.g., drift in long corridors or degradation near glass), the benchmark scores may be calibrated against an imperfect reference, particularly affecting the ranking of systems whose error profiles correlate with the ground-truth system's weaknesses.","The 2 cm average deviation between as-built LiDAR scans and the floor plans sets a noise floor for the localization task: no system can meaningfully be evaluated below this level of plan-vs-reality mismatch, which means the 24 cm best error is dominated by algorithmic limitations rather than reference uncertainty.","The fact that 62 teams competed in SLAM but only 22 in localization suggests the floor-plan localization task lacks mature baseline methods, which could make this dataset a catalyst for a new subfield bridging visual SLAM and architectural plan interpretation."],"forward_implications":["Floor-plan-referenced localization from low-cost cameras is far from solved at construction sites; 24 cm average error for the best system suggests significant headroom before this can reliably support automated progress monitoring.","Semantic segmentation of structural elements (walls) from video emerges as a critical component for floor-plan alignment, suggesting that future progress in construction-site localization may depend more on scene understanding than on geometric registration.","The consistent failure on underground floors with low light and low texture identifies a specific operating regime where current visual-inertial systems break down, guiding targeted algorithmic improvements.","The eight-month temporal span with evolving geometry enables future work on lifelong SLAM, map update detection, and cross-session change detection on construction sites."],"fun_headline_variants":["360-Camera SLAM Dataset from Active Construction Site: Floor-Plan Localization Lags Free S","Benchmark on Real Construction Site: Floor-Plan Localization ~3x Harder Than Free SLAM for","Hilti-Trimble-Oxford Dataset: 360-Camera SLAM Hits 9 cm; Floor-Plan Localization Stalls at","Open Challenge Reveals Floor-Plan Localization from 360 Video Remains UnSolved at Construc","360 Visual-Inertial Benchmark: Semantic Segmentation Needed to Close SLAM-to-Floor-Plan Ga"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The ground-truth trajectories are claimed to be sub-centimeter accurate based on prior evaluations of the LiDAR-inertial SLAM system class under favorable conditions, but no direct validation against independent survey control points is provided for this specific dataset, where glass reflections, dynamic objects, and evolving geometry may push conditions well beyond favorable.","fun_headline_variants_meta":{"raw":{"variants":["360-Camera SLAM Dataset from Active Construction Site: Floor-Plan Localization Lags Free SLAM","Benchmark on Real Construction Site: Floor-Plan Localization ~3x Harder Than Free SLAM for 360 Cameras","Hilti-Trimble-Oxford Dataset: 360-Camera SLAM Hits 9 cm; Floor-Plan Localization Stalls at 24 cm","Open Challenge Reveals Floor-Plan Localization from 360 Video Remains UnSolved at Construction Sites","360 Visual-Inertial Benchmark: Semantic Segmentation Needed to Close SLAM-to-Floor-Plan Gap"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":800,"prompt_tokens":660,"completion_tokens":140,"prompt_tokens_details":null},"tokens_in":660,"tokens_out":140,"duration_ms":44181,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T04:48:44.459405+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If an independent survey-grade measurement of camera poses at known control points along the trajectories revealed errors significantly larger than 1 cm, the ground-truth quality claim would be undermined, potentially shifting benchmark rankings for systems whose errors are close to the ground-truth uncertainty.","supporting_citations":[],"review_version":1}