{"id":"aa1147be-267c-4aad-a634-2380cd28667a","arxiv_id":"2411.14358","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"InCrowd-VI provides a realistic visual-inertial benchmark with 58 head-worn sequences in crowded indoor spaces, where state-of-the-art SLAM systems frequently fail to meet accuracy and real-time requirements.","lead":"InCrowd-VI is a new dataset of 58 visual-inertial recordings from head-mounted sensors in crowded indoor spaces, built for testing navigation systems for blind and visually impaired people. It includes ground-truth paths and 3D maps, and early tests show current navigation algorithms struggle in these conditions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground truth from the Aria SLAM service is validated only by local landmark distances on a subset of sequences; unverified long-trajectory drift in the reference could bias every ATE/DP comparison.","rationale":"The reader correctly identifies that ground truth originates from the Aria SLAM service and is validated only on a subset of sequences. My stress-test goes one step further: even on the sequences that were manually validated, the method checks only distances between landmarks in the reconstructed map, not the absolute or relative trajectory pose error that the ATE and DP metrics consume. A map can be locally scaled correctly while the camera trajectory drifts globally, especially over 100–350 m indoor paths with repetitive architecture, crowds, and motion transitions. This is precisely the regime where the paper claims modern SLAM systems fail, so the reference must be shown to be accurate in that regime. The proposed test is feasible because such survey-grade references have been used in comparable benchmarks (e.g., Hilti-Oxford, UZH-FPV). The dataset itself is still valuable and the conclusions are plausible, but they are conditional on an independent check of the reference trajectory. Thus the CONDITIONAL verdict remains appropriate, and no adjustment is needed.","tokens_in":171,"tokens_out":2949,"duration_ms":45050,"concrete_test":"Select 3–5 sequences that drive the headline conclusions: Kiko_loop (314 m), G8_cafe (294 m), Checkin2_loop (348 m), Getoff_pharmacy (159 m, motion transition), and one high-density short sequence such as Orell_fussli. Establish an independent survey-grade reference using a Leica total station or a motion-capture system, measuring a set of control points along the walked path and recording synchronized timestamps. Compute the absolute trajectory error of the Aria SLAM ground truth against this survey reference, both with and without Sim(3) alignment. If the Aria reference error exceeds ~10 cm or drifts beyond 1% on any of these sequences, the ATE/DP values reported in Table 6 for those sequences are not trustworthy and the paper's central claim should be re-evaluated. If the Aria error stays below ~5 cm, the concern is resolved and the current conclusions stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central benchmark claim depends on the accuracy of the ground-truth trajectories produced by the Meta Aria machine perception SLAM service (§3.3). The paper's validation (§3.3, Figures 1–2) consists of manually measuring distances between recognizable landmarks in the reconstructed point cloud and comparing them to real-world tape-measure distances. This validates local metric scale of the static map, not the global trajectory accuracy that ATE and DP actually measure. A trajectory can locally reconstruct a scene to 2 cm while still accumulating large drift over a 300 m loop or in scenes where the Aria SLAM service loses tracking or fails loop closure. The manual validation was performed on 'selected crowded sequences' and later on a few sequences where all evaluated systems failed; it does not cover most of the 58 sequences, and for the covered ones it only checks inter-landmark distances, not absolute pose error. The evaluation in Table 6 then compares ORB-SLAM3, SVO, DROID-SLAM, and DPV-SLAM against this reference. If the Aria reference drifts by meters in exactly the long, crowded, or motion-transition sequences (e.g., Kiko_loop, G8_cafe, Getoff_pharmacy), the reported ATEs and drift percentages are not measurements of the evaluated systems' error but of the reference's error. The conclusion that 'state-of-the-art systems fail' would be an artifact. The paper itself concedes the Aria service is an offline SLAM/VIO system, so it is not an independent ground truth. Without an independent high-accuracy reference on a representative subset of long and crowded sequences, the benchmark's headline result is not secure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces InCrowd-VI, a visual-inertial dataset for evaluating SLAM in indoor pedestrian-rich environments, recorded with Meta Aria glasses across 58 sequences totaling about 5 km and 1.5 hours. The dataset provides RGB, stereo, and IMU data, semi-dense point clouds, and ground-truth trajectories produced by the Meta Aria machine perception SLAM service, with reported accuracy of about 2 cm based on manual landmark-distance validation. The paper evaluates four SLAM/VO systems (ORB-SLAM3, SVO, DROID-SLAM, DPV-SLAM) on a subset of sequences using ATE, drift percentage, pose estimation coverage, and a real-time factor, concluding that current systems fail to meet the accuracy, robustness, and real-time requirements for visually impaired navigation.","tokens_in":19570,"tokens_out":4652,"duration_ms":48254,"significance":"If the ground-truth accuracy claim holds, InCrowd-VI would be a valuable community resource: it targets an underexplored regime (head-worn, pedestrian-rich indoor navigation), provides realistic human motion and crowd densities, and ships public data plus conversion tools. The evaluation is also useful as a stress test for current SLAM systems, and the introduced PEC and RTF metrics add practical evaluation dimensions beyond ATE alone. The main contributions are empirical and dataset-oriented rather than algorithmic, and the paper is transparent about the sensor setup and calibration. However, the benchmark's central claim depends on the reliability of the Aria SLAM service as ground truth, which is only locally validated, and on the representativeness of the evaluated sequence subset; both points need strengthening before the dataset can be adopted as a benchmark.","major_comments":[{"comment":"The ground-truth validation is not sufficient to support the claimed trajectory accuracy for all 58 sequences. The manual measurements compare distances between landmarks in the reconstructed point cloud with real-world tape-measure distances; this validates the local metric scale of the static map, not the global trajectory error that ATE and DP measure. A trajectory can reconstruct local geometry to within 2 cm while accumulating large drift over long loops or in sequences with tracking difficulty. Since the reference is itself produced by the Meta Aria machine perception SLAM service, any drift in that service (for example in Kiko_loop, G8_cafe, or Getoff_pharmacy) would be attributed to the evaluated systems. The validation was performed on selected crowded sequences and on a few failure sequences, not on most of the 58 sequences, and it checks inter-landmark distances rather than absolute pose error. The paper should either provide per-sequence trajectory uncertainty estimates, validate global consistency with an independent method (e.g., surveyed waypoints, loop-closure constraints, or a total station), or explicitly restrict the accuracy claim to the local metric scale and adjust the benchmark conclusions accordingly.","section":"§3.3, Figures 1–2"},{"comment":"The evaluation is performed on 19 of the 58 sequences, but the selection protocol is not specified beyond stating that the sequences 'represent a range of difficulty levels from easy to hard.' The main conclusions—that state-of-the-art systems fail in this dataset and that drift reaches 5–10%—are extrapolated from this convenience subset. If the subset is biased toward the hardest sequences, the dataset-level failure rate is overstated; if it is biased toward easier sequences, the dataset's challenge level is understated. The authors should evaluate all sequences or clearly document the selection criteria and show that the selected subset has the same distribution of crowd density, trajectory length, and environmental challenges as the full dataset.","section":"§4, Table 6"},{"comment":"Several reported ATE values are associated with very low pose estimation coverage, yet they are presented without flags and are used in the aggregate analysis. For example, in TH_loop the SVO row reports ATE = 0.01 m with PEC = 0.9%, and in UZH_stairs SVO reports PEC = 37%; these are effectively failed runs, and their ATE values are computed on a tiny fraction of the trajectory. The paper itself states that a notably low PEC makes ATE unreliable, but Table 6 still lists these numbers as if they were comparable to high-coverage results. These low-coverage entries should be excluded from the summary statistics or clearly marked as failures, so that the reported drift and accuracy comparisons are not distorted.","section":"§4.1, Table 6"},{"comment":"The four systems are not compared under equivalent sensor configurations: ORB-SLAM3 is evaluated with the left camera plus IMU, while SVO, DROID-SLAM, and DPV-SLAM use only the left monocular camera. This confounds the comparison between classical and deep-learning approaches, since visual-inertial SLAM has access to scale and gyroscopic measurements that monocular systems lack. The conclusion that deep-learning methods are more robust than classical methods, or that classical methods drift more, is not a head-to-head comparison of the algorithms as typically used. The authors should either evaluate all systems in both monocular and visual-inertial modes (where supported) or clearly reframe the results as per-system capability assessments rather than comparative rankings.","section":"§4, 'Experimental Evaluation'"}],"minor_comments":[{"comment":"There is a typo: 'Meta-Ariana glasses' should read 'Meta Aria glasses.' Additionally, 'approximately 1T GB' should be corrected to a proper storage unit, such as 'approximately 1 TB.'","section":"§3.4"},{"comment":"Reference [38] points to an 'edit.paperpal.com' URL, which appears to be an artifact of a writing tool rather than the actual TUM RGB-D benchmark tools page; the canonical URL for the TUM evaluation tools should be used instead.","section":"References"},{"comment":"The caption for Figure 8(a) contains a misspelling ('Midium' instead of 'Medium') and labels 'DROID-SLAM3' in the legend while the text and Table 5 refer to 'DROID-SLAM'; these should be made consistent.","section":"Figure 8"},{"comment":"The definition of PEC as '(number of estimated poses/total number of frames) × 100' should clarify whether estimated poses are counted only for frames where tracking succeeded and how the denominator handles sequences with dropped or corrupted frames; this affects the interpretation of PEC values such as 0.9%.","section":"§4.1"},{"comment":"The manuscript is formatted with an MDPI template that still contains placeholder text ('Journal Not Specified 2024, 1, 0', a placeholder DOI, and empty Received/Accepted dates); these must be resolved for the final published version.","section":"Front matter"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself is timely and likely to be useful to the SLAM community, and the authors have made a real effort to validate their ground truth. The main risk is that the benchmark's central quantitative claims rest on a ground-truth source that is itself a SLAM system, validated only locally. If the authors cannot provide global trajectory validation or per-sequence uncertainty bounds, I would be reluctant to accept the strong claim that 'state-of-the-art systems fail' on this dataset. The evaluation subset issue is fixable with more analysis, but the ground-truth question is the one that most needs attention before the dataset is widely adopted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: InCrowd-VI is a genuinely useful resource, and the paper is honest about what it does. The main caveat—ground truth comes from Meta's own SLAM service—is real and should be addressed, but it doesn't sink the dataset.\n\nWhat's new: a head-worn, indoor, pedestrian-rich visual-inertial dataset with 58 sequences and about 5 km of trajectory. That combination did not exist before; ADVIO and BPOD are handheld or not crowded, NAVER LABS comes from a mapping platform. The data are public, and the authors provide conversion tools and detailed sensor calibration. That alone is a contribution worth building on.\n\nThe evaluation is also a reasonable first pass. They pick four well-known systems, run the stochastic ones five times, and report ATE, drift percentage, pose coverage, and a real-time factor tied to walking speed. The finding that classical methods drift 5–10% in tough sequences while learning-based methods keep high coverage but miss real-time is credible and matches what I would expect in crowded scenes.\n\nThe soft spots are all about the reference. The ground truth comes from the Meta Aria machine perception SLAM service—an offline SLAM/VIO system. The validation is manual measurement of distances between landmarks in the reconstructed point cloud, compared with tape-measure distances, on a subset of sequences. That checks local metric scale, not global trajectory drift. A 300 m loop can be locally accurate and still accumulate meters of error. Since ATE and drift percentage compare against this reference, the headline numbers in Table 6 are not fully independent. I don't think this is fatal—the reported 2 cm agreement in challenging sequences is meaningful evidence, and the Aria service is known to be strong—but the paper should say clearly that the reference is not globally validated and should add independent global ground truth (surveyed points, total station, or LiDAR reference) on at least a few long sequences.\n\nSmaller issues: the evaluation subset is not precisely defined; a reader can't tell exactly which sequences were selected and why. There are no error bars or standard deviations alongside the means, though five runs is reasonable. The writing has rough spots, and the MDPI template with 'Journal Not Specified' doesn't help. None of these change the central usefulness.\n\nBottom line: this is a paper worth reviewing and a dataset worth using. A serious referee should ask for stronger ground-truth validation and a clearer protocol, but the resource itself is solid.","headline":"A genuinely useful new dataset for crowded indoor SLAM evaluation, but the headline numbers rest on a ground-truth reference that is only locally validated.","tokens_in":20161,"tokens_out":2899,"would_cite":true,"duration_ms":30596,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a head-worn visual-inertial dataset of 58 indoor pedestrian-rich sequences and shows that two classical and two deep-learning localization systems all fail the combined accuracy, drift, and real-time requirements…","keywords":["visual SLAM","visual-inertial odometry","dataset","visually impaired navigation","crowded indoor environments","benchmark","dynamic scenes","ground-truth trajectories"],"falsifier":"Measure the full 5 km of trajectories against an independent survey-grade reference, for example a total station or fixed motion-capture system covering the long and crowded sequences, and compare; if the offline SLAM ground truth deviates by well over 2 cm in precisely the scenes where the tested systems fail, the benchmark's error numbers are biased and the failure conclusion would need to be re-tested.","tokens_in":19093,"feed_emoji":"🚶","tokens_out":9345,"duration_ms":82872,"temperature":0.7,"pith_summary":"Simultaneous localization and mapping (SLAM) is a candidate technology for guiding blind and visually impaired people, but it has rarely been tested in the crowded indoor spaces where those users actually walk. InCrowd-VI is a new dataset built for exactly that test: 58 head-worn recordings totaling about 5 km and 1.5 hours across airports, train stations, museums, malls, and university buildings, with RGB, stereo, IMU data, semi-dense 3D point clouds, and trajectory ground truth accurate to roughly 2 cm. The paper evaluates two classical and two deep-learning visual odometry and SLAM systems on it and finds that none meets the combined requirements of 0.5 meter localization accuracy, under 1% drift, high pose coverage, and real-time processing at walking speed. Classical systems drift up to 5-10% of trajectory length in crowds, changing light, and motion transitions, while deep-learning systems keep pose coverage above 90% but cannot process frames fast enough. The dataset is offered as a public benchmark to push SLAM research toward robust, real-time navigation for visually impaired users.","feed_headline":"Crowded indoor scenes trip up every tested navigation system","feed_subtitle":"A head-worn, 58-sequence dataset of airports and malls shows none meets the accuracy-speed bar for blind pedestrians.","key_machinery":"The load-bearing object is the dataset itself, together with its ground-truth generation pipeline. The sensor rig is a head-worn stereo pair with wide field-of-view monocular cameras, a higher-resolution RGB camera, and dual IMUs; InCrowd-VI releases the stereo image pairs, RGB frames, IMU streams at 1000 and 800 Hz, semi-dense point clouds, and per-sequence trajectory ground truth. The ground truth is produced by an offline SLAM service that fuses all of the device's sensors with post-processing, conditions that are deliberately impractical for a real-time wearable, and the paper validates it by measuring distances between recognizable landmarks in the reconstructed point cloud and comparing them with physical measurements, reporting an average error of about 2 cm. The evaluation harness uses four metrics: absolute trajectory error as root mean square error, drift percentage defined as trajectory error divided by path length, pose estimation coverage as the fraction of frames with an estimated pose, and a real-time factor comparing the distance processed per second with the user's walking speed. These metrics are what carry the conclusion that every tested system fails at least one requirement.","core_discovery":"The central claim is that crowded indoor scenes are precisely where current visual-inertial SLAM breaks, and that the lack of a benchmark reflecting those conditions has been hiding the gap. InCrowd-VI supplies that benchmark: 58 sequences recorded from a walking person's head-mounted stereo camera pair, covering crowd densities from empty to more than ten pedestrians per frame, with occlusions, reflective and glass surfaces, flickering or dim light, stairs, escalators, and moving ramps. Ground-truth trajectories come from an offline, post-processed machine-perception SLAM service that uses the full sensor suite of the recording glasses and removes dynamic pedestrians from the map; manual comparisons of landmark distances in the reconstructed point cloud against real-world measurements give a mean deviation of about 2 cm. On the dataset, two classical and two deep-learning visual odometry and SLAM systems all miss the stated combined bar somewhere: the worst classical runs drift 21-50% of trajectory length, and the learning-based systems keep pose coverage above 90% but run several times slower than real time at walking pace. The paper concludes that no tested system is ready for reliable visually impaired navigation in complex indoor pedestrian-rich spaces.","pith_inferences":["Because the reference trajectories are themselves produced by an offline SLAM service, InCrowd-VI is strongest for comparing systems against each other; absolute accuracy claims inherit whatever errors that service has in scenes where manual validation was not performed.","The 0.5 m and 1% thresholds are application choices; a different task, such as low-speed assistive navigation in familiar buildings, might tolerate larger errors, so the failure verdict should not be over-generalized.","A natural testable extension is per-frame failure annotation, such as occlusion fraction, motion blur, and textureness, which would turn the dataset from a pass-fail benchmark into a diagnostic tool for why tracking is lost.","The dataset's average walking speed of 0.75 m/s, below typical sighted walking speed, slightly relaxes the real-time requirement; a system that fails the real-time factor here could still be adequate for slower users."],"forward_implications":["SLAM researchers get a public benchmark whose crowd densities, lighting changes, and motion transitions reproduce realistic indoor navigation rather than lab or vehicle settings.","Classical feature-based and semi-direct systems can no longer claim robustness on the basis of empty indoor or outdoor datasets, because on InCrowd-VI their drift frequently exceeds 5-10% of trajectory length.","Deep-learning systems, despite high pose coverage, must become substantially faster before they can serve as real-time navigation aids for walking users.","Application designers for visually impaired navigation should treat the 0.5 m, 1% drift, and real-time requirements as a concrete target, since no evaluated system meets them together on the hard sequences.","The dataset's lack of depth data and its indoor-only scope define natural next steps: depth-based SLAM and indoor-outdoor transitions cannot be tested with it."],"supporting_citations":[{"why":"Supplies the ground-truth trajectories and semi-dense point clouds through an offline SLAM service; every accuracy number in the paper is measured against it.","marker":"[7]"},{"why":"Motivates the manual landmark-distance validation of the ground truth via joint map-and-trajectory optimization.","marker":"[30]"},{"why":"Provides the walking-speed evidence for visually impaired pedestrians that sets the dataset's average speed and the real-time requirement.","marker":"[31]"},{"why":"One of the two deep-learning SLAM systems evaluated; its pose-coverage and speed results support the robustness-versus-realtime finding.","marker":"[32]"},{"why":"The other deep-learning SLAM system evaluated; contributes the same coverage-and-speed evidence.","marker":"[33]"},{"why":"One of the two classical systems evaluated; its drift numbers in crowds and light changes support the failure conclusion.","marker":"[34]"},{"why":"The other classical system evaluated; its drift values show classical methods are not robust on the dataset.","marker":"[35]"},{"why":"Defines the absolute trajectory error metric used to score every system against the accuracy criterion.","marker":"[37]"}],"fun_headline_variants":["New benchmark shows all SLAM systems fail in crowded indoor scenes","InCrowd-VI dataset: no tested system passes the bar for blind navigation","All current SLAM systems stumble in pedestrian-rich indoor spaces","Head-mounted dataset proves no SLAM system is ready for crowded indoor navigation","New dataset shows visual-inertial SLAM fails in pedestrian crowds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ground-truth trajectories come from an offline localization service that is itself a SLAM system, and its claimed accuracy of about 2 cm was checked by manual measurements on only part of the data; if that service's errors grow in long or heavily crowded sequences, every reported error and failure verdict would be biased.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark shows all SLAM systems fail in crowded indoor scenes","InCrowd-VI dataset: no tested system passes the bar for blind navigation","All current SLAM systems stumble in pedestrian-rich indoor spaces","Head-mounted dataset proves no SLAM system is ready for crowded indoor navigation","New dataset shows visual-inertial SLAM fails in pedestrian crowds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001239,"raw_usage":{"total_tokens":5172,"prompt_tokens":1117,"completion_tokens":4055,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":733,"completion_tokens_details":{"reasoning_tokens":3962}},"tokens_in":733,"tokens_out":4055,"duration_ms":26254,"temperature":1.0,"reasoning_tokens":3962,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:15:12.696292+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the full 5 km of trajectories against an independent survey-grade reference, for example a total station or fixed motion-capture system covering the long and crowded sequences, and compare; if the offline SLAM ground truth deviates by well over 2 cm in precisely the scenes where the tested systems fail, the benchmark's error numbers are biased and the failure conclusion would need to be re-tested.","supporting_citations":[{"cited_title":"Bad slam: Bundle adjusted direct rgb-d slam","cited_arxiv_id":null,"evidence_quote":"Motivates the manual landmark-distance validation of the ground truth via joint map-and-trajectory optimization."},{"cited_title":"Walking biomechanics and energetics of individuals with a visual impairment: a preliminary report","cited_arxiv_id":null,"evidence_quote":"Provides the walking-speed evidence for visually impaired pedestrians that sets the dataset's average speed and the real-time requirement."},{"cited_title":"Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras","cited_arxiv_id":null,"evidence_quote":"One of the two deep-learning SLAM systems evaluated; its pose-coverage and speed results support the robustness-versus-realtime finding."},{"cited_title":"Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam","cited_arxiv_id":null,"evidence_quote":"One of the two classical systems evaluated; its drift numbers in crowds and light changes support the failure conclusion."},{"cited_title":"SVO: Semidirect visual odometry for monocular and multicamera systems","cited_arxiv_id":null,"evidence_quote":"The other classical system evaluated; its drift values show classical methods are not robust on the dataset."},{"cited_title":"A benchmark for the evaluation of RGB-D SLAM systems","cited_arxiv_id":null,"evidence_quote":"Defines the absolute trajectory error metric used to score every system against the accuracy criterion."}],"review_version":1}