{"id":"a2b68ec7-dc46-4fed-b8ea-9e96a10c3311","arxiv_id":"2607.21964","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"ACME is a socially navigated robot and pedestrian trajectory dataset covering 8 sites in 5 countries with 7 robot embodiments, including human-verified BEV tracks and robot speech annotations.","lead":"This paper introduces ACME, a large multi-country, multi-robot dataset for social navigation and pedestrian trajectory prediction, with 29.35 hours of robot data and 43.5 hours of overhead pedestrian tracks. A generalist might read it as a new benchmark resource for training and evaluating robots that must move politely through crowds in different cultures.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The BEV-to-metric homographies are unvalidated; without a reprojection-error analysis, the cross-cultural and trajectory-prediction comparisons supporting the 'more challenging/broader behavior' claim are not yet falsifiable.","rationale":"The paper's central claim is that ACME is the largest and most diverse human-demonstrated social-navigation and pedestrian-trajectory dataset and that it captures more challenging scenarios and a broader distribution of pedestrian behavior than prior datasets. The dataset-size and diversity assertions are internally consistent with the reported numbers, and the qualitative descriptions are plausible. However, the evidence for 'broader distribution of pedestrian behavior' and for the quantitative benchmarking rests directly on metric-space trajectories derived from BEV homographies. The paper acknowledges homography inaccuracy but never quantifies it, and for two sites the BEV/onboard synchronization is manual. This is a genuine soft spot rather than a manufactured one: the cross-cultural analyses (passing side, personal space) and the claim that ACME is 'more challenging' for trajectory predictors both depend on the assumption that the homography-induced error is small relative to the reported behavioral differences and prediction errors. The proposed test—ground-control-point reprojection followed by a robustness rerun of the interaction analyses—would settle whether this assumption holds. If the behavioral patterns are stable under a low-error subset, the conditional acceptance is justified; if not, the metric-space contributions would need to be reframed as pixel-space annotations with approximate calibration. Because the issue is addressable and the reader already conditioned acceptance on verification, no verdict change is needed.","tokens_in":27866,"tokens_out":11359,"duration_ms":123655,"concrete_test":"At each of the 8 sites, place 15–20 surveyed ground control points spanning the BEV camera field of view (or use existing tile intersections for CMU/U-Mich), and compute the homography reprojection RMSE and maximum error in meters. Then rerun the interaction-point analyses of Figs. 17–18 using only trajectory portions where the reprojection error is below 0.3 m (or 0.5 m) and compare. If the US/Germany right-side passing preference and the reported frontal personal-space differences persist under the low-error subset, the concern is mitigated; if they shift or vanish, the homography inaccuracy is load-bearing and the metric-space utility claims require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's quantitative value hinges on the BEV-to-metric homographies (section 'Transformation from BEV Image Plane to Ground Plane') and, for Miraikan/Keio, on manual frame-selection synchronization (Appendix). The paper explicitly acknowledges 'the inherent inaccuracies of tag-based homography estimation' but provides no reprojection-error measurements, no ground-truth comparison, and no sensitivity analysis. This is load-bearing for three reasons. First, the cross-cultural comparisons in Figs. 17–18 (passing side, personal space) are computed from metric interaction points; small heading or distance errors at the image periphery could create or erase the reported left/right preferences and 0.5–1 m personal-space differences. Second, the trajectory-prediction benchmark (Table 6) uses these metric trajectories, so coordinate noise directly inflates ADE/FDE and supports the claim that ACME is 'more challenging' than ETH/UCY. Third, the contribution bullet claiming metric-space trajectories is unvalidated. Without a quantitative error characterization, the central 'broader distribution of pedestrian behavior' claim cannot be independently assessed from the released materials.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ACME, a social-navigation and pedestrian-trajectory dataset collected by 8 teams at 8 sites in 5 countries using 7 robot embodiments. It comprises 29.35 hours of onboard robot data (odometry, RGB, LiDAR, and, for some subsets, robot speech) and 43.5 hours of overhead BEV pedestrian data, with 72.1K human-verified pedestrian trajectories provided in pixel and metric coordinates. The authors compare ACME with SCAND and MuSoHu on pedestrian density, traversability, and planner failure rates; benchmark vision-navigation models (ViNT, NoMAD) and trajectory-prediction models (SocialGAN, AgentFormer, SGNet, TUTR, MoFlow); and use the BEV data to analyze cross-cultural differences in walking speed, passing side, and personal space. The headline claims are that ACME is the largest and most diverse human-demonstrated social-navigation and trajectory-prediction dataset and that it captures more challenging scenarios and a broader distribution of pedestrian behavior than prior datasets.","tokens_in":28203,"tokens_out":7909,"duration_ms":89623,"significance":"If the claims hold, ACME would be a substantial community resource. Its strengths include multi-country and multi-embodiment coverage, a large human-verified BEV trajectory set with metric coordinates, scenario and semantic tags, robot-speech annotations, and a reproducible-style pipeline with open-sourced annotation tools and processing scripts. The paper also provides initial benchmarks for both navigation and trajectory-prediction models, which can serve as baselines for future work. The multi-site IRB/ethics documentation is a positive example for dataset papers. The main risk is that the quantitative value of the dataset depends on unvalidated homography and synchronization procedures, and the 'more challenging' claim relies on proxy metrics whose definitions are not fully specified. These issues are fixable and should be addressed before publication.","major_comments":[{"comment":"Metric-space trajectories are a central contribution, and they feed the cross-cultural analyses (Figs. 17-18) and the trajectory-prediction benchmark (Table 6). However, the paper provides no quantitative validation of the BEV-to-ground homographies. The text acknowledges 'the inherent inaccuracies of tag-based homography estimation' and later limits trajectories to within ±25 m of the poster tag, but no reprojection-error statistics, ground-truth comparisons, or distance-dependent sensitivity analyses are reported. In addition, the Miraikan/Keio synchronization uses manual frame selection (Appendix), whose temporal error is not quantified. Please report per-site reprojection RMSE and validate on surveyed points, and show that coordinate errors are small relative to the reported effects: ADE/FDE gaps of 0.1-0.38 m in Table 6 and personal-space differences of roughly 0.5-1 m in Fig. 18. W","section":"Transformation from BEV Image Plane to Ground Plane"},{"comment":"The claim that ACME is 'more challenging' than SCAND and MuSoHu relies heavily on the TEB planner failure rates shown in Fig. 13. However, the manuscript never defines what counts as a planner failure, reports raw failure counts, or provides confidence intervals. The statement that 'the primary mode of planner failure is the presence of pedestrians at the sampled goal position' is asserted but not substantiated with quantitative evidence. Because the comparison across datasets depends on local planner parameters, costmap configuration, and the failure criterion, these details are load-bearing. Please state the failure definition (e.g., no plan within timeout, oscillation, collision), report counts and rates per dataset, and include a sensitivity analysis with respect to planner settings.","section":"Comparison to a Geometric Planner (Fig. 13)"},{"comment":"The traversability metric is computed with 'a finetuned SAM2 model,' but the manuscript does not describe how that model was trained, what data it was trained on, how the traversability mask is defined, or how accurate its predictions are on the egocentric images of SCAND, MuSoHu, and ACME. The cited reference (Wang et al. 2024b) appears to be a navigation system rather than the segmentation model used. Since this proxy feeds Fig. 10(d)-(g) and the robot-speech analysis in Fig. 12, please document the model and validate its outputs on a manually annotated sample from each dataset, or otherwise present the traversability results as illustrative rather than as a basis for the 'more challenging' claim.","section":"Traversability analysis (Fig. 10)"},{"comment":"The passing-side and personal-space analyses are presented as evidence that ACME captures cross-cultural behavioral differences, but no confidence intervals, sample sizes per direction bin, or statistical tests are reported. The observed left/right preferences and personal-space differences could be influenced by per-site homography errors, within-tag variation in corridor widths, or small interaction counts in some direction bins. The paper already hedges that 'lateral differences should not be attributed too strongly to culture,' but the figures and surrounding text nevertheless draw directional conclusions. Please provide uncertainty quantification and, where appropriate, statistical tests, or explicitly frame these results as descriptive observations.","section":"Comparison across Datasets (Figs. 17-18)"}],"minor_comments":[{"comment":"The data accessibility statement says the dataset is shared 'via a password-protected Box in the following link:' but the link appears to be empty in the manuscript. Please include the URL, or at minimum a clear description of how reviewers can access the data.","section":"Data Accessibility Statement"},{"comment":"The term 'finetuned SAM2' is ambiguous. Please clarify whether this refers to the SAM2 segmentation model or to a traversability model from the cited Wang et al. (2024b) work, and add the appropriate citation.","section":"Traversability section"},{"comment":"The standard-deviation values are reported as '±' numbers without being defined in the caption or text. Please state explicitly that these are standard deviations, and consider adding the number of trajectories per sub-dataset in Fig. 15.","section":"Fig. 15 and Table 5"},{"comment":"The abstract refers to 'A Cross-cultural' dataset while the title uses 'Multi-Cultural.' Please standardize the terminology across the paper.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The core dataset contribution is likely valuable, but acceptance should be contingent on the public availability of the data and on the authors providing the requested calibration-validation and metric-definition details. The TBD overlap is transparent and does not raise concerns. The manuscript's central claims are defensible but require the additional evidence described in the major comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: ACME is a real resource, not a repackaging. 29.35 hours of onboard robot data, 43.5 hours of overhead pedestrian tracking, 7 embodiments, 5 countries, human-verified trajectories, robot speech annotations, scenario tags, and semantic BEV tags. If the public release lives up to the description, this is a step change over TBD, Bi3, and THOR-MAGNI on the axes that matter for social navigation and trajectory prediction. The data collection protocol is thoughtful: goal-directed trajectories, event markings, teleoperator guidelines, and honest appendix notes about sensor failures and calibration issues. The robot speech analysis is a genuine addition, and the observation that speech use is embodiment-dependent is a real finding.\n\nThe paper has three soft spots, and they are addressable. First, the BEV-to-metric homography is load-bearing and unvalidated. The paper acknowledges tag-based homography inaccuracies but gives no reprojection error, no ground-truth comparison, no sensitivity analysis. The cross-cultural figures (passing side, personal space) and the ADE/FDE numbers all ride on those metric coordinates. This is the weakest link. Second, the 'more challenging' claim rests on TEB failure rates and SAM2 traversability masks — model-dependent proxies. The 2x/4x failure-rate difference is suggestive but deserves a robustness check with another planner or a direct crowd-density measure. Third, the cross-cultural comparisons have no error bars or significance tests. The passing-side patterns may be real, but without confidence intervals they are observations, not conclusions. I'd also like to see the dataset actually released; the current private reviewer link is standard for review, but the permanence of the HF release matters for the paper's value.\n\nMinor notes: the 'largest / most diverse' claim is defensible relative to the cited prior work, and the self-citation to TBD is appropriate — TBD is the most relevant antecedent and the comparison is done fairly. The trajectory-prediction benchmark is fine, though the static-vs-dynamic split is a good check.\n\nWho this is for: anyone working on social navigation policies or pedestrian trajectory prediction who needs a diverse benchmark. It deserves a serious referee — the data collection effort and the honest documentation are worth engaging with. The revision should add homography error analysis and at least basic significance testing on the cultural comparisons.\n\nI'd send it to review.","headline":"ACME is a genuinely large, well-organized multi-country social-navigation dataset that deserves serious referee time, but its headline 'more challenging' and cross-cultural claims need homography validation and error bars before they carry full weight.","tokens_in":28842,"tokens_out":1519,"would_cite":true,"duration_ms":21167,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces ACME, a dataset of robot and pedestrian motion gathered at eight sites in five countries with seven robot platforms, and argues that it captures more challenging and more culturally varied social-navigation scenarios t","keywords":["social navigation","pedestrian trajectory prediction","multi-cultural dataset","multi-embodiment robot data","bird's-eye view tracking","human-verified annotations","robot speech","benchmarking"],"falsifier":"Place surveyed ground-truth markers at each collection site, project them through the published homographies, and compare to their measured positions; if typical errors exceed roughly half a meter, the metric-space trajectory statistics and cross-cultural comparisons would not be trustworthy.","tokens_in":27817,"feed_emoji":"🤖","tokens_out":5391,"duration_ms":57674,"temperature":0.7,"pith_summary":"ACME is a large-scale collection of real-world robot navigation and pedestrian movement data, assembled by eight teams in five countries using seven different robot bodies. The paper's central claim is that this is the largest and most diverse human-demonstrated dataset of its kind: 29.35 hours of onboard robot sensor data and 43.5 hours of overhead, human-verified pedestrian trajectories. Its purpose is to give social-navigation researchers data that includes cultural and morphological variation, which prior single-site datasets lack, so that robots can learn context-appropriate behavior rather than a narrow, location-specific norm. The paper also shows that state-of-the-art navigation and trajectory-prediction models perform noticeably worse on ACME than on older benchmarks, which it interprets as evidence that ACME contains harder, more realistic scenarios. If that is right, ACME becomes a benchmark for testing whether models generalize across cultures, environments, and robot forms.","feed_headline":"Largest social-navigation dataset spans 5 countries, 7 robot types","feed_subtitle":"29 hours of robot sensor data and 43 hours of verified pedestrian tracks give robots a harder, more cultural test.","key_machinery":"The central object is the dataset itself: synchronized multi-modal onboard recordings (odometry, egocentric RGB, 3D LiDAR, and timestamped robot speech in some subsets) plus overhead bird's-eye-view (BEV) cameras whose pedestrian tracks are converted from pixel coordinates to ground-plane metric coordinates via homography matrices. That conversion is the load-bearing mechanism for the cross-cultural trajectory analysis, turning overhead video into quantitative distances, speeds, and passing-side measurements. A second mechanism is the annotation pipeline: automated tracking followed by human verification at 10 Hz, with tools to relabel, add, break, join, and disentangle tracks, which support","core_discovery":"On its own terms, the paper's discovery is that a coordinated multi-site collection effort can produce a social-navigation dataset with broader cultural and morphological coverage than prior efforts, and that this breadth changes what models see. ACME robot trajectories are shorter and more goal-directed than earlier in-the-wild datasets; scenes are more socially constrained, with pedestrians closer to the robot; and off-the-shelf planners fail more often per meter. Pedestrian tracks show large variance in speed, density, and duration, and measurable cross-country differences in passing side and personal space (right-side passing in US and German data, left-side tendencies in Singapore and J","pith_inferences":["A direct test the paper does not run: if ACME delivers the diversity it claims, a navigation policy trained on a subset of its countries should transfer to a held-out country better than a policy trained on a single country; that experiment would separate the value of data volume from the value of cultural coverage.","The passing-side results imply that 'socially compliant' navigation has no single ground truth; deployed robots may need to infer or adapt to local conventions, which points toward learning algorithms that condition on geography or context rather than a universal policy.","The heavy presence of static and slow pedestrians in ACME suggests that trajectory predictors should model intent (standing, conversing, queueing) explicitly; the paper's dynamic-only benchmark hints at this, but the dataset's semantic tags make a stronger intent-aware training setup possible than the paper explores."],"forward_implications":["Trajectory-prediction models trained on long-standing benchmarks lose about 0.1 m ADE and 0.15 m FDE when evaluated on ACME, and more when only moving pedestrians are considered, making ACME a harder generalization benchmark.","Off-the-shelf geometric planners fail two to four times more often per minute or per meter on ACME than on earlier social-navigation datasets, making ACME useful for stress-testing planners in socially constrained scenes.","The dataset's scenario tags (narrow corridors, blind corners, crowds, queues, overtaking) let researchers filter for specific situations, so policies can be trained or evaluated on the long tail of social navigation rather than on averaged behavior.","Robot speech annotations tied to pedestrian density and traversability provide a first real-world resource for learning when a robot should speak to negotiate space, not just how it should move.","Because trajectories are grounded in metric coordinates and verified by humans, the dataset can support quantitative study of how walking speed, passing side, and personal space vary by country and environment layout."],"fun_headline_variants":["5 countries, 7 robots: new dataset tests social navigation","Cross-cultural robot dataset: 29h of tricky social scenes","Robot navigation dataset spans cultures, shows left-right passing splits","ACME dataset: 5 cultures, 7 embodiments, harder social tests","Multi-site robot data reveals culture-specific passing habits"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's own section on transforming BEV images to ground plane acknowledges 'the inherent inaccuracies of tag-based homography estimation'; that calibration accuracy, turning overhead camera pixels into ground-plane meters, is the load-bearing premise for the metric-space trajectory statistics and cross-cultural comparisons.","fun_headline_variants_meta":{"raw":{"variants":["5 countries, 7 robots: new dataset tests social navigation","Cross-cultural robot dataset: 29h of tricky social scenes","Robot navigation dataset spans cultures, shows left-right passing splits","ACME dataset: 5 cultures, 7 embodiments, harder social tests","Multi-site robot data reveals culture-specific passing habits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000874,"raw_usage":{"total_tokens":3619,"prompt_tokens":744,"completion_tokens":2875,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":2790}},"tokens_in":488,"tokens_out":2875,"duration_ms":21335,"temperature":1.0,"reasoning_tokens":2790,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T06:11:12.209253+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Place surveyed ground-truth markers at each collection site, project them through the published homographies, and compare to their measured positions; if typical errors exceed roughly half a meter, the metric-space trajectory statistics and cross-cultural comparisons would not be trustworthy.","supporting_citations":[],"review_version":1}