{"id":"91e4ce46-6639-4a7d-8ea4-4e5ff59402bf","arxiv_id":"2608.12860","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"HumanoidVLN provides a physics-based benchmark for VLN on diverse humanoid robots, with 933 episodes and a 20-episode sim-to-real pilot showing strong alignment.","lead":"HumanoidVLN is a new physics-grounded simulator and benchmark that tests vision-language navigation (VLN) across four different humanoid robots in 87 indoor scenes. It reports JanusVLN as the best model (43.55% success rate) and a small sim-to-real study with strong correlation in navigation errors.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Only the Unitree G1 is sim-to-real validated, and only on endpoint error; H1 and the two internal robots are unvalidated, so the cross-embodiment fall-rate rankings may be simulation artifacts.","rationale":"The reader's weakest assumption identifies exactly the same risk: only G1 is validated, and only with a small pilot, while H1 and internal robots are not. My analysis agrees and sharpens it: the paper's distinctive metric is Fall Rate, and that metric is not validated even for G1. This matters because the largest cross-embodiment differences in Table III are fall rates on H1 (64.5% and 71.0%), and those differences underpin the claim that embodiment configuration materially affects physical VLN performance. Since the simulator and control stack are the only source of evidence for these numbers, a dynamics mismatch for H1 or the internal robots would change the rankings, not just a constant offset. The paper is internally coherent, and the pilot is genuine evidence for G1 endpoint behavior, so this is not a rejection-level flaw; it is an addressable validation gap. I would keep the conditional verdict: acceptance should require either an H1 sim-to-real fall-rate check or a clear release of the internal-robot assets plus a statement that their real-world fidelity is unverified. This does not change the reader's conditional disposition.","tokens_in":12251,"tokens_out":7307,"duration_ms":71665,"concrete_test":"Run a 20-episode sim-to-real pilot on the Unitree H1, the only public unvalidated embodiment, using the same DualVLN checkpoint and the same two scenes as Sec VI-D, and report per-episode fall occurrence under the paper's T1-T3 criteria. If H1's simulated fall rate is above 50% while the real robot falls in fewer than 20% of matched episodes, the Table III cross-embodiment rankings are driven by simulator dynamics rather than embodied navigation ability, and the physical-fidelity claim for H1 does not hold. If fall rates agree within a small tolerance, that specific concern is resolved; the internal robots would still require either release of specifications or hardware data before their rankings can be trusted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that HumanoidVLN provides a physically faithful testbed whose cross-embodiment rankings are meaningful. The only real-world evidence is the 20-episode G1 pilot in Sec VI-D, and it validates endpoint navigation error and trajectory nDTW for one model (DualVLN) in two scenes. It does not validate the Fall Rate metric, the most physics-dependent outcome in Table III, nor any other embodiment. Unitree H1 is reported with fall rates of 64.52% and 70.95% for StreamVLN and NaVILA, while the other three embodiments are below 10%; if H1's simulated dynamics or the simulation-trained locomotion policy are even moderately off, those dramatic differences are artifacts rather than embodied facts. Internal-A and Internal-B have no hardware counterpart, and Table II withholds their specifications for double-blind review, so their ranking (Internal-A highest average SR; JanusVLN best on it) cannot be independently checked or reproduced. The paper's own limitation paragraph lists scene diversity, annotation scaling, and compute cost, but omits the unvalidated dynamics of three of four embodiments, which is the load-bearing risk for the physical-fidelity claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HumanoidVLN, a physics-grounded simulator and benchmark for vision-language navigation with bipedal humanoid robots. Built on NVIDIA Isaac Sim, it provides a hierarchical control stack (a per-embodiment RL locomotion policy with interchangeable PD or MPC path trackers) and supports four embodiments: Unitree G1, Unitree H1, Internal-A, and Internal-B. The benchmark contains 87 scenes, filtered for at least 100 m^2 of navigable area, and 933 episodes with instructions produced by a multi-agent generator-reviewer-paraphraser pipeline with human-in-the-loop verification. Four VLN models (NaVILA, StreamVLN, DualVLN, JanusVLN) are evaluated zero-shot across the four embodiments, with JanusVLN reporting the highest mean SR of 43.55%. A 20-episode sim-to-real pilot with DualVLN on the Unitree G1 reports a strong correlation in endpoint navigation error (r = 0.935) and mean trajectory similarity of 0.782 nDTW. The central claim is that the platform provides physically executable, embodiment-aware evaluation that reveals meaningful cross-embodiment differences in navigation performance and fall rate.","tokens_in":12537,"tokens_out":6078,"duration_ms":64903,"significance":"If the physical-fidelity claim holds, HumanoidVLN addresses a genuine gap in the VLN benchmark literature: existing simulators largely rely on kinematic teleportation and do not model bipedal locomotion constraints or morphology-dependent camera dynamics. The paper has notable strengths: the evaluation is zero-shot with public checkpoints, the proposed benchmark includes human-verified instructions, the architecture is designed to be extensible to new robots and models, and the sim-to-real pilot provides a concrete, falsifiable check on the reconstructed-scene pipeline. The reported cross-embodiment fall-rate differences, if validated, would be an important new axis of evaluation for humanoid VLN. However, the current evidence supports physical fidelity for only one embodiment and only for endpoint/trajectory metrics, while the headline cross-embodiment fall rates involve robots whose simulated dynamics have not been validated against hardware. The benchmark's central contribution therefore depends on validation and analysis that is not yet present in the manuscript.","major_comments":[{"comment":"The episode sampling rule resamples any path that \"cannot be completed stably,\" but the manuscript does not report how often resampling occurs, for which embodiments, or how resampling changes the distribution of path lengths, turning frequency, and task difficulty. Because Table III compares models across four embodiments, an embodiment-specific resampling process can confound the reported rankings: the episode sets may differ across robots in difficulty, so the observed SR and fall-rate differences may reflect sampling bias rather than morphology. The authors should report per-embodiment resampling rates and compare the final episode distributions across embodiments, and ideally construct a single common episode set verified as executable by all four robots, with a sensitivity analysis showing that the Table III rankings are stable under the choice of episode set.","section":"Sec. IV-C"},{"comment":"The sim-to-real pilot validates only the Unitree G1, only endpoint navigation error and trajectory nDTW, and only two scenes with one model. It does not validate the Fall Rate metric defined in Eq. (1) or any of the other three embodiments. This is a load-bearing gap because Table III reports extreme fall rates for Unitree H1 (70.95% for NaVILA, 64.52% for StreamVLN) while the other embodiments remain below 10%; if H1's simulated dynamics, its locomotion policy, or the fall-detection thresholds are even moderately inaccurate, those dramatic differences could be simulation artifacts rather than embodied facts. The conclusion's limitation paragraph lists scene diversity, human verification, and compute cost, but omits the unvalidated dynamics of three of the four embodiments. The authors should add real-robot fall-rate validation for at least one additional embodiment, or provide a systematic dynamics sensitivity analysis (e.g., friction, mass distribution, actuator gains, control latency) with confidence intervals for the Table III fall rates, and they should release the full H1 and Internal-A/B specifications currently withheld in Table II so the experiments can be reproduced.","section":"Sec. VI-D and Table III"},{"comment":"The 3DGS reconstruction quality assessment is purely qualitative: Fig. 4 shows normal maps, but no quantitative metrics are reported. Since reconstructed scenes form part of the benchmark, and since the collision meshes derived from 3DGS directly affect physical executability, the claim that the reconstructions are \"simulation-ready\" needs quantitative support. The authors should report reconstruction accuracy (for example, depth error against reference scans or mesh accuracy), and should compare physical-execution statistics such as fall rate or footstep collision rate on reconstructed versus artist-designed scenes. Without this, Q3's conclusion is not supported by the evidence presented.","section":"Sec. VI-C"}],"minor_comments":[{"comment":"The manuscript says one instruction per episode is selected using a \"fixed, approximately balanced assignment\" across the four styles, but it does not specify the exact assignment mechanism or release the mapping; please provide the complete procedure so that the evaluation can be reproduced exactly.","section":"Sec. V-C"},{"comment":"Withholding the specifications of Internal-A and Internal-B for double-blind review is understandable during reviewing, but the final version must include full specifications (or a supplement), because a benchmark episode set and cross-embodiment comparisons cannot be independently reconstructed without them.","section":"Table II"},{"comment":"The caption states that the full pipeline \"yields the most coherent geometry,\" but the figure is qualitative; adding a quantitative caption metric or referring to a table of reconstruction errors would make the claim verifiable.","section":"Fig. 4"},{"comment":"The fall rates are computed over 933 episodes, but no confidence intervals or standard errors are reported; some of the differences highlighted in the text (for example, 7.93% versus 9.54% for G1) may be within sampling noise, so interval estimates are important for interpreting the rankings.","section":"Table III / Fig. 3"},{"comment":"The sentence \"paths that cannot be completed stably are resampled\" should state the resampling budget and the number of rejected samples, as also requested in Major Comment 1; this will help readers assess the cost and possible bias of the procedure.","section":"Sec. IV-C"},{"comment":"Please proofread the bibliography entries for consistency, as at least one entry ([24]) appears to have an inconsistent author list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable benchmark contribution, but the central physical-fidelity claim is currently supported only for one embodiment and one outcome metric. The requested revisions—resampling-bias quantification, additional real-robot or sensitivity validation of fall rates, and full disclosure of robot specifications—are necessary before the cross-embodiment rankings can be taken as evidence about morphology. I do not see a novelty problem; the concern is the gap between the strength of the claims and the evidence provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate new benchmark for VLN on humanoids, with real work behind it, and it deserves peer review. The main risk is that the cross-embodiment results—especially the H1 fall rates—may rest on unvalidated simulation dynamics.\n\nWhat's new: the combination of an Isaac Sim humanoid platform with four embodiments (G1, H1, two internal), navigability-curated 3DGS scenes, and a multi-agent annotation pipeline with human verification. No prior VLN benchmark does all three. The zero-shot evaluation with public checkpoints is clean: there's no evidence the models were fitted to this data. The MAA instruction generation and the 933-episode dataset look like solid engineering. The 20-episode G1 sim-to-real pilot is genuinely useful evidence, even if small: an r=0.935 correlation on endpoint error plus 0.78 nDTW shows the reconstructions preserve difficulty structure.\n\nSoft spots, in proportion. The biggest is the one the paper underplays: three of the four embodiments have zero real-world validation. Only G1 is checked, and only on endpoint error and trajectory similarity—not on fall rate. So the headline cross-embodiment difference (H1 falls in 64-71% of episodes, everyone else under 10%) could easily be a simulation artifact. If H1's simulated dynamics or its RL policy is off, those numbers don't transfer. The paper says this in passing in the limitations paragraph, but its own summary lists scene diversity and compute cost, not the unvalidated dynamics. That's a load-bearing omission.\n\nSecond, episode resampling (Sec IV-C: paths that cannot be completed stably are resampled) biases the reference set toward what each robot can actually do. That's defensible for building a usable benchmark, but it also means cross-embodiment comparisons are not apples-to-apples on the same route distribution. Worth disclosing the resampling rate.\n\nThird, Table III has no error bars or repeated trials. Some differences (e.g., Internal-A vs G1 SR) are within plausible noise. The 3DGS geometry quality is only qualitatively ablated. Minor point: the internal robot specs are withheld for double-blind, which makes reproduction impossible until release.\n\nNet: the central idea—physics-grounded, embodiment-diverse humanoid VLN—holds up. The execution is credible, and the public checkpoints plus the G1 pilot show the authors aren't gaming the benchmark. A serious referee could help tighten the validation story, but this should not be desk-rejected. Send it to review. Reader gets: benchmark builders, VLN researchers, legged locomotion people. I'd cite it once it's public.","headline":"A solid new humanoid VLN benchmark with real zero-shot evaluation and an honest pilot sim-to-real check; the main risk is that three of four embodiments are unvalidated, so the dramatic H1 fall rates may be simulation artifacts.","tokens_in":13049,"tokens_out":1750,"would_cite":true,"duration_ms":16643,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HumanoidVLN replaces kinematic teleportation with full bipedal physics, creating a 933-episode vision-language navigation benchmark in which every scored trajectory is physically executable by the robot under test, and shows that model…","keywords":["vision-language navigation","humanoid robots","physics simulation","embodied AI","sim-to-real transfer","3D Gaussian Splatting","leg locomotion","benchmark"],"falsifier":"Run the same evaluation episodes on the physical versions of the three robots that were never validated against reality; if their per-episode endpoint errors and fall events do not align with simulation as closely as the pilot's $r = 0.935$ and $0.68$ m gap, then the guarantee of physical executability does not transfer across embodiments.","tokens_in":12105,"feed_emoji":"🤖","tokens_out":11765,"duration_ms":112361,"temperature":0.7,"pith_summary":"This paper aims to prove that vision-language navigation (VLN) for humanoid robots — instructing a walking robot to follow natural-language directions — cannot be judged fairly on simulators that move the agent by teleporting it along a path, because a bipedal body changes which paths are passable, how much the camera shakes, and whether the robot falls. To test that claim, it builds a physics-based simulation and a 933-episode benchmark in which four humanoid robots, spanning heights of $1.17$–$1.80$ m and 10–12 lower-body degrees of freedom, execute every episode under a hierarchical controller that respects joint limits and gait stability. The paper reports that replacing kinematic stepping with real dynamics reshuffles the rankings of four VLN models across embodiments, with success rates from about 14% to 50% depending on the robot, and that a 20-episode sim-to-real pilot on one robot shows endpoint errors correlating at $r = 0.935$ with a mean absolute gap of $0.68$ m. If this holds, the benchmark offers a way to select both navigation models and robot morphologies before hardware deployment.","feed_headline":"Walking, not teleporting, reshuffles robot navigation leaderboards","feed_subtitle":"Robot height and gait change which navigation models work; simulated errors track real ones at 0.68 m.","key_machinery":"The load-bearing object is the two-level control hierarchy that replaces teleportation: a per-embodiment reinforcement-learning locomotion policy, which commands joint torques under joint limits and center-of-mass constraints, driven by interchangeable high-level path trackers (proportional–derivative control for discrete-action models, model-predictive control for continuous-action models). This stack is what the paper invokes to guarantee that every evaluated trajectory is physically executable by the specific robot under test. Around it, two supporting mechanisms carry the benchmark: a scene pipeline that reconstructs real spaces with 3D Gaussian splatting, enforces depth–normal consistency and unbiased depth, fuses collision meshes, and keeps only scenes with at least $100\\,\\mathrm{m}^2$ of navigable floor; and a multi-agent instruction pipeline in which two generators independently construct route graphs from egocentric video, a reviewer verifies them against trajectory and scene-graph priors, a paraphraser produces three style variants, and humans correct the result.","core_discovery":"On the paper's own terms, the central discovery is that physical executability is not a detail of humanoid VLN but a first-order determinant of measured performance. By running four VLN models zero-shot on identical 933 episodes across four humanoid bodies, the paper finds that model and embodiment interact strongly: the model with explicit 3D spatial memory reaches the highest mean success rate ($43.55\\%$) and path fidelity (normalized dynamic time warping, nDTW, of $48.38$), while the tallest, 10-DoF robot lowers average success rates to about $21\\%$ and produces fall rates above $70\\%$ for two models — a failure mode that cannot appear in a kinematic simulator. A second component of the discovery is that its reconstructed scenes, built from 3D Gaussian splatting with extracted collision meshes, preserve enough of the real environment that 20 paired episodes of navigation in simulation and reality give strongly correlated endpoint errors ($r = 0.935$) and mean trajectory similarity of $0.782 \\pm 0.188$ nDTW. The paper takes these two results together as evidence that its benchmark measures humanoid VLN rather than a proxy for it.","pith_inferences":["It follows, though the paper does not develop it, that the fall detector could serve as a training signal: optimizing a VLN policy jointly for instruction success and locomotion stability would likely close the gap between the most accurate model and the most stable model.","The $0.68$ m average endpoint gap could be used as a calibration offset to predict real-world success from simulated success for future robots, but only if the correlation replicates across more models and scenes.","The $100\\,\\mathrm{m}^2$ navigability filter implies that future humanoid VLN benchmarks should report traversable floor area per scene as standard metadata, or comparisons across legged platforms will be confounded by scene topology.","Scaling the pilot to multiple models and scenes is the direct test that would generalize the correlation; the paper's own stated limitation is that the current evidence is one model, two scenes, and 20 paired episodes."],"forward_implications":["Model rankings are embodiment-dependent: the same VLN checkpoint can move from best to worst across robots, so a single canonical humanoid is insufficient for benchmarking.","Kinematic simulators overstate humanoid navigation performance: they cannot register the tall robot's fall rates of $64.5\\%$ and $71.0\\%$ or the resulting loss of success.","Fall rate becomes a reportable VLN metric, capturing gait stability under language-guided control alongside success and path fidelity.","Reconstructed scenes can substitute for artist-authored environments in physical VLN evaluation, since the pilot shows endpoint errors transfer within $0.68$ m on average.","The benchmark's fixed 933-episode zero-shot evaluation set allows off-the-shelf VLN checkpoints to be compared across embodiments without any training on the benchmark."],"supporting_citations":[{"why":"defines the original VLN path-following task on viewpoint graphs that this work argues is not sufficient for legged robots.","marker":"[3]"},{"why":"provides the dominant kinematic simulator whose teleport-style action model is the physical gap the paper targets.","marker":"[10]"},{"why":"established the continuous-space VLN evaluation protocol that the benchmark adapts to walking robots.","marker":"[22]"},{"why":"first exposed locomotion-induced VLN failure modes in a rigid-body simulator, motivating embodiment-specific control.","marker":"[13]"},{"why":"supplies the full-kinematics simulator and the describer–verifier–synthesizer annotation pattern this work extends to humanoids.","marker":"[9]"},{"why":"introduces 3D Gaussian Splatting as the scene representation used for photorealistic reconstruction.","marker":"[8]"},{"why":"provides the Gaussian splatting training implementation used to build reconstructed scenes.","marker":"[26]"},{"why":"adds depth–normal consistency constraints that sharpen surfaces for collision mesh extraction.","marker":"[27]"},{"why":"contributes unbiased depth rendering, which the paper uses to extract high-quality meshes.","marker":"[30]"},{"why":"gives TSDF fusion, which converts rendered depth maps into the simulation's collision meshes.","marker":"[31]"}],"fun_headline_variants":["Humanoid body shape reshuffles VLN model rankings","Taller robots trip up navigation models; sim matches real","Physics-based simulator links robot gait to VLN success","For VLN on humans, embodiment trumps model choice","Sim-to-real correlation 0.935 in humanoid vision-language nav"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The cross-embodiment rankings rest on the simulator's dynamics matching all four robots faithfully, but the paper checks only one robot against the real world, and only on 20 episodes in two scenes.","fun_headline_variants_meta":{"raw":{"variants":["Humanoid body shape reshuffles VLN model rankings","Taller robots trip up navigation models; sim matches real","Physics-based simulator links robot gait to VLN success","For VLN on humans, embodiment trumps model choice","Sim-to-real correlation 0.935 in humanoid vision-language nav"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":2051,"prompt_tokens":1184,"completion_tokens":867,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":800,"completion_tokens_details":{"reasoning_tokens":784}},"tokens_in":800,"tokens_out":867,"duration_ms":7948,"temperature":1.0,"reasoning_tokens":784,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:48:30.793757+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same evaluation episodes on the physical versions of the three robots that were never validated against reality; if their per-episode endpoint errors and fall events do not align with simulation as closely as the pilot's $r = 0.935$ and $0.68$ m gap, then the guarantee of physical executability does not transfer across embodiments.","supporting_citations":[{"cited_title":"Vision-and-language nav- igation: Interpreting visually-grounded navigation instructions in real environments,","cited_arxiv_id":null,"evidence_quote":"defines the original VLN path-following task on viewpoint graphs that this work argues is not sufficient for legged robots."},{"cited_title":"Habitat: A platform for embodied ai research,","cited_arxiv_id":null,"evidence_quote":"provides the dominant kinematic simulator whose teleport-style action model is the physical gap the paper targets."},{"cited_title":"Beyond the nav-graph: Vision-and-language navigation in continuous environ- ments,","cited_arxiv_id":null,"evidence_quote":"established the continuous-space VLN evaluation protocol that the benchmark adapts to walking robots."},{"cited_title":"Rethinking the embodied gap in vision- and-language navigation: A holistic study of physical and visual disparities,","cited_arxiv_id":null,"evidence_quote":"first exposed locomotion-induced VLN failure modes in a rigid-body simulator, motivating embodiment-specific control."},{"cited_title":"3d gaussian splatting for real-time radiance field rendering.,","cited_arxiv_id":null,"evidence_quote":"introduces 3D Gaussian Splatting as the scene representation used for photorealistic reconstruction."},{"cited_title":"2d gaussian splatting for geometrically accurate radiance fields,","cited_arxiv_id":null,"evidence_quote":"adds depth–normal consistency constraints that sharpen surfaces for collision mesh extraction."},{"cited_title":"Pgsr: Planar-based gaussian splatting for efficient and high-fidelity surface reconstruction,","cited_arxiv_id":null,"evidence_quote":"contributes unbiased depth rendering, which the paper uses to extract high-quality meshes."},{"cited_title":"Kinectfusion: Real-time dense surface mapping and tracking,","cited_arxiv_id":null,"evidence_quote":"gives TSDF fusion, which converts rendered depth maps into the simulation's collision meshes."}],"review_version":1}