{"id":"6c1e2eb4-bd97-433a-9a02-daaa9f6f9be1","arxiv_id":"2412.01857","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"SALI adds a recurrently generated 'imagination' map of unvisited scenes to a real-observation memory, reporting state-of-the-art SPL on R2R and REVERIE.","lead":"This paper introduces SALI, a Vision-and-Language Navigation agent that stores imagined, generated images of unvisited locations in a persistent topological memory. The agent reports the best Success-weighted Path Length on R2R and REVERIE, though the authors' own ablation tables show conflicting SPL values for the headline result.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1's headline SPL=78 conflicts with ablation tables reporting 70/71 for the same configuration; the SoTA claim is unverifiable until this discrepancy is resolved.","rationale":"The paper proposes an interesting architecture: a persistent topological memory augmented by recurrently generated imagined RGB-D-semantic images, with pre-training tasks for the imagination module. The ablations in Tables 2-5 are suggestive and the qualitative results in Figure 6 are illustrative. However, the central SoTA claim in the Abstract and Table 1 is undermined by an internal numerical inconsistency. Table 1 reports val-unseen R2R SPL=78 with NE=1.92, OSR=86, SR=82, while Table 2 row 3 reports exactly the same NE, OSR, and SR but SPL=70, and Tables 3 and 4 report SPL=71 for the same or very similar configurations. This is not a subtle statistical issue; it is a direct contradiction in the paper's headline numbers. The reader's identified weakest assumption about imagination fidelity is real, but it is secondary: even if the imagined images were perfectly faithful, the claimed state-of-the-art result would be unsupported unless the SPL discrepancy is resolved. The concrete test is straightforward: evaluate the released checkpoint with standard metrics and compare. Until that is done, the appropriate verdict is conditional, matching the reader's assessment. I do not see grounds to reject the entire approach, since the ablations show a consistent internal trend (imagination helps), but the magnitude of the claimed improvement cannot be trusted. No ad hominem is intended; the issue is purely about the reproducibility of the reported numbers.","tokens_in":11838,"tokens_out":2340,"duration_ms":23753,"concrete_test":"Re-run the official R2R val-unseen evaluation for the exact fine-tuned SALI checkpoint used for Table 1, using the standard R2R metric code. Record per-trajectory SPL, SR, and TL. Check whether the produced SPL is 78 (Table 1) or 70/71 (Tables 2-4). Also verify Table 4 row 4: with Room+Waypoint and dynamic weighting, SR=80 implies SPL must be at most 80, so SPL=71 is internally plausible; confirm whether the SR in Table 1 should be 82 or 80. If the checkpoint is not public, require the authors to release code and checkpoints, or to provide per-trajectory predictions from the reported runs so the discrepancy can be recomputed independently.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that SALI achieves state-of-the-art SPL on R2R and REVERIE unseen splits. The most load-bearing evidence is Table 1, which reports val-unseen R2R SPL=78 with NE=1.92, OSR=86, and SR=82. The ablation study in Table 2 row 3 reports the same full model ('Reality + Imagination') with identical NE=1.92, OSR=86, SR=82, and TL=10.34, but SPL=70. Table 3 row 3 (M=2, Nbar=4) gives SR=82, SPL=71, and Table 4 row 4 gives SR=80, SPL=71. The same configuration cannot yield SPL 78 and 70/71 unless different checkpoints, seeds, evaluation protocols, or path-length normalizations are being used without disclosure. Since SPL is the headline metric and the claimed 8% and 4% improvements are computed from it, this internal inconsistency makes the headline result non-reproducible from the paper as written. If the true SPL is 70 or 71, SALI no longer surpasses ScaleVLN's SPL of 70 on R2R val-unseen, and the SoTA claim collapses. This concern is more decisive than the also-valid imagination-fidelity question, because even a perfect imagination module cannot support a SoTA claim built on irreconcilable reported numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SALI, a vision-and-language navigation (VLN) agent that combines a topological memory of real observations with a recurrently generated set of imagined future observations. The imagination module is composed of an inpaint model, a SPADE-based RGB generator, a room-type model, and a waypoint model, producing depth, semantic, and RGB images for potential future nodes. A multimodal transformer with graph-aware self-attention scores real and imagined nodes, and a learned dynamic fusion factor combines these scores for action selection. The authors report state-of-the-art results on R2R and REVERIE, with an 8% SPL improvement on R2R val-unseen and a 4% improvement on REVERIE unseen, and they present ablations showing consistent gains from the imagination and memory components. The central claim is that episodic simulation and episodic memory improve navigation generalization in unseen environments.","tokens_in":12188,"tokens_out":5453,"duration_ms":152168,"significance":"If the reported results are reliable, SALI would be a meaningful advance: it is among the first VLN agents to integrate recurrently generated future images into a persistent topological memory, and the ablation study suggests that both the imagination module and the memory mechanism contribute to navigation performance. The paper is also commendable for specifying its pre-training losses and for attempting quantitative evaluation of imagined images (PSNR, Pearson correlation). However, the central state-of-the-art claim currently rests on irreconcilable SPL numbers across tables, and no variance or seed information is provided. The significance of the contribution is therefore contingent on the authors resolving these reproducibility issues with exact evaluation protocols and per-seed results.","major_comments":[{"comment":"The headline R2R val-unseen SPL of 78 for SALI in Table 1 is inconsistent with the ablation tables for what appears to be the same configuration. Table 2 row 3 ('Reality + Imagination') reports NE=1.92, OSR=86, SR=82, and TL=10.34 but SPL=70. Table 3 row 3 (M=2, Nbar=4) reports SR=82 and SPL=71; Table 4 row 4 and Table 5 row 1 report SR=80 and SPL=71. All of these rows correspond to the full SALI model with imagination and dynamic fusion. Identical values of NE, OSR, and SR cannot concurrently yield SPL values of 78, 70, and 71 unless different checkpoints, seeds, or evaluation protocols are being used without disclosure. Since the abstract and introduction claim an 8% SPL improvement over ScaleVLN (SPL=70) and the ablation row implies a SPL of 70, the state-of-the-art claim is not reproducible from the paper as written. Please reconcile the numbers, state the exact configuration used for Table 1, and report per-seed results.","section":"§4.2, Table 1 vs. §4.4, Tables 2–5"},{"comment":"No error bars or multiple-seed results are reported anywhere in the experimental section. All tables present single deterministic numbers, with no standard deviation, number of runs, or statement that the benchmark evaluation is deterministic. Given the observed 8-point SPL discrepancy between Table 1 and the ablation tables, the statistical significance of the claimed improvement over prior work cannot be assessed. Please report mean and standard deviation over at least three seeds, or justify determinism of the evaluation protocol.","section":"§4.2, Tables 1–5"},{"comment":"The paper never states which environments are used to train the inpaint, room-type, SPADE, and waypoint models. The central claim is that SALI generalizes to unseen environments, but if these pre-trained imagination models were trained on the same Matterport3D scenes that appear in the R2R or REVERIE val-unseen splits, the 'unseen environment' generalization claim would be compromised. Please specify the exact train/validation/test scene split for each pre-trained component, and clarify whether any of the val-unseen environments are seen during any stage of imagination-model training.","section":"§3.3.2, Imagination Model Pre-training"},{"comment":"The claimed 4% improvement on REVERIE unseen is not verifiable from Table 1. The ScaleVLN baseline row has missing values ('-') for RGS and RGSPL, so no comparison is possible on the reported REVERIE metrics; moreover the abstract refers to an 'SPL' improvement on REVERIE, but REVERIE results in Table 1 do not include an SPL column. Please define the exact metric and baseline used for the 4% claim, and report the corresponding numbers for that baseline.","section":"Abstract and §4.2, Table 1"}],"minor_comments":[{"comment":"There is a typo in the sentence 'xisting imagination mechanisms operate in isolation'—the word 'existing' is missing its leading 'e'.","section":"§2, Related Works"},{"comment":"The weight parameter λ appears in Eqs. (7) and (8) and the loss weights λG, λF, λP appear in Eqs. (9) and (10), but their values or ranges are never specified. Please provide the hyperparameter values used in training.","section":"§3.3.2, Eqs. (7)–(10)"},{"comment":"Equation (6) writes σCS(X) + (1 − σ)CF(Xt) with σ = CG(X), but the notation is inconsistent: σ appears to be both a scalar and a gating map, and X vs. Xt is used interchangeably. Please clarify the tensor shapes and the roles of CG, CS, and CF.","section":"§3.2.2, Eq. (6)"},{"comment":"The caption states that PSNR and Pearson correlation coefficients were calculated for imagined images and that the shaded curve represents mean and variance of pixel errors, but no numerical values are reported. Please include the actual PSNR and correlation values, or remove the claim that they were computed.","section":"§4.3, Figure 6"},{"comment":"In the REVERIE rows of Table 1, the ScaleVLN baseline has dashes for RGS and RGSPL, making it impossible to assess the relative improvement on those metrics. Please fill in the values or explain why they are unavailable.","section":"§4.1, Table 1"},{"comment":"The text refers to 'visited nodes , current nodes , navigable nodes , and imagination nodes' but the node-type icons appear as blank spaces in the manuscript. Please ensure the figure symbols are rendered correctly.","section":"§3.1.1, Memory Map Representation"},{"comment":"The implementation details section reports training iterations and batch sizes but omits learning rates, optimizer settings, and any hyperparameter search details. Please provide these to facilitate reproducibility.","section":"§4.1, Implementation Details"}],"recommendation":"major_revision","confidential_remarks":"The internal SPL inconsistency between Table 1 and Tables 2–5 is the decisive issue. If the correct SPL is 70 or 71 rather than 78, the central state-of-the-art claim collapses, because SALI would no longer surpass ScaleVLN on R2R val-unseen. I would ask the authors to provide the exact evaluation configuration for Table 1, per-seed results, and the training/evaluation scene split for the imagination models. If the authors cannot reconcile the numbers, the paper should not be accepted. The contribution itself is potentially interesting, but the current presentation does not support the headline claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the central architectural idea—a persistent topological memory enriched by recurrently imagined future scenes—is genuinely new in VLN and the ablations suggest it helps. Second, the headline SoTA number is inconsistent with the paper's own ablation tables, and until that's resolved the main claim doesn't hold.\n\nWhat's new and good: the recurrent imagination tree that writes imagined RGB-D and semantic nodes into a long-term topological map, with a learned fusion factor, is a real step beyond transient imagination methods like Pathdreamer, Dreamwalker, and SOAT. The ablations in Tables 2–5 consistently show that adding imagination to reality improves SR and SPL over reality-only and imagination-only, and that dynamic weighting beats fixed. The pattern is credible and suggests the mechanism does something useful. The four pre-trained component models are borrowed, but the integration is the contribution.\n\nNow the soft spots, in proportion. The load-bearing problem: Table 1 reports val-unseen R2R SPL=78 (NE=1.92, OSR=86, SR=82). Table 2 row 3 reports the same full model with identical NE, OSR, and SR but SPL=70. Table 3 row 3 gives SR=82, SPL=71; Table 4 row 4 gives SR=80, SPL=71. The same configuration cannot yield SPL 78 and 70/71. No seeds or error bars are reported, and no code is released. If the real SPL is 70/71, SALI is roughly tied with ScaleVLN (70) rather than beating it by 8%, and the SoTA claim collapses. This needs a correction and a clear explanation of which configuration is official. This is more decisive than the also-valid imagination-fidelity question, because even a perfect imagination module cannot support a SoTA claim built on irreconcilable numbers.\n\nSecond, the imagination fidelity is only qualitatively evaluated. Figure 6 shows PSNR and error curves without a threshold, and no evidence that the imagined images are accurate enough to help. That's a real but secondary concern, because the ablations do show empirical benefit from the imagination module.\n\nMinor: no code or model release, sparse implementation details, a few typos, and no discussion of pre-training data splits for the imagination models (potential leakage if the pre-training environments overlap evaluation).\n\nBottom line: the idea deserves attention and the paper deserves a serious referee, but in its current form the central quantitative claim is not reproducible. It should go to peer review with a required major revision to fix the inconsistency, report seeds and variance, and clarify the official configuration. Recommend engaging with the work, but don't trust the headline number until the authors resolve the discrepancy.","headline":"The architecture idea is genuine and the ablations are suggestive, but the headline R2R SPL=78 contradicts the paper's own ablation tables (70/71), so the SoTA claim is unsupported as written.","tokens_in":12661,"tokens_out":2772,"would_cite":false,"duration_ms":78929,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A navigation agent that stores imagined future views alongside real ones tops R2R and REVERIE benchmarks.","keywords":["vision-and-language navigation","episodic memory","episodic simulation","topological map","generative world model","SPADE","R2R","REVERIE"],"falsifier":"Rerun the R2R validation-unseen evaluation replacing the imagined images with ground-truth future panoramas from the simulator; if SPL does not rise above the imagined-image version, the benefit is not coming from the fidelity of what the agent imagines.","tokens_in":11669,"feed_emoji":"🧠","tokens_out":5298,"duration_ms":47462,"temperature":0.7,"pith_summary":"Navigation agents often lose their way in new buildings because they only act on what they currently see. This paper argues that an agent can do better by maintaining an episodic memory that includes imagined future scenes, not just observed ones. It introduces SALI, which builds a topological map whose nodes are either real observations or images the agent generates of places it has not yet visited. On the R2R and REVERIE benchmarks, that hybrid memory raises success-weighted path length by 8% and 4% in unseen environments over prior state-of-the-art systems. The claim is that mentally simulating the next views—not just perceiving the current one—is a practical navigation resource.","feed_headline":"Agent that imagines future rooms leads VLN benchmarks","feed_subtitle":"SALI's hybrid memory of real and dreamed views lifts SPL by 8% on R2R and 4% on REVERIE.","key_machinery":"The load-bearing mechanism is the recurrent imagination tree, which generates future scenes one step at a time and writes them into the topological memory. Each expansion takes the current semantic, depth, and RGB context plus a candidate position, and produces the next semantic/depth images with an inpaint model, then the next RGB image with a SPADE generator; room-type and waypoint models add priors and propose successor positions. These imagined nodes are pruned against existing nodes by feature cosine similarity and position distance, and their navigation scores are added to the nearest real node's score through a sigmoid fusion factor. The mechanism matters because it converts transient image predictions into durable, reusable memory that the agent's graph-aware transformer can reason over jointly with real observations.","core_discovery":"SALI's central discovery is that a vision-and-language navigation agent can use a 'reality-imagination hybrid memory' to plan more effectively in unseen indoor environments. The agent keeps a topological map of the building, where visited nodes store real RGB-D and semantic images and unvisited waypoints can store images the agent generates by a recurrent imagination tree: an inpaint model produces depth and semantic maps, a SPADE-based model renders the RGB view, a room-type model injects commonsense object priors, and a waypoint model proposes the next positions. The imagined nodes are fused into the map and scored together with real nodes by a graph-aware transformer, with a learned weighting that fades imagination over time. The paper reports state-of-the-art SPL on R2R and REVERIE unseen splits, with the imagination component contributing most on instructions that combine room and object terms. In effect, the agent plans from what it can picture, not only from what it has seen.","pith_inferences":["If imagined views act like cheap exploration, the same hybrid memory could reduce collisions and dead-ends in continuous navigation, where the waypoint model is replaced by a real motion planner.","The framework suggests a principled test of generative world models for embodied AI: replace the imagined images with ground-truth future frames and measure how much of the SPL gain survives, which would isolate image fidelity from graph expansion effects.","The room-type dictionary is a hand-built commonsense prior; learning this mapping jointly with the navigation policy could remove the manual weights and adapt to novel room-object statistics.","The decreasing fusion weight resembles a memory-forgetting curve, so one could make it a function of navigation confidence rather than step index."],"forward_implications":["Imagination should be stored in long-term memory, not discarded after each step; durable imagined nodes are what lift SPL on unseen environments.","A moderate imagination horizon of two steps and at most four imagined nodes is optimal; too many imagined nodes makes different locations compete and lowers SPL.","The learned dynamic weight for imagined nodes, which shrinks as navigation proceeds, beats a fixed 0.5 blend.","Room-type and waypoint auxiliary models each add measurable SPL, confirming that commonsense priors and position prediction contribute beyond raw image generation.","The biggest gains occur on instructions mixing room and object terms, indicating that imagined scenes help the agent ground relational language."],"supporting_citations":[{"why":"Defines the R2R task, the Matterport3D simulator, and the navigation metrics used for the main evaluation.","marker":"Anderson et al. 2018b"},{"why":"Supplies the topological map representation and graph-aware self-attention that SALI inherits and extends with imagined nodes.","marker":"Chen et al. 2022"},{"why":"Pathdreamer provides the world-model approach for generating future views that the imagination tree builds upon.","marker":"Koh et al. 2021"},{"why":"SPADE is the semantic-to-RGB GAN used as the spade model within the recurrent imagination tree.","marker":"Park et al. 2019"},{"why":"The waypoint prediction model is reused to propose neighboring positions from generated RGB-D images.","marker":"Hong et al. 2022"},{"why":"REVERIE benchmark and its remote grounding metrics evaluate the coarse-instruction case.","marker":"Qi et al. 2020"},{"why":"BEVBert is the space-aware baseline whose SR and SPL SALI compares against on complex-instruction subsets.","marker":"An et al. 2023"}],"fun_headline_variants":["Imagining future rooms helps VLN agent navigate unseen paths","SALI's imagined views improve SPL on R2R and REVERIE","Dreaming scenes sharpens VLN agent's route planning","Episodic imagination and memory guide VLN agent","Agent that pictures unseen scenes tops VLN benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The imagined scenes—semantic, depth, and RGB images of unvisited places—are accurate enough that adding them to the memory map improves action selection.","fun_headline_variants_meta":{"raw":{"variants":["Imagining future rooms helps VLN agent navigate unseen paths","SALI's imagined views improve SPL on R2R and REVERIE","Dreaming scenes sharpens VLN agent's route planning","Episodic imagination and memory guide VLN agent","Agent that pictures unseen scenes tops VLN benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000262,"raw_usage":{"total_tokens":1564,"prompt_tokens":877,"completion_tokens":687,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":602}},"tokens_in":493,"tokens_out":687,"duration_ms":6984,"temperature":1.0,"reasoning_tokens":602,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:15:20.066547+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the R2R validation-unseen evaluation replacing the imagined images with ground-truth future panoramas from the simulator; if SPL does not rise above the imagined-image version, the benefit is not coming from the fidelity of what the agent imagines.","supporting_citations":[{"cited_title":"Y.; Lee, H.; Yang, Y.; Baldridge, J.; and Anderson, P","cited_arxiv_id":null,"evidence_quote":"Pathdreamer provides the world-model approach for generating future views that the imagination tree builds upon."}],"review_version":1}