Pith. sign in

Paper Citation Record · LEDGER

Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2010.07954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2010.07954 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:06:07.991028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.507471Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f48cd204-2208-4c16-884a-78557a45f7e0 · inbound

Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI cites this paper.

Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:24:57.995737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T18:24:57.956486Z digest=sha256:acb48cd6a9fc0c1bb43b161075b9c53a52724d3ac85819bd27ce92c3162eaf01

Observation a5a35ecb-9be6-480b-ac94-bb3e1f377dce · inbound

Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks cites this paper.

Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:10:26.129733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:10:26.129733Z digest=sha256:06b5858cbb89948d64fe2a91de00a548a9b1f7a0fad1b03a6fd780a1ddf832d9

Observation db037f63-c215-4897-9654-164d1191e68b · inbound

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models cites this paper.

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:47:58.974766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:47:58.974766Z digest=sha256:350c017a8cd054a34a0e0726f4f308dea219ac43ad829fc4021f317050bfd69c

Observation 0d7707cf-d0f5-4175-a900-d79c67ba6587 · inbound

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation cites this paper.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.991028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.991028Z digest=sha256:de839a2c79ff51a99520b2d5be0470d7b7c833329e7de6bcd2a3b7b49a0ad039

Observation 7dfd67f3-cb68-4b50-9467-5d061a57ead8 · inbound

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation cites this paper.

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:07.294284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:07.294284Z digest=sha256:3def57cd407d7341efc2f0495a0d036caa7645f93f6e03cc0af567ea51f44746

Observation 70223ba6-7919-4ba5-9bd5-238504b91dc6 · inbound

Efficient and Generalizable Environmental Understanding for Visual Navigation cites this paper.

Efficient and Generalizable Environmental Understanding for Visual Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:59.097188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:59.097188Z digest=sha256:cfe702b9c06d30faec0cfa198276b3f44a025adbef31d1d0e4e0cd0188596592

Observation c253d784-c227-4c25-b950-f10c9c8ce7cd · inbound

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling cites this paper.

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:05.110483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:05.110483Z digest=sha256:edc861ac05fbfe13018e2d5983295c77f5d30f3d63fda80ff5323190da2be0fd

Observation b1a397d6-37ec-422c-858d-f40c8da9034c · inbound

SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments cites this paper.

SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:40.343600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:40.343600Z digest=sha256:0ecb6b4700e64fda26bff144f71157a6aede5d47efa7eb6cc9facc55971ad5d3

Observation e30f8ce6-b283-4e04-9317-f02585409009 · inbound

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation cites this paper.

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:51:13.571137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:51:13.571137Z digest=sha256:32d324a1fe446524a948c039ebbf4262ab5fcded861f36e6cdf6ecccaf9c187e

Observation 7d223918-5734-40c5-8e9d-ad589a900e6e · inbound

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey cites this paper.

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T11:20:06.686740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:20:06.686740Z digest=sha256:f6e4c3a9b008956a9ebbf93a2c15fe6ff492d8a1b0dc7a2a7b31942d4179e8b8

Observation c09b470b-bfe6-44c1-9759-68594e9ec207 · inbound

SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps cites this paper.

SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:02.703754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:46:02.703754Z digest=sha256:bb8f65b3189fb21d9219b983a610dc492aa6a37ece9937339fdd78831f2a31fb

Observation 45615e25-8a37-4b59-951f-91762665c707 · inbound

GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation cites this paper.

GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:57:23.009185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:57:23.009185Z digest=sha256:b03125a82e9a66f951d67b9bb3fae805f1682c4332da47e17345ba3752e3ceac

Observation 710a5dbd-cc5d-443a-8bc1-23f8133a308e · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.961021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:858b6f627bfc39650843f353ff3a13140bbb8cf22af14c034a7c955e218fce47

Observation 0e0924f5-0476-4c30-a355-a4cd9a75639f · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:28:20.888687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:7e1cdb80c6dbe5a75595f459e6100c993281f9ce77bf690204ba508af9649d74

Observation 140c8ea1-5d6f-4b8e-a270-957019263e01 · inbound

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation cites this paper.

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:56.991649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:28:56.991649Z digest=sha256:2b286cb65dd68071634662d14647c61f37875a897db132188457a6e65b0f1372

Observation bedf0b36-102c-48fe-a757-0962240c6c86 · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T23:48:58.731930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:48:58.731930Z digest=sha256:f2771b1578e3b958be96e94385fb2ac732e19c6674e791662a17ed362d045d32

Observation d65e00c9-c22a-49dc-ae2a-921b19c9a4fa · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T17:15:42.404110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:15:42.404110Z digest=sha256:2327ea2b12ee6432aa9927098c3700f1f61c44cbe970fc7bb1d79e148be987d8

Observation cb750f35-117f-4b7c-81fe-bc8e0fb05753 · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T05:43:09.929718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:43:09.929718Z digest=sha256:ab755cf913b7d98672bf34f153e69e94ce1abc05b24ee7df786e5cddb718fd23

Observation 17f95598-8003-4373-a3ab-5c199862b124 · inbound

Think before Go: Hierarchical Reasoning for Image-goal Navigation cites this paper.

Think before Go: Hierarchical Reasoning for Image-goal Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.536845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:43:27.972164Z digest=sha256:672a494a3143171c214f376129cc9321d50080a8b766c14994c0722e000e655c

Observation 97613dd1-e17e-4654-98df-7407e7f5a500 · inbound

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation cites this paper.

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:45.846828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T06:50:34.310831Z digest=sha256:28be34ccf9aba7d28a59f460be25deacbb99bf98341e33eb00fcffc37e12e66c

Observation 4832bc51-6171-45b9-8807-1a5a7f2a4e5a · inbound

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation cites this paper.

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:31.296599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T05:05:33.606975Z digest=sha256:ba90d2a10297bcb74790c2bf47b20fccb70d6778205e8b3b1b611c41913f7d9d

Observation 25afc675-608f-4440-9a7f-23147d086b79 · inbound

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation cites this paper.

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:27.478293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:33:53.557357Z digest=sha256:0f17c9e332459089f30fbcf9e6e62d807b9526c6e1f9b4185b3af1ed63a82472

Observation 70cc6fd7-e32e-4d82-9d1c-a40a9c2047c9 · inbound

NavOL: Navigation Policy with Online Imitation Learning cites this paper.

NavOL: Navigation Policy with Online Imitation Learning Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.913474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:51:53.848800Z digest=sha256:61fdf668406a4fd6a346923bbca5bd10602d7078ac931a9b1bf6313a736f979d

Observation dad079b5-d37e-4a6d-9baa-ae99be1bde43 · inbound

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation cites this paper.

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:46:10.597994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T06:45:35.920991Z digest=sha256:267110d45d4e984c794daa36c5a1de7d04ac4ccaf564cad310b845051ae24a40

Observation 6b2826c3-71dc-4c98-ba20-c7d446411fbc · inbound

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation cites this paper.

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.509732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T10:24:56.607921Z digest=sha256:c409cc4333ee7ef5197fbe2ef4d7f2ac5b02985fc797c8a5200b2ffd3befb0c1

Observation 6b50e8f1-74ea-4af0-8bf9-d78277a542aa · inbound

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation cites this paper.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:54:49.926510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T05:05:56.065261Z digest=sha256:1684ffe2ea0d3f89f6e8524d86f9dda886dd98e1f1e42f63ba60f25562618c33

Observation 14c17f26-fc29-4785-aa0d-e7e187f7eae7 · inbound

Joint On-and-Off Policy Learning for Vision-and-Language Navigation cites this paper.

Joint On-and-Off Policy Learning for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:13:55.288909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:13:55.288909Z digest=sha256:cd3db218e1252d7e1608d59398c77f87765ffb420a7bb0f986298dcd69940d15

Observation 50370e0f-fdbc-4726-9b82-d38f282ff8c4 · inbound

Goal-oriented Navigation Instruction Generation with Tour Video Priors cites this paper.

Goal-oriented Navigation Instruction Generation with Tour Video Priors Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:35.115493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:35.115493Z digest=sha256:35191ed4cdd26bcdc7bb31e45d92049944d73030e5516645d4e6e2de8dca04a0