Pith. sign in

Paper Citation Record · LEDGER

Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2010.07954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2010.07954 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:06:07.991028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.507471Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f48cd204-2208-4c16-884a-78557a45f7e0 · inbound

Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI cites this paper.

Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:24:57.995737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T18:24:57.956486Z digest=sha256:f55868ca871f792fbcfa28a3983f792e706374f39630cd25fc9a947900ea9e2c

Observation a5a35ecb-9be6-480b-ac94-bb3e1f377dce · inbound

Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks cites this paper.

Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:10:26.129733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:10:26.129733Z digest=sha256:134fdead2f205d251b772ddb5897bb96661489d6a32ceb610852b7297c7e001e

Observation db037f63-c215-4897-9654-164d1191e68b · inbound

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models cites this paper.

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:47:58.974766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:47:58.974766Z digest=sha256:7d4657791c44e8f79193312f1913264919a1853a7556eff3c018bf59e3692d97

Observation 0d7707cf-d0f5-4175-a900-d79c67ba6587 · inbound

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation cites this paper.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.991028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.991028Z digest=sha256:6fd72028dcde575643823ffb947f08431d8c55820e082056e088d5e189f7e49e

Observation 7dfd67f3-cb68-4b50-9467-5d061a57ead8 · inbound

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation cites this paper.

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:07.294284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:07.294284Z digest=sha256:30df18f1b1c44c9763ec5ab31ac3ab60584dc24930e3fdeef551f7b5e250a80a

Observation 70223ba6-7919-4ba5-9bd5-238504b91dc6 · inbound

Efficient and Generalizable Environmental Understanding for Visual Navigation cites this paper.

Efficient and Generalizable Environmental Understanding for Visual Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:59.097188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:59.097188Z digest=sha256:836d6c27082b0fec66fdb43673e22e9794178df6101c1228e6ea6be5ff26af7a

Observation c253d784-c227-4c25-b950-f10c9c8ce7cd · inbound

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling cites this paper.

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:05.110483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:05.110483Z digest=sha256:3e9b1a2c393631fc9f476b62de7816a46af62a905bcf84838c67d44ac484944a

Observation b1a397d6-37ec-422c-858d-f40c8da9034c · inbound

SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments cites this paper.

SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:04:40.343600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:04:40.343600Z digest=sha256:186fb3cf4e882b5c63e8e93ce358b1229839aac5c01edf2faedd3545fb3a8609

Observation e30f8ce6-b283-4e04-9317-f02585409009 · inbound

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation cites this paper.

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:51:13.571137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:51:13.571137Z digest=sha256:b11c2206745f73b0f23514df5d33b743f8a4f8f5b22a9ff1cb36a140b92b6db5

Observation 7d223918-5734-40c5-8e9d-ad589a900e6e · inbound

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey cites this paper.

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T11:20:06.686740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:20:06.686740Z digest=sha256:dfb4a16f302058e081d3c7606617830436860e284c4b4a283a276f438136af8b

Observation c09b470b-bfe6-44c1-9759-68594e9ec207 · inbound

SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps cites this paper.

SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:02.703754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:46:02.703754Z digest=sha256:9ad599c70acf5bee501b2f2eb3858924a649991ff36b7ab82744eec591ccadac

Observation 45615e25-8a37-4b59-951f-91762665c707 · inbound

GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation cites this paper.

GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:57:23.009185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:57:23.009185Z digest=sha256:3db872105000057c06d530e348fcd3738bfba49a71efa13cfc9f2cc2856c4bca

Observation 710a5dbd-cc5d-443a-8bc1-23f8133a308e · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.961021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:b578072bd5d20665b38bf7214afc18722b148defe04ead260897e09d946102c8

Observation 0e0924f5-0476-4c30-a355-a4cd9a75639f · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:28:20.888687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:e5fb8bdb58c1677246d86fb9e7118822ebfa119222e2c49a41b9725e1103b484

Observation 140c8ea1-5d6f-4b8e-a270-957019263e01 · inbound

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation cites this paper.

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:56.991649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:28:56.991649Z digest=sha256:f722b21e49b8692e4e44879df72424c099ec11b3942e13fcd851c54dfb2f6856

Observation bedf0b36-102c-48fe-a757-0962240c6c86 · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T23:48:58.731930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:48:58.731930Z digest=sha256:0c559845c921e60988522b2dbf6f1a451ebf5f52520a8f1003ce8d344ca6efd1

Observation d65e00c9-c22a-49dc-ae2a-921b19c9a4fa · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T17:15:42.404110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:15:42.404110Z digest=sha256:2e361240dc89af17caa7cba8e67dcc68719dc03f969a355498ab3bd31d66d2fc

Observation cb750f35-117f-4b7c-81fe-bc8e0fb05753 · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T05:43:09.929718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:43:09.929718Z digest=sha256:0ace261143f437194daed84b67d3be641fea694a70e4298a1032703cdd23017c

Observation 17f95598-8003-4373-a3ab-5c199862b124 · inbound

Think before Go: Hierarchical Reasoning for Image-goal Navigation cites this paper.

Think before Go: Hierarchical Reasoning for Image-goal Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.536845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T05:43:27.972164Z digest=sha256:1838dc85507a429a5ec81171a51e148ceacb2df294d788371ca4d87189676124

Observation 97613dd1-e17e-4654-98df-7407e7f5a500 · inbound

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation cites this paper.

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:45.846828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T06:50:34.310831Z digest=sha256:0d398ef2704702702ae53da992bb18fb570748b48b66d1303edf033328edc928

Observation 4832bc51-6171-45b9-8807-1a5a7f2a4e5a · inbound

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation cites this paper.

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:31.296599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T05:05:33.606975Z digest=sha256:62c682451ef37597750638bac42077ca47032e3e1b96f92bee82ee7f9ad19a7d

Observation 25afc675-608f-4440-9a7f-23147d086b79 · inbound

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation cites this paper.

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:27.478293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:33:53.557357Z digest=sha256:f868f3f71b26572a211beba737b27955ea1e8ad05a7d0ea15a1074489db2957d

Observation 70cc6fd7-e32e-4d82-9d1c-a40a9c2047c9 · inbound

NavOL: Navigation Policy with Online Imitation Learning cites this paper.

NavOL: Navigation Policy with Online Imitation Learning Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.913474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T05:51:53.848800Z digest=sha256:711a9bb7031ad50ab0d87807045b6bd1b3c07392b11cf76336efd105df51896b

Observation dad079b5-d37e-4a6d-9baa-ae99be1bde43 · inbound

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation cites this paper.

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:46:10.597994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T06:45:35.920991Z digest=sha256:3b4e2eda5f8e83547287258f0a0ad19f063f02fafd8f2e02438246b27b874baa

Observation 6b2826c3-71dc-4c98-ba20-c7d446411fbc · inbound

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation cites this paper.

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.509732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T10:24:56.607921Z digest=sha256:4189c8d9f48ccb1a31f12327cee5f478a7e104150696ca935031b0b93f32083f

Observation 6b50e8f1-74ea-4af0-8bf9-d78277a542aa · inbound

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation cites this paper.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:54:49.926510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T05:05:56.065261Z digest=sha256:cb14d347ee1ea045932e51d868173e028dd9a2f8865a818b385bdb78f78f376b

Observation 14c17f26-fc29-4785-aa0d-e7e187f7eae7 · inbound

Joint On-and-Off Policy Learning for Vision-and-Language Navigation cites this paper.

Joint On-and-Off Policy Learning for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:13:55.288909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:13:55.288909Z digest=sha256:80fcd84e2786a0b8e232c5ede1e3401d3039302164b4af87d5822751b6f9e14f

Observation 50370e0f-fdbc-4726-9b82-d38f282ff8c4 · inbound

Goal-oriented Navigation Instruction Generation with Tour Video Priors cites this paper.

Goal-oriented Navigation Instruction Generation with Tour Video Priors Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:35.115493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:35.115493Z digest=sha256:f05c75c1a61a7da7f0028d4fe7ce6808766ed3170741158ef2968de0f7459de0