Pith. sign in

Paper Citation Record · LEDGER

$A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2308.07997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.07997 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:06:07.603723Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.680188Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f2017848-d5f4-41bb-b0da-dac0c2f3a1e2 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:55:20.433292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:804d4c6686814b9b9326e981a2130d5154a0697f3d0819b09c148c6831a91ede

Observation 99481ef7-81e6-43aa-8a80-6e604f803d84 · inbound

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences cites this paper.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:36.793861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:36.793861Z digest=sha256:0e72b458472d0a1ef5fefd9a41b7889e6df550cf61285075bd637f9b5e3ccd7b

Observation cd1733aa-5086-416d-a059-a6243cda204f · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:51:36.210974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:a9a4255c166b32d45e7e5845434946c1a929edc6ec20f3f1162e22741b3ca150

Observation 8d2e54c0-530f-4eb5-a782-e854735386fb · inbound

Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments cites this paper.

Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T16:24:58.053040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:24:58.053040Z digest=sha256:f7a211a0e620c38bcffdd0489332b94acf13eda441ea00a809fd271c206e42eb

Observation 26d07516-a634-4b97-aeb6-84e2968c057e · inbound

Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach cites this paper.

Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T20:11:13.070999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:11:13.070999Z digest=sha256:1c44f136047929d79f05b404e839c5cb75991215f3aa83eaa425342a3a73961e

Observation 07456b5e-b80a-434b-b619-dce90441f859 · inbound

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation cites this paper.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.603723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.603723Z digest=sha256:2de0ea00b85357e944994ec89014e4c4da8a39560b6e342b9c19e411310ead1c

Observation a354b9cd-0b2f-4b18-80ab-d78d4c53e4a9 · inbound

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation cites this paper.

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:04.892845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:04.892845Z digest=sha256:0b6f2cc055a5d11acc3963d46be20d1cf9eb84bca437f219b3fbb504d77a011f

Observation 5dbc5a12-f24c-4ba0-a699-c9bbb0bad00d · inbound

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity cites this paper.

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T19:31:53.491759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:31:53.491759Z digest=sha256:5db11cbda0f9a5f77a02e17e79c2f7d97a068c4e1c5502a13d5c3d42c285aa67

Observation bd6145f8-ef79-49a9-8f3a-5e4d394bf0ce · inbound

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments cites this paper.

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:23.962892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:23.962892Z digest=sha256:4295d6ab76c101558bf906ab5059d7203108eee093ea16e1b60d68825b791316

Observation a7887672-18db-407a-8b76-953b46a3a374 · inbound

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference cites this paper.

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:22.470253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:22.470253Z digest=sha256:d4606549915c4880cd8184c7637b2fa20535720cc15f2b50eba680ade714b873

Observation 7320402f-b533-4006-9ce7-6023d1c394ab · inbound

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling cites this paper.

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:06.419946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:06.419946Z digest=sha256:5a9cce89753595423b9afd9704bab9b423274f816d513547f00a0f9ce2d8df58

Observation f35be150-ebfb-4927-adf6-9618299e681b · inbound

GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation cites this paper.

GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:57:22.880205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:57:22.880205Z digest=sha256:81566af9f220bfa71064ab26bb82598e2119539a48439eaa594150d399ca00ce

Observation a7fa8149-ecd8-415c-ad33-30a5191c59d9 · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.967379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:272db3c3048d8c07515dc5356e3d52ac793c26bee0760d5b138951b201b2e5fa

Observation af29c0ea-c9a8-4f73-8142-d5139a3ddd77 · inbound

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation cites this paper.

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:22:37.623967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:13:12.029402Z digest=sha256:8b3016c70cfa9198a4c3e45bbc38d6f8007e08b927694ccf5ff246d83e46d0ef

Observation a7d11f88-d4fd-4bce-93c9-5fddcb0893e1 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.595494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:d6d0e57753ee2b33599834ae997e59d890a09d2b90cb8b537d4b634164e084ec

Observation 520c74aa-25d3-4f63-9b9f-ba79d56c1df5 · inbound

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation cites this paper.

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.371249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:33:32.671550Z digest=sha256:7c0f6d3c340f7705567aefdd1a81a67612dbc5ea544e1ff17fef6657979c5a6f

Observation e1aea1c4-aadd-4e44-b126-5a3c6a49b030 · inbound

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness cites this paper.

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.682296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T13:35:34.630980Z digest=sha256:549bb9d261be80a444a7d94808e4214af286da0709329f57c234f59f11208be8

Observation 8e2fd2c4-b015-4b8a-87a6-ae31fb7b106f · inbound

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness cites this paper.

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T11:56:28.498296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:56:28.498296Z digest=sha256:bb38f8a4a18b635da0fec77dc7b1b7095810a56771d2761c5a47e92b951e9ebd