Pith. sign in

Paper Citation Record · LEDGER

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN

As of 7 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2605.14801.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14801 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T20:53:06.964338Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:51:21.680693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bb202a4-4bd9-4b49-afec-331e877941d4 · outbound

This paper cites Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:03.885454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:94f5d1325b82d255212750bc8dd8ab0a6250aa63e13dd64ae002f54d78bdde3b

Observation 88cff52e-737d-4491-b6b4-b760b3b7fcb7 · outbound

This paper cites 2026 , eprint =.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN 2026 , eprint =

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T20:55:03.882072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:f355436a59da2f1df5cc1ccc44bd303ea37f7ab191807bee984e1f2a8610ac9a

Observation 5eb36027-246a-47f0-9930-4aa2631314ec · outbound

This paper cites Navcot: Boosting llm-based vision-and-language naviga- tion via learning disentangled reasoning,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Navcot: Boosting llm-based vision-and-language naviga- tion via learning disentangled reasoning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.006706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:3518a397a10666d76aea46f51f31fe457615769925de4574ff7d6456dd2228ff

Observation d945913c-4191-4f4f-81ff-eb58e339a8f7 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.004084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:b921bd55aed1003decb7b1dbbdd3809f29cb6613426131e3d5e9e9f9205368bb

Observation d5a66d51-96d6-4b34-8089-38848f80d46b · outbound

This paper cites NavGPT: Explicit reasoning in vision- and-language navigation with large language models,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN NavGPT: Explicit reasoning in vision- and-language navigation with large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.992815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:7c81188cfac68b15b6eb0566774da0bba02dd09610f63670b3fd9a13ff04c662

Observation cad9d981-959c-4209-beba-f681d7280729 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T20:55:03.889368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:2d7bb9aac0a26d85120c582cafbcd77baf4045702cbdae4acda480fef4f87785

Observation 57e68df1-1f6c-47ca-9bce-1ae29892a405 · outbound

This paper cites GPT-4 Technical Report.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN GPT-4 Technical Report

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T20:55:03.896309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:91122f92c651d8734772337cf791bdb33807c1162bdcbb01dba92cbddb1259de

Observation ee8a0c53-0565-499a-bb3f-dc4feb44071d · outbound

This paper cites MapGPT: Map-guided prompting with adaptive path planning for vision-and- language navigation,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN MapGPT: Map-guided prompting with adaptive path planning for vision-and- language navigation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.995798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:cb993945be61188389bafa7c609814e0b93d700409a6a36103da9b92b9d0fbcd

Observation 6603ba9b-ad7a-4b93-a40d-9422f03fd1a5 · outbound

This paper cites Open-nav: Exploring zero-shot vision-and-language naviga- tion in continuous environment with open-source llms,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Open-nav: Exploring zero-shot vision-and-language naviga- tion in continuous environment with open-source llms,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.009248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:8705032da8409ad7e28e6f2dfbc5cf46fc666944d2c6354577351cf379f9cbeb

Observation 641147b6-f487-413d-b29b-251d20a36c07 · outbound

This paper cites Spatialbot: Precise spatial understanding with vision lan- guage models,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Spatialbot: Precise spatial understanding with vision lan- guage models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.000869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:2207dae2ad4e9770a02cf5e1793063884c8523d74fef35213a82d58f351765e1

Observation e2ebdaa7-f069-4ab9-8d83-1de242203d27 · outbound

This paper cites Recognize anything: A strong image tagging model,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Recognize anything: A strong image tagging model,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.998458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:4b1204cde441e540687fd292f5c922f68571cc1ff563ee42c68b9d1a7e24b203

Observation 06bd5bf7-9ebf-4369-ba45-197aba4cd2f5 · outbound

This paper cites Fast r-cnn,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Fast r-cnn,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.011488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:5c0a358486483ecfa5d61b738e0bc9609f92d8e231ba3940646b437cbd6e232d

Observation 3ff4f8b9-bc30-4b45-a74c-c2ad14bbd16b · outbound

This paper cites EmbodiedSAM: Online Segment Any 3D Thing in Real Time.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN EmbodiedSAM: Online Segment Any 3D Thing in Real Time

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T20:55:03.893407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:4a4f0d342ff390d4860857a6f88f01a9eeabd3f3f4780ad2deb9abb5f985b6d1

Observation 6a1b0c9f-4da1-4728-9a1e-8a3b7d64e5ea · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.990190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:482042a4ac5f38d5f2c3972c1650ad760da0302e1024ed3babfa0e6b9120ef09

Observation 32e26628-699f-4284-b9bc-fdf56ed7735b · outbound

This paper cites Habitat 2.0: Training Home Assistants to Rearrange their Habitat.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Habitat 2.0: Training Home Assistants to Rearrange their Habitat

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:03.900022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:4c89ccb87f21f7ca7f05f1fa03aad4e25898403f4587608c8371a635ba467b29

Observation 05856693-ba6b-41ed-8e8d-3e04635f3e32 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Scannet: Richly-annotated 3d reconstructions of indoor scenes,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.014243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:ab8aba9b9ebcb5ce2aea7ca0d9884a29b027d5a9b5a2912486c7747c030b1ba6

Pith citing papers

Observation f5bf930a-b02c-4194-ad3c-0f7fb314fa4b · inbound

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics cites this paper.

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:51:21.680693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:51:21.680693Z digest=sha256:898c9aefb35b5fa324e649e5a6108743abe94f85cbe0843648dcfd52e2b52027