Pith. sign in

Paper Citation Record · LEDGER

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN

As of 7 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2605.14801.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14801 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T20:53:06.964338Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:51:21.680693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bb202a4-4bd9-4b49-afec-331e877941d4 · outbound

This paper cites Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:03.885454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:55d41c3597ac482b656d405a6cc5aa5a65417fb9604c97e826c6ef9ae54fc16d

Observation 88cff52e-737d-4491-b6b4-b760b3b7fcb7 · outbound

This paper cites 2026 , eprint =.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN 2026 , eprint =

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T20:55:03.882072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:c511917ad10d8bc70c5b413d94653ff4f342591da6a80d2c60e7ce62968a1667

Observation 5eb36027-246a-47f0-9930-4aa2631314ec · outbound

This paper cites Navcot: Boosting llm-based vision-and-language naviga- tion via learning disentangled reasoning,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Navcot: Boosting llm-based vision-and-language naviga- tion via learning disentangled reasoning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.006706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:5b434acecdb9ae9713e1e261a34ec39f37e203b892a225c657a2fe3d60fc7525

Observation d945913c-4191-4f4f-81ff-eb58e339a8f7 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.004084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:afc8b917d639475e0c2258a11c82ba36c465b31b3b0072c1328c3870e2430880

Observation d5a66d51-96d6-4b34-8089-38848f80d46b · outbound

This paper cites NavGPT: Explicit reasoning in vision- and-language navigation with large language models,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN NavGPT: Explicit reasoning in vision- and-language navigation with large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.992815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:f095f1511800f7cfbe47dbc8870fb6d0281647680991fc6fd8ab79ad2a29c896

Observation cad9d981-959c-4209-beba-f681d7280729 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T20:55:03.889368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:8d9353b9f1d1208c556decf68c4e594b89de9cc1dc71203a77bb8ffcb6971398

Observation 57e68df1-1f6c-47ca-9bce-1ae29892a405 · outbound

This paper cites GPT-4 Technical Report.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN GPT-4 Technical Report

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T20:55:03.896309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:0a448386fd2ceb24a6dd7cd589a54584e0cfb8e7d6a8d820ada3f20cc3cc60ad

Observation ee8a0c53-0565-499a-bb3f-dc4feb44071d · outbound

This paper cites MapGPT: Map-guided prompting with adaptive path planning for vision-and- language navigation,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN MapGPT: Map-guided prompting with adaptive path planning for vision-and- language navigation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.995798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:6730cf1be68176fccc94b5873d9120b0ea99dae8b17ec4398714a370cc6d86c2

Observation 6603ba9b-ad7a-4b93-a40d-9422f03fd1a5 · outbound

This paper cites Open-nav: Exploring zero-shot vision-and-language naviga- tion in continuous environment with open-source llms,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Open-nav: Exploring zero-shot vision-and-language naviga- tion in continuous environment with open-source llms,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.009248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:b41bd22ac42c559efcd517ce43dd62093504b35c41eb6ff884ffcea82017f4f8

Observation 641147b6-f487-413d-b29b-251d20a36c07 · outbound

This paper cites Spatialbot: Precise spatial understanding with vision lan- guage models,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Spatialbot: Precise spatial understanding with vision lan- guage models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.000869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:48731824afd425fc9a3af9425a311d73175ddde0cebbcc9d65fbad5028d6cb91

Observation e2ebdaa7-f069-4ab9-8d83-1de242203d27 · outbound

This paper cites Recognize anything: A strong image tagging model,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Recognize anything: A strong image tagging model,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.998458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:fa57bd9326d6172436d0d0252906212744c0c739dea2c55410118c64dedcd95f

Observation 06bd5bf7-9ebf-4369-ba45-197aba4cd2f5 · outbound

This paper cites Fast r-cnn,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Fast r-cnn,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.011488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:4e6bdf9507aa275b075f8cefde486151fd0de3ae72c3d62fd52fd2a2858ede0d

Observation 3ff4f8b9-bc30-4b45-a74c-c2ad14bbd16b · outbound

This paper cites EmbodiedSAM: Online Segment Any 3D Thing in Real Time.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN EmbodiedSAM: Online Segment Any 3D Thing in Real Time

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T20:55:03.893407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:7d650b0c7f5a11f5e632f8069f7a482373054fbd002e3411de279822ad3d8b7b

Observation 6a1b0c9f-4da1-4728-9a1e-8a3b7d64e5ea · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:01.990190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:f72ee21467c49ebf0e9ac3e8505272010d6df3cf920592e6c54d61f9b9f35af5

Observation 32e26628-699f-4284-b9bc-fdf56ed7735b · outbound

This paper cites Habitat 2.0: Training Home Assistants to Rearrange their Habitat.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Habitat 2.0: Training Home Assistants to Rearrange their Habitat

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:03.900022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:6f83b8362115f67c7232b54a8c2e77887f87cce21b238b308dd2967519fa50ab

Observation 05856693-ba6b-41ed-8e8d-3e04635f3e32 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes,.

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN Scannet: Richly-annotated 3d reconstructions of indoor scenes,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T18:54:02.014243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:53:06.964338Z digest=sha256:dfe256db4cbebcd19180f9ce6c3822be2f2dceb8aa2dce13e246e5b9b2b80795

Pith citing papers

Observation f5bf930a-b02c-4194-ad3c-0f7fb314fa4b · inbound

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics cites this paper.

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:51:21.680693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:51:21.680693Z digest=sha256:898c9aefb35b5fa324e649e5a6108743abe94f85cbe0843648dcfd52e2b52027