Pith. sign in

Paper Citation Record · LEDGER

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

As of 16 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 6 inbound Pith citation observations for arXiv:2505.11383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11383 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:11.164370Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:04:39.821353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:37:03.075467Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01d8955a-8656-45e7-b7e4-24c4531b860d · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.849564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.849564Z digest=sha256:33a64f77546f932daf02621e993dcb725d232b370cda18ba860dc348c584aa8a

Observation d038db71-6345-4fb3-ad16-aeb9a5715e14 · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Reverie: Remote embodied visual referring expression in real indoor environments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.903499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.903499Z digest=sha256:c9c33187d86b2113b6e59070e0460976a48d435256f8208cafa819cf508341dc

Observation d9e9caf6-4b13-4bed-8dd9-b3e6c09b2f2c · outbound

This paper cites Beyond the nav- graph: Vision-and-language navigation in continuous environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Beyond the nav- graph: Vision-and-language navigation in continuous environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.936174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.936174Z digest=sha256:af3c6ed755b93ac403330c21a868aa603b70433ee2dd7eb7e60d4e78a512ad22

Observation 45278d17-f96a-4c42-baeb-20620e5e5c34 · outbound

This paper cites NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.939678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.939678Z digest=sha256:ea5f780593535829223e6cec88362548cd962e9b7edb4375207d8467789aa8a5

Observation 705f5863-bebe-4c38-8383-8273c10a40e6 · outbound

This paper cites Navid: Video-based vlm plans the next step for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Navid: Video-based vlm plans the next step for vision-and-language navigation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.803239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.943451Z digest=sha256:756eb8c4ba4627ca1fe1308e9d8db42687720f6851b88a9f8f5b429624f85c91

Observation bcaf25e4-d524-4bf8-a3f5-77a3e0cbd0ec · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.947005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.947005Z digest=sha256:01185aec623144532e1d551812936dd76d385c6f810ca5164ae0ca07f30801a8

Observation d61b36e6-545f-431c-a197-7e3e499c7063 · outbound

This paper cites Navila: Legged robot vision-language-action model for navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Navila: Legged robot vision-language-action model for navigation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.792553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.952731Z digest=sha256:9895c85b368fdb95c1d022e228801ee4a26665cb845c761315e53331fd46fe14

Observation b3b93662-d1aa-4d52-8dc3-4e5398686caa · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.956883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.956883Z digest=sha256:e7fc91d5ba4c4ad0336af3c0ffbc7006ed04607c7273aab7403cb9fa6e48be9f

Observation 0f5cd850-a1ea-407f-85f6-db18870e19c9 · outbound

This paper cites Vila: On pre-training for visual language models.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vila: On pre-training for visual language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.961049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.961049Z digest=sha256:2ef63b5ff34b6fbd2222a23e63fa1739d41ea988a3822e6d9fee8c250ca88a92

Observation 996261e8-7526-41ee-a2bd-20522179c51d · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.964786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.964786Z digest=sha256:1dd61fbb2817da2f088d125748c164ea610d524ba72c6256f1c35041973f26d7

Observation da6b1330-cdc2-43e2-8445-7722539e026e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Learning transferable visual models from natural language supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.968497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.968497Z digest=sha256:f728f231fe3e5965b62697184154e6cdcd5e47addd3058aa46e41314340239d1

Observation 9e8723f9-ff7d-4355-be5c-e9272afdcde4 · outbound

This paper cites Fast Segment Anything.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Fast Segment Anything

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.972360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.972360Z digest=sha256:bde82e7dfb8f6f5dc8ec9d76bac3440dd20710f19d9756e4bb002896053894d8

Observation 3efa4660-06ae-47dd-9c5c-2562faa81aa3 · outbound

This paper cites Embodied- sam: Online segment any 3d thing in real time.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Embodied- sam: Online segment any 3d thing in real time

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.761723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.976611Z digest=sha256:f23c91b87a54e4c533ad816667a9d0cb6b6aa7a03e370661168ae92ee8e5d49f

Observation b51ff56e-42fa-4640-a62e-899912c30b18 · outbound

This paper cites g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:59:11.320907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.980663Z digest=sha256:f8f9d47ad2d399ae58110f7b85622b31f31844e6558777f0b338bfa759a54fec

Observation 8a0a6a91-d385-4944-8dd5-5be7d7a306d6 · outbound

This paper cites Vln bert: A recurrent vision-and-language bert for navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vln bert: A recurrent vision-and-language bert for navigation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.750903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.984771Z digest=sha256:6472271727b83577762a51398885d1605e8c40f8a9e9bc619921e3e09f821188

Observation d746555c-a006-4a4d-9404-4d5878ab8a2f · outbound

This paper cites History aware multimodal transformer for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation History aware multimodal transformer for vision-and-language navigation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.988386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.988386Z digest=sha256:fa3434e1846ed3ab65eb9d9c04e41b7804afdc7bba3ab09195129041e5655a13

Observation 131334fe-8c88-4760-9047-80e969dac600 · outbound

This paper cites Think global, act local: Dual-scale graph transformer for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Think global, act local: Dual-scale graph transformer for vision-and-language navigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.992525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.992525Z digest=sha256:17f1cf59d10470d80373bf33e6960c6e67193076506fc5f6b1a48366562c2f2f

Observation 3bdf7c7a-c675-42f4-b41c-a9caadf81c4c · outbound

This paper cites Vision- and-language navigation via causal learning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vision- and-language navigation via causal learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.995663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.995663Z digest=sha256:4c41efa897c27ae834f78567f1204076ba644dcd66279613565f1c75f36c8e0b

Observation eb7cb573-bf4a-4532-9ac4-ba7134ed269c · outbound

This paper cites Hop+: History- enhanced and order-aware pre-training for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Hop+: History- enhanced and order-aware pre-training for vision-and-language navigation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.720871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.998621Z digest=sha256:0db5c78e4d51f10e6eddb442ffc5a7f259bad2a458c97f636f95e9cab9fa37ba

Observation 2620a608-4819-4edf-b17a-1b0256c8985e · outbound

This paper cites Bird’s-eye-view scene graph for vision- language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Bird’s-eye-view scene graph for vision- language navigation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.709765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.001793Z digest=sha256:480cbb3fa9fe65bfee89a5af2dc6ee7b6c22ced305cf32ac0c333bfc3d735ef7

Observation fe999097-561a-4e17-811a-ea95375f2453 · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Matterport3d: Learning from rgb-d data in indoor environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.004945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.004945Z digest=sha256:cd07a0b56b2c3fd0a335044f41a0d3c454dbeaeda103bea69095a1be33d6a3c6

Observation ba6d9897-ac8a-498f-827d-5d122622a9d3 · outbound

This paper cites Habitat: A platform for embodied ai research.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Habitat: A platform for embodied ai research

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.008702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.008702Z digest=sha256:32b14c421b7122a2722c51c694148e3be464ccc3f2b28b7190ddfd2bbedbfb72

Observation 3826abaa-d65d-423e-90ed-837387ce4795 · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.685462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.011706Z digest=sha256:34b9d0ced09400cfd8a88f5e05f0f691892d8d088a24d4e08b135979fa4d2722

Observation 0dd1d137-ae3a-48d7-a8fd-b90898c06c12 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Gridmm: Grid memory map for vision-and-language navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.015092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.015092Z digest=sha256:7925fb4cde2cdcfe8b3463a25c992dbf01b7d9b348496a93930dfaf39680e2fb

Observation 32cf35d3-5940-4fe4-8d7c-d0e5f10a224b · outbound

This paper cites Etpnav: Evolving topological planning for vision-language navigation in continuous environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Etpnav: Evolving topological planning for vision-language navigation in continuous environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.018222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.018222Z digest=sha256:21a2abf2e4b8985c78e9d57cb26a1654350f2397122e8fa91c30e14047059a9f

Observation 9026c242-c737-4e68-bdb4-8cb23b074a4f · outbound

This paper cites Bevbert: Multimodal map pre-training for language-guided navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Bevbert: Multimodal map pre-training for language-guided navigation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.661909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.021346Z digest=sha256:ec7686f5157e80fc30fe2f994c1a785b08fa6ef2d6ba5664c7d995dff682b966

Observation 6d3d0a0b-aaeb-4744-9d4f-1dbf17bd71f3 · outbound

This paper cites Sim-to-real transfer via 3d feature fields for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Sim-to-real transfer via 3d feature fields for vision-and-language navigation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.650854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.024387Z digest=sha256:45b5863a27b2996fda68982d111e28b71e72472259dfc0b6e580269de706bdf8

Observation 3b312087-5dd7-42f7-ae1f-96e991bd4927 · outbound

This paper cites Open-nav: Exploring zero-shot vision-and-language navigation in continuous environment with open-source llms.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Open-nav: Exploring zero-shot vision-and-language navigation in continuous environment with open-source llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.640342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.028544Z digest=sha256:fa3e4373f89f4263fd36a50dc520118d2ee73763c1569d09de4aea4e266cfb7a

Observation 0cfb4c6b-3c7b-40f8-8066-1e7ab584f861 · outbound

This paper cites NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.033358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.033358Z digest=sha256:95a6264eae2ada5c51c12e9153c3d28de19598e7fe360d296674e007b1333ed9

Observation bf22fda2-2145-4d70-ab77-38def4453fd7 · outbound

This paper cites an unresolved cited work.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:59:11.629551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.037712Z digest=sha256:ad68e77e29d63d4fab509df16e1d6a5e7af05d5804707b18b13aee00b9765484

Observation 9f6a8151-3867-41fe-81e5-0f6fdbd0d5d9 · outbound

This paper cites an unresolved cited work.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:59:11.617339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.041499Z digest=sha256:eda646003e5d5308f5c07a880782d9b3fde94a175a242d29a59ec463dfffeeec

Observation 725ba98b-b1d8-4018-aa3a-29813f2fba55 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.045479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.045479Z digest=sha256:e6a370b5a5fcf38d31a4855f0b44d59bccb959e84f7771099952c65cd53ca459

Observation 42c64cff-a0f4-41e9-a8e7-87ab19b0a4f7 · outbound

This paper cites Improved baselines with visual instruction tuning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Improved baselines with visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.049094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.049094Z digest=sha256:a06d59f83f07b26eb4cd7e30bbe5cd943c487bbd22fbc6a786d92081882bd39b

Observation a916a59d-7d46-48e3-aac3-24f8236b1d8c · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.053125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.053125Z digest=sha256:3992214f2d3108fd3eaf54d485b9e5a8bf6e70764d624ee90ae87e3f17e5d22e

Observation b171e546-d8cd-4076-8a8f-7972e223fc3f · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.056801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.056801Z digest=sha256:7e72cc9dab123b3402387426d31211e1e75ce7ddc09d2277b1215ce0a1a5e7d6

Observation 8eecf82a-69cd-4065-92be-1720cb442572 · outbound

This paper cites An embodied generalist agent in 3d world.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation An embodied generalist agent in 3d world

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.583760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.061079Z digest=sha256:523087fe21866dd381e5dfc34a289aea10ab00b3248319a533efa1780304ecc0

Observation 64c95c6e-091b-4714-81cf-6abfc2239f41 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.064457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.064457Z digest=sha256:d3bfbb827e637a2b10ea77958baf97e617181693dae4e346ee0645e5205d78d8

Observation c782b4a2-c2ee-4cfe-b079-3788ba651885 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation 3d-llm: Injecting the 3d world into large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.068061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.068061Z digest=sha256:07c4fba8a8cac0228d1c62a9877794917835846a788edda57629e7d6c74b4f84

Observation 950c6378-f069-49eb-9012-56ffe44cb7b6 · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.071380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.071380Z digest=sha256:f6828c830f3bc5239a89fe3ccd089cc418bb563cc3805a0c76ad9e5c46f64cb8

Observation 346d152d-5f34-45c0-a6a0-ab946f40cebc · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.074859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.074859Z digest=sha256:2bde02e24ce497c05d94e6a0024ad0cacb293866984cfaaaa7de6464978855ab

Observation 14734b56-16dd-4f6c-a957-0ba37e13f1e0 · outbound

This paper cites Lookahead exploration with neural radiance representation for continuous vision-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Lookahead exploration with neural radiance representation for continuous vision-language navigation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.560550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.078810Z digest=sha256:bb2fc6fef59d6c1a6102ccec8941af75232ef7ec1a897307ec181fda3863a707

Observation 474084a7-5bbd-4641-92f0-e2fe4e7770e0 · outbound

This paper cites Learning Generalizable Feature Fields for Mobile Manipulation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Learning Generalizable Feature Fields for Mobile Manipulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.081771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.081771Z digest=sha256:69bed92cf5a62244b11d8784fc0a31e0e73f221c4c435c36db8b402665615d3b

Observation ecc44f67-62e1-4f06-9120-7508788c4a01 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.085792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.085792Z digest=sha256:89f929a69dc6f1cc32e3c5bece8a17fe99aa02af918e9509e0f0fcadc5907023

Observation 35e81402-52cf-4eca-a141-d8ad255f3ee7 · outbound

This paper cites Habitat- matterport 3d semantics dataset.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Habitat- matterport 3d semantics dataset

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.089780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.089780Z digest=sha256:e964a64d298973efdbab3b847e4e8b7c663f6714813105b18a409a1ec5e1ef79

Observation 1c8ec00f-2ea0-492b-9143-af31ef56b098 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Rio: 3d object instance re-localization in changing indoor environments

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.532040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.093803Z digest=sha256:c24ede7c9a98a625182fd3ea3070b81a27f9a1eeb3c5f2d5d86b11fb25f6d168

Observation c4f3ca20-cb4a-47ad-843d-610c1d23d289 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.517526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.098172Z digest=sha256:1ee3e92fa6b7c2d8e4edb25613a6ecf6966e2614bdbb0eceb643053d27679f5e

Observation 5d428d5f-01d3-4f7f-b810-584d1861ed4b · outbound

This paper cites Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.102040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.102040Z digest=sha256:b88a2c8d8ee3709e04e9e46f5e20c0b7ff772f9a5d7dc590f269e1519c01a586

Observation 80f4a096-6e24-4429-aae3-3da3c5fb0227 · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Xtuner: A toolkit for efficiently fine-tuning llm

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.106040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.106040Z digest=sha256:48a038c315ba938668dc6ed989fa8b2db6c70585fb767aa7e499ae873368a7bd

Observation c1f25e16-a2e5-4f69-a7bd-1be5f863865c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.110249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.110249Z digest=sha256:010fd086ce48dba576633911e5831fd2f991aecf5271fb855fff8e26ee67118e

Observation ee9c700a-27dd-4ec4-bc40-bd90809ce39b · outbound

This paper cites Cross-modal map learning for vision and language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Cross-modal map learning for vision and language navigation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.490076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.116507Z digest=sha256:33970ecf6999f70d68871102c3ddb80db41acdd7378eb7b325c3f0b73dafdff6

Observation 0745b914-d419-47a4-8055-e89f749e0b04 · outbound

This paper cites Weakly-supervised multi-granularity map learning for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Weakly-supervised multi-granularity map learning for vision-and-language navigation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.471819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.122152Z digest=sha256:98a6793b29f797ab07047cff3165684c2193fba6da527162fe144cca3c64c2fa

Observation 480992bc-787d-44d2-90d6-9fe4a143bf06 · outbound

This paper cites Instructnav: Zero-shot system for generic instruction navigation in unexplored environment.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Instructnav: Zero-shot system for generic instruction navigation in unexplored environment

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.456951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.126143Z digest=sha256:a195f0fb640089436a6e5007746a504ce1a03260c9166f0ebc0be86f07e14210

Observation eeb33017-9d81-4ecc-ac59-29b2f28d52c1 · outbound

This paper cites Affordances-oriented planning using foundation models for continuous vision-language nav- igation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Affordances-oriented planning using foundation models for continuous vision-language nav- igation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.444755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.129374Z digest=sha256:4ae4183e61977531c7523ac2bdd4f2571148b7f32c283b7908e23d5d5107d0ad

Observation 8a683655-5628-45f2-840f-6cbd29249200 · outbound

This paper cites Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.132492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.132492Z digest=sha256:187ffd64b036fda5e1fb94f98e0514f910266c62401a8339825f2d40a36c4b99

Observation 0e832c77-1435-4575-82fc-4d87149dd052 · outbound

This paper cites Structured3d: A large photo-realistic dataset for structured 3d modeling.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Structured3d: A large photo-realistic dataset for structured 3d modeling

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.425999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.135977Z digest=sha256:01b50c541f248b503c64143243566089788df10781484db2a773556d500c686e

Observation d0d8957f-b1f3-40df-8ac2-c8058f3bad83 · outbound

This paper cites Scaling data generation in vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Scaling data generation in vision-and-language navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.412170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.142816Z digest=sha256:39ae29b236c8e9d0f09cac4f0c8fb35ebb2fa78471221587c2c90d3d089ffd27

Observation 5bd952d9-d33d-4e06-a66b-e976435b16bf · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation A reduction of imitation learning and structured prediction to no-regret online learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.148767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.148767Z digest=sha256:329c3949b2618e346b82c9307717a004fa4f6291d5535b62bc950edaff4fd8ad

Observation 1ecbbded-bccb-4ae6-a056-5c7da6d6e675 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Adafactor: Adaptive learning rates with sublinear memory cost

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.157244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.157244Z digest=sha256:4ea6fe009b7ce0cc920a992db28e603f6f122beabc411738ff6dc124412fcf49

Observation 1f72137b-27b2-4585-9596-37037a535bad · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Training Deep Nets with Sublinear Memory Cost

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.160541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.160541Z digest=sha256:3cde085994eaf732baef36e257763205db6b8c57969c456dcd7a261276770cf2

Observation 3d4c7a5b-a3b9-4828-a977-cf88c588ed25 · outbound

This paper cites Dynamem: Online dynamic spatio-semantic memory for open world mobile manipulation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Dynamem: Online dynamic spatio-semantic memory for open world mobile manipulation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.384825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.164370Z digest=sha256:85ab3169c40a37b96230c1bacaa4d5e97cd8f5cb3db7990bdd9909f83c3b14fd

Pith citing papers

Observation 10920aae-b0b6-40f6-bfb5-c4bd0db6c11c · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.944979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:d931bd450a04959eeae5b4e15da88a3d83c8be74084586d69f98f7894c083c51

Observation a7609af3-49f1-4cfe-b0ae-76e5d5b2bd0e · inbound

What Limits Vision-and-Language Navigation ? cites this paper.

What Limits Vision-and-Language Navigation ? Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:02:32.616304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T17:59:58.891151Z digest=sha256:9327dee6a3d420a2cf2e37ee9818ec8ac957a0309e92c49917d8a6f42d7fcc26

Observation 0e70a95a-c0ab-45f8-ba18-94a5f067105f · inbound

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation cites this paper.

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:46:10.563944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T06:45:35.920991Z digest=sha256:475e67e508ed47f3316adea95e69e0952d08faf8953858b65bf163b8283e2242

Observation 60fad76a-54e2-40b7-abc8-9a427e75a6d3 · inbound

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation cites this paper.

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.357198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:33:32.671550Z digest=sha256:e5a218d44c44d97c26061e08d69fff0daadf1d825941510a5b75421cb3bd124c

Observation a7332381-e0b8-49e3-94e3-26f339981b26 · inbound

NoPA: Non-Parametric Online 3D Scene Graph Generation cites this paper.

NoPA: Non-Parametric Online 3D Scene Graph Generation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:37:03.076919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T14:36:02.425820Z digest=sha256:c72032f0eb46620d1c05368079d5ea49fb96d6f41fd4fc4da3b4d5e903aa53c3

Observation 6068e15d-7e0d-4aa9-ae6b-46132f95e96f · inbound

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation cites this paper.

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:39.821353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:04:39.821353Z digest=sha256:f183824d332adbca00a5ab568b18143f0383aca9e7497c737cd0e0cc4d58cbd9