Pith. sign in

Paper Citation Record · LEDGER

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

As of 16 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 6 inbound Pith citation observations for arXiv:2505.11383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11383 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:11.164370Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:04:39.821353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:37:03.075467Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01d8955a-8656-45e7-b7e4-24c4531b860d · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.849564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.849564Z digest=sha256:316cbfb152dae2d72ddb3c9b0d4386747f63dde1587ac3dbf44eeb78789f9496

Observation d038db71-6345-4fb3-ad16-aeb9a5715e14 · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Reverie: Remote embodied visual referring expression in real indoor environments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.903499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.903499Z digest=sha256:26603c507914703d31bb439a0bd11403ffbe6867b8447ed95368667acb5aef9c

Observation d9e9caf6-4b13-4bed-8dd9-b3e6c09b2f2c · outbound

This paper cites Beyond the nav- graph: Vision-and-language navigation in continuous environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Beyond the nav- graph: Vision-and-language navigation in continuous environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.936174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.936174Z digest=sha256:1578c4f5422ca1ad1ecfbee25777d18ade8ef1613f2ddc5ac80b4fd96feac855

Observation 45278d17-f96a-4c42-baeb-20620e5e5c34 · outbound

This paper cites NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.939678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.939678Z digest=sha256:2c0dd43f2742e5980906d41b742aec420dd37cd12af2aca7728921f6ce4a8adf

Observation 705f5863-bebe-4c38-8383-8273c10a40e6 · outbound

This paper cites Navid: Video-based vlm plans the next step for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Navid: Video-based vlm plans the next step for vision-and-language navigation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.803239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.943451Z digest=sha256:36e229168d783d64595f47feb606930b653888441c04bf6506e8901ceb847566

Observation bcaf25e4-d524-4bf8-a3f5-77a3e0cbd0ec · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.947005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.947005Z digest=sha256:ec1fe11afd5593a82d9ab012601285afd93133e8299b4acba8d1d8e1f7d632cf

Observation d61b36e6-545f-431c-a197-7e3e499c7063 · outbound

This paper cites Navila: Legged robot vision-language-action model for navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Navila: Legged robot vision-language-action model for navigation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.792553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.952731Z digest=sha256:68f8f4cac8874ce4d567778be9ee2e6cd10e0371c7583f52ea829c45ecf7edc2

Observation b3b93662-d1aa-4d52-8dc3-4e5398686caa · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.956883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.956883Z digest=sha256:5649cf255359a15f1e22e0e556c8df19daf1bfabec9b2f31656de0f1395c53b1

Observation 0f5cd850-a1ea-407f-85f6-db18870e19c9 · outbound

This paper cites Vila: On pre-training for visual language models.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vila: On pre-training for visual language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.961049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.961049Z digest=sha256:3386c0a0ec0bff43d05961991f742808e7c47b6fc4b971270f25d96429e714e7

Observation 996261e8-7526-41ee-a2bd-20522179c51d · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.964786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.964786Z digest=sha256:8ffc756d26df89d5cf71a805f4599cbe53eb0bc1c81b3432bb4e35633948ff3c

Observation da6b1330-cdc2-43e2-8445-7722539e026e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Learning transferable visual models from natural language supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.968497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.968497Z digest=sha256:d0452b3658b65f6e64fb600292a120fc44d52fd4f20d3b89679d2144c756f88a

Observation 9e8723f9-ff7d-4355-be5c-e9272afdcde4 · outbound

This paper cites Fast Segment Anything.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Fast Segment Anything

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.972360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.972360Z digest=sha256:28130e983aade755224dccf307228f2404e72922a2dd051e2d43fd7f18b6e662

Observation 3efa4660-06ae-47dd-9c5c-2562faa81aa3 · outbound

This paper cites Embodied- sam: Online segment any 3d thing in real time.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Embodied- sam: Online segment any 3d thing in real time

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.761723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.976611Z digest=sha256:93e01b234776c1ad1b7cd9f11962d836f5d1bcf39226a4cb22cc2879ac742337

Observation b51ff56e-42fa-4640-a62e-899912c30b18 · outbound

This paper cites g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:59:11.320907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.980663Z digest=sha256:870ca73c1680ca45786c4f28fc061a29a3c63d6805db653eb04b178ae7279318

Observation 8a0a6a91-d385-4944-8dd5-5be7d7a306d6 · outbound

This paper cites Vln bert: A recurrent vision-and-language bert for navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vln bert: A recurrent vision-and-language bert for navigation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.750903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.984771Z digest=sha256:da6a66baed4c8ca9dde2a90f52a2cdbdc431a33a9e36d6485cf402562e4ea3c2

Observation d746555c-a006-4a4d-9404-4d5878ab8a2f · outbound

This paper cites History aware multimodal transformer for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation History aware multimodal transformer for vision-and-language navigation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.988386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.988386Z digest=sha256:cbd7956a7c28c747fa579f67765b4b05dd4fa79eb18e65378223a6c6757c6ec4

Observation 131334fe-8c88-4760-9047-80e969dac600 · outbound

This paper cites Think global, act local: Dual-scale graph transformer for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Think global, act local: Dual-scale graph transformer for vision-and-language navigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.992525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.992525Z digest=sha256:8f5e4e9440c29ad8c7083354c68c7f6f6154b7c436ad2d44c7d21afd8289fdee

Observation 3bdf7c7a-c675-42f4-b41c-a9caadf81c4c · outbound

This paper cites Vision- and-language navigation via causal learning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Vision- and-language navigation via causal learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:10.995663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:10.995663Z digest=sha256:3b852a85f36ff2571f4b49edd2434fdd5ec56636a5d352ef8e366e1e2565c8ef

Observation eb7cb573-bf4a-4532-9ac4-ba7134ed269c · outbound

This paper cites Hop+: History- enhanced and order-aware pre-training for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Hop+: History- enhanced and order-aware pre-training for vision-and-language navigation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.720871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:10.998621Z digest=sha256:7610d82ce2dc8299726de099f28834f3b0003f1db85e45426473a202b6d39be0

Observation 2620a608-4819-4edf-b17a-1b0256c8985e · outbound

This paper cites Bird’s-eye-view scene graph for vision- language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Bird’s-eye-view scene graph for vision- language navigation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.709765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.001793Z digest=sha256:82ff23306b427e5f5d754b776304328a4c4cf5f82a91e8f6c29036c888e80846

Observation fe999097-561a-4e17-811a-ea95375f2453 · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Matterport3d: Learning from rgb-d data in indoor environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.004945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.004945Z digest=sha256:e595d4610879f23ef57c959079baa296ea17648acc15b4a60f4083076de165a0

Observation ba6d9897-ac8a-498f-827d-5d122622a9d3 · outbound

This paper cites Habitat: A platform for embodied ai research.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Habitat: A platform for embodied ai research

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.008702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.008702Z digest=sha256:3c08f82724acc1c1e30296da51832f8c4ae957fcdf8e716256aebff61381053e

Observation 3826abaa-d65d-423e-90ed-837387ce4795 · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.685462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.011706Z digest=sha256:4dc79a7b37311a142d875b0e456eeef53f02ffae1cc9782b389f0d53c4f16687

Observation 0dd1d137-ae3a-48d7-a8fd-b90898c06c12 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Gridmm: Grid memory map for vision-and-language navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.015092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.015092Z digest=sha256:7286acfefb9a68e9ec5d031f34f72286418d4bcbafccd392bf5620adfd29d823

Observation 32cf35d3-5940-4fe4-8d7c-d0e5f10a224b · outbound

This paper cites Etpnav: Evolving topological planning for vision-language navigation in continuous environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Etpnav: Evolving topological planning for vision-language navigation in continuous environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.018222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.018222Z digest=sha256:fc62ae3976eddc2291a5038f8a0d700a42b8303000efe7497bf6e676b1bbe474

Observation 9026c242-c737-4e68-bdb4-8cb23b074a4f · outbound

This paper cites Bevbert: Multimodal map pre-training for language-guided navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Bevbert: Multimodal map pre-training for language-guided navigation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.661909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.021346Z digest=sha256:cd16f28c31e878592e62c53ed77c4791c2bb3a5ddd80707191b94af2b086eb14

Observation 6d3d0a0b-aaeb-4744-9d4f-1dbf17bd71f3 · outbound

This paper cites Sim-to-real transfer via 3d feature fields for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Sim-to-real transfer via 3d feature fields for vision-and-language navigation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.650854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.024387Z digest=sha256:6843ccba47a2b459547ae5af831683766e06e8029348d3059db8ed9b20fd2a16

Observation 3b312087-5dd7-42f7-ae1f-96e991bd4927 · outbound

This paper cites Open-nav: Exploring zero-shot vision-and-language navigation in continuous environment with open-source llms.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Open-nav: Exploring zero-shot vision-and-language navigation in continuous environment with open-source llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.640342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.028544Z digest=sha256:c2bf0483d8e59c79949dd17265f80d0bca1edb38191cbcb83dbdd5c2fa84bd13

Observation 0cfb4c6b-3c7b-40f8-8066-1e7ab584f861 · outbound

This paper cites NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.033358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.033358Z digest=sha256:30a21fcc97fe992d7f90edc61ac2cc9c3c41c0b6a61d6e67cf0d11b5a15cf3ff

Observation bf22fda2-2145-4d70-ab77-38def4453fd7 · outbound

This paper cites an unresolved cited work.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:59:11.629551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.037712Z digest=sha256:8b6f7c84d0a12e47cad45710511549d3f13be15fe22f63b0bc1e9eef88cc1df1

Observation 9f6a8151-3867-41fe-81e5-0f6fdbd0d5d9 · outbound

This paper cites an unresolved cited work.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:59:11.617339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.041499Z digest=sha256:88c51a62ac26b92b5a0b2cc4c0032f066fc44ea238e92499dd3f5f3b34813d06

Observation 725ba98b-b1d8-4018-aa3a-29813f2fba55 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.045479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.045479Z digest=sha256:57c6872820ed7ab3f9567b426c090cceea425d619d42f1f6040986bab7a6ebe2

Observation 42c64cff-a0f4-41e9-a8e7-87ab19b0a4f7 · outbound

This paper cites Improved baselines with visual instruction tuning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Improved baselines with visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.049094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.049094Z digest=sha256:0b6676a0ca1b9decc8100e9c719a5bcdd046968599883a1a13779bd8a2a5f462

Observation a916a59d-7d46-48e3-aac3-24f8236b1d8c · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.053125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.053125Z digest=sha256:eff7bf26cb4ddb62127cf84d74e8aeca2aeccf86685ba3721500281c74006371

Observation b171e546-d8cd-4076-8a8f-7972e223fc3f · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.056801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.056801Z digest=sha256:4a9f621d598a8ed0c0a9390f2554dd8e3fa90dfbf4e8c6ed93ca297975469d95

Observation 8eecf82a-69cd-4065-92be-1720cb442572 · outbound

This paper cites An embodied generalist agent in 3d world.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation An embodied generalist agent in 3d world

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.583760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.061079Z digest=sha256:87e6f4b08dd6999275ba54c349181058d18d83b48d4e41f1eb1c7297494db692

Observation 64c95c6e-091b-4714-81cf-6abfc2239f41 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.064457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.064457Z digest=sha256:06c27d826c99dc07e06328d4c331b78430e0b6cfb1c02bdcff46d3684781a962

Observation c782b4a2-c2ee-4cfe-b079-3788ba651885 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation 3d-llm: Injecting the 3d world into large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.068061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.068061Z digest=sha256:876198d10f9fb367d91ccefd125ad46b32d63440f56d937c92e7b4152a2712a7

Observation 950c6378-f069-49eb-9012-56ffe44cb7b6 · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.071380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.071380Z digest=sha256:57ccfb457773a81b58a16c0586d4c5e489ddcbcb7264ec7d379be94326fe57df

Observation 346d152d-5f34-45c0-a6a0-ab946f40cebc · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.074859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.074859Z digest=sha256:0b7ee51f92104ecfaf732078a7fc82707b1b92ac86e6e0c11eb3cd75b97a2aa2

Observation 14734b56-16dd-4f6c-a957-0ba37e13f1e0 · outbound

This paper cites Lookahead exploration with neural radiance representation for continuous vision-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Lookahead exploration with neural radiance representation for continuous vision-language navigation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.560550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.078810Z digest=sha256:aa139e88205a878007ec69821ca9846a6ea19ac55780d495c1f6db26f64ce352

Observation 474084a7-5bbd-4641-92f0-e2fe4e7770e0 · outbound

This paper cites Learning Generalizable Feature Fields for Mobile Manipulation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Learning Generalizable Feature Fields for Mobile Manipulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.081771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.081771Z digest=sha256:a7e7195cd421a3ebb28830f4340220d9a819b6093c3375b3478e04ed944bd4fc

Observation ecc44f67-62e1-4f06-9120-7508788c4a01 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.085792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.085792Z digest=sha256:9d77831f40352393555eee135ffb4bef050937125e6dac6419921a6937ac7d88

Observation 35e81402-52cf-4eca-a141-d8ad255f3ee7 · outbound

This paper cites Habitat- matterport 3d semantics dataset.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Habitat- matterport 3d semantics dataset

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.089780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.089780Z digest=sha256:184cee253840780d25cc5651a49bf8591b6f901424d1dd01ff4063501b7355a1

Observation 1c8ec00f-2ea0-492b-9143-af31ef56b098 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Rio: 3d object instance re-localization in changing indoor environments

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.532040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.093803Z digest=sha256:015f44d1f16168a382d5180caf1c96dee28b8e715d0928f3a666175d7a91cf33

Observation c4f3ca20-cb4a-47ad-843d-610c1d23d289 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.517526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.098172Z digest=sha256:62a819a47a6f889a1da457b831b0f0a460476807dad10a46924efc0b4451744c

Observation 5d428d5f-01d3-4f7f-b810-584d1861ed4b · outbound

This paper cites Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.102040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.102040Z digest=sha256:9df606a7786d5a58c8183606332fbb0edef3b5153394bf4502dd9af91679273e

Observation 80f4a096-6e24-4429-aae3-3da3c5fb0227 · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Xtuner: A toolkit for efficiently fine-tuning llm

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.106040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.106040Z digest=sha256:f88f0a8a0b15f2c401e070289977690748f832708f57978c764e165c29b8fad7

Observation c1f25e16-a2e5-4f69-a7bd-1be5f863865c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.110249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.110249Z digest=sha256:318d091746cc844881a595121aff054eba0a0c0a187507e8b910bbc4012460b4

Observation ee9c700a-27dd-4ec4-bc40-bd90809ce39b · outbound

This paper cites Cross-modal map learning for vision and language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Cross-modal map learning for vision and language navigation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.490076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.116507Z digest=sha256:9400dc5116751c73513441e0959c5400bd17a8a5b65aba25eb0f2b206f005132

Observation 0745b914-d419-47a4-8055-e89f749e0b04 · outbound

This paper cites Weakly-supervised multi-granularity map learning for vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Weakly-supervised multi-granularity map learning for vision-and-language navigation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.471819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.122152Z digest=sha256:ead2c2a2cb1edb3b3f70a67b35e6219efaad5fe6c2d6169c860c41ff9052acdc

Observation 480992bc-787d-44d2-90d6-9fe4a143bf06 · outbound

This paper cites Instructnav: Zero-shot system for generic instruction navigation in unexplored environment.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Instructnav: Zero-shot system for generic instruction navigation in unexplored environment

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.456951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.126143Z digest=sha256:e10fef528b80e3489623317ee4c8628b41191ff45bcea692b663959877c34c11

Observation eeb33017-9d81-4ecc-ac59-29b2f28d52c1 · outbound

This paper cites Affordances-oriented planning using foundation models for continuous vision-language nav- igation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Affordances-oriented planning using foundation models for continuous vision-language nav- igation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.444755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.129374Z digest=sha256:04fdd36c00d046648794eb43143241fc209dbceff1065a11869c52ec489d394f

Observation 8a683655-5628-45f2-840f-6cbd29249200 · outbound

This paper cites Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.132492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.132492Z digest=sha256:29973d5e202bba2952f6788ef5473024e35cce66327010c698c2f28f58dace0b

Observation 0e832c77-1435-4575-82fc-4d87149dd052 · outbound

This paper cites Structured3d: A large photo-realistic dataset for structured 3d modeling.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Structured3d: A large photo-realistic dataset for structured 3d modeling

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.425999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.135977Z digest=sha256:1e293a1cfbed0b436cfc5a5387f833de250eb3d4ddca0325628edb9bd7696c91

Observation d0d8957f-b1f3-40df-8ac2-c8058f3bad83 · outbound

This paper cites Scaling data generation in vision-and-language navigation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Scaling data generation in vision-and-language navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.412170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.142816Z digest=sha256:9280da59454fff02bdbb11643f9a8006a0e87cf726eab8b1a2a2bc681e8976f0

Observation 5bd952d9-d33d-4e06-a66b-e976435b16bf · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation A reduction of imitation learning and structured prediction to no-regret online learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.148767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.148767Z digest=sha256:1699ea8aa4b978dbc6d8a7e86e2a66b9d5838909171cdb1113c9e5c9bd8a6637

Observation 1ecbbded-bccb-4ae6-a056-5c7da6d6e675 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Adafactor: Adaptive learning rates with sublinear memory cost

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.157244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.157244Z digest=sha256:ddbe82cbab698d4aafc7852e4c66a7d15f0ca2939b8e80e6650836673cd0441d

Observation 1f72137b-27b2-4585-9596-37037a535bad · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Training Deep Nets with Sublinear Memory Cost

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.160541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.160541Z digest=sha256:90fbffce53bb7226a6d7f81756cef264c31b01ab856d607c2e20c2e40aee9ea7

Observation 3d4c7a5b-a3b9-4828-a977-cf88c588ed25 · outbound

This paper cites Dynamem: Online dynamic spatio-semantic memory for open world mobile manipulation.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Dynamem: Online dynamic spatio-semantic memory for open world mobile manipulation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:59:11.384825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:59:11.164370Z digest=sha256:a842526426faa866587b289214535b4db9c4a522be2b87b8872d4d63c0cd3428

Pith citing papers

Observation 10920aae-b0b6-40f6-bfb5-c4bd0db6c11c · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.944979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:ecc20e0dac64dcec6ff62ca78796d502a68dd219780b185dbccae52b4a7036a2

Observation a7609af3-49f1-4cfe-b0ae-76e5d5b2bd0e · inbound

What Limits Vision-and-Language Navigation ? cites this paper.

What Limits Vision-and-Language Navigation ? Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:02:32.616304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T17:59:58.891151Z digest=sha256:0736296e5b660c6ed959f07d9299291abe4f9a8e865e5a09f50eca6a67bc673a

Observation 0e70a95a-c0ab-45f8-ba18-94a5f067105f · inbound

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation cites this paper.

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:46:10.563944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T06:45:35.920991Z digest=sha256:c2efdf2b818016e5e13786647d7ff94cbaffebfaeb27384b155c504d936856a0

Observation 60fad76a-54e2-40b7-abc8-9a427e75a6d3 · inbound

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation cites this paper.

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.357198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:33:32.671550Z digest=sha256:6939a8fa78194d199f55338630fa227dd414d03817f5ceac65a4d3a4a1d41dd0

Observation a7332381-e0b8-49e3-94e3-26f339981b26 · inbound

NoPA: Non-Parametric Online 3D Scene Graph Generation cites this paper.

NoPA: Non-Parametric Online 3D Scene Graph Generation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:37:03.076919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T14:36:02.425820Z digest=sha256:2cac87f2cd4cb547a65c20eb0b3ee3853a70b6e3ee9d374d5a518d0d0605aa31

Observation 6068e15d-7e0d-4aa9-ae6b-46132f95e96f · inbound

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation cites this paper.

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:39.821353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:04:39.821353Z digest=sha256:d238e2500afe4ffa3cab2eb8e95ffd474c1b777394efa23cf2383f51c33324b2