Pith. sign in

Paper Citation Record · LEDGER

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

As of 8 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2506.14507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14507 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:01.919344Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:39:59.878169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T23:13:15.774007Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0499a38-1ba4-43d3-b5d9-aeb475c79970 · outbound

This paper cites Getting vit in shape: Scaling laws for compute-optimal model design,.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Getting vit in shape: Scaling laws for compute-optimal model design,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.555563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:00.467925Z digest=sha256:7987053de0b03fb2cbb410f692a89f1a993feba2a9f28a231164971d86659d77

Observation 40e9d7d8-c1a9-4405-acf3-7240dd96a381 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabilities, 2024.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Spatialvlm: Endow- ing vision-language models with spatial reasoning capabilities, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.531508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:00.566662Z digest=sha256:36a3c2baad1e32bf554155b94b9dc151f76d5f63c4642dbe0026bb0c9ed47a81

Observation ed17b4aa-87dc-430f-a22e-7824d1b5cc3f · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.630568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.630568Z digest=sha256:245289c1adc325732c4288aeaf4be2f4b254fb2a56942313e1aa142e44eb601b

Observation 3c22e683-59c9-4af2-9a43-b6de6d942bcc · outbound

This paper cites CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.698604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.698604Z digest=sha256:46e482b2255848db8934d9ab665b6772194969501ba6eb68295916982cac20bc

Observation e5e2af88-9c60-4abf-99d4-6749b39fe012 · outbound

This paper cites Learning latent dynamics for planning from pixels, 2019.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Learning latent dynamics for planning from pixels, 2019

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.510339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:00.783826Z digest=sha256:9eda922d8fd392601aef2784ecf832d2e1ab7ef1efb3d69a2766c2ec3061fc8a

Observation 481ba7a2-3ddb-421e-8953-a4f958356b26 · outbound

This paper cites Mastering Diverse Domains through World Models.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Mastering Diverse Domains through World Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.842906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.842906Z digest=sha256:9ebcc1d44f00ab2a9408a7dbb08f156a65c2504a5f16c601c92c8ce94ad746bf

Observation 0040f366-669a-4909-8d08-e03002d5417c · outbound

This paper cites Visual language maps for robot navigation, 2023.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Visual language maps for robot navigation, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.486795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:00.924004Z digest=sha256:4a07e477bc0b48dc02bf4d35a577d65e67dd5f88b487859a9b5e5fa9952d5936

Observation 987b5551-2994-42c3-9cad-e7c8825761ad · outbound

This paper cites BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.963437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.963437Z digest=sha256:49311e582ee09ad1589f5561b2c6831a7c6f89133aa8d367cf4e7dc08efa597a

Observation a2d883d2-6d9b-4aea-8c14-1e0f207b4e51 · outbound

This paper cites ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.968433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.968433Z digest=sha256:e8f3b66fac5c01a2f8745eb1ac6049c8042f1aded17fc148be093aa5cf91125d

Observation c9a54b4f-259b-4157-ab69-694ee5098b09 · outbound

This paper cites Distilling Realizable Students from Unrealizable Teachers.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Distilling Realizable Students from Unrealizable Teachers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.023823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.023823Z digest=sha256:d987ce2b48f8a84b019a44c9dc1bddcf430f2f3eb9e25bfe9572b0ce7ba6fbf7

Observation 1049ce38-98fd-44b0-a912-d685f452942f · outbound

This paper cites Constrained Behavior Cloning for Robotic Learning.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Constrained Behavior Cloning for Robotic Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.104551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.104551Z digest=sha256:014aef8904399da7ac049b37a579bbf0a60c42a362d8b0f55ca63eb15ed0ae1a

Observation 8389b828-de7b-4414-97ee-71e84bdaa3ed · outbound

This paper cites ZSON: Zero-shot object-goal navigation using multimodal goal embeddings.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? ZSON: Zero-shot object-goal navigation using multimodal goal embeddings

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.465031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:01.190903Z digest=sha256:51bca8b6a2352d35cf21e886ab1ea5281691a1e69e602b541ad27c9418f714be

Observation d45a9c07-9fde-4449-b71f-79f99530b241 · outbound

This paper cites Orbit: A unified simulation framework for interactive robot learning environ- ments.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Orbit: A unified simulation framework for interactive robot learning environ- ments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.237600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.237600Z digest=sha256:bbc398128794428cd25fd751c55078c3367061e87f96d00f0a53b4369f1b11ca

Observation c28bb70d-a4d7-4b7d-93d9-cfd44f9e2958 · outbound

This paper cites Language-Conditioned Offline RL for Multi-Robot Navigation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Language-Conditioned Offline RL for Multi-Robot Navigation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.301764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.301764Z digest=sha256:899378320c632c643facadbf10300f36bfa3f4f6a1a94a937a7e8b366c1d6fd8

Observation 173f7ab1-8664-42c0-a57b-22d6abed694e · outbound

This paper cites R3m: A uni- versal visual representation for robot manipulation,.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? R3m: A uni- versal visual representation for robot manipulation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.446687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:01.359403Z digest=sha256:073626dbbea17c79c251ff7cf2828fa9a0f27c08310597b24cebb477aa449ebd

Observation 9032cfc7-51d1-4bf7-b85a-1c46bfcda798 · outbound

This paper cites Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.486633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.486633Z digest=sha256:cc5a4ecb77b715f9e86984752dc241d0dccfbbc627977454ee11eaf350cc1ec4

Observation 4dc5ed25-58df-4883-b710-8c2e4851e655 · outbound

This paper cites Isaac sim.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Isaac sim

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.425411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:01.494307Z digest=sha256:7387506a48e702e52c47656515bef1dc03aeb4c1bc9561334eb4921fb6173333

Observation 30a9500b-81ca-4169-880a-acde40926f46 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Learning Transferable Visual Models From Natural Language Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.525716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.525716Z digest=sha256:99b095383ca30031c53656459433f287b6e047658d5d4fdce377e4b621b83f7c

Observation c66757b3-508a-4717-aa66-2dc5e06524df · outbound

This paper cites Latent Plans for Task-Agnostic Offline Reinforcement Learning.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Latent Plans for Task-Agnostic Offline Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:02.127836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:01.592832Z digest=sha256:76dfc09a25dbf22ecdc3aae19ca779a2eeda819c5a56b3299f0e0040af66f977

Observation 3f7cb8c2-75c1-4d4f-921f-d2a1ed275cf4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.655232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.655232Z digest=sha256:e931d0e553b74ea945c4c7ac2f029bcc696c95e843a56be3e46d30d5ba0e2baf

Observation 3fa81a0a-5c6a-4212-8c6e-6968a2ccc98d · outbound

This paper cites Vlm-social-nav: Socially aware robot navigation through scoring using vision-language models, 2024.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Vlm-social-nav: Socially aware robot navigation through scoring using vision-language models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.404975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:01.734952Z digest=sha256:9928baab78354ccfbb62196ee3f9f042a9fea13f47551dd22b504e2d8ce930b3

Observation 775a35c4-692f-4f85-976b-691f70366aa1 · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.797762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.797762Z digest=sha256:a9028540459ee9247362e706a694ec7c9c0411b1b9cda8cf6809a7f4bba30c96

Observation c50d8b92-015b-4c1a-8568-5378e6677997 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Sigmoid Loss for Language Image Pre-Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.862440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.862440Z digest=sha256:1c41508a2317f526fef32c2cb0cd0f92f158ff24e26194d89a20784f2d5ebc69

Observation 8e4c8fb4-92e4-4e61-b02a-a8b3e0cf5ee3 · outbound

This paper cites RT-2: Vision- language-action models transfer web knowledge to robotic control.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? RT-2: Vision- language-action models transfer web knowledge to robotic control

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.385497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:01.919344Z digest=sha256:9ea6e4199553d144bb16ebd3903fb20b57e8abddd2e7631f1dce9b64958c2aff

Observation 41596763-d6ca-4f0a-ad20-f00e05c1f39b · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? R3M: A Universal Visual Representation for Robot Manipulation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.438793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.438793Z digest=sha256:ac472d089f9fb842f989904a08595c16b0bceaec5c9e6d5ce62512e1be1ae866

Observation b7bdfaf7-9860-469c-845b-9238aa34cd6f · outbound

This paper cites Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.522629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.522629Z digest=sha256:a1e50a29b8ae65057d6809f3d6a69d446627a82196ca36635770653ceee4a934

Pith citing papers

Observation bcff18d3-129c-4ce3-a3e5-0edbe42ff4a2 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

Reference 248

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:15.777167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:0f0c9fb42f7c34562ed9a20ad15281ab8b822fc9c24bf42e6ec3e219a82ae9c3

Observation b634dbf6-4d76-4ab7-98b0-f537bcf8a414 · inbound

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation cites this paper.

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

Reference 238

Resolution
unresolved
no resolver link, observed 2026-07-14T15:39:59.878169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:39:59.878169Z digest=sha256:39b10c6910aa537d0cae47955c75d0efc67d3009bd6d936c6b6e09df1d52bca2