Pith. sign in

Paper Citation Record · LEDGER

CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2211.16649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.16649 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:00.698604Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:09:59.557766Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 02e1bf55-4766-44e9-b8df-eae642b669fc · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:25:59.157170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:146deb8772659e54f6f82d5a30e0c366d608da000137a8f49b4c8094d1894f12

Observation 28af2476-28df-4681-af85-cadd6162b821 · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:51:36.288054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:4adde6db5503d680ac512dd11a377327db2f344843724c45d6d5fa82dd6a4ad6

Observation 3c22e683-59c9-4af2-9a43-b6de6d942bcc · inbound

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? cites this paper.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.698604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.698604Z digest=sha256:46e482b2255848db8934d9ab665b6772194969501ba6eb68295916982cac20bc

Observation d9b511d7-f51c-44ef-bfff-bcde7103d964 · inbound

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference cites this paper.

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:23.231918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:23.231918Z digest=sha256:8019e33c97057d27825b3f21e77c43da57dbe7136eee43ac03f25bc31b33f0e3

Observation 2647ea3d-4f34-44d7-9f22-98653abaa3d0 · inbound

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes cites this paper.

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:43:16.523905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T18:42:01.266249Z digest=sha256:3b6159aad6ca8bb03a0da122708ffbfe9232b9f0846f06dda436e0d884851ac2

Observation ae29507b-a075-4f3f-818d-cc19d812ecfb · inbound

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace cites this paper.

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:57.266116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:09:39.510770Z digest=sha256:93e4f48643c5538c7a0a23f3d52cea6687f04d1b1d9cd6de38884fbcbe4f9934

Observation 7ae10c10-5436-48a9-b4e9-38ce84c6f502 · inbound

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation cites this paper.

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:17:37.783102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:13:12.029402Z digest=sha256:ee3a6171d7db351ce897441a9a71c18c3d5416ef6af4e2e51e97e47f99ae7e21

Observation 4f670a4a-86c0-4cd4-9a19-9c12edc3ff19 · inbound

NORM-Nav: Zero-Shot Mobile Robot Navigation with Natural Language Behavioral Constraints cites this paper.

NORM-Nav: Zero-Shot Mobile Robot Navigation with Natural Language Behavioral Constraints CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:12:44.705181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T20:11:52.658283Z digest=sha256:9a440604b73c1605004ca08ebecb911dfb277d6cde1e3aedc6ae8a99403eceaa

Observation bf3fa07c-7c51-4b30-a3e2-ec3b920608f4 · inbound

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation cites this paper.

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T11:53:23.904587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T11:48:14.888295Z digest=sha256:6cc78d7d6185d947301c29883799de9b0a560c72ec02ed58db9c7c87d7f25fd8

Observation 3ba348a1-0196-4ec2-90ea-b11d30a25461 · inbound

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control cites this paper.

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.559277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T23:51:39.882029Z digest=sha256:51517b83817e471cf8561674bac02646e131ee7639b35353b857a31a1865c2c6

Observation 8d2c8d4b-e06d-4a01-93ba-54e93aa04528 · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:18.712661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:17:18.712661Z digest=sha256:580fc6db8e42a770e79571ffcad059eea7e342115a1971b003212d4362c89d91