Pith. sign in

Paper Citation Record · LEDGER

ViNT: A Foundation Model for Visual Navigation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2306.14846.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.14846 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:04:10.941000Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:08.552638Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b36a4887-8aca-4a60-b288-7bcdea25da7f · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators ViNT: A Foundation Model for Visual Navigation

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:15:18.607033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:f9d90cd575a34f79bf246f3fde3f99b419691612a752658ccb65a2735ebcc4c0

Observation 5d2783ab-23b6-40cb-9504-886ea939271e · inbound

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset cites this paper.

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset ViNT: A Foundation Model for Visual Navigation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:18.813272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:51:18.508352Z digest=sha256:6a489eac507dded7f557d85c3df60e589efe69dc18302b7d799613c83887ce96

Observation 39e852db-8e93-4270-83bb-5c330482b405 · inbound

Octo: An Open-Source Generalist Robot Policy cites this paper.

Octo: An Open-Source Generalist Robot Policy ViNT: A Foundation Model for Visual Navigation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:26:15.361975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:26:15.163358Z digest=sha256:dd98a633e467acf2c2f124d001c9714d876dde5670797f21033618c9a4fb6bf5

Observation e8386bbb-770c-46cf-8bb5-82a311ab90b0 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model ViNT: A Foundation Model for Visual Navigation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:46:36.428955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:b25e3b7e3ef0ef13bae7e41434882ffafcf634cc9f39643a0378c27b16bf0281

Observation 66941f01-0ffd-4c08-a963-5caaf9c79c81 · inbound

Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight cites this paper.

Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight ViNT: A Foundation Model for Visual Navigation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:55:25.130017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T04:53:33.503100Z digest=sha256:673c28ab7120b9799d54422407dbefef95a062cf0200fbf1aa44a48f4e60cc1a

Observation f825b08a-225c-4e22-a183-1b4b4ed3de42 · inbound

Predictive Red Teaming: Breaking Policies Without Breaking Robots cites this paper.

Predictive Red Teaming: Breaking Policies Without Breaking Robots ViNT: A Foundation Model for Visual Navigation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:10.941000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:10.941000Z digest=sha256:771951233ffac515004145305ca3e50bf63d15a39256ad1fa1ed0873cea610ad

Observation e8dedf86-1f6a-465f-977c-2584ea469877 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization ViNT: A Foundation Model for Visual Navigation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.947021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:b6a1bdbece7a1d1af4fe5fe46c34dc9a16bf5b8ef36a65fd9fe823b62cf26d60

Observation 45463947-70e1-4228-b576-c606f23ff2f8 · inbound

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments cites this paper.

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments ViNT: A Foundation Model for Visual Navigation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:04.971498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:04.971498Z digest=sha256:9fdfc64f52a87ae1e1950acc48988b39263a6b68131e59195650fa52ac1c8111

Observation a3b9629b-edf7-4985-9bd4-aee7a3b1c63a · inbound

ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments cites this paper.

ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments ViNT: A Foundation Model for Visual Navigation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:39.314742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:39.314742Z digest=sha256:2bb29a393607d94f3f3db4edc4a32e38a7a2dac1a790148132108c4d77395185

Observation d53be860-d51d-44f6-871f-14939fd49404 · inbound

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface cites this paper.

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface ViNT: A Foundation Model for Visual Navigation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:12.361685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:12.361685Z digest=sha256:0aae5f48ad1ae46f61796f3cca351f1b30c7e5c9e6d683f228298bbc2aed1ab9

Observation 3f906f59-76aa-43b6-89d6-5a93af40085e · inbound

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference cites this paper.

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference ViNT: A Foundation Model for Visual Navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:25.667794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:25.667794Z digest=sha256:324eadb44387d7e4299e7f811bfbe6338e370d4ff5990251f5ace4408c1f0035

Observation b5c086d3-636b-4df9-96a8-81861f380df6 · inbound

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding cites this paper.

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding ViNT: A Foundation Model for Visual Navigation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:16.219197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:16.219197Z digest=sha256:7b9767104afad3c2064da89a07e01872b1a2a50d4d0b8a3c594615e5db35455f

Observation 86a39f3a-db25-45be-881c-7c321934740d · inbound

PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models cites this paper.

PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models ViNT: A Foundation Model for Visual Navigation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T14:59:03.607088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:59:03.607088Z digest=sha256:9998f8ed20ef2d6c420981977e991cc71a62ddac42e45ccb1828d14c44b9acd4

Observation 0abed561-5e51-4d04-a671-9199c4446750 · inbound

From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning cites this paper.

From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning ViNT: A Foundation Model for Visual Navigation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:28.510190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:11:28.510190Z digest=sha256:c464ecded76957a65da970ac572236115d7bc8bc5b16b7ead4b3fd0311c71929

Observation c4176d68-2e40-4bfe-a753-303db470432b · inbound

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation cites this paper.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation ViNT: A Foundation Model for Visual Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.942248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.942248Z digest=sha256:5e1bb47b3a2da9e5865829982203ad0d0a376be3cd61820c5f4e799939dc7808

Observation 97f304a3-34f0-48be-b5d6-64447c219eee · inbound

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models cites this paper.

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models ViNT: A Foundation Model for Visual Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:27.032727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:27.032727Z digest=sha256:5c6f7c5c0590ed96a9d31e97de3ce963757b8cca1438d925f27e012e4517ea44

Observation 296f0711-4772-4115-b7eb-e037b94a6f41 · inbound

DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features cites this paper.

DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features ViNT: A Foundation Model for Visual Navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:21:05.054748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:21:05.054748Z digest=sha256:f9c4d154a6f23f49d211adf0a69d56924975a21576034ab9bde2232dff4b6959

Observation dec2dce5-537b-40d7-a202-44debc262802 · inbound

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning cites this paper.

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:54.645609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:37:54.645609Z digest=sha256:7cdcfb37a864d05f669a255e60a99d877ff4b0b24b6537d703cb4b6c57116dc9

Observation 47d1f422-b9d8-42dd-a79c-d98119105243 · inbound

MATT-Diff: Multimodal Active Target Tracking by Diffusion Policy cites this paper.

MATT-Diff: Multimodal Active Target Tracking by Diffusion Policy ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:42:07.427047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T21:41:44.772556Z digest=sha256:ee74c8e6d7d585351afdcad89ecf8728e011d1704e4912d9c2eaffa7f1049d02

Observation c8d781ae-5875-47a1-9c8f-d70f6e7285ea · inbound

Splatblox: Traversability-Aware Gaussian Splatting for Outdoor Robot Navigation cites this paper.

Splatblox: Traversability-Aware Gaussian Splatting for Outdoor Robot Navigation ViNT: A Foundation Model for Visual Navigation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.042373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T06:03:45.608023Z digest=sha256:1938cad6fe8bf30777a0883cd179410ff6f869912fd3849ec5f85f90e09c1b7f

Observation 90e2115f-e72e-4b12-a315-776551a94066 · inbound

CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents cites this paper.

CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:24:36.804719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:24:36.804719Z digest=sha256:ec3832bd4e6dd3d6672ff881e4ccc946c5a522f10eb6a8e4b43e686ec7022f45

Observation 02abab78-1ada-49a4-9146-9cf470127c0d · inbound

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning cites this paper.

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning ViNT: A Foundation Model for Visual Navigation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:58:55.122084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:55:11.396778Z digest=sha256:ff4d0e6fe6d7e1f70e4c46742f397207f92b64eaac2dcf72efdbdaa4c98f06b2

Observation 6f3df3c4-ee9c-4659-895a-90074b0e39e1 · inbound

Learning to Localize Reference Trajectories in Image-Space for Visual Navigation cites this paper.

Learning to Localize Reference Trajectories in Image-Space for Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:56:35.817252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:56:35.817252Z digest=sha256:30735f3eff8d6b2967c59e82446a9fc522711770cd927ebbdfb2e7314320f3f4

Observation 4e8c844d-bbf4-41cf-ad0b-672e941a9e7c · inbound

Approximate Imitation Learning for Event-based Quadrotor Flight in Cluttered Environments cites this paper.

Approximate Imitation Learning for Event-based Quadrotor Flight in Cluttered Environments ViNT: A Foundation Model for Visual Navigation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-15T13:10:42.531155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:10:42.531155Z digest=sha256:4fd5f6d48409bfe97359f49e7190c5273e9cef02a1a63aaa35505b4d421c294e

Observation b8a1c452-39d6-4731-8b11-86af9a830586 · inbound

RAE-NWM: Navigation World Model in Dense Visual Representation Space cites this paper.

RAE-NWM: Navigation World Model in Dense Visual Representation Space ViNT: A Foundation Model for Visual Navigation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-15T12:07:05.640150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:07:05.640150Z digest=sha256:c70ee75541aca1eab3841492ca6b34340e54bdd13cea6f8eaade61d3295ca328

Observation 8d205b90-6e6b-4ee9-aa56-d8a6bce10e5d · inbound

STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation cites this paper.

STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation ViNT: A Foundation Model for Visual Navigation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:13:13.493766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:11:36.717810Z digest=sha256:d98ce85f9648e884b96c938435b0837d3b5b07598e142ecc67b98224c8cb2120

Observation 89b88992-5782-4582-8245-f78c936a7f4f · inbound

Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation cites this paper.

Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:48.611135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:25:36.704959Z digest=sha256:3486f6a237f907e0a4b7124c41b2f294a70d8f7430391f5fc9e45254bb8cb5c5

Observation 9cce0e6f-dd7e-4808-a0ca-ede5a09f008e · inbound

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace cites this paper.

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace ViNT: A Foundation Model for Visual Navigation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:57.050349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:09:39.510770Z digest=sha256:c942cd273304fb4a8d0ca68dd52cc9d5f0b5a97434b553899dc447a2b56848fe

Observation 061c57a9-588c-454f-a3ff-076cadee399a · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models ViNT: A Foundation Model for Visual Navigation

Reference 147

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.201103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:7482a0275a98084469ab792e52137df7d7cd3a443dc0ca69840eca34c344255a

Observation 6ed31ce3-54cc-4393-9118-efadcf7054ab · inbound

NavOL: Navigation Policy with Online Imitation Learning cites this paper.

NavOL: Navigation Policy with Online Imitation Learning ViNT: A Foundation Model for Visual Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.907111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:51:53.848800Z digest=sha256:e43ddb920bbf66d50c3d1f234b95da321056580363766c94201dce1fa00bfb86

Observation b7040a10-19db-4a64-9b42-018e365017f4 · inbound

NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation cites this paper.

NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation ViNT: A Foundation Model for Visual Navigation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.542544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:16:06.363772Z digest=sha256:83c17838cb059b038262aa567156d428d03f28dccce62f07199295ee6432b8fe

Observation f1f397ed-7472-448b-a023-6c533351503d · inbound

Improved Baselines with Representation Autoencoders cites this paper.

Improved Baselines with Representation Autoencoders ViNT: A Foundation Model for Visual Navigation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.342413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:40:14.358108Z digest=sha256:d0110b8547b06e115ea657c0c555af6a547b741e4b319ab1f1e7705abf6551b5

Observation c3d78722-6edd-4e6f-a662-480de3c89b15 · inbound

Autonomous Frontier-Based Exploration with VLM Guidance cites this paper.

Autonomous Frontier-Based Exploration with VLM Guidance ViNT: A Foundation Model for Visual Navigation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.310430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:37:35.682778Z digest=sha256:453c670a177151e61c9a37e8616feceb5f6c9b99f13f205ceef478b2f1a7cadd

Observation 605043c3-8d2a-4628-8f39-577c7ef2a6e2 · inbound

World Models as Group Actions cites this paper.

World Models as Group Actions ViNT: A Foundation Model for Visual Navigation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.360069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:17:42.720825Z digest=sha256:2a98e1f55a69190811fd40893c3c7ff7e59cbc3778a5a55f3bb28aa5ef1a7f3a

Observation 95e3068d-ae8a-4563-b635-90a93fa1ad63 · inbound

Drift-Resistant Navigation World Model with Anchored Epipolar Guidance cites this paper.

Drift-Resistant Navigation World Model with Anchored Epipolar Guidance ViNT: A Foundation Model for Visual Navigation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.998750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:07:20.197513Z digest=sha256:15b4f8d6fc5ea14277f388f111c2afa60d4927c1642e6362e382555345be9987

Observation 95c14893-4724-4966-a61c-bbadeaceccd5 · inbound

Sentinel: Embodied Cooperative Spatial Reasoning and Planning cites this paper.

Sentinel: Embodied Cooperative Spatial Reasoning and Planning ViNT: A Foundation Model for Visual Navigation

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:04:01.451884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:58:41.050531Z digest=sha256:92f128897350e8a28616cc043e5585b7e0dc12ce475100d0da94804fc3f0b4f8

Observation 77733835-f772-40fe-b4c0-247fc1edec81 · inbound

Look Further: Socially-Compliant Navigation System in Residential Buildings cites this paper.

Look Further: Socially-Compliant Navigation System in Residential Buildings ViNT: A Foundation Model for Visual Navigation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.865324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:01:18.162750Z digest=sha256:95b2481e1451c13a81c785b76d711ccd76dbd320f2277e82d59480ae79a5c318

Observation 657cc18f-a988-48ed-b808-8972ceb6515b · inbound

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation cites this paper.

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation ViNT: A Foundation Model for Visual Navigation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:53:23.879641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T11:48:14.888295Z digest=sha256:60752a0c9d707211762d4ac3f0f016b621a0d1a7d2a1774f3b22b14e21c2ff14

Observation b9f10a16-8489-4590-a084-9ff074af788f · inbound

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation cites this paper.

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation ViNT: A Foundation Model for Visual Navigation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.766063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:39:13.803571Z digest=sha256:674b943d54651dd7277ef18601cef143b5ea22e9250bd723fd053f8e2da81738

Observation 99530b7f-5da1-4f2c-9db9-d9f6f8902c60 · inbound

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies cites this paper.

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies ViNT: A Foundation Model for Visual Navigation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.602340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:04:32.423616Z digest=sha256:5d7ef501a2a7dd808dce4061b0c8531574a80e09b007c84e4a9446a211bad12f

Observation 0a9ddb97-e826-4dbd-a98c-ce4a8a45a628 · inbound

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models cites this paper.

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models ViNT: A Foundation Model for Visual Navigation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.691295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T12:52:38.764427Z digest=sha256:1651b2425a1e50fc2bb65ab6869a8ce11260b956b0027f5fdfc8147af85a3c8b

Observation a922ce2f-ecd0-4232-8a05-eb0f98af1726 · inbound

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation cites this paper.

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation ViNT: A Foundation Model for Visual Navigation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.666292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:29:43.030058Z digest=sha256:88c683b39b63b3fb17308bee7e2db744a1bd6b6cb4be40cb816b4d58ebafcb97

Observation 661ef9d2-4339-4654-9882-fbc7c1398a43 · inbound

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation cites this paper.

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.030334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:30:17.719105Z digest=sha256:9d6774385a3d69c405ae42f1ad65a697ffe099af31afd1bfb9137d320e36b385

Observation ec350992-313e-4f4f-aab2-a37f10debbd7 · inbound

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation cites this paper.

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:33.966990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:16:52.173943Z digest=sha256:08cafb4f0729b1dd669654c533ef5d7cb62933dbbdbdb5f5720d2cf563c7bfec

Observation afbe0154-3e88-4f32-8e4c-187e97abf4f6 · inbound

NavWM: A Unified Navigation World Model for Foresight-Driven Planning cites this paper.

NavWM: A Unified Navigation World Model for Foresight-Driven Planning ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:56.754455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:42:30.237545Z digest=sha256:77f7d9f97c38b62722ab83e2312fea633fb2b27b39696f3a66615d3b62417793

Observation 0678cc9e-6dc4-469d-9dda-6fb8346f5348 · inbound

Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations cites this paper.

Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations ViNT: A Foundation Model for Visual Navigation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:08.555240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T19:06:41.116675Z digest=sha256:ce65c5b8caa0c12a440bfcc912289877814b078f55df68f3ce7fc60712bd1e59

Observation 334c8c08-cd41-4ab2-8cef-b583b016dec2 · inbound

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters cites this paper.

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T07:33:00.386358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:33:00.386358Z digest=sha256:4f3ec9023497ee7ccbe53e0e7deb6f5f3703ea60adaf534343e8451f92eadd52

Observation 6db71e94-3207-4520-889b-7cf54c88d980 · inbound

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters cites this paper.

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:14.936674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:14.936674Z digest=sha256:254c29dd50992e13126414a2c033b2cec6d27c4b8f51c8efc6f55f3760aa071e

Observation 5a79a019-96fe-4527-a696-b9995511514e · inbound

SeeSE3: Emergence of 3D Space in Vision Features cites this paper.

SeeSE3: Emergence of 3D Space in Vision Features ViNT: A Foundation Model for Visual Navigation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T02:49:56.083126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:49:56.083126Z digest=sha256:39d32832f124a1ea81ce22e941ee4d74611578ee800fe646de11468b6ea39f7f

Observation 7b5c56b6-b556-4e0b-841c-3128b91a13ed · inbound

G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation cites this paper.

G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation ViNT: A Foundation Model for Visual Navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T19:29:21.514289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:29:21.514289Z digest=sha256:ba9e2202df77b10785090484c4c61b987bf19ed5b0a755b4397f624f7f08731c

Observation a655afa5-7faf-4430-99c7-a88d41869dba · inbound

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset cites this paper.

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset ViNT: A Foundation Model for Visual Navigation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:39.811308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:39.811308Z digest=sha256:d70c387d864c567400c26eef6c69b31706c70c8ebb9e5a181ae1603324b6c763

Observation 8bf941ef-1f3b-47b7-a8bc-f503debb84b3 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills ViNT: A Foundation Model for Visual Navigation

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.206962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.206962Z digest=sha256:f49313d24abdf3d23a3fc9ef3e9b44ce5f9ba201febc16667dc7df719ea7470c

Observation d4e6f2ed-93b4-412d-815c-fe561c056ec6 · inbound

UniNav: A Unified World-Action Diffusion Model for Visual Navigation cites this paper.

UniNav: A Unified World-Action Diffusion Model for Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:35.867887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:45:35.867887Z digest=sha256:b398d5e0792125152901877fd6c42d60a8bcdecd1c8e86d66dda5f49b638b641