Pith. sign in

Paper Citation Record · LEDGER

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

As of 17 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2506.16623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16623 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:25:19.387288Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:22:17.982994Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:08:02.851997Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b997a48a-a839-4a54-8743-d6eb80b8f835 · outbound

This paper cites FrontierNet: Learning Visual Cues to Explore.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation FrontierNet: Learning Visual Cues to Explore

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.266603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.266603Z digest=sha256:d869384a127f7a021c32e7462c3a059b7084b8433161504d6bace760ff8bb301

Observation 531eb9b5-6092-41a4-8075-4e2bc68af1b1 · outbound

This paper cites Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.271735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.271735Z digest=sha256:111590e65fff2b0b3bc8b0404accc386436e398bc8a20019990b28ae8519a80a

Observation b68da96a-0c70-436d-9504-034003fcfc35 · outbound

This paper cites Object goal navigation using data regularized q-learning,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Object goal navigation using data regularized q-learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.766101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.275730Z digest=sha256:78f5008bae3fa68a26a9362dedf70823ed40a2483c1dd0c49e2a1ed73337bcde

Observation 2cce3c0d-7e79-42ba-9868-b6ede055898e · outbound

This paper cites CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.279724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.279724Z digest=sha256:21be74c2bde6a2624c1f9c9007c705f57c4a97f965ced2ba953346f1946ede89

Observation aff8b131-a2dc-41ee-858c-17f9d6e2cebe · outbound

This paper cites A survey of object goal navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation A survey of object goal navigation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.752823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.284466Z digest=sha256:5c236dc12861892fc0cdf741bfa27dfd814cc1226e0e23e9f03d9e9fb2614f34

Observation 24f71764-3c4e-4ada-8df0-5606a767304b · outbound

This paper cites Vlfm: Vision- language frontier maps for zero-shot semantic navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Vlfm: Vision- language frontier maps for zero-shot semantic navigation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.288417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.288417Z digest=sha256:cd6f99a7fa1b7f1ac92ca8917b3f5a2ac933dad05f9dada36e2dd6896be9d9ff

Observation 12ebd94b-b90d-4dcf-93c8-d8f966ce5373 · outbound

This paper cites GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:25:19.588541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.292487Z digest=sha256:78e7e9a202fc57452bba72425fc659803fd302bb353df69d52a650cde5e0a8da

Observation 3c5fc472-ac59-416c-8ad5-f5b51baf84c6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Learning transferable visual models from natural language supervision,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.296239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.296239Z digest=sha256:8f39574548415350866243273f97debf66781416e4f7086f99811f626d6961cb

Observation 325808cb-2fd8-4715-b849-d046bf7a62f0 · outbound

This paper cites CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.305443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.305443Z digest=sha256:3e9cd1e631b560cf38bafbad7fa92eb68514fc81b1909c28c5ea978228796816

Observation aba48fbe-5c5c-4b78-9bc6-c6393a52052a · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embeddings,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.722520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.310082Z digest=sha256:323dd5f6614e50933a884de5d8f56d1c981b875c6dfc21c9b3a2af606a809fc4

Observation 882e3354-cea1-4c75-8c08-2a6087dddbbf · outbound

This paper cites Habitat challenge 2023,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Habitat challenge 2023,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.318171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.318171Z digest=sha256:fb42b651b2c85d8c889c71287fdf7965441c1f000046948fd436b60b60171652

Observation a09bda92-332a-4b0e-854c-95937fe9c14f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Improved Baselines with Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.321830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.321830Z digest=sha256:ce86c9c54d79e9e814161be65929821b2b314b86bdfb83af7001f4be0295e47c

Observation 81249f90-cce7-4ccc-bd88-7f997ef1ea67 · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Pali: A jointly-scaled multilingual language-image model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.700734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.326579Z digest=sha256:bbe588d631b31d05eb89c77032e9148a847615a065e331b563b8cf03d384df36

Observation 69889f48-d61c-42ca-b982-c7acbe558e68 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation NVILA: Efficient Frontier Visual Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.335254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.335254Z digest=sha256:658c04c7569411396a801b90ab35d06482138c84b3b3edf919ea3f6c72465aec

Observation 9d0824e5-be4c-48cb-88be-48e6afc8d122 · outbound

This paper cites Frontier based exploration for autonomous robot,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Frontier based exploration for autonomous robot,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.687419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.340239Z digest=sha256:e4755cd84ad888ba64ffd4763a6aff802e2e1078041b071306d27870f3f245b8

Observation abafff3f-b900-48d2-a66c-9ab8fd40bae1 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.331049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.331049Z digest=sha256:78c9f70db62d9c6a9605dec156007ba55e4d2b5d4c13a73dd4d8cfc20404b353

Observation 62aa880c-d2f6-406f-bdd8-a7174025b588 · outbound

This paper cites On Evaluation of Embodied Navigation Agents.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation On Evaluation of Embodied Navigation Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.347842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.347842Z digest=sha256:99f2e784b9eac612bf66bf8a9cb4fab184154f250a8644b496e59c33486cc548

Observation bdf48b86-5c16-4717-9c72-edc2e8c453ec · outbound

This paper cites Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.660563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.352077Z digest=sha256:84ee7ba2cae30a36be3d0d687f6f7185611374da1fc4a45a6478fffe45b2424b

Observation 35c616ed-b541-41e4-af0b-0a3f5f25ecc6 · outbound

This paper cites Ver: Scaling on-policy rl leads to the emergence of navigation in embodied rearrangement,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Ver: Scaling on-policy rl leads to the emergence of navigation in embodied rearrangement,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.674521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:25:19.344411Z digest=sha256:176a154c02020197db27b679335535d0629ea3b2b7d86741f8b349e708c22fd5

Observation e25df732-7ba8-48c2-a1b0-c2283f830bc8 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.367137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.367137Z digest=sha256:725f5400d67abf4ef9d5eac1e006c5c473ac4acf1874a3763d23cd3818db5b2b

Observation 08d1bc69-694a-4b10-b176-bd71d2ecabb6 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.372145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.372145Z digest=sha256:553be849e59cbd07811e234f272a7906b6f98cd255f59ba14f1ad77e0f64839b

Observation c6359c9e-5c66-4754-814b-77a9e7833d43 · outbound

This paper cites L3mvn: Leveraging large language models for visual target navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation L3mvn: Leveraging large language models for visual target navigation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.377263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.377263Z digest=sha256:052c0e4c10f7649153393dce50b09674f872715aa0f9db6c29eb800faf8c987d

Observation 958275c8-8627-47b4-8b59-14b3c1fc2bc2 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Microsoft COCO: Common Objects in Context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.361565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.361565Z digest=sha256:9463140417bde739eefdb57e4c45de7a7cfb3d163aaddee376b529df459c8ced

Observation 4eb95dfd-7370-4bd8-b117-ed3fc8dacad1 · outbound

This paper cites Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.387288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.387288Z digest=sha256:f03d05f3ef557e89c0228dd5e5d23c7f6804a98dbdefe2667278f822c364fa3c

Observation eefe3ada-35d9-496f-ac6d-53d309628757 · outbound

This paper cites ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.382409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.382409Z digest=sha256:d339ae75e1a5a5a2d5991bfc8f3f96c9f70f589e929a401cda0483c1e5ea9a77

Observation ff2727f6-0e9f-4d78-9313-78087522ba10 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Learning Transferable Visual Models From Natural Language Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.300479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.300479Z digest=sha256:46928bf366972e9fb40b89928fc8b9209bd4863aae911cb6e8dce34e6cced134

Observation 85c8bf57-266c-404e-b5b8-4a197c513bfe · outbound

This paper cites YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.356806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.356806Z digest=sha256:b84bc7abdb4eb8ea9eb137e28d401a06b01351ddb6d21e23c6183fd855260118

Observation 2607f2d5-0820-4109-aa2e-629fcd08a526 · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.313862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.313862Z digest=sha256:84b7e27452572ecf36fd5bfa19e9cfae2c367daee70999d79002c7344225d024

Pith citing papers

Observation 28ed85db-802d-41e5-8563-183a2290d9eb · inbound

HOMI: Ultra-Fast EdgeAI platform for Event Cameras cites this paper.

HOMI: Ultra-Fast EdgeAI platform for Event Cameras History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:17.982994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:22:17.982994Z digest=sha256:51b76a1ed88b706a0efb80dc09db4d2f851f9885a5dd4100909edf158cac3d47

Observation c02f03f6-015c-494e-8c71-b17d72bf7d27 · inbound

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance cites this paper.

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T19:40:13.426231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:40:13.426231Z digest=sha256:ea7eb3513f666934bdfcb392f978a61d4995a18f29ab64640cce9d0f02f2a17b

Observation 82e45f01-d120-40fa-8c99-283edad080b2 · inbound

LIME: Learning Intent-aware Camera Motion from Egocentric Video cites this paper.

LIME: Learning Intent-aware Camera Motion from Egocentric Video History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:08:02.853535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-03T11:01:22.425665Z digest=sha256:8ffe52182f72edbe3379bf986090e60534bd1c162afa634eada293b443e56f31