Pith. sign in

Paper Citation Record · LEDGER

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

As of 18 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2506.16623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16623 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:25:19.387288Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:22:17.982994Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:08:02.851997Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b997a48a-a839-4a54-8743-d6eb80b8f835 · outbound

This paper cites FrontierNet: Learning Visual Cues to Explore.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation FrontierNet: Learning Visual Cues to Explore

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.266603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.266603Z digest=sha256:3d626b01bd47fa11a356db029e0033eded548afb5f4f7d3c059dee47e1d32d8b

Observation 531eb9b5-6092-41a4-8075-4e2bc68af1b1 · outbound

This paper cites Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.271735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.271735Z digest=sha256:c26bd4e1dafd0a010d66ecca9c440ff5626351137eae9f5dc9b52262571a0192

Observation b68da96a-0c70-436d-9504-034003fcfc35 · outbound

This paper cites Object goal navigation using data regularized q-learning,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Object goal navigation using data regularized q-learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.766101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.275730Z digest=sha256:f8f2970acacb93e85b1b550e98e3a56044f859ed65f83278e417115ea88331c6

Observation 2cce3c0d-7e79-42ba-9868-b6ede055898e · outbound

This paper cites CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.279724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.279724Z digest=sha256:272d6f077bd0e59de894ef02cda66883f7bf01c210f191185f2ea28622c07759

Observation aff8b131-a2dc-41ee-858c-17f9d6e2cebe · outbound

This paper cites A survey of object goal navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation A survey of object goal navigation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.752823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.284466Z digest=sha256:1e0359e0618e8522f6aac84fbf74d49b9409c3de274537664fea29a78e3b8256

Observation 24f71764-3c4e-4ada-8df0-5606a767304b · outbound

This paper cites Vlfm: Vision- language frontier maps for zero-shot semantic navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Vlfm: Vision- language frontier maps for zero-shot semantic navigation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.288417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.288417Z digest=sha256:905ded6de0f928828b60aa663bf95033def4fd88dda0a262c4d97131f0821f10

Observation 12ebd94b-b90d-4dcf-93c8-d8f966ce5373 · outbound

This paper cites GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:25:19.588541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.292487Z digest=sha256:23a7fd33e6bf9cb4bf044d95bb763ead8d15131c625eb8643b4585aa898d390b

Observation 3c5fc472-ac59-416c-8ad5-f5b51baf84c6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Learning transferable visual models from natural language supervision,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.296239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.296239Z digest=sha256:9d5b5465786cf9afd2da0cbe510804692b6720adfc920d2acb5b8a4ad0f6f37d

Observation 325808cb-2fd8-4715-b849-d046bf7a62f0 · outbound

This paper cites CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.305443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.305443Z digest=sha256:637758e3044a8757195cd91e734e9faec11d90b1092c2ad6ee41b17b1ea9d2ea

Observation aba48fbe-5c5c-4b78-9bc6-c6393a52052a · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embeddings,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.722520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.310082Z digest=sha256:a29aab75a5ea1dcc1979aae1ab8d02168008c8e5912ed95b78ce414dd044fb8f

Observation 882e3354-cea1-4c75-8c08-2a6087dddbbf · outbound

This paper cites Habitat challenge 2023,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Habitat challenge 2023,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.318171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.318171Z digest=sha256:6ab8e62ca612fef85f88ec53530722a3322284f4c3d6e2c6b18c126078910401

Observation a09bda92-332a-4b0e-854c-95937fe9c14f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Improved Baselines with Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.321830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.321830Z digest=sha256:f0410ad9e1881da799a86e075b92fecb66269ceb143f0e3610969f1f405e341d

Observation 81249f90-cce7-4ccc-bd88-7f997ef1ea67 · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Pali: A jointly-scaled multilingual language-image model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.700734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.326579Z digest=sha256:dcfb2eec42466ece4ea3c16bf9dc5421e146f9e098ca6446937750559e1dd2b0

Observation 69889f48-d61c-42ca-b982-c7acbe558e68 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation NVILA: Efficient Frontier Visual Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.335254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.335254Z digest=sha256:80de75be8048c19500c6a867f0c7b81ed728d8af8182c46c2f07f3b60f0d4057

Observation 9d0824e5-be4c-48cb-88be-48e6afc8d122 · outbound

This paper cites Frontier based exploration for autonomous robot,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Frontier based exploration for autonomous robot,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.687419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.340239Z digest=sha256:ff4ae0105dc4608e2d96897d6948a56316bd839187fa7191fce4a3669f52d501

Observation abafff3f-b900-48d2-a66c-9ab8fd40bae1 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.331049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.331049Z digest=sha256:d808ccf8adf18cb02d46bce94c5b14c34f7c6f046e20a5e4d81e3cdf5cfa8626

Observation 62aa880c-d2f6-406f-bdd8-a7174025b588 · outbound

This paper cites On Evaluation of Embodied Navigation Agents.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation On Evaluation of Embodied Navigation Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.347842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.347842Z digest=sha256:897e12c81bc74370fac3fe241924ffe99d1308b56bca1b32c716b714fada6d60

Observation bdf48b86-5c16-4717-9c72-edc2e8c453ec · outbound

This paper cites Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.660563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.352077Z digest=sha256:f6f64b73ad30b415ab0657bb37019ca39c881e893b1806a7649e894e7408761e

Observation 35c616ed-b541-41e4-af0b-0a3f5f25ecc6 · outbound

This paper cites Ver: Scaling on-policy rl leads to the emergence of navigation in embodied rearrangement,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Ver: Scaling on-policy rl leads to the emergence of navigation in embodied rearrangement,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.674521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:25:19.344411Z digest=sha256:6f2203ddcd6342aea09e7bc33c7b71a74479ca2905dd6cf1b799452f6da33db5

Observation e25df732-7ba8-48c2-a1b0-c2283f830bc8 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.367137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.367137Z digest=sha256:3853af8458d78add0f0fb0d06d4231fd502569e9e6c3fc4503881643d1fea22d

Observation 08d1bc69-694a-4b10-b176-bd71d2ecabb6 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.372145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.372145Z digest=sha256:a5f31d960b9340ced0822c8b1998c200f7dab6fb4383b29827b43244e4b07e10

Observation c6359c9e-5c66-4754-814b-77a9e7833d43 · outbound

This paper cites L3mvn: Leveraging large language models for visual target navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation L3mvn: Leveraging large language models for visual target navigation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.377263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.377263Z digest=sha256:e2bd38adfe92e84004064ab9de0676af0188187f4567f7ad7d3fe9dcbc714deb

Observation 958275c8-8627-47b4-8b59-14b3c1fc2bc2 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Microsoft COCO: Common Objects in Context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.361565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.361565Z digest=sha256:664b4ebaa4561141b64e0ac7629c0b182f7c000ac5ffb47f526847990be900f3

Observation 4eb95dfd-7370-4bd8-b117-ed3fc8dacad1 · outbound

This paper cites Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.387288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.387288Z digest=sha256:d0fc6e52cbd339a90769f473330b4cc3d0a62bc3664e44594f1ee11fed85fe9e

Observation eefe3ada-35d9-496f-ac6d-53d309628757 · outbound

This paper cites ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.382409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.382409Z digest=sha256:5a2d68848699cada27e4eab82246956f4a9d666289ba2e309a300910b2a5650e

Observation ff2727f6-0e9f-4d78-9313-78087522ba10 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Learning Transferable Visual Models From Natural Language Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.300479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.300479Z digest=sha256:5a578a2582a86c9961f4fa3c04e798983d529453b6787692121ac38fb27035da

Observation 85c8bf57-266c-404e-b5b8-4a197c513bfe · outbound

This paper cites YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.356806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.356806Z digest=sha256:a3aa957c528bc0518249a27a6c50c766da0c8cc33ce44f3e29c3ee2821df00d2

Observation 2607f2d5-0820-4109-aa2e-629fcd08a526 · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.313862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.313862Z digest=sha256:ae97a4ef692ad890cc11b447faf07379734ccab477a0eec5990f0020f22f8d66

Pith citing papers

Observation 28ed85db-802d-41e5-8563-183a2290d9eb · inbound

HOMI: Ultra-Fast EdgeAI platform for Event Cameras cites this paper.

HOMI: Ultra-Fast EdgeAI platform for Event Cameras History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:17.982994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:22:17.982994Z digest=sha256:b36369022a059b2586e2d142a66332a3cb7be9f048135e6f5a0ae3ce6cb89fed

Observation c02f03f6-015c-494e-8c71-b17d72bf7d27 · inbound

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance cites this paper.

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T19:40:13.426231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:40:13.426231Z digest=sha256:f6fd13ec29fc2d9a34043a5dbfffd8335bca4f399ee0d851cea385a97864be23

Observation 82e45f01-d120-40fa-8c99-283edad080b2 · inbound

LIME: Learning Intent-aware Camera Motion from Egocentric Video cites this paper.

LIME: Learning Intent-aware Camera Motion from Egocentric Video History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:08:02.853535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T11:01:22.425665Z digest=sha256:d1fb4ab3962edb1680cb5def2bc063bf335990e799eebbcfb1133d019a7e5606