Pith. sign in

Paper Citation Record · LEDGER

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2506.16623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16623 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:25:19.387288Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:22:17.982994Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:08:02.851997Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b997a48a-a839-4a54-8743-d6eb80b8f835 · outbound

This paper cites FrontierNet: Learning Visual Cues to Explore.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation FrontierNet: Learning Visual Cues to Explore

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.266603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.266603Z digest=sha256:e396d3afc9d8ec2fdf2f336483b83338f270b8a434dabc3a94837d5a56b328f2

Observation 531eb9b5-6092-41a4-8075-4e2bc68af1b1 · outbound

This paper cites Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.271735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.271735Z digest=sha256:1fef15ebddc9c38feda5a207ba0a5f5d83670217c7621476c25a019e63167842

Observation b68da96a-0c70-436d-9504-034003fcfc35 · outbound

This paper cites Object goal navigation using data regularized q-learning,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Object goal navigation using data regularized q-learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.766101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.275730Z digest=sha256:2640a261014749ab10289bc0be4539ee5bd8f62f5fe7a77b2c13ce6be01c3f4d

Observation 2cce3c0d-7e79-42ba-9868-b6ede055898e · outbound

This paper cites CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.279724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.279724Z digest=sha256:cbfdc00371a7794ebc7cddaaeca05d898240f2fe6615f4082edfa52128e59980

Observation aff8b131-a2dc-41ee-858c-17f9d6e2cebe · outbound

This paper cites A survey of object goal navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation A survey of object goal navigation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.752823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.284466Z digest=sha256:60c59d7536cf47dcf2c8eb991bf172696ec00f99e305d47aaf969f7a7ce6767b

Observation 24f71764-3c4e-4ada-8df0-5606a767304b · outbound

This paper cites Vlfm: Vision- language frontier maps for zero-shot semantic navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Vlfm: Vision- language frontier maps for zero-shot semantic navigation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.288417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.288417Z digest=sha256:328f8f0becc133828696a571d4a76aacb797b827aece78038a5472521f08fc8b

Observation 12ebd94b-b90d-4dcf-93c8-d8f966ce5373 · outbound

This paper cites GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:25:19.588541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.292487Z digest=sha256:ac90038b3541d6d767fbc0ae7bf024342b2513a0251492e9db189b89c5f9aa02

Observation 3c5fc472-ac59-416c-8ad5-f5b51baf84c6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Learning transferable visual models from natural language supervision,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.296239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.296239Z digest=sha256:0ce7e8f6e957ffe205814d4fe964182c10ea74d7ed9c3b491ffb617f51b773e3

Observation 325808cb-2fd8-4715-b849-d046bf7a62f0 · outbound

This paper cites CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.305443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.305443Z digest=sha256:533a97c570e4c857551ad6de730a7e86fb12b519f40136332bf2dbbeae58b59c

Observation aba48fbe-5c5c-4b78-9bc6-c6393a52052a · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embeddings,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.722520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.310082Z digest=sha256:0c2ccf436500a0fb10848fc8eaf7736e76900a201e33044ab091b56cdc21f0d6

Observation 882e3354-cea1-4c75-8c08-2a6087dddbbf · outbound

This paper cites Habitat challenge 2023,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Habitat challenge 2023,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.318171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.318171Z digest=sha256:f32f346b6bc8bba2c8e4289520621adda689c3f4787b6908e9edcc844f379428

Observation a09bda92-332a-4b0e-854c-95937fe9c14f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Improved Baselines with Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.321830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.321830Z digest=sha256:460cdf9563f0d1a0b349988f8bcbf1d22f411df9248859c7288ee4f7075a0713

Observation 81249f90-cce7-4ccc-bd88-7f997ef1ea67 · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Pali: A jointly-scaled multilingual language-image model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.700734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.326579Z digest=sha256:a0d0fb66748bac4deb5a1cc6f656af2d3ad5c3f665a5a89b96cd4f0fdee67a5b

Observation 69889f48-d61c-42ca-b982-c7acbe558e68 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation NVILA: Efficient Frontier Visual Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.335254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.335254Z digest=sha256:dbc169157488729a679c869c643b2b71894e8a81c8efde341c1d51c10d475989

Observation 9d0824e5-be4c-48cb-88be-48e6afc8d122 · outbound

This paper cites Frontier based exploration for autonomous robot,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Frontier based exploration for autonomous robot,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.687419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.340239Z digest=sha256:47a3acd3a73495d8a6d8299947d1ec0aae451929eaefaff69c53b05396e44c1b

Observation abafff3f-b900-48d2-a66c-9ab8fd40bae1 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.331049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.331049Z digest=sha256:6f158d5e971e5d15f5e12844ce95768575d7e95bee6253bf7cde856cfb9ff3af

Observation 62aa880c-d2f6-406f-bdd8-a7174025b588 · outbound

This paper cites On Evaluation of Embodied Navigation Agents.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation On Evaluation of Embodied Navigation Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.347842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.347842Z digest=sha256:787c9efbdd01cb388478cba6548c59456ed2af0f7bcb44146ecf14e2e46da3ab

Observation bdf48b86-5c16-4717-9c72-edc2e8c453ec · outbound

This paper cites Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.660563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.352077Z digest=sha256:df327d6fda7237f7425a90de3088200f97f607d265c99819d7f1a92252585e1e

Observation 35c616ed-b541-41e4-af0b-0a3f5f25ecc6 · outbound

This paper cites Ver: Scaling on-policy rl leads to the emergence of navigation in embodied rearrangement,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Ver: Scaling on-policy rl leads to the emergence of navigation in embodied rearrangement,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:25:19.674521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:25:19.344411Z digest=sha256:755ceed90c1ae7d2b25c5a57f5b8229f49679e82ab31c43ce34947e60642ccf0

Observation e25df732-7ba8-48c2-a1b0-c2283f830bc8 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.367137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.367137Z digest=sha256:f3ae1765d6f0bf13a737afa82c3fdb8f4ecb508b53eebb5d2e402688da637019

Observation 08d1bc69-694a-4b10-b176-bd71d2ecabb6 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.372145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.372145Z digest=sha256:609c457fb354a6f119a4cf8fa6dbb9d4fbd7c45d1edb37162d00f2dcfd7af019

Observation c6359c9e-5c66-4754-814b-77a9e7833d43 · outbound

This paper cites L3mvn: Leveraging large language models for visual target navigation,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation L3mvn: Leveraging large language models for visual target navigation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.377263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.377263Z digest=sha256:52098b1854938874048b28544790d6eb25e1de6f86044347dd9f85c549fc8d12

Observation 958275c8-8627-47b4-8b59-14b3c1fc2bc2 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Microsoft COCO: Common Objects in Context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.361565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.361565Z digest=sha256:5fd2c4e8d94bc82dc92dce81bb16c0f1a3bcbb252467f8021dcd0ca9b0c8f2cb

Observation 4eb95dfd-7370-4bd8-b117-ed3fc8dacad1 · outbound

This paper cites Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.387288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.387288Z digest=sha256:57a5f127737bbda1e35f479ff8b3184c6c0fe2ceb4c6e349458022480220a3b9

Observation eefe3ada-35d9-496f-ac6d-53d309628757 · outbound

This paper cites ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.382409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.382409Z digest=sha256:affb05627d5cc45b26fdf94dcc55c6c78247d750424d8372762df365df488aab

Observation ff2727f6-0e9f-4d78-9313-78087522ba10 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation Learning Transferable Visual Models From Natural Language Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.300479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.300479Z digest=sha256:eeaa768d715c7f3a0e9dc52782857b0070a09ce8785668395310805d6e2ef641

Observation 85c8bf57-266c-404e-b5b8-4a197c513bfe · outbound

This paper cites YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.356806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.356806Z digest=sha256:4496730b3cf2e97ed46fe08d128318d19f585f67d282b5ccb3084280b8eddd6c

Observation 2607f2d5-0820-4109-aa2e-629fcd08a526 · outbound

This paper cites ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings.

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:19.313862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:19.313862Z digest=sha256:030782ccbbbbc8501d228dda57aec1357e140ca667f77f3462677173bfaee309

Pith citing papers

Observation 28ed85db-802d-41e5-8563-183a2290d9eb · inbound

HOMI: Ultra-Fast EdgeAI platform for Event Cameras cites this paper.

HOMI: Ultra-Fast EdgeAI platform for Event Cameras History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:17.982994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:22:17.982994Z digest=sha256:ba6da241d10926cf338318893203990b65a83a8932d8947c8713bd4fcd71dd83

Observation c02f03f6-015c-494e-8c71-b17d72bf7d27 · inbound

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance cites this paper.

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T19:40:13.426231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:40:13.426231Z digest=sha256:aa8de6b61eb3b457961db87f871ac3f8390fbcb581ea2e166b93a84cf84ec07a

Observation 82e45f01-d120-40fa-8c99-283edad080b2 · inbound

LIME: Learning Intent-aware Camera Motion from Egocentric Video cites this paper.

LIME: Learning Intent-aware Camera Motion from Egocentric Video History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:08:02.853535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T11:01:22.425665Z digest=sha256:d57b0e72bed3ded7e786562077046f9a79bedede396674c3b16be2b706515b2e