Pith. sign in

Paper Citation Record · LEDGER

A Navigation Framework Utilizing Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2506.10172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10172 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:35:44.116335Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a7396c1-0d9e-40bb-8825-e7a11f03ed87 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

A Navigation Framework Utilizing Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:44.061335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:44.061335Z digest=sha256:fde2c69e795a01292a3937b99f935509c42072c4b1601da9504d4f62502774b9

Observation b0eb695b-3ec3-4d23-83c0-eb920ac9b74e · outbound

This paper cites Bevbert: Multimodal map pre-training for language-guided navigation.

A Navigation Framework Utilizing Vision-Language Models Bevbert: Multimodal map pre-training for language-guided navigation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.361249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.066117Z digest=sha256:f6e7c1ebf4c852934569002bb5ca4231776160850f67337f220135fab7b30f3a

Observation 87c89031-327b-4c37-bf53-6982c9064a48 · outbound

This paper cites Etpnav: Evolving topo- logical planning for vision-language navigation in continu- ous environments.

A Navigation Framework Utilizing Vision-Language Models Etpnav: Evolving topo- logical planning for vision-language navigation in continu- ous environments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.353581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.069510Z digest=sha256:d74183d785cf4095f329bad39689b80ad96873e51c43c1738080fc8a2fa0b9e2

Observation ae9f9c2d-a72e-4693-b883-b74f2095a0cd · outbound

This paper cites Vision-and-language navigation: In- terpreting visually-grounded navigation instructions in real environments.

A Navigation Framework Utilizing Vision-Language Models Vision-and-language navigation: In- terpreting visually-grounded navigation instructions in real environments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.345543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.072677Z digest=sha256:93a7d750fa2f899341eadcc82975e2d08e233b9ea4c698acc7b93f00d62de4be

Observation e5c17b20-260f-496d-8101-8280107f0c6c · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

A Navigation Framework Utilizing Vision-Language Models Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:44.076043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:44.076043Z digest=sha256:9faabdac52eb66391764f5b7465e12d23f83acc6828dc76e4e87da7b00c376bd

Observation 936c4e2c-9221-4ed4-81a8-5ccd91f526a8 · outbound

This paper cites Vlmaps: Extracting world knowledge from pre-trained vision-language models for zero-shot visual navigation.

A Navigation Framework Utilizing Vision-Language Models Vlmaps: Extracting world knowledge from pre-trained vision-language models for zero-shot visual navigation

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:35:44.250933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.079261Z digest=sha256:00cec57a2d1afc2b6fb888d78377c61dd89d94f2a569ce9a35aff7efe4e35c80

Observation da448fd3-ff8e-4d80-96dd-13d502b1cc5f · outbound

This paper cites Etp- nav: Efficient trajectory planning for instruction-following embodied navigation.

A Navigation Framework Utilizing Vision-Language Models Etp- nav: Efficient trajectory planning for instruction-following embodied navigation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.336757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.082560Z digest=sha256:b78e79929ff19e3fb971971ac86c24e736c9ee6b467be9e78a4995ef23a515ae

Observation 23ce9c76-8dfc-45bb-b596-8432493960d2 · outbound

This paper cites Beyond the nav-graph: Vision and language navigation in continuous environments.

A Navigation Framework Utilizing Vision-Language Models Beyond the nav-graph: Vision and language navigation in continuous environments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.328776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.085721Z digest=sha256:73bd4745b8833ab2dede2d018841dd5a49eded7ef4d572b18f8d607d0c49ff2d

Observation a482e6f2-a4d3-46e5-832c-7021e9a29022 · outbound

This paper cites Mem2ego: Empowering vision-language models with global-to-ego memory for long-horizon embodied naviga- tion.

A Navigation Framework Utilizing Vision-Language Models Mem2ego: Empowering vision-language models with global-to-ego memory for long-horizon embodied naviga- tion

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.320319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.088445Z digest=sha256:e815119a0342a14f2c8654a78809b3bfbcf732ce20a1c99c698229bbeb6fb48f

Observation fc4e8794-8d35-4cce-adb8-cc6c05251b86 · outbound

This paper cites Habitat 3.0: A co-habitat for humans, avatars and robots, 2023.

A Navigation Framework Utilizing Vision-Language Models Habitat 3.0: A co-habitat for humans, avatars and robots, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.310581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.091762Z digest=sha256:b82d8d74c12a2d017619b363e1ef6ae8bda808321f3c459a136fa7442482ea91

Observation 11cb8350-025a-4fb3-ba18-f8987d9972b9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

A Navigation Framework Utilizing Vision-Language Models Learning transferable visual models from natural language supervision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.301878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.094571Z digest=sha256:630420c075ff16ba85ea874cb266b3348ed2fd931b91b0d63eab4ef4b588d866

Observation 704ac10f-7ae7-493f-99d2-a5007bf1a98a · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

A Navigation Framework Utilizing Vision-Language Models Habitat 2.0: Training home assistants to rearrange their habitat

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:44.097436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:44.097436Z digest=sha256:54296c6ba65776901ae54d22690404fc9a5a3c96659fa7cc163e2a3dbf76584c

Observation fdd1c6ff-27bd-4226-895f-0d7496838633 · outbound

This paper cites Qwen2.5-vl, 2025.

A Navigation Framework Utilizing Vision-Language Models Qwen2.5-vl, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.286952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.100566Z digest=sha256:ad5fae167921c323df036d962faa92c52cda35a7e3286b49411894bcc7f2c39a

Observation ab0e8cad-3a21-403f-be9e-fc073ee3e897 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Navigation Framework Utilizing Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:44.103599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:44.103599Z digest=sha256:ff337a7fb8c6e78ce66508f802f7d60a07e64e3379bd108d64f21ad592c56720

Observation c9d96a78-ad85-4a4c-ad37-2c7d4daa377d · outbound

This paper cites FractalAD: A simple industrial anomaly detection method using fractal anomaly generation and backbone knowledge distillation.

A Navigation Framework Utilizing Vision-Language Models FractalAD: A simple industrial anomaly detection method using fractal anomaly generation and backbone knowledge distillation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:35:44.175007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.106715Z digest=sha256:4485e2f9d7030ba82e944bf9fd556891af6119fdbabf88ada09b1ea07571c421

Observation 2efebbbf-7dd9-4d39-94e4-4d1d694df264 · outbound

This paper cites Mapnav: A novel memory represen- tation via annotated semantic maps for vlm-based vision- and-language navigation.

A Navigation Framework Utilizing Vision-Language Models Mapnav: A novel memory represen- tation via annotated semantic maps for vlm-based vision- and-language navigation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:35:44.278301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.109946Z digest=sha256:d1c583b33f1503e9b90b8d8b4edd39e3b7848091f147d754bc1da1c3c5907449

Observation 152ad8c4-6154-422c-8ff9-2bacbb4c0391 · outbound

This paper cites Certified Deductive Reasoning with Language Models.

A Navigation Framework Utilizing Vision-Language Models Certified Deductive Reasoning with Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:35:44.163249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.112895Z digest=sha256:8ffc3dae299cdbca33a5291c3b290f5ed6c1de92d43b6a19bbce012c42f1876a

Observation 19bb3a7b-f60b-43ea-9941-1923827e6e34 · outbound

This paper cites MPE4G: Multimodal Pretrained Encoder for Co-Speech Gesture Generation.

A Navigation Framework Utilizing Vision-Language Models MPE4G: Multimodal Pretrained Encoder for Co-Speech Gesture Generation

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:35:44.150653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T04:35:44.116335Z digest=sha256:57a6dd16d5de1e1c8ed9b34c92be73601efc207abd9223371cecc504e99b8698

Pith citing papers

No inbound Pith citation observations are available.