Pith. sign in

Paper Citation Record · LEDGER

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model

As of 20 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2607.17806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17806 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:36:48.962835Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5af842a2-c52c-4452-829e-26a8932de0e3 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.301451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.887459Z digest=sha256:b8f3ad90529d2a9869afc6e5dc96bdb6d96bb16a9cb3eb28a7cb82966c0b6e39

Observation 78f01b40-0cbe-4d14-96cb-48988ad53d10 · outbound

This paper cites The revolution of multimodal large language models: A survey.arXiv preprint arXiv:2402.12451, 2024.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model The revolution of multimodal large language models: A survey.arXiv preprint arXiv:2402.12451, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:48.893401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:48.893401Z digest=sha256:e226da7ebf10e570c5879745b2527483a5078b15e8d1e897ccf6e1d5f368dd33

Observation 2eeb13b3-751e-4ea1-8be8-78a2756f696c · outbound

This paper cites Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:48.898310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:48.898310Z digest=sha256:b93f642f4985559886ce2a1a7ae741adf35de8bcd96b352b237ba8d3677180cd

Observation 35bea90c-f91c-4bb2-9675-3108bb84a2b6 · outbound

This paper cites History aware multimodal transformer for vision-and-language navi- gation.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model History aware multimodal transformer for vision-and-language navi- gation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.285803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.904078Z digest=sha256:7fe534a8a76eff9e571e7f181267663db6bd5cd0fecb182769a27991c8fdda49

Observation c3454965-681f-42d0-8b6e-55d18d7856e6 · outbound

This paper cites Airbert: In-domain pretraining for vision-and-language navigation.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Airbert: In-domain pretraining for vision-and-language navigation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.269869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.908715Z digest=sha256:509d9b481697d6b9dacbf28aabad5092d06f14e50df4ba284002e59c8cd47417

Observation 2f87adb4-2c89-4686-b828-c5cb2b9c2fbb · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, et al.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Hu, Yelong Shen, Phillip Wallis, et al

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.253007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.913724Z digest=sha256:132f9a21953f8489b2c17a3c4361d459d4ae91c8901d5a60176502c3fed6a804

Observation efeab6e3-47db-4772-a9ec-d298f7985cc6 · outbound

This paper cites Navillm: Towards generalizable embodied agents via navigation instruction tuning.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Navillm: Towards generalizable embodied agents via navigation instruction tuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.238026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.918693Z digest=sha256:757d1a56f71343b99f9443c6da9474288cc619eac01be1649cdc87296e0d617a

Observation a67ece24-f377-4077-8f97-0fedfaa7524b · outbound

This paper cites Beyond the nav- graph: Vision-and-language navigation in continuous environments.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Beyond the nav- graph: Vision-and-language navigation in continuous environments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.222335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.924582Z digest=sha256:e5d569130b7f806d8f7a4ae2a186e961f8419686039129fb4b76ad1ac52fdfbb

Observation d24ee6a2-ade3-4ede-bafa-cf7efb6893dc · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International Journal of Computer Vision, 2017.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Visual genome: Connecting language and vision using crowdsourced dense image annotations.International Journal of Computer Vision, 2017

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.204355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.929341Z digest=sha256:46191a41fcc8d7b3e36dbdfdde2bb6dae3189646980ee9a6a424564b3e61c17c

Observation d12a2831-7626-4a6d-a19b-f439e59d6470 · outbound

This paper cites Room-across- room: Multilingual vision-and-language navigation with dense spatio- temporal grounding.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Room-across- room: Multilingual vision-and-language navigation with dense spatio- temporal grounding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.186336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.934072Z digest=sha256:d174b0354993dcef941bb0f2218bff2514b64e411c23aeb641656f832e3ec416

Observation 5b832416-ed24-4647-a5cc-cbe1ccbcffc1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:48.939535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:48.939535Z digest=sha256:1a388242c0da72f2bbfdaf927845e383e5c8df69b24037c40f6e45b57468b52f

Observation bd4e02dd-e080-4d02-9ef6-f1434ed50959 · outbound

This paper cites Microsoft coco: Common objects in context.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Microsoft coco: Common objects in context

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.170507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.944118Z digest=sha256:3f6942921bec3ecfd6a099967058203482475e4f20b6c7f2d15e84e1b86dbb67

Observation 81ba54dd-e2da-41aa-a774-17665823e8f6 · outbound

This paper cites Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 2024.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.154526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.948642Z digest=sha256:2a006d9a121445ae9d360bd3b1eb0b58e6c3525a4bac1ab857070bbdcda173e7

Observation 750c8415-03ad-4bf4-bd33-82e553c560e5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 2022.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:36:49.138854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:36:48.953499Z digest=sha256:a3b73f8f511c1904c99eba7731eabee8aa98fbe133d561d2cb46247cece707e8

Observation adefadf2-3e05-4eec-a49a-2af3fb7e1af2 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:48.958063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:48.958063Z digest=sha256:8d96e6e9ea16583614eff7b621989e00eab3de6e6d1602dbba9e461433979d09

Observation 1aebbcfd-b34e-4ecb-a785-0269cc8888b5 · outbound

This paper cites NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models.

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:48.962835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:48.962835Z digest=sha256:cfa6de5cd82482469a7da76115bb18f73f1e2c2ee2db0f4b511b6c18ba6b1315

Pith citing papers

No inbound Pith citation observations are available.