Pith. sign in

Paper Citation Record · LEDGER

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning

As of 6 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2606.01205.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01205 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:13:10.713616Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:42:59.130907Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact12
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a511e124-18f5-4c2d-bb81-95e4b1993152 · outbound

This paper cites Vision-language navigation: a survey and taxonomy,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Vision-language navigation: a survey and taxonomy,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:1778bfe1a7444416d9babb7510782af88d37bd7e7644adcfea8bebdd32f3951c

Observation 12f331b0-714a-4dcd-bab0-fda0463b39fa · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.903772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:998cabc62724e65271eb05557cef851788ebb9cac1c8c3860617eb9787943446

Observation 077a25af-c10c-4a41-a26b-932756bfa018 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning OpenVLA: An Open-Source Vision-Language-Action Model

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.932419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:7f14563d1e11756d2c5726cd122ea10dc8c37c48cc65c3920c64d6ae4c915c45

Observation eb772f8d-d6fc-4a93-80ac-47f6cf900cae · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Wan: Open and Advanced Large-Scale Video Generative Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.927426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:c850c705f993d47d643aab0f5f9fc0f8eefe5ad7a5d570fe9b47ce978bdb2b57

Observation 1419e32d-799c-4364-80ad-af4bb8e4f0d4 · outbound

This paper cites Airscape: An aerial generative world model with motion controllability,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Airscape: An aerial generative world model with motion controllability,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:5782c366dacc8916179877b66f4387d140f97f35f5c60ebe24033197d752a2b7

Observation ee8f230c-3bc8-4717-886e-bc65c53ae5c0 · outbound

This paper cites Uav-flow colosseo: A real-world benchmark for flying-on- a-word uav imitation learning,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Uav-flow colosseo: A real-world benchmark for flying-on- a-word uav imitation learning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:2172a3dcb40f8ae7cadd6afc040c8a5c36ab8ca6167d32c63768d3301214b54c

Observation 3c0328b4-d0d3-4b33-8ac0-2588e829c96a · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environ- ments,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environ- ments,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:dcf921a4b02e59656ed0134779e3ddcd83c6b196f72efa7d3ddd21a6719c7120

Observation 8adc6f3c-911b-472b-bf66-a2555dbc182f · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Beyond the nav-graph: Vision-and-language navigation in continuous environments,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:51046c28cfc54a8ab36d25bd179e5220c634d93626daea070c6d819c27f64ba3

Observation 76fe94af-1ed2-4b42-9597-4b7f01a80223 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.929924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:ba748fbb90c74383ca8821bbfd5008741c8b7423d3edca0cdb606e4fab93d7c4

Observation 5a63de00-3d62-44da-a519-25ea82d086cc · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Rt-1: Robotics transformer for real-world control at scale,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:e98d6f68c8f2971ff06a934aee4725641a14174dbaa47ff43cedf3c4f7c9fbdc

Observation f8957b9c-2727-4a97-afc6-f8da44ab2be6 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:631edea5abf3a105342ebb67a91d28c5e96739fa72a63eb76fa1efb799567669

Observation 40fcf209-6c6e-4e10-bddd-daa7c2b88177 · outbound

This paper cites Vision-language- action models: Concepts, progress, applications and chal- lenges.arXiv preprint arXiv:2505.04769.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Vision-language- action models: Concepts, progress, applications and chal- lenges.arXiv preprint arXiv:2505.04769

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.909681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:d5cc85af77e7b4f252336433e15a718d2283d8e05e8a1982ebf6c0db6ae2e348

Observation 5d13ee26-ab66-439c-9985-b2505a19cdf1 · outbound

This paper cites OpenFly: A comprehensive platform for aerial vision-language navigation.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning OpenFly: A comprehensive platform for aerial vision-language navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.906709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:b0cf3bb484b16efcd3fcf5f0402ec670a3c47687b327d5e1beb1dabbcae3aa97

Observation 09e6d90a-6011-4a9e-a0c1-7b764b8f5ec5 · outbound

This paper cites Mastering diverse control tasks through world models,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Mastering diverse control tasks through world models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:017fe4253cf30d5ca1cc9d3bcb8e61f8efb7897d538b97f179f98c1756dd3f6e

Observation 73838e77-2f03-46ef-a675-af443e7171d3 · outbound

This paper cites World Action Models: The Next Frontier in Embodied AI.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning World Action Models: The Next Frontier in Embodied AI

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.900379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:003f5aef47e66c68f170761aaf220fff236520e1ae7ad8e4bbcdc61eefbf98d3

Observation 9ce8d1c6-0ba0-4301-a052-d5281ee024cc · outbound

This paper cites World Action Models are Zero-shot Policies.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning World Action Models are Zero-shot Policies

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.924351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:ecedc6de3af40d2aecdc02c05aee3b3d60bad5d3732d59cf042b56869b2acf04

Observation 98cc0af8-91e8-48b2-a608-94dc48f01a5e · outbound

This paper cites Navigation world models,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Navigation world models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:b95c5066cb71edc5fda3c3adc8e1bab7e03277e047e303f9f8344f87e65c8901

Observation bdd9825d-d69b-46b2-b33d-3e207ee110fa · outbound

This paper cites Mowm: Mixture-of-world-models for embodied planning via latent-to-pixel feature modulation.arXiv preprint arXiv:2509.21797.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Mowm: Mixture-of-world-models for embodied planning via latent-to-pixel feature modulation.arXiv preprint arXiv:2509.21797

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.912624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:77229ca1222fd1c27723db268e8cc5d76639f27222d15f514cfffb366f4d93cb

Observation 6c0e09e0-e1ed-4289-ae1a-639aa5503cb9 · outbound

This paper cites Navdreamer: Video models as zero-shot 3d navigators.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Navdreamer: Video models as zero-shot 3d navigators

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.915770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:d980f7b2148775e1e8b248157470cb9afc69123ddb4f4ec13ae89fbe35d093b1

Observation 39397bc1-c891-43ad-b7d8-2b4e06011b49 · outbound

This paper cites Robust real-time uav replanning using guided gradient-based optimization and topological paths,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Robust real-time uav replanning using guided gradient-based optimization and topological paths,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:1449e52d36a073940f1b4b21fab6a44682d9c147fbcc97fcf900fd8b2ed4ace4

Observation c17b6962-04dd-4eb0-aafb-e8fd3068be14 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.918474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:47f100629457af631912237af926fe523beca7032e48448bbba324ba883503f1

Observation 36ae7bfd-361c-4b23-9888-3cd66a20d07a · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer,.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning Cogvideox: Text-to-video diffusion models with an expert transformer,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T17:13:10.713616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:84523ed385ed34761574d390020cca9e842d5c194b49b0d11c016a0aded6224c

Observation 3687c8b7-7fb8-4512-b102-88a8532d02cb · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning HunyuanVideo 1.5 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.921309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:cc6438aa9761a4949cdbc4592ad1a275517f793b4c06eb461327fbb403c6ebbd

Pith citing papers

Observation 76bfd3cf-f3cc-47e7-9305-ad386db48e4e · inbound

SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents cites this paper.

SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-04T08:54:16.286985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-04T08:50:11.584806Z digest=sha256:5468c8f33e2f304088b4947f19778df92f451c9f2637a19973fa3a7148de44af

Observation 761234f3-623d-4c88-9c48-ae353740f63d · inbound

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation cites this paper.

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:59.130907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:42:59.130907Z digest=sha256:5d9440a5d82ea247d6fd8e66abbd4fe8555bf65eac5722a8237fa74445786329