Pith. sign in

Paper Citation Record · LEDGER

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2505.20718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20718 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:46.426406Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54c4669a-2f11-403f-baaa-d4c2a298f02a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.474536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.474536Z digest=sha256:e885da5c3df047c2b7f332167dff810e6e7d16a0c4cb6539ec0b92f8a85ecc11

Observation 6ef6bfa1-9b6e-4479-8d86-5dd33924f34b · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:43.540009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:43.540009Z digest=sha256:9484a3ea24aed78a4e7152d7cb228a000258f19cf7595784f8dd7db4e2b3da14

Observation 31cb6773-2e01-4d29-8031-b6364f4d8287 · outbound

This paper cites Tracking anything with decoupled video segmenta- tion.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking anything with decoupled video segmenta- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.969592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:43.644671Z digest=sha256:1a3703a4ec3aeda4370cefb0f1e8f0135d6560722b80639a9e2bdcec2da3b6d7

Observation 5b26f5f9-35cb-44d4-9ce4-4137ccf72add · outbound

This paper cites an unresolved cited work.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:51:50.855505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:43.761418Z digest=sha256:e76cd5a502c34f2587da572339cfaefb03c2a267249d7ec7d4cad7a1aff9989c

Observation 704b213d-417b-4f9a-bad5-8c4c6730ff8b · outbound

This paper cites Proactive multi-camera collaboration for 3d human pose estimation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Proactive multi-camera collaboration for 3d human pose estimation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.692694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:43.901893Z digest=sha256:7d455536d066edb8bcb8df50839b44d4b3272d099ccb5d1823c6579aaba328e4

Observation b18b71a2-8040-4c01-aef6-96f5211b6d6d · outbound

This paper cites Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Enhancing continuous control of mobile robots for end-to-end visual active tracking.Robotics and Autonomous Systems, 142:103799, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.605703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.059949Z digest=sha256:5215cac6748fa745f8748149ca0abdc02a19829d670bf06c209ab1ed46a7482f

Observation f6d464b0-ed64-428e-939c-1a569fac0b7d · outbound

This paper cites E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models E-vat: An asymmetric end-to-end approach to visual active exploration and tracking.IEEE Robotics and Automation Letters, 7(2):4259–4266, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.455798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.177272Z digest=sha256:eefb1f03b48827585504b36a3b4a4c63e1b37d585a14124183258f40020242be

Observation cbf17c7f-be72-4e1e-949f-5826f5d84ac8 · outbound

This paper cites D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models D-vat: End-to-end visual active tracking for micro aerial vehicles.IEEE Robotics and Automation Letters, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.324369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.299789Z digest=sha256:cd1210a9f1e29e07b5e748ef8dd8f6e801b523b4340994fd5ca4986b2deb0101

Observation 16a841a7-1f69-462c-b24b-69874cc65de2 · outbound

This paper cites Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Memory sharing for large language model based agents.Arxiv Preprint Arxiv:2404.09982, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:44.418085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:44.418085Z digest=sha256:b8789d27054759b5410c533d21bf43e49f3f413113a85ef93ab6bb077bc8e50c

Observation 93964576-1afd-41f8-b39a-fd8b7be5de22 · outbound

This paper cites An embodied generalist agent in 3d world.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models An embodied generalist agent in 3d world

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.500183Z digest=sha256:4a32757f4192014bac772ab1d01035f57613f6a326305354145bca63b99f96ae

Observation 6abaf5df-2d3c-4371-b70e-d389b36ce974 · outbound

This paper cites Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Conquering Ghosts: Relation Learning for Information Reliability Representation and End-to-End Robust Navigation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:51:46.630634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.571948Z digest=sha256:75882d52d288d20745ba0f5a7ec174b63a81b134cf01a5302d1dbfa4f5bdaef8

Observation 55664fc4-dc56-47dd-83c1-d2bd62cba002 · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models OpenVLA: An open-source vision-language-action model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:50.001978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.643669Z digest=sha256:9490231980c8318e2ec98a5421d8892c7e3182b17e8e5ad5e96e3244ca6976e2

Observation 7f395053-766f-41b7-ad25-7da0c9e89a57 · outbound

This paper cites A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models A novel performance evaluation methodology for single-target trackers.IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(11):2137–2155, Nov 2016

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.824518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.737506Z digest=sha256:c0d3a1756cc28b0f9b6b64a5f6353399a050b55851f9836e6432462fd1538039

Observation 97273674-23de-4b91-ae61-801372c2e4ee · outbound

This paper cites Person following robot based on real time single object tracking and rgb-d image.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Person following robot based on real time single object tracking and rgb-d image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.727647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.829613Z digest=sha256:d93be268eb8c035eedc3bc95d5ab1b2cc9a54e391f1485b090d2861865e4bea1

Observation 225dd818-5576-43b6-8d35-e74c33fcbff8 · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.536692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:44.927307Z digest=sha256:080c09d726c3ed420f769fe300ceb70345d626146f47905253fb6a1e9740ff2c

Observation 4417db72-92aa-45cc-8bf4-bc54c8dfe350 · outbound

This paper cites Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vi- sual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.380606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.026284Z digest=sha256:c5279490a9bb65c2a9893a2c356f56582442eaa2b0295ef594d0d9c5caa87af1

Observation 6f02cfff-2347-4e62-8510-e8c41d2acedc · outbound

This paper cites End-to-end active object tracking via reinforcement learning.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models End-to-end active object tracking via reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:49.120498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.123410Z digest=sha256:a6a0551a4ef0950133d45d5034d520a476da52069d244a11fb49acf8236d9e16

Observation 093d7d60-5442-4e70-97ac-08a00379108d · outbound

This paper cites Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Curious george: An attentive semantic robot.Robotics and Autonomous Systems, 56(6):503–511, 2008

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.945402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.249150Z digest=sha256:45ea64d89863e5c4bd6b3dec06b96d11c98629f2bb662cc38003728b93f48f55

Observation 2ccb200f-9763-4a12-b6f4-52931b22ea36 · outbound

This paper cites The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models The hands-free push-cart: Autonomous following in front by predicting user trajectory around obstacles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.787461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.348257Z digest=sha256:6bc3e353f03f725f18ad372e30d4e0b3563d103ea8926938aea255329b58a995

Observation fa7968d5-f715-4a43-bfc7-33a0b93f4ed8 · outbound

This paper cites Unrealcv: Virtual worlds for computer vision.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Unrealcv: Virtual worlds for computer vision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.557921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.423214Z digest=sha256:0da68c49c99d15565c86536c513cee42b95e8a37ff68b1ae74c5475fba861a63

Observation a0c6ebb1-b3d7-45f0-a4e9-0f8eadb2f2ef · outbound

This paper cites Tracking multiple moving targets with a mobile robot using particle filters and statistical data association.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Tracking multiple moving targets with a mobile robot using particle filters and statistical data association

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.406896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.501658Z digest=sha256:ad8535c51cde6ea00605fa275df747832d07da64586c0e457480f720afb695dd

Observation f929f2e8-5a5c-43f3-b8ed-543c40b5b99e · outbound

This paper cites Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:48.185872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.573732Z digest=sha256:a6b883ea2445ef21f7369805d9230df0f332d8cca5c2efcc108de909bf22ad28

Observation f698e522-b59a-4366-81c7-c3897d833e31 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.971967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.633871Z digest=sha256:359a52f38a04832313c34947d8973f52eda765af98ca88c5c8a744842af92700

Observation bf9ea878-827c-4e9a-801d-5334361f5f9f · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.712882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.712882Z digest=sha256:73ec54133eb1b142ae33ee0f609a5799f46b6c4d6c2cab4d0769dc023ad4f40e

Observation ae97a3f0-c3f1-4b96-9694-406729479488 · outbound

This paper cites Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Vision- language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:45.745565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:45.745565Z digest=sha256:997958f5e3861304e0583b790829af6c63e2dca97491b685910fc97001c52bd6

Observation b8c1f7a8-18fb-4977-97da-77328d528d46 · outbound

This paper cites Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Ad-vat+: An asymmetric dueling mechanism for learning and understanding visual active tracking.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5):1467–1482, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.772963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.848692Z digest=sha256:14b6f334590e8a3199c423560c71ee19c9c78101084de19ff371a442b4971b57

Observation 03053d95-a475-4b71-9abf-f544c9f0f1a0 · outbound

This paper cites AD-V AT: An asymmetric dueling mechanism for learning visual active tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models AD-V AT: An asymmetric dueling mechanism for learning visual active tracking

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.591371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.913847Z digest=sha256:c45864ab39f2799f7759f252d26301c5d1e6ed96319eeeb0abb9e527cc474d19

Observation e7dcbe83-8bb0-40c4-ad19-73d0dcef778a · outbound

This paper cites Towards distraction-robust active visual tracking.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Towards distraction-robust active visual tracking

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.403448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:45.993230Z digest=sha256:80adafdab8525d8eb59d89a7b0bcf08ceb4b3cb693c2109f60ba7439f5b6f37c

Observation 4ecb89fa-7d86-472b-9434-d00fa02bbbab · outbound

This paper cites Empowering embodied visual tracking with visual foundation models and offline rl.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Empowering embodied visual tracking with visual foundation models and offline rl

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.194940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:46.111382Z digest=sha256:ac72b6e3262d4211e190d58598fc7883b426dede6d5552dc3cbe544063511c11

Observation c0649b78-027a-409b-add2-5304a117d34a · outbound

This paper cites UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:46.216740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:46.216740Z digest=sha256:50aba12e7c63a76827c9bee473f88a02748646f72a6a6b5124f7cf87c7d51472

Observation 7286d023-c9cd-4e30-b802-7eceb1b93786 · outbound

This paper cites On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models On deep recurrent reinforcement learning for active visual tracking of space noncooperative objects.IEEE Robotics and Automation Letters, 8(8):4418–4425, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:47.054337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:46.295861Z digest=sha256:e4079cb078d9a5f907eaec367d747b2defc2e3d9312957e9d2a6e7fa4968553a

Observation 571ac838-6b22-4f3d-9760-004a84501875 · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:51:46.915857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:51:46.426406Z digest=sha256:69bf42b0d513e92a0c7c8280168b87ff3afec7594cf82d353c2ade521f5be181

Pith citing papers

No inbound Pith citation observations are available.