Pith. sign in

Paper Citation Record · LEDGER

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

As of 15 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2412.19648.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19648 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:09:21.286253Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ce8b7bc-cf57-4c1c-88d0-a949b62bddb3 · outbound

This paper cites Online object tracking: A benchmark,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Online object tracking: A benchmark,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.922959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.099080Z digest=sha256:68e8fa7eb94eee060c5857f01264e3a47f2ca08d20a538b960fe9cf3b780867e

Observation 8d0f460c-1756-4474-8eec-a2ce41200fad · outbound

This paper cites Sotverse: A user-defined task space of single object tracking,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Sotverse: A user-defined task space of single object tracking,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.909695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.104964Z digest=sha256:3ffeca0bbab5a73dc730efbe2bce9b50b1177a23e38074cecfdef6a2d9a66cb5

Observation d9851940-fb6b-4c24-81b0-f8910b57ce6e · outbound

This paper cites Global instance tracking: Locating target more like humans,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Global instance tracking: Locating target more like humans,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.895254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.112347Z digest=sha256:a35dda8f34be27b85e6652cbfa1d2931788070ae539bb92c76c4460a3857702d

Observation 3500bf63-3a49-4cca-8927-a1dd9ab70992 · outbound

This paper cites Biodrone: A bionic drone-based single object tracking benchmark for robust vision,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Biodrone: A bionic drone-based single object tracking benchmark for robust vision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.882730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.116056Z digest=sha256:9581d80ee0efccafb5216d7530d8b357e28347f4705c98e2784eb820a5f0aa84

Observation 9399313f-73b0-4606-b4b3-d8027d4c8ec6 · outbound

This paper cites Tracking by natural language specification,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Tracking by natural language specification,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.868600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.120361Z digest=sha256:346299daeca63de379c64794061b6948a3b7b7de9e735b68fe2aa42f6f2d80b7

Observation 4e213500-0984-49da-ae46-608e92f176cc · outbound

This paper cites A multi-modal global instance tracking benchmark (mgit): Better locating target in complex spatio-temporal and causal relationship,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues A multi-modal global instance tracking benchmark (mgit): Better locating target in complex spatio-temporal and causal relationship,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.855129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.125309Z digest=sha256:41ff02462ae5b197e98d64a1750e1c7871ed29afc38fecf431fd7a3b4db857da

Observation 0dc5b86e-cb77-4273-bfa4-8e437ec7594a · outbound

This paper cites Dtllm-vlt: Diverse text generation for visual language tracking based on llm,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Dtllm-vlt: Diverse text generation for visual language tracking based on llm,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.131846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.131846Z digest=sha256:06a902060c6f6f18577d03960e2f422afd197678a1a8210306cfb289428170cc

Observation f9b351c5-372e-4071-9993-8bc2101de048 · outbound

This paper cites How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.137235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.137235Z digest=sha256:89883a7913252c222097ade0bbe8d038227baefbd448da8160f87e6637ffa0dc

Observation 6d03555c-6e21-48d2-902d-7c5898d56d63 · outbound

This paper cites Memvlt: Vision-language tracking with adaptive memory- based prompts,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Memvlt: Vision-language tracking with adaptive memory- based prompts,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.831870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.142371Z digest=sha256:026846f4a58e97a0b69a6ae507495fd18c4b25a5cb052e6a2ff5a4367ecafb65

Observation 72a449c3-916d-483c-b773-c6748222424e · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Transvg: End-to-end visual grounding with transformers,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.147386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.147386Z digest=sha256:12f69e606c7f01789cffbaf1bd3fcaa07d28fd036596e6def311a23559d2b3d5

Observation 37070dac-d77c-4256-8bec-02a2926166d7 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.151547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.151547Z digest=sha256:c588ceae488c303f9ce4d737c14b23eae2bb0c917884b383b016743a44479399

Observation 4459fc2c-02ae-49cf-8bc9-f7edce7d603f · outbound

This paper cites Grounded language-image pre- training,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounded language-image pre- training,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.157276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.157276Z digest=sha256:895612875dd147aed6b05f165f43905076e5fca09861240d02f79fec4b9c1b76

Observation fdcddb25-e650-4ee8-b3b4-c0535bc4e394 · outbound

This paper cites Towards more flexible and accurate object tracking with natural lan- guage: Algorithms and benchmark,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Towards more flexible and accurate object tracking with natural lan- guage: Algorithms and benchmark,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.798433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.165684Z digest=sha256:ae62bbea020aafed7216b57d90c689c799ad15f7c97c367d0aa653d73eb173a0

Observation 9ca28498-f976-42c1-98f6-0dc8e360ed7d · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single object tracking,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Lasot: A high-quality benchmark for large-scale single object tracking,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.785969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.169718Z digest=sha256:b222d67daec26bc2c5b81adf009b150a725702df8e5426c4cf3f19841bfe6cc3

Observation a2327595-219d-46f3-acd5-b2ac41e0e6ff · outbound

This paper cites Siamese natural lan- guage tracker: Tracking by natural language descriptions with siamese trackers,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Siamese natural lan- guage tracker: Tracking by natural language descriptions with siamese trackers,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.773385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.174292Z digest=sha256:e74a5486c3bba76eafd4f0a4c0fb2ac6ee058e841a254c6332c7f15295b2b604

Observation 1f01f6b6-3c2c-4cd6-bb55-af78f3de0794 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.178184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.178184Z digest=sha256:e4eae903a6a04c9e313cf1cab859715175452b11f005c4aa2f85dbe22707b49c

Observation f7654db1-5777-4a16-b11b-f047a781db9c · outbound

This paper cites All in one: Exploring unified vision-language tracking with multi-modal alignment,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues All in one: Exploring unified vision-language tracking with multi-modal alignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.759758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.183263Z digest=sha256:5300b4852bc759f68ec3fb3857a0ed0eff63fad40cde00d976b8eea384f7b4ed

Observation 8741e183-681a-4312-8d27-596395c369a9 · outbound

This paper cites Context-Aware Integration of Language and Visual References for Natural Language Tracking.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Context-Aware Integration of Language and Visual References for Natural Language Tracking

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:09:21.414752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.188117Z digest=sha256:000c35ce1baf90a79afffda35f909e4413520792640feabf49bdbf8961afecae

Observation ee9c9992-ebd6-4346-884d-cf09f4950604 · outbound

This paper cites Textual tokens classification for multi-modal alignment in vision-language tracking,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Textual tokens classification for multi-modal alignment in vision-language tracking,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.743502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.193551Z digest=sha256:180037444f18cf918273be95e4dfef12862fbec1d1803f7e3e17a6cb249e11d4

Observation f584d9bd-a673-4b29-802b-e4c400ab1dcf · outbound

This paper cites One-stream stepwise decreasing for vision-language tracking,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues One-stream stepwise decreasing for vision-language tracking,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.729869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.197462Z digest=sha256:724fdda06d1a007b4e272543ebbcab9d457724183620bf4c7add27f55494c153

Observation 9346192e-2e6a-484a-afe4-8d8d037553f0 · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Joint feature learning and relation modeling for tracking: A one-stream framework,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.714147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.203237Z digest=sha256:f37e6bd27074df4805ed584c51267c3905ae5617bf64b50d7da09f91fe027e56

Observation 423a5b03-a574-48bc-91cd-d1d894be3c7d · outbound

This paper cites Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.208743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.208743Z digest=sha256:6ef68cda1b2351ee7e22a51ec04396770231ee61791971d748be858f24fc4368

Observation 3f00a9b8-4f42-4e0a-8986-817c701d7ea0 · outbound

This paper cites ODTrack: Online Dense Temporal Token Learning for Visual Tracking.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues ODTrack: Online Dense Temporal Token Learning for Visual Tracking

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.214232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.214232Z digest=sha256:6e1abc14b18874c0def67d74ce9ce64e9061621bdfd09cded945bfd5ade48054

Observation 98dc6def-3d1b-48ec-b923-a889b466b1fc · outbound

This paper cites Beyond accuracy: Tracking more like human via visual search,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Beyond accuracy: Tracking more like human via visual search,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.699965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.219980Z digest=sha256:7a79986862021281d174ed1312b7efae6f70e827850572cc0f9754fe932eb123

Observation f5ec92d6-af42-4fbd-9d16-bac37de21266 · outbound

This paper cites Attention is all you need,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Attention is all you need,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.685836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.224233Z digest=sha256:0f720a4635bedcd799de3dd6775dc187755d5939107645bf1fd5900750c36b6e

Observation 1e8ee257-f920-4471-a286-46d129370c3b · outbound

This paper cites A hierarchical theme recognition model for sandplay therapy,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues A hierarchical theme recognition model for sandplay therapy,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.671566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.227797Z digest=sha256:804be2be8dde1484f4c5ddcfb99eebf26d73ed84001760c88f3cee7a324755e6

Observation 6e1b57d6-ec9e-400f-8d42-fa7985f4347c · outbound

This paper cites Emergent open-vocabulary semantic segmentation from off-the-shelf vision-language models,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Emergent open-vocabulary semantic segmentation from off-the-shelf vision-language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.657452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.231253Z digest=sha256:81e0b9d6f63cdb55b46cd36c16d609050b04f0f5fa620ea8eafde2e166d3ea94

Observation 4970cd11-8fb6-4364-83cd-c3125459b98d · outbound

This paper cites Grounding everything: Emerging localization properties in vision-language trans- formers,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounding everything: Emerging localization properties in vision-language trans- formers,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.646507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.234927Z digest=sha256:23d25df3c509b493ab0e1bd892a9e422c65a9664f7e5a57812fd21974e0684f6

Observation 3437ab9c-df7c-483d-bc71-ca6163217b25 · outbound

This paper cites Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:09:21.353626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.239084Z digest=sha256:d994c157a46377b8f77d05177fcba7edf86a8c9b9d4b951cac7639daabbd3461

Observation 6451c08e-91d8-4db6-8142-d283ad44ae7f · outbound

This paper cites Real-time visual object tracking with natural language description,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Real-time visual object tracking with natural language description,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.633423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.242815Z digest=sha256:1efbf1568699f5e9c16ecc2b6b2765f818cee810d1cba6afbea92cc773b8d68d

Observation d93aa292-e641-4391-a64c-e0c56caf243f · outbound

This paper cites Grounding-tracking- integration,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounding-tracking- integration,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.618925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.246189Z digest=sha256:4d449ef83896d354659840c5dd97df2543f503c8032ef3bb35b42733c9136210

Observation f376ada5-b882-4924-aaa1-bec3c8103258 · outbound

This paper cites Cross-modal target retrieval for tracking by natural language,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Cross-modal target retrieval for tracking by natural language,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.604664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.250783Z digest=sha256:4e5f8ebabc24a3656e63ba4a1cbad9abb3c5e73c886083b8c2889a6fcaa4b437

Observation a4b03346-ff1d-41d5-a88f-e6be11836e7a · outbound

This paper cites Divert more attention to vision-language tracking,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Divert more attention to vision-language tracking,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.591510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.254618Z digest=sha256:0a8053f87c1eedd70a35f1843bb1dce32fcac04b3072e7ce788e85421a70c38f

Observation 34ecee71-2ae1-41c3-8713-6c111625ee60 · outbound

This paper cites Transformer vision- language tracking via proxy token guided cross-modal fusion,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Transformer vision- language tracking via proxy token guided cross-modal fusion,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.575307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.259156Z digest=sha256:93766473f777a51ff3aeca82b32da2036773e3aae067627e28cdbf0eba977245

Observation 099272b9-5f48-4a98-9f3c-0ed22c46eda5 · outbound

This paper cites Joint visual grounding and tracking with natural language specification,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Joint visual grounding and tracking with natural language specification,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.546758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.263416Z digest=sha256:77e394b50ad67ecb7ffaaf2d9c95c482c1fbdeda6fcd78fa96c2bfe4487738e9

Observation fa10191e-8c75-4359-bc34-d7f56c52edfb · outbound

This paper cites Unified transformer with isomorphic branches for natural language tracking,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Unified transformer with isomorphic branches for natural language tracking,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.529467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.267373Z digest=sha256:63c5b5fd26e08c31cc73ea24781b213612c58e1de0e23edccbf7c6fb63d95267

Observation 5e8e7f31-8eb6-4c01-a740-49ce5db7ebb8 · outbound

This paper cites Tracking by natural language specification with long short-term context decoupling,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Tracking by natural language specification with long short-term context decoupling,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.514100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.271905Z digest=sha256:fcce9c88f6743da80a67a268ce4443a4f5dd5ffdd8d1a6cb503310fa3610ab55

Observation c07f8fe7-5b1e-48ac-8749-690803189283 · outbound

This paper cites Towards unified token learning for vision-language tracking,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Towards unified token learning for vision-language tracking,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.499643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.276444Z digest=sha256:5dc4e5778c923ab8f7d3de4469fec6cd42b607f28d9086607102d38c2fb48fcb

Observation 5e406d4f-84ec-4d62-9e4f-bdb7ca81c7ee · outbound

This paper cites Onetracker: Unifying visual object tracking with foundation models and efficient tuning,.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Onetracker: Unifying visual object tracking with foundation models and efficient tuning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:09:21.484872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T00:09:21.281342Z digest=sha256:0355b501a96e2103839922c6fdf64bb5ad575c689b1b20e4452382feda77c3fc

Observation ba124da0-7489-43ba-88d1-7a98ea6d5dff · outbound

This paper cites Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness.

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:09:21.286253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:09:21.286253Z digest=sha256:9352073fe212c86d6dffc5ab7ea1fa216d7e052a938a322564c0f6bc9bcfeba3

Pith citing papers

No inbound Pith citation observations are available.