Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:09:21.286253Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2412.19648.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:09:21.286253Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9ce8b7bc-cf57-4c1c-88d0-a949b62bddb3 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Online object tracking: A benchmark,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8d0f460c-1756-4474-8eec-a2ce41200fad · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Sotverse: A user-defined task space of single object tracking,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d9851940-fb6b-4c24-81b0-f8910b57ce6e · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Global instance tracking: Locating target more like humans,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3500bf63-3a49-4cca-8927-a1dd9ab70992 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Biodrone: A bionic drone-based single object tracking benchmark for robust vision,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9399313f-73b0-4606-b4b3-d8027d4c8ec6 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Tracking by natural language specification,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4e213500-0984-49da-ae46-608e92f176cc · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues A multi-modal global instance tracking benchmark (mgit): Better locating target in complex spatio-temporal and causal relationship,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0dc5b86e-cb77-4273-bfa4-8e437ec7594a · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Dtllm-vlt: Diverse text generation for visual language tracking based on llm,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b351c5-372e-4071-9993-8bc2101de048 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d03555c-6e21-48d2-902d-7c5898d56d63 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Memvlt: Vision-language tracking with adaptive memory- based prompts,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 72a449c3-916d-483c-b773-c6748222424e · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Transvg: End-to-end visual grounding with transformers,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37070dac-d77c-4256-8bec-02a2926166d7 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4459fc2c-02ae-49cf-8bc9-f7edce7d603f · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounded language-image pre- training,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcddb25-e650-4ee8-b3b4-c0535bc4e394 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Towards more flexible and accurate object tracking with natural lan- guage: Algorithms and benchmark,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9ca28498-f976-42c1-98f6-0dc8e360ed7d · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Lasot: A high-quality benchmark for large-scale single object tracking,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a2327595-219d-46f3-acd5-b2ac41e0e6ff · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Siamese natural lan- guage tracker: Tracking by natural language descriptions with siamese trackers,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f01f6b6-3c2c-4cd6-bb55-af78f3de0794 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7654db1-5777-4a16-b11b-f047a781db9c · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues All in one: Exploring unified vision-language tracking with multi-modal alignment,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8741e183-681a-4312-8d27-596395c369a9 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Context-Aware Integration of Language and Visual References for Natural Language Tracking
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ee9c9992-ebd6-4346-884d-cf09f4950604 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Textual tokens classification for multi-modal alignment in vision-language tracking,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f584d9bd-a673-4b29-802b-e4c400ab1dcf · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues One-stream stepwise decreasing for vision-language tracking,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9346192e-2e6a-484a-afe4-8d8d037553f0 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Joint feature learning and relation modeling for tracking: A one-stream framework,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 423a5b03-a574-48bc-91cd-d1d894be3c7d · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f00a9b8-4f42-4e0a-8986-817c701d7ea0 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues ODTrack: Online Dense Temporal Token Learning for Visual Tracking
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98dc6def-3d1b-48ec-b923-a889b466b1fc · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Beyond accuracy: Tracking more like human via visual search,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f5ec92d6-af42-4fbd-9d16-bac37de21266 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Attention is all you need,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e8ee257-f920-4471-a286-46d129370c3b · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues A hierarchical theme recognition model for sandplay therapy,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6e1b57d6-ec9e-400f-8d42-fa7985f4347c · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Emergent open-vocabulary semantic segmentation from off-the-shelf vision-language models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4970cd11-8fb6-4364-83cd-c3125459b98d · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounding everything: Emerging localization properties in vision-language trans- formers,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3437ab9c-df7c-483d-bc71-ca6163217b25 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6451c08e-91d8-4db6-8142-d283ad44ae7f · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Real-time visual object tracking with natural language description,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d93aa292-e641-4391-a64c-e0c56caf243f · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Grounding-tracking- integration,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f376ada5-b882-4924-aaa1-bec3c8103258 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Cross-modal target retrieval for tracking by natural language,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a4b03346-ff1d-41d5-a88f-e6be11836e7a · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Divert more attention to vision-language tracking,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 34ecee71-2ae1-41c3-8713-6c111625ee60 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Transformer vision- language tracking via proxy token guided cross-modal fusion,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 099272b9-5f48-4a98-9f3c-0ed22c46eda5 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Joint visual grounding and tracking with natural language specification,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fa10191e-8c75-4359-bc34-d7f56c52edfb · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Unified transformer with isomorphic branches for natural language tracking,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5e8e7f31-8eb6-4c01-a740-49ce5db7ebb8 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Tracking by natural language specification with long short-term context decoupling,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c07f8fe7-5b1e-48ac-8749-690803189283 · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Towards unified token learning for vision-language tracking,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5e406d4f-84ec-4d62-9e4f-bdb7ca81c7ee · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Onetracker: Unifying visual object tracking with foundation models and efficient tuning,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ba124da0-7489-43ba-88d1-7a98ea6d5dff · outbound
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.