Pith. sign in

Paper Citation Record · LEDGER

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2507.00454.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00454 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:08.038670Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:20:54.642591Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:24:21.277633Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0634cba0-2421-4b87-88c0-09f2b4f9763e · outbound

This paper cites Hiptrack: Visual tracking with historical prompts.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Hiptrack: Visual tracking with historical prompts

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.929103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:03.495167Z digest=sha256:9905105779b509316fe83a1abca9ff85e9b8789326099dab10eb04ecdd9723ae

Observation eabbfb60-8b03-41bb-91cc-f8cc805ca021 · outbound

This paper cites Robust object modeling for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Robust object modeling for visual tracking

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.915407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:03.562427Z digest=sha256:64a6f7b283e7aa1ac03ba03ce6b90fc96a8557af5b9aa76c1507608c7e163e60

Observation b2115ddc-afda-49a4-93aa-6452e2a3f57e · outbound

This paper cites Ost: Refining text knowledge with optimal spatio-temporal descriptor for general video recognition.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Ost: Refining text knowledge with optimal spatio-temporal descriptor for general video recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.901038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:03.635785Z digest=sha256:173d1851713c3adf05dbaa576825c6eaa94bbd7e662008f03a51701a113bec09

Observation d6c5bd46-d76c-4189-8199-7bc98634a03e · outbound

This paper cites Seqtrack: Sequence to sequence learning for visual ob- ject tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Seqtrack: Sequence to sequence learning for visual ob- ject tracking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.887068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:03.711114Z digest=sha256:c53188b5d28139b27644628083016c0a941eb950edf93dfe849801e084e08fbd

Observation a36d3e90-8321-4c59-9d61-30a9f50cbdd6 · outbound

This paper cites Mixformer: End-to-end tracking with iterative mixed atten- tion.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Mixformer: End-to-end tracking with iterative mixed atten- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.873201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:03.807241Z digest=sha256:438937048f6d521a523a23d9b0d1e249a835660d8225e262857c377d0c897c0f

Observation b4c15925-262e-43c1-93f4-db01dc76931c · outbound

This paper cites MixFormerV2: Efficient Fully Transformer Tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales MixFormerV2: Efficient Fully Transformer Tracking

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:21:08.313273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:03.895699Z digest=sha256:0ff7dc9a15a8ca93e87a5db6bee66615defd58f37ac6c269f01b9141c4d736a7

Observation 2b70bf93-1333-489f-aee9-e2bb68d051e6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:04.010181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:04.010181Z digest=sha256:3ab39c115def65e7569aeaa02acf95265001e7deb53f9be61096222dc03dbc4d

Observation 2da48220-6dc6-4af6-9e3a-21ca566761d3 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.859467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.110836Z digest=sha256:b1556164374b916f537a8177c8a5cacc5ce479ba9c2e5b8d6804589fc36fa6c3

Observation c932f6be-67eb-488a-8651-65a9c8cf0862 · outbound

This paper cites Siamese natural language tracker: Tracking by nat- ural language descriptions with siamese trackers.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Siamese natural language tracker: Tracking by nat- ural language descriptions with siamese trackers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.846051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.177793Z digest=sha256:a0b8607061f7a09ffa2e335adf746d447828bca760f419af932ec32649408144

Observation f43e7444-8140-4443-8336-05a45de9ae8a · outbound

This paper cites Generalized relation modeling for transformer tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Generalized relation modeling for transformer tracking

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.831765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.242868Z digest=sha256:2a3e20f6e448574839fafe5ae32bbda4c898a366c4d8d2cba591d459d1a1938f

Observation a702d729-0758-4471-aacc-60cb5f5e99a8 · outbound

This paper cites Siamcar: Siamese fully convolutional classification and regression for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Siamcar: Siamese fully convolutional classification and regression for visual tracking

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.818032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.345312Z digest=sha256:cc178c00ab98231e0c0964c51f65f510b5ec832ea3313fb5d78f8ca88c434665

Observation b178a3a2-e6b5-4675-8a29-79d9a77bd249 · outbound

This paper cites Divert more attention to vision-language tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Divert more attention to vision-language tracking

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.804426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.494381Z digest=sha256:ad69484dfd7c223ffc4bb77b43ae4c060201ecc0a6aaad0e1fc80de7914acd6e

Observation d1e1ea23-8165-4bd7-b684-830a03a2a0bc · outbound

This paper cites Masked autoencoders are scalable vision learners.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Masked autoencoders are scalable vision learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.790001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.590862Z digest=sha256:76cf71488cb89a7fdf8af393c42cb85545b08c22b66ad6dff5f363efa85deaca

Observation c21eb719-5b92-4c1f-ba85-e0dc9d398176 · outbound

This paper cites Target-aware tracking with long-term context at- tention.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Target-aware tracking with long-term context at- tention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.776016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.749674Z digest=sha256:783e4fce6c720101bf8053f2a3da0e93a74ac6d24bb133b45078d51ce7b76781

Observation db567757-5c6d-4cbf-b351-a7e00a7646fe · outbound

This paper cites A multi-modal global in- stance tracking benchmark (mgit): better locating target in complex spatio-temporal and causal relationship.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales A multi-modal global in- stance tracking benchmark (mgit): better locating target in complex spatio-temporal and causal relationship

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.761688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.878689Z digest=sha256:ed9e6c65567709876a562c75829f8a0d70cdb53a12c34f7f0395df8b4464e2af

Observation 70a604bd-57af-43ad-af90-1cebdadefccb · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.748053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:04.990174Z digest=sha256:17698b9f52f42278c472eb8a1011c1aea44600d5d1e04b3c6837a0935c203419

Observation 1a1abc37-1e63-44cd-bd3c-b35a96fa3a58 · outbound

This paper cites Rtracker: Recoverable tracking via pn tree structured memory.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Rtracker: Recoverable tracking via pn tree structured memory

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.733707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.106071Z digest=sha256:c9a63a1f7b84ab83d2117a268cc3ad6b626ef4fc2bfd5c592dbc370591330a53

Observation 391f9ee8-b9d4-4ba3-b8d6-722e483c4b3f · outbound

This paper cites Towards sequence-level training for vi- sual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Towards sequence-level training for vi- sual tracking

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.718635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.230930Z digest=sha256:062dd2dc67cd64018bbd1cd3b5e5ddaadb6a4197daffc79f734032e9877db8c9

Observation c26405f6-b8d3-403b-b6a2-eac472e374ce · outbound

This paper cites Citetracker: Correlating image and text for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Citetracker: Correlating image and text for visual tracking

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.635446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.358470Z digest=sha256:b4e1457d52ffb23e2937dfaa36e98c0d3915ec898faf933821b43c2728f0b09a

Observation 886f1d5f-ed15-4466-93f3-ba4143f64452 · outbound

This paper cites Dtllm-vlt: Diverse text generation for visual language tracking based on llm.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Dtllm-vlt: Diverse text generation for visual language tracking based on llm

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.302631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.489098Z digest=sha256:9a2c373d0f69303ce303139968ec14b094ea7c6ae1af58e7ceacf7ef3961916a

Observation 4f11bd9e-500a-4cdd-8786-fb5beb1d4da9 · outbound

This paper cites Beyond MOT: Semantic Multi-Object Tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Beyond MOT: Semantic Multi-Object Tracking

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:21:08.163075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.618383Z digest=sha256:da15902663e7206bce89bd9511374783440c92e6624d54fb261182acf5671935

Observation 2b0a9cbc-00e2-4fba-bcc1-13c1fb6f4899 · outbound

This paper cites an unresolved cited work.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:21:13.057046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.728136Z digest=sha256:d163a4acd4d3885b03b29362adc5984d0ac02f87a3212b9c33899334d2b9fac3

Observation 5011503f-7aa2-4a2d-8056-94d7647f901b · outbound

This paper cites Swintrack: A simple and strong baseline for trans- former tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Swintrack: A simple and strong baseline for trans- former tracking

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.777830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.856313Z digest=sha256:d42be165092a141d86fec217d7168b816f3fb321fcccb4d18316465e17013d04

Observation eaf084b2-d4ad-448d-9ec6-fbf7829d4b75 · outbound

This paper cites Tracking meets lora: Faster training, larger model, stronger performance.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Tracking meets lora: Faster training, larger model, stronger performance

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.550238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:05.995947Z digest=sha256:2e460fba4dd004a9832a732cbbbe0e4f88667dc0bcf3418c8865e0bc6252d1df

Observation 77c8ef6c-dc8b-43a3-bc8a-b950ded5fc8e · outbound

This paper cites Tracking by natural language specification with long short-term context decoupling.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Tracking by natural language specification with long short-term context decoupling

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.343384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:06.105323Z digest=sha256:558957858d338afb8d9c8ac61d17b1f58cb0b833405c29a1045b09bd9cb0a076

Observation 4f3ab851-5150-4666-9679-f7bf9a13dc81 · outbound

This paper cites Unifying visual and vision-language tracking via contrastive learning.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Unifying visual and vision-language tracking via contrastive learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.123095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:06.246551Z digest=sha256:7571631c024e3a115990771fb5b15de08ca8e8c1763f7c4aea6c72d255b5ca1c

Observation 4afffe9b-0de5-45a7-a5a3-9e8e41b1db7b · outbound

This paper cites Trackingnet: A large-scale dataset and benchmark for object tracking in the wild.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Trackingnet: A large-scale dataset and benchmark for object tracking in the wild

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:11.866219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:06.412887Z digest=sha256:dfb9e2da3dcca5486a2a9f0b9c4081c5d472f3dadc11496d2d495f2f9819da87

Observation 11ce511f-a597-4d94-8a2d-43bd5d5a5c90 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Learning transferable visual models from natural language supervi- sion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:06.545414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:06.545414Z digest=sha256:a1759c75e7f811dd5cac76c7c3452bd1b77fb68c0e0f5b3f9895faa633000117

Observation 2ccac787-eac7-48fc-b362-5d96d258a98a · outbound

This paper cites Context-aware integration of lan- guage and visual references for natural language tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Context-aware integration of lan- guage and visual references for natural language tracking

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:11.513979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:06.691730Z digest=sha256:7fcb3bf5202263a2121c59589d39b17c8eb52df2d81ddae84065873e20f20766

Observation 47e8c3f8-087a-489d-842b-f5551db649d3 · outbound

This paper cites Chat- tracker: Enhancing visual tracking performance via chatting with multimodal large language model.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Chat- tracker: Enhancing visual tracking performance via chatting with multimodal large language model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:11.202632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:06.808290Z digest=sha256:219eaf882ed8ff2abf518209daed3e8af84f8f6e7552f090923b382b76f60067

Observation 6905c2cb-a44a-4f2a-a8e6-199fbbf32e9a · outbound

This paper cites Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:10.992361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:06.922319Z digest=sha256:3089ddd436c1da01c14cac0fb6680b579bea47be0aca08ae4941766701d0883d

Observation 0440d920-f877-4a4a-995d-03cc064e50ca · outbound

This paper cites Autoregressive visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Autoregressive visual tracking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:10.682233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.042125Z digest=sha256:5c7ffb52f23cb25bf6ac33fc1889174177f81584325456ee909caada15b6208d

Observation b970d63e-fa50-44fe-a6dd-600895141981 · outbound

This paper cites Improving visual grounding with multi-scale discrep- ancy information and centralized-transformer.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Improving visual grounding with multi-scale discrep- ancy information and centralized-transformer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:10.402450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.202804Z digest=sha256:7fa0a4d9f56cf67abd3469c67124e5717da766bd7e9a9112f255f06019eea6cc

Observation 70f2c872-bded-448e-bab3-b82b5fd045f4 · outbound

This paper cites an unresolved cited work.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:21:10.155071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.347044Z digest=sha256:c07fbf3f89a1ab6cd8525ada606e4b4435c861088710ec3748e5be5daf93404f

Observation 0e722b8f-2db2-4702-a60c-816d6726111a · outbound

This paper cites Object track- ing benchmark.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Object track- ing benchmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.720456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.447212Z digest=sha256:0bc2e905cbab86bf25ec3dcbb4c02e2238e693157fe7f4275b6e7fe8eb4f1946

Observation 4ea25646-625c-4125-a4bb-cd649532ef37 · outbound

This paper cites Correlation-aware deep tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Correlation-aware deep tracking

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.516437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.587470Z digest=sha256:f1368ce7f7ab65a28f89c8745c2c0e10bbdcc3557916a639fe9ba88353b9d85a

Observation 502c8dee-b8f0-4046-acd8-0faf4c29afeb · outbound

This paper cites Autore- gressive queries for adaptive tracking with spatio-temporal transformers.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Autore- gressive queries for adaptive tracking with spatio-temporal transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.363218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.678542Z digest=sha256:53bf7cedc2a012708239510166e766733e01dcec5788404e6d17a0e482f61081

Observation 54f03118-3210-4a38-92eb-94896ca203ec · outbound

This paper cites Learning spatio-temporal transformer for vi- sual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Learning spatio-temporal transformer for vi- sual tracking

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.150372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.754690Z digest=sha256:117cf8642f5d8b5a7d43f4cd08be390c918cd67a46351aba69267b93d701754f

Observation 33cdc7ef-33c0-4e99-984a-0690f2c19783 · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Joint feature learning and relation modeling for tracking: A one-stream framework

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.924328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.813342Z digest=sha256:bbd2c02cb737ae7501f1648aff79d4b4e7658b9d6e9c346aee870a842bf84cfa

Observation 09e8bdca-f269-459f-a03f-01a7f7d42358 · outbound

This paper cites Exploring the feature extraction and relation modeling for light-weight transformer tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Exploring the feature extraction and relation modeling for light-weight transformer tracking

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.809214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.876163Z digest=sha256:71aea722edc5b98f1c66b595ac5e68d0677b6f634eb8a37367d2111ed9e6f940

Observation 536240b9-8b56-4e9f-b962-66ad43d28bc0 · outbound

This paper cites Ji, and Xianxian Li.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Ji, and Xianxian Li

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.669724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.946025Z digest=sha256:ce84578e5e96b4365913af8709f94777de1163f3b9d271ea44760ae28cc77f2f

Observation 329e24f2-28f4-4ac2-8d79-c3daba8fe15b · outbound

This paper cites Odtrack: Online dense temporal token learning for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Odtrack: Online dense temporal token learning for visual tracking

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.554193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:07.998946Z digest=sha256:f5b6b6a571d19f500bb1769a9c5c7c0cc1d0596746ac43eba9bf5d9501a0bc18

Observation 587798bc-d420-4fb2-8e7a-8fb7a4e99f4c · outbound

This paper cites Joint visual grounding and tracking with natural language specifi- cation.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Joint visual grounding and tracking with natural language specifi- cation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.433120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:21:08.038670Z digest=sha256:38db8e94f54a3ec0d0743ad3ec38b30adad84bc2fe410cf2a04b8f397e94cf25

Pith citing papers

Observation 9bd60c18-a9cc-4ce4-9dc2-387bbc900887 · inbound

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking cites this paper.

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.279204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:20:54.642591Z digest=sha256:370f45d5166135107d22196cec5c324feef91175230eae62272acf4a9932020e