Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:35:54.922566Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2505.12606.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:35:54.922566Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-22T06:33:36.846345Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T06:34:40.918877Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aad54390-de3a-4a5a-95df-27ef567da53e · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Fully-convolutional siamese networks for object tracking
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 705244ed-0bea-4c29-9a10-8ae7a3b910aa · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Transformer tracking
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ff937f3d-1851-40dc-a209-2e67632d671a · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking ECO: Efficient convolution operators for tracking
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d32eeea5-3ddd-402c-990e-eae624078da9 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Probabilistic regression for visual tracking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bbdb12e2-af53-4c6c-a499-2a072b2a8863 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking An image is worth 16x16 words: Transformers for image recognition at scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b366b8f-9586-46cb-b145-ea7b20d19f56 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking LaSOT: A high-quality benchmark for large-scale single object tracking
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5d47ed8a-5bf9-45a7-ba58-da819f8d2924 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Siamese natural language tracker: Tracking by natural language descriptions with siamese trackers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0c37997d-1ebc-4549-900c-e8e6b166bac6 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d49ff0b4-d774-4e4f-840f-c88c7f39e096 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Deep adaptive fusion network for high performance RGBT tracking
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9039c009-ee7f-48db-a826-6a966bce9efc · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Are we ready for autonomous driving? the kitti vision benchmark suite
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1ffbb455-5557-4523-a904-ec6a96719e70 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Generative adversarial nets
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4918259-53c2-4678-a055-1bb0006665bc · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking High-speed tracking with kernelized correlation filters
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eba6c020-68c9-48e2-9887-ed7ed8bdf3cb · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Denoising diffusion probabilistic models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d6b52a0-b017-4c7d-a44d-af85e2f9eaa5 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Onetracker: Unifying visual object tracking with foundation models and efficient tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 53c087ad-0fe2-465d-b88e-b8c54f4eac34 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Sdstrack: Self-distillation symmetric adapter learning for multi-modal visual object tracking
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7656bdf6-b8b2-493f-9ca7-898a8fb3eb33 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Lora: Low-rank adaptation of large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1adf717-bd3c-4510-8c8a-cf3da1e4ee0a · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Got-10k: A large high-diversity benchmark for generic object tracking in the wild
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 94682704-5871-4cbd-b884-8674ebbbdcc9 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Visual prompt tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2659a8e-a6cb-462d-9fb0-da0f2ba095cd · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Repurposing diffusion-based image generators for monocular depth estimation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f2d0d2-827c-467c-bc63-55a762686fc0 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking The sixth visual object tracking VOT2018 challenge results
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 72efa579-c028-4ee4-8ef2-e4bac786af99 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking The tenth visual object tracking vot2022 challenge results
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8603501e-d3cf-4def-8a14-f77785fe6045 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Stable diffusion image variations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9e1e1931-c5e5-455c-82c0-1243549691d4 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Cornernet: Detecting objects as paired keypoints
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbfdce2a-3218-4e45-86ab-382b5af5191b · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking SiamRPN++: Evolution of siamese visual tracking with very deep networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7898066-aa07-482e-848e-69bfec4f8702 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking RGB-T object tracking: Benchmark and baseline
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 51b3b29f-a612-4c80-9ec4-dbac26752a47 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Lasher: A large-scale high-diversity benchmark for RGBT tracking
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation afd604ce-d5c5-4fc2-b98d-a5d1a548c31d · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Tracking by natural language specification
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d5e3f1a-0909-4de1-8a25-91d78f85cd3f · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Belongie, Lubomir D
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 20745626-58d5-41df-ba74-4100dcb190d8 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Decoupled Weight Decay Regularization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f1e6d9-600f-474a-b2df-63c900074710 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Narrowing the Synthetic-to- Real Gap for Thermal Infrared Semantic Image Segmentation Using Diffusion-based Conditional Image Synthesis
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7107b7b0-9967-4248-be31-125e7f6def6a · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8eab2a34-a4e0-4a9c-822f-5b1e9f89ef7f · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking TrackingNet: A large-scale dataset and benchmark for object tracking in the wild
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3f34eb64-95e4-4350-b1f9-8edd91376981 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Learning multi–domain convolutional neural networks for visual tracking
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06572793-932b-4812-b7fd-c1d15de05a9c · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Scalable diffusion models with transformers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6becfe6-a078-4424-9369-12ee2167dc43 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Zero: Memory optimizations toward training trillion parameter models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f25372a8-6b09-4899-afca-a59a657b5e8a · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Generalized intersection over union: A metric and a loss for bounding box regression
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e5bb35ee-c3c2-4347-807e-5f6d2efc9523 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking High-resolution image synthesis with latent diffusion models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b38bd6-4934-457f-9702-677e1ba3f9fb · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3c7583-527e-41e0-94d5-960d5c0a388a · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Transformer rgbt tracking with spatio- temporal multimodal tokens
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 120df161-bacc-49dd-a92f-0424c5d79b4b · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking XTrack: Multimodal Training Boosts RGB-X Video Object Trackers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdac64c7-f6e5-43b9-813e-aca8e782377d · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Emergent correspondence from image diffusion
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f5b567f-4e7b-44ea-94b3-516d87f59976 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16d02971-0acc-4a4c-b962-87c7ff61f642 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking VisEvent: Reliable Object Tracking via Collaboration of Frame and Event Flows
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82250ec-9399-40c9-ac99-5907c4d2d57d · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fcbd34df-7175-43d4-ba72-8521d8f12138 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking E-motion: Future motion simulation via event sequence diffusion
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e43429da-1f73-46ba-a141-007fcc74b46e · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Object tracking benchmark
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 474d023e-4cb3-4778-bce1-3d90c0db9e8f · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Single-model and any-modality for video object tracking
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2d125a98-848b-4338-952e-4f2c3f4aa086 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Multiple human tracking based on multi-view upper-body detection and discriminative learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b1c44179-b610-4b3c-988a-7d78f257b2e7 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Learning spatio-temporal transformer for visual tracking
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2afebe87-e411-425d-b9e1-4da6c7381c7c · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Depth- track: Unveiling the power of RGBD tracking
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eb323a9a-8769-4d4f-81fa-bbf922ec5b0c · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Prompting for multi-modal tracking
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b13b2ce1-f5f4-4055-93b5-5819dc2d2041 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Joint feature learning and relation modeling for tracking: A one-stream framework
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e90c5164-1e3e-4642-8441-2d4729a1f1a1 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Adding conditional control to text-to-image diffusion models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70bd387-45d9-42bc-8baa-ff59514fe79d · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Paste, Inpaint and Harmonize via Denoising: Subject-Driven Image Editing with Pre-Trained Diffusion Model
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43dfd1b4-8c22-44de-ada1-37721fc376c7 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Diff-tracker: Text-to-image diffusion models are unsupervised trackers
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7b831a14-542d-44f0-9bfc-93cd53ab7521 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Unleashing text-to-image diffusion models for visual perception
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b8efab1-7067-49d6-930a-7556cc8aef71 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Joint visual grounding and tracking with natural language specification
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fde0bfcc-c433-4c65-a027-5eba4f70193c · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Visual prompt multi-modal tracking
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6b2969ce-dac8-4cc4-ae6b-85040f083fb5 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking RGBD1K: A large-scale dataset and benchmark for RGB-D object tracking
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c8ef7b8-09b6-48c3-80aa-c1cb5f9aba70 · outbound
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Exploring pre-trained text-to-video diffusion models for referring video object segmentation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ac86f337-d99a-41c3-a8ad-45ee4d7ba08d · inbound
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.