Pith. sign in

Paper Citation Record · LEDGER

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking

As of 23 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2505.12606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12606 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:35:54.922566Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T06:33:36.846345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T06:34:40.918877Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aad54390-de3a-4a5a-95df-27ef567da53e · outbound

This paper cites Fully-convolutional siamese networks for object tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Fully-convolutional siamese networks for object tracking

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.642494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.702365Z digest=sha256:47c39bf85844b3851a971b28d7492e428cea1e0ef151d2368737ae6d4f78e3a1

Observation 705244ed-0bea-4c29-9a10-8ae7a3b910aa · outbound

This paper cites Transformer tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Transformer tracking

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.629296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.707054Z digest=sha256:dc16514499fdd3ce0182fffe55117aaec52b67227f7462e1c975a5612f2fc28b

Observation ff937f3d-1851-40dc-a209-2e67632d671a · outbound

This paper cites ECO: Efficient convolution operators for tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking ECO: Efficient convolution operators for tracking

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.616611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.710838Z digest=sha256:559bbed0df8076ef7f012fffba468db5a201f3f9b252fd385e3ab5831d62cae9

Observation d32eeea5-3ddd-402c-990e-eae624078da9 · outbound

This paper cites Probabilistic regression for visual tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Probabilistic regression for visual tracking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.604160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.715055Z digest=sha256:c208bff116937196b9ecd5fe3055e95b22ed14e3ba22b82416a3a6f0f6721218

Observation bbdb12e2-af53-4c6c-a499-2a072b2a8863 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking An image is worth 16x16 words: Transformers for image recognition at scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.718873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.718873Z digest=sha256:61c754fe828e2f45a30db2f35d9134edcefe838c659175f119fdc7db8d68ef98

Observation 9b366b8f-9586-46cb-b145-ea7b20d19f56 · outbound

This paper cites LaSOT: A high-quality benchmark for large-scale single object tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking LaSOT: A high-quality benchmark for large-scale single object tracking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.583742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.722647Z digest=sha256:e7e5b9576b17f8fea6825342173af594984d4a292d52ffaecfea048fca29fec0

Observation 5d47ed8a-5bf9-45a7-ba58-da819f8d2924 · outbound

This paper cites Siamese natural language tracker: Tracking by natural language descriptions with siamese trackers.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Siamese natural language tracker: Tracking by natural language descriptions with siamese trackers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.571231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.726526Z digest=sha256:1b45dbf2f784c96e311e2fc7ca37df47a8fe217fc164d19e0778026f48679908

Observation 0c37997d-1ebc-4549-900c-e8e6b166bac6 · outbound

This paper cites Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.557463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.729996Z digest=sha256:e1c279c92fd45c0dde72be3c5c26643966ce30f87531377ee947855733ea8a18

Observation d49ff0b4-d774-4e4f-840f-c88c7f39e096 · outbound

This paper cites Deep adaptive fusion network for high performance RGBT tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Deep adaptive fusion network for high performance RGBT tracking

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.544300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.733720Z digest=sha256:c37d322186975dc15c4b0c8e9f158fb2675c2843c6badeaaca7fd18a7fbd0821

Observation 9039c009-ee7f-48db-a826-6a966bce9efc · outbound

This paper cites Are we ready for autonomous driving? the kitti vision benchmark suite.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Are we ready for autonomous driving? the kitti vision benchmark suite

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.532514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.737151Z digest=sha256:c1c86ca6229a57af48ab722f42f88894e51f4d64721bba09fd1d1b1583e97593

Observation 1ffbb455-5557-4523-a904-ec6a96719e70 · outbound

This paper cites Generative adversarial nets.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Generative adversarial nets

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.740686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.740686Z digest=sha256:f321a961f436553b1f393a14a718d6cf593ad4e84d8af11b80f6d725dc069b23

Observation d4918259-53c2-4678-a055-1bb0006665bc · outbound

This paper cites High-speed tracking with kernelized correlation filters.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking High-speed tracking with kernelized correlation filters

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.513657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.744379Z digest=sha256:1d9b63b5eb79694d5e9265731a1a01427973932802bacd33bb61ad36385b85a2

Observation eba6c020-68c9-48e2-9887-ed7ed8bdf3cb · outbound

This paper cites Denoising diffusion probabilistic models.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Denoising diffusion probabilistic models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.502132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.748061Z digest=sha256:d975c8a3907e015f1cc72a2105cd9b6cbdf4aff7d4d291e400fee3c6545f7108

Observation 0d6b52a0-b017-4c7d-a44d-af85e2f9eaa5 · outbound

This paper cites Onetracker: Unifying visual object tracking with foundation models and efficient tuning.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Onetracker: Unifying visual object tracking with foundation models and efficient tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.490749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.751624Z digest=sha256:2a36fec19dedd49e686cf31c05e5731e1d0467d6de31361ed9795d714e91debd

Observation 53c087ad-0fe2-465d-b88e-b8c54f4eac34 · outbound

This paper cites Sdstrack: Self-distillation symmetric adapter learning for multi-modal visual object tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Sdstrack: Self-distillation symmetric adapter learning for multi-modal visual object tracking

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.478126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.755267Z digest=sha256:b738a65e0010cf5fa6c7fa4b710042a04f9eee466b6cef3c3570ae9a6eb2603f

Observation 7656bdf6-b8b2-493f-9ca7-898a8fb3eb33 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Lora: Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.758961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.758961Z digest=sha256:03202f624d49865b17cfdde8d7bb2964208eef4be17e8e28864754a4991a01d3

Observation d1adf717-bd3c-4510-8c8a-cf3da1e4ee0a · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.458454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.762572Z digest=sha256:ecd5bdcbef123b57dfca0cb7de041177f9adc49b7da018442d7216012009ced7

Observation 94682704-5871-4cbd-b884-8674ebbbdcc9 · outbound

This paper cites Visual prompt tuning.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Visual prompt tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.765744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.765744Z digest=sha256:c181f85592ee2757247ab0ab24404ec800f3396c706fa6bab174d64b901c61c4

Observation d2659a8e-a6cb-462d-9fb0-da0f2ba095cd · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Repurposing diffusion-based image generators for monocular depth estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.769585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.769585Z digest=sha256:63d25757ee0bedf525e15aea9f0129e8a5a7fefcbea1b63cc4f7acff8127fde5

Observation 14f2d0d2-827c-467c-bc63-55a762686fc0 · outbound

This paper cites The sixth visual object tracking VOT2018 challenge results.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking The sixth visual object tracking VOT2018 challenge results

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.430270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.773584Z digest=sha256:27c4d3df48b259258cde87a0d39d9d4dfafe0928f3527896105ab571c9a3a8c4

Observation 72efa579-c028-4ee4-8ef2-e4bac786af99 · outbound

This paper cites The tenth visual object tracking vot2022 challenge results.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking The tenth visual object tracking vot2022 challenge results

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.417077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.777090Z digest=sha256:16522c4bd26c3f8c23a76b2fa22dfe629cff63a263517e1b514f4ef27208c2b6

Observation 8603501e-d3cf-4def-8a14-f77785fe6045 · outbound

This paper cites Stable diffusion image variations.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Stable diffusion image variations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.404686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.780579Z digest=sha256:9b88aa31480816f119a27f343427c1708a56b65ee1ec345523f2ca1a136b9b3e

Observation 9e1e1931-c5e5-455c-82c0-1243549691d4 · outbound

This paper cites Cornernet: Detecting objects as paired keypoints.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Cornernet: Detecting objects as paired keypoints

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.784051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.784051Z digest=sha256:d4bbc74d9bbaf03f6d7c98cd052b6bde0ab98770b7bb6295d2f98b961a719478

Observation fbfdce2a-3218-4e45-86ab-382b5af5191b · outbound

This paper cites SiamRPN++: Evolution of siamese visual tracking with very deep networks.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking SiamRPN++: Evolution of siamese visual tracking with very deep networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.787540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.787540Z digest=sha256:ae2bf01038a241150d7b9df0646e7df36fd04bd94eff5c60804fe56d12fd62f9

Observation d7898066-aa07-482e-848e-69bfec4f8702 · outbound

This paper cites RGB-T object tracking: Benchmark and baseline.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking RGB-T object tracking: Benchmark and baseline

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.378038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.791235Z digest=sha256:9e93bf369ae225b5f1ebeb058cf912288793a69e9c628031db86af1e47d98ae2

Observation 51b3b29f-a612-4c80-9ec4-dbac26752a47 · outbound

This paper cites Lasher: A large-scale high-diversity benchmark for RGBT tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Lasher: A large-scale high-diversity benchmark for RGBT tracking

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.365790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.795026Z digest=sha256:a2323b8d618ff6f763282b1da22710dd17072830b22c11443d9f065d7c1c2b52

Observation afd604ce-d5c5-4fc2-b98d-a5d1a548c31d · outbound

This paper cites Tracking by natural language specification.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Tracking by natural language specification

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.351744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.798878Z digest=sha256:2e57b360dff8e4c237bcd60ab74d7d1b57c103a2a525008d00cf9aa6253b0f15

Observation 3d5e3f1a-0909-4de1-8a25-91d78f85cd3f · outbound

This paper cites Belongie, Lubomir D.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Belongie, Lubomir D

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.337727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.802644Z digest=sha256:4a07f9a13a8e7b954433933ba9b9f85b5130df570ba0a1a7ac789114470f0434

Observation 20745626-58d5-41df-ba74-4100dcb190d8 · outbound

This paper cites Decoupled Weight Decay Regularization.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.805883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.805883Z digest=sha256:00546d5b4b2bf6f7a197144073556f2fca41762b37ad8beb6b7dd252ae24a71a

Observation d6f1e6d9-600f-474a-b2df-63c900074710 · outbound

This paper cites Narrowing the Synthetic-to- Real Gap for Thermal Infrared Semantic Image Segmentation Using Diffusion-based Conditional Image Synthesis.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Narrowing the Synthetic-to- Real Gap for Thermal Infrared Semantic Image Segmentation Using Diffusion-based Conditional Image Synthesis

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.324922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.809651Z digest=sha256:02ee7610f36554b7bd6fa788d75a71030eaddbf0767f84edf19ea6272c66c672

Observation 7107b7b0-9967-4248-be31-125e7f6def6a · outbound

This paper cites T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.312290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.813162Z digest=sha256:4670d14d5f3215312a0444d74aff95884d1824c876ff05cf03000ab166f13844

Observation 8eab2a34-a4e0-4a9c-822f-5b1e9f89ef7f · outbound

This paper cites TrackingNet: A large-scale dataset and benchmark for object tracking in the wild.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking TrackingNet: A large-scale dataset and benchmark for object tracking in the wild

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.299169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.816548Z digest=sha256:6959a2c942c3b317769a89102a5c8488995ff5c3c70f4556108e1d22b8c562a5

Observation 3f34eb64-95e4-4350-b1f9-8edd91376981 · outbound

This paper cites Learning multi–domain convolutional neural networks for visual tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Learning multi–domain convolutional neural networks for visual tracking

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.285935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.820084Z digest=sha256:c97f694744202dd8d021f4ac48cba14cf9a99ffbd956897fd0acff4e568c14b4

Observation 06572793-932b-4812-b7fd-c1d15de05a9c · outbound

This paper cites Scalable diffusion models with transformers.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Scalable diffusion models with transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.823611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.823611Z digest=sha256:688ee1bb9de627169f665b1028e1af794aa39a30c9ed63283bd2274dcccc8787

Observation b6becfe6-a078-4424-9369-12ee2167dc43 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Zero: Memory optimizations toward training trillion parameter models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.827538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.827538Z digest=sha256:8bf9e2dfed5f9559cd82fc3c85b30028dccab0d5c6143347c5f5d5aee90c7468

Observation f25372a8-6b09-4899-afca-a59a657b5e8a · outbound

This paper cites Generalized intersection over union: A metric and a loss for bounding box regression.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Generalized intersection over union: A metric and a loss for bounding box regression

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.256208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.831495Z digest=sha256:7c54504e61a9f4f03f142a5e519588f236425add4737fccfee25a9001bf7eafb

Observation e5bb35ee-c3c2-4347-807e-5f6d2efc9523 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.835071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.835071Z digest=sha256:9e301ac2c0e718a664d39e4a8db76c94adb5e3084a8b33efd83b5ee982d6025d

Observation d4b38bd6-4934-457f-9702-677e1ba3f9fb · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.838626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.838626Z digest=sha256:e8557bbb883865715210cd57c027e92966e01980eb9b96489263873a1598b1c7

Observation de3c7583-527e-41e0-94d5-960d5c0a388a · outbound

This paper cites Transformer rgbt tracking with spatio- temporal multimodal tokens.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Transformer rgbt tracking with spatio- temporal multimodal tokens

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.228589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.842442Z digest=sha256:9af0556a2f252874392fa52c89abbdbb23ec1c70e84dc6838c9967f88911b33e

Observation 120df161-bacc-49dd-a92f-0424c5d79b4b · outbound

This paper cites XTrack: Multimodal Training Boosts RGB-X Video Object Trackers.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking XTrack: Multimodal Training Boosts RGB-X Video Object Trackers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.846053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.846053Z digest=sha256:637dcbf58b7b7304b459eded8fe81cd138de8a5b73fcb4b5061906ec60614aea

Observation fdac64c7-f6e5-43b9-813e-aca8e782377d · outbound

This paper cites Emergent correspondence from image diffusion.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Emergent correspondence from image diffusion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.216727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.850354Z digest=sha256:05d1b4f59a03bd944c1cbd97b9b24e9bcd58a1e3e660936a4010773675147ff3

Observation 0f5b567f-4e7b-44ea-94b3-516d87f59976 · outbound

This paper cites Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.204396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.853891Z digest=sha256:f8ae8ecb4532457c559b73536c3c92ea6330a207e97099cfd39ad3aeaeb0910e

Observation 16d02971-0acc-4a4c-b962-87c7ff61f642 · outbound

This paper cites VisEvent: Reliable Object Tracking via Collaboration of Frame and Event Flows.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking VisEvent: Reliable Object Tracking via Collaboration of Frame and Event Flows

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.857390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.857390Z digest=sha256:efed651a1bb69817837487c6372659ac337119c032bba1ca5a98fa7f34d311de

Observation c82250ec-9399-40c9-ac99-5907c4d2d57d · outbound

This paper cites Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.192788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.861771Z digest=sha256:345f0262c4253b498c1627c5c554325e950e07cd1d837ad6d8d6a34b19ac0b32

Observation fcbd34df-7175-43d4-ba72-8521d8f12138 · outbound

This paper cites E-motion: Future motion simulation via event sequence diffusion.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking E-motion: Future motion simulation via event sequence diffusion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.179838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.865423Z digest=sha256:b6e0426d621cf862b296a8675c8e40a3b7311a3f46933476c9e1fa5fb5b3fe9d

Observation e43429da-1f73-46ba-a141-007fcc74b46e · outbound

This paper cites Object tracking benchmark.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Object tracking benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.167779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.869292Z digest=sha256:38c0445e368fe932d6f3c07c5e66cb25d93e9df706741993c59144f42b4ef42a

Observation 474d023e-4cb3-4778-bce1-3d90c0db9e8f · outbound

This paper cites Single-model and any-modality for video object tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Single-model and any-modality for video object tracking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.154893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.873290Z digest=sha256:b5159dbf87aed70b1c4c8b09fec56aff20d01445fd55745df897a2355ca3eb74

Observation 2d125a98-848b-4338-952e-4f2c3f4aa086 · outbound

This paper cites Multiple human tracking based on multi-view upper-body detection and discriminative learning.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Multiple human tracking based on multi-view upper-body detection and discriminative learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.140124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.877114Z digest=sha256:d0f8f1301fd71186780467a634c5d2925d81f4434f362e8ac88466e95778b84e

Observation b1c44179-b610-4b3c-988a-7d78f257b2e7 · outbound

This paper cites Learning spatio-temporal transformer for visual tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Learning spatio-temporal transformer for visual tracking

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.127115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.880879Z digest=sha256:d93c08ada91b2e41212d0c99e32003ec3fa6db1c8607da437f81904a3e9a369d

Observation 2afebe87-e411-425d-b9e1-4da6c7381c7c · outbound

This paper cites Depth- track: Unveiling the power of RGBD tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Depth- track: Unveiling the power of RGBD tracking

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.114657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.884703Z digest=sha256:1c05d6c2461a79d054e04751044c8a82aafee1301373ba7d42e44a43e45df07a

Observation eb323a9a-8769-4d4f-81fa-bbf922ec5b0c · outbound

This paper cites Prompting for multi-modal tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Prompting for multi-modal tracking

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.101082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.888531Z digest=sha256:e6de954aae2e9c7a61946cfbaa88e1dd1a6d02fb3031672dfeab0210da0453ca

Observation b13b2ce1-f5f4-4055-93b5-5819dc2d2041 · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Joint feature learning and relation modeling for tracking: A one-stream framework

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.089329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.892596Z digest=sha256:b2441dd9a58d1c0168c4fcccb252c6777823fad8e44a8c9715b5fa5b8fa7132c

Observation e90c5164-1e3e-4642-8441-2d4729a1f1a1 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Adding conditional control to text-to-image diffusion models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.896362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.896362Z digest=sha256:6a0e3155e1ef79a8aaf369d37c9c488beedfd4d05203f21e9b7069627283f15a

Observation c70bd387-45d9-42bc-8baa-ff59514fe79d · outbound

This paper cites Paste, Inpaint and Harmonize via Denoising: Subject-Driven Image Editing with Pre-Trained Diffusion Model.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Paste, Inpaint and Harmonize via Denoising: Subject-Driven Image Editing with Pre-Trained Diffusion Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.900052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.900052Z digest=sha256:e6afde320e9ab5c2d870ede9dcdcbb99097fcadcaff3fb1a9eb37591d969fae5

Observation 43dfd1b4-8c22-44de-ada1-37721fc376c7 · outbound

This paper cites Diff-tracker: Text-to-image diffusion models are unsupervised trackers.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Diff-tracker: Text-to-image diffusion models are unsupervised trackers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.067828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.904191Z digest=sha256:a137f0afc66bc2c98443c7a98253133c30ac095e289758709ff78e8045df7999

Observation 7b831a14-542d-44f0-9bfc-93cd53ab7521 · outbound

This paper cites Unleashing text-to-image diffusion models for visual perception.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Unleashing text-to-image diffusion models for visual perception

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.054664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.908175Z digest=sha256:6e8b4cb035b90accf80ebc238dba2831ddedeca76218d3787d956fb14f67b88b

Observation 0b8efab1-7067-49d6-930a-7556cc8aef71 · outbound

This paper cites Joint visual grounding and tracking with natural language specification.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Joint visual grounding and tracking with natural language specification

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.041125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.911798Z digest=sha256:af24aaff726dde5d5b90fe594b71b5b29780e5c84d9f4ca8da68301be2e09fd0

Observation fde0bfcc-c433-4c65-a027-5eba4f70193c · outbound

This paper cites Visual prompt multi-modal tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Visual prompt multi-modal tracking

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.027591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.915541Z digest=sha256:11d425a8296ca7bb31b3815c58ed567ee276cd31f055a37c8a9f14ee70fd7bad

Observation 6b2969ce-dac8-4cc4-ae6b-85040f083fb5 · outbound

This paper cites RGBD1K: A large-scale dataset and benchmark for RGB-D object tracking.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking RGBD1K: A large-scale dataset and benchmark for RGB-D object tracking

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.015232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.918920Z digest=sha256:ea4d5af0b2cd9171dff42cb989064ec09ee4ae9fb7e1a1a2e6db2de061d8b525

Observation 3c8ef7b8-09b6-48c3-80aa-c1cb5f9aba70 · outbound

This paper cites Exploring pre-trained text-to-video diffusion models for referring video object segmentation.

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking Exploring pre-trained text-to-video diffusion models for referring video object segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:35:55.001458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:35:54.922566Z digest=sha256:64b01c579984821266cf69b307b064d57000089c679bc580cb2cece48c0ce422

Pith citing papers

Observation ac86f337-d99a-41c3-a8ad-45ee4d7ba08d · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:40.921134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:f79f593e530e7b0e8ec858af2651e14de6b1e3bafd82c9ea75c750043f288b5b