Pith. sign in

Paper Citation Record · LEDGER

Emergent Temporal Correspondences from Video Diffusion Transformers

As of 17 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 5 inbound Pith citation observations for arXiv:2506.17220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17220 v2

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:13:40.206348Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:51:02.404616Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:28:38.483958Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d667174f-0174-4d1c-8067-ee8edf5577a8 · outbound

This paper cites GPT-4 Technical Report.

Emergent Temporal Correspondences from Video Diffusion Transformers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.587882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.587882Z digest=sha256:b7f017ea21c02a2378272472560c8a6bfdcfc879ed9535eaa4fd48f4799898b4

Observation 9aa4e928-019a-4e95-b902-deff9ee5bf74 · outbound

This paper cites Self-rectifying diffusion sampling with perturbed-attention guidance.

Emergent Temporal Correspondences from Video Diffusion Transformers Self-rectifying diffusion sampling with perturbed-attention guidance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.618161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.618161Z digest=sha256:9bea6e5a5d33bd10807160f7bdebc600e7798dcb737ad507abe0fcef218d0b9d

Observation 720025fd-2594-439a-b772-c81b13b132b4 · outbound

This paper cites Cross-View Completion Models are Zero-shot Correspondence Estimators.

Emergent Temporal Correspondences from Video Diffusion Transformers Cross-View Completion Models are Zero-shot Correspondence Estimators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.624015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.624015Z digest=sha256:0f2205564905adddbd8d955a20ec3e36781ea128fec2e814a2acc85597a58687

Observation 3157be52-84b8-44d4-8a19-d22a302d83ec · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.630126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.630126Z digest=sha256:0f99b12309b407116828575993e08a25d1633d25a7b1455e095cbf39fbaca929

Observation 11ae0dc1-4dd1-446a-babc-ed4975cb84be · outbound

This paper cites Can Visual Foundation Models Achieve Long-term Point Tracking?.

Emergent Temporal Correspondences from Video Diffusion Transformers Can Visual Foundation Models Achieve Long-term Point Tracking?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.635364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.635364Z digest=sha256:2dc4c5908b9636d774384e52e4e207725a2ada40fe9f298beebff2f3d9dca9ad

Observation 2a0cbfe5-1465-4cfa-a541-fc5924e52bc5 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Emergent Temporal Correspondences from Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.640365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.640365Z digest=sha256:da3856ffebf797205188b5ddd708238c39a6fdaf63cf02380835d843913329f5

Observation ef334008-58c0-411b-a554-718ab610f3d6 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Emergent Temporal Correspondences from Video Diffusion Transformers Align your latents: High-resolution video synthesis with latent diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.691658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.691658Z digest=sha256:09e7cf2dd98ca56927f1f49c4559fbbee16ee1e69825cbcbe22872291660629e

Observation 80b36b94-06f3-443d-8076-d2abc153124a · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.703145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.703145Z digest=sha256:7be48072697593332f1c83cbc80af24ad58558bd31744a1770aeabaf26a9b471

Observation 335da015-e5d6-4308-bdd9-aaed01a63ee9 · outbound

This paper cites Can Generative Video Models Help Pose Estimation?.

Emergent Temporal Correspondences from Video Diffusion Transformers Can Generative Video Models Help Pose Estimation?

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T19:13:41.000631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.707597Z digest=sha256:e1dfa70cfeccfd5f85f695bcbf9b1bb53a8dec1bb26e7987658c1e6233707f00

Observation 372abf9e-fd1f-441e-9c52-77ac313b2de7 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Emergent Temporal Correspondences from Video Diffusion Transformers Emerging properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.713091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.713091Z digest=sha256:9b5db16e0223c7b8b3dbc999f336eeb88fcf62892b3ef842925ae22c705b3fdc

Observation 0fd63fc9-ed9e-40b9-9262-05af6cc1665e · outbound

This paper cites VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models.

Emergent Temporal Correspondences from Video Diffusion Transformers VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.718442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.718442Z digest=sha256:3f2d7e90d4a38aa5e885ae016cd30ac5a57d149f05660d0b273633ec793a739b

Observation d4d86ce4-71c2-4fc2-9461-af681ef6d2de · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.722682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.722682Z digest=sha256:f608793a6fe8c2a9e8322a91818536825dc92a070e8d2aa1f1a1009ed0ed9b2e

Observation bcdcf13c-4af2-49d1-94e9-10333f1a0245 · outbound

This paper cites CATs: Cost aggregation transformers for visual correspondence.NeurIPS, 34:9011–9023, 2021.

Emergent Temporal Correspondences from Video Diffusion Transformers CATs: Cost aggregation transformers for visual correspondence.NeurIPS, 34:9011–9023, 2021

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.959581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.802609Z digest=sha256:32e7c0ee62d3e4ba076c252f37910cc49454a8985e43fb9d48a01cf188208d7d

Observation fe13b774-46d7-4f59-bffd-1e143e5c33d8 · outbound

This paper cites CATs++: Boosting cost aggregation with convolutions and transformers.IEEE TPAMI, 45(6):7174–7194, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers CATs++: Boosting cost aggregation with convolutions and transformers.IEEE TPAMI, 45(6):7174–7194, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.939150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.848794Z digest=sha256:36235a3ede3c82f835e316acdeb9858a2367ae0f85713568d969ca8923ef91b2

Observation f68ddba8-690b-4958-9c7b-a8d7b7341831 · outbound

This paper cites Seurat: From Moving Points to Depth.

Emergent Temporal Correspondences from Video Diffusion Transformers Seurat: From Moving Points to Depth

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.852878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.852878Z digest=sha256:a9ce2e878e0eb74d74c4b3f7ffebc80bfdec7bb07e456498557580053f7be4e0

Observation 7a7c5bfc-ea72-47f3-a3b7-ad22484a8d66 · outbound

This paper cites Local all-pair correspondence for point tracking.

Emergent Temporal Correspondences from Video Diffusion Transformers Local all-pair correspondence for point tracking

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.822039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.858121Z digest=sha256:6f4a91fd25d0b6f1e92584eb047a9b74deff66534e65b579376835955ce1c37d

Observation f3395634-0928-47da-80e6-ccc2e5a1dfaa · outbound

This paper cites Vision Transformers Need Registers.

Emergent Temporal Correspondences from Video Diffusion Transformers Vision Transformers Need Registers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.862232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.862232Z digest=sha256:9ba705fb9753e0c0e6098f619723ebd643fc9431b7b2258b8822a12b1838142d

Observation 3c986a74-a6b8-4b9a-9803-2689d37ccb7a · outbound

This paper cites TAP-Vid: A benchmark for tracking any point in a video.NeurIPS, 35:13610–13626, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers TAP-Vid: A benchmark for tracking any point in a video.NeurIPS, 35:13610–13626, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.695873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.867585Z digest=sha256:2400b5d23bbbc3d6c14c0aa47dd278ac7462bbfe6cafba02dd380b41cddd629c

Observation a82932e7-36cb-41f1-8319-32a377a32d9a · outbound

This paper cites TAPIR: Tracking any point with per-frame initialization and temporal refinement.

Emergent Temporal Correspondences from Video Diffusion Transformers TAPIR: Tracking any point with per-frame initialization and temporal refinement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.611186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.934649Z digest=sha256:a9e99ac8bb8778f6c4aedff507b274f80d66d874f0f2a8b1331f84fef6802653

Observation 3b2003ff-6fbb-4af5-abe0-8df9aa9156a3 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Emergent Temporal Correspondences from Video Diffusion Transformers Scaling rectified flow transformers for high-resolution image synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.961511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.961511Z digest=sha256:a55469c1ba671d66702ecb3fa059f98f891ef6699429727c14f846a6c4b2bca0

Observation 33f5eb86-6c04-40e5-8b71-137af0883265 · outbound

This paper cites Perceptual quality assessment of smartphone photography.

Emergent Temporal Correspondences from Video Diffusion Transformers Perceptual quality assessment of smartphone photography

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.452787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.967025Z digest=sha256:9f1c6c4b548576a37c443c1fba3b7bb16022fa3313c1e44df0676fc66962ec2a

Observation 3e921545-138c-4ded-b054-d57a6a7305d5 · outbound

This paper cites CAT3D: Create Anything in 3D with Multi-View Diffusion Models.

Emergent Temporal Correspondences from Video Diffusion Transformers CAT3D: Create Anything in 3D with Multi-View Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.971715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.971715Z digest=sha256:416d39279f34d7cb14f4aae2cbe9ffcebaa70036bbefbf52334f70707e024ba6

Observation 6066171c-b810-4915-9367-280e54f8c5e5 · outbound

This paper cites Motion Prompting: Controlling Video Generation with Motion Trajectories.

Emergent Temporal Correspondences from Video Diffusion Transformers Motion Prompting: Controlling Video Generation with Motion Trajectories

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.975658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.975658Z digest=sha256:b046b60ab493290c4d788330b24876320688be5c84524c039a38085e2f451d10

Observation 8eb2c98d-a25a-4595-8494-b293fdb143bf · outbound

This paper cites Mochi 1: A new SOTA in open-source video generation models, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers Mochi 1: A new SOTA in open-source video generation models, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.390377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.980746Z digest=sha256:2028e3dacb1b8998ab9e98125396a393e3eef3fa03969f4cde68528e5544a697

Observation 66fcffc3-04b9-4896-a1ae-75e4fe170bab · outbound

This paper cites SparseCtrl: Adding sparse controls to text-to-video diffusion models.

Emergent Temporal Correspondences from Video Diffusion Transformers SparseCtrl: Adding sparse controls to text-to-video diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.328877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:38.985936Z digest=sha256:91dde1e58c3edf1cf8ec4168cd1b276f8d06203e95d3cee9f4b7805393de433c

Observation f442abc9-75f3-4fd7-aab7-3498a74b3651 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Emergent Temporal Correspondences from Video Diffusion Transformers AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.990688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.990688Z digest=sha256:bb023fc099835d75a2870a9e481c15a9151a391ca584042dc22559ccf5cb671a

Observation 2e4c1763-768f-4813-940e-81297155648d · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Emergent Temporal Correspondences from Video Diffusion Transformers LTX-Video: Realtime Video Latent Diffusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.011253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.011253Z digest=sha256:f55f36dd2488114b35e9da61d2f11a2af36b2ce23970150831ce2537e0b56d25

Observation 5a0beaf0-fc71-4364-83ec-60bcfb33eb53 · outbound

This paper cites Harley, Zhaoyuan Fang, and Katerina Fragkiadaki.

Emergent Temporal Correspondences from Video Diffusion Transformers Harley, Zhaoyuan Fang, and Katerina Fragkiadaki

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.289520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.083506Z digest=sha256:92c39966976d1b8c31a53d664fbdf3e029fb6682b44659aca1bc64e5a396db99

Observation 39821cab-3ba0-4fd2-bd7a-cb7c7ce76a6e · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.114343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.114343Z digest=sha256:d68bc2bfdeedc11c071ac00f84acc0650f968b7981a118e9c0cb436a3d183379

Observation fa35a892-0df5-42c6-a7dc-19ae6d5ac23a · outbound

This paper cites Unsupervised semantic correspondence using stable diffusion.NeurIPS, 36:8266–8279, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers Unsupervised semantic correspondence using stable diffusion.NeurIPS, 36:8266–8279, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.273165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.119451Z digest=sha256:12bd5a2816ad5e3110fe22f6cc01c7ef55081a62f8d1842e4f4ffa75b945b18c

Observation f6413a89-f3da-4fe7-9c0a-af2769bdb820 · outbound

This paper cites Denoising diffusion probabilistic models.NeurIPS, 33:6840– 6851, 2020.

Emergent Temporal Correspondences from Video Diffusion Transformers Denoising diffusion probabilistic models.NeurIPS, 33:6840– 6851, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.173878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.124777Z digest=sha256:45a7018863f46ddc97293781db7671f1cfc93b04bfe14993b44f7848bae55b91

Observation a2b514bb-3489-4d12-aec4-a5e0e0e5c9c6 · outbound

This paper cites Integrative Feature and Cost Aggregation with Transformers for Dense Correspondence.

Emergent Temporal Correspondences from Video Diffusion Transformers Integrative Feature and Cost Aggregation with Transformers for Dense Correspondence

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.128864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.128864Z digest=sha256:41be30ddfe4f35ee94ebfe31e573c8c28a90cef3c12cbaa7418d1f732f950877

Observation bc77ebad-7518-4716-baae-0ec19155cd71 · outbound

This paper cites Cost aggregation with 4D convolutional swin transformer for few-shot segmentation.

Emergent Temporal Correspondences from Video Diffusion Transformers Cost aggregation with 4D convolutional swin transformer for few-shot segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.119740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.133584Z digest=sha256:9dea774bb11026e4bebdd67a4ef680a96a0bed48a951fd8b04d68c0392277cec

Observation 2ac08c4b-4c11-499f-b871-f4435aab7cf2 · outbound

This paper cites Unifying correspondence pose and nerf for generalized pose-free novel view synthesis.

Emergent Temporal Correspondences from Video Diffusion Transformers Unifying correspondence pose and nerf for generalized pose-free novel view synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.025900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.138210Z digest=sha256:b8740a0fce10df2fed0310e8e25d2458e2524692d1832cbbc2f5e9f739e43b6a

Observation cc1d3be2-47d3-45a3-9939-c21ebe4ca8cf · outbound

This paper cites Deep matching prior: Test-time optimization for dense correspon- dence.

Emergent Temporal Correspondences from Video Diffusion Transformers Deep matching prior: Test-time optimization for dense correspon- dence

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.008580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.142769Z digest=sha256:85fb91986376a7f970ef63d4013273c0aaca249e27380cf4742b6cd1016d64d9

Observation 811b445e-ace6-4c39-b735-28bb56fe30a3 · outbound

This paper cites Neural matching fields: Implicit representation of matching fields for visual correspondence.NeurIPS, 35:13512–13526, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers Neural matching fields: Implicit representation of matching fields for visual correspondence.NeurIPS, 35:13512–13526, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.825966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.147536Z digest=sha256:9bba4da5333fa487bb6eba38c83615b880e0346b953bb1ab7bc4c2cabe8d3126

Observation e6b06d83-a966-4503-8475-c1bdbed5c6da · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative models.

Emergent Temporal Correspondences from Video Diffusion Transformers VBench: Comprehensive benchmark suite for video generative models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.808548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.197835Z digest=sha256:6bbb8c6c65e178339e998491e56e9a0290938920618472e51bcb6e06ecfe4a67

Observation e511ab6b-bbc3-4577-9c92-8c89d679f741 · outbound

This paper cites Space-time correspondence as a contrastive random walk.

Emergent Temporal Correspondences from Video Diffusion Transformers Space-time correspondence as a contrastive random walk

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.789974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.235267Z digest=sha256:9bea47fca47454738d0f5e72472aa4e732f5dfdc862d5b16cbb8fa2777b19b7b

Observation f1b61264-22c8-40db-9f42-360642c5195a · outbound

This paper cites Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.257714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.257714Z digest=sha256:34e8e3e87784336270385cf5b9474a034beeac5881def5567c23c9635e862d87

Observation 97805d44-ad69-4e2d-9d20-55c463834da5 · outbound

This paper cites Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild.

Emergent Temporal Correspondences from Video Diffusion Transformers Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:13:40.767937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.262818Z digest=sha256:95faec2511e42c790374ac884452befa6140b200572a702b3e5f13dd50a2aaaa

Observation ba5f69c1-c466-4e7d-b68a-8fe2c5615902 · outbound

This paper cites CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos.

Emergent Temporal Correspondences from Video Diffusion Transformers CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.267934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.267934Z digest=sha256:7da5df1d73b9edb2886e9502816edab82b36943e941c19f29b1bb29a7435b9f5

Observation 143386ba-99a7-4f30-8a62-e8a69ddb0e33 · outbound

This paper cites CoTracker: It is better to track together.

Emergent Temporal Correspondences from Video Diffusion Transformers CoTracker: It is better to track together

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.766778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.274246Z digest=sha256:662054fc4f0b28bb5dee703429d071912c3c7be0cb23d1546a86df72451f4fd6

Observation 26c5ddf7-c618-4537-88d2-59aec05fd1b8 · outbound

This paper cites MUSIQ: Multi-scale image quality transformer.

Emergent Temporal Correspondences from Video Diffusion Transformers MUSIQ: Multi-scale image quality transformer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.750973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.278977Z digest=sha256:13f8594f6fb1bf44d55070905020c60877af5e7fedbdded06d938893ae1922db

Observation 81b736a7-1451-4379-9744-d565a8ab7b78 · outbound

This paper cites Exploring Temporally-Aware Features for Point Tracking.

Emergent Temporal Correspondences from Video Diffusion Transformers Exploring Temporally-Aware Features for Point Tracking

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:13:40.678647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.283076Z digest=sha256:41454076d74122875f25344399d85a101ee687d2b505cbc8a92ce33247519f8a

Observation 4058f211-7f87-4e58-bc0b-f97b9417d45a · outbound

This paper cites MoDiTalker: Motion-disentangled diffusion model for high-fidelity talking head generation.

Emergent Temporal Correspondences from Video Diffusion Transformers MoDiTalker: Motion-disentangled diffusion model for high-fidelity talking head generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.663115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.288344Z digest=sha256:b6e2e844878ec73a0421db26298b30741039fde2397514854141e9830195f98a

Observation 4d731154-c6f7-4fd7-a74d-711eefd5279b · outbound

This paper cites Berg, Wan-Yen Lo, et al.

Emergent Temporal Correspondences from Video Diffusion Transformers Berg, Wan-Yen Lo, et al

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.648010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.330619Z digest=sha256:8c0303028d73b5ea616aaf552dbd2a45947f790a4dfa6fc63039e1241cfe86b6

Observation 6a6e781b-c4ef-4a0f-81ae-4189cd7c8c4c · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Emergent Temporal Correspondences from Video Diffusion Transformers HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.411683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.411683Z digest=sha256:63d6e7859986be0fe8cbff953f910b758fa6b16fbb46ac03e79775149eb887ae

Observation 596a9e4f-4bc8-49c1-bb5d-4b7a3741329a · outbound

This paper cites Kling: Video generation by kuaishou, 2024.

Emergent Temporal Correspondences from Video Diffusion Transformers Kling: Video generation by kuaishou, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.573897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.487395Z digest=sha256:6078b325b19c7ffd50dde0386a15c531f23efc615ded7035fdffd0527ba61d9a

Observation 0fbe97fc-80df-44af-ac1b-57d4c8b2c169 · outbound

This paper cites Efficient spatially sparse inference for conditional gans and diffusion models.NeurIPS, 35:28858–28873, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers Efficient spatially sparse inference for conditional gans and diffusion models.NeurIPS, 35:28858–28873, 2022

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.499942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.491851Z digest=sha256:3b2823944e8cbffe2aad413eb35061e40f101f3861a9024073ab5af2de528e26

Observation ccd121ab-a919-48be-b891-2ad6528f30c2 · outbound

This paper cites Spatial-then-temporal self-supervised learning for video correspondence.

Emergent Temporal Correspondences from Video Diffusion Transformers Spatial-then-temporal self-supervised learning for video correspondence

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.443079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.500695Z digest=sha256:6260422e5bb4b921838ee391df21bc6256231b45bb09d8857c005e8b7f9e2358

Observation 90a6b606-ee26-4627-a0b7-53dd9999f42f · outbound

This paper cites ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model.

Emergent Temporal Correspondences from Video Diffusion Transformers ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.506088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.506088Z digest=sha256:15b702a4f11ceec8250e742844ede1934ff0809edf1fc5b2df9038679cf564cf

Observation 7d77039e-0df3-4b17-ab88-82f11ccd6125 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Emergent Temporal Correspondences from Video Diffusion Transformers Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.510944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.510944Z digest=sha256:0741449afa91767c58a4eeebe24fddd420e18051cd933a1410c5c434156e2641

Observation 9990f26c-fd8a-4df9-9c8f-f7930005e5c7 · outbound

This paper cites Not all diffusion model activations have been evaluated as discriminative features.NeurIPS, 37:55141–55177, 2025.

Emergent Temporal Correspondences from Video Diffusion Transformers Not all diffusion model activations have been evaluated as discriminative features.NeurIPS, 37:55141–55177, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.322221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.515483Z digest=sha256:08a832a86d0f8abc1570cea527d7303db1305e3ebfd36a7377a087c7caffc012

Observation bdfb7234-6734-47f1-a680-c4c1a4c41309 · outbound

This paper cites DreamMatcher: appearance matching self-attention for semantically-consistent text-to-image personalization.

Emergent Temporal Correspondences from Video Diffusion Transformers DreamMatcher: appearance matching self-attention for semantically-consistent text-to-image personalization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.305310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.583898Z digest=sha256:fc1cc28a78ae0b16bc7869baa1bd45b84ab787b509249b20a98e7f53101a2d71

Observation e796b727-caed-4871-bbb2-6f58a84a91ed · outbound

This paper cites Diffusion Model for Dense Matching.

Emergent Temporal Correspondences from Video Diffusion Transformers Diffusion Model for Dense Matching

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.669281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.669281Z digest=sha256:827c1eb2a75b3a48b3a995af2df78d643ff26c67400b61c2d72f7e173c1ceb03

Observation 191a8596-d505-4d96-8833-0028ceff3908 · outbound

This paper cites Visual Persona: Foundation Model for Full-Body Human Customization.

Emergent Temporal Correspondences from Video Diffusion Transformers Visual Persona: Foundation Model for Full-Body Human Customization

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:13:40.557927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.674774Z digest=sha256:27fb7af6da7375a4c5230253d15fd901416f35a746ed15af2cc79d2f6c3e2f25

Observation 03198abf-ae6d-4fdc-92e2-e155a9838b2c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Emergent Temporal Correspondences from Video Diffusion Transformers DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.679605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.679605Z digest=sha256:2f75d01391f2640464d71186bd0da3c0599609b4f66c954ba8293b827ab42691

Observation 8a8467e4-cf6f-4b84-ad33-ef018b3f5a5e · outbound

This paper cites Scalable diffusion models with transformers.

Emergent Temporal Correspondences from Video Diffusion Transformers Scalable diffusion models with transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.685451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.685451Z digest=sha256:419005df6a91a2ee958bbc0eca7a332127075c93059bcf6be139d3d528b54902

Observation 84d46051-60bd-412a-8cee-78d597477bd1 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Emergent Temporal Correspondences from Video Diffusion Transformers SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.691810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.691810Z digest=sha256:aa57fcb62da45d2dd6ff72962324e71336f7c615441ac135fcffda640ea0c588

Observation 750dd959-4eea-4cbf-9a44-9decd7879ccf · outbound

This paper cites Movie Gen: A cast of media foundation models, 2025.

Emergent Temporal Correspondences from Video Diffusion Transformers Movie Gen: A cast of media foundation models, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.212793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.699881Z digest=sha256:2e814b9cdb7969e53e31bd48b89767acdf0d7b8c7627abe9c5581185934074eb

Observation 4397c251-e36b-43cf-bd37-2209c9cc56f1 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Emergent Temporal Correspondences from Video Diffusion Transformers The 2017 DAVIS Challenge on Video Object Segmentation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.788381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.788381Z digest=sha256:13ed33fe13f80519c42a3362fdbc49dac0cf7c6da75b3133e05ea42aec242507

Observation b6960061-cde2-4743-af19-2e8370c0d9b4 · outbound

This paper cites Semantics meets temporal correspondence: Self- supervised object-centric learning in videos.

Emergent Temporal Correspondences from Video Diffusion Transformers Semantics meets temporal correspondence: Self- supervised object-centric learning in videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.195229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.817315Z digest=sha256:341eeeab96a745080b749d8ca2d9f2678729937bdec1802615036ea8d054ab80

Observation 594e3340-0fdb-4d73-b40a-217b69f0ee06 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Emergent Temporal Correspondences from Video Diffusion Transformers Learning transferable visual models from natural language supervision

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.823600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.823600Z digest=sha256:3caead43c6871fdde1629b5f382bfd05142c1d72255269c13167eeaf19704cb6

Observation 72d7111e-42c4-46b6-98cf-274ae5656bbe · outbound

This paper cites an unresolved cited work.

Emergent Temporal Correspondences from Video Diffusion Transformers Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:13:42.139594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.828902Z digest=sha256:24469a0eacc172745292d77378b03c2398714e41089a543b32e4a10daf00e6bb

Observation efd69bfd-d978-403d-83cb-c3d94b1630e5 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Emergent Temporal Correspondences from Video Diffusion Transformers High-resolution image synthesis with latent diffusion models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.833861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.833861Z digest=sha256:945d6a89f9774a8bf7383b91913313942e89822e18665420ad3473d144b0d431

Observation 8780f1eb-469a-457e-a5f2-f2445ab3bb97 · outbound

This paper cites Introducing Gen-3 Alpha, 2024.

Emergent Temporal Correspondences from Video Diffusion Transformers Introducing Gen-3 Alpha, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.114305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.839671Z digest=sha256:fda699653ae4cd128770268393540ef74c5ec2812d159e873477bd1378238411

Observation e6912f63-954a-4a6f-a455-71737afcc15e · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Emergent Temporal Correspondences from Video Diffusion Transformers Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.844362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.844362Z digest=sha256:1f6c3b49afa9c742f38942a0f35e7cebbc17419dd952d383323eb7a360db102f

Observation 80d30d88-9235-4a58-9c54-838715681181 · outbound

This paper cites Denoising Diffusion Implicit Models.

Emergent Temporal Correspondences from Video Diffusion Transformers Denoising Diffusion Implicit Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.881217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.881217Z digest=sha256:8187a1db2d74fde78bfcf0d9ade670a5bed3f680148fb9ec8f77435cabff98c9

Observation 0633a4b2-2e23-4b0e-b10c-82f32bd4e20d · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Emergent Temporal Correspondences from Video Diffusion Transformers Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.933730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.933730Z digest=sha256:cff6a870b957c4aa445f255a4f265ff4c975fdf694a5010a1668b6f740c5230b

Observation aa8f1fe9-27d9-4348-a6a6-401f47abb153 · outbound

This paper cites DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion.

Emergent Temporal Correspondences from Video Diffusion Transformers DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.939649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.939649Z digest=sha256:27c80d23167148c0b2c44d8dba98a397297d65f763b8f7fec5b5ca7fbc8513e7

Observation 92a5006e-b65d-4fbd-b4c1-dc643758d51d · outbound

This paper cites Emergent correspondence from image diffusion.NeurIPS, 36:1363–1389, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers Emergent correspondence from image diffusion.NeurIPS, 36:1363–1389, 2023

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.089298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.945033Z digest=sha256:fcfc037ea9ab72000c87ddae475cae510580b805734059a9b21b3771658ce0e7

Observation 35976cb2-6e83-47ba-831d-4df7642ab946 · outbound

This paper cites RAFT: Recurrent all-pairs field transforms for optical flow.

Emergent Temporal Correspondences from Video Diffusion Transformers RAFT: Recurrent all-pairs field transforms for optical flow

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.950837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.949243Z digest=sha256:f6461fac92dbedc5fd8592e05bcb3a8916e478bf92e9a7142f1a1ea7c06fbbfd

Observation 7f3bdeda-b99e-4963-974e-65f05a75398c · outbound

This paper cites GLU-Net: Global-local universal network for dense flow and correspondences.

Emergent Temporal Correspondences from Video Diffusion Transformers GLU-Net: Global-local universal network for dense flow and correspondences

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.890337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:39.953490Z digest=sha256:d01858b1437bf16c671f6a668556c2ca34a8da6596681710620dcb0c333c3166

Observation f088dc25-a18d-4988-b2cc-7efb65d16c2f · outbound

This paper cites Learning accurate dense correspon- dences and when to trust them.

Emergent Temporal Correspondences from Video Diffusion Transformers Learning accurate dense correspon- dences and when to trust them

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.770708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:40.018213Z digest=sha256:6b66002e8ebd68a5f701653aa56ba3d7df25ff772392d7eba5b6dc89cb835617

Observation 16e739dc-0e71-4631-a8ad-9e3b5aad1253 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Emergent Temporal Correspondences from Video Diffusion Transformers Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.621054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:40.045840Z digest=sha256:bb2348718c02104f5d362b8c4f40dce43939710894aa3360c912ca731fbd9ef5

Observation 56f38322-c7d3-47af-ba9d-df34764ed62e · outbound

This paper cites CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models.

Emergent Temporal Correspondences from Video Diffusion Transformers CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.050744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.050744Z digest=sha256:9475ff1969798171a9eacb73beab745b26e741df32cc60ccdf051fd41757ada3

Observation 71215075-cecb-4c25-8509-20f0bfe36853 · outbound

This paper cites Video Diffusion Models are Training-free Motion Interpreter and Controller.

Emergent Temporal Correspondences from Video Diffusion Transformers Video Diffusion Models are Training-free Motion Interpreter and Controller

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.056138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.056138Z digest=sha256:a0dfd359b9b14df542c3d10770cf652afba030d0ec63c7545362913b7bf54974

Observation 434068e2-d2e7-4a0e-966c-968658c06bd0 · outbound

This paper cites DynamiCrafter: Animating open-domain images with video diffusion priors.

Emergent Temporal Correspondences from Video Diffusion Transformers DynamiCrafter: Animating open-domain images with video diffusion priors

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.496368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:40.061077Z digest=sha256:4a87cb4f839accae6761780ee3cb163d98b8a0e03bc7b7b19fb8d8cb4c10ca9a

Observation 5f0bf968-ae86-41b8-81db-55c7770c86f0 · outbound

This paper cites Rethinking self-supervised correspondence learning: A video frame-level similarity perspective.

Emergent Temporal Correspondences from Video Diffusion Transformers Rethinking self-supervised correspondence learning: A video frame-level similarity perspective

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.342982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:40.065360Z digest=sha256:6e40a37b1625e6166223e09804e0e8f769d4d44e7d539871389fa1b32cc98c57

Observation 326fb110-3160-487c-9171-65cf122024e0 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Emergent Temporal Correspondences from Video Diffusion Transformers CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.190379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.190379Z digest=sha256:646e61a0e8174cab869c213f360e809d124660becc82318ae7ca6fbb9ced2959

Observation 5616778c-5335-475e-84ff-dc368a19dc30 · outbound

This paper cites A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence.NeurIPS, 36:45533–45547, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence.NeurIPS, 36:45533–45547, 2023

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.248103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:13:40.196336Z digest=sha256:c7f842e9f4433b330e7f8c2f5d86dc08600bcd04fb8df0f93e1584edcd1f37a7

Observation bab06ec1-95da-4441-aae9-cdc3817c1259 · outbound

This paper cites World-consistent Video Diffusion with Explicit 3D Modeling.

Emergent Temporal Correspondences from Video Diffusion Transformers World-consistent Video Diffusion with Explicit 3D Modeling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.201882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.201882Z digest=sha256:0466656621a3602b59f94f92e01ae37ec82dfe5c4ed4e1498d174a529ff40f49

Observation 2a288842-bd80-40f9-ab84-029626272ace · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Emergent Temporal Correspondences from Video Diffusion Transformers Open-Sora: Democratizing Efficient Video Production for All

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.206348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.206348Z digest=sha256:ae248d0189d3418bac26268c8508fa8d23880ce8c42e816f589599d8293d4e2e

Pith citing papers

Observation 1a61c6d2-5b04-48cc-8f92-960d16369481 · inbound

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking cites this paper.

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.981190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T21:28:12.151547Z digest=sha256:60df8d1c5d49a41c75577ccc43e66b511e45ca859533a6dd2259a8e93275cb7d

Observation 080a97df-cd2c-4677-92d0-e4204c1d6f08 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.652234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:43a619658f39d7c5f45510825c5e6d0758bdc22188b4529d5a96aa1769f063c1

Observation 4aaee8e3-0ee6-43c6-a899-01dbf60d87fa · inbound

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers cites this paper.

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:28:38.485470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-03T16:19:51.348169Z digest=sha256:15fb4847d176d9bc42e8039810b03b2317f51b2740b0cdbdfdc8d19154af9fef

Observation f2c825ce-0016-4433-b2f9-1252f9b3732f · inbound

Controlling Motion Transfer in Diffusion Transformers via Attention Heads cites this paper.

Controlling Motion Transfer in Diffusion Transformers via Attention Heads Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T07:11:45.493916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:11:45.493916Z digest=sha256:db62f1eec128bbc86dfa13e609127d39867caeed93d36e5407ea3344657f009a

Observation 08d76611-dc7a-47f3-b541-efa44cb57232 · inbound

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment cites this paper.

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T16:51:02.404616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:51:02.404616Z digest=sha256:70f597ebb401ea3a15989ae28abb20860a2289f3244495f0c9e9e36a67ffa6de