Pith. sign in

Paper Citation Record · LEDGER

Emergent Temporal Correspondences from Video Diffusion Transformers

As of 17 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 5 inbound Pith citation observations for arXiv:2506.17220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17220 v2

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:13:40.206348Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:51:02.404616Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:28:38.483958Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d667174f-0174-4d1c-8067-ee8edf5577a8 · outbound

This paper cites GPT-4 Technical Report.

Emergent Temporal Correspondences from Video Diffusion Transformers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.587882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.587882Z digest=sha256:819cd9ceabe1394d5925448871fa817521b26a24f35bd30d15e8d6dbdafe1016

Observation 9aa4e928-019a-4e95-b902-deff9ee5bf74 · outbound

This paper cites Self-rectifying diffusion sampling with perturbed-attention guidance.

Emergent Temporal Correspondences from Video Diffusion Transformers Self-rectifying diffusion sampling with perturbed-attention guidance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.618161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.618161Z digest=sha256:9bea6e5a5d33bd10807160f7bdebc600e7798dcb737ad507abe0fcef218d0b9d

Observation 720025fd-2594-439a-b772-c81b13b132b4 · outbound

This paper cites Cross-View Completion Models are Zero-shot Correspondence Estimators.

Emergent Temporal Correspondences from Video Diffusion Transformers Cross-View Completion Models are Zero-shot Correspondence Estimators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.624015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.624015Z digest=sha256:0f2205564905adddbd8d955a20ec3e36781ea128fec2e814a2acc85597a58687

Observation 3157be52-84b8-44d4-8a19-d22a302d83ec · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.630126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.630126Z digest=sha256:0f99b12309b407116828575993e08a25d1633d25a7b1455e095cbf39fbaca929

Observation 11ae0dc1-4dd1-446a-babc-ed4975cb84be · outbound

This paper cites Can Visual Foundation Models Achieve Long-term Point Tracking?.

Emergent Temporal Correspondences from Video Diffusion Transformers Can Visual Foundation Models Achieve Long-term Point Tracking?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.635364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.635364Z digest=sha256:2dc4c5908b9636d774384e52e4e207725a2ada40fe9f298beebff2f3d9dca9ad

Observation 2a0cbfe5-1465-4cfa-a541-fc5924e52bc5 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Emergent Temporal Correspondences from Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.640365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.640365Z digest=sha256:e546051bf4147302dded5007e838da23334f8e6fc76fcc57182cb9f18a1bde5a

Observation ef334008-58c0-411b-a554-718ab610f3d6 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Emergent Temporal Correspondences from Video Diffusion Transformers Align your latents: High-resolution video synthesis with latent diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.691658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.691658Z digest=sha256:09e7cf2dd98ca56927f1f49c4559fbbee16ee1e69825cbcbe22872291660629e

Observation 80b36b94-06f3-443d-8076-d2abc153124a · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.703145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.703145Z digest=sha256:7be48072697593332f1c83cbc80af24ad58558bd31744a1770aeabaf26a9b471

Observation 335da015-e5d6-4308-bdd9-aaed01a63ee9 · outbound

This paper cites Can Generative Video Models Help Pose Estimation?.

Emergent Temporal Correspondences from Video Diffusion Transformers Can Generative Video Models Help Pose Estimation?

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T19:13:41.000631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.707597Z digest=sha256:159a97900cab3f17dbc465f360fb3cf517f1363756b10b90e843f46ab3890700

Observation 372abf9e-fd1f-441e-9c52-77ac313b2de7 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Emergent Temporal Correspondences from Video Diffusion Transformers Emerging properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.713091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.713091Z digest=sha256:9b5db16e0223c7b8b3dbc999f336eeb88fcf62892b3ef842925ae22c705b3fdc

Observation 0fd63fc9-ed9e-40b9-9262-05af6cc1665e · outbound

This paper cites VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models.

Emergent Temporal Correspondences from Video Diffusion Transformers VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.718442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.718442Z digest=sha256:3f2d7e90d4a38aa5e885ae016cd30ac5a57d149f05660d0b273633ec793a739b

Observation d4d86ce4-71c2-4fc2-9461-af681ef6d2de · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.722682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.722682Z digest=sha256:f608793a6fe8c2a9e8322a91818536825dc92a070e8d2aa1f1a1009ed0ed9b2e

Observation bcdcf13c-4af2-49d1-94e9-10333f1a0245 · outbound

This paper cites CATs: Cost aggregation transformers for visual correspondence.NeurIPS, 34:9011–9023, 2021.

Emergent Temporal Correspondences from Video Diffusion Transformers CATs: Cost aggregation transformers for visual correspondence.NeurIPS, 34:9011–9023, 2021

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.959581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.802609Z digest=sha256:3b9d708c701f58caadd6ab1a94d42b6d98caa11f7ab3023612a6edbc8cc03bfe

Observation fe13b774-46d7-4f59-bffd-1e143e5c33d8 · outbound

This paper cites CATs++: Boosting cost aggregation with convolutions and transformers.IEEE TPAMI, 45(6):7174–7194, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers CATs++: Boosting cost aggregation with convolutions and transformers.IEEE TPAMI, 45(6):7174–7194, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.939150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.848794Z digest=sha256:239559d9b5f01477eeb7a0d54e40028324ce333eeed7bdc3534f05014c3f086f

Observation f68ddba8-690b-4958-9c7b-a8d7b7341831 · outbound

This paper cites Seurat: From Moving Points to Depth.

Emergent Temporal Correspondences from Video Diffusion Transformers Seurat: From Moving Points to Depth

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.852878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.852878Z digest=sha256:a9ce2e878e0eb74d74c4b3f7ffebc80bfdec7bb07e456498557580053f7be4e0

Observation 7a7c5bfc-ea72-47f3-a3b7-ad22484a8d66 · outbound

This paper cites Local all-pair correspondence for point tracking.

Emergent Temporal Correspondences from Video Diffusion Transformers Local all-pair correspondence for point tracking

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.822039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.858121Z digest=sha256:60574cc7f5516463b3c9e7db29c59717d6c09311141bdbedb08048024601481d

Observation f3395634-0928-47da-80e6-ccc2e5a1dfaa · outbound

This paper cites Vision Transformers Need Registers.

Emergent Temporal Correspondences from Video Diffusion Transformers Vision Transformers Need Registers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.862232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.862232Z digest=sha256:9ba705fb9753e0c0e6098f619723ebd643fc9431b7b2258b8822a12b1838142d

Observation 3c986a74-a6b8-4b9a-9803-2689d37ccb7a · outbound

This paper cites TAP-Vid: A benchmark for tracking any point in a video.NeurIPS, 35:13610–13626, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers TAP-Vid: A benchmark for tracking any point in a video.NeurIPS, 35:13610–13626, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.695873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.867585Z digest=sha256:7bee26abdfeb5f2c7026e64cd324022afab492971a5b296a9841f780bb84d162

Observation a82932e7-36cb-41f1-8319-32a377a32d9a · outbound

This paper cites TAPIR: Tracking any point with per-frame initialization and temporal refinement.

Emergent Temporal Correspondences from Video Diffusion Transformers TAPIR: Tracking any point with per-frame initialization and temporal refinement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.611186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.934649Z digest=sha256:4a55eb3bb8e7c6d28e0516d3dba2b969991eb511c8f967ee05a9248673ff6f21

Observation 3b2003ff-6fbb-4af5-abe0-8df9aa9156a3 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Emergent Temporal Correspondences from Video Diffusion Transformers Scaling rectified flow transformers for high-resolution image synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.961511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.961511Z digest=sha256:a55469c1ba671d66702ecb3fa059f98f891ef6699429727c14f846a6c4b2bca0

Observation 33f5eb86-6c04-40e5-8b71-137af0883265 · outbound

This paper cites Perceptual quality assessment of smartphone photography.

Emergent Temporal Correspondences from Video Diffusion Transformers Perceptual quality assessment of smartphone photography

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.452787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.967025Z digest=sha256:303545053f8c9605692baa045145f0c9b1a991fc89d52e7dd3e810c2cbea7a7f

Observation 3e921545-138c-4ded-b054-d57a6a7305d5 · outbound

This paper cites CAT3D: Create Anything in 3D with Multi-View Diffusion Models.

Emergent Temporal Correspondences from Video Diffusion Transformers CAT3D: Create Anything in 3D with Multi-View Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.971715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.971715Z digest=sha256:416d39279f34d7cb14f4aae2cbe9ffcebaa70036bbefbf52334f70707e024ba6

Observation 6066171c-b810-4915-9367-280e54f8c5e5 · outbound

This paper cites Motion Prompting: Controlling Video Generation with Motion Trajectories.

Emergent Temporal Correspondences from Video Diffusion Transformers Motion Prompting: Controlling Video Generation with Motion Trajectories

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.975658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.975658Z digest=sha256:b046b60ab493290c4d788330b24876320688be5c84524c039a38085e2f451d10

Observation 8eb2c98d-a25a-4595-8494-b293fdb143bf · outbound

This paper cites Mochi 1: A new SOTA in open-source video generation models, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers Mochi 1: A new SOTA in open-source video generation models, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.390377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.980746Z digest=sha256:720590989f55baea7f88f361b6b07079f98033b73e9718f996ba0cf3847069d9

Observation 66fcffc3-04b9-4896-a1ae-75e4fe170bab · outbound

This paper cites SparseCtrl: Adding sparse controls to text-to-video diffusion models.

Emergent Temporal Correspondences from Video Diffusion Transformers SparseCtrl: Adding sparse controls to text-to-video diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.328877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:38.985936Z digest=sha256:66f07b29a30199aa40aca86acbe97c2fdbf4f344ca0424f2c02df64bf6bdd2a4

Observation f442abc9-75f3-4fd7-aab7-3498a74b3651 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Emergent Temporal Correspondences from Video Diffusion Transformers AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:38.990688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:38.990688Z digest=sha256:aec7cd1e91f9cfb2afad3a8acfa48d99b16fde400d66367faed31f1d4b0f2cf8

Observation 2e4c1763-768f-4813-940e-81297155648d · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Emergent Temporal Correspondences from Video Diffusion Transformers LTX-Video: Realtime Video Latent Diffusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.011253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.011253Z digest=sha256:f55f36dd2488114b35e9da61d2f11a2af36b2ce23970150831ce2537e0b56d25

Observation 5a0beaf0-fc71-4364-83ec-60bcfb33eb53 · outbound

This paper cites Harley, Zhaoyuan Fang, and Katerina Fragkiadaki.

Emergent Temporal Correspondences from Video Diffusion Transformers Harley, Zhaoyuan Fang, and Katerina Fragkiadaki

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.289520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.083506Z digest=sha256:92c5a67925a28004419c42e43c6a673f8dffcb6157172100b645321f1db1c0b4

Observation 39821cab-3ba0-4fd2-bd7a-cb7c7ce76a6e · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.114343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.114343Z digest=sha256:d68bc2bfdeedc11c071ac00f84acc0650f968b7981a118e9c0cb436a3d183379

Observation fa35a892-0df5-42c6-a7dc-19ae6d5ac23a · outbound

This paper cites Unsupervised semantic correspondence using stable diffusion.NeurIPS, 36:8266–8279, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers Unsupervised semantic correspondence using stable diffusion.NeurIPS, 36:8266–8279, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.273165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.119451Z digest=sha256:700483b2a1b29a710b2ec935df81d30b1b98e76523940b0e8a4dddc48e705425

Observation f6413a89-f3da-4fe7-9c0a-af2769bdb820 · outbound

This paper cites Denoising diffusion probabilistic models.NeurIPS, 33:6840– 6851, 2020.

Emergent Temporal Correspondences from Video Diffusion Transformers Denoising diffusion probabilistic models.NeurIPS, 33:6840– 6851, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.173878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.124777Z digest=sha256:78a0e749e033a3e894aab31890884786d867b45e33cfde748cf95f5796559613

Observation a2b514bb-3489-4d12-aec4-a5e0e0e5c9c6 · outbound

This paper cites Integrative Feature and Cost Aggregation with Transformers for Dense Correspondence.

Emergent Temporal Correspondences from Video Diffusion Transformers Integrative Feature and Cost Aggregation with Transformers for Dense Correspondence

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.128864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.128864Z digest=sha256:41be30ddfe4f35ee94ebfe31e573c8c28a90cef3c12cbaa7418d1f732f950877

Observation bc77ebad-7518-4716-baae-0ec19155cd71 · outbound

This paper cites Cost aggregation with 4D convolutional swin transformer for few-shot segmentation.

Emergent Temporal Correspondences from Video Diffusion Transformers Cost aggregation with 4D convolutional swin transformer for few-shot segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.119740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.133584Z digest=sha256:70d7c84ca473d15c831febd8577221fe83cebefafe2c9b71684365a23bab6d9f

Observation 2ac08c4b-4c11-499f-b871-f4435aab7cf2 · outbound

This paper cites Unifying correspondence pose and nerf for generalized pose-free novel view synthesis.

Emergent Temporal Correspondences from Video Diffusion Transformers Unifying correspondence pose and nerf for generalized pose-free novel view synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.025900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.138210Z digest=sha256:a22cacc819e6218a7fd1be7aa5d7ac97400c417d5511be5b607a3f455a66214b

Observation cc1d3be2-47d3-45a3-9939-c21ebe4ca8cf · outbound

This paper cites Deep matching prior: Test-time optimization for dense correspon- dence.

Emergent Temporal Correspondences from Video Diffusion Transformers Deep matching prior: Test-time optimization for dense correspon- dence

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:43.008580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.142769Z digest=sha256:d0bed56a4d3e975edbe766fb47a65e42c5229c917b1824308d128e7340f71dbb

Observation 811b445e-ace6-4c39-b735-28bb56fe30a3 · outbound

This paper cites Neural matching fields: Implicit representation of matching fields for visual correspondence.NeurIPS, 35:13512–13526, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers Neural matching fields: Implicit representation of matching fields for visual correspondence.NeurIPS, 35:13512–13526, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.825966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.147536Z digest=sha256:44c5401968c55481e6ea7cdc0c2e252365892f13aab4ae8abfae0de1e12bbcf0

Observation e6b06d83-a966-4503-8475-c1bdbed5c6da · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative models.

Emergent Temporal Correspondences from Video Diffusion Transformers VBench: Comprehensive benchmark suite for video generative models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.808548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.197835Z digest=sha256:f4340ba6e2ba5381d82d8d4c0fc0b61aaa54a24ea29bbf729f237aee2a190119

Observation e511ab6b-bbc3-4577-9c92-8c89d679f741 · outbound

This paper cites Space-time correspondence as a contrastive random walk.

Emergent Temporal Correspondences from Video Diffusion Transformers Space-time correspondence as a contrastive random walk

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.789974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.235267Z digest=sha256:4c89f51e7a164346304d642cf3e70a95a2f71cf2a9911bd1a6d73793b0d60e40

Observation f1b61264-22c8-40db-9f42-360642c5195a · outbound

This paper cites Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation.

Emergent Temporal Correspondences from Video Diffusion Transformers Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.257714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.257714Z digest=sha256:34e8e3e87784336270385cf5b9474a034beeac5881def5567c23c9635e862d87

Observation 97805d44-ad69-4e2d-9d20-55c463834da5 · outbound

This paper cites Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild.

Emergent Temporal Correspondences from Video Diffusion Transformers Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:13:40.767937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.262818Z digest=sha256:b19aaff60301b5f28b5a3f009829ccd4dd8349d32a5518381ca681c4a565078c

Observation ba5f69c1-c466-4e7d-b68a-8fe2c5615902 · outbound

This paper cites CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos.

Emergent Temporal Correspondences from Video Diffusion Transformers CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.267934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.267934Z digest=sha256:7da5df1d73b9edb2886e9502816edab82b36943e941c19f29b1bb29a7435b9f5

Observation 143386ba-99a7-4f30-8a62-e8a69ddb0e33 · outbound

This paper cites CoTracker: It is better to track together.

Emergent Temporal Correspondences from Video Diffusion Transformers CoTracker: It is better to track together

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.766778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.274246Z digest=sha256:39f0d46227b0a5126b780ce4aec9ae8a6d8d4b8ff41c413182f78c271d071b9a

Observation 26c5ddf7-c618-4537-88d2-59aec05fd1b8 · outbound

This paper cites MUSIQ: Multi-scale image quality transformer.

Emergent Temporal Correspondences from Video Diffusion Transformers MUSIQ: Multi-scale image quality transformer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.750973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.278977Z digest=sha256:0a8e598a70fe50b79f58af0c8bc8ce9a96ee4d3439a891ab94396977b5117eca

Observation 81b736a7-1451-4379-9744-d565a8ab7b78 · outbound

This paper cites Exploring Temporally-Aware Features for Point Tracking.

Emergent Temporal Correspondences from Video Diffusion Transformers Exploring Temporally-Aware Features for Point Tracking

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:13:40.678647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.283076Z digest=sha256:654b1c94c504a3cb779ff8c9332c8df4bb67485559192328f20f441a2944ac5b

Observation 4058f211-7f87-4e58-bc0b-f97b9417d45a · outbound

This paper cites MoDiTalker: Motion-disentangled diffusion model for high-fidelity talking head generation.

Emergent Temporal Correspondences from Video Diffusion Transformers MoDiTalker: Motion-disentangled diffusion model for high-fidelity talking head generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.663115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.288344Z digest=sha256:022cc3d2f84497da6a028ec1d681d462fd0b57c1b097ec9ffedfd0017cdb9d81

Observation 4d731154-c6f7-4fd7-a74d-711eefd5279b · outbound

This paper cites Berg, Wan-Yen Lo, et al.

Emergent Temporal Correspondences from Video Diffusion Transformers Berg, Wan-Yen Lo, et al

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.648010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.330619Z digest=sha256:ce8dcde484f8062e6f2779845252d1cda9cdccf6ab5137602f986ee80c8f4d95

Observation 6a6e781b-c4ef-4a0f-81ae-4189cd7c8c4c · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Emergent Temporal Correspondences from Video Diffusion Transformers HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.411683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.411683Z digest=sha256:a447a0ab4d48a3b35fbd7077d2b291547fc6dcbdd9edeb5048187a4a98c0d2de

Observation 596a9e4f-4bc8-49c1-bb5d-4b7a3741329a · outbound

This paper cites Kling: Video generation by kuaishou, 2024.

Emergent Temporal Correspondences from Video Diffusion Transformers Kling: Video generation by kuaishou, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.573897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.487395Z digest=sha256:50770855835143a87e27703038a79d179063e35e873a950983654565fabca8ce

Observation 0fbe97fc-80df-44af-ac1b-57d4c8b2c169 · outbound

This paper cites Efficient spatially sparse inference for conditional gans and diffusion models.NeurIPS, 35:28858–28873, 2022.

Emergent Temporal Correspondences from Video Diffusion Transformers Efficient spatially sparse inference for conditional gans and diffusion models.NeurIPS, 35:28858–28873, 2022

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.499942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.491851Z digest=sha256:6b696512d975dd9cb3170fae454684a9724021dbde91423fb2022fada4b08449

Observation ccd121ab-a919-48be-b891-2ad6528f30c2 · outbound

This paper cites Spatial-then-temporal self-supervised learning for video correspondence.

Emergent Temporal Correspondences from Video Diffusion Transformers Spatial-then-temporal self-supervised learning for video correspondence

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.443079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.500695Z digest=sha256:857ad5e29d4368025c16eaa4ce48bc142f45cafa3d7231178dfae4adec8c96d7

Observation 90a6b606-ee26-4627-a0b7-53dd9999f42f · outbound

This paper cites ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model.

Emergent Temporal Correspondences from Video Diffusion Transformers ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.506088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.506088Z digest=sha256:15b702a4f11ceec8250e742844ede1934ff0809edf1fc5b2df9038679cf564cf

Observation 7d77039e-0df3-4b17-ab88-82f11ccd6125 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Emergent Temporal Correspondences from Video Diffusion Transformers Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.510944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.510944Z digest=sha256:0741449afa91767c58a4eeebe24fddd420e18051cd933a1410c5c434156e2641

Observation 9990f26c-fd8a-4df9-9c8f-f7930005e5c7 · outbound

This paper cites Not all diffusion model activations have been evaluated as discriminative features.NeurIPS, 37:55141–55177, 2025.

Emergent Temporal Correspondences from Video Diffusion Transformers Not all diffusion model activations have been evaluated as discriminative features.NeurIPS, 37:55141–55177, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.322221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.515483Z digest=sha256:27547fe6c9d387c3032c9b69896e9a0f2c6615641fc53ba60ee9aa8fc4563fe2

Observation bdfb7234-6734-47f1-a680-c4c1a4c41309 · outbound

This paper cites DreamMatcher: appearance matching self-attention for semantically-consistent text-to-image personalization.

Emergent Temporal Correspondences from Video Diffusion Transformers DreamMatcher: appearance matching self-attention for semantically-consistent text-to-image personalization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.305310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.583898Z digest=sha256:aa5d6481df44ad6cff050e82471e66c826a1d0e794d3ff39d427c03ce6398597

Observation e796b727-caed-4871-bbb2-6f58a84a91ed · outbound

This paper cites Diffusion Model for Dense Matching.

Emergent Temporal Correspondences from Video Diffusion Transformers Diffusion Model for Dense Matching

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.669281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.669281Z digest=sha256:827c1eb2a75b3a48b3a995af2df78d643ff26c67400b61c2d72f7e173c1ceb03

Observation 191a8596-d505-4d96-8833-0028ceff3908 · outbound

This paper cites Visual Persona: Foundation Model for Full-Body Human Customization.

Emergent Temporal Correspondences from Video Diffusion Transformers Visual Persona: Foundation Model for Full-Body Human Customization

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:13:40.557927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.674774Z digest=sha256:9ff17ffba995a693b05b259f478d23920ec08f580262a0840d0aecb9ac662d6b

Observation 03198abf-ae6d-4fdc-92e2-e155a9838b2c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Emergent Temporal Correspondences from Video Diffusion Transformers DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.679605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.679605Z digest=sha256:1ed807674a63271ce41766a8c19f96819283f59239c4eecbce943c87aec9af0e

Observation 8a8467e4-cf6f-4b84-ad33-ef018b3f5a5e · outbound

This paper cites Scalable diffusion models with transformers.

Emergent Temporal Correspondences from Video Diffusion Transformers Scalable diffusion models with transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.685451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.685451Z digest=sha256:419005df6a91a2ee958bbc0eca7a332127075c93059bcf6be139d3d528b54902

Observation 84d46051-60bd-412a-8cee-78d597477bd1 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Emergent Temporal Correspondences from Video Diffusion Transformers SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.691810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.691810Z digest=sha256:aa57fcb62da45d2dd6ff72962324e71336f7c615441ac135fcffda640ea0c588

Observation 750dd959-4eea-4cbf-9a44-9decd7879ccf · outbound

This paper cites Movie Gen: A cast of media foundation models, 2025.

Emergent Temporal Correspondences from Video Diffusion Transformers Movie Gen: A cast of media foundation models, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.212793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.699881Z digest=sha256:4d7c62e7b9e762600a2f82d72b9e06a77ae17e16ad202ef101241a65d59734d9

Observation 4397c251-e36b-43cf-bd37-2209c9cc56f1 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Emergent Temporal Correspondences from Video Diffusion Transformers The 2017 DAVIS Challenge on Video Object Segmentation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.788381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.788381Z digest=sha256:13ed33fe13f80519c42a3362fdbc49dac0cf7c6da75b3133e05ea42aec242507

Observation b6960061-cde2-4743-af19-2e8370c0d9b4 · outbound

This paper cites Semantics meets temporal correspondence: Self- supervised object-centric learning in videos.

Emergent Temporal Correspondences from Video Diffusion Transformers Semantics meets temporal correspondence: Self- supervised object-centric learning in videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.195229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.817315Z digest=sha256:31ca5184844ff62f39f9d1a62542ef01dced74c76249807ffe2c276659156fb8

Observation 594e3340-0fdb-4d73-b40a-217b69f0ee06 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Emergent Temporal Correspondences from Video Diffusion Transformers Learning transferable visual models from natural language supervision

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.823600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.823600Z digest=sha256:3caead43c6871fdde1629b5f382bfd05142c1d72255269c13167eeaf19704cb6

Observation 72d7111e-42c4-46b6-98cf-274ae5656bbe · outbound

This paper cites an unresolved cited work.

Emergent Temporal Correspondences from Video Diffusion Transformers Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:13:42.139594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.828902Z digest=sha256:ffda81f02f1f98a3283c4b7fd6bcad67e53c0d9fae4abe6a2ee178d98daff445

Observation efd69bfd-d978-403d-83cb-c3d94b1630e5 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Emergent Temporal Correspondences from Video Diffusion Transformers High-resolution image synthesis with latent diffusion models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.833861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.833861Z digest=sha256:945d6a89f9774a8bf7383b91913313942e89822e18665420ad3473d144b0d431

Observation 8780f1eb-469a-457e-a5f2-f2445ab3bb97 · outbound

This paper cites Introducing Gen-3 Alpha, 2024.

Emergent Temporal Correspondences from Video Diffusion Transformers Introducing Gen-3 Alpha, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.114305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.839671Z digest=sha256:e28106ceb53ca9764f9adc58fa398b16656c8d7fcf2614d8cadf83434ac8c08f

Observation e6912f63-954a-4a6f-a455-71737afcc15e · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Emergent Temporal Correspondences from Video Diffusion Transformers Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.844362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.844362Z digest=sha256:1f6c3b49afa9c742f38942a0f35e7cebbc17419dd952d383323eb7a360db102f

Observation 80d30d88-9235-4a58-9c54-838715681181 · outbound

This paper cites Denoising Diffusion Implicit Models.

Emergent Temporal Correspondences from Video Diffusion Transformers Denoising Diffusion Implicit Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.881217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.881217Z digest=sha256:8187a1db2d74fde78bfcf0d9ade670a5bed3f680148fb9ec8f77435cabff98c9

Observation 0633a4b2-2e23-4b0e-b10c-82f32bd4e20d · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Emergent Temporal Correspondences from Video Diffusion Transformers Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.933730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.933730Z digest=sha256:cff6a870b957c4aa445f255a4f265ff4c975fdf694a5010a1668b6f740c5230b

Observation aa8f1fe9-27d9-4348-a6a6-401f47abb153 · outbound

This paper cites DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion.

Emergent Temporal Correspondences from Video Diffusion Transformers DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:39.939649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:39.939649Z digest=sha256:27c80d23167148c0b2c44d8dba98a397297d65f763b8f7fec5b5ca7fbc8513e7

Observation 92a5006e-b65d-4fbd-b4c1-dc643758d51d · outbound

This paper cites Emergent correspondence from image diffusion.NeurIPS, 36:1363–1389, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers Emergent correspondence from image diffusion.NeurIPS, 36:1363–1389, 2023

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:42.089298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.945033Z digest=sha256:9f8b0747dac1b2221426e1e7ea0c4d29608afaa0dd6e15251a0d206c4256a67e

Observation 35976cb2-6e83-47ba-831d-4df7642ab946 · outbound

This paper cites RAFT: Recurrent all-pairs field transforms for optical flow.

Emergent Temporal Correspondences from Video Diffusion Transformers RAFT: Recurrent all-pairs field transforms for optical flow

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.950837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.949243Z digest=sha256:d380993b16990bd6667ea4ed42875a78d9ac336cb7e42bd4de04b1c6e6149ada

Observation 7f3bdeda-b99e-4963-974e-65f05a75398c · outbound

This paper cites GLU-Net: Global-local universal network for dense flow and correspondences.

Emergent Temporal Correspondences from Video Diffusion Transformers GLU-Net: Global-local universal network for dense flow and correspondences

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.890337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:39.953490Z digest=sha256:dfbf2b9319971c00b52903eabe31f13a69791c85f3a8555bfbb7bbe1438f5604

Observation f088dc25-a18d-4988-b2cc-7efb65d16c2f · outbound

This paper cites Learning accurate dense correspon- dences and when to trust them.

Emergent Temporal Correspondences from Video Diffusion Transformers Learning accurate dense correspon- dences and when to trust them

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.770708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:40.018213Z digest=sha256:7f0e134ee07bf2f7d727307f42de10b6e3778821ea247707f3e2b84ec285d6b9

Observation 16e739dc-0e71-4631-a8ad-9e3b5aad1253 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Emergent Temporal Correspondences from Video Diffusion Transformers Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.621054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:40.045840Z digest=sha256:0fa591f562109ca3bee5f3304a9795e7fa2826eb5d4a2e13eae749f1f02b8bca

Observation 56f38322-c7d3-47af-ba9d-df34764ed62e · outbound

This paper cites CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models.

Emergent Temporal Correspondences from Video Diffusion Transformers CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.050744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.050744Z digest=sha256:66de44f2e56fa17855d47b24290306bffbe44d274e68e4b704a8e7e642b5ce80

Observation 71215075-cecb-4c25-8509-20f0bfe36853 · outbound

This paper cites Video Diffusion Models are Training-free Motion Interpreter and Controller.

Emergent Temporal Correspondences from Video Diffusion Transformers Video Diffusion Models are Training-free Motion Interpreter and Controller

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.056138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.056138Z digest=sha256:a0dfd359b9b14df542c3d10770cf652afba030d0ec63c7545362913b7bf54974

Observation 434068e2-d2e7-4a0e-966c-968658c06bd0 · outbound

This paper cites DynamiCrafter: Animating open-domain images with video diffusion priors.

Emergent Temporal Correspondences from Video Diffusion Transformers DynamiCrafter: Animating open-domain images with video diffusion priors

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.496368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:40.061077Z digest=sha256:0101d6c1a7096d45d854bf0bd4f432e04dfb0de1285ee0229459dbd4c76f9923

Observation 5f0bf968-ae86-41b8-81db-55c7770c86f0 · outbound

This paper cites Rethinking self-supervised correspondence learning: A video frame-level similarity perspective.

Emergent Temporal Correspondences from Video Diffusion Transformers Rethinking self-supervised correspondence learning: A video frame-level similarity perspective

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.342982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:40.065360Z digest=sha256:2cd49ea70f3121e9760b86c80e9c941f7409290aa7f2773ca9f46b9dd8b62318

Observation 326fb110-3160-487c-9171-65cf122024e0 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Emergent Temporal Correspondences from Video Diffusion Transformers CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.190379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.190379Z digest=sha256:646e61a0e8174cab869c213f360e809d124660becc82318ae7ca6fbb9ced2959

Observation 5616778c-5335-475e-84ff-dc368a19dc30 · outbound

This paper cites A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence.NeurIPS, 36:45533–45547, 2023.

Emergent Temporal Correspondences from Video Diffusion Transformers A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence.NeurIPS, 36:45533–45547, 2023

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:13:41.248103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:13:40.196336Z digest=sha256:b584f7f55beaa9be5747b67844aa2ffc262d8fb2ca3c2e3fae9ce420233fe6c7

Observation bab06ec1-95da-4441-aae9-cdc3817c1259 · outbound

This paper cites World-consistent Video Diffusion with Explicit 3D Modeling.

Emergent Temporal Correspondences from Video Diffusion Transformers World-consistent Video Diffusion with Explicit 3D Modeling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.201882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.201882Z digest=sha256:0466656621a3602b59f94f92e01ae37ec82dfe5c4ed4e1498d174a529ff40f49

Observation 2a288842-bd80-40f9-ab84-029626272ace · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Emergent Temporal Correspondences from Video Diffusion Transformers Open-Sora: Democratizing Efficient Video Production for All

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.206348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.206348Z digest=sha256:ae248d0189d3418bac26268c8508fa8d23880ce8c42e816f589599d8293d4e2e

Pith citing papers

Observation 1a61c6d2-5b04-48cc-8f92-960d16369481 · inbound

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking cites this paper.

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.981190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T21:28:12.151547Z digest=sha256:ea1de037ef2cf21edce48d70fe39c7c84e277596f03ecf0e434ee23ad322b666

Observation 080a97df-cd2c-4677-92d0-e4204c1d6f08 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.652234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:7e1a4e07403a2d588d5a47d924012cef7d2834acae42a89c6a05761cbcb2366c

Observation 4aaee8e3-0ee6-43c6-a899-01dbf60d87fa · inbound

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers cites this paper.

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:28:38.485470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T16:19:51.348169Z digest=sha256:8baf4a60a7aa3d5c16445c47df8554e9885e0fd84cf9e663cd5371b5d5467b99

Observation f2c825ce-0016-4433-b2f9-1252f9b3732f · inbound

Controlling Motion Transfer in Diffusion Transformers via Attention Heads cites this paper.

Controlling Motion Transfer in Diffusion Transformers via Attention Heads Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T07:11:45.493916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:11:45.493916Z digest=sha256:db62f1eec128bbc86dfa13e609127d39867caeed93d36e5407ea3344657f009a

Observation 08d76611-dc7a-47f3-b541-efa44cb57232 · inbound

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment cites this paper.

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment Emergent Temporal Correspondences from Video Diffusion Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T16:51:02.404616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:51:02.404616Z digest=sha256:70f597ebb401ea3a15989ae28abb20860a2289f3244495f0c9e9e36a67ffa6de