Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:13:40.206348Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 5 inbound Pith citation observations for arXiv:2506.17220.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:13:40.206348Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:51:02.404616Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T16:28:38.483958Z
83 of 83 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d667174f-0174-4d1c-8067-ee8edf5577a8 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aa4e928-019a-4e95-b902-deff9ee5bf74 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Self-rectifying diffusion sampling with perturbed-attention guidance
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720025fd-2594-439a-b772-c81b13b132b4 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Cross-View Completion Models are Zero-shot Correspondence Estimators
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3157be52-84b8-44d4-8a19-d22a302d83ec · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ae0dc1-4dd1-446a-babc-ed4975cb84be · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Can Visual Foundation Models Achieve Long-term Point Tracking?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0cbfe5-1465-4cfa-a541-fc5924e52bc5 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef334008-58c0-411b-a554-718ab610f3d6 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Align your latents: High-resolution video synthesis with latent diffusion models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b36b94-06f3-443d-8076-d2abc153124a · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335da015-e5d6-4308-bdd9-aaed01a63ee9 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Can Generative Video Models Help Pose Estimation?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 372abf9e-fd1f-441e-9c52-77ac313b2de7 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Emerging properties in self-supervised vision transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd63fc9-ed9e-40b9-9262-05af6cc1665e · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4d86ce4-71c2-4fc2-9461-af681ef6d2de · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcdcf13c-4af2-49d1-94e9-10333f1a0245 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers CATs: Cost aggregation transformers for visual correspondence.NeurIPS, 34:9011–9023, 2021
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fe13b774-46d7-4f59-bffd-1e143e5c33d8 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers CATs++: Boosting cost aggregation with convolutions and transformers.IEEE TPAMI, 45(6):7174–7194, 2022
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f68ddba8-690b-4958-9c7b-a8d7b7341831 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Seurat: From Moving Points to Depth
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a7c5bfc-ea72-47f3-a3b7-ad22484a8d66 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Local all-pair correspondence for point tracking
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f3395634-0928-47da-80e6-ccc2e5a1dfaa · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Vision Transformers Need Registers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c986a74-a6b8-4b9a-9803-2689d37ccb7a · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers TAP-Vid: A benchmark for tracking any point in a video.NeurIPS, 35:13610–13626, 2022
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a82932e7-36cb-41f1-8319-32a377a32d9a · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers TAPIR: Tracking any point with per-frame initialization and temporal refinement
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b2003ff-6fbb-4af5-abe0-8df9aa9156a3 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Scaling rectified flow transformers for high-resolution image synthesis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f5eb86-6c04-40e5-8b71-137af0883265 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Perceptual quality assessment of smartphone photography
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3e921545-138c-4ded-b054-d57a6a7305d5 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers CAT3D: Create Anything in 3D with Multi-View Diffusion Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6066171c-b810-4915-9367-280e54f8c5e5 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Motion Prompting: Controlling Video Generation with Motion Trajectories
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eb2c98d-a25a-4595-8494-b293fdb143bf · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Mochi 1: A new SOTA in open-source video generation models, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 66fcffc3-04b9-4896-a1ae-75e4fe170bab · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers SparseCtrl: Adding sparse controls to text-to-video diffusion models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f442abc9-75f3-4fd7-aab7-3498a74b3651 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e4c1763-768f-4813-940e-81297155648d · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers LTX-Video: Realtime Video Latent Diffusion
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a0beaf0-fc71-4364-83ec-60bcfb33eb53 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Harley, Zhaoyuan Fang, and Katerina Fragkiadaki
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 39821cab-3ba0-4fd2-bd7a-cb7c7ce76a6e · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Latent Video Diffusion Models for High-Fidelity Long Video Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa35a892-0df5-42c6-a7dc-19ae6d5ac23a · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Unsupervised semantic correspondence using stable diffusion.NeurIPS, 36:8266–8279, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f6413a89-f3da-4fe7-9c0a-af2769bdb820 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Denoising diffusion probabilistic models.NeurIPS, 33:6840– 6851, 2020
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a2b514bb-3489-4d12-aec4-a5e0e0e5c9c6 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Integrative Feature and Cost Aggregation with Transformers for Dense Correspondence
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc77ebad-7518-4716-baae-0ec19155cd71 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Cost aggregation with 4D convolutional swin transformer for few-shot segmentation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2ac08c4b-4c11-499f-b871-f4435aab7cf2 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Unifying correspondence pose and nerf for generalized pose-free novel view synthesis
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cc1d3be2-47d3-45a3-9939-c21ebe4ca8cf · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Deep matching prior: Test-time optimization for dense correspon- dence
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 811b445e-ace6-4c39-b735-28bb56fe30a3 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Neural matching fields: Implicit representation of matching fields for visual correspondence.NeurIPS, 35:13512–13526, 2022
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6b06d83-a966-4503-8475-c1bdbed5c6da · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers VBench: Comprehensive benchmark suite for video generative models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e511ab6b-bbc3-4577-9c92-8c89d679f741 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Space-time correspondence as a contrastive random walk
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f1b61264-22c8-40db-9f42-360642c5195a · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97805d44-ad69-4e2d-9d20-55c463834da5 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ba5f69c1-c466-4e7d-b68a-8fe2c5615902 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143386ba-99a7-4f30-8a62-e8a69ddb0e33 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers CoTracker: It is better to track together
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 26c5ddf7-c618-4537-88d2-59aec05fd1b8 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers MUSIQ: Multi-scale image quality transformer
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 81b736a7-1451-4379-9744-d565a8ab7b78 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Exploring Temporally-Aware Features for Point Tracking
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4058f211-7f87-4e58-bc0b-f97b9417d45a · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers MoDiTalker: Motion-disentangled diffusion model for high-fidelity talking head generation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4d731154-c6f7-4fd7-a74d-711eefd5279b · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Berg, Wan-Yen Lo, et al
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a6e781b-c4ef-4a0f-81ae-4189cd7c8c4c · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596a9e4f-4bc8-49c1-bb5d-4b7a3741329a · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Kling: Video generation by kuaishou, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0fbe97fc-80df-44af-ac1b-57d4c8b2c169 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Efficient spatially sparse inference for conditional gans and diffusion models.NeurIPS, 35:28858–28873, 2022
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ccd121ab-a919-48be-b891-2ad6528f30c2 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Spatial-then-temporal self-supervised learning for video correspondence
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 90a6b606-ee26-4627-a0b7-53dd9999f42f · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d77039e-0df3-4b17-ab88-82f11ccd6125 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9990f26c-fd8a-4df9-9c8f-f7930005e5c7 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Not all diffusion model activations have been evaluated as discriminative features.NeurIPS, 37:55141–55177, 2025
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bdfb7234-6734-47f1-a680-c4c1a4c41309 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers DreamMatcher: appearance matching self-attention for semantically-consistent text-to-image personalization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e796b727-caed-4871-bbb2-6f58a84a91ed · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Diffusion Model for Dense Matching
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191a8596-d505-4d96-8833-0028ceff3908 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Visual Persona: Foundation Model for Full-Body Human Customization
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 03198abf-ae6d-4fdc-92e2-e155a9838b2c · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers DINOv2: Learning Robust Visual Features without Supervision
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8467e4-cf6f-4b84-ad33-ef018b3f5a5e · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Scalable diffusion models with transformers
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d46051-60bd-412a-8cee-78d597477bd1 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 750dd959-4eea-4cbf-9a44-9decd7879ccf · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Movie Gen: A cast of media foundation models, 2025
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4397c251-e36b-43cf-bd37-2209c9cc56f1 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers The 2017 DAVIS Challenge on Video Object Segmentation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6960061-cde2-4743-af19-2e8370c0d9b4 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Semantics meets temporal correspondence: Self- supervised object-centric learning in videos
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 594e3340-0fdb-4d73-b40a-217b69f0ee06 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Learning transferable visual models from natural language supervision
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72d7111e-42c4-46b6-98cf-274ae5656bbe · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation efd69bfd-d978-403d-83cb-c3d94b1630e5 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers High-resolution image synthesis with latent diffusion models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8780f1eb-469a-457e-a5f2-f2445ab3bb97 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Introducing Gen-3 Alpha, 2024
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6912f63-954a-4a6f-a455-71737afcc15e · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d30d88-9235-4a58-9c54-838715681181 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Denoising Diffusion Implicit Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0633a4b2-2e23-4b0e-b10c-82f32bd4e20d · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8f1fe9-27d9-4348-a6a6-401f47abb153 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a5006e-b65d-4fbd-b4c1-dc643758d51d · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Emergent correspondence from image diffusion.NeurIPS, 36:1363–1389, 2023
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 35976cb2-6e83-47ba-831d-4df7642ab946 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers RAFT: Recurrent all-pairs field transforms for optical flow
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7f3bdeda-b99e-4963-974e-65f05a75398c · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers GLU-Net: Global-local universal network for dense flow and correspondences
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f088dc25-a18d-4988-b2cc-7efb65d16c2f · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Learning accurate dense correspon- dences and when to trust them
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 16e739dc-0e71-4631-a8ad-9e3b5aad1253 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 56f38322-c7d3-47af-ba9d-df34764ed62e · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71215075-cecb-4c25-8509-20f0bfe36853 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Video Diffusion Models are Training-free Motion Interpreter and Controller
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434068e2-d2e7-4a0e-966c-968658c06bd0 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers DynamiCrafter: Animating open-domain images with video diffusion priors
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5f0bf968-ae86-41b8-81db-55c7770c86f0 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Rethinking self-supervised correspondence learning: A video frame-level similarity perspective
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 326fb110-3160-487c-9171-65cf122024e0 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5616778c-5335-475e-84ff-dc368a19dc30 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence.NeurIPS, 36:45533–45547, 2023
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bab06ec1-95da-4441-aae9-cdc3817c1259 · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers World-consistent Video Diffusion with Explicit 3D Modeling
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a288842-bd80-40f9-ab84-029626272ace · outbound
Emergent Temporal Correspondences from Video Diffusion Transformers Open-Sora: Democratizing Efficient Video Production for All
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a61c6d2-5b04-48cc-8f92-960d16369481 · inbound
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking Emergent Temporal Correspondences from Video Diffusion Transformers
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 080a97df-cd2c-4677-92d0-e4204c1d6f08 · inbound
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning Emergent Temporal Correspondences from Video Diffusion Transformers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4aaee8e3-0ee6-43c6-a899-01dbf60d87fa · inbound
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers Emergent Temporal Correspondences from Video Diffusion Transformers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2c825ce-0016-4433-b2f9-1252f9b3732f · inbound
Controlling Motion Transfer in Diffusion Transformers via Attention Heads Emergent Temporal Correspondences from Video Diffusion Transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d76611-dc7a-47f3-b541-efa44cb57232 · inbound
ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment Emergent Temporal Correspondences from Video Diffusion Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.