Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:40:20.184805Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 3 inbound Pith citation observations for arXiv:2412.10783.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:40:20.184805Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:34:44.785950Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T22:25:27.572440Z
100 of 111 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation df516353-8b38-448b-bb77-75dc42ad6061 · outbound
Video Diffusion Transformers are In-Context Learners All are worth words: A vit backbone for diffusion models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5721a23-8a25-4756-a88a-252a2cddc35e · outbound
Video Diffusion Transformers are In-Context Learners LatentWarp: Consistent Diffusion Latents for Zero-Shot Video-to-Video Translation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c9f385-7793-4fe8-b89a-38c9cdb48665 · outbound
Video Diffusion Transformers are In-Context Learners Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4214ad67-5436-41d8-b120-98a0bc21505a · outbound
Video Diffusion Transformers are In-Context Learners Align your latents: High-resolution video synthesis with latent diffusion models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0474b52-7601-4ff3-aee3-667722e043fb · outbound
Video Diffusion Transformers are In-Context Learners Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f8e4c9-4dbf-42bb-9635-130797b3cc0e · outbound
Video Diffusion Transformers are In-Context Learners End-to-end object detection with transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a591a4e3-f5d8-4806-afa7-f7ea8ba55713 · outbound
Video Diffusion Transformers are In-Context Learners PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7906ac26-21f7-4c7c-b41c-0397b13a86ab · outbound
Video Diffusion Transformers are In-Context Learners Seine: Short-to-long video diffusion model for generative transition and prediction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7b2f696-fc0d-492e-ac2d-9b93a5891f94 · outbound
Video Diffusion Transformers are In-Context Learners Adversarial Video Generation on Complex Datasets
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf002f81-5853-4e0a-a5c4-622c955c1daf · outbound
Video Diffusion Transformers are In-Context Learners A Survey on In-context Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec81b35b-563f-4ac8-bf3e-50267266e727 · outbound
Video Diffusion Transformers are In-Context Learners An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85092cd7-a4cc-4500-a4c9-766ab1a3568d · outbound
Video Diffusion Transformers are In-Context Learners Scaling rectified flow transform- ers for high-resolution image synthesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44767c6-4ac6-4720-99d5-29fe249c2e36 · outbound
Video Diffusion Transformers are In-Context Learners Taming transformers for high-resolution image synthesis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ae5ba6-8300-4f4d-b6fb-5a9e9469d1dd · outbound
Video Diffusion Transformers are In-Context Learners Stable Audio Open
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4e82a7-08df-4cc9-9209-44d59d2a8579 · outbound
Video Diffusion Transformers are In-Context Learners Motioncharacter: Identity- preserving and motion controllable human video generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aaa000f-d6f8-465c-ac06-a0341f061851 · outbound
Video Diffusion Transformers are In-Context Learners Fast Image Caption Generation with Position Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278fce09-db7f-425b-b2ee-9ad43b492ddd · outbound
Video Diffusion Transformers are In-Context Learners Partially non-autoregressive image captioning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 023e501f-7cd6-4341-bcd1-f7d7e4cc8a23 · outbound
Video Diffusion Transformers are In-Context Learners A-JEPA: Joint-Embedding Predictive Architecture Can Listen
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4156a2-6c0d-48f0-93b0-88d1e907f563 · outbound
Video Diffusion Transformers are In-Context Learners Gradient-free textual inversion
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d87e7c-0302-48c2-8555-8e73365f504a · outbound
Video Diffusion Transformers are In-Context Learners Music Consistency Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ef7702-669e-4c6c-9e80-511acb3ea22c · outbound
Video Diffusion Transformers are In-Context Learners FLUX that Plays Music
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 810582bd-2f29-484f-88d4-2820193e5a48 · outbound
Video Diffusion Transformers are In-Context Learners Scalable Diffusion Models with State Space Backbone
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30639231-2dc9-415e-9a3a-90e553dcfe31 · outbound
Video Diffusion Transformers are In-Context Learners Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44c58d3-54e9-41c7-b26a-b0acff7cb46d · outbound
Video Diffusion Transformers are In-Context Learners Scaling Diffusion Transformers to 16 Billion Parameters
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e50a9c7-566d-4ff8-be67-05236ea380f2 · outbound
Video Diffusion Transformers are In-Context Learners Dimba: Transformer-Mamba Diffusion Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef893ee3-7034-47d7-9950-38b2cb1a957a · outbound
Video Diffusion Transformers are In-Context Learners Progressive Text-to-Image Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81edbeb9-d745-4065-9fab-eb8cb75be443 · outbound
Video Diffusion Transformers are In-Context Learners Masked auto-encoders meet generative adversarial networks and beyond
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef1f1ffd-9425-4b4e-b090-ce610b0dc933 · outbound
Video Diffusion Transformers are In-Context Learners Ingredients: Blending Custom Photos with Video Diffusion Transformers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7cfabdc-c88b-49c0-8fb5-facd0519b68d · outbound
Video Diffusion Transformers are In-Context Learners Deecap: Dynamic early exiting for efficient image captioning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db81a31-98dc-4f48-8d09-ecce494f230e · outbound
Video Diffusion Transformers are In-Context Learners I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec00dce-c7da-4fcc-b437-011fb662f956 · outbound
Video Diffusion Transformers are In-Context Learners An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c5fffb-27ba-46fc-ba1a-9f59906940ed · outbound
Video Diffusion Transformers are In-Context Learners Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f0497a-402d-4c67-b6b0-49014d72adc0 · outbound
Video Diffusion Transformers are In-Context Learners The unreasonable effectiveness of few-shot learning for machine translation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c3d6d94-b945-458c-b211-343982de1c5e · outbound
Video Diffusion Transformers are In-Context Learners Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f279d4a0-cad4-437a-86ef-763b147b5f75 · outbound
Video Diffusion Transformers are In-Context Learners Vector quantized diffusion model for text-to-image synthesis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20bb7940-050c-409e-8240-2a2f7ab8b0d4 · outbound
Video Diffusion Transformers are In-Context Learners CameraCtrl: Enabling Camera Control for Text-to-Video Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecbc1d73-dac0-4707-9129-ccc726984f2d · outbound
Video Diffusion Transformers are In-Context Learners Masked autoencoders are scalable vision learners
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14db66ce-19a2-4fb2-bcb7-2e18f0f97ad5 · outbound
Video Diffusion Transformers are In-Context Learners Latent Video Diffusion Models for High-Fidelity Long Video Generation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4dca82c-3813-49b1-bff2-8101e3b86387 · outbound
Video Diffusion Transformers are In-Context Learners StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adccf00f-d031-4ffb-b4ff-1a5cf11c78d8 · outbound
Video Diffusion Transformers are In-Context Learners Imagen Video: High Definition Video Generation with Diffusion Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f5c521-5dfe-494f-ba01-f36ee48c0c3d · outbound
Video Diffusion Transformers are In-Context Learners Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28e7520-f271-4732-87c7-b347c667cc14 · outbound
Video Diffusion Transformers are In-Context Learners CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fda646e-00e3-4950-a03a-ae226fb9f2d7 · outbound
Video Diffusion Transformers are In-Context Learners LoRA: Low-Rank Adaptation of Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead37ca4-8608-4ed9-ad51-8bf89d698a0f · outbound
Video Diffusion Transformers are In-Context Learners DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a731d4-7d01-49e0-b1ce-fc465856f0d9 · outbound
Video Diffusion Transformers are In-Context Learners VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d85853-4654-4459-8b91-95cc6c1e34eb · outbound
Video Diffusion Transformers are In-Context Learners In-Context LoRA for Diffusion Transformers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43aa58fa-c530-4a8e-80a5-c2a14999d1d5 · outbound
Video Diffusion Transformers are In-Context Learners The Platonic Representation Hypothesis
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5b00a5-8f49-4cc4-82f5-5b6abb5bce10 · outbound
Video Diffusion Transformers are In-Context Learners Panoptic segmentation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa196681-f38d-486b-acd6-8e477be472a3 · outbound
Video Diffusion Transformers are In-Context Learners VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9fbc2c-9b6f-4bef-aa15-46ec3e60883c · outbound
Video Diffusion Transformers are In-Context Learners Controlnet ++ : Improving conditional controls with efficient consistency feedback
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dae9957-9eb6-4bc2-9a9e-8f7c14f916a1 · outbound
Video Diffusion Transformers are In-Context Learners Few-shot In-context Learning for Knowledge Base Question Answering
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8613954c-0e30-40fc-ba28-3d48f7efcca2 · outbound
Video Diffusion Transformers are In-Context Learners MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ec997f-2674-4277-b120-5ac0cf44457e · outbound
Video Diffusion Transformers are In-Context Learners ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0c2c4f-dc28-4d36-8c38-c357a45b48f6 · outbound
Video Diffusion Transformers are In-Context Learners Swin transformer: Hierarchical vision transformer using shifted windows
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45e4b19-251f-4a57-84fb-06e57d81d9ba · outbound
Video Diffusion Transformers are In-Context Learners Video swin transformer
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfea9a48-193f-4814-a19e-181fdeb859fc · outbound
Video Diffusion Transformers are In-Context Learners VDT: General-purpose Video Diffusion Transformers via Mask Modeling
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa28ded-b2b7-4467-8c55-55f5323cde43 · outbound
Video Diffusion Transformers are In-Context Learners Latte: Latent Diffusion Transformer for Video Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17a4a83-3dd2-426b-ae01-071436f88d75 · outbound
Video Diffusion Transformers are In-Context Learners Adaptive Machine Translation with Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159bcfc3-5ef0-4a03-91fa-56f55ba03fdb · outbound
Video Diffusion Transformers are In-Context Learners T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac03540-4608-4609-b467-84498e3e7339 · outbound
Video Diffusion Transformers are In-Context Learners MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa139526-954e-438c-9817-382508f35773 · outbound
Video Diffusion Transformers are In-Context Learners Scalable diffusion models with transformers
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6926e07a-7644-4a5c-a992-4f5415a6b54a · outbound
Video Diffusion Transformers are In-Context Learners ControlNeXt: Powerful and Efficient Control for Image and Video Generation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b0b877-64f7-48ac-a3f7-b6561bdac91d · outbound
Video Diffusion Transformers are In-Context Learners Movie Gen: A Cast of Media Foundation Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609ca3bd-0352-4dfc-83c0-074e2188de08 · outbound
Video Diffusion Transformers are In-Context Learners In-Context Learning with Iterative Demonstration Selection
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311acede-8374-44c4-842c-75552205808c · outbound
Video Diffusion Transformers are In-Context Learners SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca23c518-fe40-46ce-8004-4c58fe96c55f · outbound
Video Diffusion Transformers are In-Context Learners Learning transferable visual models from natural language supervision, 2021
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0944f881-de7b-41b4-a391-310a3d0c5160 · outbound
Video Diffusion Transformers are In-Context Learners Improving language understanding with unsupervised learning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741fc7cc-1f1a-45d2-9ebe-0872c0d7eec1 · outbound
Video Diffusion Transformers are In-Context Learners Language models are unsupervised multitask learners
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36132a8-1dc3-4d37-928f-a2cb3462717a · outbound
Video Diffusion Transformers are In-Context Learners Unresolved cited work
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8032b96b-9897-4ce2-9f76-830bd4689975 · outbound
Video Diffusion Transformers are In-Context Learners Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81feab7c-dc66-408c-8e1a-1492f0dfc7aa · outbound
Video Diffusion Transformers are In-Context Learners Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0133ca99-2169-4821-8c03-5d296e6bb75d · outbound
Video Diffusion Transformers are In-Context Learners High-resolution image synthesis with latent diffusion models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5ff2e1-061d-4d92-a5f1-c2f454fe2e1f · outbound
Video Diffusion Transformers are In-Context Learners U-net: Convolutional networks for biomedical image segmentation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74fe4163-53e2-4d1f-b64f-ba1bbb053b20 · outbound
Video Diffusion Transformers are In-Context Learners Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27e29002-0c80-45e9-a6e0-d793a5f9f26f · outbound
Video Diffusion Transformers are In-Context Learners Palette: Image-to-image diffusion models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a115041d-4f99-48cb-a834-a4519ba79178 · outbound
Video Diffusion Transformers are In-Context Learners Photorealistic text-to-image diffusion models with deep language understanding
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24974e9-836f-4812-8f36-c9cfbe80b68b · outbound
Video Diffusion Transformers are In-Context Learners Temporal generative adversarial nets with singular value clipping
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a42254-658a-4394-8b08-c02b6e861d05 · outbound
Video Diffusion Transformers are In-Context Learners Denoising Diffusion Implicit Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659e2927-d455-4de2-b46f-adff3b125b3f · outbound
Video Diffusion Transformers are In-Context Learners Segmenter: Transformer for semantic segmentation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad21d50-5ca6-4938-84f5-a634b4f1cfbb · outbound
Video Diffusion Transformers are In-Context Learners Video-Infinity: Distributed Long Video Generation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9dc9846-6cb8-4d28-8322-5cee112e3008 · outbound
Video Diffusion Transformers are In-Context Learners Training data-efficient image transformers & distillation through attention
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36a444ba-594b-44ce-916d-4fd060175df1 · outbound
Video Diffusion Transformers are In-Context Learners Mocogan: Decomposing motion and content for video generation
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6631a51e-40af-4834-8eed-844498cf6879 · outbound
Video Diffusion Transformers are In-Context Learners Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 410c6557-ead3-477a-9328-285a0b3ccb24 · outbound
Video Diffusion Transformers are In-Context Learners Phenaki: Variable length video generation from open domain textual descriptions
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5f63bfb7-c163-4629-9e30-2d95271eeb4f · outbound
Video Diffusion Transformers are In-Context Learners Generating videos with scene dynamics
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b01888a-52b5-423c-bad4-5dec7185c72f · outbound
Video Diffusion Transformers are In-Context Learners Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362d2a9d-f9e3-40f7-bd9a-d3e631fdc7cd · outbound
Video Diffusion Transformers are In-Context Learners Boximator: Generating Rich and Controllable Motions for Video Synthesis
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3899b909-679c-4e44-9f48-c4431af1c89f · outbound
Video Diffusion Transformers are In-Context Learners Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22998e89-8373-4d7b-bb70-434c96bfd573 · outbound
Video Diffusion Transformers are In-Context Learners Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d05f96-8dad-4d96-8b89-c6da3f786633 · outbound
Video Diffusion Transformers are In-Context Learners Pvt v2: Improved baselines with pyramid vision transformer
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a51d953-4eaf-4ae6-bf10-afade32b0df1 · outbound
Video Diffusion Transformers are In-Context Learners MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c6b44e5-50bd-4f0d-88b3-236d4c09d3d1 · outbound
Video Diffusion Transformers are In-Context Learners Draganything: Motion control for anything using entity representation
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d89c89dd-ea31-43b4-86df-35b5cf8f8bba · outbound
Video Diffusion Transformers are In-Context Learners Segformer: Simple and efficient design for semantic segmentation with transformers.Advances in Neural Information Processing Systems, 34:12077–12090, 2021
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9b24e87c-55a5-47fb-9bd0-412225713398 · outbound
Video Diffusion Transformers are In-Context Learners Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3ed8637-5429-4f84-9f6b-0bf4a82bbc12 · outbound
Video Diffusion Transformers are In-Context Learners CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32337659-badf-4393-a3bc-349918561f04 · outbound
Video Diffusion Transformers are In-Context Learners VideoGPT: Video Generation using VQ-VAE and Transformers
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd6c1b9-fff3-4f99-a5b6-406c5005b8df · outbound
Video Diffusion Transformers are In-Context Learners Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 819e559c-5e17-4fce-87d7-bdc50efe1857 · outbound
Video Diffusion Transformers are In-Context Learners Rerender a video: Zero-shot text-guided video-to-video translation
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9df9fd9-270a-4cdc-a4c5-d1460d3fcbe5 · outbound
Video Diffusion Transformers are In-Context Learners CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f151cbcd-bb8c-44b9-925f-dedd05f3d2a7 · outbound
Video Diffusion Transformers are In-Context Learners Space-time diffusion features for zero-shot text-driven motion transfer
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b195bde-8c54-47c7-97c0-eb2a18f64425 · inbound
Ingredients: Blending Custom Photos with Video Diffusion Transformers Video Diffusion Transformers are In-Context Learners
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ed0d0cab-40ff-4c6f-ab86-bf5d9103ce55 · inbound
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer Video Diffusion Transformers are In-Context Learners
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea42b9c8-20c1-4395-a209-de4a99b83469 · inbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Video Diffusion Transformers are In-Context Learners
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.