Pith. sign in

Paper Citation Record · LEDGER

Video Diffusion Transformers are In-Context Learners

As of 19 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 3 inbound Pith citation observations for arXiv:2412.10783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10783 v3

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:40:20.184805Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:34:44.785950Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T22:25:27.572440Z

Reference resolution

100 of 111 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df516353-8b38-448b-bb77-75dc42ad6061 · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

Video Diffusion Transformers are In-Context Learners All are worth words: A vit backbone for diffusion models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.635053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.635053Z digest=sha256:a73adcd6525923d6f8cd79448b02e6d490979f9b81e499d5a2e29e0b3fb78990

Observation b5721a23-8a25-4756-a88a-252a2cddc35e · outbound

This paper cites LatentWarp: Consistent Diffusion Latents for Zero-Shot Video-to-Video Translation.

Video Diffusion Transformers are In-Context Learners LatentWarp: Consistent Diffusion Latents for Zero-Shot Video-to-Video Translation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.641163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.641163Z digest=sha256:573a3ee0799a59f76e201f232114d0379e51730def4b5b9c48225b6f673dfeb7

Observation 39c9f385-7793-4fe8-b89a-38c9cdb48665 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Video Diffusion Transformers are In-Context Learners Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.646865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.646865Z digest=sha256:1f25c02102822e3be248157e52fdf12529dfaaf9727971795d08327a93a91e4e

Observation 4214ad67-5436-41d8-b120-98a0bc21505a · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Video Diffusion Transformers are In-Context Learners Align your latents: High-resolution video synthesis with latent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.652389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.652389Z digest=sha256:0afb3bb66a416285057bd4ae7d1cc4dceaac7ca867a1d520889a4a83f8ed9a75

Observation c0474b52-7601-4ff3-aee3-667722e043fb · outbound

This paper cites Language Models are Few-Shot Learners.

Video Diffusion Transformers are In-Context Learners Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.657686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.657686Z digest=sha256:82d8f68b4f3225098acf7ffedf45dee4210d29e51dcbbadfbe14ab3b1224d7f7

Observation 78f8e4c9-4dbf-42bb-9635-130797b3cc0e · outbound

This paper cites End-to-end object detection with transformers.

Video Diffusion Transformers are In-Context Learners End-to-end object detection with transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.663139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.663139Z digest=sha256:5caf2dd0301f01bfc0ed046b80ee7d5d11867dfffcf6d18d9ad1eae787a0dedd

Observation a591a4e3-f5d8-4806-afa7-f7ea8ba55713 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Video Diffusion Transformers are In-Context Learners PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.669022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.669022Z digest=sha256:0a6ec83e7e0f60446edaf8d5029b4464d50591badd0cac9b856d1cfa2fc25c22

Observation 7906ac26-21f7-4c7c-b41c-0397b13a86ab · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

Video Diffusion Transformers are In-Context Learners Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.674419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.674419Z digest=sha256:04f31a09e02216f72526307067b1bf92e5e64db886cc3f8d6e9f3271d72c23d5

Observation f7b2f696-fc0d-492e-ac2d-9b93a5891f94 · outbound

This paper cites Adversarial Video Generation on Complex Datasets.

Video Diffusion Transformers are In-Context Learners Adversarial Video Generation on Complex Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.679097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.679097Z digest=sha256:67346dee019525dd0510800d4246d7786ffcb36c11faab236e1c1967354d9bf4

Observation bf002f81-5853-4e0a-a5c4-622c955c1daf · outbound

This paper cites A Survey on In-context Learning.

Video Diffusion Transformers are In-Context Learners A Survey on In-context Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.684635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.684635Z digest=sha256:ca234ff9db77c2b6b1fa4eef5bbba1509b36c899b7125d8918703af3188176ca

Observation ec81b35b-563f-4ac8-bf3e-50267266e727 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Video Diffusion Transformers are In-Context Learners An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.689881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.689881Z digest=sha256:4eebc1e5012ead7201af1181e605d5dd1d1b8aec81b70c96cee3afc56bf5d844

Observation 85092cd7-a4cc-4500-a4c9-766ab1a3568d · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

Video Diffusion Transformers are In-Context Learners Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.694948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.694948Z digest=sha256:0afb78dee4b651bc45f718da64b94dbf49f9152790a8950d129563e53b5debf6

Observation d44767c6-4ac6-4720-99d5-29fe249c2e36 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Video Diffusion Transformers are In-Context Learners Taming transformers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.699796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.699796Z digest=sha256:c2a13e1bbe36fee28ac3f6986b4e6eef15a682cb7c1173ec27de19ba43ec0e4a

Observation a0ae5ba6-8300-4f4d-b6fb-5a9e9469d1dd · outbound

This paper cites Stable Audio Open.

Video Diffusion Transformers are In-Context Learners Stable Audio Open

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.704569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.704569Z digest=sha256:630f09904530a92fb2b631ff62ec901503aff97cd4fc989dbdba5dd0b9719de1

Observation 5a4e82a7-08df-4cc9-9209-44d59d2a8579 · outbound

This paper cites Motioncharacter: Identity- preserving and motion controllable human video generation.

Video Diffusion Transformers are In-Context Learners Motioncharacter: Identity- preserving and motion controllable human video generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.709507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.709507Z digest=sha256:adbbde54675b8fdfab441ff4bb8cadfe0b7cf50ec761bcfe74fc46d8773851fe

Observation 2aaa000f-d6f8-465c-ac06-a0341f061851 · outbound

This paper cites Fast Image Caption Generation with Position Alignment.

Video Diffusion Transformers are In-Context Learners Fast Image Caption Generation with Position Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.714326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.714326Z digest=sha256:9a6dadf21a7c31f7a2ebae9b3ffc7324f417561118e30fe65e9667eb527bd0f6

Observation 278fce09-db7f-425b-b2ee-9ad43b492ddd · outbound

This paper cites Partially non-autoregressive image captioning.

Video Diffusion Transformers are In-Context Learners Partially non-autoregressive image captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.719378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.719378Z digest=sha256:50d5d01f204da02200d1892548698ed2c0225ba3c5a3917d894cf0dc6a259304

Observation 023e501f-7cd6-4341-bcd1-f7d7e4cc8a23 · outbound

This paper cites A-JEPA: Joint-Embedding Predictive Architecture Can Listen.

Video Diffusion Transformers are In-Context Learners A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.724388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.724388Z digest=sha256:fb3abc1d5f2541d42a15f91e83701177cf570dd2f55443bf1684c2051bc7c4c5

Observation 1f4156a2-6c0d-48f0-93b0-88d1e907f563 · outbound

This paper cites Gradient-free textual inversion.

Video Diffusion Transformers are In-Context Learners Gradient-free textual inversion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.729569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.729569Z digest=sha256:1bcda6b3b22cdf14328545dc1c4b1f3c401892e8eb252c11f642e7a0b02ff15a

Observation f5d87e7c-0302-48c2-8555-8e73365f504a · outbound

This paper cites Music Consistency Models.

Video Diffusion Transformers are In-Context Learners Music Consistency Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.734960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.734960Z digest=sha256:be52907aa014e8cdc3d9c728f6028c771694ea8d6ca4b15cd8b7e941ac7aa7b1

Observation 94ef7702-669e-4c6c-9e80-511acb3ea22c · outbound

This paper cites FLUX that Plays Music.

Video Diffusion Transformers are In-Context Learners FLUX that Plays Music

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.740228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.740228Z digest=sha256:c0117d9c5b3110db3f830ad308c3307a697937c15adb090e850849fb13589eba

Observation 810582bd-2f29-484f-88d4-2820193e5a48 · outbound

This paper cites Scalable Diffusion Models with State Space Backbone.

Video Diffusion Transformers are In-Context Learners Scalable Diffusion Models with State Space Backbone

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.745469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.745469Z digest=sha256:60cd79957b30c5edac6ef47f43b5b709fa78b65cb527d3a7b3105599bb5a6a85

Observation 30639231-2dc9-415e-9a3a-90e553dcfe31 · outbound

This paper cites Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models.

Video Diffusion Transformers are In-Context Learners Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.750605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.750605Z digest=sha256:f6213d6394cd89ff77a06d66f2fa08eba0599d3e28826391bb44f75c0ba47e91

Observation a44c58d3-54e9-41c7-b26a-b0acff7cb46d · outbound

This paper cites Scaling Diffusion Transformers to 16 Billion Parameters.

Video Diffusion Transformers are In-Context Learners Scaling Diffusion Transformers to 16 Billion Parameters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.755839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.755839Z digest=sha256:475ffb57550f7a2bd6de7b7d4c39d8d8fc81404f115ba8db9bf35f2a0f116376

Observation 5e50a9c7-566d-4ff8-be67-05236ea380f2 · outbound

This paper cites Dimba: Transformer-Mamba Diffusion Models.

Video Diffusion Transformers are In-Context Learners Dimba: Transformer-Mamba Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.761103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.761103Z digest=sha256:9f1fa6d92e83aecd61895d340ac00e1e50625c9efbff75e199270dbd5efec1f7

Observation ef893ee3-7034-47d7-9950-38b2cb1a957a · outbound

This paper cites Progressive Text-to-Image Generation.

Video Diffusion Transformers are In-Context Learners Progressive Text-to-Image Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.766441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.766441Z digest=sha256:78b1794cd6c5721e2400a17746e2ca102a8044a1b25337e9be01280cf6add0e5

Observation 81edbeb9-d745-4065-9fab-eb8cb75be443 · outbound

This paper cites Masked auto-encoders meet generative adversarial networks and beyond.

Video Diffusion Transformers are In-Context Learners Masked auto-encoders meet generative adversarial networks and beyond

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.771771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.771771Z digest=sha256:2a5799e93761611afc798507b4b459a50970fbe9d902f8ccea026672ac9543b9

Observation ef1f1ffd-9425-4b4e-b090-ce610b0dc933 · outbound

This paper cites Ingredients: Blending Custom Photos with Video Diffusion Transformers.

Video Diffusion Transformers are In-Context Learners Ingredients: Blending Custom Photos with Video Diffusion Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.776786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.776786Z digest=sha256:7c3af5676fcdf3bf859a64ee21352965d9ed7352653e1fadb2ba17e3c69f99f2

Observation b7cfabdc-c88b-49c0-8fb5-facd0519b68d · outbound

This paper cites Deecap: Dynamic early exiting for efficient image captioning.

Video Diffusion Transformers are In-Context Learners Deecap: Dynamic early exiting for efficient image captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.782139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.782139Z digest=sha256:2de477f36260fc4afb0f34531146733acf1114f04268751e8d90ecb1d3d63b74

Observation 6db81a31-98dc-4f48-8d09-ecce494f230e · outbound

This paper cites I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength.

Video Diffusion Transformers are In-Context Learners I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.787193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.787193Z digest=sha256:c33bf13e68df7730b4be52c9e4513ac70f5ba53cfc326cade81667f5ed041a57

Observation dec00dce-c7da-4fcc-b437-011fb662f956 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Video Diffusion Transformers are In-Context Learners An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.792563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.792563Z digest=sha256:e087cc65c232ce4d8641c89223774259a064598788b611cffd83f60e6b99e2cd

Observation 60c5fffb-27ba-46fc-ba1a-9f59906940ed · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

Video Diffusion Transformers are In-Context Learners Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.798497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.798497Z digest=sha256:8e757468883c937638512e4eccfb2ad36a8011aad34efd8e89965bf03ee80a27

Observation 97f0497a-402d-4c67-b6b0-49014d72adc0 · outbound

This paper cites The unreasonable effectiveness of few-shot learning for machine translation.

Video Diffusion Transformers are In-Context Learners The unreasonable effectiveness of few-shot learning for machine translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.804073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.804073Z digest=sha256:a9d0c4ebc6c97af5d9f5070576c5797b43650951b090bd43de3ca284707d0d9c

Observation 9c3d6d94-b945-458c-b211-343982de1c5e · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Video Diffusion Transformers are In-Context Learners Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.814698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.814698Z digest=sha256:33bb22e39121f6b1c44df12ae28b63238550a5791e9c487618f9893742769736

Observation f279d4a0-cad4-437a-86ef-763b147b5f75 · outbound

This paper cites Vector quantized diffusion model for text-to-image synthesis.

Video Diffusion Transformers are In-Context Learners Vector quantized diffusion model for text-to-image synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.820483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.820483Z digest=sha256:46c23b51175fbe84c3a9ea15ca2b2d04220c2d44ddde91ac465d317caf1a8d05

Observation 20bb7940-050c-409e-8240-2a2f7ab8b0d4 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Video Diffusion Transformers are In-Context Learners CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.825814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.825814Z digest=sha256:2d415e3908097cfc702f77da4ce1d89baa05f252535113f0ee77dd7160a72a99

Observation ecbc1d73-dac0-4707-9129-ccc726984f2d · outbound

This paper cites Masked autoencoders are scalable vision learners.

Video Diffusion Transformers are In-Context Learners Masked autoencoders are scalable vision learners

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.831153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.831153Z digest=sha256:7213c6cb9f603b5c53e06107935bc57df6d98a7899a674e669be7e06d44da35a

Observation 14db66ce-19a2-4fb2-bcb7-2e18f0f97ad5 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Video Diffusion Transformers are In-Context Learners Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.835947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.835947Z digest=sha256:1956ce1f468aa22e068bd544f076d4b54e819158eb97471a7dc5c54bd66cd640

Observation c4dca82c-3813-49b1-bff2-8101e3b86387 · outbound

This paper cites StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text.

Video Diffusion Transformers are In-Context Learners StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.841363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.841363Z digest=sha256:28f4fb034afd9f8c920855ae9127188a40ce1efabb71b9790809a884c9387a9b

Observation adccf00f-d031-4ffb-b4ff-1a5cf11c78d8 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Video Diffusion Transformers are In-Context Learners Imagen Video: High Definition Video Generation with Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.846704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.846704Z digest=sha256:226ac80de0fa57e44fdc5fac27c098483dd25e3bf65a115b68c899eba5cf24c2

Observation b5f5c521-5dfe-494f-ba01-f36ee48c0c3d · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Video Diffusion Transformers are In-Context Learners Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.852327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.852327Z digest=sha256:79c83758afb494740c9958bfbbe1f8bf8b801e78862441f071f1de94359a5235

Observation f28e7520-f271-4732-87c7-b347c667cc14 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Video Diffusion Transformers are In-Context Learners CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.857283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.857283Z digest=sha256:1f4dd76b068df110613e95958e636a05e8ecc5b72d9bb4feda0c594bd3c454de

Observation 3fda646e-00e3-4950-a03a-ae226fb9f2d7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Video Diffusion Transformers are In-Context Learners LoRA: Low-Rank Adaptation of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.862321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.862321Z digest=sha256:4b8cdf869278a2ecf91687fe89ebbb4d8a9d012ecdfafbe57fad5506c189c0c3

Observation ead37ca4-8608-4ed9-ad51-8bf89d698a0f · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

Video Diffusion Transformers are In-Context Learners DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.867264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.867264Z digest=sha256:5649a0d5bad943b36beefe9ad8debe0e6cab76ebb12c5b98faa36e59fb9a5b62

Observation b8a731d4-7d01-49e0-b1ce-fc465856f0d9 · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

Video Diffusion Transformers are In-Context Learners VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.872352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.872352Z digest=sha256:c35f15e95945d03d5538d195c22cd80407ca855a38b0ed781a2c571187503cff

Observation 24d85853-4654-4459-8b91-95cc6c1e34eb · outbound

This paper cites In-Context LoRA for Diffusion Transformers.

Video Diffusion Transformers are In-Context Learners In-Context LoRA for Diffusion Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.877452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.877452Z digest=sha256:7981640d03ebe47b82976d8766379e8704591b2eba0823ce7bce98800e73d122

Observation 43aa58fa-c530-4a8e-80a5-c2a14999d1d5 · outbound

This paper cites The Platonic Representation Hypothesis.

Video Diffusion Transformers are In-Context Learners The Platonic Representation Hypothesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.882761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.882761Z digest=sha256:fad2ded3c44c06642ec15327292747c2c32b585b125a09c51556cfff04766c12

Observation ec5b00a5-8f49-4cc4-82f5-5b6abb5bce10 · outbound

This paper cites Panoptic segmentation.

Video Diffusion Transformers are In-Context Learners Panoptic segmentation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.887858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.887858Z digest=sha256:aaac1523e497434d43ae19131840afeddd9d7c323ffb0724c099204d68b74829

Observation fa196681-f38d-486b-acd6-8e477be472a3 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Video Diffusion Transformers are In-Context Learners VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.892921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.892921Z digest=sha256:5a0d41b6266a7cd54a7200e6292205d3aa025f57c8a80449d644e529c6baaed3

Observation 6a9fbc2c-9b6f-4bef-aa15-46ec3e60883c · outbound

This paper cites Controlnet ++ : Improving conditional controls with efficient consistency feedback.

Video Diffusion Transformers are In-Context Learners Controlnet ++ : Improving conditional controls with efficient consistency feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.898070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.898070Z digest=sha256:a79aef954380a8616227f3f39adc6ceca45f1b633cd114d621180b030c399d54

Observation 8dae9957-9eb6-4bc2-9a9e-8f7c14f916a1 · outbound

This paper cites Few-shot In-context Learning for Knowledge Base Question Answering.

Video Diffusion Transformers are In-Context Learners Few-shot In-context Learning for Knowledge Base Question Answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.902936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.902936Z digest=sha256:9c1291080c8ecea8e24a50301068903022333c367c8b9f93c25aa48277040dda

Observation 8613954c-0e30-40fc-ba28-3d48f7efcca2 · outbound

This paper cites MarDini: Masked Autoregressive Diffusion for Video Generation at Scale.

Video Diffusion Transformers are In-Context Learners MarDini: Masked Autoregressive Diffusion for Video Generation at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.908050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.908050Z digest=sha256:8a49dea21df73ccbe5ebcad3c6dfbb9ec2a2c93df8d73c324dcac4983a155558

Observation 27ec997f-2674-4277-b120-5ac0cf44457e · outbound

This paper cites ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer.

Video Diffusion Transformers are In-Context Learners ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.913235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.913235Z digest=sha256:1ed147df36279a9635a1b74058841b309cbdd9f0681b0691c19c6915b8162849

Observation 6e0c2c4f-dc28-4d36-8c38-c357a45b48f6 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Video Diffusion Transformers are In-Context Learners Swin transformer: Hierarchical vision transformer using shifted windows

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.918443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.918443Z digest=sha256:b334bc82e95c5654851c20db825cf5c1e024400db39dc996b927558cd049f38c

Observation e45e4b19-251f-4a57-84fb-06e57d81d9ba · outbound

This paper cites Video swin transformer.

Video Diffusion Transformers are In-Context Learners Video swin transformer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.923047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.923047Z digest=sha256:3fb6d3ea1002b6eaf60a47ec13af98efb9a717507d46f48b8d9917c18b35abbc

Observation dfea9a48-193f-4814-a19e-181fdeb859fc · outbound

This paper cites VDT: General-purpose Video Diffusion Transformers via Mask Modeling.

Video Diffusion Transformers are In-Context Learners VDT: General-purpose Video Diffusion Transformers via Mask Modeling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.928134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.928134Z digest=sha256:68c191b6f026eb38826bf4cc1c44b9c9cb923782a5ca092d31b0bb3e092b7960

Observation eaa28ded-b2b7-4467-8c55-55f5323cde43 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Video Diffusion Transformers are In-Context Learners Latte: Latent Diffusion Transformer for Video Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.933825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.933825Z digest=sha256:5b994644995ab2d9671be3de4a8fdfc4eaf0f6a412752850cbc2da682d4bcdab

Observation b17a4a83-3dd2-426b-ae01-071436f88d75 · outbound

This paper cites Adaptive Machine Translation with Large Language Models.

Video Diffusion Transformers are In-Context Learners Adaptive Machine Translation with Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.939173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.939173Z digest=sha256:473eab5f7ae45f8b3a71447bb0384a5b49c281638264508979e6b1b01825ba19

Observation 159bcfc3-5ef0-4a03-91fa-56f55ba03fdb · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Video Diffusion Transformers are In-Context Learners T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.944426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.944426Z digest=sha256:7e7fb16e5a6c32463de17e6615d220ddf36cc18170b5b3b89c90f7661bb996eb

Observation 0ac03540-4608-4609-b467-84498e3e7339 · outbound

This paper cites MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model.

Video Diffusion Transformers are In-Context Learners MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.949699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.949699Z digest=sha256:9b9d7a1038f0a7a39122a529c50dd36ea2beb710e601911420253e721b6649f6

Observation fa139526-954e-438c-9817-382508f35773 · outbound

This paper cites Scalable diffusion models with transformers.

Video Diffusion Transformers are In-Context Learners Scalable diffusion models with transformers

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.959829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.959829Z digest=sha256:2be7c7cf6d4440fed2038a8d3dba14e275d3b29ca3c1f152e0a9122e5a659ddc

Observation 6926e07a-7644-4a5c-a992-4f5415a6b54a · outbound

This paper cites ControlNeXt: Powerful and Efficient Control for Image and Video Generation.

Video Diffusion Transformers are In-Context Learners ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.964749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.964749Z digest=sha256:7dc27db350e53393ae0dba73917168d3a43dc1ca19eb3240b2122189a58116a6

Observation 26b0b877-64f7-48ac-a3f7-b6561bdac91d · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Video Diffusion Transformers are In-Context Learners Movie Gen: A Cast of Media Foundation Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.970117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.970117Z digest=sha256:4cc24ae1b20992ce2dc146940b9ef394f1fd533b7f1c3d4479e7a3c868982fd0

Observation 609ca3bd-0352-4dfc-83c0-074e2188de08 · outbound

This paper cites In-Context Learning with Iterative Demonstration Selection.

Video Diffusion Transformers are In-Context Learners In-Context Learning with Iterative Demonstration Selection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.975980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.975980Z digest=sha256:51d87adc244014ebe90358dfa689a30fd07853431230f969d6ea86cf2ae262f2

Observation 311acede-8374-44c4-842c-75552205808c · outbound

This paper cites SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

Video Diffusion Transformers are In-Context Learners SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.981858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.981858Z digest=sha256:9f2888cc6327241b46bf4faefd7f5dc2f3a82592e0a870488b2b317753965fc8

Observation ca23c518-fe40-46ce-8004-4c58fe96c55f · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Video Diffusion Transformers are In-Context Learners Learning transferable visual models from natural language supervision, 2021

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.987326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.987326Z digest=sha256:52599f336e803686dc340772f339ebdfb6a851816263ef28c24baaa742d1bf6e

Observation 0944f881-de7b-41b4-a391-310a3d0c5160 · outbound

This paper cites Improving language understanding with unsupervised learning.

Video Diffusion Transformers are In-Context Learners Improving language understanding with unsupervised learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.992668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.992668Z digest=sha256:f539a8e5fecbfb8688a9318323f5b1a5a671836f8305e3960ef4761c48226f39

Observation 741fc7cc-1f1a-45d2-9ebe-0872c0d7eec1 · outbound

This paper cites Language models are unsupervised multitask learners.

Video Diffusion Transformers are In-Context Learners Language models are unsupervised multitask learners

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.997771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.997771Z digest=sha256:8ddb15dec4f71ad84111a63c4e1f4a91c1a836c7122ca3707f90cf79cc484edc

Observation e36132a8-1dc3-4d37-928f-a2cb3462717a · outbound

This paper cites an unresolved cited work.

Video Diffusion Transformers are In-Context Learners Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.003196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.003196Z digest=sha256:4239edbaa4a7e313ec3af60d973ed19234c8ff9952438ad5375b4e8f0bababcb

Observation 8032b96b-9897-4ce2-9f76-830bd4689975 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Video Diffusion Transformers are In-Context Learners Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.008387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.008387Z digest=sha256:97325ae1693bf9facf5de0ced5413af909e5ef401fba3f481603102a31f4d17c

Observation 81feab7c-dc66-408c-8e1a-1492f0dfc7aa · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Video Diffusion Transformers are In-Context Learners Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.014339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.014339Z digest=sha256:25d48b8b28779cf6310b824f7a920b4cc39f7970c6073698e11111d8d24c6d82

Observation 0133ca99-2169-4821-8c03-5d296e6bb75d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Video Diffusion Transformers are In-Context Learners High-resolution image synthesis with latent diffusion models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.021755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.021755Z digest=sha256:abcf8751ff94d322ce9897edd3f86e1d363869010b4471f2da9f6c9959b0a3a3

Observation 9d5ff2e1-061d-4d92-a5f1-c2f454fe2e1f · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Video Diffusion Transformers are In-Context Learners U-net: Convolutional networks for biomedical image segmentation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.027407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.027407Z digest=sha256:72fec27bb0642937a1ba08f362c82cf874f60b84776ee541b646e3bfb7d78cc1

Observation 74fe4163-53e2-4d1f-b64f-ba1bbb053b20 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Video Diffusion Transformers are In-Context Learners Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.033296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.033296Z digest=sha256:5f09f093be3b5c2844cd5a20300aa246b7029f5a175f1c1044ada60e1b724a99

Observation 27e29002-0c80-45e9-a6e0-d793a5f9f26f · outbound

This paper cites Palette: Image-to-image diffusion models.

Video Diffusion Transformers are In-Context Learners Palette: Image-to-image diffusion models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.038429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.038429Z digest=sha256:59bedc06de0946d03bf8d0be299e7b418890c859c32112b1146b1373f27eb809

Observation a115041d-4f99-48cb-a834-a4519ba79178 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Video Diffusion Transformers are In-Context Learners Photorealistic text-to-image diffusion models with deep language understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.043662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.043662Z digest=sha256:b5a3644e01f7da1706ac1661e44d6422cf3ecfdb4adb5a0c9c342f7579951fdd

Observation b24974e9-836f-4812-8f36-c9cfbe80b68b · outbound

This paper cites Temporal generative adversarial nets with singular value clipping.

Video Diffusion Transformers are In-Context Learners Temporal generative adversarial nets with singular value clipping

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.049433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.049433Z digest=sha256:c12d8f0bf3caf9ead4fa97055c1adfa46f38512ea2b4b0be406e5a8b1faa2974

Observation b5a42254-658a-4394-8b08-c02b6e861d05 · outbound

This paper cites Denoising Diffusion Implicit Models.

Video Diffusion Transformers are In-Context Learners Denoising Diffusion Implicit Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.055466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.055466Z digest=sha256:48ec8af9c87a1312e493ff6dd78c1b638ef153623cc0f443d3d1f4563a3f0fda

Observation 659e2927-d455-4de2-b46f-adff3b125b3f · outbound

This paper cites Segmenter: Transformer for semantic segmentation.

Video Diffusion Transformers are In-Context Learners Segmenter: Transformer for semantic segmentation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.061905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.061905Z digest=sha256:cf48877f28aab8c8bfa95ed7335b2d9c1fb2ee0a29cb701990a33b688d14a0d7

Observation aad21d50-5ca6-4938-84f5-a634b4f1cfbb · outbound

This paper cites Video-Infinity: Distributed Long Video Generation.

Video Diffusion Transformers are In-Context Learners Video-Infinity: Distributed Long Video Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.068973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.068973Z digest=sha256:92b21fdcfbc9b7e59b0b698732ece32917920d23c18373affa1ff1e91d4bcd57

Observation a9dc9846-6cb8-4d28-8322-5cee112e3008 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Video Diffusion Transformers are In-Context Learners Training data-efficient image transformers & distillation through attention

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.076395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.076395Z digest=sha256:725e0f5dea72994594b2e716b9f6557344f003a0e96decd64f989da82ede6727

Observation 36a444ba-594b-44ce-916d-4fd060175df1 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Video Diffusion Transformers are In-Context Learners Mocogan: Decomposing motion and content for video generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.083720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.083720Z digest=sha256:317154dda04ecfaee32048edea362b0cc0c8670b5c3c4d1881fbb0102f826335

Observation 6631a51e-40af-4834-8eed-844498cf6879 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Video Diffusion Transformers are In-Context Learners Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.090340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.090340Z digest=sha256:9255a777dee6cbe6ff898a1a613dd5df960e3b1a169ad3d79be3d135fd1e6e19

Observation 410c6557-ead3-477a-9328-285a0b3ccb24 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Video Diffusion Transformers are In-Context Learners Phenaki: Variable length video generation from open domain textual descriptions

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:40:21.865559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:40:20.095493Z digest=sha256:3616c12c84ca5402e36e0333bb2e8a016544b0b4614b53c23a276d5056072599

Observation 5f63bfb7-c163-4629-9e30-2d95271eeb4f · outbound

This paper cites Generating videos with scene dynamics.

Video Diffusion Transformers are In-Context Learners Generating videos with scene dynamics

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.100726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.100726Z digest=sha256:f0ead65f4685cd9912b568ae195f7b0d5dd0695bf8f5a99b9a92f876112567ca

Observation 9b01888a-52b5-423c-bad4-5dec7185c72f · outbound

This paper cites Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising.

Video Diffusion Transformers are In-Context Learners Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.106573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.106573Z digest=sha256:0899f99d523fa47c2f425d9d464c67aa206982bad14597b8226f9e3e03f5b5cc

Observation 362d2a9d-f9e3-40f7-bd9a-d3e631fdc7cd · outbound

This paper cites Boximator: Generating Rich and Controllable Motions for Video Synthesis.

Video Diffusion Transformers are In-Context Learners Boximator: Generating Rich and Controllable Motions for Video Synthesis

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.112074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.112074Z digest=sha256:396371a57214bd58857ed2e8922b1e8e5a1eb3a151a79990b32c7dd766b4a2ca

Observation 3899b909-679c-4e44-9f48-c4431af1c89f · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning.

Video Diffusion Transformers are In-Context Learners Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.118393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.118393Z digest=sha256:15d04c26f4450802f219d579fa53242150cebc5c240fb5b21814cc12be459d9b

Observation 22998e89-8373-4d7b-bb70-434c96bfd573 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Video Diffusion Transformers are In-Context Learners Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.123911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.123911Z digest=sha256:497869f4d4da550268b1f3acff65323e824b68e31e4fde24db1d5c0a06a47820

Observation e2d05f96-8dad-4d96-8b89-c6da3f786633 · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

Video Diffusion Transformers are In-Context Learners Pvt v2: Improved baselines with pyramid vision transformer

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.130021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.130021Z digest=sha256:1399a3b0f535e1f6360ef07f59a8278a92ead59a69fea9c146939248fc00c14a

Observation 1a51d953-4eaf-4ae6-bf10-afade32b0df1 · outbound

This paper cites MotionCtrl: A Unified and Flexible Motion Controller for Video Generation.

Video Diffusion Transformers are In-Context Learners MotionCtrl: A Unified and Flexible Motion Controller for Video Generation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.136351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.136351Z digest=sha256:8eda75e437768f96efe34bf7ed2dc4c12752dd5d5353c6a2b76060149ff5522c

Observation 7c6b44e5-50bd-4f0d-88b3-236d4c09d3d1 · outbound

This paper cites Draganything: Motion control for anything using entity representation.

Video Diffusion Transformers are In-Context Learners Draganything: Motion control for anything using entity representation

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:40:21.811418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:40:20.141705Z digest=sha256:c0378b4183f732111293d775c19f4a3e24f275948cdfc2201373d59203f864cf

Observation d89c89dd-ea31-43b4-86df-35b5cf8f8bba · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers.Advances in Neural Information Processing Systems, 34:12077–12090, 2021.

Video Diffusion Transformers are In-Context Learners Segformer: Simple and efficient design for semantic segmentation with transformers.Advances in Neural Information Processing Systems, 34:12077–12090, 2021

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:40:21.790804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T15:40:20.146730Z digest=sha256:e772d172bcdb1906f59dd4a9f430b5e6457a9bab7782ecc34f4826b3dc4725b5

Observation 9b24e87c-55a5-47fb-9bd0-412225713398 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Video Diffusion Transformers are In-Context Learners Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.151919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.151919Z digest=sha256:0740c48936f4b715e5abab3ac569f6e8af2a6ee55f32f63e322ce4ae0d3ee90b

Observation f3ed8637-5429-4f84-9f6b-0bf4a82bbc12 · outbound

This paper cites CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation.

Video Diffusion Transformers are In-Context Learners CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.157664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.157664Z digest=sha256:4042299127a3319030e8e6738957720857d844aaa2faae3d7c3cbada329ec053

Observation 32337659-badf-4393-a3bc-349918561f04 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Video Diffusion Transformers are In-Context Learners VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.163747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.163747Z digest=sha256:d90eb5baf8bb7821195f97712b43e2e72b0f6dddb406176e438f26453adf622f

Observation 9dd6c1b9-fff3-4f99-a5b6-406c5005b8df · outbound

This paper cites Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion.

Video Diffusion Transformers are In-Context Learners Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.169346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.169346Z digest=sha256:b8655a858d0ed66059fbc0a12da5775432f4aef46a1d1d922510449616fb9cb3

Observation 819e559c-5e17-4fce-87d7-bdc50efe1857 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

Video Diffusion Transformers are In-Context Learners Rerender a video: Zero-shot text-guided video-to-video translation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.174556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.174556Z digest=sha256:1634dca3b86e73b9ff5d0a3375d4c94d4bbfc5f8ed87950d4cd75c1967d3b9b3

Observation e9df9fd9-270a-4cdc-a4c5-d1460d3fcbe5 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Video Diffusion Transformers are In-Context Learners CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.179499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.179499Z digest=sha256:5ff1b6f4cf1d089b6ec4635f3494edfb71b26fd535491d855d3e271ef9b978d2

Observation f151cbcd-bb8c-44b9-925f-dedd05f3d2a7 · outbound

This paper cites Space-time diffusion features for zero-shot text-driven motion transfer.

Video Diffusion Transformers are In-Context Learners Space-time diffusion features for zero-shot text-driven motion transfer

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:20.184805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:20.184805Z digest=sha256:67dc4180a5e0668951d91e76d936d38732e17dd322371d2890650f962f79aa00

Pith citing papers

Observation 3b195bde-8c54-47c7-97c0-eb2a18f64425 · inbound

Ingredients: Blending Custom Photos with Video Diffusion Transformers cites this paper.

Ingredients: Blending Custom Photos with Video Diffusion Transformers Video Diffusion Transformers are In-Context Learners

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:25:27.581607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:25:25.580466Z digest=sha256:987cc52f5a72ba0d293ff8d6aeea91006d10691645241ff9c90f96ed74ae2cb4

Observation ed0d0cab-40ff-4c6f-ab86-bf5d9103ce55 · inbound

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer cites this paper.

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer Video Diffusion Transformers are In-Context Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:30:28.157199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:30:28.157199Z digest=sha256:c390721f25614c25d665f7138b31b64165626d9e08293a56554bbe9623582855

Observation ea42b9c8-20c1-4395-a209-de4a99b83469 · inbound

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos cites this paper.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Video Diffusion Transformers are In-Context Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.785950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.785950Z digest=sha256:3f49e5d6771d99e4a8de2c60d8c0cce242d5382d8a9eb0df215468941270bd58