{"work":{"id":"0c67bec8-6388-449f-9600-64c8ea9bc79f","openalex_id":null,"doi":null,"arxiv_id":null,"raw_key":"raw:db4edb184e5b2fd0a15ef77e","title":"Scalable diffusion models with transformers","authors":null,"authors_text":"William Peebles and Saining Xie","year":2023,"venue":null,"abstract":null,"external_url":null,"cited_by_count":null,"metadata_source":"raw_reference","metadata_fetched_at":"2026-07-11T02:27:53.186413+00:00","pith_arxiv_id":null,"created_at":"2026-05-11T02:07:59.344986+00:00","updated_at":"2026-07-11T02:27:53.186413+00:00","title_quality_ok":true,"display_title":"Scalable diffusion models with transformers","render_title":"Scalable diffusion models with transformers"},"hub":{"state":{"work_id":"0c67bec8-6388-449f-9600-64c8ea9bc79f","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":101,"external_cited_by_count":null,"distinct_field_count":10,"first_pith_cited_at":"2025-01-23T18:55:41+00:00","last_pith_cited_at":"2026-07-09T17:58:29+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-19T05:19:56.944607+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"method","n":8},{"context_role":"background","n":6},{"context_role":"baseline","n":3}],"polarity_counts":[{"context_polarity":"use_method","n":8},{"context_polarity":"background","n":6},{"context_polarity":"baseline","n":3}],"runs":{"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-20T21:12:18.389741+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851","work_id":"82ba805b-3e59-43c6-b37f-3aa1940eea68","shared_citers":22},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":22},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":19},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":14},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":13},{"title":"High- resolution image synthesis with latent diffusion models","work_id":"5427867b-47ba-4d43-a415-2912684a2d41","shared_citers":13},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":12},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":11},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":11},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":11},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":11},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":11},{"title":"U-net: Convolutional networks for biomedical image segmentation","work_id":"11e30d44-6e35-4e94-b299-57b524187090","shared_citers":11},{"title":"Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063","work_id":"b5fd43ff-336f-45fd-933c-ffddf3003880","shared_citers":10},{"title":"Diffusion models beat gans on image synthesis","work_id":"632b8c2b-98f4-4211-8cd8-c9bb7c5b4b8b","shared_citers":9},{"title":"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion","work_id":"53e58ef9-7932-4b83-b757-34ac14db3e0f","shared_citers":9},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":8},{"title":"Attention is all you need.Advances in neural information processing systems, 30","work_id":"751efe07-5e91-415c-b3d1-f4734aa26960","shared_citers":8},{"title":"Flux.https://github.com/black-forest-labs/flux","work_id":"ff476bd7-afa4-451d-b45d-54207aeb1545","shared_citers":8},{"title":"MAGI-1: Autoregressive Video Generation at Scale","work_id":"25e8bd3d-e51c-43ae-8126-4ea6ecdb3321","shared_citers":8},{"title":"Vbench: Comprehensive benchmark suite for video generative models","work_id":"2383dbc4-0e5d-40ed-a102-5aac37332eae","shared_citers":8},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":7},{"title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30","work_id":"32c02792-a0f3-4b5e-8609-8e836c289269","shared_citers":7},{"title":"Open-Sora: Democratizing Efficient Video Production for All","work_id":"8b29ba7b-3d84-4281-85b7-9eaf905afd7f","shared_citers":7}],"time_series":[{"n":13,"year":2025},{"n":46,"year":2026}],"dependency_candidates":[{"n":1,"role":"baseline","polarity":"baseline","paper_title":"Lance: Unified Multimodal Modeling by Multi-Task Synergy","primary_cat":"cs.CV","context_text":"44 57.35 58.66 93.07 AnimateDiff-V2 [35] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.22 92.55 CogVideoX [140] 5B 82.75 77.04 96.23 96.52 98.66 96.92 70.97 61.98 62.90 85.23 Kling [50] - 83.39 75.68 98.33 97.60 99.30 99.40 46.94 61.21 65.62 87.24 Open-Sora-2.0 [90] - 82.10 80.14 98.75 98.00 99.40 99.49 20.74 64.33 65.62 94.50 Gen-3 [96] - 84.11 75.17 97.10 96.62 98.61 99.23 60.14 63.34 66.82 87.81 Step-Video-T2V [79] 30B 84.46 71.28 98.05 97.67 99.40 99.08 53.06 61.23 70.63 80.56 HunyuanVideo [121] - 85.07 76.88 97.22 97.60 99.39 99.05 71.94 60.28 67.24 83.48 Wan2.1-T2V [109] 14B 85.59 76.11 97.52 98.09 99.46 98.","citing_arxiv_id":"2605.18678"},{"n":1,"role":"method","polarity":"use_method","paper_title":"RotVLA: Rotational Latent Action for Vision-Language-Action Model","primary_cat":"cs.RO","context_text":"DINOv2 [46] to extract frame-wise visual features. These features are further processed by a spatial-temporal transformer [ 47]. The decoder is implemented with standard transformer [ 48]. The latent action dimension n is set to 16. For RotVLA, the VLM backbone is initialized from InternVL3.5-1B [38], followed by an action expert implemented as a 24-layer Diffusion Transformer (DiT) [49]. The full model contains approximately 1.7B, including 304M in the vision encoder, 752M in the language model, 305M in the action head, and 290M in the latent action model. Training details.We pretrain both LAM and RotVLA for 200k steps, with a batch size of 256. During downstream finetuning, RotVLA is trained with a batch size of 128. For robot action representation","citing_arxiv_id":"2605.13403"},{"n":1,"role":"method","polarity":"use_method","paper_title":"CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives","primary_cat":"cs.CV","context_text":"3); and (iii) distill the resulting full-step causal model into a four-step generator for interactive synthesis (Sec. 3.4). 3.1 Preliminaries Flow-matching video diffusion.We operate in the video V AE latent space, where a clean video latent x0 ∈R F×C×H×W and Gaussian noise ϵ∼ N(0,I) are interpolated as xt = (1−σ t)x0 +σ tϵ under a shifted schedule [ 9]. A DiT [ 33] velocity field vθ(xt, t,c) is trained with the rectified flow-matching loss LFM =E t,x0,ϵ vθ(xt, t,c)−(ϵ−x 0) 2 ,(1) and sampling integratesdx/dt=v θ with a few-step Euler solver. Distribution matching distillation.DMD [ 54, 53] compresses a pretrained teacher into a few-step student Gϕ by minimizing a reverse KL between the student and teacher distributions at every noise","citing_arxiv_id":"2605.12496"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Generating Symmetric Materials using Latent Flow Matching","primary_cat":"cs.LG","context_text":"For both G and O, we can use the empirical training data distribution ˆpdata(G) and ˆpdata(O|G). By definingX :=(A,W,F, ℓ), our model can be written as pθ(G, O,Z,X) = ˆpdata(G) ˆpdata(O|G)p θFM(Z|G, O)p θD(X|Z, G, O),(5) where θ= (θ FM, θD) are the parameters of the flow matching and decoder neural networks. Similar to ADiT, we use a Diffusion Transformer (DiT) [30] as our denoiser network F, where conditional information gets added through adaptive layer norm. Moreover, we use self-conditioning [ 31], in which the denoiser's prediction from the previous timestep is concatenated with the current input, with a 50% dropout probability during training. We condition on the space group via classifier-free guidance [32] to incorporate the conditioning on space group into the latent generative process (see","citing_arxiv_id":"2605.10115"},{"n":1,"role":"method","polarity":"use_method","paper_title":"SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation","primary_cat":"cs.CV","context_text":"at inference time, whereas most existing multi-person video generation methods require dedicated multi-person training data and task-specific fine-tuning. 2.3 Cross-Attention Control in Diffusion Models In diffusion models, cross-attention serves as the key mechanism for conditioning generation on external signals such as text prompts and reference images, applicable across both U-Net [56, 55] and Diffusion Transformer [50] backbones. Cross-attention control has been widely adopted for text- to-image generation [26, 6, 54], image editing [19, 72, 1], layout-guided generation [41, 9], subject- driven synthesis [65, 63, 68], noise optimization [ 17, 47], and identity preservation [ 79, 69, 76]. Building on these works, SocialDirector applies cross-attention masking and reweighting to multi-","citing_arxiv_id":"2605.10079"},{"n":1,"role":"method","polarity":"use_method","paper_title":"NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models","primary_cat":"cs.RO","context_text":"formulation ties prediction and control to the same denoising/flow trajectory: the future latents provide an explicit imagined context, while the action tokens are refined in the same process. 4 The backbone realizes this W AM with a Mixture-of-Transformers (MoT) [10] architecture, integrating a chunk-bidirectional video DiT (Wan 2.2-TI2V-5B [8]) with an Action-Expert DiT [40]. The clean observation latent v0 is pinned at t0 = 0, while predicted latent frames {vtf f }F f=1 and action tokens ata are processed within a joint sequence where all modalities share self-attention blocks but utilize modality-specific feed-forward layers. As established in §3.2, this shared attention serves as the physical layer for information gating: the","citing_arxiv_id":"2605.07794"},{"n":1,"role":"baseline","polarity":"baseline","paper_title":"Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction","primary_cat":"cs.CV","context_text":"ConvNeXt-UNet Backbone for Microscopic Locality and Multi-Scale Structure The diffusion backbone determines how effectively masked-diffusion pretraining captures pathology morphology. For cell-level dense prediction, features must preserve nuclear contours, chromatin texture, thin boundaries, and multi-scale tissue context. We therefore compare DiT [32], Attention U-Net [33], and ConvNeXt-UNet under the same pretraining and evaluation protocol. DiT offers scalable global modeling, but patch tokenization may weaken small structures and intra- patch boundaries. Attention U-Net is naturally suited to dense prediction, yet its conventional blocks may limit scalability. ConvNeXt-UNet combines U-Net-style multi-scale feature reuse with modern","citing_arxiv_id":"2605.08276"},{"n":1,"role":"method","polarity":"use_method","paper_title":"SDFlow: Similarity-Driven Flow Matching for Time Series Generation","primary_cat":"cs.AI","context_text":"Why SDFlow Generalizes.The kernel-smoothed anchor prior is a continuous initializer over low- rank VQ-latent coordinates, not a raw-data generator; Appendix C provides the theoretical perspective, while Section 5.3 reports held-out latent-flow stress tests. 4.4 Architecture and Training Flow Network Architecture.We employ a Diffusion Transformer (DiT) [ 20] adapted for sequential data. The input is the interpolated latent zt ∈R L×dc treated as a sequence of L tokens. Time conditioning uses adaptive layer normalization (AdaLN): sinusoidal embeddings of t are projected to scale and shift parameters that modulate layer statistics. A learnable global token is prepended to aggregate sequence-level information.","citing_arxiv_id":"2605.05736"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Seedance 1.0: Exploring the Boundaries of Video Generation Models","primary_cat":"cs.CV","context_text":"reconstruction by enforcing finer supervision on local textures and detailed structures. Taking into account appearance and motion modeling simultaneously, we apply a hybrid discriminator with an architecture similar to that used in PatchGAN [11]. 2.2 Diffusion Transformer With the visual tokens encoded by VAE and text tokens generated by a text encoder, we employ the transformer as our diffusion backbone [20], where a fine-tuned decoder-only LLM as the text encoder. The visual tokens are then concatenated with textual tokens and fed into the transformer blocks. Decoupled Spatial and Temporal Layers. Considering both training and inference efficiency, we build the diffusion transformer with decoupled spatial and temporal layers, where the spatial layers perform attention","citing_arxiv_id":"2506.09113"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Improving Video Generation with Human Feedback","primary_cat":"cs.CV","context_text":"Moreover, despite being trained on a disjoint dataset (see Fig. 10), it still achieves comparable performance on GenAI-Bench, indicating robust generalization across different eras of T2V models.Ablation studies are provided in Appendix E. 5.2 Video Alignment Training Setting.Our pretrained model pref is an internal, research-purpose video generation model based on Transformer architecture [53], which is trained using rectified flow (see Appendix 7 for details). Following SD3 [15], all alignment experiments fine-tune the Transformer using LoRA [24]. For training-based alignment methods, including SFT, Flow-DPO, and Flow-RWR, we adopt Video- Reward to provide the reward signals. For reward guidance, we employ thelatentreward model to generate rewards.","citing_arxiv_id":"2501.13918"}]},"error":null,"updated_at":"2026-05-20T21:12:18.439168+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-20T21:12:15.923639+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Scalable diffusion models with transformers","claims":[{"claim_text":"44 57.35 58.66 93.07 AnimateDiff-V2 [35] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.22 92.55 CogVideoX [140] 5B 82.75 77.04 96.23 96.52 98.66 96.92 70.97 61.98 62.90 85.23 Kling [50] - 83.39 75.68 98.33 97.60 99.30 99.40 46.94 61.21 65.62 87.24 Open-Sora-2.0 [90] - 82.10 80.14 98.75 98.00 99.40 99.49 20.74 64.33 65.62 94.50 Gen-3 [96] - 84.11 75.17 97.10 96.62 98.61 99.23 60.14 63.34 66.82 87.81 Step-Vi","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The code is available at https://github.com/ywlq/MotionCache. 1 Introduction Video generation models [12, 18, 24, 26, 34, 41, 47] have achieved remarkable success, facilitating applications ranging from autonomous driving [10, 11, 37] and cinematic creation [6, 38] to social media [3]. While architectures have evolved from U-Nets [2, 27, 29] to scalable Diffusion Transformers (DiTs) [25], practical deployment is hindered by the prohibitive costs of iterative denoising. Moreover, the quadratic co","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"DINOv2 [46] to extract frame-wise visual features. These features are further processed by a spatial-temporal transformer [ 47]. The decoder is implemented with standard transformer [ 48]. The latent action dimension n is set to 16. For RotVLA, the VLM backbone is initialized from InternVL3.5-1B [38], followed by an action expert implemented as a 24-layer Diffusion Transformer (DiT) [49]. The full model contains approximately 1.7B, including 304M in the vision encoder, 752M in the language model","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"reconstruction by enforcing finer supervision on local textures and detailed structures. Taking into account appearance and motion modeling simultaneously, we apply a hybrid discriminator with an architecture similar to that used in PatchGAN [11]. 2.2 Diffusion Transformer With the visual tokens encoded by VAE and text tokens generated by a text encoder, we employ the transformer as our diffusion backbone [20], where a fine-tuned decoder-only LLM as the text encoder. The visual tokens are then c","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ConvNeXt-UNet Backbone for Microscopic Locality and Multi-Scale Structure The diffusion backbone determines how effectively masked-diffusion pretraining captures pathology morphology. For cell-level dense prediction, features must preserve nuclear contours, chromatin texture, thin boundaries, and multi-scale tissue context. We therefore compare DiT [32], Attention U-Net [33], and ConvNeXt-UNet under the same pretraining and evaluation protocol. DiT offers scalable global modeling, but patch toke","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"For both G and O, we can use the empirical training data distribution ˆpdata(G) and ˆpdata(O|G). By definingX :=(A,W,F, ℓ), our model can be written as pθ(G, O,Z,X) = ˆpdata(G) ˆpdata(O|G)p θFM(Z|G, O)p θD(X|Z, G, O),(5) where θ= (θ FM, θD) are the parameters of the flow matching and decoder neural networks. Similar to ADiT, we use a Diffusion Transformer (DiT) [30] as our denoiser network F, where conditional information gets added through adaptive layer norm. Moreover, we use self-conditioning","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Scalable diffusion models with transformers because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (8 contexts).","role_counts":[{"n":8,"context_role":"method"},{"n":6,"context_role":"background"},{"n":2,"context_role":"baseline"}]},"error":null,"updated_at":"2026-05-20T21:12:11.738485+00:00"}},"summary":{"title":"Scalable diffusion models with transformers","claims":[{"claim_text":"44 57.35 58.66 93.07 AnimateDiff-V2 [35] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.22 92.55 CogVideoX [140] 5B 82.75 77.04 96.23 96.52 98.66 96.92 70.97 61.98 62.90 85.23 Kling [50] - 83.39 75.68 98.33 97.60 99.30 99.40 46.94 61.21 65.62 87.24 Open-Sora-2.0 [90] - 82.10 80.14 98.75 98.00 99.40 99.49 20.74 64.33 65.62 94.50 Gen-3 [96] - 84.11 75.17 97.10 96.62 98.61 99.23 60.14 63.34 66.82 87.81 Step-Vi","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The code is available at https://github.com/ywlq/MotionCache. 1 Introduction Video generation models [12, 18, 24, 26, 34, 41, 47] have achieved remarkable success, facilitating applications ranging from autonomous driving [10, 11, 37] and cinematic creation [6, 38] to social media [3]. While architectures have evolved from U-Nets [2, 27, 29] to scalable Diffusion Transformers (DiTs) [25], practical deployment is hindered by the prohibitive costs of iterative denoising. Moreover, the quadratic co","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"DINOv2 [46] to extract frame-wise visual features. These features are further processed by a spatial-temporal transformer [ 47]. The decoder is implemented with standard transformer [ 48]. The latent action dimension n is set to 16. For RotVLA, the VLM backbone is initialized from InternVL3.5-1B [38], followed by an action expert implemented as a 24-layer Diffusion Transformer (DiT) [49]. The full model contains approximately 1.7B, including 304M in the vision encoder, 752M in the language model","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"reconstruction by enforcing finer supervision on local textures and detailed structures. Taking into account appearance and motion modeling simultaneously, we apply a hybrid discriminator with an architecture similar to that used in PatchGAN [11]. 2.2 Diffusion Transformer With the visual tokens encoded by VAE and text tokens generated by a text encoder, we employ the transformer as our diffusion backbone [20], where a fine-tuned decoder-only LLM as the text encoder. The visual tokens are then c","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ConvNeXt-UNet Backbone for Microscopic Locality and Multi-Scale Structure The diffusion backbone determines how effectively masked-diffusion pretraining captures pathology morphology. For cell-level dense prediction, features must preserve nuclear contours, chromatin texture, thin boundaries, and multi-scale tissue context. We therefore compare DiT [32], Attention U-Net [33], and ConvNeXt-UNet under the same pretraining and evaluation protocol. DiT offers scalable global modeling, but patch toke","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"For both G and O, we can use the empirical training data distribution ˆpdata(G) and ˆpdata(O|G). By definingX :=(A,W,F, ℓ), our model can be written as pθ(G, O,Z,X) = ˆpdata(G) ˆpdata(O|G)p θFM(Z|G, O)p θD(X|Z, G, O),(5) where θ= (θ FM, θD) are the parameters of the flow matching and decoder neural networks. Similar to ADiT, we use a Diffusion Transformer (DiT) [30] as our denoiser network F, where conditional information gets added through adaptive layer norm. Moreover, we use self-conditioning","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Scalable diffusion models with transformers because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (8 contexts).","role_counts":[{"n":8,"context_role":"method"},{"n":6,"context_role":"background"},{"n":2,"context_role":"baseline"}]},"graph":{"co_cited":[{"title":"Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851","work_id":"82ba805b-3e59-43c6-b37f-3aa1940eea68","shared_citers":22},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":22},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":19},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":14},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":13},{"title":"High- resolution image synthesis with latent diffusion models","work_id":"5427867b-47ba-4d43-a415-2912684a2d41","shared_citers":13},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":12},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":11},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":11},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":11},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":11},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":11},{"title":"U-net: Convolutional networks for biomedical image segmentation","work_id":"11e30d44-6e35-4e94-b299-57b524187090","shared_citers":11},{"title":"Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063","work_id":"b5fd43ff-336f-45fd-933c-ffddf3003880","shared_citers":10},{"title":"Diffusion models beat gans on image synthesis","work_id":"632b8c2b-98f4-4211-8cd8-c9bb7c5b4b8b","shared_citers":9},{"title":"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion","work_id":"53e58ef9-7932-4b83-b757-34ac14db3e0f","shared_citers":9},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":8},{"title":"Attention is all you need.Advances in neural information processing systems, 30","work_id":"751efe07-5e91-415c-b3d1-f4734aa26960","shared_citers":8},{"title":"Flux.https://github.com/black-forest-labs/flux","work_id":"ff476bd7-afa4-451d-b45d-54207aeb1545","shared_citers":8},{"title":"MAGI-1: Autoregressive Video Generation at Scale","work_id":"25e8bd3d-e51c-43ae-8126-4ea6ecdb3321","shared_citers":8},{"title":"Vbench: Comprehensive benchmark suite for video generative models","work_id":"2383dbc4-0e5d-40ed-a102-5aac37332eae","shared_citers":8},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":7},{"title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30","work_id":"32c02792-a0f3-4b5e-8609-8e836c289269","shared_citers":7},{"title":"Open-Sora: Democratizing Efficient Video Production for All","work_id":"8b29ba7b-3d84-4281-85b7-9eaf905afd7f","shared_citers":7}],"time_series":[{"n":13,"year":2025},{"n":46,"year":2026}],"dependency_candidates":[{"n":1,"role":"baseline","polarity":"baseline","paper_title":"Lance: Unified Multimodal Modeling by Multi-Task Synergy","primary_cat":"cs.CV","context_text":"44 57.35 58.66 93.07 AnimateDiff-V2 [35] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.22 92.55 CogVideoX [140] 5B 82.75 77.04 96.23 96.52 98.66 96.92 70.97 61.98 62.90 85.23 Kling [50] - 83.39 75.68 98.33 97.60 99.30 99.40 46.94 61.21 65.62 87.24 Open-Sora-2.0 [90] - 82.10 80.14 98.75 98.00 99.40 99.49 20.74 64.33 65.62 94.50 Gen-3 [96] - 84.11 75.17 97.10 96.62 98.61 99.23 60.14 63.34 66.82 87.81 Step-Video-T2V [79] 30B 84.46 71.28 98.05 97.67 99.40 99.08 53.06 61.23 70.63 80.56 HunyuanVideo [121] - 85.07 76.88 97.22 97.60 99.39 99.05 71.94 60.28 67.24 83.48 Wan2.1-T2V [109] 14B 85.59 76.11 97.52 98.09 99.46 98.","citing_arxiv_id":"2605.18678"},{"n":1,"role":"method","polarity":"use_method","paper_title":"RotVLA: Rotational Latent Action for Vision-Language-Action Model","primary_cat":"cs.RO","context_text":"DINOv2 [46] to extract frame-wise visual features. These features are further processed by a spatial-temporal transformer [ 47]. The decoder is implemented with standard transformer [ 48]. The latent action dimension n is set to 16. For RotVLA, the VLM backbone is initialized from InternVL3.5-1B [38], followed by an action expert implemented as a 24-layer Diffusion Transformer (DiT) [49]. The full model contains approximately 1.7B, including 304M in the vision encoder, 752M in the language model, 305M in the action head, and 290M in the latent action model. Training details.We pretrain both LAM and RotVLA for 200k steps, with a batch size of 256. During downstream finetuning, RotVLA is trained with a batch size of 128. For robot action representation","citing_arxiv_id":"2605.13403"},{"n":1,"role":"method","polarity":"use_method","paper_title":"CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives","primary_cat":"cs.CV","context_text":"3); and (iii) distill the resulting full-step causal model into a four-step generator for interactive synthesis (Sec. 3.4). 3.1 Preliminaries Flow-matching video diffusion.We operate in the video V AE latent space, where a clean video latent x0 ∈R F×C×H×W and Gaussian noise ϵ∼ N(0,I) are interpolated as xt = (1−σ t)x0 +σ tϵ under a shifted schedule [ 9]. A DiT [ 33] velocity field vθ(xt, t,c) is trained with the rectified flow-matching loss LFM =E t,x0,ϵ vθ(xt, t,c)−(ϵ−x 0) 2 ,(1) and sampling integratesdx/dt=v θ with a few-step Euler solver. Distribution matching distillation.DMD [ 54, 53] compresses a pretrained teacher into a few-step student Gϕ by minimizing a reverse KL between the student and teacher distributions at every noise","citing_arxiv_id":"2605.12496"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Generating Symmetric Materials using Latent Flow Matching","primary_cat":"cs.LG","context_text":"For both G and O, we can use the empirical training data distribution ˆpdata(G) and ˆpdata(O|G). By definingX :=(A,W,F, ℓ), our model can be written as pθ(G, O,Z,X) = ˆpdata(G) ˆpdata(O|G)p θFM(Z|G, O)p θD(X|Z, G, O),(5) where θ= (θ FM, θD) are the parameters of the flow matching and decoder neural networks. Similar to ADiT, we use a Diffusion Transformer (DiT) [30] as our denoiser network F, where conditional information gets added through adaptive layer norm. Moreover, we use self-conditioning [ 31], in which the denoiser's prediction from the previous timestep is concatenated with the current input, with a 50% dropout probability during training. We condition on the space group via classifier-free guidance [32] to incorporate the conditioning on space group into the latent generative process (see","citing_arxiv_id":"2605.10115"},{"n":1,"role":"method","polarity":"use_method","paper_title":"SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation","primary_cat":"cs.CV","context_text":"at inference time, whereas most existing multi-person video generation methods require dedicated multi-person training data and task-specific fine-tuning. 2.3 Cross-Attention Control in Diffusion Models In diffusion models, cross-attention serves as the key mechanism for conditioning generation on external signals such as text prompts and reference images, applicable across both U-Net [56, 55] and Diffusion Transformer [50] backbones. Cross-attention control has been widely adopted for text- to-image generation [26, 6, 54], image editing [19, 72, 1], layout-guided generation [41, 9], subject- driven synthesis [65, 63, 68], noise optimization [ 17, 47], and identity preservation [ 79, 69, 76]. Building on these works, SocialDirector applies cross-attention masking and reweighting to multi-","citing_arxiv_id":"2605.10079"},{"n":1,"role":"method","polarity":"use_method","paper_title":"NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models","primary_cat":"cs.RO","context_text":"formulation ties prediction and control to the same denoising/flow trajectory: the future latents provide an explicit imagined context, while the action tokens are refined in the same process. 4 The backbone realizes this W AM with a Mixture-of-Transformers (MoT) [10] architecture, integrating a chunk-bidirectional video DiT (Wan 2.2-TI2V-5B [8]) with an Action-Expert DiT [40]. The clean observation latent v0 is pinned at t0 = 0, while predicted latent frames {vtf f }F f=1 and action tokens ata are processed within a joint sequence where all modalities share self-attention blocks but utilize modality-specific feed-forward layers. As established in §3.2, this shared attention serves as the physical layer for information gating: the","citing_arxiv_id":"2605.07794"},{"n":1,"role":"baseline","polarity":"baseline","paper_title":"Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction","primary_cat":"cs.CV","context_text":"ConvNeXt-UNet Backbone for Microscopic Locality and Multi-Scale Structure The diffusion backbone determines how effectively masked-diffusion pretraining captures pathology morphology. For cell-level dense prediction, features must preserve nuclear contours, chromatin texture, thin boundaries, and multi-scale tissue context. We therefore compare DiT [32], Attention U-Net [33], and ConvNeXt-UNet under the same pretraining and evaluation protocol. DiT offers scalable global modeling, but patch tokenization may weaken small structures and intra- patch boundaries. Attention U-Net is naturally suited to dense prediction, yet its conventional blocks may limit scalability. ConvNeXt-UNet combines U-Net-style multi-scale feature reuse with modern","citing_arxiv_id":"2605.08276"},{"n":1,"role":"method","polarity":"use_method","paper_title":"SDFlow: Similarity-Driven Flow Matching for Time Series Generation","primary_cat":"cs.AI","context_text":"Why SDFlow Generalizes.The kernel-smoothed anchor prior is a continuous initializer over low- rank VQ-latent coordinates, not a raw-data generator; Appendix C provides the theoretical perspective, while Section 5.3 reports held-out latent-flow stress tests. 4.4 Architecture and Training Flow Network Architecture.We employ a Diffusion Transformer (DiT) [ 20] adapted for sequential data. The input is the interpolated latent zt ∈R L×dc treated as a sequence of L tokens. Time conditioning uses adaptive layer normalization (AdaLN): sinusoidal embeddings of t are projected to scale and shift parameters that modulate layer statistics. A learnable global token is prepended to aggregate sequence-level information.","citing_arxiv_id":"2605.05736"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Seedance 1.0: Exploring the Boundaries of Video Generation Models","primary_cat":"cs.CV","context_text":"reconstruction by enforcing finer supervision on local textures and detailed structures. Taking into account appearance and motion modeling simultaneously, we apply a hybrid discriminator with an architecture similar to that used in PatchGAN [11]. 2.2 Diffusion Transformer With the visual tokens encoded by VAE and text tokens generated by a text encoder, we employ the transformer as our diffusion backbone [20], where a fine-tuned decoder-only LLM as the text encoder. The visual tokens are then concatenated with textual tokens and fed into the transformer blocks. Decoupled Spatial and Temporal Layers. Considering both training and inference efficiency, we build the diffusion transformer with decoupled spatial and temporal layers, where the spatial layers perform attention","citing_arxiv_id":"2506.09113"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Improving Video Generation with Human Feedback","primary_cat":"cs.CV","context_text":"Moreover, despite being trained on a disjoint dataset (see Fig. 10), it still achieves comparable performance on GenAI-Bench, indicating robust generalization across different eras of T2V models.Ablation studies are provided in Appendix E. 5.2 Video Alignment Training Setting.Our pretrained model pref is an internal, research-purpose video generation model based on Transformer architecture [53], which is trained using rectified flow (see Appendix 7 for details). Following SD3 [15], all alignment experiments fine-tune the Transformer using LoRA [24]. For training-based alignment methods, including SFT, Flow-DPO, and Flow-RWR, we adopt Video- Reward to provide the reward signals. For reward guidance, we employ thelatentreward model to generate rewards.","citing_arxiv_id":"2501.13918"}]},"authors":[]}}