Pith. sign in

super hub Mixed citations

Scalable diffusion models with transformers

Mixed citation behavior. Most common role is method (47%).

101 Pith papers citing it
Method 47% of classified citations

hub tools

citation-role summary

method 8 background 6 baseline 3

citation-polarity summary

claims ledger

  • baseline 44 57.35 58.66 93.07 AnimateDiff-V2 [35] - 82.90 69.75 95.30 97.68 98.75 97.76 40.83 67.16 70.10 90.90 VideoCrafter-2.0 [10] - 82.20 73.42 96.85 98.22 98.41 97.73 42.50 63.13 67.22 92.55 CogVideoX [140] 5B 82.75 77.04 96.23 96.52 98.66 96.92 70.97 61.98 62.90 85.23 Kling [50] - 83.39 75.68 98.33 97.60 99.30 99.40 46.94 61.21 65.62 87.24 Open-Sora-2.0 [90] - 82.10 80.14 98.75 98.00 99.40 99.49 20.74 64.33 65.62 94.50 Gen-3 [96] - 84.11 75.17 97.10 96.62 98.61 99.23 60.14 63.34 66.82 87.81 Step-Vi
  • background The code is available at https://github.com/ywlq/MotionCache. 1 Introduction Video generation models [12, 18, 24, 26, 34, 41, 47] have achieved remarkable success, facilitating applications ranging from autonomous driving [10, 11, 37] and cinematic creation [6, 38] to social media [3]. While architectures have evolved from U-Nets [2, 27, 29] to scalable Diffusion Transformers (DiTs) [25], practical deployment is hindered by the prohibitive costs of iterative denoising. Moreover, the quadratic co
  • method DINOv2 [46] to extract frame-wise visual features. These features are further processed by a spatial-temporal transformer [ 47]. The decoder is implemented with standard transformer [ 48]. The latent action dimension n is set to 16. For RotVLA, the VLM backbone is initialized from InternVL3.5-1B [38], followed by an action expert implemented as a 24-layer Diffusion Transformer (DiT) [49]. The full model contains approximately 1.7B, including 304M in the vision encoder, 752M in the language model
  • method reconstruction by enforcing finer supervision on local textures and detailed structures. Taking into account appearance and motion modeling simultaneously, we apply a hybrid discriminator with an architecture similar to that used in PatchGAN [11]. 2.2 Diffusion Transformer With the visual tokens encoded by VAE and text tokens generated by a text encoder, we employ the transformer as our diffusion backbone [20], where a fine-tuned decoder-only LLM as the text encoder. The visual tokens are then c
  • baseline ConvNeXt-UNet Backbone for Microscopic Locality and Multi-Scale Structure The diffusion backbone determines how effectively masked-diffusion pretraining captures pathology morphology. For cell-level dense prediction, features must preserve nuclear contours, chromatin texture, thin boundaries, and multi-scale tissue context. We therefore compare DiT [32], Attention U-Net [33], and ConvNeXt-UNet under the same pretraining and evaluation protocol. DiT offers scalable global modeling, but patch toke
  • method For both G and O, we can use the empirical training data distribution ˆpdata(G) and ˆpdata(O|G). By definingX :=(A,W,F, ℓ), our model can be written as pθ(G, O,Z,X) = ˆpdata(G) ˆpdata(O|G)p θFM(Z|G, O)p θD(X|Z, G, O),(5) where θ= (θ FM, θD) are the parameters of the flow matching and decoder neural networks. Similar to ADiT, we use a Diffusion Transformer (DiT) [30] as our denoiser network F, where conditional information gets added through adaptive layer norm. Moreover, we use self-conditioning

co-cited works

representative citing papers

GEAR: Guided End-to-End AutoRegression for Image Synthesis

cs.CV · 2026-06-30 · unverdicted · novelty 7.0

GEAR jointly trains VQ tokenizer and AR generator end-to-end via dual hard/soft read-out and representation alignment, achieving up to 10x faster ImageNet gFID convergence than LlamaGen-REPA while generalizing across quantizers and to text-to-image.

Vision-Language Binding in In-Context Image Generation

cs.CV · 2026-05-23 · unverdicted · novelty 7.0

Text tokens in FLUX.2 absorb reference image properties like color and style to influence outputs while pixel-exact details bypass them, localized to padding tokens via causal interventions.

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization

cs.CV · 2026-05-11 · unverdicted · novelty 7.0 · 2 refs

DRoRAE adaptively fuses multi-layer features from vision encoders via energy-constrained routing to enrich visual tokens, cutting rFID from 0.57 to 0.29 and generation FID from 1.74 to 1.65 on ImageNet-256 while revealing a log-linear scaling law with fusion capacity.

Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation

cs.CV · 2026-04-12 · unverdicted · novelty 7.0

Immune2V immunizes images against dual-stream I2V generation by enforcing temporally balanced latent divergence and aligning generative features to a precomputed collapse trajectory, yielding stronger persistent degradation than image-level baselines.

Training Agents Inside of Scalable World Models

cs.AI · 2025-09-29 · conditional · novelty 7.0

Dreamer 4 is the first agent to obtain diamonds in Minecraft from only offline data by reinforcement learning inside a scalable world model that accurately predicts game mechanics.

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

cs.CV · 2026-07-07 · conditional · novelty 6.5

A feed-forward feature-Gaussian plus one-step geometry-aware pixel-flow simulator converts large image collections into 20K interactive scenes and 10M+ navigation samples that improve zero-shot Habitat and real-robot performance.

OpenCoF: Learning to Reason Through Video Generation

cs.CV · 2026-07-09 · conditional · novelty 6.0

Fine-tuning a video generator on a new 17K reasoning-video dataset improves Chain-of-Frame reasoning, and adding learnable visual/textual reasoning tokens yields further gains on external benchmarks.

Generative OOD-regularized Model-based Policy Optimization

cs.LG · 2026-05-23 · unverdicted · novelty 6.0

GORMPO uses generative models for density-based regularization in model-based offline RL, outperforming baselines by 17% on a medical dataset while providing theoretical guarantees under mild assumptions.

citing papers explorer

Showing 50 of 101 citing papers.