Pith. sign in

REVIEW 14 cited by

ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02540 v3 pith:KOB76DQG submitted 2024-06-04 cs.CV

classification cs.CV
keywords quantizationvideochallengesdiffusiongenerationmemorytransformersvidit-q
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion transformers have demonstrated remarkable performance in visual generation tasks, such as generating realistic images or videos based on textual instructions. However, larger model sizes and multi-frame processing for video generation lead to increased computational and memory costs, posing challenges for practical deployment on edge devices. Post-Training Quantization (PTQ) is an effective method for reducing memory costs and computational complexity. When quantizing diffusion transformers, we find that existing quantization methods face challenges when applied to text-to-image and video tasks. To address these challenges, we begin by systematically analyzing the source of quantization error and conclude with the unique challenges posed by DiT quantization. Accordingly, we design an improved quantization scheme: ViDiT-Q (Video & Image Diffusion Transformer Quantization), tailored specifically for DiT models. We validate the effectiveness of ViDiT-Q across a variety of text-to-image and video models, achieving W8A8 and W4A8 with negligible degradation in visual quality and metrics. Additionally, we implement efficient GPU kernels to achieve practical 2-2.5x memory saving and a 1.4-1.7x end-to-end latency speedup.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A temporal-spatial LSB mask over one shared weight buffer lets diffusion models use lower bit precision in less sensitive denoising stages, cutting compute by 25-50% on bit-serial hardware with no loss in image quality.

  2. QuantWAMs: Calibrating at the Right Granularity for World Action Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Aligning PTQ decisions to WAM structure, closed-loop rollouts, and the joint video–action objective yields W4A4 policies within 0.2–0.7 pp of FP16 on simulation benchmarks with ~29% block memory.

  3. DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A post-training quantization method that combines learned channel scaling and power-of-two scaling keeps diffusion image quality high at 4-bit weight, 6-bit activation precision.

  4. Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Q-VDiT quantizes video diffusion transformers to 3-4 bit weights by adding a learned rank-1 error correction (TQE) and a temporal distribution distillation loss (TMD), nearly doubling VBench scene consistency at W3A6 ...

  5. MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MARche accelerates masked autoregressive image generation by caching stable token projections and refreshing only attention-selected tokens, reaching up to 1.72x speedup with some loss in FID.

  6. On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A training-free pipeline makes diffusion text-to-video generation run on an iPhone 15 Pro with quality close to GPU output, at the cost of slower generation.

  7. Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Sparse VideoGen accelerates video diffusion transformers by about 2.3x with only small quality loss by classifying attention heads into spatial and temporal sparse patterns and using hardware-friendly layouts.

  8. CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A token-pruning cache method cuts diffusion model computation by roughly half while keeping image quality, using noise magnitude, spatial clustering, and selection balance.

  9. 1.58-bit FLUX

    cs.CV 2024-12 reject novelty 6.0 of 10

    A post-training method reduces 99.5% of FLUX.1-dev's transformer weights to ternary values and reports roughly comparable text-to-image quality with large storage and memory savings.

  10. FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FADA distills a diffusion-based talking avatar model into a 6-step student that mimics multi-condition classifier-free guidance with learnable tokens, achieving 4.17 to 12.5 times NFE speedup with comparable quality.

  11. MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MPQ-DM combines kurtosis-based intra-layer mixed-precision weight quantization with time-smoothed relation distillation to keep diffusion models accurate at 2 to 4 bit widths.

  12. HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Training-free sparse attention for video DiTs cuts latency up to ~2.1× via 3D local-window clustering, hybrid step updates, and hardware-aware cluster merging while improving fidelity over prior sparse methods.

  13. Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.

  14. LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization

    cs.CV 2024-11 conditional novelty 4.0 of 10

    LiteVAR compresses VAR image generation via multi-diagonal windowed attention, CFG output sharing, and mixed-precision quantization, reporting up to 85% attention savings and 50% memory reduction with minimal FID change.

Pith tools