REVIEW 3 major objections 5 minor 60 references
DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read DSTAR claims to make DiT inference up to 7.33x faster and 41.89x more energy-efficient than a leading GPU by combining fine-grained differential quantization with block-wise attention-score reuse, with no accuracy degradation.
desk verdict A genuine co-design with a real unanswered memory-traffic question about cached attention scores; the headline speedups may survive, but the paper doesn't show it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two algorithmic mechanisms carry the argument. FMDQ (fine-grained mixed-precision differential quantization) splits the differential activation tensor into sub-channels (token tiles × channel), quantizes each using power-of-two scales at 2/4/6/8 bits, decomposes all sub-channels into uniform 2-bit slices, and reorders slices by FPoT—a fused precision-and-scale tag—so that partial-sum accumulation can be aligned with bit shifts instead of expensive requantization. A static base scale, profiled per model, avoids a second data pass. SAR (sparse attention reuse) uses a statically profiled block-wise mask over attention score matrices, prunes low-importance blocks (spatial reduction), caches the
What would settle it
Run the DSTAR accelerator (or a cycle-accurate FPGA prototype at a comparable technology node) on the same seven DiT workloads with the same per-model thresholds, and compare end-to-end latency and energy against an A100-class GPU under identical prompts and seeds. If per-model speedups fall outside the reported 3.54x–7.33x range, or energy savings fall short of 18.89x, the simulator's models are the weak link. For the algorithm half, re-profile the attention masks on new data and check whether FID/IS/CLIP remain within the stated 5%/1-point tolerance.
Extended reading notes
Core claim
The paper's central claim is that temporal and spatial redundancy in DiTs can be attacked together without accuracy loss. For linear layers, FMDQ quantizes the differential activation between adjacent timesteps at a sub-channel granularity—splitting the token dimension into tiles and each tile's channels separately—assigning 2, 4, 6, or 8 bits per sub-channel with power-of-two scales, then decomposing everything into uniform 2-bit slices so the arithmetic unit stays dense. For attention, SAR exploits an observed block-wise sparsity pattern in score matrices that recurs across timesteps: a static mask prunes low-importance blocks, and the sparse score matrix is cached and reused for several s
Load-bearing premise
The headline speedup and energy numbers are generated by a custom discrete-event simulator, not a fabricated chip; if the simulator's timing, memory, or power models are optimistic, the 7.33x and 41.89x figures will not materialize on real hardware.
Editorial extensions
If this is right
- If DSTAR's numbers hold, the same seven DiT workloads would run at 3.54x–7.33x lower latency and 18.89x–41.89x lower energy than an A100-class GPU, roughly preserving FID/IS/CLIP.
- FMDQ alone would bring average activation bit-width to 4.3 bits (3.2 bits for highly redundant video models), compared with 5.4 bits for the closest prior quantization scheme—meaning linear-layer compute falls roughly in proportion.
- SAR alone would cut attention-layer computation by up to half, and because the sparse mask is static it avoids runtime prediction overheads and skips QK^T/softmax work during reuse intervals.
- The framework adapts across sampling schedules: in an ablation from 10 to 50 timesteps, speedup stays in a narrow 2.45x–2.76x band while average bit-width decreases as timesteps increase—so gains persist for both few-step and many-step settings.
Reading between the lines
- Editorial: the 7.33x/41.89x figures come from a discrete-event simulation of the accelerator, not silicon; until the design is emulated on FPGA or fabricated, the practical ceiling is best read as an upper bound.
- Editorial: the same differential-quantization trick could transfer to other iterative generative models (flow matching, consistency models) where adjacent solver steps produce similar hidden states, provided their outlier distributions also cluster intra-channel.
- Editorial: because SAR's mask and FMDQ's base scale are profiled per model weights and step count, deployment to new resolutions, prompts, or fine-tuned weights may require re-profiling; the paper's sensitivity study covers seeds and step counts but not all distribution shifts.
- Editorial: the energy comparison normalizes accelerators to equal TOPS and same technology node; a full-system accounting including host CPUs and DRAM refresh would likely reduce the absolute savings, though the relative ordering may stand.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DSTAR, a software-hardware co-design framework for accelerating diffusion transformer (DiT) inference by exploiting temporal and spatial redundancy. At the algorithm level, it introduces fine-grained mixed-precision differential quantization (FMDQ) for linear layers and sparse attention reuse (SAR) for attention layers. The hardware design is a 32-core accelerator with mixed-precision systolic MPUs, a quantization unit (QPU), and a vector unit (VPU). Evaluation on seven DiTs claims up to 7.33x latency speedup and 41.89x energy savings versus an NVIDIA A100, and up to 2.54x latency speedup versus prior accelerators, all described as 'without accuracy degradation.' The central speedup and energy numbers are produced by a custom SimPy/Ramulator simulator with power scaled from 45nm synthesis via DeepScaleTool.
Significance. If the claimed results hold, DSTAR would be a meaningful contribution to DiT accelerator design: it jointly targets FFN and attention bottlenecks, proposes a hardware-friendly sub-channel mixed-precision scheme with slice decomposition, and evaluates on a broader set of DiTs than prior UNet-focused accelerators. The slice-decomposition arithmetic in §3.1 is internally consistent, and the GPU-SAR Triton kernel provides some real measured evidence that attention-score reuse helps. However, the headline latency/energy numbers rest entirely on a simulator whose treatment of cached attention-score traffic is unspecified, and the 'no accuracy degradation' claim is supported by per-model tuned thresholds without confidence intervals. The potential is high, but the current evidence is not yet sufficient for the strength of the abstract claims.
major comments (3)
- [§3.2, §4.3, Table 3] SAR caches sparse attention-score matrices, but the manuscript never states where these matrices reside. On-chip storage is only about 3.1MB total (512KB global scratchpad plus 32 MPC unified buffers per Table 3), while a single 4096-token, 16-head score matrix is 536MB in FP16 (or ~268MB in INT8) before pruning; even after 37% block pruning, a 28-layer model needs gigabytes. These matrices must therefore live in HBM, causing hundreds of MB to GB of reads/writes per reused step. The simulator description in §5.1 does not state whether this score-cache spill traffic and HBM contention are modeled. Since DSTAR's HBM2 (1500GB/s) has lower bandwidth than the A100 (2039GB/s), the skipped QK^T/softmax compute may be offset by DRAM traffic. This is load-bearing for the claimed 7.33x/41.89x speedups/energy savings. Please report the cache capacity and location, and provide a sensitivity analysis
- [§5.2, Table 2, §3.1/§3.2] The 'without accuracy degradation' claim is not currently supported with sufficient rigor. Each model uses per-model SAR Thr/Interval and FMDQ ScaleBase values (Table 2), and §3.1 states that Thr0 is tuned to encourage low precision. The reported 4.3-bit average and 37% pruning are therefore partly produced by threshold selection rather than being parameter-free predictions. Evaluation uses a single 1K-sample run per model with no confidence intervals, and the sensitivity study in Fig. 18 reports PSNR, not the FID/IS/CS metrics used in Table 2. The manuscript's own text acknowledges that SAR requires re-profiling whenever model weights or step counts change (§3.2). Please provide FID/IS/CS means and confidence intervals over multiple seeds and dataset partitions, fix thresholds before evaluation, and report how sensitive the quality metrics are to Thr/Interval.
- [§5.1, §5.3] The central hardware results are produced entirely by a custom discrete-event simulator (SimPy + Ramulator) with no validation against RTL, FPGA, or a real chip, and no artifact is released. Power is scaled from 45nm synthesis via DeepScaleTool, and all accelerator baselines are also implemented inside the same simulator. This makes the 7.33x latency and 41.89x energy claims difficult to verify. The measured GPU and GPU-SAR baselines are real, but DSTAR-ALL is not. At minimum, please validate the simulator timing/energy against a small RTL or FPGA prototype, show per-component cycle breakdowns for the SAR reuse phase, and release the simulator/traces. In the absence of validation, the abstract should state that these are simulator projections, not measured results.
minor comments (5)
- [Figures 16/17 captions] The abbreviation is inconsistent: the text and some captions say DSTAR-DiffQ while others say DSTAR-FMDQ. Unify the naming.
- [Table 2] Typo: 'Thre=1e-4' should be 'Thr=1e-4'. Also clarify whether SAR Thr is a fraction of scores pruned or an absolute importance threshold.
- [§3.1, Eq. (4)] The formula for FPoT is typeset ambiguously ('PCeil/base 2×2^PoTi'); please rewrite with clear parentheses and define the rounding in Eq. (3) explicitly.
- [Fig. 1] The y-axis labels and stacked-bar values are difficult to read; consider a table or larger font with explicit percentages.
- [§5.3] The figure labels 'x4.67x' and 'x7.33x' are redundant or inconsistent; choose a single speedup annotation convention.
Circularity Check
No significant circularity: reported results are empirical outcomes of tuned design parameters, and self-citations are not load-bearing.
full rationale
The derivation chain is not circular. FMDQ's static base and thresholds are explicitly profiled/tuned design parameters (Sec. 3.1: 'using a predefined static base, selected through dataset profiling' and 'By tuning Thr0, we can encourage the selection of lower precision while maintaining negligible accuracy loss'), so the reported 4.3-bit average and 37% pruning are measured operating points under an externally stated accuracy constraint (Table 2: <5% FID/FVD degradation, <1 IS drop), not predictions that reduce by construction to their inputs. SAR's Thr/Interval and block-wise masks are likewise per-model hyperparameters chosen by static profiling, with a separate train/test sensitivity analysis (Fig. 18). The quantization equations (2)-(5) are self-contained arithmetic; no claimed output is defined in terms of itself. The self-citations present ([13] OliVe, [59] Oltron) are background/motivation references for known problems, not load-bearing justification of DSTAR's contributions. Whether the SimPy/Ramulator simulator accurately charges for cached attention-score memory traffic is a verification/correctness risk, not a circularity. No specific reduction of a claimed result to its own inputs was found.
Assumptions & free parameters
free parameters (4)
- FMDQ quantization threshold Thr0 =
not reported (tuned)
- SAR pruning threshold Thr =
DiT-XL: 2e-3; other models: 1e-4
- SAR recompute interval =
2 (all models)
- FMDQ static base ScaleBase =
64/256, model/layer dependent
assumptions (6)
- domain assumption Differential activations between adjacent timesteps are small enough that low-bit quantization (2-8 bits) introduces negligible generation error
- domain assumption DiT attention score matrices exhibit block-wise spatial sparsity and temporal similarity so a static per-model mask can be reused for n steps
- domain assumption Static profiling on training prompts transfers to held-out prompts and to all denoising steps
- domain assumption The SimPy/Ramulator discrete-event simulator and DeepScaleTool power scaling faithfully model DSTAR and the baseline accelerators
- domain assumption Softmax attention scores in [0,1] are less sensitive to reuse error than value matrices
- standard math Linearity y' = y + kΔx and the PoT slice-decomposition arithmetic are correct
Cite this review
Pith. "Pith review of DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction." pith.science (2026). https://pith.science/paper/ZBRRLI45
@misc{pith2026260715846,
author = {Pith},
title = {Pith review of: DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZBRRLI45}},
note = {Machine review of arXiv:2607.15846}
}
read the original abstract
Diffusion Transformers (DiTs) have been widely used in many tasks, including image synthesis, video generation, and content editing. However, their multi-iteration inference process leads to performance inefficiency and high energy consumption. Existing acceleration methods primarily focus on reducing temporal redundancy between adjacent timesteps, but often overlook the specific features of DiTs. As a result, these approaches either suffer from great accuracy degradation or fail to achieve high efficiency. We present DSTAR, a software-hardware co-design framework that accelerates DiT inference by reducing spatial and temporal redundancy. At the algorithmic level, DSTAR introduces a fine-grained mixed-precision quantization method for differential activations in linear operations, significantly increasing the proportion of low-bit computations. Additionally, DSTAR incorporates a sparse attention reuse mechanism to minimize redundant computation in attention layers. For architectural support, we design a specialized hardware accelerator which achieves high efficiency in both latency and energy consumption. Evaluation on seven typical DiTs demonstrates that DSTAR achieves up to 7.33x latency speedup and 41.89x energy savings compared to an NVIDIA A100 GPU, and achieves up to 2.54x latency speedup and 3.68x energy savings compared to SOTA accelerators, without accuracy degradation.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
[n. d.]. NVIDIA A100 PCIe 40 GB Specs — techpowerup.com. https://www. techpowerup.com/gpu-specs/a100-pcie-40-gb.c3623. [Accessed 09-04-2025]
2025
-
[2]
Shubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam, Koyel Mukherjee, and Shiv Kumar Saini. 2024. Approximate Caching for Ef- ficiently Serving Text-to-Image Diffusion Models. In21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). USENIX Associa- tion, Santa Clara, CA, 1173–1189. https://www.usenix.org/confer...
2024
-
[3]
Zhenyu Bai, Pranav Dangi, Huize Li, and Tulika Mitra. 2024. SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs. In Proceedings of the 61st ACM/IEEE Design Automation Conference(San Francisco, CA, USA)(DAC ’24). Association for Computing Machinery, New York, NY, USA, Article 93, 6 pages. https://doi.org/10.1145/3649329.3658488
arXiv 2024
-
[4]
Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas
Rajeev Balasubramonian, Andrew B. Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas. 2017. CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories.ACM Trans. Archit. Code Optim.14, 2, Article 14 (June 2017), 25 pages. https://doi.org/10.1145/3085572
doi:10.1145/3085572 2017
- [5]
-
[6]
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rombach. 2023. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv:2311.15127 [cs.CV] https://arxiv.org/abs/2311.15127
arXiv 2023
-
[7]
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhong- dao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. 2023. PixArt-𝛼: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis. arXiv:2310.00426 [cs.CV] https://arxiv.org/abs/2310.00426
arXiv 2023
-
[8]
Lei Chen, Yuan Meng, Chen Tang, Xinzhu Ma, Jingyan Jiang, Xin Wang, Zhi Wang, and Wenwu Zhu. 2025. Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers. InProceedings of the Computer Vision and Pattern Recognition Conference (CVPR). 28306–28315
2025
Show all 60 references
-
[9]
Jack Choquette and Wish Gandhi. 2020. NVIDIA A100 GPU: Performance & Innovation for GPU Computing. In2020 IEEE Hot Chips 32 Symposium (HCS). 1–43. https://doi.org/10.1109/HCS49909.2020.9220622
2020
-
[10]
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAt- tention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), ...
2022
-
[11]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale. InAd- vances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. C...
2022
-
[12]
Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu, Yanjing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, and Baochang Zhang. 2023. Implicit Diffu- sion Models for Continuous Super-Resolution. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP...
2023
-
[13]
Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. 2023. OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization. InProceedings of the 50th Annual International Symposium on Com...
2023
-
[14]
Ruiqi Guo, Lei Wang, Xiaofeng Chen, Hao Sun, Zhiheng Yue, Yubin Qin, Huiming Han, Yang Wang, Fengbin Tu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2024. 20.2 A 28nm 74.34TFLOPS/W BF16 Heterogenous CIM-Based Accelerator Exploiting Denoising-Similarity for Diffusion Models. In2024 I...
2024
-
[15]
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W. Lee. 2021. ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks. In2021 ACM/IEEE 48th Annual International Symposium on Computer Arch...
2021
-
[17]
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. 2023. T2i- compbench: A comprehensive benchmark for open-world compositional text-to- image generation.Advances in Neural Information Processing Systems36 (2023), 78723–78747
2023
-
[18]
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023. Imagic: Text-Based Real Image Editing With Diffusion Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6007–6017
2023
-
[19]
Sungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park, and Won Woo Ro. 2025. Ditto: Accelerating Diffusion Model via Temporal Value Similarity. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). 338–352. https://doi.org/10.1109/HPCA61900.2025.00035
2025
-
[20]
Sangjin Kim, Jungjun Oh, Jeonggyu So, Yuseon Choi, Sangyeob Kim, Dongseok Im, Gwangtae Park, and Hoi-Jun Yoo. 2026. EdgeDiff: Energy-Efficient Multi- Modal Few-Step Diffusion Model Accelerator Using Mixed-Precision and Re- ordered Group Quantization.IEEE Journal of Solid-State...
2026
-
[22]
Raghuraman Krishnamoorthi. 2018. Quantizing deep convolutional networks for efficient inference: A whitepaper. arXiv:1806.08342 [cs.LG] https://arxiv.org/ abs/1806.08342
2018 arXiv
-
[23]
Black Forest Labs. 2024. FLUX. https://github.com/black-forest-labs/flux
2024
-
[24]
Jungi Lee, Wonbeom Lee, and Jaewoong Sim. 2024. Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). 1048–1062. https://doi.org/10.1109/ISCA59077.2024.00080
2024
-
[25]
Huize Li, Zhaoying Li, Zhenyu Bai, and Tulika Mitra. 2024. ASADI: Accelerating Sparse Attention Using Diagonal-based In-Situ Computing. In2024 IEEE Interna- tional Symposium on High-Performance Computer Architecture (HPCA). 774–787. https://doi.org/10.1109/HPCA57654.2024.00065
2024
-
[26]
Xiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang, Zhen Dong, Daniel Kang, Shang- hang Zhang, and Kurt Keutzer. 2023. Q-Diffusion: Quantizing Diffusion Models. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 17535–17545
2023
-
[27]
Yuhang Li, Xin Dong, and Wei Wang. 2020. Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural Networks. arXiv:1909.13144 [cs.LG] https://arxiv.org/abs/1909.13144
2020 arXiv
-
[28]
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, Tanghui Jia, Junwu Zhang, Zhenyu Tang, Yatian Pang, Bin She, Cen Yan, Zhiheng Hu, Xiaoyi Dong, Lin Chen, Zhang Pan, Xing Zhou, Shaoling Dong, Yonghong Tian, ...
2024 arXiv
-
[29]
Belongie, Lubomir D
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Gir- shick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll’a r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context.CoRRabs/1405.0312 (2014). arXiv:1405.0312 http://arxiv.org/...
2014 arXiv
-
[30]
Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. 2023. FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation.arXiv preprint arXiv: 2311.01813(2023)
2023 arXiv
-
[31]
Jinming Lou, Wenyang Luo, Yufan Liu, Bing Li, Xinmiao Ding, Weiming Hu, Jiajiong Cao, Yuming Li, and Chenguang Ma. 2024. Token Caching for Diffusion Transformer Acceleration. arXiv:2409.18523 [cs.LG] https://arxiv.org/abs/2409. 18523
2024
-
[32]
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan LI, and Jun Zhu. 2022. DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps. InAdvances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Ch...
2022
-
[33]
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021. Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture. InMICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture(Virtual Event,...
2021
-
[34]
Nisa Bostancı, Ataberk Olgun, A
Haocong Luo, Yahya Can Tuğrul, F. Nisa Bostancı, Ataberk Olgun, A. Giray Yağlıkçı, and Onur Mutlu. 2024. Ramulator 2.0: A Modern, Modular, and Extensi- ble DRAM Simulator.IEEE Computer Architecture Letters23, 1 (2024), 112–116. https://doi.org/10.1109/LCA.2023.3333759
2024
-
[35]
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2024. DeepCache: Accelerating Diffusion Models for Free. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 15762–15772
2024
-
[36]
Xin Ma, Yaohui Wang, Xinyuan Chen, Gengyun Jia, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. 2025. Latte: Latent Diffusion Transformer for Video Generation.Transactions on Machine Learning Research(2025)
2025
-
[37]
OpenAI. 2021. Triton: An Open-Source GPU Programming Language. https: //github.com/openai/triton. Accessed: 2026-03-06
2021
-
[38]
Pavlov, A
I. Pavlov, A. Ivanov, and S. Stafievskiy. 2023. Text-to-Image Benchmark: A benchmark for generative models. https://github.com/boomb0om/text2image- benchmark. Version 0.1.0
2023
-
[39]
William Peebles and Saining Xie. 2023. Scalable Diffusion Models with Trans- formers. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 4195–4205
2023
-
[40]
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019. Transfer Learning in Biomedical Natural Language Processing: An Evaluation of BERT and ELMo on Ten Bench- marking Datasets. arXiv:1906.05474 [cs.CL] https://arxiv.org/abs/1906.05474
2019 arXiv
-
[41]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv:2307.01952 [cs.CV] https://arxiv.org/abs/2307.01952
2023 arXiv
-
[42]
Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2023. Fact: Ffn-attention co-optimized transformer architecture with eager correlation prediction. InProceedings of the 50th Annual International Symposium on Compu...
2023
-
[43]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. InProceedings ...
2021
-
[44]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.Journal of Machine Learning Research21, 140 (2020), 1–67. http:...
2020
-
[45]
Akshat Ramachandran, Souvik Kundu, and Tushar Krishna. 2025. MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quanti- zation. arXiv:2411.05282 [cs.AR] https://arxiv.org/abs/2411.05282
2025 arXiv
-
[46]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10684–10695
2022
-
[47]
Tim Salimans and Jonathan Ho. 2022. Progressive Distillation for Fast Sampling of Diffusion Models. arXiv:2202.00512 [cs.LG] https://arxiv.org/abs/2202.00512
2022 arXiv
-
[48]
Satyabrata Sarangi and Bevan Baas. 2021. DeepScaleTool: A Tool for the Accurate Estimation of Technology Scaling in the Deep-Submicron Era. In2021 IEEE International Symposium on Circuits and Systems (ISCAS). 1–5. https://doi.org/ 10.1109/ISCAS51556.2021.9401196
2021
-
[49]
Team SimPy. 2025. SimPy: Discrete-event simulation for Python. https://github. com/SimPy/SimPy. Accessed: 2025-08-20
2025
-
[50]
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2022. Denoising Diffusion Implicit Models. arXiv:2010.02502 [cs.LG] https://arxiv.org/abs/2010.02502
2022 arXiv
-
[51]
Stine, Ivan Castellanos, Michael Wood, Jeff Henson, Fred Love, W
James E. Stine, Ivan Castellanos, Michael Wood, Jeff Henson, Fred Love, W. Rhett Davis, Paul D. Franzon, Michael Bucher, Sunil Basavarajaiah, Julie Oh, and Ravi Jenkal. 2007. FreePDK: An Open-Source Variation-Aware Design Kit. In2007 IEEE International Conference on Microelect...
2007 doi
-
[52]
Desen Sun, Henry Tian, Tim Lu, and Sihang Liu. 2024. FlexCache: Flexible Approximate Cache System for Video Diffusion. arXiv:2501.04012 [cs.MM] https://arxiv.org/abs/2501.04012
2024 arXiv
-
[53]
Desen Sun, Zepeng Zhao, and Yuke Wang. 2025. PATCHEDSERVE: A Patch Man- agement Framework for SLO-Optimized Hybrid Resolution Diffusion Serving. arXiv:2501.09253 [cs.DC] https://arxiv.org/abs/2501.09253
2025
-
[54]
Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, Dhruv Nair, Sayak Paul, William Berman, Yiyi Xu, Steven Liu, and Thomas Wolf. 2022. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers
2022
-
[55]
Huizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue, Jinxi Li, Yubin Qin, Sihan Guan, Qinze Yang, Yang Wang, Chao Li, Yang Hu, and Shouyi Yin. 2024. SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling. In2024 57th IEEE/ACM International Sym...
2024
-
[56]
Yu Emma Wang, Gu-Yeon Wei, and David Brooks. 2019. Benchmarking TPU, GPU, and CPU Platforms for Deep Learning. arXiv:1907.10701 [cs.LG] https: //arxiv.org/abs/1907.10701
2019 arXiv
-
[57]
Yuchen Xia, Divyam Sharma, Yichao Yuan, Souvik Kundu, and Nishil Talati. 2025. MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models. arXiv:2503.11972 [cs.DC] https://arxiv.org/abs/2503.11972
2025 arXiv
-
[58]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. arXiv:2309.17453 [cs.CL] https://arxiv.org/abs/2309.17453
2024 arXiv
-
[59]
Chenhao Xue, Chen Zhang, Xun Jiang, Zhutianya Gao, Yibo Lin, and Guangyu Sun. 2024. Oltron: Algorithm-Hardware Co-design for Outlier-Aware Quanti- zation of LLMs with Inter-/Intra-Layer Adaptation. InProceedings of the 61st ACM/IEEE Design Automation Conference(San Francisco, ...
2024
-
[60]
Zhihang Yuan, Hanling Zhang, Pu Lu, Xuefei Ning, Linfeng Zhang, Tianchen Zhao, Shengen Yan, Guohao Dai, and Yu Wang. 2024. DiTFastAttn: Atten- tion Compression for Diffusion Transformer Models. InAdvances in Neu- ral Information Processing Systems, A. Globerson, L. Mackey, D. ...
2024
-
[61]
Jintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei, Haocheng Xi, Jun Zhu, and Jianfei Chen. 2025. SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference. arXiv:2502.18137 [cs.LG] https: //arxiv.org/abs/2502.18137 14 DSTAR: Accelerating ...
2025
-
[62]
Zihan Zou, Xinming Yan, Shun Zhang, Peng Zheng, Guang Yang, Hao Cai, and Bo Liu. 2025. S-DMA: Sparse Diffusion Models Acceleration via Spatiality- Aware Prediction and Dimension-Adaptive Dataflow. InProceedings of the 58th IEEE/ACM International Symposium on Microarchitecture ...
2025
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.