REVIEW 3 major objections 4 minor 52 references
A mostly-linear attention mix can match full-softmax video DiTs in quality while running at linear-attention speed—this paper shows how, at 5B and 14B scales.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:05 UTC pith:WAE23BSP
load-bearing objection A well-executed, honest systems paper on hybrid linear/softmax attention for video DiTs; the efficiency story holds up, but the 'matches full-softmax' quality claim rests on a small proxy and needs a production-scale check. the 3 major comments →
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
SANA-Video 2.0 is a scratch-trained hybrid-attention video diffusion transformer (5B and 14B) that matches full-softmax quality while keeping the O(N) scaling of linear attention. The architecture interleaves gated linear attention with periodic softmax anchors at a 3:1 ratio (25% softmax), and Block Attention Residuals route completed block summaries into later linear layers. Proxy studies at 256p found 25% softmax to be the best quality-efficiency knee, not the lowest-loss point (50% was lower loss but much slower). The 5B model reaches VBench Total 84.30 at 480x832x81 in 13.2s on one H100; its compiled DiT forward is 3.2x faster than full softmax at 720p/60s, a gap that widens with durati
What carries the argument
Hybrid Linear-Softmax Attention: gated linear attention for O(N)-dominated mixing, with periodic gated-softmax anchors every fourth layer. Block Attention Residuals (AttnRes): routes the completed block summaries and the current partial sum into later layers via a depth-shared learned query per branch, exposing the softmax-refreshed representations to the linear majority.
Load-bearing premise
Decisions made from short proxy runs at 256p, such as the 25% softmax ratio and the AttnRes block span, transfer to the much larger 5B and 14B production models at high resolution and long duration.
What would settle it
Run a production-scale ablation that varies the softmax ratio and AttnRes on/off, training long enough to produce a final checkpoint, and compare VBench or similar quality at matched latency. If 25% softmax or AttnRes does not yield comparable quality to full softmax at that scale, the central claim fails.
If this is right
- Long-video generation becomes much cheaper at a given quality: the efficiency advantage grows with duration. With a mostly-linear backbone, the cost of generating 60s 720p clips is substantially reduced relative to full-softmax models at the same scale.
- Scaling to longer horizons is more feasible: the O(N) backbone shifts the bottleneck away from the token-mixing cost, allowing training and inference to extend beyond current 8s horizons.
- Hardware-friendly backbones combine with deployment optimization: the conv-free SwiGLU FFN and fixed anchors map directly to fused kernels and sparse attention, yielding a measured 3.58x end-to-end speedup on B200.
- Low-precision quantization is possible without quality loss: QAT with MXFP4 weights and MXFP8 activations matches BF16 VBench scores, cutting static storage by 68%.
- The design transfers to physical AI: fine-tuning on robot and egocentric video produces realistic manipulation clips competitive with models several times larger.
Where Pith is reading between the lines
- If hybrid attention is adopted broadly in video DiTs, quality-vs-cost trade-offs could shift for the whole field: the paper suggests a recipe to make video generation widely accessible on single GPUs, potentially democratizing long-form video creation.
- The same hybrid-plus-AttnRes recipe might be carried into causal generation: the linear operator drops the delta-rule update, making it a natural seed for a causal Gated DeltaNet, which could bring the efficiency and quality to streaming and world-model settings.
- The proxy-study methodology (short runs at 256p to select architecture) implies a testable extension: check whether the 25% softmax ratio and S=8 block span remain optimal at production scales and much longer sequences, or whether the Pareto knee shifts toward more softmax.
- The authors' observation that AttnRes adds only a small quality edge but boosts deep-layer rank suggests a potential standalone diagnostic: measuring effective-rank recovery could become a general tool for evaluating whether 'expressiveness' of hybrid architectures is genuinely improving.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SANA-Video 2.0, a video diffusion transformer that replaces most softmax attention with gated linear attention, inserts periodic softmax anchors at a 3:1 linear-to-softmax ratio, and adds Block Attention Residuals (AttnRes) to route block summaries across depth. The model is trained from scratch at 5B and 14B scales. Proxy experiments at 256p select 25% softmax as a quality-efficiency knee. The 5B model achieves VBench Total 84.30 at 480x832x81 in 13.2s on an H100, and compiled DiT forwards are up to 3.2x faster than a matched full-softmax baseline at 720p/60s. A Sol-Engine deployment yields a further 3.58x speedup. The paper also reports mechanistic analyses (state effective rank, routing patterns) and a detailed training and evaluation protocol.
Significance. If the central claim holds, this is a significant result: it demonstrates that a mostly-linear hybrid attention can approach full-softmax video DiT quality at a fraction of the long-sequence cost, extending recent LLM hybrid-attention designs to bidirectional video diffusion with a from-scratch training recipe. The efficiency claims are carefully profiled with matched compiled kernels and a consistent no-AttnRes protocol, and the paper is unusually transparent about the proxy basis of the architecture selection and about the lack of a measured quality advantage for AttnRes. The explicit reporting of training stages, timestep-stratified validation, and evaluation protocols is a strength. The main weakness is the absence of a matched production-scale full-softmax control for quality, which leaves the central 'matches full-softmax' claim partially unsupported.
major comments (3)
- [§5.3.1, Fig. 4, Table 2] The central quality claim ('matches full-softmax video DiTs in quality') lacks a matched production-scale full-softmax control. The 3:1 ratio is locked in from a 256p/10K proxy (depth-28, width-3072) in which 50% softmax has the lowest loss (0.897 vs 0.905 for 25% and 0.945 for all-softmax). Production models differ in depth/width, resolution, data, and the multi-stage curriculum, and the VBench baselines in Table 2 are differently trained and at different resolutions (e.g., Wan 2.2 A14B is the official 720p score; the 5B is 480x832x81). The margin over Wan 2.2 is 0.07, within likely noise. If the ratio knee does not transfer, the central claim collapses. Please provide a same-recipe production-scale full-softmax or at least 50% softmax control, or explicitly qualify the claim to proxy-based selection.
- [§5.3.2, Table 3a] AttnRes, a named contribution, shows no measured quality benefit in the only controlled quality probe (0.4851 vs 0.4855; 17/20 buckets slightly favorable), and the text says 'we do not read a quality advantage.' The supporting evidence is rank recovery (~12%, same-checkpoint) and routing-mass analyses, which are representation-level and not linked to output quality. AttnRes adds +3.1% latency and +2.1% memory (Table 3d). Thus its inclusion is not justified by the evidence as a component of the efficiency-quality trade-off. Please demonstrate a downstream benefit (quality, convergence, or sampling) or present AttnRes as an optional mechanism and adjust the title/abstract accordingly.
- [§5.2, Appendix D] No quality evaluation is reported for the 14B configuration; Table 2 and Table 8 contain only 5B VBench results, while 14B appears only in latency profiles (§G.2, §G.3). The abstract states the model is 'instantiated at 5B and 14B scales under a unified architecture' with quality parity, but this is unsupported for 14B. Please report at least one 14B quality measurement (VBench or a controlled production-scale proxy) or explicitly scope the quality claim to the 5B model.
minor comments (4)
- [Table 2] The superscript markers (†, ⋆, ‡) after baseline names are defined only in Appendix D; please define them in the caption or near the table for readability.
- [§5.4 / Fig. 5] The speedup figures are quoted for a 'no-AttnRes' protocol, whereas the production quality results include AttnRes. State explicitly that the reported speedups exclude the +3.1% AttnRes latency overhead so readers do not conflate the two settings.
- [Fig. 1(b), §6] The '120x faster than Wan 2.2-A14B' headline should state in the caption that it includes the full Sol-Engine stack and the 40-step protocol; currently this is only implicit.
- [Conclusion] The self-acknowledged limitation that 'the longest-duration results are tensor-shape profiles' is important and should appear in the abstract or introduction to align expectations with the claim 'unlocking scalable long, high resolution video generation.'
Circularity Check
No significant circularity: architecture choices are empirically selected on proxy validation loss and the headline quality/efficiency numbers are measured against external benchmarks, not derived from the selected parameters.
full rationale
SANA-Video 2.0's central claims are empirical rather than derived, and I found no step where an output is equivalent to its input by construction. The 25% softmax ratio is chosen from a held-out proxy sweep (Sec. 5.3.1, Fig. 4) and explicitly re-swept for video rather than imported as an assumption: the paper states it 'confirm[s] the 25% anchor ratio as a quality–efficiency knee for the video regime by sweeping it from scratch rather than assuming the language-model value.' The ratio is then fixed and the final VBench and latency numbers are measured independently. This is model selection, not a fitted parameter renamed as a prediction. Similarly, AttnRes is adopted for cross-depth reuse, and the paper explicitly declines to claim a quality gain from its narrow loss margin ('We do not read a quality advantage from this narrow margin'), so the rank/routing analysis is mechanism evidence, not a circular quality proof. Self-citations to SANA-Video and Sol-Engine are prior system components and deployment tooling, not load-bearing proofs; no uniqueness theorem or ansatz is smuggled in via self-citation. The admitted reliance on short 256p/10K-step proxy studies and tensor-shape profiles for long durations is a transfer-validity risk, not a circularity, and the paper flags it in its conclusion.
Axiom & Free-Parameter Ledger
free parameters (7)
- Softmax anchor ratio (3:1) =
25% softmax (8/32 layers in 5B, 10/40 in 14B)
- AttnRes block span S =
8 layers
- Token-count flow-shift endpoints =
shift 3 at 4,290 tokens; shift 6 at 23,000 tokens
- TQD bias magnitudes and thresholds =
±1.1 logit; UniMatch >30; DOVER >0.91
- ReFL reward weights =
4:4:1 HPSv3++:DeQA-Score:UniPercept
- Sampling guidance and flow-shift for VBench =
CFG 6.0/shift 6.0 (81f); CFG 8.0/shift 12.0 (121f/193f)
- Self-Flow distillation schedule =
student/teacher 9/25 (5B), weight 0.8, R_M=0.1
axioms (5)
- domain assumption VBench Total is a valid proxy for video generation quality.
- domain assumption Short reduced-resolution proxy studies transfer to production scales.
- domain assumption Linear-state effective rank is a meaningful measure of expressiveness.
- domain assumption Compiled best-kernel DiT-forward profiles represent realistic deployment speedups.
- domain assumption Mixed-source VBench baseline scores are comparable.
read the original abstract
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.
Reference graph
Works this paper leans on
-
[1]
Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026
Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026
Pith/arXiv arXiv 2026
-
[2]
Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré. Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff. InICML, 2024
2024
-
[3]
Bernini: Latent Semantic Planning for Video Diffusion
Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, and Zehuan Yuan. Bernini: Latent Semantic Planning for Video Diffusion. arXiv preprint arXiv:2605.22344, 2026
Pith/arXiv arXiv 2026
-
[4]
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
Shuo Cao, Jiayang Li, Xiaohui Li, Yuandong Pu, Kaiwen Zhu, Yuanting Gao, Siqi Luo, Yi Xin, Qi Qin, Yu Zhou, Xiangyu Chen, Wenlong Zhang, Bin Fu, Yu Qiao, and Yihao Liu. UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture. InICML, 2026
2026
-
[5]
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, and Robin Rombach. Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis. arXiv preprint arXiv:2603.06507, 2026
arXiv 2026
-
[6]
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu, Junyu Chen, Shuai Yang, Xianbang Wang, Yicheng Pan, Daquan Zhou, Huan Ling, Haozhe Liu, Hongwei Yi, Hao Zhang, Muyang Li, Yukang Chen, Han Cai, Sanja Fidler, Ping Luo, Song Han, and Enze Xie. SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer. arXiv preprint arXiv:2509.24695, 2025
arXiv 2025
-
[7]
Breaking the Low-Rank Dilemma of Linear Attention
Qihang Fan, Huaibo Huang, and Ran He. Breaking the Low-Rank Dilemma of Linear Attention. InCVPR, 2025
2025
-
[8]
Gemma 2: Improving Open Language Models at a Practical Size
Gemma Team. Gemma 2: Improving Open Language Models at a Practical Size. arXiv preprint arXiv:2408.00118, 2024
Pith/arXiv arXiv 2024
-
[9]
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
Mohsen Ghafoorian, Denis Korzhenkov, and Amirhossein Habibian. Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer. arXiv preprint arXiv:2509.24899, 2025
arXiv 2025
-
[10]
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert Gu and Tri Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. InConference on Language Modeling (COLM), 2024
2024
-
[11]
LTX-Video: Realtime Video Latent Diffusion
Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, et al. LTX-Video: Realtime Video Latent Diffusion. arXiv preprint arXiv:2501.00103, 2025
Pith/arXiv arXiv 2025
-
[12]
VBench: Comprehensive Benchmark Suite for Video Generative Models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive Benchmark Suite for Video Generative Models. InCVPR, 2024
2024
-
[13]
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. InICML, 2020
2020
-
[14]
Kimi Linear: An Expressive, Efficient Attention Architecture
Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv preprint arXiv:2510.26692, 2025
Pith/arXiv arXiv 2025
-
[15]
Kimi K3: Open Frontier Intelligence
Kimi Team. Kimi K3: Open Frontier Intelligence. Moonshot AI technical blog, https://www.kimi.com/blog/ kimi-k3, 2026
2026
-
[16]
Kimi Team. Attention Residuals. arXiv preprint arXiv:2603.15031, 2026
Pith/arXiv arXiv 2026
-
[17]
HunyuanVideo: A Systematic Framework For Large Video Generative Models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, et al. HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603, 2024
Pith/arXiv arXiv 2024
-
[18]
PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
Haopeng Li, Shitong Shao, Wenliang Zhong, Zikai Zhou, Lichen Bai, Hui Xiong, and Zeke Xie. PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers. arXiv preprint arXiv:2602.01077, 2026
arXiv 2026
-
[19]
VideoMamba: State Space Model for Efficient Video Understanding
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. VideoMamba: State Space Model for Efficient Video Understanding. InECCV, 2024. 14 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
2024
-
[20]
Radial Attention: 𝑂(𝑛log𝑛) Sparse Attention with Energy Decay for Long Video Generation
Xingyang Li et al. Radial Attention: 𝑂(𝑛log𝑛) Sparse Attention with Energy Decay for Long Video Generation. arXiv preprint arXiv:2506.19852, 2025
arXiv 2025
-
[21]
Yitong Li, Junsong Chen, Haopeng Li, Haozhe Liu, Jincheng Yu, Ligeng Zhu, Ping Luo, Song Han, and Enze Xie. Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation. arXiv preprint arXiv:2606.23743, 2026
Pith/arXiv arXiv 2026
-
[22]
Toward A Prac- tical Perceptual Video Quality Metric
Zhi Li, Anne Aaron, Ioannis Katsavounidis, Anush Moorthy, and Megha Manohara. Toward A Prac- tical Perceptual Video Quality Metric. Netflix Technology Blog, https://netflixtechblog.com/ toward-a-practical-perceptual-video-quality-metric-653f208b9652, 2016
2016
-
[23]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow Matching for Generative Modeling. InICLR, 2023
2023
-
[24]
HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities
Yijun Liu, Jie Huang, Zeyue Xue, Yuming Li, Ruizhe He, Haoran Li, Shijia Ge, and Siming Fu. HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities. arXiv preprint arXiv:2606.14657, 2026
arXiv 2026
-
[25]
VMamba: Visual State Space Model
Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, Jianbin Jiao, and Yunfan Liu. VMamba: Visual State Space Model. InNeurIPS, 2024
2024
-
[26]
Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training
Xiangyang Luo, Qingyu Li, Yuming Li, Guanbo Huang, Yongjie Zhu, Wenyu Qin, Meng Wang, Pengfei Wan, and Shao-Lun Huang. Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training. InCVPR, 2026
2026
-
[27]
Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang, et al. Scaling mixture-of-experts video pretraining for embodied intelligence.arXiv preprint arXiv:2607.07675, 2026
Pith/arXiv arXiv 2026
-
[28]
Scalable Diffusion Models with Transformers
William Peebles and Saining Xie. Scalable Diffusion Models with Transformers. InICCV, 2023
2023
-
[29]
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, and Junyang Lin. Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free. arXiv preprint arXiv:2505.06708, 2025
Pith/arXiv arXiv 2025
-
[30]
Qwen3-Next: Towards Ultimate Training & Inference Efficiency
Qwen Team. Qwen3-Next: Towards Ultimate Training & Inference Efficiency. Qwen Team blog, https: //qwen.ai/blog?id=qwen3-next, 2025
2025
-
[31]
MAGI-1: Autoregressive Video Generation at Scale
Sand.ai, Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, et al. MAGI-1: Autoregressive Video Generation at Scale. arXiv preprint arXiv:2505.13211, 2025
Pith/arXiv arXiv 2025
-
[32]
Seedance 2.0: Advancing Video Generation for World Complexity
Team Seedance et al. Seedance 2.0: Advancing Video Generation for World Complexity. arXiv preprint arXiv:2604.14148, 2026
Pith/arXiv arXiv 2026
-
[33]
TransNet V2: An Effective Deep Network Architecture for Fast Shot Transition Detection
Tomáš Souˇcek and Jakub Lokoˇc. TransNet V2: An Effective Deep Network Architecture for Fast Shot Transition Detection. arXiv preprint arXiv:2008.04838, 2020
Pith/arXiv arXiv 2008
-
[34]
RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024
2024
-
[35]
DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis
Yao Teng, Yue Wu, Han Shi, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu. DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis. arXiv preprint arXiv:2405.14224, 2024
Pith/arXiv arXiv 2024
-
[36]
Diffusion Model Alignment Using Direct Preference Optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion Model Alignment Using Direct Preference Optimization. InCVPR, 2024
2024
-
[37]
Wan2.2: Open and Advanced Large-Scale Video Generative Models
Wan Team. Wan2.2: Open and Advanced Large-Scale Video Generative Models. Official code and model release, https://github.com/Wan-Video/Wan2.2, 2025
2025
-
[38]
Wan: Open and Advanced Large-Scale Video Generative Models
Wan Team, Ang Wang, Baole Ai, Bin Wen, et al. Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314, 2025. 15 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
Pith/arXiv arXiv 2025
-
[39]
Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives
Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives. InICCV, 2023
2023
-
[40]
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
Haocheng Xi et al. Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity. InICML, 2025. arXiv:2502.01776
Pith/arXiv arXiv 2025
-
[41]
Unifying Flow, Stereo and Depth Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11): 13941–13958, 2023
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying Flow, Stereo and Depth Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11): 13941–13958, 2023
2023
-
[42]
ImageRe- ward: Learning and Evaluating Human Preferences for Text-to-Image Generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. ImageRe- ward: Learning and Evaluating Human Preferences for Text-to-Image Generation. InNeurIPS, 2023
2023
-
[43]
Gated Linear Attention Transformers with Hardware-Efficient Training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated Linear Attention Transformers with Hardware-Efficient Training. InICML, 2024
2024
-
[44]
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, et al. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. InICLR, 2025
2025
-
[45]
Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution
Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, and Chao Dong. Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution. InCVPR, 2025
2025
-
[46]
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Jingyang Yuan et al. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. arXiv preprint arXiv:2502.11089, 2025
Pith/arXiv arXiv 2025
-
[47]
Sigmoid Loss for Language Image Pre-Training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid Loss for Language Image Pre-Training. InICCV, 2023
2023
-
[48]
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
Jintao Zhang et al. SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference. InICML, 2025. arXiv:2502.18137
arXiv 2025
-
[49]
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu, Junsong Chen, Tian Ye, Haozhe Liu, Enze Xie, and Song Han. SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer. arXiv preprint arXiv:2605.30409, 2026
Pith/arXiv arXiv 2026
-
[50]
Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k
Zangwei Zheng, Xiangyu Peng, Chenhui Shen, et al. Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k. arXiv preprint arXiv:2503.09642, 2025
Pith/arXiv arXiv 2025
-
[51]
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, and Enze Xie. SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer. arXiv preprint arXiv:2605.15178, 2026
Pith/arXiv arXiv 2026
-
[52]
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
Lianghui Zhu, Zilong Huang, Bencheng Liao, Jun Hao Liew, Hanshu Yan, Jiashi Feng, and Xinggang Wang. DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention. InCVPR, 2025. 16 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation A. Related Work A.1. Video Diffusion and Efficient Sequence Modeling ...
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.