Pith. sign in

REVIEW 3 major objections 4 minor 52 references

A mostly-linear attention mix can match full-softmax video DiTs in quality while running at linear-attention speed—this paper shows how, at 5B and 14B scales.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:05 UTC pith:WAE23BSP

load-bearing objection A well-executed, honest systems paper on hybrid linear/softmax attention for video DiTs; the efficiency story holds up, but the 'matches full-softmax' quality claim rests on a small proxy and needs a production-scale check. the 3 major comments →

arxiv 2607.21553 v1 pith:WAE23BSP submitted 2026-07-23 cs.CV

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

classification cs.CV
keywords video diffusion transformerhybrid linear attentionsoftmax anchorsattention residualsefficient video generationflow matchingVBenchlinear attention
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that a video diffusion transformer can keep full-softmax-quality generation while replacing most of its quadratic attention with linear attention. The core claim is that a 3:1 mix—75% gated linear attention for ordinary token mixing, plus periodic softmax layers that restore full-rank interactions—is a practical Pareto-optimal trade-off for video. To propagate the rank-restoring updates across depth, they add Block Attention Residuals that route summaries of completed blocks into later layers. The paper's 5B model reports VBench 84.30, on par with far larger softmax models, while its compiled forward pass is up to 3.2x faster at long 720p durations, and the advantage grows with sequence length.

Core claim

SANA-Video 2.0 is a scratch-trained hybrid-attention video diffusion transformer (5B and 14B) that matches full-softmax quality while keeping the O(N) scaling of linear attention. The architecture interleaves gated linear attention with periodic softmax anchors at a 3:1 ratio (25% softmax), and Block Attention Residuals route completed block summaries into later linear layers. Proxy studies at 256p found 25% softmax to be the best quality-efficiency knee, not the lowest-loss point (50% was lower loss but much slower). The 5B model reaches VBench Total 84.30 at 480x832x81 in 13.2s on one H100; its compiled DiT forward is 3.2x faster than full softmax at 720p/60s, a gap that widens with durati

What carries the argument

Hybrid Linear-Softmax Attention: gated linear attention for O(N)-dominated mixing, with periodic gated-softmax anchors every fourth layer. Block Attention Residuals (AttnRes): routes the completed block summaries and the current partial sum into later layers via a depth-shared learned query per branch, exposing the softmax-refreshed representations to the linear majority.

Load-bearing premise

Decisions made from short proxy runs at 256p, such as the 25% softmax ratio and the AttnRes block span, transfer to the much larger 5B and 14B production models at high resolution and long duration.

What would settle it

Run a production-scale ablation that varies the softmax ratio and AttnRes on/off, training long enough to produce a final checkpoint, and compare VBench or similar quality at matched latency. If 25% softmax or AttnRes does not yield comparable quality to full softmax at that scale, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Long-video generation becomes much cheaper at a given quality: the efficiency advantage grows with duration. With a mostly-linear backbone, the cost of generating 60s 720p clips is substantially reduced relative to full-softmax models at the same scale.
  • Scaling to longer horizons is more feasible: the O(N) backbone shifts the bottleneck away from the token-mixing cost, allowing training and inference to extend beyond current 8s horizons.
  • Hardware-friendly backbones combine with deployment optimization: the conv-free SwiGLU FFN and fixed anchors map directly to fused kernels and sparse attention, yielding a measured 3.58x end-to-end speedup on B200.
  • Low-precision quantization is possible without quality loss: QAT with MXFP4 weights and MXFP8 activations matches BF16 VBench scores, cutting static storage by 68%.
  • The design transfers to physical AI: fine-tuning on robot and egocentric video produces realistic manipulation clips competitive with models several times larger.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If hybrid attention is adopted broadly in video DiTs, quality-vs-cost trade-offs could shift for the whole field: the paper suggests a recipe to make video generation widely accessible on single GPUs, potentially democratizing long-form video creation.
  • The same hybrid-plus-AttnRes recipe might be carried into causal generation: the linear operator drops the delta-rule update, making it a natural seed for a causal Gated DeltaNet, which could bring the efficiency and quality to streaming and world-model settings.
  • The proxy-study methodology (short runs at 256p to select architecture) implies a testable extension: check whether the 25% softmax ratio and S=8 block span remain optimal at production scales and much longer sequences, or whether the Pareto knee shifts toward more softmax.
  • The authors' observation that AttnRes adds only a small quality edge but boosts deep-layer rank suggests a potential standalone diagnostic: measuring effective-rank recovery could become a general tool for evaluating whether 'expressiveness' of hybrid architectures is genuinely improving.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes SANA-Video 2.0, a video diffusion transformer that replaces most softmax attention with gated linear attention, inserts periodic softmax anchors at a 3:1 linear-to-softmax ratio, and adds Block Attention Residuals (AttnRes) to route block summaries across depth. The model is trained from scratch at 5B and 14B scales. Proxy experiments at 256p select 25% softmax as a quality-efficiency knee. The 5B model achieves VBench Total 84.30 at 480x832x81 in 13.2s on an H100, and compiled DiT forwards are up to 3.2x faster than a matched full-softmax baseline at 720p/60s. A Sol-Engine deployment yields a further 3.58x speedup. The paper also reports mechanistic analyses (state effective rank, routing patterns) and a detailed training and evaluation protocol.

Significance. If the central claim holds, this is a significant result: it demonstrates that a mostly-linear hybrid attention can approach full-softmax video DiT quality at a fraction of the long-sequence cost, extending recent LLM hybrid-attention designs to bidirectional video diffusion with a from-scratch training recipe. The efficiency claims are carefully profiled with matched compiled kernels and a consistent no-AttnRes protocol, and the paper is unusually transparent about the proxy basis of the architecture selection and about the lack of a measured quality advantage for AttnRes. The explicit reporting of training stages, timestep-stratified validation, and evaluation protocols is a strength. The main weakness is the absence of a matched production-scale full-softmax control for quality, which leaves the central 'matches full-softmax' claim partially unsupported.

major comments (3)
  1. [§5.3.1, Fig. 4, Table 2] The central quality claim ('matches full-softmax video DiTs in quality') lacks a matched production-scale full-softmax control. The 3:1 ratio is locked in from a 256p/10K proxy (depth-28, width-3072) in which 50% softmax has the lowest loss (0.897 vs 0.905 for 25% and 0.945 for all-softmax). Production models differ in depth/width, resolution, data, and the multi-stage curriculum, and the VBench baselines in Table 2 are differently trained and at different resolutions (e.g., Wan 2.2 A14B is the official 720p score; the 5B is 480x832x81). The margin over Wan 2.2 is 0.07, within likely noise. If the ratio knee does not transfer, the central claim collapses. Please provide a same-recipe production-scale full-softmax or at least 50% softmax control, or explicitly qualify the claim to proxy-based selection.
  2. [§5.3.2, Table 3a] AttnRes, a named contribution, shows no measured quality benefit in the only controlled quality probe (0.4851 vs 0.4855; 17/20 buckets slightly favorable), and the text says 'we do not read a quality advantage.' The supporting evidence is rank recovery (~12%, same-checkpoint) and routing-mass analyses, which are representation-level and not linked to output quality. AttnRes adds +3.1% latency and +2.1% memory (Table 3d). Thus its inclusion is not justified by the evidence as a component of the efficiency-quality trade-off. Please demonstrate a downstream benefit (quality, convergence, or sampling) or present AttnRes as an optional mechanism and adjust the title/abstract accordingly.
  3. [§5.2, Appendix D] No quality evaluation is reported for the 14B configuration; Table 2 and Table 8 contain only 5B VBench results, while 14B appears only in latency profiles (§G.2, §G.3). The abstract states the model is 'instantiated at 5B and 14B scales under a unified architecture' with quality parity, but this is unsupported for 14B. Please report at least one 14B quality measurement (VBench or a controlled production-scale proxy) or explicitly scope the quality claim to the 5B model.
minor comments (4)
  1. [Table 2] The superscript markers (†, ⋆, ‡) after baseline names are defined only in Appendix D; please define them in the caption or near the table for readability.
  2. [§5.4 / Fig. 5] The speedup figures are quoted for a 'no-AttnRes' protocol, whereas the production quality results include AttnRes. State explicitly that the reported speedups exclude the +3.1% AttnRes latency overhead so readers do not conflate the two settings.
  3. [Fig. 1(b), §6] The '120x faster than Wan 2.2-A14B' headline should state in the caption that it includes the full Sol-Engine stack and the 40-step protocol; currently this is only implicit.
  4. [Conclusion] The self-acknowledged limitation that 'the longest-duration results are tensor-shape profiles' is important and should appear in the abstract or introduction to align expectations with the claim 'unlocking scalable long, high resolution video generation.'

Circularity Check

0 steps flagged

No significant circularity: architecture choices are empirically selected on proxy validation loss and the headline quality/efficiency numbers are measured against external benchmarks, not derived from the selected parameters.

full rationale

SANA-Video 2.0's central claims are empirical rather than derived, and I found no step where an output is equivalent to its input by construction. The 25% softmax ratio is chosen from a held-out proxy sweep (Sec. 5.3.1, Fig. 4) and explicitly re-swept for video rather than imported as an assumption: the paper states it 'confirm[s] the 25% anchor ratio as a quality–efficiency knee for the video regime by sweeping it from scratch rather than assuming the language-model value.' The ratio is then fixed and the final VBench and latency numbers are measured independently. This is model selection, not a fitted parameter renamed as a prediction. Similarly, AttnRes is adopted for cross-depth reuse, and the paper explicitly declines to claim a quality gain from its narrow loss margin ('We do not read a quality advantage from this narrow margin'), so the rank/routing analysis is mechanism evidence, not a circular quality proof. Self-citations to SANA-Video and Sol-Engine are prior system components and deployment tooling, not load-bearing proofs; no uniqueness theorem or ansatz is smuggled in via self-citation. The admitted reliance on short 256p/10K-step proxy studies and tensor-shape profiles for long durations is a transfer-validity risk, not a circularity, and the paper flags it in its conclusion.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 0 invented entities

No new physical entities are introduced; the only new artifacts are architectural (AttnRes block summaries, shared-query router), which have no independent falsifiable handle outside this paper's own measurements. The central claim rests on several domain assumptions and hand-chosen hyperparameters, most importantly the proxy-to-production transfer of the 25% softmax ratio.

free parameters (7)
  • Softmax anchor ratio (3:1) = 25% softmax (8/32 layers in 5B, 10/40 in 14B)
    Selected from a 256p/81f proxy sweep as the Pareto knee; 50% has lower validation loss (0.897 vs 0.905) but higher latency, so 25% is a trade-off choice, not a minimum-loss point (Section 5.3.1).
  • AttnRes block span S = 8 layers
    Chosen as an engineering default because S in {4,8,16} give similar loss point estimates (0.962–0.965) and memory varies slightly; not an optimized value (Table 3d).
  • Token-count flow-shift endpoints = shift 3 at 4,290 tokens; shift 6 at 23,000 tokens
    Interpolated in log-shift space by latent token count; endpoints chosen by hand to match curriculum resolutions/durations (Section 4.2, Table 1).
  • TQD bias magnitudes and thresholds = ±1.1 logit; UniMatch >30; DOVER >0.91
    Content-aware timestep sampling biases from [26]; the paper sets magnitudes/thresholds without a reported sensitivity study (Appendix B.3).
  • ReFL reward weights = 4:4:1 HPSv3++:DeQA-Score:UniPercept
    Hand-chosen combination for online RL; no ablation or sensitivity analysis is reported (Section 4.4.2).
  • Sampling guidance and flow-shift for VBench = CFG 6.0/shift 6.0 (81f); CFG 8.0/shift 12.0 (121f/193f)
    Inference hyperparameters chosen per operating point; baseline models keep their defaults, so the comparison is not protocol-matched (Appendix D.1).
  • Self-Flow distillation schedule = student/teacher 9/25 (5B), weight 0.8, R_M=0.1
    Auxiliary training objective coefficients taken from prior work; not ablated here (Appendix B.2).
axioms (5)
  • domain assumption VBench Total is a valid proxy for video generation quality.
    Headline quality claims (84.30) are VBench scores compared against baselines with different protocols and no human evaluation; if VBench is not predictive of real quality, the 'matches full-softmax' claim weakens.
  • domain assumption Short reduced-resolution proxy studies transfer to production scales.
    The 25% softmax ratio, block span S=8, shared-query router, and timestep-free router are locked in from 256p, 10K-step, depth-28/width-3072 runs (Section 5.3.1, Table 3); no production-scale ablation of these choices is reported.
  • domain assumption Linear-state effective rank is a meaningful measure of expressiveness.
    AttnRes is justified primarily by an 11.7% deep-layer effective-rank increase and routing-mass analyses (Section 5.5, Figure 6), while the direct quality probe is a tie; the assumption connects rank to final quality.
  • domain assumption Compiled best-kernel DiT-forward profiles represent realistic deployment speedups.
    The headline 3.2x/3.58x speedups come from forward-only profiles with AttnRes disabled or from the Sol-Engine pipeline; end-to-end quality/latency at 60s duration is a tensor-shape profile, not a trained long-video deployment (Sections 5.4, 6, Conclusion).
  • domain assumption Mixed-source VBench baseline scores are comparable.
    Table 2 mixes official (star), author-reported (double-dagger), and measured (dagger) scores at different resolutions/frame counts; the paper's 81-frame result is set against 93–129-frame baselines, and score-source markers are the only adjustment.

pith-pipeline@v1.3.0-alltime-deepseek · 29246 in / 17518 out tokens · 155508 ms · 2026-08-01T07:05:16.499792+00:00 · methodology

0 comments
read the original abstract

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2x faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58x, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120x faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 20 linked inside Pith

  1. [1]

    Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026

    Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026

  2. [2]

    Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff

    Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré. Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff. InICML, 2024

  3. [3]

    Bernini: Latent Semantic Planning for Video Diffusion

    Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, and Zehuan Yuan. Bernini: Latent Semantic Planning for Video Diffusion. arXiv preprint arXiv:2605.22344, 2026

  4. [4]

    UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

    Shuo Cao, Jiayang Li, Xiaohui Li, Yuandong Pu, Kaiwen Zhu, Yuanting Gao, Siqi Luo, Yi Xin, Qi Qin, Yu Zhou, Xiangyu Chen, Wenlong Zhang, Bin Fu, Yu Qiao, and Yihao Liu. UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture. InICML, 2026

  5. [5]

    Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

    Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, and Robin Rombach. Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis. arXiv preprint arXiv:2603.06507, 2026

  6. [6]

    SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

    Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu, Junyu Chen, Shuai Yang, Xianbang Wang, Yicheng Pan, Daquan Zhou, Huan Ling, Haozhe Liu, Hongwei Yi, Hao Zhang, Muyang Li, Yukang Chen, Han Cai, Sanja Fidler, Ping Luo, Song Han, and Enze Xie. SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer. arXiv preprint arXiv:2509.24695, 2025

  7. [7]

    Breaking the Low-Rank Dilemma of Linear Attention

    Qihang Fan, Huaibo Huang, and Ran He. Breaking the Low-Rank Dilemma of Linear Attention. InCVPR, 2025

  8. [8]

    Gemma 2: Improving Open Language Models at a Practical Size

    Gemma Team. Gemma 2: Improving Open Language Models at a Practical Size. arXiv preprint arXiv:2408.00118, 2024

  9. [9]

    Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer

    Mohsen Ghafoorian, Denis Korzhenkov, and Amirhossein Habibian. Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer. arXiv preprint arXiv:2509.24899, 2025

  10. [10]

    Mamba: Linear-Time Sequence Modeling with Selective State Spaces

    Albert Gu and Tri Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. InConference on Language Modeling (COLM), 2024

  11. [11]

    LTX-Video: Realtime Video Latent Diffusion

    Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, et al. LTX-Video: Realtime Video Latent Diffusion. arXiv preprint arXiv:2501.00103, 2025

  12. [12]

    VBench: Comprehensive Benchmark Suite for Video Generative Models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive Benchmark Suite for Video Generative Models. InCVPR, 2024

  13. [13]

    Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

    Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. InICML, 2020

  14. [14]

    Kimi Linear: An Expressive, Efficient Attention Architecture

    Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv preprint arXiv:2510.26692, 2025

  15. [15]

    Kimi K3: Open Frontier Intelligence

    Kimi Team. Kimi K3: Open Frontier Intelligence. Moonshot AI technical blog, https://www.kimi.com/blog/ kimi-k3, 2026

  16. [16]

    Attention Residuals

    Kimi Team. Attention Residuals. arXiv preprint arXiv:2603.15031, 2026

  17. [17]

    HunyuanVideo: A Systematic Framework For Large Video Generative Models

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, et al. HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603, 2024

  18. [18]

    PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers

    Haopeng Li, Shitong Shao, Wenliang Zhong, Zikai Zhou, Lichen Bai, Hui Xiong, and Zeke Xie. PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers. arXiv preprint arXiv:2602.01077, 2026

  19. [19]

    VideoMamba: State Space Model for Efficient Video Understanding

    Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. VideoMamba: State Space Model for Efficient Video Understanding. InECCV, 2024. 14 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

  20. [20]

    Radial Attention: 𝑂(𝑛log𝑛) Sparse Attention with Energy Decay for Long Video Generation

    Xingyang Li et al. Radial Attention: 𝑂(𝑛log𝑛) Sparse Attention with Energy Decay for Long Video Generation. arXiv preprint arXiv:2506.19852, 2025

  21. [21]

    Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation

    Yitong Li, Junsong Chen, Haopeng Li, Haozhe Liu, Jincheng Yu, Ligeng Zhu, Ping Luo, Song Han, and Enze Xie. Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation. arXiv preprint arXiv:2606.23743, 2026

  22. [22]

    Toward A Prac- tical Perceptual Video Quality Metric

    Zhi Li, Anne Aaron, Ioannis Katsavounidis, Anush Moorthy, and Megha Manohara. Toward A Prac- tical Perceptual Video Quality Metric. Netflix Technology Blog, https://netflixtechblog.com/ toward-a-practical-perceptual-video-quality-metric-653f208b9652, 2016

  23. [23]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow Matching for Generative Modeling. InICLR, 2023

  24. [24]

    HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities

    Yijun Liu, Jie Huang, Zeyue Xue, Yuming Li, Ruizhe He, Haoran Li, Shijia Ge, and Siming Fu. HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities. arXiv preprint arXiv:2606.14657, 2026

  25. [25]

    VMamba: Visual State Space Model

    Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, Jianbin Jiao, and Yunfan Liu. VMamba: Visual State Space Model. InNeurIPS, 2024

  26. [26]

    Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training

    Xiangyang Luo, Qingyu Li, Yuming Li, Guanbo Huang, Yongjie Zhu, Wenyu Qin, Meng Wang, Pengfei Wan, and Shao-Lun Huang. Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training. InCVPR, 2026

  27. [27]

    Scaling mixture-of-experts video pretraining for embodied intelligence.arXiv preprint arXiv:2607.07675, 2026

    Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang, et al. Scaling mixture-of-experts video pretraining for embodied intelligence.arXiv preprint arXiv:2607.07675, 2026

  28. [28]

    Scalable Diffusion Models with Transformers

    William Peebles and Saining Xie. Scalable Diffusion Models with Transformers. InICCV, 2023

  29. [29]

    Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

    Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, and Junyang Lin. Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free. arXiv preprint arXiv:2505.06708, 2025

  30. [30]

    Qwen3-Next: Towards Ultimate Training & Inference Efficiency

    Qwen Team. Qwen3-Next: Towards Ultimate Training & Inference Efficiency. Qwen Team blog, https: //qwen.ai/blog?id=qwen3-next, 2025

  31. [31]

    MAGI-1: Autoregressive Video Generation at Scale

    Sand.ai, Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, et al. MAGI-1: Autoregressive Video Generation at Scale. arXiv preprint arXiv:2505.13211, 2025

  32. [32]

    Seedance 2.0: Advancing Video Generation for World Complexity

    Team Seedance et al. Seedance 2.0: Advancing Video Generation for World Complexity. arXiv preprint arXiv:2604.14148, 2026

  33. [33]

    TransNet V2: An Effective Deep Network Architecture for Fast Shot Transition Detection

    Tomáš Souˇcek and Jakub Lokoˇc. TransNet V2: An Effective Deep Network Architecture for Fast Shot Transition Detection. arXiv preprint arXiv:2008.04838, 2020

  34. [34]

    RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024

  35. [35]

    DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis

    Yao Teng, Yue Wu, Han Shi, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu. DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis. arXiv preprint arXiv:2405.14224, 2024

  36. [36]

    Diffusion Model Alignment Using Direct Preference Optimization

    Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion Model Alignment Using Direct Preference Optimization. InCVPR, 2024

  37. [37]

    Wan2.2: Open and Advanced Large-Scale Video Generative Models

    Wan Team. Wan2.2: Open and Advanced Large-Scale Video Generative Models. Official code and model release, https://github.com/Wan-Video/Wan2.2, 2025

  38. [38]

    Wan: Open and Advanced Large-Scale Video Generative Models

    Wan Team, Ang Wang, Baole Ai, Bin Wen, et al. Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314, 2025. 15 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

  39. [39]

    Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives

    Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives. InICCV, 2023

  40. [40]

    Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

    Haocheng Xi et al. Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity. InICML, 2025. arXiv:2502.01776

  41. [41]

    Unifying Flow, Stereo and Depth Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11): 13941–13958, 2023

    Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying Flow, Stereo and Depth Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11): 13941–13958, 2023

  42. [42]

    ImageRe- ward: Learning and Evaluating Human Preferences for Text-to-Image Generation

    Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. ImageRe- ward: Learning and Evaluating Human Preferences for Text-to-Image Generation. InNeurIPS, 2023

  43. [43]

    Gated Linear Attention Transformers with Hardware-Efficient Training

    Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated Linear Attention Transformers with Hardware-Efficient Training. InICML, 2024

  44. [44]

    CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, et al. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. InICLR, 2025

  45. [45]

    Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution

    Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, and Chao Dong. Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution. InCVPR, 2025

  46. [46]

    Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

    Jingyang Yuan et al. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. arXiv preprint arXiv:2502.11089, 2025

  47. [47]

    Sigmoid Loss for Language Image Pre-Training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid Loss for Language Image Pre-Training. InICCV, 2023

  48. [48]

    SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference

    Jintao Zhang et al. SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference. InICML, 2025. arXiv:2502.18137

  49. [49]

    SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

    Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu, Junsong Chen, Tian Ye, Haozhe Liu, Enze Xie, and Song Han. SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer. arXiv preprint arXiv:2605.30409, 2026

  50. [50]

    Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

    Zangwei Zheng, Xiangyu Peng, Chenhui Shen, et al. Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k. arXiv preprint arXiv:2503.09642, 2025

  51. [51]

    SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

    Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, and Enze Xie. SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer. arXiv preprint arXiv:2605.15178, 2026

  52. [52]

    DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention

    Lianghui Zhu, Zilong Huang, Bencheng Liao, Jun Hao Liew, Hanshu Yan, Jiashi Feng, and Xinggang Wang. DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention. InCVPR, 2025. 16 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation A. Related Work A.1. Video Diffusion and Efficient Sequence Modeling ...