{"total":14,"items":[{"citing_arxiv_id":"2607.06631","ref_index":37,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation","primary_cat":"cs.CV","submitted_at":"2026-07-07T13:14:33+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.5,"formal_verification":"none","one_line_summary":"Joint few-step distillation and step-specific structural pruning turns a video diffusion model into a compact Mixture-of-Models that cuts 24% extra FLOPs per step and reaches 30× speedup on Wan-14B.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.25473","ref_index":76,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models","primary_cat":"cs.CV","submitted_at":"2026-06-24T06:58:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Causal-rCM unifies teacher-forcing and self-forcing distillation for autoregressive video diffusion, delivering a 2-step model with VBench-T2V score 84.63 and enabling interactive world models on Cosmos 3 using only synthetic data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.30409","ref_index":42,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer","primary_cat":"cs.CV","submitted_at":"2026-05-28T17:59:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SANA-Streaming delivers 1280x704 streaming video editing at 24 FPS end-to-end on an RTX 5090 using hybrid DiT blocks, cycle-reverse training, and mixed-precision quantization.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.28691","ref_index":36,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning","primary_cat":"cs.CV","submitted_at":"2026-05-27T16:19:45+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"OSP-Next reports 83.73% VBench score and up to 2.27x speedup via hybrid sparse attention, SSP parallelism, HiF8 quantization, and Mix-GRPO on diffusion transformers.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.26632","ref_index":74,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models","primary_cat":"cs.LG","submitted_at":"2026-05-26T07:09:49+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.14513","ref_index":45,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion","primary_cat":"cs.CV","submitted_at":"2026-05-14T07:57:55+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.10198","ref_index":42,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Empty SPACE: Cross-Attention Sparsity for Concept Erasure in Diffusion Models","primary_cat":"cs.LG","submitted_at":"2026-05-11T08:46:29+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"SPACE induces sparsity in cross-attention parameters via closed-form iterative updates to erase target concepts more effectively than dense baselines in large diffusion models.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Empty SPACE: Cross-Attention Sparsity for Concept Erasure in Diffusion Models improve inference efficiency or noise robustness by sparsifying the attention activation maps. SPACE differentiates itself by targeting parameter sparsity rather than activation sparsity. Furthermore, while some computer vision architectures employ parameter-level attention sparsity [42, 43], they use standard backpropagation. SPACE is distinct as it achieves the sparsity of the cross-attention parameters without relying on backpropagation. Unlike the methods analyzed in this paragraph, SPACE focuses on concept erasure. Sparsity-based machine unlearningThe intersection of sparsity and machine unlearning is an emerging topic. Some"},{"citing_arxiv_id":"2605.04956","ref_index":12,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels","primary_cat":"cs.LG","submitted_at":"2026-05-06T14:18:36+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"KernelBenchX benchmark shows task category explains nearly three times more variance in LLM kernel correctness than method choice, iterative refinement boosts correctness but reduces performance, and quantization remains unsolved.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.22575","ref_index":39,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference","primary_cat":"cs.LG","submitted_at":"2026-04-24T14:07:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"SpikingBrain2.0 is a 5B hybrid spiking-Transformer that recovers most base model performance while delivering 10x TTFT speedup at 4M context and supporting over 10M tokens on limited GPUs via dual sparse attention and dual quantization paths.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.17397","ref_index":20,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Speculative Decoding for Autoregressive Video Generation","primary_cat":"cs.CV","submitted_at":"2026-04-19T12:01:57+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"A training-free speculative decoding method for block-based autoregressive video diffusion uses a quality router on worst-frame ImageReward scores to accept drafter proposals, achieving up to 2.09x speedup at 95.7% quality retention.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.15911","ref_index":176,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Efficient Video Diffusion Models: Advancements and Challenges","primary_cat":"cs.CV","submitted_at":"2026-04-17T10:11:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"A survey that groups efficient video diffusion methods into four paradigms—step distillation, efficient attention, model compression, and cache/trajectory optimization—and outlines open challenges for practical use.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Publication date: April 2026. Efficient Video Diffusion Models: Advancements and Challenges•35 [175] Lvmin Zhang, Shengqu Cai, Muyang Li, Chong Zeng, Beijia Lu, Anyi Rao, Song Han, Gordon Wetzstein, and Maneesh Agrawala. 2026. Pretraining Frame Preservation in Autoregressive Video Memory Compression. arXiv:2512.23851 https://arxiv.org/abs/2512.23851 [176] Liuzhou Zhang, Jiarui Ye, Yuanlei Wang, Ming Zhong, Mingju Cao, Wanke Xia, Bowen Zeng, Zeyu Zhang, and Hao Tang. 2025. EgoLCD: Egocentric Video Generation with Long Context Diffusion. arXiv:2512.04515 https://arxiv.org/abs/2512.04515 [177] Peiyuan Zhang, Yongqi Chen, Haofeng Huang, Will Lin, Zhengzhong Liu, Ion Stoica, Eric Xing, and Hao Zhang. 2025."},{"citing_arxiv_id":"2604.12219","ref_index":26,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Ride the Wave: Precision-Allocated Sparse Attention for Smooth Video Generation","primary_cat":"cs.CV","submitted_at":"2026-04-14T02:51:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"PASA uses curvature-aware dynamic budgeting, grouped approximations, and stochastic attention routing to accelerate video diffusion transformers while eliminating temporal flickering from sparse patterns.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"For example, SLA [25] stratifies attention weights into crit- ical, marginal, and negligible tiers. It allocates dense computation to critical weights, appliesO (𝑁) linear attention to marginal ones, and prunes the negligible connections entirely. Through lightweight fine-tuning, SLA drastically amplifies sparsity while preserving generation fidelity. Its successor, SLA2 [26], further advances this by introducing a differentiable routing mechanism. Concurrently, to sustain end-to-end performance under ultra-sparse conditions, SpargeAttention2 [24] proposes a hybrid top-𝑘 and top-𝑝 dynamic masking criterion. This is coupled with a velocity-level distillation fine-tuning objective, which utilizes a frozen full-attention model"},{"citing_arxiv_id":"2604.10103","ref_index":56,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation","primary_cat":"cs.CV","submitted_at":"2026-04-11T08:54:07+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Hybrid Forcing combines linear temporal attention for long-range retention, block-sparse attention for efficiency, and decoupled distillation to achieve real-time unbounded 832x480 streaming video generation at 29.5 FPS.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Errors introduced in earlier frames accumulate over time, leading to degraded temporal coherence and unstable long-form generation. Sparse Attention:Recent works have explored the sparsity in attention maps for video generation [47,48,56,58]. Leveraging this property, Sparse Video Gen- eration [47] methods define a fixed sparse attention mask to reduce the computa- tional cost of attention. AdaSpa [56] introduces a blockified pattern to efficiently capture the hierarchical sparsity in DiTs, significantly reducing attention com- Fast Stream Video Generation with Hybrid Attention 5 plexity. VSA [58] adapts the DeepSeek NSA [55] to video DiTs, finetuning the model with newly introduced VSA modules to achieve better performance com- pared to training-free methods."},{"citing_arxiv_id":"2604.04451","ref_index":10,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse","primary_cat":"cs.CV","submitted_at":"2026-04-06T05:55:13+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Chorus accelerates video DiT serving up to 45% via inter-request caching reuse in a three-stage denoising strategy with token-guided attention amplification.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}