{"paper":{"title":"Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Pyramid Forcing assigns different KV cache lengths to three attention head types to reduce error accumulation in long autoregressive video generation.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Guojie Luo, Jiawei Yang, Jiayi Luo, Jiayu Chen, Junbei Tang, Maoliang Li, Wenbiao Zhao, Xiang Chen, Zihao Zheng","submitted_at":"2026-05-13T07:23:02Z","abstract_excerpt":"Autoregressive video generation enables streaming and open-ended long video synthesis, but still suffers from long-term degradation caused by accumulated errors. Existing KVCache strategies usually apply unified historical-frame retention, implicitly assuming homogeneous historical dependencies across attention heads. We revisit historical-frame attention and reveal three distinct head types: Anchor Heads require broad long-range context, Wave Heads exhibit periodic temporal dependencies, and Veil Heads focus on initial and adjacent frames. Based on this finding, we propose Pyramid Forcing, a "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Pyramid Forcing consistently improves long-horizon generation quality on VBench-Long, increasing the 60-second Self Forcing score from 77.87 to 81.21 while enhancing motion dynamics, visual fidelity, and semantic consistency.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The three head types (Anchor, Wave, Veil) are stable across different models and datasets and can be reliably identified offline without retraining or runtime overhead.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Pyramid Forcing classifies attention heads into Anchor, Wave, and Veil types and applies type-specific KV cache policies to improve long-horizon autoregressive video generation quality.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Pyramid Forcing assigns different KV cache lengths to three attention head types to reduce error accumulation in long autoregressive video generation.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"0034745f686999cf8d3a57fd914a790d63029398a276cab7a25fe2884107ad44"},"source":{"id":"2605.13111","kind":"arxiv","version":1},"verdict":{"id":"a47233f5-e04b-434f-bff1-a76a4301e02d","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-14T19:45:28.528497Z","strongest_claim":"Pyramid Forcing consistently improves long-horizon generation quality on VBench-Long, increasing the 60-second Self Forcing score from 77.87 to 81.21 while enhancing motion dynamics, visual fidelity, and semantic consistency.","one_line_summary":"Pyramid Forcing classifies attention heads into Anchor, Wave, and Veil types and applies type-specific KV cache policies to improve long-horizon autoregressive video generation quality.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The three head types (Anchor, Wave, Veil) are stable across different models and datasets and can be reliably identified offline without retraining or runtime overhead.","pith_extraction_headline":"Pyramid Forcing assigns different KV cache lengths to three attention head types to reduce error accumulation in long autoregressive video generation."},"references":{"count":37,"sample":[{"doi":"","year":2025,"title":"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion","work_id":"53e58ef9-7932-4b83-b757-34ac14db3e0f","ref_index":1,"cited_arxiv_id":"2506.08009","is_internal_anchor":true},{"doi":"","year":2025,"title":"MAGI-1: Autoregressive Video Generation at Scale","work_id":"25e8bd3d-e51c-43ae-8126-4ea6ecdb3321","ref_index":2,"cited_arxiv_id":"2505.13211","is_internal_anchor":true},{"doi":"","year":2026,"title":"Causal forcing: Autoregressive diffusion distillation done right for high-quality real- time interactive video generation","work_id":"04f67f5b-e79a-4ad9-8e90-763e5e54bd3d","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2023,"title":"Scalable diffusion models with transformers","work_id":"3e203719-f1d3-4517-959c-98e6121f3e23","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2025,"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","ref_index":5,"cited_arxiv_id":"2503.20314","is_internal_anchor":true}],"resolved_work":37,"snapshot_sha256":"0a92017860eae31d2d678829bc712e5e95c682d41cfe1ca05185fbb2b7ccdfd3","internal_anchors":10},"formal_canon":{"evidence_count":2,"snapshot_sha256":"cdf72e11bb671562156e1ed3f09472f4931f15d2709f30fe7a0712cb7d21041a"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}