Pith. sign in

Paper Citation Record · LEDGER

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 11 inbound Pith citation observations for arXiv:2506.03065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03065 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:15:18.581949Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:50:03.947151Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:49:44.825271Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved42
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afe33d08-fe92-432b-ae1a-471921e8a8b4 · outbound

This paper cites Longformer: The Long-Document Transformer.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Longformer: The Long-Document Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.392823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.392823Z digest=sha256:15f356bb76a74e2993db491ec444dff0f644c7f9512a50276ff627467471e49e

Observation 11ee65f5-0dba-4961-9e6f-9d909588bd2a · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.494535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.494535Z digest=sha256:b640bbb1036dc8b5ebaec0302e765388063f3b3a6056bfafc61295378a34865b

Observation 4292088d-975d-4d55-9069-e1445ce8c2d5 · outbound

This paper cites Ld-pruner: Ef- ficient pruning of latent diffusion models using task-agnostic insights.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Ld-pruner: Ef- ficient pruning of latent diffusion models using task-agnostic insights

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.604960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.604960Z digest=sha256:0330e473bd656de1a3609e2d054a2aab036d01f2e58af67cbb3957bed219757d

Observation 2827ffec-b6e0-424d-9e2f-d85e9e39bd39 · outbound

This paper cites $\Delta$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers $\Delta$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.714375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.714375Z digest=sha256:e58cda616946576d637c5c565e31f1c0c0e6ab9652f1a8d33b1a15df6caf6bbc

Observation fb21d783-e0b0-42f4-bf03-c471abc5243d · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Generating Long Sequences with Sparse Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.783778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.783778Z digest=sha256:d91df1ab24c7db78bf7f93ada11b2b8275acc2477d914363a99ed00f998b2798

Observation 008a59f2-17ab-48b3-91c7-ab475280a28d · outbound

This paper cites Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.861189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.861189Z digest=sha256:42626f5c32b19ff47d37ad1444832c84b447c14650261e8220fc8749d6e25b56

Observation 80444ff7-37e0-495c-8dc2-99dce13b4358 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.919834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.919834Z digest=sha256:1d61f4704547b052d69f0f6f89c3e9291629785566411514197a612a6fe15254

Observation 57ae08fa-4ebf-4722-935b-c2d05eb92350 · outbound

This paper cites Structural pruning for diffusion models.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Structural pruning for diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:20.118445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:15.053622Z digest=sha256:743502655f99e9102bd7d2004f1144820d80ad995b3e1b85b05f1e2aec87600b

Observation 816893ad-16f6-4cfb-9959-e4b0df2676f7 · outbound

This paper cites Neighborhood attention transformer.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Neighborhood attention transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.131954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.131954Z digest=sha256:c171302231a32f84651062c93a1c37c83dfcd37526e70f8b870aca01e8b48062

Observation 2ec664b1-d70f-464c-b996-844981a4de2a · outbound

This paper cites Pre-Trained Video Generative Models as World Simulators.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Pre-Trained Video Generative Models as World Simulators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.192689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.192689Z digest=sha256:b1dfe3585e0bf24f38e7c2d5484b3fb4e2b9c63e9283be1446df83f91347e1b5

Observation fa396819-df35-4c58-9cf3-c4ebd7ea91d6 · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.291377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.291377Z digest=sha256:0209fc10f3596639ac32543ab83127bd258694c66e246b85acbd968183b498d0

Observation 5d05a333-f744-4c2a-b9d0-f5d926f63550 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:19.977505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:15.393406Z digest=sha256:50df67e46db34547bec2ef66cdf721481d3d69e3a8d77caa2a59bcbb6e48073a

Observation 01a894a9-e7c8-42f0-a2dc-423e4c1a28aa · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Vbench: Comprehensive benchmark suite for video generative models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.493451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.493451Z digest=sha256:ef071f052df2d966d4c9660dcbbfe71ba3d4d00b936500f9cde1565bb34b99b9

Observation 3a30a21f-b308-4d8a-a49f-d2c786c475e7 · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515, 2024.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.586946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.586946Z digest=sha256:e1751feef3c1acf50ceb99e580af8295586a949bf760f69acc25cd93b6dd7aa6

Observation 82fa8275-d35d-4172-8333-c53ecb4f40be · outbound

This paper cites Adaptive Caching for Faster Video Generation with Diffusion Transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Adaptive Caching for Faster Video Generation with Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.704247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.704247Z digest=sha256:cd14a2b0129177fdf18ffba73961f29e546326a43c0ec18037a79abe9ef69f41

Observation 31a7549c-480a-45fa-9364-fa2da046fab1 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.801169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.801169Z digest=sha256:42a95120bdec725c4f64c8f21154dd0852f666e4ef58426c46ab613ae3b8f14a

Observation 5bfbc421-98c1-467d-a47e-8c7bdcfe314a · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.864790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.864790Z digest=sha256:36ffa31e138c7ccc78aa7230c7d6ff963739177cffe2a74ab149fc5eb14505d6

Observation 9328eb4b-3f0a-4c56-bef5-d61349337019 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Open-Sora Plan: Open-Source Large Video Generation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.930372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.930372Z digest=sha256:0878f7d9db6510779223dbbf6a869c04233cc98abaa59ee8125a368a85e7a8af

Observation 55920021-a95f-42b0-80e6-661be7bc7920 · outbound

This paper cites Diffusion adversarial post-training for one-step video generation.arXiv preprint arXiv:2501.08316, 2025.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Diffusion adversarial post-training for one-step video generation.arXiv preprint arXiv:2501.08316, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.001509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.001509Z digest=sha256:ea7e340366bfdb0067f6086edfdeb0a42e8f807c194b94bd5851e8977309b773

Observation ecff04f6-bdca-4c98-b052-00923f6c1cef · outbound

This paper cites Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.056909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.056909Z digest=sha256:dba386edf2d1959c44947492f00b4f6e6c15ee18980ef138b952583223d3598f

Observation 819697d8-6b6f-453f-afef-8326d1911db9 · outbound

This paper cites CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.170658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.170658Z digest=sha256:8f2f180675fcb45e7331ab7b558deb262f416b5d7b9a1aee42bd7a4df8b6e53e

Observation 4619a122-399c-4073-8e37-cd9583cd1fac · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.222163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.222163Z digest=sha256:d271adfc9767a5dd813992d7f267cb620fef0fe03947050fc73cfe3b4b632b8c

Observation 8b0d94a6-e522-45d3-b7c0-b8109eec9798 · outbound

This paper cites FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.311954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.311954Z digest=sha256:0f8131c5bce0591fb7116b6b3760788fc77f4bed4853a297edf118cdcd3c48be

Observation ce59189e-b970-4f39-a222-fb9665cb49f9 · outbound

This paper cites Deepcache: Accelerating diffusion models for free.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Deepcache: Accelerating diffusion models for free

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.389068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.389068Z digest=sha256:cdd87ade5adec78ec2914c9493d4d85ee254e2efbe0729bc1f4edad8211ec6b9

Observation ec6d1b80-03d1-47a5-919d-6c805ff1082f · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.497799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.497799Z digest=sha256:59c18a8ff32fdf557a67dd5c4179db351000fdd881fd0fe679a81d342b780d12

Observation 357eba18-fac0-4ec5-8fb1-aff71f81e27e · outbound

This paper cites Scalable diffusion models with transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Scalable diffusion models with transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.586188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.586188Z digest=sha256:1c2d7519b4060a1efc7ba828fb40c113204f20a4ca5e7b2c093c36558d82d687

Observation a617c6d5-21a8-4ede-bfda-125ba76e3be3 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers High- resolution image synthesis with latent diffusion models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.680560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.680560Z digest=sha256:828d048803c6de97344d5cd3d9a849926463a09996cad1da54ce348a6c9582ff

Observation 4d3c533e-c998-4540-ac8f-a6575e1435fc · outbound

This paper cites Post-training quantization on diffusion models.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Post-training quantization on diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:19.773754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:16.734593Z digest=sha256:47842b6ca79bc28aae6d81b037719416f9fb74918822792110310c0264cd1bc9

Observation 66961187-1e0a-4f09-9bc2-81c586e7b42e · outbound

This paper cites MD-dit: Step-aware mixture-of-depths for efficient diffusion transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers MD-dit: Step-aware mixture-of-depths for efficient diffusion transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:19.657568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:16.828763Z digest=sha256:d8e44a306a03579f643e65b2d97c52b2602e3005bc0a9e410e676768ff95b3d0

Observation a1d97e77-f827-405d-92dd-6aaf73342a61 · outbound

This paper cites REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.889183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.889183Z digest=sha256:38d5bd8b4c8b0a79d699b2596d071c3012783271d1458a7e65031f93636191cd

Observation 65cb2c5f-47ee-4a1e-954e-87c25a2dba23 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Triton: an intermediate language and compiler for tiled neural network computations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:16.969476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:16.969476Z digest=sha256:4fcf126e72b8b345ec722aa4b68c35dd932300c4f504fe1c415773a7b033c6dc

Observation 0b95db35-90b3-4bd5-8322-65f8cd92703d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.100784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.100784Z digest=sha256:16519d0ebe75cc62b2e0e78cd1ffd99a7f4b78d6a9b5d5df40433569bfe63efe

Observation b5f83350-da94-45a8-ae81-389529d36905 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.174838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.174838Z digest=sha256:fd6509ab7bf1b1ba1d11514adf6205a4a0432f2031c29c9554eb56ccfa9e510b

Observation 9508745a-2707-407d-bf3c-c7cb80e309aa · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Wan: Open and Advanced Large-Scale Video Generative Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.312148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.312148Z digest=sha256:048e752cabc2e23f9fedd038b4807552aa8b90e6b1270eba625fba62f6a63976

Observation 01a36b3c-7a2b-4268-b0cd-e644168d0c1a · outbound

This paper cites Taming Rectified Flow for Inversion and Editing.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Taming Rectified Flow for Inversion and Editing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.421351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.421351Z digest=sha256:3977b436e134bd59010515b9f2a828c390c4135d420eb5d16c77be3ac05dac0d

Observation f901e63f-848b-410c-aa41-fcdc43698073 · outbound

This paper cites A universal image quality index.IEEE signal processing letters, 9(3):81–84, 2002.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers A universal image quality index.IEEE signal processing letters, 9(3):81–84, 2002

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:19.454645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:17.526500Z digest=sha256:b82dd05c26d915175e089beae7132f365427be55c936422f2594ea43c6dc0d52

Observation 48ef1737-aea0-4d16-8b0f-7eff42736ed0 · outbound

This paper cites PTQ4DiT: Post-training Quantization for Diffusion Transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers PTQ4DiT: Post-training Quantization for Diffusion Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.600341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.600341Z digest=sha256:a2d2ecf37bca9469021231fedee1bbfe7c46223854598a9e3eb46f27b11ece3d

Observation e7e096eb-6989-47c4-b97d-ae330f0af6cf · outbound

This paper cites Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.692975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.692975Z digest=sha256:b9e288383f5d4613a434b654125f6aec8670f7470ada0cbe5cf3b977fdd00038

Observation 0a2310c5-0489-47ab-9507-9fb286ca6155 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.818811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.818811Z digest=sha256:8924fae6dd131a667f7171a8b26c06a3560d55b175427a9da6ebfcb1c5dba2ae

Observation 89422a6f-e3c9-4110-8a11-52789c4060a8 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Efficient Streaming Language Models with Attention Sinks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.899959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.899959Z digest=sha256:fedea1546f215c8b0b9ea10d669845fd3d4490730b407a38476d8f3e73731f40

Observation 61c2594e-de15-4db9-9916-604867f466f5 · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:17.947369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:17.947369Z digest=sha256:567d356ec5de125d240e63029c4f86867ce9f16ab05d9d769fdcfc0a6fce67ce

Observation d8ee225c-35fb-4915-afcb-57e13eb2f44d · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:18.000721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:18.000721Z digest=sha256:9d7bc401d67934e647096cbd1a55e3c6a39dbdd3fcce4489ae837dadedfdf733

Observation b69334d9-a790-4463-ab78-c6d493652f41 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:18.074114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:18.074114Z digest=sha256:2d51562efe30f68425bd838ce12bbb35b85103b809435d505ffa6fa0ddd482b4

Observation 6d5b2940-0fed-4dcc-b58a-aab97b4b341c · outbound

This paper cites Ditfastattn: Attention compression for diffusion transformer models.Advances in Neural Information Processing Systems, 37:1196–1219, 2024.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Ditfastattn: Attention compression for diffusion transformer models.Advances in Neural Information Processing Systems, 37:1196–1219, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:19.301120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:18.137279Z digest=sha256:95fcc03146766f1ad62633123fa51498da177cf396fc53ff544e9a35539eb85e

Observation 638d182d-b4c7-476b-88c6-873c9e37be02 · outbound

This paper cites Motion consistency model: Accelerating video diffusion with disentangled motion-appearance distillation.Advances in Neural Information Processing Systems, 37:111000–111021, 2024.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Motion consistency model: Accelerating video diffusion with disentangled motion-appearance distillation.Advances in Neural Information Processing Systems, 37:111000–111021, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:19.171343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:18.198285Z digest=sha256:ae901c3a3645be1c102a62e2a2780049dc33e72cbbc0c95da3f89c2f44bbe111

Observation 5daa6921-079d-4f24-943a-1ee2dcc981b3 · outbound

This paper cites InstructVEdit: A Holistic Approach for Instructional Video Editing.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers InstructVEdit: A Holistic Approach for Instructional Video Editing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:18.258170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:18.258170Z digest=sha256:9a6595f87d38942e8a43cd0b90fc09a6dd9739d8fd4e2e7fc7433cd446bb110a

Observation 55d919ff-f73d-4561-87b5-7e1836b9f89d · outbound

This paper cites DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:18.309915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:18.309915Z digest=sha256:edbf481ed7cb781b5001e78fc51261c0e11288262db604489a4c435802387c83

Observation 7c4bc236-ea8a-4fbf-8cc6-e223d5ca6900 · outbound

This paper cites The unrea- sonable effectiveness of deep features as a perceptual metric.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers The unrea- sonable effectiveness of deep features as a perceptual metric

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:18.375024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:18.375024Z digest=sha256:bdbf441ca38955ca647efd4e080e6828dfb89120f62d1a5da82856a4096d2009

Observation 9e98e2f2-69ff-4ace-a17c-31c298d5ede7 · outbound

This paper cites Pioneering 4-bit fp quantization for diffusion models: Mixup-sign quantization and timestep-aware fine-tuning, 2025.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Pioneering 4-bit fp quantization for diffusion models: Mixup-sign quantization and timestep-aware fine-tuning, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:19.024778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:15:18.430630Z digest=sha256:48e4f3fa61875e750790a40c9b8d5881f5ac810a157253ef6786a5b043321c29

Observation 7785b33f-34ad-4b07-b343-b05afabecfbd · outbound

This paper cites Real-Time Video Generation with Pyramid Attention Broadcast.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Real-Time Video Generation with Pyramid Attention Broadcast

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:18.510621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:18.510621Z digest=sha256:4883d95da264c2e41b561dca07798bbfa87cece844cf0fc01f5771556f3e41f4

Observation 6d3c4f4e-e987-4af4-a5db-a2c5a4471d8f · outbound

This paper cites DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:15:18.581949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:18.581949Z digest=sha256:8334d7c9878e26235921f843d70dfdee9f20767871a48654340159a53b3c602a

Pith citing papers

Observation 5588daa0-1b4c-4c7b-a408-50218720e742 · inbound

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer cites this paper.

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:30:26.075671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:30:26.075671Z digest=sha256:d3c93be1a112ab3e5b3ff7060ebd6995d1266569b80d5fc8ceb9b78feaead89a

Observation 67a6f03d-8ad4-450c-a056-0638d32f6739 · inbound

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers cites this paper.

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T15:35:13.213540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:35:13.213540Z digest=sha256:ada708d7575d1620b609043ed2e585d104c55c7693207c45f9758e918366d715

Observation 1a628508-83c3-47b5-a8ac-a9d5b7d00b11 · inbound

Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers cites this paper.

Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T10:37:47.271073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:37:47.271073Z digest=sha256:2daca320a62064e575c763118209e8c7a19563793e1b92c9b26f7cae2f7a8cb3

Observation 72a290b8-0d04-49ef-85a3-f30d74098ea3 · inbound

Ride the Wave: Precision-Allocated Sparse Attention for Smooth Video Generation cites this paper.

Ride the Wave: Precision-Allocated Sparse Attention for Smooth Video Generation Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:02.769647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:20:29.350915Z digest=sha256:ceb6938780975778898bc7acd103fb157ac1db6d4c8442d0b627cc8746a6ab41

Observation fcf2776e-34b9-4b9b-96c2-553a28bf05a7 · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:25.984657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:e801de25a0e8b345938a746c07b184697cff35b3cc05d960a7942dfe25bdc1c5

Observation c509e37d-a7a3-44e1-bc6f-91d1390c36b5 · inbound

HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention cites this paper.

HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:19:46.210669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:17:48.403406Z digest=sha256:4df672c213876b2dc59956e4a1fa6a8f0c555273d9a19d6116e18984c17919b2

Observation 1f7c32e5-9011-4239-949e-e4e86e825e2c · inbound

Veda: Scalable Video Diffusion via Distilled Sparse Attention cites this paper.

Veda: Scalable Video Diffusion via Distilled Sparse Attention Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.729248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:54:13.390911Z digest=sha256:f91a3d6d8d9e8ae51a8913ffc741b00f3bd5797ff3716c06b3859f2cd01dcd5c

Observation 53268387-2a8c-422e-81a0-a77bba15a33d · inbound

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion cites this paper.

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.715737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:14:23.336848Z digest=sha256:9b36f8f7e417c6095063ccc5d638c1e4ff2c2cbcd614f1d9d685699e58bbcd71

Observation 3cb0ef35-9ca6-45fa-9632-d9a643d66168 · inbound

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization cites this paper.

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.913950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T13:07:53.001390Z digest=sha256:34d71d4e8d2a287d751647c806b9ba4b11f03cc4357b7f56c623183299485967

Observation 2250f170-1136-4b46-9d7c-fa8567d37eb4 · inbound

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers cites this paper.

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:49:44.826596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:22:55.731241Z digest=sha256:4094c5bc58d3ce9e95ac69cd4c786fd3219c840faf5446a0493d52254dd4f6ed

Observation 5468d60f-563f-43f9-a3f7-aeeccec62463 · inbound

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference cites this paper.

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:03.947151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:50:03.947151Z digest=sha256:60d069ec8afa89ede4a15cd84f6ce910f172f228954acabd381172be99762167