Pith. sign in

Paper Citation Record · LEDGER

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

As of 1 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.21553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21553 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:09:10.589942Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b1e1ac5-9340-47c9-9240-c51e842ff138 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.165440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.165440Z digest=sha256:2dc1ae19db0c7f2d0cf6b2641e2e04f20f2874469079a52da61dff09be2638cc

Observation cbcb147a-5dfb-477b-ace1-7a7c0e394126 · outbound

This paper cites Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.210677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.210677Z digest=sha256:977ebc2de2af95ff2455c5984c3d3843b86d7ef0a48e555352631ce388bbcb2c

Observation 19d8d957-ddf4-4adb-a4e4-6a5fc2817c4b · outbound

This paper cites Bernini: Latent Semantic Planning for Video Diffusion.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Bernini: Latent Semantic Planning for Video Diffusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.295403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.295403Z digest=sha256:c7a114b46c313e3f424dcf36847615b03b05f5d7622631f56724dda55aada251

Observation 73899bf4-96e4-450c-b1a2-0d5e30f248c9 · outbound

This paper cites UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.408160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.408160Z digest=sha256:3c5c7784592a5a2bcf58df70c057dc28203a754a97834d3aa07aaeda5a6f411f

Observation d2c70687-a060-4975-9f38-cccde0324b59 · outbound

This paper cites Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.542402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.542402Z digest=sha256:84d0f38a1e52d51adffe8b4f2911bb1f19617f7ac6343c680e713377686d8354

Observation 082ea584-5be5-4446-9ef3-89016abe5e69 · outbound

This paper cites SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.684694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.684694Z digest=sha256:67343e5b183fd53339da4e9aba5e4401bc1132d4ccaaa3c624588a8d3c253dfc

Observation 9b61acad-4356-4f24-abae-18cbd0c4f276 · outbound

This paper cites Breaking the Low-Rank Dilemma of Linear Attention.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Breaking the Low-Rank Dilemma of Linear Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.828094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.828094Z digest=sha256:88cba1b819987f11ea47af2b7cf801c6b55213fa4db46b42f877f260cbe9d080

Observation ef5954d5-8d0e-4d5c-8c8d-8cea12cb2f78 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Gemma 2: Improving Open Language Models at a Practical Size

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.934288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.934288Z digest=sha256:a203c175f8a292267924942a4e92590979e28b129bcd3e5daa5c155bb6b8361a

Observation a052d4ae-dd25-40af-9300-421992159c62 · outbound

This paper cites Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.084210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.084210Z digest=sha256:34c6f179c9373b68e71bf11ca55050008694695830afff6f94c5f52d9ce79e34

Observation 83e788a5-4197-45ad-b62f-15d98788b0d2 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.188820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.188820Z digest=sha256:ea1d24e73ac9ffe270e8033e2edb30a3798b3e17c7fc58baeadc75c7261f4cda

Observation f9872eb2-77b7-401c-99be-2203668466c1 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation LTX-Video: Realtime Video Latent Diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.305026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.305026Z digest=sha256:98f6ca1a30b925ab351fa1725413433ec43ebcb8e20bc9c866e5f15f86c8c497

Observation 9996f1d9-74bc-4cfa-bf1f-f29249bb040f · outbound

This paper cites VBench: Comprehensive Benchmark Suite for Video Generative Models.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation VBench: Comprehensive Benchmark Suite for Video Generative Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.421147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.421147Z digest=sha256:cee153eb30c72baab216dfb3bb73ae782c88fab4f841be2e0904fca695ed7614

Observation f800d22a-e588-4f25-ba32-62d101f1e751 · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.517153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.517153Z digest=sha256:b5ec220b19ab7036bb07b456c5fe9bc00a44449c53dad540a0af180e8627a7d4

Observation d3fed15d-513b-4eb8-acb4-9120230e8599 · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.662964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.662964Z digest=sha256:b1801ade8352aaa6b7d1907769f5ed3277781e0f51e2a11d37eaf7a81b24db94

Observation 5ef96763-d99f-4034-94c4-9aaf620025be · outbound

This paper cites Kimi K3: Open Frontier Intelligence.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Kimi K3: Open Frontier Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.821367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.821367Z digest=sha256:e5b6baba9217ebf969f142b8d9b907669ee7e706e42c694c89bbe1d74260c12f

Observation 022e5892-adbc-4a92-9637-75959bf95652 · outbound

This paper cites Attention Residuals.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Attention Residuals

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:05.956746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:05.956746Z digest=sha256:98161b103db7a1d8fa02d982a0d61ead5f2cdf8b2844102583967dadbbed1bfb

Observation c9ca3f93-369e-4c1e-8150-804e07b3b727 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.066938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.066938Z digest=sha256:9a8fc71e21128a35dfbd648dbd48958d3766bba4a6849922ac2ee833ff4a2f9a

Observation 9a203054-2a57-4b2d-bf8e-d8f943ac781b · outbound

This paper cites PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.173473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.173473Z digest=sha256:a94a9326a33df0b6ca22925614e827c9a38e055afac2999720b7c3c516439805

Observation e72ad94b-33b6-4ffe-8d0f-ea30a0ce162e · outbound

This paper cites VideoMamba: State Space Model for Efficient Video Understanding.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation VideoMamba: State Space Model for Efficient Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.301385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.301385Z digest=sha256:99c33c75c051604355e1b01cc6e2052c03cea7d05d260f156638a4f6347eca4f

Observation ada065ee-a314-427c-a7e2-ad15f2412de9 · outbound

This paper cites Radial Attention: 𝑂(𝑛log𝑛) Sparse Attention with Energy Decay for Long Video Generation.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Radial Attention: 𝑂(𝑛log𝑛) Sparse Attention with Energy Decay for Long Video Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.452147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.452147Z digest=sha256:2a48185e3d7cd0caa63688c371716b198385877c967826f355b1b96d55debbec

Observation d840381f-f287-40b2-94af-15806ed26d88 · outbound

This paper cites Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.585681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.585681Z digest=sha256:89a2cd772b57d38ef2a28c204e8fdfe383de790c08f1568ac71c5c9b1a0fd262

Observation 6b267aca-3c10-4a8a-b317-02f155a496f4 · outbound

This paper cites Toward A Prac- tical Perceptual Video Quality Metric.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Toward A Prac- tical Perceptual Video Quality Metric

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.696742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.696742Z digest=sha256:24c089c95c8e9e9732d9e734b997fe983e09c4d692ef31c5fd33feb792875c5a

Observation c5fd60a1-ef5c-4b34-ae1e-d24eddf3b9d8 · outbound

This paper cites an unresolved cited work.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.815351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.815351Z digest=sha256:160df7729236613e13daa6ca138b72fe81f1202ab50a4584c627974c369640f9

Observation 51d1d801-8c1e-49db-a11c-4c54f1947083 · outbound

This paper cites HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:06.929517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:06.929517Z digest=sha256:dd43ae7b0da4af5d3080440da96d869519e54e9c0cd4cf642ae91e514be31277

Observation 3e02b234-06f3-4820-8891-4dc30a4d5f13 · outbound

This paper cites VMamba: Visual State Space Model.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation VMamba: Visual State Space Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.057918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.057918Z digest=sha256:97d9958254617f48bb01c6e5f06680b98a82a77c19f87acba97407fcc46d0da7

Observation 887082f7-fe31-4548-837c-3e0827a0d247 · outbound

This paper cites Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.203444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.203444Z digest=sha256:83d698513434e063837c1cf6f4d7c8ee7cdda35794cdce7b576197e9ad178e29

Observation 2119e194-4fbd-4a60-b3f8-fc2cb5f98de2 · outbound

This paper cites Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.343135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.343135Z digest=sha256:b8afa3c01eb7523312a849fca04630828e5ddb6070747cd6f22a4382daff5d80

Observation 66c7c2a2-659e-4565-89fe-d6d83a225b28 · outbound

This paper cites Scalable Diffusion Models with Transformers.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Scalable Diffusion Models with Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.456884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.456884Z digest=sha256:84b0c9cd6e11c1c2b03f638e6d1520945bd8189467566ffcebd5bbe21e8ae4b2

Observation cced871c-cbd8-48f4-bf59-aee4cac671e4 · outbound

This paper cites Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.598100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.598100Z digest=sha256:f85fb01528e42d12f79e10cf8f3f362f36f67fcd4ddc4ac719fcc919ca9425ba

Observation 999d7a27-f956-4b8d-bfc8-10436d5d4cbd · outbound

This paper cites Qwen3-Next: Towards Ultimate Training & Inference Efficiency.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Qwen3-Next: Towards Ultimate Training & Inference Efficiency

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.715164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.715164Z digest=sha256:c531b16b9b34964468e39fc9054e92fe1b51fd7fadb1eed27e09c0fc59923a49

Observation f604e2e5-8fec-46e6-b833-afdcf770b5f4 · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation MAGI-1: Autoregressive Video Generation at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.832039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.832039Z digest=sha256:e81987402eea90c1a044d2367d70673ba3da855d5afb7720c2b2cdef2106fbbf

Observation b73745d6-0306-48bd-b60a-eaca92f8ca2e · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Seedance 2.0: Advancing Video Generation for World Complexity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.970233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.970233Z digest=sha256:1dcbe24a779e8bb74b926af921e31afea75b36e2c4ff2b3d20cc913050157124

Observation 640c4dcc-79f7-46cf-b6cf-7c238bbc1559 · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.081225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.081225Z digest=sha256:2bd99f4cd91f511715b37000341414d3d1d53222160a92024de626c532186276

Observation 1b8fee61-361d-430c-90d0-733631c88fea · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing, 568:127063, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.266614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.266614Z digest=sha256:f55475b128d95d3bb0dd647ff4413482a871f8e9c891a1fe6c8d286ff5d9166b

Observation fd2842a6-f9b5-4354-aa62-8f3e9b4e5626 · outbound

This paper cites DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.398370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.398370Z digest=sha256:ad2e06ef6f0608719d7ea1443e4f38800f00682bc716711b6365ee2646c1dea1

Observation f250d9a7-6ee0-4c57-b19a-8015c93c36cc · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Diffusion Model Alignment Using Direct Preference Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.542920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.542920Z digest=sha256:94bd794f0349c2d9367170c3cba009fcf2da22da31bded13f5abf0ae0438421a

Observation 44ca3871-0407-4b82-9035-bb26fe2bbc1d · outbound

This paper cites Wan2.2: Open and Advanced Large-Scale Video Generative Models.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Wan2.2: Open and Advanced Large-Scale Video Generative Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.630618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.630618Z digest=sha256:bd8c66faa4af9460c47b0f7dad23b34d50ae4cf831dae83828c38437d3a0d389

Observation 58cd1b90-7d58-4106-b8f4-44bbaad2dd34 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.710736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.710736Z digest=sha256:e9e9fca8825ca73adb3ae7ff7a2c7eb3ed2468827785bb2b7f452c9629bb37a0

Observation 61bbe66a-4ef4-4dcc-be18-67897e00162d · outbound

This paper cites Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.848484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.848484Z digest=sha256:064e2ad7e542c289b1f36aa77ed2beaa34fe259baa8fe5752fae0882f31f26e8

Observation 9653e097-2e9c-4e36-a90a-2cbb82949270 · outbound

This paper cites Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:08.982968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:08.982968Z digest=sha256:bed3fb17805288ad730018d7994540e15f772b99689d24c6b0dab47c42d220a0

Observation 54b66d7c-87ba-4a79-af70-eb4703016be7 · outbound

This paper cites Unifying Flow, Stereo and Depth Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11): 13941–13958, 2023.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Unifying Flow, Stereo and Depth Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11): 13941–13958, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:09.153544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:09.153544Z digest=sha256:cad3e1f12c59e6ec30914b28adbcc7ecfc780983cfe32b269c0375f69dd6ef4f

Observation 64856d84-6dcd-4441-8e23-464b8cd55285 · outbound

This paper cites ImageRe- ward: Learning and Evaluating Human Preferences for Text-to-Image Generation.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation ImageRe- ward: Learning and Evaluating Human Preferences for Text-to-Image Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:09.259598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:09.259598Z digest=sha256:346b419d4f1abf6f3016ed9637f4984bc8cd80b1b12004c281a79f221505360d

Observation 110fe622-34d5-4823-9dc4-28cc3066df91 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:09.367772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:09.367772Z digest=sha256:3195a2d697d34cce8e3f1205f260a961fbefcd19d96a76d090948285b16df5ab

Observation ae2c8a2a-1f1c-45cf-a701-a40ac93977ab · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:09.504622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:09.504622Z digest=sha256:ffc28f4591b9201b50afd61eb6c4089defc8af1f4c9254396935978e93b5c5a4

Observation 42e049af-2592-4825-8dc4-ec39913eac64 · outbound

This paper cites Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:09.627087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:09.627087Z digest=sha256:403cfbef38e3efbd5d9f61814423a53465ca8f80e39199b03baa81ced37223c3

Observation 53766f05-37c6-4596-bbab-07456923d508 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:09.744027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:09.744027Z digest=sha256:4bc3e98e4d45b218b7c7c25e339f322944b8bd4fdf64b713ccd6037027947b52

Observation e918a28b-d180-4ba4-9e46-985de5cee056 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Sigmoid Loss for Language Image Pre-Training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:09.905050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:09.905050Z digest=sha256:99ac54b16e8d274d852608805ed4de1e9dd1739dc0a31683ba1d83f84a002e07

Observation 6c359b24-7a7b-42ff-9b5c-c1796c5181d0 · outbound

This paper cites SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:10.076055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:10.076055Z digest=sha256:af834d246732abf767608c4ddb5236a3e630e545c55f926ed5447c34003f32b1

Observation 5537ff05-11a8-416a-bb93-5fef02adcac6 · outbound

This paper cites SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:10.254223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:10.254223Z digest=sha256:f500f051f11a9b7195db8f38b6f803baee2751f8c035966d4edc58a0a712f7df

Observation 9f939f6d-c5fb-4c6e-92dd-4cef37750e66 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:10.371272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:10.371272Z digest=sha256:2c90ab5c733399d176c82cb28072348291208b8cd4b5c970ab109b1ab4bd949e

Observation 147a1451-2749-4bef-9a9f-1f90c6b49484 · outbound

This paper cites SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:10.471087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:10.471087Z digest=sha256:85670b9cc717d4b830b264779d43cd6aecc794b678d9d47c0397d842f01de35c

Observation c70410cd-7a54-425b-89bc-1f033054650b · outbound

This paper cites DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-01T07:09:10.589942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:10.589942Z digest=sha256:b0372999ef0f278450f75c404d2719652ff133343ff7f510ae6ea02d189df27e

Pith citing papers

No inbound Pith citation observations are available.