Pith. sign in

Paper Citation Record · LEDGER

Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2305.10874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10874 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:04:21.817269Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:34:37.702024Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 532c46f3-6ed1-43f3-8103-8712c7bd60d1 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.704069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:87bea2054a51f585ecdf80dd63fe61cb0ee66a763655883e13d0e9c32ac1b47f

Observation d380d28c-b9b5-4b58-8966-571beb6a8667 · inbound

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation cites this paper.

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T21:04:21.817269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:04:21.817269Z digest=sha256:dbdd4ca369907739e02d02eff27e3b9e36623d722b029761b63c7cbabb062456

Observation 98a14837-41f5-4e78-b2cc-08baec0a695e · inbound

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation cites this paper.

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:53:11.758478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:53:11.758478Z digest=sha256:0787425f6e195a4546914a601528f29f9c6fb15ee9b4c10d1f723760a7c8804c

Observation 203d4ee9-6bf2-441b-9a17-d75a069808db · inbound

Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention cites this paper.

Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T10:28:07.661616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:28:07.661616Z digest=sha256:69b88e85eeeb56c34447f4c2775fe27cf7b69b3bc975d2b51fd669e010a9ddcc

Observation 486c7642-7c9b-4c10-9a4b-dcb47fe89ae7 · inbound

Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation cites this paper.

Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:32:44.504172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:32:44.504172Z digest=sha256:ce0f91212d6a69c19f3d166aada8e67525ed2114c422229e014a69bfcef18078

Observation 52e689b0-0447-45a6-8c07-7e6606a3d111 · inbound

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention cites this paper.

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.190603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.190603Z digest=sha256:5d150e902883bd55dbeaf2afaa160d6b0ac4b258047477be5372550a5e3cd80e

Observation 81e4186e-3df2-4fd5-bd42-aae55404c3c4 · inbound

VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation cites this paper.

VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T18:05:13.592846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:05:13.592846Z digest=sha256:5fc6fccf89a6175e1d889666c7dbff93662945e45df4fa6518733433cf72879e

Observation 89e4b106-5620-4235-93ac-6db5cc6376bf · inbound

Grid: Omni Visual Generation cites this paper.

Grid: Omni Visual Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:33.540440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:33.540440Z digest=sha256:32c5254f39d51c88d5f651a525b1ec101d6fb20fe612acb7ee0c4ef1e33fa402

Observation c7f3d49c-d3db-4c1b-b2c0-4b2634614128 · inbound

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation cites this paper.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:15.128034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:15.128034Z digest=sha256:fb456965d5905bb3cc28e83d9494525439e04a18ecf359543524dcb5eb260a45

Observation fed664e7-9746-4563-81b5-26a7dbc4ec9d · inbound

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models cites this paper.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.696450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.696450Z digest=sha256:195cbeec4d6a21deede229516e984025204f436d9af3f05480612577c1dbfaeb

Observation 0060c327-9d8f-41d8-adbe-7add5337125d · inbound

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation cites this paper.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.074328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.074328Z digest=sha256:852617fac5c98ede4799a960732a84013fbe3a4235250cf2087d15a94eed7f0e

Observation c5788791-343a-4a42-b610-91b72e515481 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.266101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.266101Z digest=sha256:c654e369c4d252d1164bcb11e6c6a469d0e80eb4c48554f4fb0f2b5fd953cfc4

Observation d06e90c2-55dd-4d97-b791-7ce4f5559953 · inbound

TextMesh4D: Zero-shot Text-to-4D Mesh Generation cites this paper.

TextMesh4D: Zero-shot Text-to-4D Mesh Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:50.065412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:50.065412Z digest=sha256:1021747fea14ad2cdb8ff22242966d45eca6627675cf8eed3166d3d203556421

Observation d299757f-a3c3-43ad-86ee-fdca42af45cb · inbound

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion cites this paper.

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:18.331975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:18.331975Z digest=sha256:eeb48831c7c6eaee6f080857fa65fdd6b9e8bc4ef1ceddec4d903f0a2ef97d66

Observation cfc808d4-8e26-44b4-83a7-5c2aba9e5c35 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.322880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.322880Z digest=sha256:4f7e1dc2e549814fbd75d834a9ec11a88f38c31d1c581bdd433050291a4946d4

Observation c1512bc3-fb6e-4207-aee3-43338d15919a · inbound

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences cites this paper.

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:40.408942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:20:40.408942Z digest=sha256:e85a7232ada636306f515690fcb63d7ff9f93a052a644e2a7e6f67cf1b6ede99

Observation 4b8f6296-d361-4a00-84b4-4ba3751f0150 · inbound

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency cites this paper.

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:30.901204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:30.901204Z digest=sha256:5da58cef47760b24e43e758158fefbab053b1224838c23360b147072cee6c041

Observation 0c99a30f-fd9e-415b-ab35-3235bac5faa3 · inbound

UNICA: A Unified Neural Framework for Controllable 3D Avatars cites this paper.

UNICA: A Unified Neural Framework for Controllable 3D Avatars Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:48:11.255188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T19:47:13.052204Z digest=sha256:71c834380e8655f6ed00b000cf9645314b56f5d799c2d56dc13ddfcd65b7d026

Observation 4a9695ad-c41e-40c6-b1a4-01a26f35a35b · inbound

Detecting AI-Generated Videos with Spiking Neural Networks cites this paper.

Detecting AI-Generated Videos with Spiking Neural Networks Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:12.328250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T16:15:42.113951Z digest=sha256:e4b929e29c2d8f26d1934fb5c9a75824e9c57a21924a4559b1b34ef018a1257c

Observation 824d2951-fae7-4812-972e-8320785a452d · inbound

Detecting AI-Generated Videos with Spiking Neural Networks cites this paper.

Detecting AI-Generated Videos with Spiking Neural Networks Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T02:25:01.702234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:25:01.702234Z digest=sha256:e1d4f88f4f9d18f4df5268cfded8737aedcf402b1abdb41cfa35bae0fa1e59cc

Observation 1f389f1e-0f8a-4008-899c-9103a5bf502d · inbound

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation cites this paper.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:48.054091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:23:48.054091Z digest=sha256:337284dd35e369cf97609931ab4a2afcb7f9fc7e5322099adb408adb89de87c5

Observation e7c544ed-a4d6-4341-b448-2587a11b1bdc · inbound

Retrieval-Driven Training-Free AI-Generated Video Attribution cites this paper.

Retrieval-Driven Training-Free AI-Generated Video Attribution Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:36:48.967450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:36:48.967450Z digest=sha256:a3197fbf0d11254be3dce31a552a57cca83ad40e1abdb98a892029faa4f30980

Observation 8c187cd9-4c89-44ba-8d5c-d7695cb93257 · inbound

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images cites this paper.

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:20:04.584508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:20:04.584508Z digest=sha256:9972725c7501ae626e55dde5bb4d2442e116be3fec328aae667a1636d7c924b8

Observation c23c8670-7780-4b06-a1a1-8747df9e688f · inbound

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection cites this paper.

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:24:41.080064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:24:41.080064Z digest=sha256:f4624ef5fb275bddaf90927fc307fbbd3fdd6382acd4d0ca1d0206e8de0ab60f