Pith. sign in

Paper Citation Record · LEDGER

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2604.19473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19473 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T02:45:10.577070Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T01:52:14.874049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-15T01:53:28.740566Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact15
  • verified fuzzy8
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cdde0bd5-d641-4a8c-bfaa-0d05ea1d7fb5 · outbound

This paper cites TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:53:29.905801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:d79bffdab89dab00d8de432f1afb9c889482b3f53e055cad01ad8eac9aee0396

Observation e3723677-5e7a-466c-b7cb-8ddf88060a44 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation SkyReels-V2: Infinite-length Film Generative Model

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:23:04.486750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:abc35aa1f2a07c3c40939412473038f1172f675ca97bceb0aa4704441eca6edf

Observation 41c64761-3b6b-4330-9bf9-988ed60ec2d6 · outbound

This paper cites VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.915192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:0cbaf5e4a3be0c8f9c89b3b9c649227d6d4415ad6c9ef52c48426e08eb7a0dcb

Observation a6452147-1bd9-46b1-b879-ee5b15d83f5b · outbound

This paper cites CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.912925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:e5f23454183f1ceaa6ccf0c514e66e69ea7f81bde2a51ea93371e4e97769bf17

Observation 1f880598-6788-4466-8203-a223b55a83ae · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.919945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:0d2644665d5fc04f57491abef8bf6ede5fed1bfbd0e3dcd9a8a51c97c705ea51

Observation 9ced2ea6-9867-401f-a814-493c6c86721a · outbound

This paper cites I2v-adapter: A general image-to-video adapter for diffusion models.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation I2v-adapter: A general image-to-video adapter for diffusion models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.540880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:25fb04acb7a298827f3c379daf14fe043a053baee5754eb843988ce703e5ab9c

Observation d78a4c90-b8cd-4a0c-9c6b-8605e9222a88 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation LTX-Video: Realtime Video Latent Diffusion

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:12.849504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:56ce340b3d606465056bad99c5ad2ad4740c8846ea2914984c9461ef823d9728

Observation 96d0743c-6afa-47d6-8b75-e84bd5c558fb · outbound

This paper cites Hailuo.https://hailuoai.video/.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Hailuo.https://hailuoai.video/

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.542672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:8b43788a9f9d5d8d4e9949fb0f3eed409c3a5fd0b8bfae5275b3b4b584e16645

Observation 014dd85b-2817-4cc2-be1c-e2a8fccbdcff · outbound

This paper cites ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.908134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:be73abb4765138c45d5755f3489cef759ce6dd440cd515e1d8d860337c81fca4

Observation 65333d83-84de-4bec-87d3-da71d12a5d45 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:53:29.903427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:a11e312b441db2c678636a2f1992065159e53a4b242a1acbec68b2c887322bbc

Observation 012f6468-cda6-44dc-b980-755992f6f31b · outbound

This paper cites Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.924663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:d9443d396d6d257f680f8f92f7b819789f419f2c1f942d98f2126cb16572fc0a

Observation 09e337e9-9e14-4424-b69c-ba81df138223 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-10T02:53:29.910521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:27d2f44ce494d7821e96ccb6097412ea8ac22639f79c0ed46f93c85d4230edc9

Observation 79254b7a-101c-4b9b-9afb-a7f5e1eca917 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.943899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:904272dc304faf8eb020b1f67e1d84cad99fa7b54b4fc5cc98c6cd6733704fda

Observation 117af888-d99e-4b7e-97e1-8f0468250913 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Open-Sora Plan: Open-Source Large Video Generation Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.936813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:d44b7974645779a8f5778b2f2b1a47a69b279316b083cbf2fbc9d38c7c39a028

Observation 0c383d40-0285-442d-817d-50c420810adc · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.929385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:f18e239605d1ea44dce38eeabc3665b0e441741d7e3f815c724fd7805d36082e

Observation 118c969b-57b7-4cf7-9781-c1842ef96b7f · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:45:35.835056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:d42ea5e2b82301a87d3d86c1a718bc4467031abeafa72da95d5ff7732df36a1c

Observation f0753719-1b15-42aa-bd84-05af895bd6ca · outbound

This paper cites U-net: Convolutional networks for biomed- ical image segmentation.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation U-net: Convolutional networks for biomed- ical image segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.538759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:c02a15c8b93463473ff70a6e99df2d2168eca18f47e36c40874e2bfef56fed1b

Observation 8816a3ba-8772-4f67-a00a-840ad7beff4d · outbound

This paper cites Longcat-video technical report.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Longcat-video technical report

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.544546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:ea0f80d8eb9d614db21f5ba4c5dd885d5792ad0251635cc28c19d4fdc1eed551

Observation be861c36-4715-4d8a-ac58-27b095d2df79 · outbound

This paper cites Longcat-video technical report.arXiv preprint arXiv:2510.22200.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Longcat-video technical report.arXiv preprint arXiv:2510.22200

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:53:29.946447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:2ab78acbb370e2dec5beffd87a383f2b9c49893cd53301a221410fa23d7c0217

Observation 0ac2aefb-0eba-43fc-bc0f-32c5a17f98c8 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-10T02:53:29.941501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:9d0a691108e8d8b93410878de20c65cadbe83a985d74edd44709dcde5b4e3a4b

Observation e1ad8102-c819-4722-b346-2bd5ea659a37 · outbound

This paper cites Dreamrunner: Fine-grained storytelling video gen- eration with retrieval-augmented motion adaptation.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Dreamrunner: Fine-grained storytelling video gen- eration with retrieval-augmented motion adaptation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:53:29.939209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:42ea9cfcc9e87957ee4fc36e35d45367185c9a6200418538ecd2b943575709e9

Observation a7ec3c7b-1855-4682-b364-2575f024a8d2 · outbound

This paper cites MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.934274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:b7a1bac009916cce531e5dab644dd79c70cb11b7018ba1475b1550562e334e2d

Observation a177a1d8-693b-419d-ab36-0002b88977d9 · outbound

This paper cites Magiccomp: Training-free dual-phase refinement for compositional video generation.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Magiccomp: Training-free dual-phase refinement for compositional video generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.948877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:85d0a01237b4a277a6d978688baa00d83ecc568cb7209e52b729e9b136d97ca2

Observation e7c25fe8-c48f-4fc5-9108-b62c57f85319 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Open-Sora: Democratizing Efficient Video Production for All

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:01:52.219350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:0e16b55e814ba8952338732b540a4e5eec99211cfab56e2cafdbee70eb5dcdda

Observation bd3be9df-652d-4734-b421-f5b95a01abc9 · outbound

This paper cites Therefore, this section provides a supplementary explanation for scenarios involving multiple subjects.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Therefore, this section provides a supplementary explanation for scenarios involving multiple subjects

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.534664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:7120d5b8305fc80b769c690e2b8bb0a8b6398d1663d855a02fcbdd493d9fc34d

Observation b2d55bc2-49f2-49c4-857a-5ba6ff2cbc8e · outbound

This paper cites Relying solely on attention reinforcement reduces TS-Attn to a mere attention enhancement mechanism for event tokens, lacking temporal correspondence.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Relying solely on attention reinforcement reduces TS-Attn to a mere attention enhancement mechanism for event tokens, lacking temporal correspondence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.528712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:1ed9b83f6e195d8fbef3d671995650ed48181e6a5dba7d91f6384790309b67ab

Observation ea8da723-6fe1-4088-a6db-519756208752 · outbound

This paper cites As shown in Figure 9, the attention distributions of different actions in TS-Attn are clearly separated along the temporal sequence.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation As shown in Figure 9, the attention distributions of different actions in TS-Attn are clearly separated along the temporal sequence

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.530673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:d7d3b1cb2da2c851257b4317a98a0b4837fb9ea94715e9ee9abe5404fea9aca2

Observation 6a32f843-e629-4881-8b50-f69b4f2038f2 · outbound

This paper cites an unresolved cited work.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-22T19:37:01.532735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:a576fcf4cebb6b16df408568a8a4a5f1830f671725d42419d601e065d383c4d4

Observation af187422-e27d-44cc-a545-b67b002f14ab · outbound

This paper cites (2025), which natively supports video continuity.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation (2025), which natively supports video continuity

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T19:37:01.536588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:1ab51144d9a0566c4402d6141b4df62fa74aa84804b43faf13b7bd7f56089945

Observation e0d31b98-f067-4839-940e-38e4f967d706 · outbound

This paper cites These results highlight the potential of TS-Attn for both interactive and long-form video generation.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation These results highlight the potential of TS-Attn for both interactive and long-form video generation

Reference 30

Resolution
malformed identifier
raw_fallback, observed 2026-05-22T19:37:01.526466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:e79393462b367f03c5735a6382fe3d2f78a0e7d2d6d451c0f8cb6a0ad2532b3d

Pith citing papers

Observation ea3d4ba2-8682-4ae4-b034-0258bf30f9f5 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:53:28.741906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:52:14.874049Z digest=sha256:e371b6dd52c8794fa705755cf47bf58ce0e72e86ddacfc89decaea3182d75bca