Pith. sign in

Paper Citation Record · LEDGER

SkyReels-A2: Compose Anything in Video Diffusion Transformers

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2504.02436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.02436 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:29.507056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.697793Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39046ad6-2539-40e1-b5a4-c817f2aee04a · inbound

SkyReels-V2: Infinite-length Film Generative Model cites this paper.

SkyReels-V2: Infinite-length Film Generative Model SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:23:04.222693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:23:04.022599Z digest=sha256:dd682cbe5ce4d420385532bd774eccefae9a4213da29bb25ebfc8d330e0436d1

Observation 87b78b0f-fefc-47bf-a394-e5f5bcd6680b · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.507056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.507056Z digest=sha256:34b57f2733ec9047953884324a7eb0ef119a56a4a62a53d3748cb29f7ace2956

Observation ab65c0ea-f8a6-4c34-b126-694ca2bb456c · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.579680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.579680Z digest=sha256:c44f7303c6f7217a20441b321a18fc916247f6537164fc9a13ad0e323c792665

Observation d13d751c-5cc6-4351-86f9-1e9a92164b10 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.911278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.911278Z digest=sha256:7971fda39e8d26224385a0c82d2f5f34c2dd645ef809ee6584f5680e18c729b5

Observation 3578d72a-dae4-4e0f-9541-0faced2c218e · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:08.358935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:08.358935Z digest=sha256:e7ceb9f0ba37e88fe9b0ce331862a5bbdca3dfe3d12551df4a03fe1d4242a5a0

Observation d6acd82e-d179-47d7-9f00-b4b7cc9edfba · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:49.621536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:49.621536Z digest=sha256:fe60cbb33f9b80614619816b3ea6fc5e2d174c2931392eb552234baf1e496632

Observation 1b6d3fdb-d546-415e-8772-7bc04b4a3446 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:11.237810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:11.237810Z digest=sha256:009024e9e4792783b4ad977f4717b09cd9840045ca91cc8d8e390c32365d5479

Observation b1f8a611-a649-4508-8766-71186b1dff45 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:1894cfac5d55ae972a4c671386f4ebc8948a034245acf3f0912a0f854cf5fc86

Observation 2fc26fd8-3a6e-41f9-848c-bfa68a93da58 · inbound

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation cites this paper.

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T12:23:05.876881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:23:05.876881Z digest=sha256:3dc3fdb2698bed78046db32b1484a78bb204fa9a557bda5b6f4758a94b5d755b

Observation 72e0e95b-1a83-4e3c-9ff6-e486b320c67c · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.736553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:69f4e5808e124abafe460b1e92b5f839ea9de63e1e919db55ef20fd423d6c77a

Observation 6d430c7f-fd44-475e-80bd-e20e4a92d472 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.589537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:0f7a6f64e776508a2dc36ee2bc3dafa5dd7698fa54e6970f4499e44d103b5a8e

Observation abfaf6de-042a-4f2f-a06f-c4a18735ba4f · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:27.920254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:0aac87f02bbc3bd2ef2dd7ef90bd4c2fc2686d0993d172ff0432d346e2a4e78f

Observation f809b158-7281-40e1-a246-02951b9a356a · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.822714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:17e06520c1db397f449fe6e1721973d74847b5e7247819e5556962ed8b4aae9f

Observation c9566795-c710-47f4-a259-cb5f6d37f777 · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.513027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:3802fec879c42ee48d8bd6b19d623415fd1590512f83f9ee620c8c800a3ec7e3

Observation 214c9d33-2862-44ea-a967-9c7f20332a3b · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:57.606468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:18:58.996355Z digest=sha256:489902ac1a5d0f18e8189ff7a4a33ae8520cc1cc00c40f9b7c2de487d1677600

Observation 9aa3b2cc-8c15-4372-9f55-8689061fc327 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:09.728834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:31c9212af06e711e23c3a75f066baff429e343991f3a4f55a193b925286346e8

Observation 34ff1e3e-f549-4184-b7b4-3404c3260b75 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.556258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:55aead9025c778f0599b2c74e0ce385b19d4757fa3d0ff806b31a1dc59216eb1

Observation 8923338d-5b4a-4ab9-a032-b86c1e9aa642 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:25:00.575779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:133034ebf2c78f2cc9fed1d737421b0ec22506b2395d2e82e626a6de51112a8c

Observation 2d1a34fd-fe81-4098-a5e2-3b1286bbf6a7 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.376116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:8976a100d20a84500bd33ae57a8be970c74ad379df3857404324199c37ab9eba

Observation 49ff3ff5-4d70-4335-814f-50e9c371f5c2 · inbound

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation cites this paper.

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:31:09.909379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T06:30:18.961069Z digest=sha256:52e2089b693126dc81005775e6c08f23a9748a601fc476979bb6c4324cb90858

Observation 641621f0-0387-4d94-b33e-9082869dd837 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.408596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:db8c5801e4240fc7b17ed30c9fde0929303e37047abd9eba15ca0cbdf75cae01

Observation 4243e34e-c835-4773-91d9-83118c6f6ee1 · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.490012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:dc436cbbb112a3d49de909f03207594e3f24f9e6ae25777da0283bc4dbe6118f

Observation 5283a361-62e2-45a8-b4eb-f8d9ed9e2e13 · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.380960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:590707a28175080881bbc207ef68578246d216d1fcf9869611031ebf5c701ed7

Observation 0afa77da-e7c3-4a5a-a63a-19e530758b4b · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.951440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:aea5fe284b60ee154bdf623ba6b6fb703834a136749a581bba6d9fb609aac6ad

Observation 929fbd2a-cce8-4503-a2d1-03045acf5f06 · inbound

CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation cites this paper.

CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:32.689770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T06:51:31.625416Z digest=sha256:7d302bf06fcd4553d8d8ec71c9057f70dcafc5f2acdefbdc13550b846163c40c

Observation 08d03e3d-dd47-4ca5-88ce-e75f00e6a344 · inbound

DramaDirector: Geometry-Guided Short Drama Generation cites this paper.

DramaDirector: Geometry-Guided Short Drama Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:59:55.923797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T01:57:15.550335Z digest=sha256:5f64d6e35cc8b4d7ad04fe4b82bce46fbd6fff140a21ed43b867d96ae9991ace

Observation b53a9ff1-507c-4eca-b576-3096239f3fab · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.699283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:f3fd1b6cb81c4dfd5f07c785c8ced9acc47b2d751099e11a7c1990a7a9116e3d

Observation e285619c-f356-468f-af29-8b43e4a227d8 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:c831c9f4981dedbb6e0e21d579457bebd3c6344f9861bb5c1a2b35b5027dc5cd

Observation 78d4f218-73d2-4042-bb9a-9381a80e5282 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:00.510556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:00.510556Z digest=sha256:e4159bd6d09b24fabc26695e5e8b5a49822dcc32518c112751114ae6045aceb6

Observation c532743d-3314-4c6a-bf32-0df7f6622b84 · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:19.331303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:19.331303Z digest=sha256:1711cd1e23a66ea8c7839a42e2ac8213331836332d39d24fe582181901a9261b

Observation 36060856-f20e-4f5f-8a43-15ad641130d3 · inbound

Vera: Identity-Faithful Human Subject-to-Video Generation cites this paper.

Vera: Identity-Faithful Human Subject-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T10:24:51.437202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:24:51.437202Z digest=sha256:ad53f18963d6a32bca0a1478f18306720c5679c166e1b8877f2ff3492f6bbfcc

Observation 893e6074-5bbd-470d-bdf8-5352e25aa5b9 · inbound

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling cites this paper.

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T13:56:37.838490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:56:37.838490Z digest=sha256:5fd1e6af58608fccaef012b061633a79991ec8feef48a32f5113f30a14de7210