Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:34:44.921475Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2608.11752.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:34:44.921475Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e0a2bd94-66fd-41ab-82c9-92f5384b7982 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Emerg- ing properties in self-supervised vision transformers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 131cc951-9f92-4247-b836-801728c96427 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Diffusion forcing: Next-token prediction meets full-sequence diffu- sion
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 80487082-2489-480d-aafb-69e303f873f6 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Simswap: An efficient framework for high fidelity face swapping
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4b5412f7-ae97-47f5-a7b4-c92da1f5e8f4 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Wan-animate: Unified character animation and replacement with holistic replica- tion
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41245199-c981-429f-b2e9-d543c56e4958 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Out of time: Auto- mated lip sync in the wild
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4928abc3-0ffe-454e-b5c4-bfa568fe2e5a · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos High fidelity neural audio compression
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1e77cd2b-6f56-4497-96f8-6ca223385895 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce80ddc6-bd10-4157-8407-f312b4e86e03 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Freeman, and Michael Rubinstein
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ea42b9c8-20c1-4395-a209-de4a99b83469 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Video Diffusion Transformers are In-Context Learners
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13ae0261-a23d-4d0b-b3b5-7e85f30771a4 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Infoswap: Information bottleneck disentanglement for identity swapping
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2593158-ce82-4d20-a5fd-2699eebc81c1 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos LTX-2: Efficient Joint Audio-Visual Foundation Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d79091-442b-4a65-9d08-8155f7354844 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0750c1-638e-483d-8947-82f86d380629 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1aad171a-0f5d-4b06-a868-d2a11f8e1f47 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c9d40e-5cc9-42ac-aa28-f3ff1e0ebb9f · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos In-Context LoRA for Diffusion Transformers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca726dfa-3ba5-4a90-8266-6570709a4c23 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44c1d740-c104-4cd3-b809-a566076147bc · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48a0d8a-00f4-4601-bf67-0ed00ea735c8 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos VACE: All-in-One Video Creation and Editing
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6483e68a-ab90-48ca-88f1-d0ec1eea21fa · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Faceshifter: Towards high fidelity and occlusion aware face swapping
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6515f221-9296-4b83-9bc6-6fba7e6d4e29 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Rolling forcing: Autoregressive long video diffusion in real time
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1de3ba52-1241-4d3e-888c-d7004100fa8e · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos JavisDiT: Joint audio-video diffusion trans- former with hierarchical spatio-temporal prior synchroniza- tion
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 02750d3b-f5d3-46e8-8522-9a31d5734a70 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Zero-shot Voice Conversion with Diffusion Transformers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f19a9fc5-caa5-49ba-9085-a4bbcfb44244 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Scalable diffusion models with transformers
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0e569ce-d332-4391-a1f8-166049140720 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos OpenVoice: Versatile Instant Voice Cloning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fea2ab3-3036-4c28-b478-673e796c16bc · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Sam 2: Segment anything in images and videos
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4b19e922-6906-4475-b323-9d24855ffcfd · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53991683-2ff0-40f4-9991-609ffbaceea2 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 257e803d-5781-4ce5-9c0b-659b4e0d74de · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Consistency models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd3a51e-658e-4533-bf93-f464a06f43ef · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Roformer: Enhanced transformer with rotary position embedding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 99ab2679-559d-454a-a029-d90548a7a6b3 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Omniforcing: Unleashing real-time joint audio- visual generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89528c1c-f894-4ead-a3dc-9ca6444e4839 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Wan: Open and Advanced Large-Scale Video Generative Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe34031a-b2e8-43af-b62a-73bcf18a7490 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Generalized end-to-end loss for speaker verification
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 58307b3f-a855-4825-a056-cf32436d8b5c · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Bovik, Hamid R
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cf83fc06-4075-4d70-aa5f-d8ce9cba181e · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Williams and David Zipser
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb0d50fc-89bc-41b0-a29b-4eedf4bd61e7 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Q-Align: Teaching LMMs for vi- sual scoring via discrete text-defined levels
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9a94d4ea-e773-4e0f-b6b3-5873782d1241 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Efficient streaming language models with attention sinks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation afb43cf0-ccee-4f73-add5-d82decaa9806 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Vit- pose: Simple vision transformer baselines for human pose estimation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd0df51c-5499-4fe8-b510-0302c302c4da · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Mocha: End-to-end video character re- placement without structural guidance
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d99a8a-e375-42f7-846f-178ac4403970 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos SCAIL-2: Unifying Controlled Character Animation with End-to-End In-Context Conditioning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 075238d9-9d63-4afa-9e59-aabaf7e7e446 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a404c2-381d-4a3c-ba4d-5b6e4d836660 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ff9bd18d-c2e3-4c90-b0a7-b8a03b906b33 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Freeman, and Taesung Park
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4492e9ed-aa5a-494e-b2ff-17a36c0cc83b · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Free- man, Fr´edo Durand, Eli Shechtman, and Xun Huang
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0c368f77-fe29-47d4-9c41-cd7a5dc58c3e · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The ablated variants exhibit increasing identity drift and vi- sual artifacts in later segments, whereas the full model re- mains more consistent
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation de58b078-ca90-434e-af6a-44373ebb30cc · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The reference cache persists throughout generation, source keys and values are tem- porary, and completed target blocks are committed to the clean-history cache
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 896a7ddd-61ad-43c7-bf44-608589ad03ca · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The study compared UniSwap with four video-replacement baselines, each paired with Seed-VC following the cascade protocol in Table 1
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bc9a1f1d-3ab5-4c1e-a6e3-b7bcdeca3c31 · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos In all figures, each example con- tains a reference image and reference voice clip, a source video and its audio, and the joint audio-video output pro- duced by UniSwap
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a15be8c3-6fd1-496e-9ff5-4214f51def9c · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd69542b-0ccd-45b2-8e7a-6dea3a7af39e · outbound
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Deployment should require consent and provenance mechanisms, visible disclosure where ap- propriate, access controls, and compatibility with forensic detection tools
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.