Pith. sign in

Paper Citation Record · LEDGER

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2608.11752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11752 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:34:44.921475Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0a2bd94-66fd-41ab-82c9-92f5384b7982 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Emerg- ing properties in self-supervised vision transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.617921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.758406Z digest=sha256:bddc1291b629e1125c41a13a7f247317579fb7f9bfebc92c9ec9804cac09f8cb

Observation 131cc951-9f92-4247-b836-801728c96427 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Diffusion forcing: Next-token prediction meets full-sequence diffu- sion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.608749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.762274Z digest=sha256:57f85c2b885ce960d7a138d27362c14876a66f7dc154465429406e75e5f3adf6

Observation 80487082-2489-480d-aafb-69e303f873f6 · outbound

This paper cites Simswap: An efficient framework for high fidelity face swapping.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Simswap: An efficient framework for high fidelity face swapping

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.599587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.765808Z digest=sha256:83a564c6b3e0f1c05c2ea1e223281b2aa5157523fbb2eba3c045b1c0144bbfd4

Observation 4b5412f7-ae97-47f5-a7b4-c92da1f5e8f4 · outbound

This paper cites Wan-animate: Unified character animation and replacement with holistic replica- tion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Wan-animate: Unified character animation and replacement with holistic replica- tion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.769542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.769542Z digest=sha256:d64af122e88e6045646f662f033e6db432779524ca14e722eed083fca4b56022

Observation 41245199-c981-429f-b2e9-d543c56e4958 · outbound

This paper cites Out of time: Auto- mated lip sync in the wild.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Out of time: Auto- mated lip sync in the wild

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.590449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.772765Z digest=sha256:4c65c9f789f1abd8c801a04557d857f3f9157ff3f0f47df2a58a2dc89e3b457b

Observation 4928abc3-0ffe-454e-b5c4-bfa568fe2e5a · outbound

This paper cites High fidelity neural audio compression.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos High fidelity neural audio compression

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.581075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.776021Z digest=sha256:0e924f23eb3486d707947cc9bea5a24af09a33c7810d4c4a8c94a5a43a83a75e

Observation 1e77cd2b-6f56-4497-96f8-6ca223385895 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.779593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.779593Z digest=sha256:c24fa7b3c0d4f5713c360a7c2b1526e37078e56bb8c3b7682849798bca2dcb65

Observation ce80ddc6-bd10-4157-8407-f312b4e86e03 · outbound

This paper cites Freeman, and Michael Rubinstein.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Freeman, and Michael Rubinstein

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.571614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.783139Z digest=sha256:76e771390b4febaeeb352b3dcb57b4209293ecf84a10a07eadc6c265adc1b2de

Observation ea42b9c8-20c1-4395-a209-de4a99b83469 · outbound

This paper cites Video Diffusion Transformers are In-Context Learners.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Video Diffusion Transformers are In-Context Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.785950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.785950Z digest=sha256:a62c4de2f00acb711d830fb8ce164547f79563a6a04cf0b9b6c1f13f9248d573

Observation 13ae0261-a23d-4d0b-b3b5-7e85f30771a4 · outbound

This paper cites Infoswap: Information bottleneck disentanglement for identity swapping.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Infoswap: Information bottleneck disentanglement for identity swapping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.561843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.789439Z digest=sha256:07d696c5ae1704dd8d68061ff7136266d687098d6106334ebff2ccd21a8370b9

Observation d2593158-ce82-4d20-a5fd-2699eebc81c1 · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.792702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.792702Z digest=sha256:8fd3106f97ecca821ac1ce5d9fc87dc56a972357a665156babbd73ebe47dbd42

Observation 84d79091-442b-4a65-9d08-8155f7354844 · outbound

This paper cites FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.795982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.795982Z digest=sha256:92dc2045ba54bc7a0aceff97e607d93bb220d5ef813fc21c073d75ff839aeee4

Observation ef0750c1-638e-483d-8947-82f86d380629 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.552326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.799943Z digest=sha256:337a9fa660dc356e3f7e916be0394223ab8f1509430da36b25ce7b3240faf5a9

Observation 1aad171a-0f5d-4b06-a868-d2a11f8e1f47 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.802842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.802842Z digest=sha256:9feaa526676ebed100d97fddd4f181215b84f4c6246ff1953aac3dd96c51bd9e

Observation 45c9d40e-5cc9-42ac-aa28-f3ff1e0ebb9f · outbound

This paper cites In-Context LoRA for Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos In-Context LoRA for Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.806956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.806956Z digest=sha256:9d2a7bce4a1d38539905abc01aae7a1d9dd3c3bdd6a4a1fa5db296de828a3a3a

Observation ca726dfa-3ba5-4a90-8266-6570709a4c23 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.810403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.810403Z digest=sha256:9b11071db726073fa54a075a8045240385485baa72e9622d74a77f9d7e22b908

Observation 44c1d740-c104-4cd3-b809-a566076147bc · outbound

This paper cites REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.813743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.813743Z digest=sha256:711e6e9bc4dcaa1312ba3a4bc170100a669199098d26e0e94ba1264daaacdfab

Observation c48a0d8a-00f4-4601-bf67-0ed00ea735c8 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos VACE: All-in-One Video Creation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.817988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.817988Z digest=sha256:e360b1517edf460b75f16a4bb75d265bc4cb3163ae6b2fcca4216b34a72d6573

Observation 6483e68a-ab90-48ca-88f1-d0ec1eea21fa · outbound

This paper cites Faceshifter: Towards high fidelity and occlusion aware face swapping.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Faceshifter: Towards high fidelity and occlusion aware face swapping

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.543317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.821578Z digest=sha256:070181b37dfaa27e8776ccb3558ea56d9592a1721a1abc6a451ac7095134d2cf

Observation 6515f221-9296-4b83-9bc6-6fba7e6d4e29 · outbound

This paper cites Rolling forcing: Autoregressive long video diffusion in real time.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Rolling forcing: Autoregressive long video diffusion in real time

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.533176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.824578Z digest=sha256:f79ec2ea2f9745fbbee80828a7210a243ea1674fbbfe70239e9ba02dc12a424e

Observation 1de3ba52-1241-4d3e-888c-d7004100fa8e · outbound

This paper cites JavisDiT: Joint audio-video diffusion trans- former with hierarchical spatio-temporal prior synchroniza- tion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos JavisDiT: Joint audio-video diffusion trans- former with hierarchical spatio-temporal prior synchroniza- tion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.521525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.827507Z digest=sha256:f4953761530cfeb8582cfed6da94ebfa6fc19b2f3b5a29f482053121ee814eba

Observation 02750d3b-f5d3-46e8-8522-9a31d5734a70 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Zero-shot Voice Conversion with Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.830514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.830514Z digest=sha256:e427d6a95caccde011ba519d06d8370eb60a412257c105f739522a3f79e46b84

Observation f19a9fc5-caa5-49ba-9085-a4bbcfb44244 · outbound

This paper cites Scalable diffusion models with transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Scalable diffusion models with transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.511740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.833696Z digest=sha256:702073a79c7c930874eace97d3b92101cc8af1276bd5cb987b01162d8e6a1777

Observation c0e569ce-d332-4391-a1f8-166049140720 · outbound

This paper cites OpenVoice: Versatile Instant Voice Cloning.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos OpenVoice: Versatile Instant Voice Cloning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.836725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.836725Z digest=sha256:97d45d3cd595b037d70c3d6d8a3ea322afc3346e41cf9653986e86c29d51cbd4

Observation 7fea2ab3-3036-4c28-b478-673e796c16bc · outbound

This paper cites Sam 2: Segment anything in images and videos.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Sam 2: Segment anything in images and videos

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.501215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.840178Z digest=sha256:62d8a3cca8c50c2b825dbae4049714fe415c30354e2f005161256358c902f177

Observation 4b19e922-6906-4475-b323-9d24855ffcfd · outbound

This paper cites an unresolved cited work.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:34:45.491605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.843641Z digest=sha256:7fca4ce126376d386d32d76e782c7aeb49159e50181f7fbc54b03bba7351db73

Observation 53991683-2ff0-40f4-9991-609ffbaceea2 · outbound

This paper cites MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.481419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.847306Z digest=sha256:a7b4eff5928a2c00afcb31824d05dc5f8527c9e87ad2ee0fb6e091cbdbd2e7ad

Observation 257e803d-5781-4ce5-9c0b-659b4e0d74de · outbound

This paper cites Consistency models.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Consistency models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.850720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.850720Z digest=sha256:d84c56259e4d6c4312c3500fbd5e8ec6a418ec31a9562ce43d389fec7cf11981

Observation 9fd3a51e-658e-4533-bf93-f464a06f43ef · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Roformer: Enhanced transformer with rotary position embedding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.466367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.854000Z digest=sha256:3cc6fad0ff3760cc07dddbe588cc6ac53abb49a1aeed6f9136db728c1f910edc

Observation 99ab2679-559d-454a-a029-d90548a7a6b3 · outbound

This paper cites Omniforcing: Unleashing real-time joint audio- visual generation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Omniforcing: Unleashing real-time joint audio- visual generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.857252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.857252Z digest=sha256:2a96f851079290fcf7167a6da0bda9ba64725873836feddb3154fab432278504

Observation 89528c1c-f894-4ead-a3dc-9ca6444e4839 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Wan: Open and Advanced Large-Scale Video Generative Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.860514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.860514Z digest=sha256:b752b30691ef2b8349943708e9b32a036d5a6198bab6971b33196791959249a0

Observation fe34031a-b2e8-43af-b62a-73bcf18a7490 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Generalized end-to-end loss for speaker verification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.456924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.864312Z digest=sha256:53ad293d2111104e309322b4b1652d21b8502fd5c512f759364c5367fbd9c1a0

Observation 58307b3f-a855-4825-a056-cf32436d8b5c · outbound

This paper cites Bovik, Hamid R.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Bovik, Hamid R

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.447706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.867781Z digest=sha256:7d2256d0f702f2335f6abf406c505f79123c4b0bf527c660d4b215295390a321

Observation cf83fc06-4075-4d70-aa5f-d8ce9cba181e · outbound

This paper cites Williams and David Zipser.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Williams and David Zipser

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.438474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.871050Z digest=sha256:00c429386734581045aa4fa83b9cec03d36b57747ab8ec528abe4abd0ad9822b

Observation bb0d50fc-89bc-41b0-a29b-4eedf4bd61e7 · outbound

This paper cites Q-Align: Teaching LMMs for vi- sual scoring via discrete text-defined levels.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Q-Align: Teaching LMMs for vi- sual scoring via discrete text-defined levels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.429404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.875183Z digest=sha256:f55818fbb546ac1b2c9e3c5b879d01b9a7d9afcc647df6334331376174a6e390

Observation 9a94d4ea-e773-4e0f-b6b3-5873782d1241 · outbound

This paper cites Efficient streaming language models with attention sinks.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Efficient streaming language models with attention sinks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.419609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.878806Z digest=sha256:6fe1d7870934a5950a56f11d55c53ea49fe83cc3c91be3df2fb843dd8911c1e3

Observation afb43cf0-ccee-4f73-add5-d82decaa9806 · outbound

This paper cites Vit- pose: Simple vision transformer baselines for human pose estimation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Vit- pose: Simple vision transformer baselines for human pose estimation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.409786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.882068Z digest=sha256:1fd8fa55407486d6637c7c6ce78c85ac78f764b013bd88c3519bfa07a0015bf8

Observation fd0df51c-5499-4fe8-b510-0302c302c4da · outbound

This paper cites Mocha: End-to-end video character re- placement without structural guidance.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Mocha: End-to-end video character re- placement without structural guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.885408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.885408Z digest=sha256:5f251a13aaf64fda47bf233028bc2866bf418caafee870bfd51e077d349a3373

Observation 34d99a8a-e375-42f7-846f-178ac4403970 · outbound

This paper cites SCAIL-2: Unifying Controlled Character Animation with End-to-End In-Context Conditioning.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos SCAIL-2: Unifying Controlled Character Animation with End-to-End In-Context Conditioning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.888795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.888795Z digest=sha256:96b8ddd8c32fd557bf9b866ee9fe512127e8c2fd5cbdc557bcfa681842cb3718

Observation 075238d9-9d63-4afa-9e59-aabaf7e7e446 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.892364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.892364Z digest=sha256:7c0e8380ba4d87244af05c812d5ecac57f9edad2b112851e0616f3cc56118b3c

Observation 92a404c2-381d-4a3c-ba4d-5b6e4d836660 · outbound

This paper cites an unresolved cited work.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:34:45.399087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.895952Z digest=sha256:107ed2d8655be1928ed8514182b6e44cf9f3573196c3ee010e63f735e245a85e

Observation ff9bd18d-c2e3-4c90-b0a7-b8a03b906b33 · outbound

This paper cites Freeman, and Taesung Park.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Freeman, and Taesung Park

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.388288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.899017Z digest=sha256:7d456a4e0fa3fa1733aef3b6fef4dfe5a74ba3576849de37abd17640a988dcd4

Observation 4492e9ed-aa5a-494e-b2ff-17a36c0cc83b · outbound

This paper cites Free- man, Fr´edo Durand, Eli Shechtman, and Xun Huang.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Free- man, Fr´edo Durand, Eli Shechtman, and Xun Huang

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.377929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.902081Z digest=sha256:6c5d397ed2d05864dc8ba05d2bcbaa4058b742e0932855e476116471ea6956c6

Observation 0c368f77-fe29-47d4-9c41-cd7a5dc58c3e · outbound

This paper cites The ablated variants exhibit increasing identity drift and vi- sual artifacts in later segments, whereas the full model re- mains more consistent.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The ablated variants exhibit increasing identity drift and vi- sual artifacts in later segments, whereas the full model re- mains more consistent

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.367030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.905261Z digest=sha256:0060d90a4c388e543d2166f55a4438c35aeee88d9f12b62483b5aa3b7209c400

Observation de58b078-ca90-434e-af6a-44373ebb30cc · outbound

This paper cites The reference cache persists throughout generation, source keys and values are tem- porary, and completed target blocks are committed to the clean-history cache.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The reference cache persists throughout generation, source keys and values are tem- porary, and completed target blocks are committed to the clean-history cache

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.355170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.908491Z digest=sha256:6fdb1e2f8d52cdb70a716a263563f1adcae0e8be7cbfec0f58f97aa1f3ae94a5

Observation 896a7ddd-61ad-43c7-bf44-608589ad03ca · outbound

This paper cites The study compared UniSwap with four video-replacement baselines, each paired with Seed-VC following the cascade protocol in Table 1.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The study compared UniSwap with four video-replacement baselines, each paired with Seed-VC following the cascade protocol in Table 1

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.345225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.911900Z digest=sha256:51ffb0fe3888b946a6d5066d06d4d8a6918480daec0c0d8d7286840f92d51940

Observation bc9a1f1d-3ab5-4c1e-a6e3-b7bcdeca3c31 · outbound

This paper cites In all figures, each example con- tains a reference image and reference voice clip, a source video and its audio, and the joint audio-video output pro- duced by UniSwap.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos In all figures, each example con- tains a reference image and reference voice clip, a source video and its audio, and the joint audio-video output pro- duced by UniSwap

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.335051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.915155Z digest=sha256:f44438afead22de5e2436c4747be1d573bdb86f6e29e3f6507c305fa988fa55f

Observation a15be8c3-6fd1-496e-9ff5-4214f51def9c · outbound

This paper cites an unresolved cited work.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:34:45.323486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.918493Z digest=sha256:ed3caf05113e550dfeff4ef2d40abff2ffd925735903943166780e1282314a62

Observation fd69542b-0ccd-45b2-8e7a-6dea3a7af39e · outbound

This paper cites Deployment should require consent and provenance mechanisms, visible disclosure where ap- propriate, access controls, and compatibility with forensic detection tools.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Deployment should require consent and provenance mechanisms, visible disclosure where ap- propriate, access controls, and compatibility with forensic detection tools

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.312445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:44.921475Z digest=sha256:f24d18962585f4250f07397107f0cb0748a32a6d60c1dbf50f2684b352fc21ef

Pith citing papers

No inbound Pith citation observations are available.