Pith. sign in

Paper Citation Record · LEDGER

MuteSwap: Visual-informed Silent Video Identity Conversion

As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2507.00498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00498 v3

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:53.157881Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact6
  • verified fuzzy3
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9291ee37-5d5e-4583-9f71-c7208b032c31 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

MuteSwap: Visual-informed Silent Video Identity Conversion , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.147531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.147531Z digest=sha256:7e796c20f1251776ff0db75f943a02891a42091c772985a2a6956b76b3f23bac

Observation 5706bf15-168a-4b5c-b8d3-606e51028c32 · outbound

This paper cites write newline.

MuteSwap: Visual-informed Silent Video Identity Conversion write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.171962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.171962Z digest=sha256:3863bdf8408834e7339ffca00a63922137049f4570dc9c757b6779056f2e244d

Observation 9d74eef3-79db-482a-95fc-45a46764d337 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MuteSwap: Visual-informed Silent Video Identity Conversion LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.211857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.211857Z digest=sha256:c78809b5a3177a8e0e290355cc5dac6aec148b29f7cc6a0946992ca874684141

Observation fe53fff1-4baf-44a5-878b-030fb9ff06e0 · outbound

This paper cites CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information.

MuteSwap: Visual-informed Silent Video Identity Conversion CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.154005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.301107Z digest=sha256:b0a96d278a448a9f8cefe0da28943c7a03a60074011294903b3893aa07b45638

Observation f1bd61be-3db6-4f8d-9f3e-756e67d0be9a · outbound

This paper cites DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding.

MuteSwap: Visual-informed Silent Video Identity Conversion DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.028479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.343362Z digest=sha256:5b6009ca997eb5a5291882df605542c220ad09539b6dcc411e2e8574661a7f18

Observation d0dcf465-cbdc-4f90-8470-770e0df21ade · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

MuteSwap: Visual-informed Silent Video Identity Conversion Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.402047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.402047Z digest=sha256:e9f18163d3053845a9e401322f8ef94f921060e86ccf9f651cf838be33aa8597

Observation cdd914df-a71d-42fe-a688-a570af45917b · outbound

This paper cites S.; Nagrani, A.; and Zisserman, A.

MuteSwap: Visual-informed Silent Video Identity Conversion S.; Nagrani, A.; and Zisserman, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.919670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.480178Z digest=sha256:c905df5c2e3d38349dcd22970274d74b35ca8d36a41fdb3dfd30113e9bc57677

Observation 9c492056-dde3-4364-b09b-f71204d78236 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

MuteSwap: Visual-informed Silent Video Identity Conversion Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.504954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.504954Z digest=sha256:ebbb3d676b502a9f27ba077503017516e17a1522f9778a3c8eb48b5353529e87

Observation cd2426c6-f426-4597-933d-b22d65292576 · outbound

This paper cites PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association.

MuteSwap: Visual-informed Silent Video Identity Conversion PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:19:53.919077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.555712Z digest=sha256:5760385b68321dd40b30819a1644d69747fb6ef2206f8fda59e38844ad77d32b

Observation 56aaea29-bba8-449d-a66d-fb3c40cdc900 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.720531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.633979Z digest=sha256:ea61e9ed4020a25e59a1d6e92429859d0db77c6cca9338d5fb14fb7eefc5c464

Observation 3dd5f74d-27b0-4753-bda5-0a530fe82af2 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.479759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.670519Z digest=sha256:f7bf1f1a2db119b52551aff8dc527274fdf55f0909c5ff20378b6986ca887aa3

Observation 72acfcf5-f43a-4759-97bd-069837665ac5 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.294575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.730095Z digest=sha256:33c84ccfe5bd9811a45a5f6eafb525db50c4b4ada747b8efce07be83d16c1573

Observation ead1e5f4-0116-4f3a-9204-a73e4d863fef · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.800306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.800306Z digest=sha256:0d700110addca5be72ca3a09fb645055c571df5ad486a5a4b71de9fe10c0440c

Observation d44b3b0f-7301-4d36-8f68-2b9e7b233f15 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.104956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.840171Z digest=sha256:54059fb296d477ba15904e3fdbaa030b48f7a6ca46720d9f8eccf7907f8b5ab6

Observation 56964d59-c821-4aad-80d6-d7417c6160e2 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.887753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.901215Z digest=sha256:019e7d9dc0ce5122cc162b8af58495e314c004beb399688ff8f6200d19765636

Observation 3f95c228-93b5-4a97-a1a6-c4029cf95e2f · outbound

This paper cites DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility.

MuteSwap: Visual-informed Silent Video Identity Conversion DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.799591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.982678Z digest=sha256:935545893491c035d762037d34f18b971be44b4c2243bda932e2381ce90ee68b

Observation b0ef3fd9-2231-4fd7-81ed-c7e16396d90c · outbound

This paper cites Decoupled Weight Decay Regularization.

MuteSwap: Visual-informed Silent Video Identity Conversion Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.026803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.026803Z digest=sha256:26fa4361226d05a12cf6754072cd73e6da9bac2565dbe388b696ff79a39c3080

Observation 42454f16-fcba-4cd1-8889-8dc3634c7aaa · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.650955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.102249Z digest=sha256:5bbb8de8ea9d0e9938737b3fe30db3ad739fd5d71ecbd6460bd8e12e3addc3f6

Observation 2349b402-a25c-4f30-8019-de0503723d04 · outbound

This paper cites SVTS: Scalable Video-to-Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion SVTS: Scalable Video-to-Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.154124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.154124Z digest=sha256:50ede691b2034027002704b2dcef6ee117eb2a42299b0fe550dd88265dcbda0f

Observation 1ddb64d0-3d9b-4bab-bbd4-67d91f2768f6 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.361477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.238451Z digest=sha256:4fd821e31013762dd90519e4f67178e6a12eba6dcb27e82e37e6a32948f705c4

Observation 8c51ab8f-2738-4bce-b9e7-13c903ae60b7 · outbound

This paper cites Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:19:53.678456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.300850Z digest=sha256:bd2bd4f9eaa21e1f78dcf6bad7d37a35ce368e2e45aa0955d853e9b7b2ac36a6

Observation aa72b93c-a3f2-4fc3-b831-72dac0197c3a · outbound

This paper cites R.; Mukhopadhyay, R.; Namboodiri, V.

MuteSwap: Visual-informed Silent Video Identity Conversion R.; Mukhopadhyay, R.; Namboodiri, V

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.151833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.373040Z digest=sha256:d9a40ef203e6f8c4a28177f5802b3b9476e5411125f4fa6828901f6398ed4be1

Observation fc331158-e15a-4a77-a0e5-f0ce5bbd55f1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.448733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.448733Z digest=sha256:c97fcb5e669fe6d5e5f5c4e7f46179126fe5e7fdd4299d854cb1328fc40045f5

Observation 74b17056-4bc1-417d-8f27-8f6839385594 · outbound

This paper cites Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion.

MuteSwap: Visual-informed Silent Video Identity Conversion Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.571796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.494181Z digest=sha256:3d30c64a88adcd3ad9e63decc065690987301c030405688320513337d1f4b8a3

Observation e1c57262-d27c-4f3f-85df-687f200d9f9c · outbound

This paper cites Fusion and Orthogonal Projection for Improved Face-Voice Association.

MuteSwap: Visual-informed Silent Video Identity Conversion Fusion and Orthogonal Projection for Improved Face-Voice Association

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.437781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.564281Z digest=sha256:c1850d85cebed9326454ed23ff741105914a4059e32b0d89b5d0eb7cdac09cf6

Observation b0cc4085-4273-41c9-9173-25cbc1c6259d · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.976495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.620787Z digest=sha256:88066299a36ae95fb30afd26e6d2632f52ab6f53933b1c34999b924620da98b9

Observation 3fcc39cd-9f1e-49b8-8764-0adcdd1345be · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.685389Z digest=sha256:3a6716c8d552cd391db9b47615b03eece6d7a502a9ab509dc01dd49bec5cf51e

Observation 8cb0660a-527a-42db-9ce8-43d9d600d0a5 · outbound

This paper cites Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.752908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.752908Z digest=sha256:2e2b25b4b9eb10f3cdbb7115bb262e2707f77ea8eaa08163f4038e3e5c1429d7

Observation 883da9ac-52a1-46e5-a25f-59b1be6e6c70 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MuteSwap: Visual-informed Silent Video Identity Conversion MUSAN: A Music, Speech, and Noise Corpus

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.819201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.819201Z digest=sha256:6742a24a8363abbba909df8897594e4938402ca381a4a2ca599b2d183e791e39

Observation d725aa85-d2f9-4f6f-98c3-7f58e0c27f87 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.737554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.880819Z digest=sha256:bf4683e253707c751bb5e2228022ec6d40d059d61a768b4839de6394fefc2c3f

Observation 867fd800-a343-4e75-b653-fd511c8127b0 · outbound

This paper cites T.; Chen, X.; Liu, X.; and Meng, H.

MuteSwap: Visual-informed Silent Video Identity Conversion T.; Chen, X.; Liu, X.; and Meng, H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.506179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.983030Z digest=sha256:ad1778ef93dffa32770cf469752fae1957689181deaf74e6f91dde7bc8decb39

Observation 0ff7a1f0-01db-4f74-b9cc-9a39661c1e70 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.402901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.056506Z digest=sha256:7ba72a09f9240d6ae4695c6fa60df03aad9d2215fad84d569c962ecd26049d3c

Observation ef9ea61d-c62b-4b58-aeae-317c2391e166 · outbound

This paper cites LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading.

MuteSwap: Visual-informed Silent Video Identity Conversion LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.314800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.084783Z digest=sha256:b557eaddacf1fd7dcc9b1290881e30c99b8c3d5292fe9c2e8ba8a627087f182b

Observation 9e976be1-c64c-4263-a750-d09811322450 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.271812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.157881Z digest=sha256:50774614303766056aec9a53130b35235cf07902c7c00c01e93c0e233e207669

Pith citing papers

No inbound Pith citation observations are available.