Pith. sign in

Paper Citation Record · LEDGER

MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2410.10122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10122 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:44:49.755514Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.011253Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d1d351a7-3f33-4d31-946a-2e7787663330 · inbound

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model cites this paper.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.080042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.080042Z digest=sha256:b841700d1cdc2b9962f8a0a7a7d682d20d886eeac24880bc91c6798ad93b9010

Observation 797b51f3-e5ae-4cab-ad66-8478e9a0117a · inbound

LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision cites this paper.

LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T17:11:02.250238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:11:02.250238Z digest=sha256:430dfc88822a92b4badc37f69a91b4fae47ef0863a5a0b8bb42ddf8813e683c3

Observation c0353239-26a7-493e-b74b-63e4103a90a6 · inbound

KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution cites this paper.

KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:44:49.755514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:44:49.755514Z digest=sha256:1deee7e4a4e726c1e3fe3fa5db2b0572af06d25600ccd7fe17a7a28d00631940

Observation d25bc7b7-a3f2-4947-9dd9-cf59745c94b0 · inbound

Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network cites this paper.

Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:13:31.294970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:13:31.294970Z digest=sha256:f29e220c137b2debda0b49c06a50e3e3dc18944720d0df7ef17d866df0f5ec3f

Observation 011b60a5-55dd-45ba-9a0e-fc1025065282 · inbound

Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos cites this paper.

Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:46.730807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:46.730807Z digest=sha256:aa85a999db0c3dca004d7f9666308f97de666fb4250ab1e8151f428f52ff48cb

Observation 3bf6e830-fa96-46aa-8213-d2887cacab16 · inbound

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture cites this paper.

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.390763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.390763Z digest=sha256:15ecf9f6ae9661e478bc926e780008bd56d930c7583d62737b73a1cd243bba8a

Observation c045272e-bab1-4bf1-a1f3-b3c58fd6f7c3 · inbound

Fine-Grained Zero-Shot Object Detection cites this paper.

Fine-Grained Zero-Shot Object Detection MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:41:02.903909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:41:02.903909Z digest=sha256:09c028f5e69ec99ad038d23716e19dae01ff7e80162f2f80f4ed77453b1a6c51

Observation 74d21d0f-12d3-4872-b90f-f88063bff5ab · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:57.033054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:57.033054Z digest=sha256:96f73b4df4e325866f97469500c973cf8ef2493f534d04303c919486af09d158

Observation 9811c709-e290-485d-8a75-120a73b6c8ff · inbound

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync cites this paper.

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:35.346930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:35.346930Z digest=sha256:71d4927a2f1bc2a82b15e83f4de3beb09793c02c2e10bcf5103c1c1d5ae9b5c9

Observation 554011ca-c2db-4c6d-bc9a-557ade5b7d1b · inbound

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads cites this paper.

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:19.154197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:19.154197Z digest=sha256:1c0749ef0f7f2c40e46c79a1dfaf446f23d4f371f613b19664d750f6672ab0eb

Observation 4a19a2d5-4538-4199-8f15-632ebf821315 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.311819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.311819Z digest=sha256:80ba92b68a9e0f1aa4c7957a6baeec6016d204efe6192d7c47b28c0605f29255

Observation 90ff5ada-57db-41fe-bf4f-0ecb1f85de0d · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.818772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:3c923bc035e9b08a097cfbd026b27867d5fd1f563e49b4fd9690bdf29a5b8297

Observation 88af652b-816c-426b-9966-cf5ed06e921f · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:52.881855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:52.881855Z digest=sha256:4620276f6d188586c169d97a3edadabd8830806a29593c0837c7585311acb16c

Observation a37a3846-f6ce-4a5e-98c0-87cdc70ff99f · inbound

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence cites this paper.

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.231717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T08:27:41.839123Z digest=sha256:39b0c9fa0d98969c3ffaf426511f956b14cd9b3dac9a1178b848e18266031f25

Observation 2ac94527-4d4f-4665-8b36-b66dede390be · inbound

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation cites this paper.

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:12.895871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T06:56:19.795651Z digest=sha256:cb8a1c0a46cd6ffd771060fdd815386bc51d33906873a4eba1cf3c5612b6a867

Observation 8a48b7fb-336e-4f81-b0eb-43feeb2e5363 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:03:50.705507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T23:01:15.886434Z digest=sha256:6a4ab7032ff457acb7763a0677319b02c01d960b29de966e184d65373a39e3d1

Observation f0bcc245-6bbf-47ea-985c-c15129de1490 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:54.590138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:54.590138Z digest=sha256:20149e827e092e9e25aeb2b9ecb80811a43532ed8dcbf954a322a9836b257297

Observation 5914fffe-a6e0-4506-8cd8-579769c5f438 · inbound

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models cites this paper.

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:17:48.632095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T21:14:25.606097Z digest=sha256:801a26986025f7a96dce4e3270377c2c85db0e59c6d27526ccac12c7a66d85ca

Observation 5d5100b1-824e-4d2e-83d1-65e4672194d3 · inbound

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization cites this paper.

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.012668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T13:40:14.506288Z digest=sha256:bc8af7317a062044553adcd743df26d57c7ed0c94c520707211c1b6bffdc3962

Observation 661c553b-0133-47e6-8d93-e98fdc407a11 · inbound

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations cites this paper.

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:33:51.213900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T05:04:31.366845Z digest=sha256:79f5c9078446aa7e3fb077daa3ff9845c6f01c8477ebd2f52c271462c6f890ff

Observation 511abe14-d690-4b64-8f87-673f607e6ae4 · inbound

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization cites this paper.

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.569495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T01:07:49.064263Z digest=sha256:a23b6a5e5eb641a8357361fcb039881e299d04972f89dead88320075a3c518b2

Observation 820bd89c-a5ba-4416-a4d3-cb2c25c59470 · inbound

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications cites this paper.

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:52:51.792816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:52:51.792816Z digest=sha256:db549d5e9c68a75ebeb1d350a3402d808b2e8f7396b9f13532ba375f6a63e63e

Observation 21d2069e-e574-49a4-8cdf-c505334d2d27 · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:12.897470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:12.897470Z digest=sha256:35717d2601e65da2b33654f9cca9681513965655ce77fae31515d50ad664cc4d

Observation 2a6fab00-b4cf-4027-81d2-52a40deb78b8 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:13.376570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:13.376570Z digest=sha256:c23cc817b112af81b173412ddbe9bd5d05efe8ea15a35bf0423d07c96d6c892e