Pith. sign in

Paper Citation Record · LEDGER

MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2410.10122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10122 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:20:13.390763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.011253Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3bf6e830-fa96-46aa-8213-d2887cacab16 · inbound

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture cites this paper.

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.390763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.390763Z digest=sha256:ff89a7b40d04d8197ffb75354287311ddc65c5a70b94434dd46c495580f3a3df

Observation c045272e-bab1-4bf1-a1f3-b3c58fd6f7c3 · inbound

Fine-Grained Zero-Shot Object Detection cites this paper.

Fine-Grained Zero-Shot Object Detection MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:41:02.903909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:41:02.903909Z digest=sha256:7cead6d143e792b4ca8f698bb425c2c6c9bb2c81430cafd1e9ef2e90897d20ba

Observation 74d21d0f-12d3-4872-b90f-f88063bff5ab · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:57.033054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:57.033054Z digest=sha256:dd44cc31691b4e35661c81c4c7e10b6d6c6fb92d566fa01e863d238b90dfd500

Observation 9811c709-e290-485d-8a75-120a73b6c8ff · inbound

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync cites this paper.

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:35.346930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:35.346930Z digest=sha256:133ed1b7faa5b3f654c5442f02e488c65fdd5998ac8b77bd2cce566738cba521

Observation 554011ca-c2db-4c6d-bc9a-557ade5b7d1b · inbound

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads cites this paper.

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:19.154197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:19.154197Z digest=sha256:96a88971bcc64cfbc761c0dd6d47c38b76b582b7ba7eae24518a8e5bb1061f42

Observation 4a19a2d5-4538-4199-8f15-632ebf821315 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.311819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.311819Z digest=sha256:b4553253409b8f46283ca0f58a2faa565fe38cfe3e72f662c6fec1dd0c6dac58

Observation 90ff5ada-57db-41fe-bf4f-0ecb1f85de0d · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.818772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:859b67370955c06248641f9f15812dc401abd89d3639652f4c06c317efe6adbf

Observation 88af652b-816c-426b-9966-cf5ed06e921f · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:52.881855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:52.881855Z digest=sha256:490e28cd8dc1b951e703ed25dd73875fca8afc474899a72372d29b90935593dd

Observation a37a3846-f6ce-4a5e-98c0-87cdc70ff99f · inbound

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence cites this paper.

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.231717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:27:41.839123Z digest=sha256:c9e63348b14de3254b413c376a69144f5345b49f89a59bb0e09e192a63ad3177

Observation 2ac94527-4d4f-4665-8b36-b66dede390be · inbound

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation cites this paper.

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:12.895871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T06:56:19.795651Z digest=sha256:fbb6ea4b1467323f28bc4d9f0d8eaa8d053d1f487724e76fdd76ba0339df2311

Observation 8a48b7fb-336e-4f81-b0eb-43feeb2e5363 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:03:50.705507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T23:01:15.886434Z digest=sha256:d99bdd73eddf64d0763113d56aa9df6d0dd7d2d6fb503c5e7d6e3d78a5a9f52e

Observation f0bcc245-6bbf-47ea-985c-c15129de1490 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:54.590138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:54.590138Z digest=sha256:3ddac6d65f8aeff33f3f8766f2371e9913a89a5547cdf0670a09c58a5c5353b6

Observation 5914fffe-a6e0-4506-8cd8-579769c5f438 · inbound

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models cites this paper.

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:17:48.632095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:14:25.606097Z digest=sha256:2a3d1806388310840ec363989cd11dd0406b041ecd174858ea7cc6a06c546f34

Observation 5d5100b1-824e-4d2e-83d1-65e4672194d3 · inbound

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization cites this paper.

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.012668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:40:14.506288Z digest=sha256:2b4d23604d27269c3112e7a72a81c37177113b633feff39831ccecbbe88b4e81

Observation 661c553b-0133-47e6-8d93-e98fdc407a11 · inbound

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations cites this paper.

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:33:51.213900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T05:04:31.366845Z digest=sha256:d38d916447706275108f76525121a3faac7a44c3bd75e3083288478e1a800a40

Observation 511abe14-d690-4b64-8f87-673f607e6ae4 · inbound

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization cites this paper.

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.569495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T01:07:49.064263Z digest=sha256:429b8be4889768994392b7cbfb7741137b474580393fd97e756dcebc076871ba

Observation 820bd89c-a5ba-4416-a4d3-cb2c25c59470 · inbound

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications cites this paper.

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:52:51.792816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:52:51.792816Z digest=sha256:79e6adb3f1845253612bbdd4eedc412a8d9500ad8f156ed6e95bb1c9fdaff7b1

Observation 21d2069e-e574-49a4-8cdf-c505334d2d27 · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:12.897470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:12.897470Z digest=sha256:87f5e318e9166187c8dff6e640f0b7ddd4ef6ecfa8fb7435b915c25d1e3cae06

Observation 2a6fab00-b4cf-4027-81d2-52a40deb78b8 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:13.376570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:13.376570Z digest=sha256:a82e0e9bad57bb249ab24a8c0330e50f87d7da3027eb7aff677c7b008130308e