Pith. sign in

Paper Citation Record · LEDGER

MoCha: Towards Movie-Grade Talking Character Synthesis

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.23307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23307 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:47:49.015470Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.815940Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec5c9838-5c65-4c45-b545-a00c7f0754d3 · inbound

Seeing Voices: Generating A-Roll Video from Audio with Mirage cites this paper.

Seeing Voices: Generating A-Roll Video from Audio with Mirage MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:16.914776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:16.914776Z digest=sha256:bb5c2d2f6fe6b3dae52181a24192af7356f92dec71907e0b918115c105994a83

Observation 31529f22-2b38-45ac-bb66-f124081eb677 · inbound

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation cites this paper.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.890139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.890139Z digest=sha256:390acaaf8810b41518ef71485ccc14c69f252fdef021348e644883b1b2025065

Observation cd64ce50-fb39-4b1b-9e39-8cc6ab8c717f · inbound

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation cites this paper.

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:49.015470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:47:49.015470Z digest=sha256:1ce0eb563f8db5b2855bfac6966a2723147eff631069e64ca2e9e75f4c87a10f

Observation eb926631-9f96-4f14-965e-53cf9066cb53 · inbound

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router cites this paper.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.134588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.134588Z digest=sha256:9fda0f47a1c6af3fa2292e762f8bd6f099ae882ee4149d86d9ece4e8313532cd

Observation 452247b6-40bc-43e3-8f01-2e4c146252fa · inbound

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation cites this paper.

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:46.885872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:46.885872Z digest=sha256:5f9ad8eb04244d456e0e29d13f7cc9183d375f6763101d7e40c8a301450d7f08

Observation 89c8071b-2f94-44ca-9b2a-8b25c0f7a1bb · inbound

Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach cites this paper.

Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-15T17:41:20.620092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:41:20.620092Z digest=sha256:432a570242f26c9a6e3dcde6f5c6112e34b6f2422f57232df6ee46a5c5511a00

Observation 9416b590-3d2b-47f4-9035-2dc81d9e28f6 · inbound

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering cites this paper.

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:49.072437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:05:49.072437Z digest=sha256:378f8b5c20280f9b9b89a38e48d26507a3e36120544825686fd1d00edd0913a5

Observation 5f33149e-4ab6-4e21-a093-41c8dfc4dc2f · inbound

ShoulderShot: Generating Over-the-Shoulder Dialogue Videos cites this paper.

ShoulderShot: Generating Over-the-Shoulder Dialogue Videos MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:24.880230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:24.880230Z digest=sha256:abf0816107f0ad980fe93c0e4f068a5850709c6813d0b8abc5ccbcf692aa4694

Observation e3481134-6850-4cd3-8ffa-015982987b44 · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.790252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.790252Z digest=sha256:ed955a2c39f7a1f199d0949a6ef8025adcf5b6f5681b766e3d3dc4747cb375f8

Observation 54d53228-2a20-42c5-9b26-de6fdfe00a92 · inbound

LongLive: Real-time Interactive Long Video Generation cites this paper.

LongLive: Real-time Interactive Long Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:52:59.524719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T03:52:59.287555Z digest=sha256:c6f38a66dc95b5f301f8009117fcdbe86b6bcfce5699b29b7bce0e530f72092d

Observation 0c831259-bb3b-4e91-bec4-e4b1de6d4d7a · inbound

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation cites this paper.

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T13:04:43.447804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:04:43.447804Z digest=sha256:df2817b8f2b09c9fa130dd07422ac78b151d3c1a888960695ac202bef576ed30

Observation 37e0ceed-b7c1-4a73-b5d4-090dac5edc33 · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.281694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:da84ef8b385d75e404ba7820707b5c5902de3bd66e43dd85a2570bf50407df87

Observation eb69789b-82a6-45ab-bb6c-617887a56a9c · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.055665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:e96f0e0f0fc3c87ac21039f18d23366dafb69c0a30d7252839732fef7c7b4e2f

Observation a130bcbf-8fcc-4e42-9d4f-0144bfca2b0e · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.817537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:85377803e43e693c0c7e4e7b3f598a58edf714052218a6b4e589218f2e0eda42

Observation 80430d9c-9cc8-4bf8-b34f-840614c7b596 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:49.263380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:49.263380Z digest=sha256:23d259067741f806d09c79642f6e64b41a922f98604db0006dde13d8aedabd06