Pith. sign in

Paper Citation Record · LEDGER

MoCha: Towards Movie-Grade Talking Character Synthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.23307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23307 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:16.914776Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.815940Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec5c9838-5c65-4c45-b545-a00c7f0754d3 · inbound

Seeing Voices: Generating A-Roll Video from Audio with Mirage cites this paper.

Seeing Voices: Generating A-Roll Video from Audio with Mirage MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:16.914776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:16.914776Z digest=sha256:157259f88af6c5c8e14399c40f56219301671688e4c18a18e9590f339fe178e8

Observation 31529f22-2b38-45ac-bb66-f124081eb677 · inbound

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation cites this paper.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.890139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.890139Z digest=sha256:dcbcb81d12ed5f45012c59865bc418677cca47d92f8495b38f76ae1100392806

Observation 452247b6-40bc-43e3-8f01-2e4c146252fa · inbound

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation cites this paper.

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:46.885872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:46.885872Z digest=sha256:ab308a6094b673f959222494176dffda0bf329d58d4626f14fd7992f9f47f1a4

Observation 9416b590-3d2b-47f4-9035-2dc81d9e28f6 · inbound

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering cites this paper.

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:49.072437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:05:49.072437Z digest=sha256:7df7045d7ef8607f13e9fec7eb55d7ce6264133d80d48f5fb03e954eaff08587

Observation 5f33149e-4ab6-4e21-a093-41c8dfc4dc2f · inbound

ShoulderShot: Generating Over-the-Shoulder Dialogue Videos cites this paper.

ShoulderShot: Generating Over-the-Shoulder Dialogue Videos MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:24.880230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:24.880230Z digest=sha256:3a26fa9818ebb03a7f325dcaac59216df6368339a2fb860ed7ba27b9c9ac89d9

Observation e3481134-6850-4cd3-8ffa-015982987b44 · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.790252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.790252Z digest=sha256:b1d04764eb70ac82697ecd6a626b5c468393f851da367669a88dbdd310df14bb

Observation 54d53228-2a20-42c5-9b26-de6fdfe00a92 · inbound

LongLive: Real-time Interactive Long Video Generation cites this paper.

LongLive: Real-time Interactive Long Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:52:59.524719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T03:52:59.287555Z digest=sha256:71cad4dfd1ab3bf17a07dd2313ff3411c6a1c948b84d1748501a3f77f876fad6

Observation 0c831259-bb3b-4e91-bec4-e4b1de6d4d7a · inbound

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation cites this paper.

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T13:04:43.447804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:04:43.447804Z digest=sha256:946456bf04b5e6a2360df618cbf73ec76b45317f32e711943f92b9eb6fc5e342

Observation 37e0ceed-b7c1-4a73-b5d4-090dac5edc33 · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.281694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:785966128f682694f739a8f56fea25f2914d4f1660183e328cb13087309e647f

Observation eb69789b-82a6-45ab-bb6c-617887a56a9c · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.055665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:af77724686fd9553f0623b91e8c4d607686c140fe5058fde951ee5baea396bcd

Observation a130bcbf-8fcc-4e42-9d4f-0144bfca2b0e · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.817537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:9e037719bd472e4d065c66ab12de325a8385f223f03a9723119739f811062b85

Observation 80430d9c-9cc8-4bf8-b34f-840614c7b596 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:49.263380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:49.263380Z digest=sha256:fcb1f17ed1a0e554e7d5a58d3303d46e500527fb75b3b8894782c27ba91287e5