Pith. sign in

Paper Citation Record · LEDGER

EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2501.10687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10687 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:47:13.040011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:08:24.974505Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5afc09ef-6566-4fdd-991a-13e24d1e0f52 · inbound

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models cites this paper.

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T16:47:13.040011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:47:13.040011Z digest=sha256:7500b70160162e144fad4952d4ad236a13837d69a55bf7291deb76007339ced2

Observation 3db3c4ac-a43b-4926-b51d-162747b2aa91 · inbound

DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation cites this paper.

DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:37.744863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:39:37.744863Z digest=sha256:17cf67a6a55ae40e286b3d9bacf8a808842267280d901cac57635801f0d452d7

Observation f2aa37e3-d4ee-4e5b-8ad2-d4f0164a6dd9 · inbound

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation cites this paper.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.372279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.372279Z digest=sha256:557465261c00d7bb351d85db1f614e758fbd0fd79b6bff4fa25b31d65e699a48

Observation 71ee290a-41c2-40a6-a35c-52be9b948044 · inbound

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation cites this paper.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.850843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.850843Z digest=sha256:59a7976acaf7d87bf428cdb907e41308efe6688549ce9e366646dce05b9c414b

Observation 6edac27d-81d2-4a35-8790-f58beb70f529 · inbound

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation cites this paper.

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:39.085973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:39.085973Z digest=sha256:cb6fff4124240d79a6c3b33cba34087ec323c83c43276375dfae625d38e009c8

Observation 66570f13-5e7a-4442-9bca-2b84f46e50b1 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.273895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.273895Z digest=sha256:af9c00a9a0ab8160d4d6b7bfa48246a7575f6b854ff832a54323910447de5351

Observation 6897fdfb-9ade-43ff-a776-5b1078b0f010 · inbound

Wan-S2V: Audio-Driven Cinematic Video Generation cites this paper.

Wan-S2V: Audio-Driven Cinematic Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.025790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.025790Z digest=sha256:96d98db0b32d68ef60319930fc39284808e402aff558b623f9e50a05b9700bbb

Observation c3902634-2cbb-4768-91eb-efa0626f97cf · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.783917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.783917Z digest=sha256:b5418bcc9898f0156acb09ced73f034bfed8e17af16081b2f2377ef6c922b72b

Observation 87f6a9f6-5adb-4993-9e9b-a852732d3937 · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.300349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:c9c67377b4184cb343bbc1a841288c54db853714c4b7fb95cdb409800dd89b56

Observation 05e49444-5004-4e83-a3f8-573441607464 · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:57.529763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:18:58.996355Z digest=sha256:af5087f58d9474f0335fb1092f8bb9bdd1fd7e4564ab2f3d7a5406c094591ac6

Observation 2935ab38-1ffd-44af-a3e3-0a45cc479da8 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.976159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:04b65670abfbcb1512f953f8b724951230415ef8578877ed3e1394db27fa2f66

Observation 10763ea5-837b-46f1-9cae-d67df8a3e3ae · inbound

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing cites this paper.

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:38:14.731854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:34:32.558440Z digest=sha256:77b6a7ca3a67232aa67798419012aa7cc43408f37853f675a1ed9229bd812822

Observation 74502ed6-8987-428d-8011-e947ddedcbc0 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:48.907820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:48.907820Z digest=sha256:bc67941f378d24d0bdaffdb2d501ba37892ff3d9e0b797443fc3c6e64740165d