Pith. sign in

Paper Citation Record · LEDGER

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation

As of 31 July 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2411.09209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.09209 v5

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T16:59:57.727035Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:36:53.138354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T21:46:15.570278Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact12
  • verified fuzzy33
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43cf5105-36a2-49f2-809a-1a0fde3c3c3a · outbound

This paper cites GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.603956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:43496f0c92b0ed9de2095245d47c0c0548943c0ea3402c4d32b21aee29476809

Observation af259182-f475-49fb-915f-2c93f8bd620b · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.594476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:99222394ad58ec44b928950e24fca22afceda812c4912b4ce06e812402c8ab0c

Observation 01855afa-757d-419b-9cb5-40e207cc27d3 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.558326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:2b6d0b00e01d3485eab8b9bb361caa8144582172b1e8c172418662c60974668e

Observation 81bf1d80-3a4d-45d6-a9a3-59376834d9ae · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.612108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:ab2b246daed78efb57bce572fe468f2b877e1ae6642b2d2f4607db2a02ab116b

Observation 477d88a1-bf97-4697-8fbd-6fd9f95d32fd · outbound

This paper cites VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.574515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:f9d3c5de27e30f6190a98d2b316dd4ab6b0ce076b6897be3a6133c27bcf1ce6f

Observation f8279f83-7197-421d-8218-d6473a9fc138 · outbound

This paper cites Emo: Emote portrait alive - generating expressive portrait videos with audio2video diffusion model under weak conditions.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Emo: Emote portrait alive - generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.909596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:20d6de1e24c56272e928eef3b9d03cee7ffc7b4d7d3ee6eb8d0f0e1930ba9415

Observation 62cf65c2-c35f-40bb-9443-75208361bb64 · outbound

This paper cites Emotalker: Emotionally editable talking face generation via diffusion model.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Emotalker: Emotionally editable talking face generation via diffusion model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.905886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:ebe509c9f38f67e5d7facffc5ccf9a96fd51ca10a4f4c49492845829a278daad

Observation 8a2da2d1-d68f-4003-bc52-d5b271ef16f5 · outbound

This paper cites Digital avatars: Promoting independent living for older adults.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Digital avatars: Promoting independent living for older adults

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.899634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:1ea144608cd22194ccc62f0c4aaeb1b491fbf6c0963acd53310fe6640310d504

Observation 31b53fbe-6680-4f4c-992f-8c1f8100db27 · outbound

This paper cites Talking face generation with multilingual tts.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Talking face generation with multilingual tts

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.895958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:f57fa5e391e50bb31013dac07c95f39deddb585f79a1e29d7199ccfedfec7784

Observation a73e94c4-44a2-441e-b089-acddd406dab4 · outbound

This paper cites Improving user experience of virtual health assistants: scoping review.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Improving user experience of virtual health assistants: scoping review

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.892170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:ce4b9ef14def58e083668423bc0e02cb31a40bc1d5dadd663b503212b60b45af

Observation 77549e31-efcf-478c-baae-7171dcfea7db · outbound

This paper cites ChatAnything: Facetime Chat with LLM-Enhanced Personas.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation ChatAnything: Facetime Chat with LLM-Enhanced Personas

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.590051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:7025af0dd8bd5247b9228e26dbb28c7e3f9f26bdab0a007f86750d430a7597c1

Observation 9e0a6530-afb2-4dbb-bad0-9a656d3be864 · outbound

This paper cites Building llm-based ai agents in social virtual reality.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Building llm-based ai agents in social virtual reality

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.888288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:57b5f4db59f98c6f8925248400cb54db081c7cb6d52da5b4dd808fe2e0d8d3cc

Observation fd48d80c-16f9-4d81-ac23-1e70175746e2 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based generator.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Stylesync: High-fidelity generalized and personalized lip sync in style-based generator

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.884876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:79ffd0f9518c071a767b807f3f67d0f08d5eb26d6572fc1c4771a22cc19e881a

Observation f5270fe4-727a-4699-9929-6bb284221350 · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.881553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:025165f42bfa15cfc266939655038895d276adc9d0adb619b6820a28a0da0c3a

Observation 30f4f98c-fe0d-49f9-b4eb-6835ef806755 · outbound

This paper cites Hallo: Hierarchical audio-driven visual synthesis for portrait image animation.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Hallo: Hierarchical audio-driven visual synthesis for portrait image animation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.878139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:906e596720416932c3fe26d4639ea04cb55a1aece0208884b8c12235b21f5c52

Observation 2aa6725c-b2e9-4778-b13f-4ad69909356d · outbound

This paper cites JoyHallo: Digital human model for Mandarin.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation JoyHallo: Digital human model for Mandarin

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.568862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:7386b97ccf10a3df2c0a24f236b0c7e662481dcb70613bc045118392a2a65a49

Observation cfae7aa5-7904-4cff-bc95-e877c8cf2f0e · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.585478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:3bcab223e1ceb9e79f721f94a1e70d0da32a0f7bf4e4deedd8006e79274b3782

Observation 12aaf60a-78cf-4c43-a023-e3d238f0f981 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation A lip sync expert is all you need for speech to lip generation in the wild

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.874641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:80ca7defe2b66e46b1df09d44a2cb19f2d277ebc7dcd403fda385808dea61df7

Observation 7d09c244-92c5-4f6c-bd39-34f9fb9228a2 · outbound

This paper cites Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.871117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:50886669d63159e98fa234a5e8cd11093683ed35f9df4d941b36693613ce2d14

Observation e1b02d9b-12f9-41d0-995a-fbf9e51a4e9a · outbound

This paper cites Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.867400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:a0701104ad65dcafe5fdbabc3dd5fbf65a92b8552ea3e913bd5834aa3560ee0a

Observation 2c4941b5-1469-48a7-b556-811492413afb · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.863874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:3e1607cce89e13cb5a43a017657e7b63d0fb7586d162799b233463a59bacee86

Observation 59433fa3-7ab4-4237-965d-87ed000e5eda · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.563841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:ed177d9bfd76a29e453ce1138af0bbd2041dc5c2f1ca257933300be7788204b4

Observation 8d8f8fd9-c39e-4a0a-8df3-f130a291b359 · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.608008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:b5018f4258ed18db2e79d934ddefb97e75549396c4bb470962a73a02afe47c39

Observation 0da630c4-6b96-4aa9-b628-0cf05ca64b60 · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.860487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:d72ec223c0faa7de1035f9ce9aac6525b209c7eb113a24c33660ada5df862e13

Observation 9ceb3f55-3ffc-48d0-acb2-f738f8a39389 · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.857296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:73574423d95929b026d7b80ad025185450308f4e08de2b321fe36e21fbc32e0f

Observation 62d81ef9-87ab-4874-bf70-2c088bf901d5 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Neural voice puppetry: Audio-driven facial reenactment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.854028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:078e80ee0a2e45e701a03f81fb55f9b785f12726539a82e9f7b1889ee06d0de2

Observation 69ae3649-7449-4d81-a742-2c54dbfbdcb1 · outbound

This paper cites Audio-driven talking face video generation with learning-based personalized head pose.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Audio-driven talking face video generation with learning-based personalized head pose

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.850618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:9d798d4ed71016ad0294d9f2a21daf99f51d0978fdc2b13271bca828d29bab83

Observation 299dbd96-e1f9-4b0b-88db-deb5615c11e6 · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferencing.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation One-shot free-view neural talking-head synthesis for video conferencing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.847600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:c045f2e5c53a765122b8306ff492919b2e55117b0d4ddd9b208cfe3eab04ffb2

Observation 66e2a98f-19a4-46dc-b7f3-67fa40c502f4 · outbound

This paper cites Pirenderer: Controllable portrait image generation via semantic neural rendering.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Pirenderer: Controllable portrait image generation via semantic neural rendering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.844484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:cda858b917ea0305afffa8d686ab1b249ba9fc2dd76a5faeb61160fd45f68153

Observation 8c25cec8-be35-4e04-b8e7-7ca8fb02e773 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.581463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:9fcd337cbc10223dfce045ff27a2de0f4fbb26935b4e73d88059a66ad4dc9ba4

Observation 1c1fad3e-f6ec-48fc-8c33-d3e92b41a702 · outbound

This paper cites Megaportraits: One-shot megapixel neural head avatars.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Megaportraits: One-shot megapixel neural head avatars

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.840982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:b3d9c2c32ec46a9004fc666d091e2de7c18e94ca665d1f6811de0a35589a125d

Observation 30927a77-d69d-49af-8725-db7155d20958 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation A morphable model for the synthesis of 3d faces

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.837720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:1cbb71cbef3ccdd696ff16b968d8b53e1901497f1eb79844cdbd3e674b957f76

Observation 0ce78366-adbd-4f5c-a5a0-c9de4532e56c · outbound

This paper cites Expression invariant 3d face recognition with a morphable model.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Expression invariant 3d face recognition with a morphable model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.834119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:8b07d15085633c2f6717982b11f8777cb486b4ae8a8a75df77a2bdeff8fc58c5

Observation 385b1e27-a574-4399-b5d4-5c70fe15ebe5 · outbound

This paper cites Learning a model of facial shape and expression from 4d scans.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Learning a model of facial shape and expression from 4d scans

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.830479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:6205bb6c3c2c4ee9dd1afbbd11c32bd413b520065b688c52aea476f24491e4dc

Observation 7bc07450-5619-4044-b2a2-6b978f60549a · outbound

This paper cites Disentangled representation learning for 3d face shape.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Disentangled representation learning for 3d face shape

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.826849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:a3697cd8b2cdb0eefae1029187a85b589085ab82286b51860440100a66502cbf

Observation 2ba51db7-f8db-4ecf-a5d5-bada6296affd · outbound

This paper cites Emoportraits: Emotion-enhanced multimodal one-shot head avatars.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Emoportraits: Emotion-enhanced multimodal one-shot head avatars

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.823622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:f2cfb31f1483698e9fa645f0cd3dd8e8e2d6b31b7374a5020b787548be105a43

Observation c2fd8d3c-062e-4d0a-82f3-54287005e5f7 · outbound

This paper cites First order motion model for image animation.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation First order motion model for image animation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.820600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:640b179ad652395391f4b45850134a42b7622656b2bdea6754d58b70d2728bce

Observation 0558fa06-ef45-4d92-bd5f-a18654b09ab0 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.817648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:84ecd62a61bde300753b4fbf947940195d752c2701487da562959a80e7c0d798

Observation 37cbf5a4-ace7-4ad4-8c16-64a4645074b5 · outbound

This paper cites Capture, learning, and synthesis of 3d speaking styles.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Capture, learning, and synthesis of 3d speaking styles

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.813654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:e0ccee190838b2138c658f3be24f162617a3df832df251be74dcb11b24c0d75b

Observation 45762e3c-f687-47d4-abdd-244a28f27ca4 · outbound

This paper cites Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.809792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:021b9945bb659a8c7c3f5e52046ad8a7877253f0a29001115226225db2fe5867

Observation 90c00899-965b-4375-bcca-2ab093d52a15 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.806013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:9080185473d2fb8c49f9e3273f6d5d8d4df40806fd805c15fe992c44d2caf094

Observation 7856fde6-9a8d-437b-8894-b349f5c50022 · outbound

This paper cites Celebv-hq: A large-scale video facial attributes dataset.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Celebv-hq: A large-scale video facial attributes dataset

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.801894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:d763ccd6673ab58b99330f139d0643558b6bdb05999c73ff5aa4304a4d823ff0

Observation 3bc999ad-3cc3-4771-8b01-0342b21c69d1 · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:03:12.598662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:0d948a1fc2da44b3b469890cba492f9a70d410100a225cc77f254392f8a69cfb

Observation 06b2aab1-77c6-4825-aa34-141c87ba6343 · outbound

This paper cites Out of time: Automated lip sync in the wild.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation Out of time: Automated lip sync in the wild

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.798558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:881d849e5ddd8f965c1d16a9a411f4ba8d8b93d585177f19e299224a385dc21a

Observation 52fae500-32fa-40aa-a408-26fd8d746434 · outbound

This paper cites FVD: A new metric for video generation.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation FVD: A new metric for video generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T17:03:13.795055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:748d59189274e2eb26a9bddf784687807c381aa98da2311b890a4d85e5a9ffc3

Pith citing papers

Observation 2a1d3b0f-91b5-4891-9e23-c231d5e89308 · inbound

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation cites this paper.

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.684711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-06-29T22:36:53.138354Z digest=sha256:59a86e4833117789fad6874c32af026003002f10dbda65cf885507e38814e377

Observation bad1223a-4c00-45a6-b6e1-3c8a4d901742 · inbound

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation cites this paper.

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:46:15.571777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-06-28T16:19:55.048831Z digest=sha256:3f66a1516e588fecf7f0071cd7a81c4bedb86ccd2fa34fbbc5410da921cb5088