Pith. sign in

Paper Citation Record · LEDGER

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 5 inbound Pith citation observations for arXiv:2412.04037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04037 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:11.966188Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:29:41.247052Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T04:46:05.518160Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9dfbd778-91a0-4e1a-b756-f41b373cabdb · outbound

This paper cites Talknet: Fully-convolutional non-autoregressive speech syn- thesis model, 2020.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Talknet: Fully-convolutional non-autoregressive speech syn- thesis model, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.662393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.736835Z digest=sha256:89bbeb3ee61f2f98579fe443197a9d53bdc0d24c9b246609edc861845eb464f8

Observation e01a7de3-3a67-43a5-80b4-4f2f332e41cb · outbound

This paper cites EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.741339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.741339Z digest=sha256:21c3bc38188e842aac194855dac929d205bec232b1fd2ee934bebf87432eecac

Observation 8c939fb3-5021-474a-a984-8d0821648288 · outbound

This paper cites an unresolved cited work.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:53:12.651792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.745431Z digest=sha256:aab8437a3fafafdfd7075858a7a4a36ffeb1350433e44168b374f06d801f6cfb

Observation 1d51c5e5-81bf-4137-86c4-16bb6e96f1b8 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Arcface: Additive angular margin loss for deep face recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.641213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.749257Z digest=sha256:b3854e0aa10abb5467f5860acae23b32d16d7b065140eae8948e2b740bd07d3e

Observation 08c07cb8-55ad-46f0-b586-bd5b8030d8f4 · outbound

This paper cites Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.629328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.753360Z digest=sha256:6e20b790c42d232fb868407fb6f1ad469e4b9d0ae867fccfa4adcd4bc1443a97

Observation af9b6e5d-18ec-444a-84a9-a2c8769f56a4 · outbound

This paper cites Megaportraits: One-shot megapixel neural head avatars.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Megaportraits: One-shot megapixel neural head avatars

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.617980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.757387Z digest=sha256:b26a8cc0aaa2dd5dc6deb5e8174b947057d01449094f6a6cee0b68be529f3818

Observation e4c6fd79-5ff4-4c1e-80f9-c40506c89119 · outbound

This paper cites Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.761550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.761550Z digest=sha256:0384810741f31711616251fc57dedcb1dfed1cc91bec4bd4f19bc68ff8b191c0

Observation 3f5d3a74-cdb1-4e34-8cbe-fce70d047100 · outbound

This paper cites Affective faces for goal-driven dyadic communication.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Affective faces for goal-driven dyadic communication

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.605474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.765905Z digest=sha256:54602db61638b868b391fd27834c9d13299ce2ec583fbfede58587ce5e1653e5

Observation d0a58d21-4ac4-4bf8-a592-1fdf781060f1 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.769746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.769746Z digest=sha256:9f9ae74b2de4f03472871407ca4147da3d6abd5802168b5e37f4428a2defde35

Observation 64457f55-f057-4b5d-9b20-366deb89aea7 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.587350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.773705Z digest=sha256:670944acc383286d62780fd1e399d1d6fe8a87d8ccc2245eff14e1fb49b1272c

Observation d46729b8-d32c-4717-9bde-47efb366501e · outbound

This paper cites Denoising dif- fusion probabilistic models.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Denoising dif- fusion probabilistic models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.777648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.777648Z digest=sha256:5d6b13492ac6a9b692bc3dbf9a1ac9ad816f1a5e160df5d88884becaebaaf654

Observation 3ea19e3d-1b9a-444c-84f8-3b943d9c3b02 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.570776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.781659Z digest=sha256:16dbb6d5b5c3a38f972558721500d009aaa5c958ce8baa26775b9b2860b1aa02

Observation 4cd2147a-360d-47ac-a1ff-c919f54cc49c · outbound

This paper cites Percep- tual conversational head generation with regularized driver and enhanced renderer.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Percep- tual conversational head generation with regularized driver and enhanced renderer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.560307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.786102Z digest=sha256:7507f4bf7d3820d37b823cc8d7466237becdfa2e7ba83ce134f91fa59988be6f

Observation f876d396-533d-4fa7-a2af-2d1b9e52ea9b · outbound

This paper cites Interact: Capture and modelling of realistic, ex- pressive and interactive activities between two persons in daily scenarios.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Interact: Capture and modelling of realistic, ex- pressive and interactive activities between two persons in daily scenarios

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.549397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.790062Z digest=sha256:38f49cefeda753ee5352db03ec6e19e06fe776aa1be52e36acf3a54624455d68

Observation 4c5b979d-c718-4557-a079-32305fd51330 · outbound

This paper cites Analyzing and improving the image quality of StyleGAN.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Analyzing and improving the image quality of StyleGAN

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.537303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.793901Z digest=sha256:f7846f574fb8d3e15b2ec6d3f9ec112fc947f3d8a7b73657211cc639a801309b

Observation 80ee6fb2-eae4-4aa6-93ef-5b5a6110388a · outbound

This paper cites Mfr-net: Multi-faceted responsive listening head generation via denoising diffusion model.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Mfr-net: Multi-faceted responsive listening head generation via denoising diffusion model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.524818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.798431Z digest=sha256:d865975e9590c0391d119b7372ea1f21b174d4605d755015b17bf7c77baa4b2e

Observation 0a198528-c71e-4090-b58a-2a5be7fb51e0 · outbound

This paper cites Anitalker: Animate vivid and di- verse talking faces through identity-decoupled facial motion encoding, 2024.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Anitalker: Animate vivid and di- verse talking faces through identity-decoupled facial motion encoding, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.512443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.802718Z digest=sha256:c9f95c7f1411591749b06602e2c794fb1bdc2c03097f8f7f0ae774341aecbf16

Observation 1bcd68d9-ed2a-4e36-b3a7-d0d0cf83888f · outbound

This paper cites Customlistener: Text-guided responsive inter- action for user-friendly listening head generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Customlistener: Text-guided responsive inter- action for user-friendly listening head generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.500099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.806502Z digest=sha256:766f0d9e22717da5a3cec31539486be92c5a023e00c40282bd0dd7f6ebaccc6e

Observation 92457ba9-b2f2-426b-b2e2-a57c29c27bdf · outbound

This paper cites Decoupled weight decay regularization, 2019.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Decoupled weight decay regularization, 2019

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.810244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.810244Z digest=sha256:255bb720b0eb07aac40c3feaac1e57b30ac5f111c996a340fc3fa723aa8bdc8f

Observation 4d2cc443-95cc-4ba6-b0e8-8f657a0ff0fa · outbound

This paper cites Mediapipe: A framework for perceiving and processing reality.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Mediapipe: A framework for perceiving and processing reality

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.478745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.814002Z digest=sha256:72180364f84db0d67c34d94a66d09ad4070c73f3c4761dc0ed5b1c4cbea4f58a

Observation 17a75fd3-f8c2-4b18-9369-64d289ad0d98 · outbound

This paper cites Styletalk: one-shot talking head generation with controllable speaking styles.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Styletalk: one-shot talking head generation with controllable speaking styles

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.466465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.817967Z digest=sha256:79cfa45c3629d616605f18db4f43206e6963880cb84819e8ea6461338519fe18

Observation 90b8a417-58e2-4082-a8b7-fb75f5d16a88 · outbound

This paper cites Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.821517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.821517Z digest=sha256:ac608c42221134f63f19f855e8f30acbed036dd0cfdcedeecc9cbd846a177dd4

Observation c3b53312-fc9d-4762-b338-396a61ab3e84 · outbound

This paper cites Learning to listen: Modeling non-deterministic dyadic facial motion.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Learning to listen: Modeling non-deterministic dyadic facial motion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.455133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.826076Z digest=sha256:1f63c9b55c0571e3597b62a662299ba811d5f8dba0fbedca37a8de03395c159b

Observation 34ff912f-bca7-4d07-8d0f-a8975091555f · outbound

This paper cites Can language 9 models learn to listen? In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 10083– 10093, 2023.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Can language 9 models learn to listen? In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 10083– 10093, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.445047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.829510Z digest=sha256:4914119457ad385c03b496bb023766180310400e43f3b79489f3ca30818b1a45

Observation 1bb1ca10-fd45-4e29-b21e-510fe3a8ceb7 · outbound

This paper cites From audio to photoreal embodiment: Synthesizing humans in conversations.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations From audio to photoreal embodiment: Synthesizing humans in conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.832950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.832950Z digest=sha256:c260a82653bb1daf298f93d6b8f2173c711d7bb67c9ff6a9fae72287b1c33c4e

Observation 828926df-fee3-4fa5-aa54-cc8362e75f70 · outbound

This paper cites Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.836010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.836010Z digest=sha256:fa5c7b7160bbb557625bbbdc4a490956e6e8c3076e87a71a38530d474877f12f

Observation 26962c05-95d5-4fc1-ab7c-e7f69bf8a61a · outbound

This paper cites Nambood- iri, and C.V.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Nambood- iri, and C.V

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.429424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.839810Z digest=sha256:2b94e0b3052c8a2c4661c88b045d920a5103d8f112284bb69e20baad7a3e9a59

Observation 4167e2e0-7650-4ac9-ae74-1d60ebefa28e · outbound

This paper cites Li, and Shan Liu.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Li, and Shan Liu

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.419924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.843204Z digest=sha256:095e3bff8a143f6b5fff6b8b965218297612628de31141724da1f841346aa358

Observation 00fd7ca6-e607-405b-8a26-32954fd78f03 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations High-resolution image syn- thesis with latent diffusion models, 2021

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.407735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.846509Z digest=sha256:f53fc35d12f3c548824ebc680fbb8da3116002678ac01b58a92233e83ed9abde

Observation c1e83505-e230-48f0-9ace-d61588619bd3 · outbound

This paper cites Denoising Diffusion Implicit Models.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Denoising Diffusion Implicit Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.849442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.849442Z digest=sha256:d9164af32c669288c38750448c179329d0b6c7495350252d47381be88665b347

Observation 5ba9a542-de87-4e9b-8dc8-2bb4b7183f42 · outbound

This paper cites Emotional listener portrait: Realistic listener motion simulation in conversation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Emotional listener portrait: Realistic listener motion simulation in conversation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.394845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.853496Z digest=sha256:64d8c5b121b0b4c0ade326ed85d14f0cf05f124897bae9e5c43724e9572c86d0

Observation f49d2435-db0a-451f-8d7c-0425748ea3b4 · outbound

This paper cites React 2024: the second multiple appropriate facial reaction generation challenge, 2024.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations React 2024: the second multiple appropriate facial reaction generation challenge, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.383278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.857028Z digest=sha256:c89b0987ba97df235aaee718976fa00e073f3fcd43cb92c4a06c9dc45d1acf35

Observation 44fce604-1b28-4b16-a51f-e5b43513cf80 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.370436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.860695Z digest=sha256:d73710d6b44ed9e07acbc26625a260b7f7c333997885b7e28fc4dc34d8eb86f2

Observation bc4d08c7-6a15-489d-845c-6610fe9d8394 · outbound

This paper cites Beyond Talking -- Generating Holistic 3D Human Dyadic Motion for Communication.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Beyond Talking -- Generating Holistic 3D Human Dyadic Motion for Communication

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:53:12.138969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.864164Z digest=sha256:8f7304e708017afead8f2be98aaddb181ad72660ed0bd7f8f2e8175e54c06d6e

Observation c36eb33e-eb22-470d-aa2f-1244e748e648 · outbound

This paper cites Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.358339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.868282Z digest=sha256:be1aa5a6835a4f85b337214efeba1f8dc25618c665ab4fcee2a28cba47157594

Observation b7ae3069-786c-4293-940e-d229b14366d4 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.871802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.871802Z digest=sha256:5cadc96fe023b818ac9664847c505f7aed897dce5e3705dba07ab44977b3b915

Observation 5422135d-daf7-4ae8-9c3f-f71349a2abcb · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.875876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.875876Z digest=sha256:1e5cce213c10535a16092573ec75ab313b9a23c9d32f079230d71b154a17507b

Observation 44299887-66e1-4064-9b30-a94d1897ee86 · outbound

This paper cites The information bottleneck method.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations The information bottleneck method

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.879569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.879569Z digest=sha256:74ab4691e13738ca5c5f9511caa5be94cd10ba3f87734444d21164dcfaae3d13

Observation 15856dd0-0eb4-4200-8c0a-0fe06384bf69 · outbound

This paper cites Dyadic Interaction Modeling for Social Behavior Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Dyadic Interaction Modeling for Social Behavior Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.883377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.883377Z digest=sha256:ed0e654300647835266e4838132a53a082b995c4aba4cc49beb8e34538f4bb3a

Observation f55b42de-9ac8-47a2-a1f2-d080df0dae48 · outbound

This paper cites Neural discrete representation learning,.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Neural discrete representation learning,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.887127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.887127Z digest=sha256:33e6db01d4f9803ecaa0dd1f93f663c200ec8c3352b5f4a657ced617eed67d9e

Observation 38bd2266-8725-4b43-b102-90c96203977f · outbound

This paper cites V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.890901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.890901Z digest=sha256:408fdcaacaab39d290cf94a42d2081ec6aeddd416eb1a344327f4f14d9339d0e

Observation 5b61c8f8-b7cb-4adb-beca-738db3cbe68f · outbound

This paper cites Dis- entangling planning, driving and rendering for photorealistic avatar agents.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Dis- entangling planning, driving and rendering for photorealistic avatar agents

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.341725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.895026Z digest=sha256:cac28a1f672c5fb3fd19c9bd28577a30033bf1deb2c451bd6957529d3596d479

Observation fb60fb04-14c2-42f7-a499-1b56bb6f36da · outbound

This paper cites Progressive disentangled representation learning for fine-grained controllable talking head synthesis.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Progressive disentangled representation learning for fine-grained controllable talking head synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.331269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.898857Z digest=sha256:782a8b3cd4837b70189092fe7bac31fe0732d48f3a11b2057d14afd092d206aa

Observation c4bf814d-f4fb-4ae2-a766-52e77c0130c6 · outbound

This paper cites Latent Image Animator: Learning to Animate Images via Latent Space Navigation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Latent Image Animator: Learning to Animate Images via Latent Space Navigation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.902478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.902478Z digest=sha256:ad6d236ff02b3f80880c850446094e8868d1172b6759b0f0769ea43c3a7ebbdf

Observation 3dc37f7c-6fb2-4cbf-b667-81018ea390d6 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.906895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.906895Z digest=sha256:c40d6ab488e53cc4a2eaa17eee71256d3346ec8d3ef54aae611f242f172e96dc

Observation 190e1e3a-50df-4b69-a36a-75126bd6c944 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.911780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.911780Z digest=sha256:cbe3890f54a1b781cdc608a5af5706a75228460579d59fc9ef3680caed9861fc

Observation 2f065508-d3ba-46c9-9ade-ff4086ac8fe5 · outbound

This paper cites Dialoguenerf: Towards realistic avatar face- to-face conversation video generation, 2023.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Dialoguenerf: Towards realistic avatar face- to-face conversation video generation, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.320642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.916197Z digest=sha256:bb5c9fa4fab23f58bcd01014ab4b051f89c9e01eab1745d4443c4b89adf0e97e

Observation d274d88d-da5d-4bc9-a55b-d7a30f776d41 · outbound

This paper cites DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.920401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.920401Z digest=sha256:073ab974779336f5b2f8018638c4130392e9fbefed13ad0e6b52159462981094

Observation d59b9a30-b0f5-4193-9e94-cdbed22f3d6d · outbound

This paper cites PersonaTalk: Bring Attention to Your Persona in Visual Dubbing.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations PersonaTalk: Bring Attention to Your Persona in Visual Dubbing

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.924761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.924761Z digest=sha256:29afd710b90fee3fb2729005f04300bd3a2ec67d96fbb58d4b9ed4a825e7913e

Observation 337fa33e-41e9-4670-bb57-838f5fff805f · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations The unreasonable effectiveness of deep features as a perceptual metric

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.928742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.928742Z digest=sha256:2a802859d915b0c4106ac94b9e0a7b61bb7a013d2ca602b4be0a6232a41b2122

Observation d69ba8ac-5856-4051-abbf-0a4f49d87ea5 · outbound

This paper cites an unresolved cited work.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:53:12.302455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.932444Z digest=sha256:bb54f0eb68d3992768e054b937186c423f033d22f5b1a5e38f37e0c4adfa90a5

Observation ec525900-31ea-4a90-a3fc-7c55d91ef699 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.291723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.935710Z digest=sha256:5b470e01423b9f1180040bcdefcc7ab60ca51f66b5329caed3656374ee5bce79

Observation 9d72ee49-1f4e-44a9-88b4-b06b855aac9a · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.280781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.939319Z digest=sha256:4b972f3e2b22d223fca37e2efcf5da23815e68ea3e7e86c328b7cbb20a2dee73

Observation cba50b78-9cf3-4440-8865-b82f94135e34 · outbound

This paper cites Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation, 2023.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.270963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.942072Z digest=sha256:bcb4ce5d32a24b64c3879cb81c764943577d24b84fd2922a8ceb59da19816d00

Observation e01301c8-85e0-43fd-95db-9deb7568ab80 · outbound

This paper cites Semantic-aware responsive listener head syn- thesis.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Semantic-aware responsive listener head syn- thesis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.258543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.944923Z digest=sha256:18df0427942da2889a061da621dcad2b34cc511542142e0c480664fa25a21c33

Observation ea97939c-5eeb-4b4e-96d6-562d765352ff · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.246872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.948116Z digest=sha256:5156b90942b5a4545065d5ad77ad5f87c7b908f49b0e4597bf9f69a1811d098f

Observation 7a6df66d-d046-4ea9-97dd-b41d5db23498 · outbound

This paper cites Vico-x: Multimodal conversation dataset.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Vico-x: Multimodal conversation dataset

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.235650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.951252Z digest=sha256:6c2edecc794bb16c457bfb40a9cfe85b2d3e468cae216ce96eaa4e8f6e8d1232

Observation 38ddad94-5703-4d7c-a4fe-5dbafca72a3e · outbound

This paper cites Responsive listening head generation: A benchmark dataset and baseline.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Responsive listening head generation: A benchmark dataset and baseline

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.213548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.958153Z digest=sha256:91e87e59732c586df2ba953d5719630b86e385ce305a71779685c93a7fc9851a

Observation aea57f22-9287-4d71-8a75-dc7fa33b229c · outbound

This paper cites Interactive Conversational Head Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Interactive Conversational Head Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.961583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.961583Z digest=sha256:705533cc6773be9565e2ad354332aaefc567a1145e1b7f3bdf9afc43cc786c2d

Observation 21ee5012-1755-4aec-a448-072998ad928d · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Makelttalk: speaker-aware talking-head animation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.200475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.966188Z digest=sha256:5148433e600d9ac2381e706ae0b7f2f69aeac3487da18c22b2dc9bb7e6043932

Observation b748deca-3023-4fc7-a8a8-84f33a166cc1 · outbound

This paper cites an unresolved cited work.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:53:12.224929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:53:11.954498Z digest=sha256:0cccda50c85c0a481426f9268c897a991491abc895c685c807dec66df3175018

Pith citing papers

Observation bc068e31-25cc-4c12-bd41-c1c26721d8de · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.943246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.943246Z digest=sha256:90fb7af0fb33869eb270a1129fa98728a8ae00b30a5e498d95213b09054f7207

Observation 65f06d46-f32a-4f29-bb42-726d23356af8 · inbound

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router cites this paper.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.247052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.247052Z digest=sha256:824f1b7c8576d876f764973fbcd82c5e591e16c434a8b89068fbe6eb02603dc5

Observation 3239a036-a110-470e-84cc-4cd7c2e1fd95 · inbound

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations cites this paper.

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:17.092084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:17.092084Z digest=sha256:9e99ca8c3346cbbc6df1873745450ef4388e34ff25458474eee09eff099b39bd

Observation 466d4ee9-2b49-4c9d-871e-bd7dd98597c9 · inbound

Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System cites this paper.

Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:58:01.712367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:58:01.712367Z digest=sha256:2f1ad8b97852076e0c8faf91fc3b0e1df22b0aaba97219d2c891bc9b7fca92cd

Observation 8569be9f-6540-496b-9667-ed88b5a6cfde · inbound

Multi-human Interactive Talking Dataset cites this paper.

Multi-human Interactive Talking Dataset INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:46:05.601149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T04:46:05.382405Z digest=sha256:bae530d3ebd107c641c4dd40838f997f3e31af480ec06467e417f93d15ce5629