Pith. sign in

Paper Citation Record · LEDGER

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction

As of 19 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2504.21718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21718 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:00:11.354369Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11052c50-5b9c-4368-bb20-e7ddcd658c03 · outbound

This paper cites Facetalk: Audio-driven motion diffusion for neural parametric head models.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Facetalk: Audio-driven motion diffusion for neural parametric head models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:14.196598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.647115Z digest=sha256:af4e177800811adcdaa67528727428cf7393124a11fd822b5aa0bac6f1072b37

Observation 69e10463-1475-40a0-ad0b-77dd261f720c · outbound

This paper cites On the challenges and opportu- nities of physically situated dialog.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction On the challenges and opportu- nities of physically situated dialog

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:14.166038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.664607Z digest=sha256:b6a6794fea4e97edcaac77d8cbd655c8adc295b557ca15f8afb44df7cb5a1768

Observation c2be071f-9a08-423a-9386-390d1e2e0156 · outbound

This paper cites Facilitating multiparty dia- log with gaze, gesture, and speech.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Facilitating multiparty dia- log with gaze, gesture, and speech

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:14.136820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.683881Z digest=sha256:8c4f0fcb0193c6ef90db3ec5ce1487228cdae2c82dac77b74440ecf0fc9c45a8

Observation 86bafd4e-8131-45ba-8035-8ad5f9291e39 · outbound

This paper cites Libreface: An open-source toolkit for deep facial expression analysis.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Libreface: An open-source toolkit for deep facial expression analysis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:14.086776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.710214Z digest=sha256:c92213b903aa8e028ba74b19fb51cf65cb307f3ef3bc67acae6564ee52f19a2e

Observation 57a75beb-773c-4579-9035-08f171b1aee6 · outbound

This paper cites Us- ability and responsiveness of artificial intelligence chatbot on online customer experience in e-retailing.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Us- ability and responsiveness of artificial intelligence chatbot on online customer experience in e-retailing

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:14.023429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.733865Z digest=sha256:f51ea710a8505a91e93bbd043837bc62ff25c82b014745dedd0757f0e73c0b59

Observation 54fc0aa6-d150-42d5-91bb-3ad2321cbdd8 · outbound

This paper cites an unresolved cited work.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:00:13.986684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.745386Z digest=sha256:ff800a0a4c9621b7847ce9b96c113dd1e690cff84867574f94a04a58b87f21b1

Observation ead78a83-1883-46d7-9677-c313c7e0eb26 · outbound

This paper cites Out of time: auto- mated lip sync in the wild.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Out of time: auto- mated lip sync in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.936563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.761303Z digest=sha256:b24beec69fd4703ec6e0d0ca164397c18676e114c7c2aafa5d34f0ef5d808bad

Observation 2d0f400a-ff1e-4e5d-bdf3-d8635e6dd794 · outbound

This paper cites Capture, learning, and synthe- sis of 3d speaking styles.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Capture, learning, and synthe- sis of 3d speaking styles

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.904124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.773734Z digest=sha256:d9e6cb6ce410e612f689a0403377d542875ee928c1a8927e7de434246918fc6e

Observation a7c7ff17-a182-449c-b765-2322aea258bb · outbound

This paper cites Emoca: Emotion driven monocular face capture and animation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Emoca: Emotion driven monocular face capture and animation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.878069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.792908Z digest=sha256:d13b7f5307b7dc1f881d7e20d274d2462e7c10d3c7ee0e4b1a4eef098ab6c589

Observation b87afe44-ba04-4c41-bf33-09365436be6e · outbound

This paper cites Emotional speech-driven animation with content-emotion disentangle- ment.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Emotional speech-driven animation with content-emotion disentangle- ment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.850023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.802904Z digest=sha256:114f077bf54fd17c47916a6ad090b6296c67b89099003a905458d1646cad0d8b

Observation 26caec1e-9e9c-4276-8c62-622d8c0d74ea · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction A survey of embodied ai: From simulators to research tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:10.819204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:10.819204Z digest=sha256:a03d95e839515a30cb4deb038ad7fc57206feb9c5813b1cf22de6be9ff61f742

Observation bf9e8299-0a87-44d7-b202-a13943f2e38a · outbound

This paper cites Facial action coding system.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Facial action coding system

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:10.831297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:10.831297Z digest=sha256:3d3f1d8155d4d0d73c275d503138ecd64a0fa765dd0a0258f708fd0df07a4b78

Observation 52528e96-88be-4ad7-844f-0747ac8ef416 · outbound

This paper cites Studying human robot interaction and its characteristics.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Studying human robot interaction and its characteristics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.726650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.845930Z digest=sha256:edead0ea75b0f0167becb5f2d01c77b731539b35871b5cde22a058db98d8db33

Observation a9ea1f5f-b612-4ccd-bb78-bade98106322 · outbound

This paper cites for real.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction for real

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.702772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.860070Z digest=sha256:66c5f3999030953386e79762bd3f7482d5fe8ec83b8064b9bca3369a201dc0ec

Observation f0b2af32-f95b-4aed-ae26-52916c115213 · outbound

This paper cites Affective Faces for Goal-Driven Dyadic Communication.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Affective Faces for Goal-Driven Dyadic Communication

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:10.877799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:10.877799Z digest=sha256:4082173f07aaac6396a41e7e8e27deca7d8b83aa2a404abc078a22553aea6592

Observation 77269df6-7368-4592-971f-d714923738aa · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.647531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.887789Z digest=sha256:e0703f12535ef86a29268e68a857c0e5b5f93574b337daffbd1b0906f4308186

Observation 5dc20c7c-26c7-49fd-b525-61c7c9b53e20 · outbound

This paper cites Denoising dif- fusion probabilistic models.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Denoising dif- fusion probabilistic models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:10.909601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:10.909601Z digest=sha256:bf82be420af44e2d299db54d46673fddcbaf66b777562bbaa8635fd87b2e302a

Observation 91e157f1-14bc-4d86-82cd-8ec793210913 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Arbitrary style transfer in real-time with adaptive instance normalization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.576870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.917650Z digest=sha256:5c89a84b56547167108c9a5db8061cfba6184c8abfe80da2f61e09310098d785

Observation 2a50c6e1-45f4-4a6b-89a6-30b08b94cda9 · outbound

This paper cites Dyadgan: Generating fa- cial expressions in dyadic interactions.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Dyadgan: Generating fa- cial expressions in dyadic interactions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.534549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.935579Z digest=sha256:41b23d2511ec3a77ae432d6f876c06b810954966299d01405685f79144aac2ff

Observation 2ab53fd7-b2d7-4a7f-b795-1bd658e8f52e · outbound

This paper cites Gradient-based learning applied to document recog- nition.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Gradient-based learning applied to document recog- nition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:10.956819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:10.956819Z digest=sha256:67bdc0efa5600151536d804ea9c75638a757ab1ebaeb33d1cd7544fee01dd2c8

Observation 6008833b-199a-47e2-b794-d8e1a8325f3a · outbound

This paper cites Learning a model of facial shape and expression from 4d scans.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Learning a model of facial shape and expression from 4d scans

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.473076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.972613Z digest=sha256:b389a2b96b167c4c79060ebf486627837af47de8dd64e57536d0488a5c744cb7

Observation d4f4a457-e7c1-4798-9576-3f183f0bffe4 · outbound

This paper cites Visual instruction tuning.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Visual instruction tuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.392886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.983099Z digest=sha256:539c9a4a46f30ac5c3ea63ffec12cd5bae6e874f3a9913e56500190f8490216f

Observation fa5f69c3-c821-4684-b61d-749a683c9cc0 · outbound

This paper cites Mfr-net: Multi-faceted responsive listening head generation via denoising diffusion model.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Mfr-net: Multi-faceted responsive listening head generation via denoising diffusion model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:10.989657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:10.989657Z digest=sha256:a3e911013d0554c9e3f8dd4b5309327160ac95b0d401794ac0a9741acdccb058

Observation dc73003a-6921-4756-b0e2-c7c3ab2edf7b · outbound

This paper cites Customlistener: Text-guided responsive inter- action for user-friendly listening head generation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Customlistener: Text-guided responsive inter- action for user-friendly listening head generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.345395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:10.997165Z digest=sha256:2feace241e7ef89f573c12754c1e151e3146649f1a55eb12858ddb672abb51ee

Observation 133d583f-9512-415e-be7b-307973c3d222 · outbound

This paper cites Decoupled Weight Decay Regularization.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Decoupled Weight Decay Regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.006310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.006310Z digest=sha256:f39393da491b454b8fe0f27bef2500a09ccb049a306ef93ca77c9a28c897e8b5

Observation e7d314fb-6f26-4d0d-9ccb-32397977bbe3 · outbound

This paper cites Automated measurement of facial ex- pression in infant–mother interaction: A pilot study.Infancy, 14(3):285–305, 2009.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Automated measurement of facial ex- pression in infant–mother interaction: A pilot study.Infancy, 14(3):285–305, 2009

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.293449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.015129Z digest=sha256:4c071e42e114e2d3cd89ffa54d06e325fa6d2094b0f30990df86e662685603d4

Observation e220ded1-7e18-4229-8ef8-44f873769488 · outbound

This paper cites Learning to listen: Modeling non-deterministic dyadic facial motion.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Learning to listen: Modeling non-deterministic dyadic facial motion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.253271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.024041Z digest=sha256:da3ce5a336dbf29a25b1ee3d46814f01d2540dc406e9b26b85a95384ea4f5611

Observation 94cf6cff-8b11-42a1-bb6e-e1627ca17049 · outbound

This paper cites Can language models learn to listen? In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 10083– 10093, 2023.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Can language models learn to listen? In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 10083– 10093, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.037246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.037246Z digest=sha256:9bcb81c67ac52f823d78cabdae9bdd00e78302af8e3c260d058bd173e41491c6

Observation f4c9052f-61f4-4195-88c1-57fca4cb7668 · outbound

This paper cites Improved denoising diffusion probabilistic models.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Improved denoising diffusion probabilistic models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.062905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.062905Z digest=sha256:85ffe2c304f4a1ceb4d556583a1ca1f5fdd5633fd3cd54456e888a4b05b807f2

Observation 37a20fcd-a049-4c31-854b-caa0b7d91809 · outbound

This paper cites Interactive Generative Adversarial Networks for Facial Expression Generation in Dyadic Interactions.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Interactive Generative Adversarial Networks for Facial Expression Generation in Dyadic Interactions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.071744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.071744Z digest=sha256:466939540084584ce8f84668a93f6ae3ec201035d694d2a61ed9319788644b45

Observation e5d4972b-7802-4113-b013-61614c9231bc · outbound

This paper cites Scalable diffusion models with transformers.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.078129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.078129Z digest=sha256:650866392c675c83e6db89ed1b56baee4252a9d8bc83b2aa3520edeb25b8b419

Observation 27edae4c-e825-44a0-b3c6-5fe5752aaba3 · outbound

This paper cites Deepfake generation and detection: A benchmark and survey.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Deepfake generation and detection: A benchmark and survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.085275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.085275Z digest=sha256:79c594894338b97eed6e77c1f654124af5e91a8790d6cf327d85648038e031e6

Observation 6831dd92-b20f-4167-8e62-8155f217adc7 · outbound

This paper cites Emotalk: Speech-driven emotional disentanglement for 3d face anima- tion.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Emotalk: Speech-driven emotional disentanglement for 3d face anima- tion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.092405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.092405Z digest=sha256:9cb27b0e0e725d95a1c5a27bfcbdc66af44a7dd70cfb6070bf1d2382e9c9bb54

Observation ba468625-a194-4145-8412-3092fe7b882a · outbound

This paper cites Emotiongesture: Audio-driven diverse emo- tional co-speech 3d gesture generation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Emotiongesture: Audio-driven diverse emo- tional co-speech 3d gesture generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.082449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.098680Z digest=sha256:45e45f010538802709856923d27b7a229b3776b8c1f0c6faf560978ee9dde096

Observation 27cedaf2-6692-4ad9-ae37-ac8677dacbc5 · outbound

This paper cites Weakly-supervised emotion transition learning for diverse 3d co-speech gesture generation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Weakly-supervised emotion transition learning for diverse 3d co-speech gesture generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:13.010525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.106685Z digest=sha256:eb96a740cc86e2bb898d0489770cb5e75d78f370bd273ea9f1f7e84e7b109ac7

Observation 268a7236-f29e-49dc-bdbd-ac9ac9699f48 · outbound

This paper cites CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.115102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.115102Z digest=sha256:075232645a5f5b8ce9f4104b0c4617119706bb2b788d72d87be6dd8a6df50e86

Observation 41ca3b71-2927-4f6c-9bfa-e7d9d473fb5b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Learning transferable visual models from natural language supervi- sion

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.978281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.124172Z digest=sha256:e92835716712025e4d9b5103808572e348461b51086a0065e0ebff84f32a43d8

Observation 344704e7-b287-48db-af3e-f3bfff8d9971 · outbound

This paper cites Meshtalk: 3d face an- imation from speech using cross-modality disentanglement.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Meshtalk: 3d face an- imation from speech using cross-modality disentanglement

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.933022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.133319Z digest=sha256:a98525a5c7e0e5be9b7b4a8b65c15d070246d7b2ed08dffd99b0695ee3c7a6dd

Observation b35b30f4-0945-4fbe-baf1-87ce0f0e405e · outbound

This paper cites Quantifying facial expression synchrony in face-to- face dyadic interactions: Temporal dynamics of simultane- ously recorded facial emg signals.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Quantifying facial expression synchrony in face-to- face dyadic interactions: Temporal dynamics of simultane- ously recorded facial emg signals

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.894686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.142188Z digest=sha256:57aeed7ee81c0d9b94309ed0d7c0b6ce8668d22e4bc6be50bf514b1ea9deae26

Observation 455c3b56-d068-4fa0-81e8-2624ea53531a · outbound

This paper cites A circumplex model of affect.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction A circumplex model of affect

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.855289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.149565Z digest=sha256:eb1d6675443ce1ff6593bc05d56a0dfbf53f89e27152889cfae0e59b8e5a79be

Observation 9603134e-b518-4dc5-b07c-62ed96edf06e · outbound

This paper cites Denoising Diffusion Implicit Models.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Denoising Diffusion Implicit Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.157904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.157904Z digest=sha256:13e7540887eb84563d4fa5e49b33b107d728b06821848df2361e51f237f11a86

Observation fb0d6c2c-861c-48a9-bed6-22f8d908cd8a · outbound

This paper cites Emotional listener portrait: Neural lis- tener head generation with emotion.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Emotional listener portrait: Neural lis- tener head generation with emotion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.820123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.169039Z digest=sha256:3225a24b49183a439a1f478dc74a7cdb65b61e7f8d1f688b159dd343fcbadbc7

Observation 77716b6b-958a-4cde-a796-326162da239f · outbound

This paper cites A con- versational agent framework with multi-modal personality expression.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction A con- versational agent framework with multi-modal personality expression

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.772221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.178397Z digest=sha256:eedbb1fd73d01df4e90c1ec795102caa81d40338756160c7e3c9a3af20129d13

Observation cd226d1b-3117-4fc2-b790-4394809254a9 · outbound

This paper cites Artificial intelligence, machine learning and deep learning in advanced robotics, a review.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Artificial intelligence, machine learning and deep learning in advanced robotics, a review

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.721129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.185761Z digest=sha256:aac3eb7b4bde301383bf56b8bf81d7c60c09b1c2327730626253b55ae76cab34

Observation 5d05e84d-d8e5-473f-bdc2-64f04132c080 · outbound

This paper cites Analyzing human– human interactions: A survey.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Analyzing human– human interactions: A survey

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.664218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.192971Z digest=sha256:2699b200166538d7270106de182fd6fad504419a5e4dd850eba5fc2ec34cb138

Observation 7e81b087-0169-4800-845e-a8c71ebf0a97 · outbound

This paper cites Avi-talking: Learning audio-visual in- structions for expressive 3d talking face generation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Avi-talking: Learning audio-visual in- structions for expressive 3d talking face generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.635576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.202100Z digest=sha256:21c7966c26901653e9674008b06ca8ff8af2e6c5d0c6d2acbc8b9b79cb64378a

Observation 7da6a430-c4c4-4e48-96c1-7f3d181a41fe · outbound

This paper cites Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.377967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.209316Z digest=sha256:120521b511d6897f7edb0a272cbc29ad01b25690f85403e11071c4313360971a

Observation a4dd25e5-b010-4179-b6d1-4437b600232b · outbound

This paper cites Listen- ing head motion generation for multimodal dialog system.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Listen- ing head motion generation for multimodal dialog system

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.341819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.217844Z digest=sha256:aa798c7a7778d6d2aaecb670ee3ab6b362c576c1c11ff15036b198d8290e0750

Observation a870529d-112d-4360-96d1-949cdd43fcfc · outbound

This paper cites Imitator: Personalized speech-driven 3d facial animation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Imitator: Personalized speech-driven 3d facial animation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.300406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.234959Z digest=sha256:fa5c37dcfcd56b0cf379da3632fa008371177862e9bf14d43dbaef01f1b885b6

Observation de91a6b4-a88a-4e2f-9b9b-8e95adfd8bfd · outbound

This paper cites Estimation of continuous va- lence and arousal levels from faces in naturalistic conditions.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Estimation of continuous va- lence and arousal levels from faces in naturalistic conditions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.269947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.243101Z digest=sha256:529463f829bb3aced12e751d7c284eef740505f50a5756d90188bfa13c4b4e21

Observation c74135e3-8fca-4eb4-b87f-22644a5a08db · outbound

This paper cites Dim: Dyadic interaction modeling for social be- havior generation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Dim: Dyadic interaction modeling for social be- havior generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.230779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.258044Z digest=sha256:1603082a481308e165a9c1a60427c76ace58c435d2642698f0df4d2ca78d3246

Observation c1d27f79-3330-4a5f-bc40-e99f6c04e60a · outbound

This paper cites Dyadic Interaction Modeling for Social Behavior Generation.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Dyadic Interaction Modeling for Social Behavior Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.269099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.269099Z digest=sha256:2b8e9e4dd64baafca3b35742ea8951acfd372fafd07e2b76a14696c85167bb9a

Observation 9d27516c-26b4-42bb-bcb7-90b98cf80d85 · outbound

This paper cites Neural discrete representation learning.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Neural discrete representation learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.281840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.281840Z digest=sha256:c8f9abcf9eff63c8e50911928feab66b9a7b6bbab2e262a3ddb4de5c4be401bf

Observation 3a8875ff-8087-443f-b852-4ba440d245cd · outbound

This paper cites Effects of inter- acting with a crowd of emotional virtual humans on users’ affective and non-verbal behaviors.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Effects of inter- acting with a crowd of emotional virtual humans on users’ affective and non-verbal behaviors

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.163347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.296724Z digest=sha256:023a877377c875a90bc6fec01ccec72be5bd3e8d2036bcaf24d4068641a4b05a

Observation 8b475221-0252-4c55-aebe-8f6a5fa49fd2 · outbound

This paper cites Versa- tile face animator: Driving arbitrary 3d facial avatar in rgbd space.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Versa- tile face animator: Driving arbitrary 3d facial avatar in rgbd space

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.120344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.312863Z digest=sha256:d7300c59c43a3bfddadda37c4003e3b666140ce6626248a712a365e1c108c861

Observation e92475be-5550-4bf5-88f6-47b6d3c6bf26 · outbound

This paper cites Speech-driven 3d face animation with com- posite and regional facial movements.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Speech-driven 3d face animation with com- posite and regional facial movements

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.086190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.322039Z digest=sha256:42e9c89f984802aae4603a067311a3c2b9822f2ecb7021f88e11a8f82318516a

Observation 8d9af6a0-db0e-4e47-b410-6dc2ec775601 · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:11.332706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:11.332706Z digest=sha256:9523d76307b043d8f4b90ce74f86b8d5c0720031f1502ba91761e10239bb50a6

Observation a73e35c7-7bd1-4e45-b386-f14cf3ef253d · outbound

This paper cites Media2face: Co-speech facial animation gen- eration with multi-modality guidance.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Media2face: Co-speech facial animation gen- eration with multi-modality guidance

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.049094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.338244Z digest=sha256:4e80f6b4d14f0dc68fb25aa313edea29a7b139583d3c22d9ddbe01e170756352

Observation b698f257-03c0-4a82-b767-8aff0c9e65b7 · outbound

This paper cites Responsive listening head generation: a benchmark dataset and baseline.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction Responsive listening head generation: a benchmark dataset and baseline

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:12.004407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.345532Z digest=sha256:ac14bd3c9ecb4170a2c7753cb322750de4e9ce1ec29a70313b69b7537619ee04

Observation adbedc43-3381-4d89-88c3-8d6aa603ec1b · outbound

This paper cites To your family and nation, being a trickster is destructive, but it’s like finding God in freedom.

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction To your family and nation, being a trickster is destructive, but it’s like finding God in freedom

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:00:11.979278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:00:11.354369Z digest=sha256:6b18da3409c752c8f310fd400c29dfc91a48fe5b0668c2ed18aad73ef2c88a1a

Pith citing papers

No inbound Pith citation observations are available.