Pith. sign in

Paper Citation Record · LEDGER

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

As of 20 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2506.05806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05806 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:27.810835Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:45:49.611770Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:58:51.404766Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca17bc80-f60d-4a04-98ef-1412b9481454 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models A lip sync expert is all you need for speech to lip generation in the wild

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.196349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.196349Z digest=sha256:4f528775e5727962ad6f2e9fa2bafd2fe23dd5d56ccedbb7555d0bcbfd7e3d69

Observation 414366dc-3bdd-410c-bb28-36aa5f59796e · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.303790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.303790Z digest=sha256:07122ad81e9de33f8414a1527daf16659c63fc815faf104f67e941db1d092edf

Observation f43dcb10-69d6-49e7-bdeb-7d8b0775adde · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time.Advances in Neural Information Processing Systems, 37:660–684, 2024.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Vasa-1: Lifelike audio-driven talking faces generated in real time.Advances in Neural Information Processing Systems, 37:660–684, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.353953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.353953Z digest=sha256:320e8f3ed137f79a57f04c00061630ec5fac78d1176332482f628c062ce688fb

Observation bff6e7e8-659e-4164-a4ff-5516565eb2fe · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.735534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.468957Z digest=sha256:d04c0871ace8b5c81a43447eec3133acf83a78a8e01b89d996dd9160c80607eb

Observation 67a55ddd-daf9-4f89-b61c-93d999784cdc · outbound

This paper cites Denoising Diffusion Implicit Models.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Denoising Diffusion Implicit Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.581989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.581989Z digest=sha256:157fae7dd6ec2ea1b90bfdf0a165bd57ef86150a5cb36833ea0eccd2d3eb26d3

Observation c55c8c15-75b7-4eff-bc7f-ef6a04219bd3 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.702045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.702045Z digest=sha256:c5584bf23b365ea62c8e0b6213c7ea109e37ff7f79893f5524d47eb98f183933

Observation 3dd28e69-e6fd-4311-9756-8f0bc0ac0d65 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.849703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.849703Z digest=sha256:367fc8caa952123513ab256aa4e354caa4ee434fcf8f0b5c7bfa577302ffeaec

Observation 7d26a6e7-3c58-4ad6-a627-9edd312330b6 · outbound

This paper cites Megaportraits: One-shot megapixel neural head avatars.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Megaportraits: One-shot megapixel neural head avatars

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.479575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.901840Z digest=sha256:92ea4db5afd2532d828b8caa1e024140c4d87830499cbc478d55cc42c6377eff

Observation bc068e31-25cc-4c12-bd41-c1c26721d8de · outbound

This paper cites INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.943246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.943246Z digest=sha256:90fb7af0fb33869eb270a1129fa98728a8ae00b30a5e498d95213b09054f7207

Observation 81f47936-4d82-4270-8276-134aa87531d0 · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.000147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.000147Z digest=sha256:679ccdc117cb2b5e802549700e4e76c81bd3abd2e64cb24c814525168221ee08

Observation ddb92de9-fd10-43c3-beb7-8a690711cdf2 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.119456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.119456Z digest=sha256:78f35e66106e51e431100e768a5959885adcf7836379a6d924ca80ba5d323f6c

Observation be5e88b4-c072-4605-8b1e-ab752b2d2f3e · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.230525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.230525Z digest=sha256:410c7e940a028263a4e55699f893573286fa57e83b364354958d57b6a3341a4f

Observation e92e6c34-72c1-457f-94e0-37c71a4b7bc5 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.346992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.346992Z digest=sha256:848491abd22ce5a2831afb441b0f826b1ef14a0955e535b7a42dc12f3098de3f

Observation 8dce6ce4-191d-456a-94f1-149be2ffcbab · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.450551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.450551Z digest=sha256:cbdc4cfb1eaaa049ddc54e04278c32d66ea54461964e17caaccb41892a09d5ea

Observation 4a76ebc3-cbf2-4246-8763-529b4b596b24 · outbound

This paper cites CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.579865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.579865Z digest=sha256:2dbc3dc143098c2b281e7c9341a9b57710cb09f49f14763cb57c0d71c4025418

Observation 33133a8a-1ce5-4dcc-8d69-a991d324d416 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.712720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.712720Z digest=sha256:bc19ef9e90ee500cb96e75f7325a11ebdcf0df9a11a580493e154c87708bf1ae

Observation 697f31a6-831e-4965-8661-c50adda35c8e · outbound

This paper cites Consistency models.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Consistency models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.846533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.846533Z digest=sha256:e801bb8f86c2b4699e4666400b5687be65c5cc41cdae5208ec27f3b28ed24421

Observation bc55c5a8-c0ba-4577-ae14-afadc0db8298 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.944161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.944161Z digest=sha256:6996176e73c894152f1cb454df73994dc788ce2407a73dfa4b7616643725e1e1

Observation e5eaeb69-6529-43ee-972c-34fcf1bf1952 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.039278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.039278Z digest=sha256:680fa68e14a973fe15c2195679f2a09805b1c8ef4c698192856355d29fb24f29

Observation 39e9e265-0112-47c7-9195-558aaf3c522d · outbound

This paper cites X-portrait: Expressive portrait animation with hierarchical motion attention.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models X-portrait: Expressive portrait animation with hierarchical motion attention

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.186798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:25.147513Z digest=sha256:7566fa7ad182ad7c9c42c0055abcf4bf40300199ef1ab3c07f8951d9bdc1675d

Observation f43c7580-7b67-4ca7-ace9-e864facf826f · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.256345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.256345Z digest=sha256:c87fdeb848b5a22aea74abfa9372ba45e7f589ca1d3cfeb16675c5eda85fbfbf

Observation 1dc0c154-2846-459b-86c5-c2fdb5b91be2 · outbound

This paper cites First order motion model for image animation.Advances in neural information processing systems, 32, 2019.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models First order motion model for image animation.Advances in neural information processing systems, 32, 2019

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.377340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.377340Z digest=sha256:5b9ed39472d76ec2692b2c686c42efe11f38560b34eda00e1575290d622a5621

Observation 8c23bccc-b356-4b6b-bf42-54674df3b85a · outbound

This paper cites OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.487102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.487102Z digest=sha256:d35b2ad479722d58b876cf99f998853b3f1782ed2cb3a04e2837ec93ce493762

Observation adf63a13-6dec-4a02-bbb7-b95d647a2d93 · outbound

This paper cites ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.637600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.637600Z digest=sha256:bf08fc39b1cc69a4a79e4f2f6200168cd27a32d169a17f645d412e7cd28fc930

Observation 717ece2f-dd5b-439a-aad0-e5076817d8f4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.735445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.735445Z digest=sha256:8b98c0e9d000ae749c96aa74da50bc7b52bd5689c8b3e3963ebe46fbfd1c9a5d

Observation 0717dae3-1cfa-4ab1-995d-90e61e9b060a · outbound

This paper cites GPT-4o System Card.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.836802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.836802Z digest=sha256:ae6ea1ad015aeabd132bfe0455dc980a0baaf734a9dd3259336a886d7399f7c3

Observation 968c5f37-a31d-4231-9f4e-a29f2ae3147a · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models wav2vec 2.0: A framework for self-supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.969650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.969650Z digest=sha256:2507c236a076f8aafd621a99749685b93c7f242d9e0e5e5bcb845aaed023067e

Observation fbf8f96b-d270-49c2-93b7-92681762c5ff · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.115489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.115489Z digest=sha256:99da684540e31d76936bf54c237793aba3b5a1f59de29d99d92fbfcd2adae399

Observation 5943c728-876a-4f5d-874f-dffb4ad6e053 · outbound

This paper cites chinese speech pretrain, 2022.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models chinese speech pretrain, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.032142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:26.276961Z digest=sha256:3f98e321ec5ce57d29c46af3ef4af5e0bc0fccb291dbfb3401ec8c0ed6b75055

Observation 2b9f9eef-7a4e-4e63-a5ac-b45a97da4f5b · outbound

This paper cites Hsemotion: High-speed emotion recognition library.Software Impacts, 14:100433, 2022.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hsemotion: High-speed emotion recognition library.Software Impacts, 14:100433, 2022

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.831253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:26.409346Z digest=sha256:42c488a00bbb94ca585bdfb25ebe7a76d88fbfdc9b81e155dd9e97252c4b6193

Observation a47c9991-4a64-4092-987f-714aa37a07b6 · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.580353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.580353Z digest=sha256:536e1c5d1cb0a151fe58c48ed9a9ae546a8706b85bce7512e61880809d279edb

Observation 50b578e2-7349-4c67-9925-0e2ecbd5ec06 · outbound

This paper cites SDXL-Lightning: Progressive Adversarial Diffusion Distillation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models SDXL-Lightning: Progressive Adversarial Diffusion Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.743525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.743525Z digest=sha256:22ba75d8a38292c0aa918f98b6e5e37d8a38b6cb88840bceab5bb29bc12a6595

Observation b2e01aa1-ed59-4949-a6f3-8843792ae7a4 · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.967027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.967027Z digest=sha256:ad891bcc050f51ba48fb8e62f1ee286f1609ed77c96d74fc6ff790b6af7b0d7f

Observation 7dacbc6b-1c76-4e54-a9b1-191da6f89350 · outbound

This paper cites Animatelcm: Computation-efficient personalized style video generation without personalized video data.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Animatelcm: Computation-efficient personalized style video generation without personalized video data

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.603473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:27.146314Z digest=sha256:9cd4be4ea6b91557ac8556b3e6c1ca27230918ca824ac15cba87f290db1c3724

Observation e31a5c23-a23b-4ce4-9c00-b9aefb053ddf · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.295219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.295219Z digest=sha256:fb85405dbe319666d53483b1ac3bb5b138babd0f774d9f3009ad652d04908cd5

Observation dac69101-cd27-46dd-8147-7781b68b1b9a · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Mead: A large-scale audio-visual dataset for emotional talking-face generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.369641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:27.412988Z digest=sha256:2b29c3e897eb7495f749ba1866689f360d50e9bdc69f3278cb74ba149a50a522

Observation 6fab25fc-0325-4226-9fef-b30f63d9c48c · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models VoxCeleb2: Deep Speaker Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.509209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.509209Z digest=sha256:84d9ac4431c2fbedf89cfd270cb57c6814b61b6b8bd1a9d6709cef069ad7dc5f

Observation e97bcf03-adde-49d0-9043-b325031729a4 · outbound

This paper cites Celebv-hq: A large-scale video facial attributes dataset.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Celebv-hq: A large-scale video facial attributes dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.601101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.601101Z digest=sha256:6480852a186cb65b1e1d6f21f890c36e506b570e3b82be9d9899798d498c9db6

Observation 548058e6-ad0b-41b6-8251-ab0ed6dd69a3 · outbound

This paper cites YOLOv6 v3.0: A Full-Scale Reloading.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models YOLOv6 v3.0: A Full-Scale Reloading

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.700080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.700080Z digest=sha256:c76ddabda02228d57ef0e3b6c42c5fc9addd9a105076207078b0957dc9562ef0

Observation b2da14e8-311a-41ca-a8be-501cc43112c5 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models MediaPipe: A Framework for Building Perception Pipelines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.810835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.810835Z digest=sha256:70419e233924a6b0f292d361a58b4e7b45e70f7bc78362c16acc4ae8bb722d02

Pith citing papers

Observation d188d4fa-de9f-4e3e-9ba6-660af4482247 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.407987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:5d9b82238a86b19cb2d4f58a0c56ec7372c91d28cf44ed642905b9f9a35429c1

Observation fb09a11e-2136-4d79-978a-df66026497ea · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:43.646679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:43.646679Z digest=sha256:00b6ac04f5b873bc0ba8ed5a1d13c6518013a575602bc7c1bcf7c9be921d44a6

Observation 72ceac6e-82b4-470d-b963-349472f310d9 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:49.611770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:49.611770Z digest=sha256:b868c40715c0c3f3829098f7d7cd6ee73a587033754a7b00dcd5286f5cbc2ada