Pith. sign in

Paper Citation Record · LEDGER

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

As of 7 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2606.01620.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01620 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:41:59.683979Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact26
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c813f42c-f684-4bae-b51c-48dfe9a889d5 · outbound

This paper cites an unresolved cited work.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:210c8cc4a4ec3266d6f09d9294875df715fecfe62c829bc327f53f518ba5d2a7

Observation 83ca6265-f657-48e3-a336-382951465e1c · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.013473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:8869a3a0d7e05393cf7a62369fb283187c22111639e6f0025de02c606f78fcf9

Observation 4c9f3cf7-807c-4417-a283-666403fc1a0a · outbound

This paper cites V oice puppetry.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs V oice puppetry

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:de00de871efcfa5fe357497b58ec300c5baabd79d624b2c0ec0c3b6f9e56ae96

Observation 7c8bda84-1a96-47c1-935f-cb4ea9608c49 · outbound

This paper cites Video rewrite: Driving visual speech with audio.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Video rewrite: Driving visual speech with audio

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:c35a51d3b111a0de2b62ac01fe66913cf2d541326e99a72b8d014594d127d168

Observation 5739b4dc-38b7-45ed-b468-484c6a477c62 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.Advances in Neural Information Processing Systems, 37:24081–24125, 2024.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.Advances in Neural Information Processing Systems, 37:24081–24125, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:7006ae77b8af75b874dd49d45c85177a7d91012dbb2594b6905de814cb38867f

Observation cf2927df-d842-44d5-ab01-d491843454d1 · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.016488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:a43653f4821fa5285a62fa328b70e3c2d3976a5525bc26f06d50c4e9080454fc

Observation d0c54751-da2e-4982-928b-52405f182036 · outbound

This paper cites Dc-videogen: Efficient video gen- eration with deep compression video autoencoder.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Dc-videogen: Efficient video gen- eration with deep compression video autoencoder

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.019820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:5ed2fdccd6156519d8f210d1e263d0cc5c65b287d675e61cd5e80af4568eca66

Observation 84883b82-5ab2-4acc-9875-908523c39f90 · outbound

This paper cites Lip movements generation at a glance.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Lip movements generation at a glance

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:56ec40a36ee03de7a347e96b2049bce3d0bc40abfdccb8ae24fa831a503e93de

Observation e8fb1432-b4cb-4bfd-9ad5-3c106087b9c5 · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:804c3c3baace2a70101de60e2012d0f3213f2284486445e96cf47f5ee74beeb6

Observation e3bff521-6d9e-447b-93e2-9a13aa0a305f · outbound

This paper cites Out of time: auto- mated lip sync in the wild.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Out of time: auto- mated lip sync in the wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:d54e634f9bcb5fbc58ff7a10fe7ce81bc5832ef6b005d62c087c8afc5d4bb050

Observation abad4ece-1d32-4592-a098-be24efa0a62a · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs VoxCeleb2: Deep Speaker Recognition

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.027913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:1a74d90b15303b9143ddf28952c034b8ccd260426cd387d8f1409a941d077619

Observation bd568a8e-1ed5-45e7-b378-c40b135ec2fe · outbound

This paper cites Hallo2: Long-duration and high-resolution audio-driven portrait im- age animation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Hallo2: Long-duration and high-resolution audio-driven portrait im- age animation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:3952a69f6c2e82197861076610c654b393b7135f2cbe0cc6a9e34d5287f28a28

Observation d182d93f-391d-4f3c-9894-f5b9293c93d7 · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:89dbf62f26b99c8862fdcfecd7f9884e745eccb343ce0d24ed209417ce344523

Observation 68a15157-24ca-439d-ab10-cf9f7192b1a6 · outbound

This paper cites EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.978617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:6f02eb716b729c7691155fff97113aeda2de05cfae9b3844b55003cbd8cd6fbc

Observation 78929906-de10-4040-bc24-d20d1778260e · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:ef12c2c9cdada87ca7ea3eea3b5604d553037aa4ee58512987a192befb0013cc

Observation faa4bb31-951f-4461-b8bc-bbda93ad11fe · outbound

This paper cites Introducing gemini 2.5 flash image: Our state-of-the-art image model.https://developers.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Introducing gemini 2.5 flash image: Our state-of-the-art image model.https://developers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:d09c2304e95ead1f9325c78300397a0e60b84a07f7f7b82c59fa38523b355604

Observation 0b3e1e87-c56e-44ce-a5b9-47ae08da7fc7 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.968008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:9ca5e652a1adde0aa814bf3eaac37efbb18ff545efc4958e72fba646679e8ac4

Observation 128d84e6-e17b-41dd-8ebd-9da7e4358dd6 · outbound

This paper cites Space: Speech-driven portrait an- imation with controllable expression.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Space: Speech-driven portrait an- imation with controllable expression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:a38ea14e39655740328188746d17d1045b87904df9ade5fdeeb0aba6abfe44b4

Observation 0bfd662d-e545-45f4-bd5b-b2c13df345b6 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs LTX-Video: Realtime Video Latent Diffusion

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:16.975818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:134a6de01d193f4d0e754bf00ffa907cde1721637c0bd2434bc59caafcf1376a

Observation c7086045-68a1-4c48-a328-e741cc56d020 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Classifier-Free Diffusion Guidance

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.022220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:7be7db5429e51a2e9b56a362bc2ee468a4b727ad134a1c4a01d2021368448487

Observation 404d936d-962b-45dd-bf0b-9a73af6c4195 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:e66c96048a4a012e55d4311f4dd6c4179d6622b1f09888a0588abf0f78862cd5

Observation 32d6b69a-ae24-4da6-bd77-39fd3d3a2434 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.025068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:fab6fa81b57430ec50f436732ba96c73639be5aa7fd8a40603cc7d40d83c677d

Observation 348273fd-297f-4451-865d-48fcf9465b53 · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Eamm: One-shot emotional talking face via audio-based emotion-aware motion model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:2e464449498a6f87d449ac235ace1a8a1edca7188151194f37124b7f2fd60379

Observation c439a47e-ae37-440c-992f-82413fc84d00 · outbound

This paper cites Sonic: Shifting focus to global audio perception in portrait anima- tion.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Sonic: Shifting focus to global audio perception in portrait anima- tion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:4a7ee9889a421934451f36077b533e6cb22af395dbc33b5d6639128f38d10220

Observation 7af8d6f1-de50-408a-88c2-75fbd8775aab · outbound

This paper cites Auto-encoding vari- ational bayes, 2013.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Auto-encoding vari- ational bayes, 2013

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:ef58af0818159a89496d1ea8c3066f74434abc8f437656c2812809b5a44a8b5e

Observation dd95b62a-58e9-4853-94f9-a9075aad1e5a · outbound

This paper cites Tokenmotion: Decoupled motion control via token disentanglement for human-centric video generation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Tokenmotion: Decoupled motion control via token disentanglement for human-centric video generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:f42abe96bcb77013f7c8b2d634061741773dca5408d5ffb779cee09cbcd38479

Observation 77dd19b0-02c9-4605-a2ad-aed09b59b4e4 · outbound

This paper cites Expressive talking head generation with granular audio-visual control.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Expressive talking head generation with granular audio-visual control

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:0a4b2ed2883fc76d3a9bef2fce55413e8be98fb91c9553f443d80d4f90f12019

Observation d8f06aec-259c-4c30-9bd3-b12057a11376 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.031391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:aeb78312465f794e4d6a9cdd26dc44dbbbaf63e5eef063a6adbea2d89b202c27

Observation 28fe0d3c-60e9-4a85-8542-8a889451a53a · outbound

This paper cites Diffusion adversarial post-training for one-step video generation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Diffusion adversarial post-training for one-step video generation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.042850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:f85e4d8b4bab3cf71d67a818e92c0f8562f28f41715b435de7c18858bc213767

Observation d20be360-94c1-42cf-9542-e8a98fa6d7ef · outbound

This paper cites Autoregressive adversarial post- training for real-time interactive video generation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Autoregressive adversarial post- training for real-time interactive video generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.037239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:fd3a79ad48383767b9abc49a68ead786638fc76ca90f7a83528f658af1d209e6

Observation 42eb07bb-d5c3-464e-a546-3a0bc46992d7 · outbound

This paper cites Flow Matching for Generative Modeling.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Flow Matching for Generative Modeling

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.042323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:0820c12da3b5d587a336fda706a22739498ecaf469425004b241f42607b27c0a

Observation 2578d23f-a9fa-472f-9cc1-2a4f51157076 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.036750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:f45b88b79e444195ef112097f56e5ad86d182996625b984409b2313c6ce6e6c3

Observation 8824b76c-9096-4e16-acbf-911c9d25e250 · outbound

This paper cites TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.028231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:451c0749954bfe3f8f93e40b8dd106320b8aeab61b37438269a41bb99bc0a9a8

Observation 63907d27-5bd1-4706-acdd-6f4e2e0b7191 · outbound

This paper cites Scalable diffusion models with transformers.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Scalable diffusion models with transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:01298a19ebcbb3010ee1f85e01348c25560ea75bd5268cc38b4627aeca44e00e

Observation 837f0016-0877-4393-a5be-724c0ca2cb0b · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.007451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:1c54286f105f56b29028c4b6e788109b0500a726629e38590949370da1bd11d7

Observation 080d370c-711e-4619-8f20-38259de1b6f4 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs A lip sync expert is all you need for speech to lip generation in the wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:a7868ba64566bbf77eaec048b5bf11b054d78071b56d3a150e0ac1c61d35c25e

Observation 07d79fc0-424d-4c60-9975-3ef5356e3887 · outbound

This paper cites ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.004633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:95d0c5075fa526e57b37f0bbc6509211b26006b4c91e4758af7529db71a882d2

Observation ca902b09-28de-436b-8798-6240138994d0 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs High-resolution image synthesis with latent diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:5cd645c5a320f443fe96fffc53a086add83e0b55a9c2f837a30427b2daa2e6c2

Observation b3f0d381-da78-4aa0-ad2b-15a1ccd747ec · outbound

This paper cites Fast high- resolution image synthesis with latent adversarial diffusion distillation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Fast high- resolution image synthesis with latent adversarial diffusion distillation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:1fa2409e34028e4f5a950d29469d79d55d47ff9c815fe738edf3612e2d016573

Observation d598ce8a-f2a1-459b-b5bc-5b82075bb1c6 · outbound

This paper cites Adversarial diffusion distillation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Adversarial diffusion distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:fcc1014abac433d93cfde0ee5ccb16c799fa42a2d8637f6ee5c3a3230b290c0e

Observation 7c11c598-c407-470b-837c-4e473c7f0831 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Score-Based Generative Modeling through Stochastic Differential Equations

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.010726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:4b64ab919cb7688d6457765882f56483411eaa92e5c23e37f8b4e1a38015948d

Observation 382db98f-2f1e-4f0e-82f9-a28b56fe37d3 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:0b42396a6dc1b4c83e8567055eb619b907d83b678a7073ea196d3edfc474fe09

Observation ed3414aa-3d56-43ed-86b7-623ea378a884 · outbound

This paper cites Synthesizing obama: learn- ing lip sync from audio.ACM Transactions on Graphics, 36 (4):1–13, 2017.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Synthesizing obama: learn- ing lip sync from audio.ACM Transactions on Graphics, 36 (4):1–13, 2017

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:247d92326c03ab4a586d6bc73720558ebed072d1587afba55d8c0bb33f9533f3

Observation 1207205d-4708-4eba-b7f2-6443f6133cc3 · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs MAGI-1: Autoregressive Video Generation at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.002298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:9a166b67d53ff0239564890a6146fb6a8af08538c8b3d2526392fc22854eaccb

Observation 6c77681b-e29f-4049-af64-1e9cda9efae0 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:7606a76559f15c1339dd8e459e806d48f35d94983bd933fa188fb498abc3f08f

Observation e098ef0b-d711-436a-85db-e978b75d2a8b · outbound

This paper cites Reducio! generating 1k video within 16 seconds using extremely compressed mo- tion latents.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Reducio! generating 1k video within 16 seconds using extremely compressed mo- tion latents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:48019f4df57555ff24730ad59ff8f4a7c8715f33a3c508c19cd9d489d906efb3

Observation 1ad01a55-9503-45e5-a82e-a7492f7f1d9e · outbound

This paper cites Fvd: A new metric for video generation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Fvd: A new metric for video generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:aa2c6f3e3c3a1abfe01bca44a39cb3466b0ad151b1adb0c60ee782d21f21f0d9

Observation e25ec6e8-9d5f-40f5-b002-445fa2f40815 · outbound

This paper cites Conditional image genera- tion with pixelcnn decoders.Advances in neural information processing systems, 29, 2016.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Conditional image genera- tion with pixelcnn decoders.Advances in neural information processing systems, 29, 2016

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:abad44e20e5dc658feb1101af94c3ca89517f7ca762f9eb65300c9fd9c2ce566

Observation c4c3dd92-4805-48e8-88e7-cb56beb32e2a · outbound

This paper cites Pixel recurrent neural networks.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Pixel recurrent neural networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:f138a035898e0d7b96e636090913d5b83a91ba0e1fd2cb9ab3e349294cbddf77

Observation c06e113e-91ef-48c1-b6f4-fef71baf7dec · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Wan: Open and Advanced Large-Scale Video Generative Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:17.004638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:5ee46c450829887efc5cab0140fd731553ad70110bf5dfcb3c5d67fac58a735c

Observation 19201b67-4902-4a91-93fe-d0e67f4cf536 · outbound

This paper cites Progressive disentangled representation 10 learning for fine-grained controllable talking head synthesis.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Progressive disentangled representation 10 learning for fine-grained controllable talking head synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:892f5412257a187e4984c837d4bc00f277f7108296d080d4ddd647f158a1e57f

Observation fab4fd07-4d3b-4af9-9191-04be2ccc08eb · outbound

This paper cites Echoshot: Multi-shot portrait video generation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Echoshot: Multi-shot portrait video generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:1ad62e79cd51bb1b88aea258c3839ae6490391c532394d0f92805e4914d2d6e1

Observation 274846f4-57e9-4428-ae7d-daf5365ca4d3 · outbound

This paper cites Fanta- sytalking: Realistic talking portrait generation via coherent motion synthesis.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Fanta- sytalking: Realistic talking portrait generation via coherent motion synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:800dd42dbddf281e320ff733a02409583045e89ee9c80aa0eab64efa75c8d635

Observation 98847c52-c98e-4f91-a1a1-970789955024 · outbound

This paper cites Audio2head: Audio-driven one-shot talking-head gener- ation with natural head motion.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Audio2head: Audio-driven one-shot talking-head gener- ation with natural head motion

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:9cbf2c1cca1b2e0cb308458e3fda129cf39310677aa83902d12e4ea60f53e889

Observation de4bb315-8d87-4484-ac21-c4228792c5b2 · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferenc- ing.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs One-shot free-view neural talking-head synthesis for video conferenc- ing

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:cdd7bf8fcf37d186019399b5b61a83dad83730241200fc7bbfffe1859d93dd65

Observation 679e79a2-4c15-4b52-8c6a-2414065c442f · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.010720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:f6e7aa270cc977ba32f7c56d657a3c281c288f4f8867e52a6ec41b9b91c5c4e0

Observation 0019adac-4292-42ae-8420-c522e548afc7 · outbound

This paper cites A learning algorithm for continually running fully recurrent neural networks.Neu- ral computation, 1(2):270–280, 1989.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs A learning algorithm for continually running fully recurrent neural networks.Neu- ral computation, 1(2):270–280, 1989

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:ca7e5041e8c8d7ce7dd33e6857a5962516d96087f9820c94f6d6809d943e66f4

Observation 1a0bf6b9-050b-4aae-9500-3c2f8ad4609f · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.998665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:281c904c6eccb63a2e6601c1356e2401a59a2ed840c2f9f11303dd568dd08fa4

Observation 56ac3d45-0d76-49e0-af35-1bfce50c0b49 · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time.Advances in Neural Information Pro- cessing Systems, 37:660–684, 2024.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Vasa-1: Lifelike audio-driven talking faces generated in real time.Advances in Neural Information Pro- cessing Systems, 37:660–684, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:d000ea7c52b2f4ea6b8ad30d2af8663ac2f6357e35bebb79a5ede63782481ef5

Observation b1d6057f-fa6e-4fce-81af-0455903a8b65 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:16.995433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:46af438e6e4c0fd9a11b6f8fd9acba7a1da7e08008762706fc501c47263a9bb8

Observation c6b6c8bd-bace-48d2-b730-5f653fedd40a · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:16.983406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:66b106f15f8254d402aea2ef5bcc3c1fd6ec8e0e752bb9c09896cb78d4c1aa96

Observation b1aca6fa-de05-425d-afa0-29fcaae31a06 · outbound

This paper cites Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.987357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:43fba66f62c54d3a01ffea8d3ba282e5f596241f4d308ceec8eddfa20c04dcf0

Observation 6893e448-5f43-4a37-a0aa-e5f020447700 · outbound

This paper cites Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:508b049146506bb0b1bdbc23935a4196a7c512bd55fb2eb3adebb562cab19293

Observation b2fe0017-c206-4946-95fc-d2ba933d0ca8 · outbound

This paper cites Im- proved distribution matching distillation for fast image syn- thesis.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Im- proved distribution matching distillation for fast image syn- thesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:19c87caf00d3b99231436388d1bd580676a3adb04bd1a58a9f86a11f1ee0584a

Observation 92b13d9e-0fe2-4156-af8f-6e531913be1a · outbound

This paper cites One-step diffusion with distribution matching distillation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs One-step diffusion with distribution matching distillation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:f78f1bb8eab91c6e4413dc23f4c4dc1d71fcded9becb430ba07101f391c4b685

Observation b6e4521e-524d-46cf-8bbf-a039197b53eb · outbound

This paper cites One-step diffusion with distribution matching distillation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs One-step diffusion with distribution matching distillation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:0f6da9045fb897e645f472763e36585755d931d20dc784a441408089f1eaf251

Observation 3ba160af-7d76-4209-ba35-a4e79a86cff1 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion mod- els.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs From slow bidirectional to fast autoregressive video diffusion mod- els

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:5fad75ff18accc2ef5cde5ac17a1dacc9b2734c91c565de55f2d490edce68e11

Observation 3a6c1562-5294-4ff8-9492-01538a1ad6dc · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940– 128966, 2024.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940– 128966, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:dee46cd06b609b21ee3eb03411106df3cb7ca56cb17e1ad76a83508255f85f4e

Observation 26b29c8c-6b41-4f2b-aa04-8677dde9e0c0 · outbound

This paper cites Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.990094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:650221343e785bfbc3bc49af7454a3b1c113f71b7e5c009c48ff4b6858f41089

Observation e97a59e1-a265-4adf-a3b4-1f20ce1bd1fc · outbound

This paper cites Talking head generation with probabilistic audio-to-visual diffusion priors.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Talking head generation with probabilistic audio-to-visual diffusion priors

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:ecfc63f519115d309ec269fe1e0e180a387779eaa65439bf363eb61a24a44f74

Observation a44d403c-9d15-47d6-a899-e0b09b9fea29 · outbound

This paper cites Identity- preserving text-to-video generation by frequency decompo- sition.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Identity- preserving text-to-video generation by frequency decompo- sition

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:9c751fc8ba35e328eb3c91f9c646b14882b8e233dc68f60bd93506e73e685087

Observation fce1d9d8-c618-4f57-84b6-70d7487ca5c9 · outbound

This paper cites Root mean square layer nor- malization.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Root mean square layer nor- malization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:a0d12c9042c6daf0d437ae19cac2300c292c9aabeb851c87f6c90b04c47adabc

Observation b0d344b0-bbb0-4bf3-97e8-3b1cdd83a99c · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs The unreasonable effectiveness of deep features as a perceptual metric

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:1c488eb1e1759b99094f975db248b3ed9484a3172d6d6fd79d8005a22e8eeb14

Observation 21d60297-8308-45c1-b5be-8c109b7f30a0 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:4a09ac61dfdecc0b3118af2837204abdeebb81816a4e1528a9f121b7003860ee

Observation c498c47c-c7e2-4d2e-8720-73b1c5d9b11a · outbound

This paper cites Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:b94b09c26d0998f9fb46c265abfbd2583835bb6cd09c5af252b436560b6e12ce

Observation 42f90887-6782-47b2-92ca-d20516d65fbf · outbound

This paper cites 11 Taming teacher forcing for masked autoregressive video gen- eration.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs 11 Taming teacher forcing for masked autoregressive video gen- eration

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:2315cac83ef84d50aaa8e30c182ae62625ce4749d541f8eeee9d198181053d96

Observation 0cef9d99-0165-4744-a646-493882917e41 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:0bdef2fbe0d79466f32c016b7605115b6e5fe8f23d68064892d56d3cdaa25e80

Observation 9f1b0e0d-86c6-43cf-bbb9-693e66a7be1d · outbound

This paper cites split-first.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs split-first

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-28T15:41:59.683979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:0b4cd00a0d0f6ba34a9a07cfe27f2628e163297a610ba0f94355407f121bf189

Pith citing papers

No inbound Pith citation observations are available.