Pith. sign in

Paper Citation Record · LEDGER

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

As of 5 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2605.30230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.30230 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:50:56.947671Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact13
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 158878bb-2764-4cb4-bc92-b4db6567b7a6 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation A morphable model for the synthesis of 3d faces

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:fa278a138ce8af6c8016f397c66b1da4f2e4b353eef9953af6ce8e5f89bbf95b

Observation 54dc6a8c-bd38-49c6-92c4-6b7f186af7da · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.352553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:d4b548f184e427757854feb62f2b56906cf42d195dd5bca3aae2db9173e3c5f8

Observation 29a8aa5a-91fc-4ec1-bdb9-454dee1a07d5 · outbound

This paper cites A no reference image blur detection using cu- mulative probability blur detection (cpbd) metric.Interna- tional Journal of Science and Modern Engineering, 1(5),.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation A no reference image blur detection using cu- mulative probability blur detection (cpbd) metric.Interna- tional Journal of Science and Modern Engineering, 1(5),

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:a71764a55e28e4d6d21e0bdf7a13d95215273eab28c7e34d159c9e3ac98cc316

Observation e4abc6e1-f003-421a-9047-bc0b296d80f9 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset.IEEE transactions on affective computing, 5(4):377–390, 2014.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Crema-d: Crowd-sourced emotional multimodal actors dataset.IEEE transactions on affective computing, 5(4):377–390, 2014

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:973b98f01a62ac440d8fee7e1f60764cacfeb41d4b2823acac7632e620cc8c92

Observation 22e9ea92-5e1f-41fa-8633-8fe1fa90e322 · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:e5152962e9915ae75896329cbebd2d738acd6cec8819455a6651a43d47869057

Observation 12454229-ab24-45cb-92d3-ac8b8d3a7b10 · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.381159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:1f53f59edd4ee6e273938124e06afa9ea283d28ba004c0501b90de05047f16e0

Observation 899595d1-4933-4ae9-9dbc-dee115267ad7 · outbound

This paper cites Out of time: auto- mated lip sync in the wild.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Out of time: auto- mated lip sync in the wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:df713504a136929a362be79902b905d4d3ea2e5343f7aab8361b52243c82b2e2

Observation 54d5e6ca-0fb2-4633-ad75-d2ef26479d5b · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation VoxCeleb2: Deep Speaker Recognition

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.378912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:9bcf71087b2c7e4b67ba6725b5dbf4c3f11dfc4b93866009db1268d13ec25ac0

Observation 49b7f15e-7bf3-4ab7-8cb5-9e5594e07e6d · outbound

This paper cites Hallo2: Long-duration and high-resolution audio-driven portrait im- age animation.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Hallo2: Long-duration and high-resolution audio-driven portrait im- age animation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:333d2b9cf97aa2aed5af2c4ed8a07a84c43d49c8c68ec2700b2a7f5f1417d85d

Observation abc18808-f875-4add-af7c-f0872658280c · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Arcface: Additive angular margin loss for deep face recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:91aaccec909b5371f3cfc22ccda628fd736cead9c1cc3bdcf466780036f58453

Observation faa24f3c-407e-4809-b3b2-906bb85cf988 · outbound

This paper cites Generative adversarial networks.Commu- nications of the ACM, 63(11):139–144, 2020.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Generative adversarial networks.Commu- nications of the ACM, 63(11):139–144, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:0c5b69c1932fb4612e9627c6b4451f798f899b438cb3a5c7a90f82310c1e59c2

Observation 1732ee76-5d22-4a22-af68-2e08474b8a90 · outbound

This paper cites Generative adversarial nets.Advances in neural information processing systems, 27, 2014.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Generative adversarial nets.Advances in neural information processing systems, 27, 2014

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:022eb1b81afc16a41b6c6fc3f61a24a8b7a1a5888decd2120bd4b4711a828764

Observation d3510382-f6c7-4375-a2e5-5839afc77837 · outbound

This paper cites Generalized procrustes analysis.Psychome- trika, 40(1):33–51, 1975.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Generalized procrustes analysis.Psychome- trika, 40(1):33–51, 1975

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:2ff3f05c9a7533ead88e6a6924ffecd68fb7d38c8ff37dcf74720d4700faf040

Observation 91e3457b-b9d6-41af-917f-f184338a3d7b · outbound

This paper cites Animatediff: Animate your personalized text- to-image diffusion models without specific tuning.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Animatediff: Animate your personalized text- to-image diffusion models without specific tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:25d9e183b42757c82b544f543f19862300f751a06f1283b39275e76c8a4456ae

Observation ac4ba313-a31d-47e7-bbf2-21b0a158f264 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:d1904071d93aafc0ab8bd511e68c1f04c511513e60794eac1a526bf6bcfd23fc

Observation 9cb31306-7fc8-4cac-8e1d-f619742ab85f · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:56901fd2a31a28f10961fe0b0fd9ce4195f8c53c273ee84320f61e9b0fec5fa3

Observation 825201cf-afaa-4dc9-b08b-227329b464f9 · outbound

This paper cites Long short-term memory.Neural computation, 9(8):1735–1780, 1997.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Long short-term memory.Neural computation, 9(8):1735–1780, 1997

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:6b97838acafcfcdc3777ced12325dd6f6a86e07e9072b65627db7f672db2cf7b

Observation f1ec6658-00a4-4536-b232-d4c3684102b6 · outbound

This paper cites A multiresolution 3d morphable face model and fitting framework.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation A multiresolution 3d morphable face model and fitting framework

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:134e57d0ea6f21eff277513c6cf822d20e0e91287ad9b9f4fe3cb7e5d958f47b

Observation 86b0bb51-247f-4be6-b50b-23564fcd333a · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:ed3d08c974122f97f96ae2dac8b432b5467f10d62b8525d7a1f462f2b4279792

Observation 1324c729-c8a7-4d03-9a81-5639a3c0ff12 · outbound

This paper cites Sonic: Shifting focus to global au- dio perception in portrait animation.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Sonic: Shifting focus to global au- dio perception in portrait animation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:bfc22bc250e9f2d95086d6cc3ead67ed144645cb27ef953d7316ecdfe825d7a0

Observation c65b835d-a955-46ef-ab41-04cc87cc4e8a · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.364202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:640a569d5cf65ff120e8a8505f2f6d83cda7cee61cb40dc079ddaedd1af428f1

Observation c0de5e89-5438-4e5c-b229-753cc5c940f6 · outbound

This paper cites Modular primitives for high-performance differentiable rendering.ACM Transac- tions on Graphics (ToG), 39(6):1–14, 2020.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Modular primitives for high-performance differentiable rendering.ACM Transac- tions on Graphics (ToG), 39(6):1–14, 2020

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:0a52bccb872c5a35ea040f586708a8de7ac75336f206845deb23a2f39962697a

Observation 87a6e748-6fea-43b1-9e31-d89be474d3d0 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.366600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:e2bd4c437da5ac3f9e9bc5b05dc1148fb6bf81544e086fcd51d306b22f5d5804

Observation 958d26e1-4378-4c3c-b3a8-3964269a2059 · outbound

This paper cites Dpm-solver++: Fast solver for guided sam- pling of diffusion probabilistic models.Machine Intelligence Research, pages 1–22, 2025.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Dpm-solver++: Fast solver for guided sam- pling of diffusion probabilistic models.Machine Intelligence Research, pages 1–22, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:0ffbb9cc7a349ff8c00e66de27c1c89d6ea5c10418dd072ba5bde5584ef6e77f

Observation 1b359f18-ea41-495e-9e7c-22a1b1da809e · outbound

This paper cites Umap: Uniform manifold approximation and projection.Journal of Open Source Software, 3(29):861,.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Umap: Uniform manifold approximation and projection.Journal of Open Source Software, 3(29):861,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:42365d9e874ad2da3e1b2abf0062111f7fe577bd3dc4dd2ed58b3ff7a636aeac

Observation 0ec4de4c-67f6-4b5a-a0fd-369d1b190162 · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi- body human animation.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Echomimicv2: Towards striking, simplified, and semi- body human animation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:6c87d71db357d8a784499c82af710ed27072cc1b274d6c14f8b5f66ec76e57cb

Observation e18987d0-999a-4853-be91-4b091b81f5c7 · outbound

This paper cites Conditional Generative Adversarial Nets.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Conditional Generative Adversarial Nets

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.368861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:dd1cedf5cced0f288486150e64bdfc19feb51e94a1ca4ccdf34e13aca10eb162

Observation 14d73753-57cc-477f-855b-5ecb54c3358d · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation A lip sync expert is all you need for speech to lip generation in the wild

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:cc36a54b9a2b879605954ff32e287ec18a8298eebfbf729a50784bcb0e12af75

Observation de306a7a-bf1d-4535-89a3-a0ba5edf4ff5 · outbound

This paper cites Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.379708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:4c284650bd55f2e6f4d6dfcf57692f6e45da473abdf1f17f432f6c3c943c2aa1

Observation 2041faa7-1452-424f-aa01-70dabb446a85 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Learning transferable visual models from natural language supervi- sion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:eb717c6009b60bf6b814e97f6b6cdd1f36fd45f9a35741a937ee362883b6cf5b

Observation 7640b7a7-f78c-4850-9f31-ed7ebbfd1d37 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:8232a3650bc8f6a44202803acc0e1ec5e281edcc06d67821ef41c06654a68759

Observation a36dab9e-f5d3-4350-84ae-876b1a80dfb2 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation U- net: Convolutional networks for biomedical image segmen- tation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:39514f4f9a4449b90a4f2bed1d6e67ca4dc10e87e73a12db30b2536278af561c

Observation 3d50fb21-688d-412c-8616-a011a17c0a94 · outbound

This paper cites Long Short-Term Memory Based Recurrent Neural Network Architectures for Large Vocabulary Speech Recognition.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Long Short-Term Memory Based Recurrent Neural Network Architectures for Large Vocabulary Speech Recognition

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.371163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:af4fdfbd6d3f6ac796f9943de6fe6171eb55b3e5842bb26b241159664f5caeb0

Observation 3bf900ef-22cc-45c1-bca6-d3e078ac4fff · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:e60fe2929722c356b1d5cf8305819713648e4e252b8ea96eb74380dd84a1d918

Observation 0eda510d-6b2d-4a06-bc63-9f6a0f2a7c8c · outbound

This paper cites An analysis of variance test for normality.Biometrika, 52(3):591–611, 1965.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation An analysis of variance test for normality.Biometrika, 52(3):591–611, 1965

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:6e821170fc63a7e4ae7cba804b8192758c96e4acb11f58b2630922a63f6bbab4

Observation 43be2e79-f7fd-4218-a05e-a28cbd0d7956 · outbound

This paper cites Denoising Diffusion Implicit Models.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Denoising Diffusion Implicit Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.376258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:dcb02d79a3dfbd7f083df0867c51d0ef05c2ffd448db4122396305dfe0ccbeeb

Observation 7d64a7ae-ef43-4e05-9a3e-840dca4938c4 · outbound

This paper cites Seeing what you said: Talking face gen- eration guided by a lip reading expert.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Seeing what you said: Talking face gen- eration guided by a lip reading expert

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:84439892795c8864eea627365d779e144955d458ee83a38b7859ee99832251e7

Observation 84f34697-78db-4053-be6a-b312f29cfdf7 · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.382287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:1e70c5d62da20289e43fb5fafea4f96ab60374291d02a7cf61c72e049db40c2d

Observation 0d5d366b-99c1-4658-99fe-968a99dfc9e6 · outbound

This paper cites Audio2head: Audio-driven one-shot talking-head generation with natural head motion.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Audio2head: Audio-driven one-shot talking-head generation with natural head motion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:582b43597a2eab3f68091157cbf98e5b6d09be51388c1c8e21cab0ec406e1cc1

Observation 35c1c419-e08c-47c2-a075-f516d9e9a118 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.369443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:c3095ba990ee6eb273597eb1f969b63f362efde97ff964f368e64f6992cb2d2e

Observation 7954f15d-9872-4304-8dfb-54357796104d · outbound

This paper cites Vfhq: A high-quality dataset and bench- mark for video face super-resolution.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Vfhq: A high-quality dataset and bench- mark for video face super-resolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:8141cf4c1c0fb618a3d711de9c7220ce8d13b308a7294e1a1f3b12ca1cec3500

Observation a89fcbd1-5187-4cb5-abf3-3dc878572ccf · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.366896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:2058ac366c09a9ddc0a6b5b37a6a6a1a867af2832476af9f1cacce7e8a60351c

Observation 44aa7573-a938-4cfb-87c0-57509c3ed80d · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.374710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:17735acc0ff3909619bb0049450c3f829ca11cd2d9824932593f8d5fa0ce71c6

Observation ec0d4b40-2d66-4a74-aaec-5e2397ea623a · outbound

This paper cites Celebv-text: A large-scale facial text-video dataset.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Celebv-text: A large-scale facial text-video dataset

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:89176e4a6a187e0acf2c58d1a134246bdafebbde68a6b2d3ab3b081e7f212a21

Observation eb170dfb-26c2-47c9-a195-a6a3f357cf84 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Adding conditional control to text-to-image diffusion models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:bc1e31b9123c0236fcd4fe345cfa20de391c4b37c06992bea3b9e70a9ee4a12e

Observation c41f292c-c68d-4452-b17e-01258a1a3d45 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation The unreasonable effectiveness of deep features as a perceptual metric

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:8b64b1fc7401af8b2c31400089fdffa21d8680185ef4c9971d222b72e08c5fb1

Observation ddf0cf62-4118-40f6-9bc9-0bbcc1961ac1 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:d7eec0e063d043b622f9f43612fcf428c99f64a868117fbb70caf88a1345ad27

Observation 266d3d54-885e-401e-bd75-37703e4ad51a · outbound

This paper cites Musetalk: Real-time high quality lip synchronization with latent space inpainting.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Musetalk: Real-time high quality lip synchronization with latent space inpainting

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:bef206b31f324a1e2c40b3e7b79a8b9c380a065e3ea2ac3fad8786641e146414

Observation f3043fff-9c33-4714-b040-20261cb01415 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:a95368d5d720a5ffa6941d492280a395f9d910096cbe383abf3bf2ca122acfcd

Observation 8005b1dc-1148-4234-8fe3-1bb1d904a57d · outbound

This paper cites Hearing lips: Improving lip reading by dis- tilling speech recognizers.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Hearing lips: Improving lip reading by dis- tilling speech recognizers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:9212c8e43184be886ffdde9b2955ad2bface0832220e712299f2d9341e3a2d68

Observation a31039a5-3168-416c-8c8a-3157c46de69e · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.ACM Transactions On Graphics (TOG), 39(6):1–15, 2020.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Makelttalk: speaker-aware talking-head animation.ACM Transactions On Graphics (TOG), 39(6):1–15, 2020

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:4e47c644e0e72a605c472239db697a3af677f8f52c17a89ee75ff46db8118f36

Observation f7de0adb-7e20-4a56-8a25-a46e3d37bafe · outbound

This paper cites Celebv- hq: A large-scale video facial attributes dataset.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Celebv- hq: A large-scale video facial attributes dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:2b9b5c62bed89a5f0515fb7179bb9eccd5f18911c34cbf14f64304b7bbfc814b

Observation 67fb0735-a5c3-43e8-838e-69fb746e4977 · outbound

This paper cites an unresolved cited work.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:9945576b788b6cccc3faf39d8c1d22e532a4c9324c778ad19a00c0416e24bec8

Observation 440b30a9-b72f-4d3e-951c-2b8e31a6a603 · outbound

This paper cites All derivations and intermediate steps are included to ensure completeness and clarity.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation All derivations and intermediate steps are included to ensure completeness and clarity

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:b37c4d840bfd9ab50276d78ffbaa437f78c6c1b92ef899b64dc05a33a25a95c9

Observation 70ac0295-6604-4af3-a33f-3279e4b586ee · outbound

This paper cites AnimateDiff + IP-Adapter.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation AnimateDiff + IP-Adapter

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:21675d5973ad29e058b0d50f93bdfc487e53263b4d2b5d26d44ff64f8764642d

Observation 7a7085f9-c6d8-4fe3-9f1f-2d2d7113c691 · outbound

This paper cites To mitigate such risks, all generated videos in our study can be clearly marked as synthetic (Fig.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation To mitigate such risks, all generated videos in our study can be clearly marked as synthetic (Fig

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T07:50:56.947671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:c39a53cf1f1ebfa2892b099aa72eef12ea3fb3ef32e7f61e2046d3eca507f37a

Pith citing papers

No inbound Pith citation observations are available.