Pith. sign in

Paper Citation Record · LEDGER

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2412.03430.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03430 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:28:08.105440Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:47:47.534590Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T11:47:52.259298Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2a840d3-675c-4e10-978c-5398166d3d32 · outbound

This paper cites A wavelet neural net- work conjunction model for groundwater level forecasting.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model A wavelet neural net- work conjunction model for groundwater level forecasting

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:09.083172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.796856Z digest=sha256:482a029d411752b9ff596e33076c191b577dd33d1d5411be5a8f85f29e0f1864

Observation 94c28e0e-3bef-474f-b25c-0e4e374ad584 · outbound

This paper cites an unresolved cited work.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:28:09.065353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.804403Z digest=sha256:f05ed8d2d82143fa3f2f95669592f90d882b401f2b6d2eac9febcd46eeadf6cf

Observation 7fcec83d-8225-44dd-9e13-a16c8290985d · outbound

This paper cites Speech enhancement in the STFT domain.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Speech enhancement in the STFT domain

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:09.047403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.811183Z digest=sha256:53d22069b99d07e50d2a4da3904cfd2dbd7b1c2f9ccb39dad643fea6a132b535

Observation 930cc9c0-4c95-4225-a63a-f4414b1fdbbd · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:09.031296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.817585Z digest=sha256:f70c6abd147798f8c4847c82cad5ea9cddb9ae025bb8055e2fbd9ad53e957524

Observation 4f77072e-d387-402c-a9f8-7de2d33cacef · outbound

This paper cites EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:07.823769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:07.823769Z digest=sha256:f940c2922ba196d9bd35f4e1d3d22aedca17a70887964e09c16b53c0b67644fe

Observation 7178f086-5cd4-4ca1-b208-e55a2d258e8a · outbound

This paper cites Audio surveillance: A systematic review.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Audio surveillance: A systematic review

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:09.016038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.831063Z digest=sha256:7c355a915c41bc117c6f6b413e8da8578d2f79d552c7098f00bf3448835e0eae

Observation 1ce5cb17-e64a-4fe8-a0be-c328b9684fbe · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:07.837803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:07.837803Z digest=sha256:3e70bd3732be3efab3f4e2f8e6cadf4d84cf82c494a8eed6d87cd407a01a3e66

Observation 459fc4d1-c1c5-435b-a121-6367282cb448 · outbound

This paper cites Diffusion models beat gans on image synthesis.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Diffusion models beat gans on image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:07.843423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:07.843423Z digest=sha256:aca7347b45f3f223056849735f8f9a8edc04c5aaa3b6e6de0a48cedde6443462

Observation e4fe5288-3f80-42fd-aadd-c5739330a0d3 · outbound

This paper cites Detection of the valvular split within the second heart sound using the reassigned smoothed pseudo wigner–ville distribution.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Detection of the valvular split within the second heart sound using the reassigned smoothed pseudo wigner–ville distribution

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.989675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.851167Z digest=sha256:c40b4448371f496a27fcb6cdb2fad0a025fd736a0becfbe631cf47c6cfeb8aba

Observation 05ca4531-a75f-4116-accc-c58e13dfa8f2 · outbound

This paper cites Wavelet multiresolution anal- ysis based speech emotion recognition system using 1d cnn lstm networks.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Wavelet multiresolution anal- ysis based speech emotion recognition system using 1d cnn lstm networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.973211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.857207Z digest=sha256:05ad563cb3fa4324b8c505764ccdb78d1687c61209d9b736e1b34ad847c8352f

Observation 357dddd9-0c18-4656-ac17-e3046727142f · outbound

This paper cites Sound quality evaluation of vehicle suspension shock absorber rattling noise based on the wigner–ville dis- tribution.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Sound quality evaluation of vehicle suspension shock absorber rattling noise based on the wigner–ville dis- tribution

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.955356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.865220Z digest=sha256:c8ffff8207d20e3c483488cdd64191345ae284d23232c77c994a845c399f9f67

Observation 81e3f114-36f8-4ba4-8bb3-3128c25ed44c · outbound

This paper cites Song2face: Synthesizing singing facial animation from audio.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Song2face: Synthesizing singing facial animation from audio

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.938736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.872878Z digest=sha256:5f62e60be6d0daca29b76cfab2d57c02dea64d963370675d45190e49758da1e4

Observation cf95e60e-54d2-4e75-a6fe-80824f2a2ad1 · outbound

This paper cites Audio-driven emotional video portraits.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Audio-driven emotional video portraits

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:07.881478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:07.881478Z digest=sha256:d19283ee2a4e3b5c69a70203315119bb33cd893ac1d3b61afccf4b40f0970385

Observation 156a8704-9c3a-48a1-bbfe-4640e059ccac · outbound

This paper cites Exploring spatial- temporal multi-frequency analysis for high-fidelity and temporal-consistency video prediction.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Exploring spatial- temporal multi-frequency analysis for high-fidelity and temporal-consistency video prediction

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.910156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.888363Z digest=sha256:d5eb82890d9a18d64909574a972f66e71214e57ac683a655e8c1631181b420d6

Observation 50864aed-67f9-4c80-8a75-e2262a578770 · outbound

This paper cites A novel deep wavelet convolutional neural network for actual ecg signal denoising.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model A novel deep wavelet convolutional neural network for actual ecg signal denoising

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.893580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.895439Z digest=sha256:369ca3fcf3d3d46422481911faa25a8fc20d20322229e5303967eca48951227f

Observation 6d2dcbbf-9cbe-4f77-9bca-4f37db4d9b3b · outbound

This paper cites Fre-GAN: Adversarial Frequency-consistent Audio Synthesis.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Fre-GAN: Adversarial Frequency-consistent Audio Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:07.902985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:07.902985Z digest=sha256:d403368b95b4e18c676875db0b9660c779c35cec7c288f05c57a819ada4885d9

Observation 9910adbd-86d3-4e2f-b66b-3aebd915277f · outbound

This paper cites Lipsync3d: Data-efficient learning of per- sonalized 3d talking faces from video using pose and light- ing normalization.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Lipsync3d: Data-efficient learning of per- sonalized 3d talking faces from video using pose and light- ing normalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:07.909756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:07.909756Z digest=sha256:ce422f3ccc8b135f86e921d4532b1cad3224706e3ad3b60630d2930af9f745f7

Observation 1f2d40c5-3527-4a4a-bab8-7f7120c844f8 · outbound

This paper cites Classification of audio signals using statis- tical features on time and wavelet transform domains.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Classification of audio signals using statis- tical features on time and wavelet transform domains

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.866189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.916420Z digest=sha256:ca07e06bbaced816e30c0b8772d1fdc98c682742a14bacabded8a225fcaad7a5

Observation 2c1c5fe7-6fe9-43a6-bc07-0c7ecf2a2e6d · outbound

This paper cites On the generalization properties of diffusion models.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model On the generalization properties of diffusion models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.850018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.922121Z digest=sha256:7ef1b76c1fed3f84de0cab4d024d8f9140de5d10d872b16f065b34cd24efb361

Observation 8f0945c6-00df-447c-be26-a61ba405cdd7 · outbound

This paper cites Expressive talking head generation with granular audio-visual control.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Expressive talking head generation with granular audio-visual control

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.834160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.928152Z digest=sha256:3107ff9cecc8398efca37e0a281019f3dd2cbe573c7d3889489875252f0d5481

Observation 7d1dab60-0420-45ec-80ff-56ab632b9ee5 · outbound

This paper cites Rolling bearing fault diagnosis based on stft-deep learning and sound signals.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Rolling bearing fault diagnosis based on stft-deep learning and sound signals

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.817657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.933749Z digest=sha256:089978c777a4cef76b2474eeee0cc1f72627f8aa3342ae9fdbf5f1f821ee9f2f

Observation 38beb26a-07e3-4689-87c3-f2dd236ffd04 · outbound

This paper cites Music- face: Music-driven expressive singing face synthesis.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Music- face: Music-driven expressive singing face synthesis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.799967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.939579Z digest=sha256:444b46f467c4955870497e7290ffce44650e18db520eaebd0e7a2415fe1cba9d

Observation 755cf35e-f517-4877-8c38-ae4af9dc09e7 · outbound

This paper cites The ryer- son audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model The ryer- son audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.781918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.944115Z digest=sha256:c418c19930f920cd9221cbdbfec1157b9116c85c8bb5393c7be98af49623fc4f

Observation dfc3292b-8baa-45a6-9cfa-349acaa1fa0e · outbound

This paper cites Formant estimation of speech and singing voice by combining wavelet with lpc and cepstrum techniques.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Formant estimation of speech and singing voice by combining wavelet with lpc and cepstrum techniques

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.762016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.949724Z digest=sha256:212d0c32f33362ddb1c92bc453938933218b26b4b70a93f89bc9b1abe7daa90f

Observation b35ec0c2-4ede-4bbe-87a0-f29b2d2be03c · outbound

This paper cites A no-reference im- age blur metric based on the cumulative probability of blur detection (cpbd).

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model A no-reference im- age blur metric based on the cumulative probability of blur detection (cpbd)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.744029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.955655Z digest=sha256:8b045617c62c3aabb647b7baebec71ef9d3ada08f834b2a2bd4eb94b0cba4707

Observation 5ee19b88-2cc9-4b50-8131-6062252ec535 · outbound

This paper cites Wnet: Audio-guided video object segmentation via wavelet-based cross-modal denoising networks.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Wnet: Audio-guided video object segmentation via wavelet-based cross-modal denoising networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.724215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.961429Z digest=sha256:4cf72081417a66ddb7f9702143889f1ec9e719ee332b61f2dfed88865b3c83fd

Observation b0b82e77-f6b7-4987-9d0e-425f4b958b7f · outbound

This paper cites Wavelet diffusion models are fast and scalable image generators.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Wavelet diffusion models are fast and scalable image generators

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.705211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.966113Z digest=sha256:82415e239c72dc7929d86eeab0ee9549751b68f7e17c844f484bc56eba36e524

Observation c5885b82-da4b-4d2f-b768-a7a3a2f61b82 · outbound

This paper cites Diagnostic Biomedical Sig- nal and Image Processing Applications with Deep Learning Methods.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Diagnostic Biomedical Sig- nal and Image Processing Applications with Deep Learning Methods

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.686071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.971052Z digest=sha256:39a1a25dd22f8d806cfdf8a2b25518a552deffa941b63dc3ac9fd4e8326e7fdc

Observation 9ae781a0-3a25-4c37-989d-7c1bfc398f29 · outbound

This paper cites Lung sound signal denoising us- ing discrete wavelet transform and artificial neural net- work.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Lung sound signal denoising us- ing discrete wavelet transform and artificial neural net- work

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.669687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.975899Z digest=sha256:d6aca7cab4153b9a0281d80c063bcc94bdcf0104a4b24cf8ae21b5c74397d674

Observation b31809e1-8be4-4af5-ac00-6a22b3615882 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model A lip sync expert is all you need for speech to lip generation in the wild

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.650120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.981145Z digest=sha256:b259234abf21414d08ffb6393a0f50ae70cc0f0d3a9dac4dc54f9fcac9a20335

Observation 2b6a15b4-fe1c-4a75-b609-56de6464e588 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.632234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.987320Z digest=sha256:da33e72653272e8bfa84a428e5e9b2c0317d2f78620881c854479b2a0fd749f0

Observation 55083178-73b8-4813-8eef-09bcec344ace · outbound

This paper cites Bailando: 3d dance generation by actor-critic gpt with choreographic memory.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Bailando: 3d dance generation by actor-critic gpt with choreographic memory

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.612453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:07.992406Z digest=sha256:5fd14b406ff23f78d1f61f4e0eccc7b612441c158da0a77ddbffeb467ade35a8

Observation 408969d4-33b3-49ca-849e-26d00a75d5f7 · outbound

This paper cites Everybody’s talkin’: Let me talk as you want.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Everybody’s talkin’: Let me talk as you want

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:07.997752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:07.997752Z digest=sha256:99d57f1510c30560f4ba93f0acdf622a8595b1252f5938e4cc9243ef0ab1c4e0

Observation 41490fa0-13b5-4922-92ab-dfdfcff36f88 · outbound

This paper cites Signal reconstruc- tion from stft magnitude: A state of the art.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Signal reconstruc- tion from stft magnitude: A state of the art

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.582608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.003790Z digest=sha256:6af10ad2de9b835a61606a6d8da7dcb45a36e079bf9e3e8b707477853707259f

Observation 3d7f366d-dafc-4bde-9298-db42b2963cb3 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.008908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.008908Z digest=sha256:56fe2814d9ebcc2fe82c1029a1b6b2fb6ea3612064223f2b3171298bd06f2b27

Observation f540ddec-4eec-4104-ae47-3538868451d7 · outbound

This paper cites Audio anal- ysis using the discrete wavelet transform.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Audio anal- ysis using the discrete wavelet transform

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.552175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.014915Z digest=sha256:4576e010e39263a42da7403d4f4cdac92e271c5fe61d4e9bffcd22abe319d992

Observation bc266121-c6dd-4492-a31e-9ca3ce087adb · outbound

This paper cites Fvd: A new metric for video generation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Fvd: A new metric for video generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.534783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.020142Z digest=sha256:268a377592224baeac35e8d0f39e353d4652814297359d480879bc1b41e7114b

Observation 092cd4b9-3e9a-48f8-8523-da2de40250cf · outbound

This paper cites Seeing what you said: Talking face gen- eration guided by a lip reading expert.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Seeing what you said: Talking face gen- eration guided by a lip reading expert

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.509939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.026139Z digest=sha256:0fef390dd69cf4894e1195407e5998097c8f688890f03f10bf34a7d178b97966

Observation e54efc39-3ba1-4a82-93b9-1890e0c57b89 · outbound

This paper cites Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.031619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.031619Z digest=sha256:aecaab122d9df893a908b09eda3d2bbc12acbcfa6a75c5c3754ed6d760d7d59a

Observation c16f206f-8548-4ab2-b3b7-41a838c01033 · outbound

This paper cites One- shot talking face generation from single-speaker audio-visual correlation learning.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model One- shot talking face generation from single-speaker audio-visual correlation learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.484692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.037092Z digest=sha256:dbea1aa0c533da764c6863022ea45e6dbc381ccb87710f74d8ddc3c8ffef9546

Observation 3ffdca56-b916-4914-94cf-d752134a3cb8 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Image quality assessment: from error visibility to structural similarity

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.465235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.042483Z digest=sha256:ff328e5e8c915ee8e02c8d5cb111b0819ebab3d940c7332f35717b8aa5462304

Observation 9834dd9a-38ab-4947-b585-c9801ceea50b · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.047726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.047726Z digest=sha256:8415acedf909d6c5f85652a65c0d2dd54f233674885bffe46ff0019bc40bd7f4

Observation 85ec3a90-b9b2-48a9-b13f-d49f99fb151a · outbound

This paper cites SingingHead: A Large-scale 4D Dataset for Singing Head Animation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model SingingHead: A Large-scale 4D Dataset for Singing Head Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.053712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.053712Z digest=sha256:469ac02661f36d0e100732234c48e80d389a0433f589ce41241355b497eb510c

Observation 13840db1-d430-42ca-afd3-3d1c4c990361 · outbound

This paper cites Facechain-imagineid: Freely crafting high- fidelity diverse talking faces from disentangled audio.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Facechain-imagineid: Freely crafting high- fidelity diverse talking faces from disentangled audio

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.443032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.059408Z digest=sha256:77a19ab3ee9e9b24f85b60368c93d6de546ee445a374c4567599f3da29dafdd5

Observation ae3d08e6-8f1a-4775-be60-6eb94c765fac · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.064506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.064506Z digest=sha256:78c4fe6b493fff286b96e11c6c8bcc7b5d9b36bec523d73b2b10880869e36aaf

Observation ebec744e-c540-4b68-9ae7-0fd5de71305f · outbound

This paper cites One-shot domain adaptation for face generation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model One-shot domain adaptation for face generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.420316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.069809Z digest=sha256:78caf27c8b1e8e8a056170a025f6cc1bfbeb92d70f1a293ba960685790a38ae7

Observation a08fd800-91f7-4138-af18-4d53073c8b08 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.400233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.074569Z digest=sha256:a817c0cbd6b66ca5e81ffb06ccea5bdc70795c4bb986f9f813be71229d943bd3

Observation d1d351a7-3f33-4d31-946a-2e7787663330 · outbound

This paper cites MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.080042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.080042Z digest=sha256:b841700d1cdc2b9962f8a0a7a7d682d20d886eeac24880bc91c6798ad93b9010

Observation 27a5750a-37f7-4539-824c-c316b8679a98 · outbound

This paper cites Identity- preserving talking face generation with landmark and ap- pearance priors.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Identity- preserving talking face generation with landmark and ap- pearance priors

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.086696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.086696Z digest=sha256:2e4ba2e0c24b5017f9aadebe6711e4c749f205a6247ea53b0626d3e32708102e

Observation 3c8c2218-00d4-4a81-b6e7-b0fc918adf45 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:28:08.092300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:28:08.092300Z digest=sha256:169e3ca8de19e299940e7a06984855a5a26a8cbd96868afc1f753eea355e437d

Observation e1993e18-87e0-41a3-a333-e7e0a786a237 · outbound

This paper cites Subj.”, “BGM.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Subj.”, “BGM

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:28:08.332580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.105440Z digest=sha256:7855e4d306dcb0aedd4e3aaceb0d96d9eaacbeabb0b93eb18316860cc954a2eb

Observation f1098c32-d98c-41b5-abe0-ac1ce88a85c7 · outbound

This paper cites an unresolved cited work.

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:28:08.350834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T22:28:08.099358Z digest=sha256:53f29395e8b9de04862cfaf92fe004e2bb086b733306b14e31eda7a73fdec5bc

Pith citing papers

Observation b1dae37c-e687-4d06-9f3e-5c93c144d6d0 · inbound

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation cites this paper.

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:47:52.264043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T11:47:47.534590Z digest=sha256:a810a1c7a35023e8a6a95065a41dda3503fbda04be8036211ef3bf61aa838706