Pith. sign in

Paper Citation Record · LEDGER

Deepfake Detection of Singing Voices With Whisper Encodings

As of 21 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2501.18919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18919 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:01:29.676332Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7884cc9-152b-401b-85ea-c26c7b257bf8 · outbound

This paper cites ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale.

Deepfake Detection of Singing Voices With Whisper Encodings ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.624531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.624531Z digest=sha256:9c0f1045be0a4369c161ab030bcde99a689d8de0969847d2503a150f37dd34ee

Observation 17e5790d-b201-40ae-aa4b-df60e846e6de · outbound

This paper cites Vulnerability issues in automatic speaker verification (ASV) systems,.

Deepfake Detection of Singing Voices With Whisper Encodings Vulnerability issues in automatic speaker verification (ASV) systems,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.997777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.628731Z digest=sha256:9382b9b87f6fc1db1b7c473f13c4a70cf80b7611b71d31116bdace6ed2da8af8

Observation 6f0203a1-c935-4c7e-8830-3f8b620a150a · outbound

This paper cites Visinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,.

Deepfake Detection of Singing Voices With Whisper Encodings Visinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.987883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.632041Z digest=sha256:1543ccf70add4b1838fb7b2cfeea86e7e6eaa0b5bb2ce5ed2fe29d56fd84ee6d

Observation 0894bf25-0541-4c78-890f-6c3a5e419a5d · outbound

This paper cites Diffsinger: Singing voice synthesis via shallow diffusion mechanism,.

Deepfake Detection of Singing Voices With Whisper Encodings Diffsinger: Singing voice synthesis via shallow diffusion mechanism,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.977291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.635653Z digest=sha256:b9e0d90ff6ab571eb5ffb8ad2432a791c1b2a5dc4d5b91f54adf6d3c62449f8b

Observation cc59e713-0837-4127-887d-d165d73241cc · outbound

This paper cites Midi-voice: Expressive zero-shot singing voice synthesis via midi-driven priors,.

Deepfake Detection of Singing Voices With Whisper Encodings Midi-voice: Expressive zero-shot singing voice synthesis via midi-driven priors,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.968300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.639190Z digest=sha256:b72ea8277247ea6368787c2c2897d28685f15568dc79716ad2708954e34f121b

Observation 3e5aa955-04b0-4e02-8120-77dbf1698e28 · outbound

This paper cites Sintechsvs: A singing technique controllable singing voice synthesis system,.

Deepfake Detection of Singing Voices With Whisper Encodings Sintechsvs: A singing technique controllable singing voice synthesis system,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.959180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.642435Z digest=sha256:5b4ec2f436ef8a2e9a803e4c9c9dd714cc5164835101598e90e74aabb68bda52

Observation bfad7d8b-22f3-4776-87d5-c8bd9e6fd41f · outbound

This paper cites Singfake: Singing voice deepfake detection,.

Deepfake Detection of Singing Voices With Whisper Encodings Singfake: Singing voice deepfake detection,

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-09T22:01:29.859370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.645756Z digest=sha256:4a19c3699d6b60e663b7b1a44cd9fbc3e71e798e61465914388822d5825c59d9

Observation 13a450de-71bd-4c75-8848-df4c53b6df59 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Deepfake Detection of Singing Voices With Whisper Encodings Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.648510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.648510Z digest=sha256:52f897b7a646be31f229367214592b7b2c32461b10474c52f13d5e89806f9440

Observation eaa66595-6426-46d5-b8db-de5e8d06bdc5 · outbound

This paper cites Whisper-AT: Noise- robust automatic speech recognizers are also strong general audio event taggers,.

Deepfake Detection of Singing Voices With Whisper Encodings Whisper-AT: Noise- robust automatic speech recognizers are also strong general audio event taggers,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.943528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.651539Z digest=sha256:54993df1e5b8cefd21487c3b38b4c42931dfbbc02b0cd6440781661d34697930

Observation 2627ff73-870c-4990-97fb-dfb158136340 · outbound

This paper cites Attention Is All You Need.

Deepfake Detection of Singing Voices With Whisper Encodings Attention Is All You Need

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.654338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.654338Z digest=sha256:bcbba4d2ae616fd670953e5726c8d9d2851b77e9591c44d57ae4664018a3f440

Observation 1ba5303e-4e87-4166-beb2-8fc9a90a7da8 · outbound

This paper cites Invariant representations for noisy speech recognition,.

Deepfake Detection of Singing Voices With Whisper Encodings Invariant representations for noisy speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.933211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.657540Z digest=sha256:efd8d84cb6691b510c97daba1b9eaaeac443c52433aaf41b985fc0931f339aba

Observation 46a6252e-bcc1-4b11-8e32-979ce0d63482 · outbound

This paper cites Learning noise-invariant representations for robust speech recognition,.

Deepfake Detection of Singing Voices With Whisper Encodings Learning noise-invariant representations for robust speech recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.922339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.660236Z digest=sha256:ac425ee6d99ee3fe7b308a8e509f9a46d05a349e900cf3608ea49731639510da

Observation d0e43df8-43b4-4eeb-abcf-53661500b285 · outbound

This paper cites A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition,.

Deepfake Detection of Singing Voices With Whisper Encodings A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.911605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.663171Z digest=sha256:0623f8cab62080ff691f4996e4286382f3727a2c4a5f3bd0042ee4ff8bd91e32

Observation 6d92f38f-6561-4450-9781-ea6bb4005d2e · outbound

This paper cites Unsupervised learning of time–frequency patches as a noise-robust representation of speech,.

Deepfake Detection of Singing Voices With Whisper Encodings Unsupervised learning of time–frequency patches as a noise-robust representation of speech,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.900155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.666193Z digest=sha256:34ada73f6cf7e9481e84324eca2a721ea6991d86435c9ea8c2ed5df5363b0006

Observation c629cc56-827a-4de2-b089-638f5822f1a2 · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

Deepfake Detection of Singing Voices With Whisper Encodings Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:01:29.669479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:01:29.669479Z digest=sha256:2a76ca2caf1921cbec47e2b9c606d7adfb1cc5523059134908ca041285d414b0

Observation e410df5b-a605-4a62-8673-6709c5502b9a · outbound

This paper cites Computationally-efficient voice activity detection based on deep neural networks,.

Deepfake Detection of Singing Voices With Whisper Encodings Computationally-efficient voice activity detection based on deep neural networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.890043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.673284Z digest=sha256:4cf30cc526df9c11fe825907dc1fa96d90240e700ab20145ea22b19a95996769

Observation d9913635-5c40-472f-98cf-451174f351c7 · outbound

This paper cites Pyannote. audio: neu- ral building blocks for speaker diarization,.

Deepfake Detection of Singing Voices With Whisper Encodings Pyannote. audio: neu- ral building blocks for speaker diarization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:01:29.880003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-09T22:01:29.676332Z digest=sha256:a271af59ab17d081d0ed2089a4c9a63949062e2c1b916c977be2c25e7a1ffc95

Pith citing papers

No inbound Pith citation observations are available.