Pith. sign in

Paper Citation Record · LEDGER

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

As of 7 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2507.00324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00324 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:23:53.153699Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:23:45.554040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:26:22.148036Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact4
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec41e0e6-1bdb-490c-9c9d-924a5e4a3f48 · outbound

This paper cites This high quality of synthesized speech and the ability to distribute it through so- cial media platforms are giving rise to manipulated information in the digital ecosystem.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges This high quality of synthesized speech and the ability to distribute it through so- cial media platforms are giving rise to manipulated information in the digital ecosystem

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:59.161684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:45.377493Z digest=sha256:5a0b05ade72ef6f1d1daedfa0a33ca3141a081aee068eafa7caa69e80156af45

Observation f6badd89-b359-41b2-abf5-00ce9b129525 · outbound

This paper cites Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:45.554040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:45.554040Z digest=sha256:1549ce9739b2ecadec9a3ac39852704525069037e287eddf2072b023f2137cd0

Observation 34a76118-16bc-4390-b374-e4e12431aa93 · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:58.763526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:45.820734Z digest=sha256:946a2a1f1595ca0507c33d8d623f7cf937ec8a82225fc07485e3dd0c52b18a70

Observation ef2e435e-1bf9-447f-b5bf-7d23e788857f · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:58.124999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:46.517361Z digest=sha256:c11e9c26cf02fed0408e4c8f9dd5e1d94cc266ffeddec924feb527863fd82372

Observation 0cb07594-6495-4010-8f5f-3ddfbfd46961 · outbound

This paper cites It uses ffmpeg to download the best audio available in the W A V format, and resample it 16 kHz.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges It uses ffmpeg to download the best audio available in the W A V format, and resample it 16 kHz

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.593590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:45.966902Z digest=sha256:55fa82cdf7ab06f0e4067ee9de805345a0310c3dc611766373c669b905295126

Observation 1b89340b-2c8d-421c-b373-261544e841bf · outbound

This paper cites This step eliminates cross-talk and background speakers.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges This step eliminates cross-talk and background speakers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.415398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:46.148015Z digest=sha256:fa90ba4faf20749a7d101ea9ec06dad0c5b4bd36d32e9f3cc7c444dd71abb298

Observation 538b5387-e11c-4005-ad38-1da8854d1d6d · outbound

This paper cites We also experimented with the Google speech recognition api package and other commer- cial tools; however, they generated text with less accuracy and without proper punctuation.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges We also experimented with the Google speech recognition api package and other commer- cial tools; however, they generated text with less accuracy and without proper punctuation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.280280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:46.307871Z digest=sha256:edb92b3c869dff7cb53367970e2c9032005b2a1e0003205b0dfce196b1c1465b

Observation 4c6b9b41-3a0c-48ba-b550-de8cd6d9fb70 · outbound

This paper cites The Biden Deepfake Robocall Is Only the Beginning,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges The Biden Deepfake Robocall Is Only the Beginning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.327407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:48.600644Z digest=sha256:ae9372a840de45382eba94e525ec83b6f7f5fc627ebbf089c1f13bc5c21f222d

Observation a2d27f96-3a9b-4336-98aa-fb4c88df9e0b · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:57.977394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:46.655038Z digest=sha256:d9ec55f6ab8163307f0894556e33e2674bcfd2079e113578a7f83429611fbcc6

Observation 4e2b14f0-ca40-4164-8436-31dbb8330c30 · outbound

This paper cites Al- though early models produced robotic-sounding speech despite extensive training data, SSL-based methods significantly im- proved quality through large-scale pre-training.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Al- though early models produced robotic-sounding speech despite extensive training data, SSL-based methods significantly im- proved quality through large-scale pre-training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.786761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:46.800698Z digest=sha256:e8013bdf863fb384c36ad7f716ee77e67e499e732f3f74f28b0dfad2177390d2

Observation db3e056c-b005-4bf6-be5d-33f6bd23d504 · outbound

This paper cites We train StyleTTS2 [22] only using this approach.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges We train StyleTTS2 [22] only using this approach

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.612401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:46.967101Z digest=sha256:18c50d2d18be2cc4a845c55eb789757825b7ee66f9db3ef9fda51276351eec75

Observation 68170591-9c61-42c8-965d-3b9d503f8dbd · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:57.454318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:47.155081Z digest=sha256:a7ce2ab89061fed94334084ab42f45fcb35d1b90a02490f50afbfb9042cdbaac

Observation d3b137ea-e398-49a9-8790-0dbbb28585d5 · outbound

This paper cites This approach preserved linguistic coherence and improved phoneme alignment, leading to reduced noise and more ac- curate spectrogram generation.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges This approach preserved linguistic coherence and improved phoneme alignment, leading to reduced noise and more ac- curate spectrogram generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.305606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:47.266100Z digest=sha256:96eaa66eb00fe39bcbb8ddd88dcdfbaa74d44b3019ec25aadf442a5226e5c29b

Observation c11b0556-b91d-47d2-b848-ecabd713c468 · outbound

This paper cites an unresolved cited work.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:57.171255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:47.416615Z digest=sha256:abff0454be50e085307fa61fed957457bcc34a777bab3bb72be5c754bdd29daa

Observation 7963db6e-35d6-4ed1-8b50-8ea095bc9efe · outbound

This paper cites For subjective evaluation, we implemented a web-based listening test with 32 unique participants.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges For subjective evaluation, we implemented a web-based listening test with 32 unique participants

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:57.026356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:47.567771Z digest=sha256:b8826d3fb85f75c9f47b722466045867ea8c247938860735e8463acf2b25c1de

Observation 7c026615-d60c-4f4e-b76d-e32cb31a84dd · outbound

This paper cites Both datasets are derived from the VCTK corpus, which comprises high-quality speech recordings collected in a controlled laboratory environment.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Both datasets are derived from the VCTK corpus, which comprises high-quality speech recordings collected in a controlled laboratory environment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:58.981915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:45.684201Z digest=sha256:422f12b252c93cba81a1066ebd6c827d8999302606699831ce57d3b5e87ae790

Observation b33136b0-dcd8-4c67-b4de-efbb7d05da97 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges A Survey on Neural Speech Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.667605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.667605Z digest=sha256:9f34be99fbd2dc76dcb8cb0740635dff88f82558ba27e3307a9e42fa86725d1b

Observation 8341b7aa-fa69-4608-9424-d0d851b6f797 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.831204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.831204Z digest=sha256:3cb77ea2a066b3be1bec3666d53649171173ad9df4a598ff43a1da779e38e3f9

Observation 42a44163-1117-4c28-a7f1-e42f8587f90f · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.947919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.947919Z digest=sha256:b95838082762370071197f27952e1be7e8d128a58f71b8eeb19ebdcca20bb6c8

Observation 013f9f70-da0e-40f7-ad16-7ecb4c6027bb · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:48.037703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:48.037703Z digest=sha256:d1d2d0479704361c964ea7fe30773c866d0f9ca71a72b0d2c2d90819657676e1

Observation b4c32037-b3df-447c-a59e-bf9d1295845a · outbound

This paper cites Zse-vits: A zero-shot expressive voice cloning method based on vits,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Zse-vits: A zero-shot expressive voice cloning method based on vits,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.862210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:48.196907Z digest=sha256:40808a1a877c718b17d4535c6abb2c8c4ba79272b8f9f93404c319e55e029035

Observation f633927c-eb74-431a-9d15-b2a2f1f499a7 · outbound

This paper cites Global Risks Report 2024.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Global Risks Report 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.696718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:48.363816Z digest=sha256:cfeb28548ed0f470cb888e7475f7f28f3bf7ff564a73d41388cc9f53ea2d293c

Observation 0646b28d-b632-4307-8403-ed1575d8c012 · outbound

This paper cites Russian War Report: Hacked news program and deep- fake video spread false Zelenskyy claims,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Russian War Report: Hacked news program and deep- fake video spread false Zelenskyy claims,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.520757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:48.508515Z digest=sha256:9d9b72c82b726e8de943009c9a7b8da9c9cbd99e462ba469259f1aa2ca043d78

Observation e58cadbc-45e2-4815-8c6e-4857dd46ce3a · outbound

This paper cites Sadiq Khan says fake AI audio of him nearly led to serious disorder,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Sadiq Khan says fake AI audio of him nearly led to serious disorder,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:56.151468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:48.786228Z digest=sha256:ac69fb4bd80b2515b0e8e3400887aca1ad64850aeab0c81fa445fbd77066887d

Observation ec2c84af-dc38-4b29-94ad-3daabec6004d · outbound

This paper cites Asvspoof 2019: A large-scale public database of synthe- sized, converted and replayed speech,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Asvspoof 2019: A large-scale public database of synthe- sized, converted and replayed speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.981769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:49.011538Z digest=sha256:c70d16daabf6b6064929b2492f89af703e011f08203b07ab40e8e647adafe78e

Observation 9881999c-8420-449f-ab10-23fe9167081d · outbound

This paper cites ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild,

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:23:54.270034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:49.236797Z digest=sha256:df26cb989de209241e27dd69c8e922bcf43590b20d8aa832c9101b116ecd23ee

Observation 6d8b812b-97c3-4995-af76-0fcc54dedc96 · outbound

This paper cites Asvspoof 2021: accelerating progress in spoofed and deep- fake speech detection,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Asvspoof 2021: accelerating progress in spoofed and deep- fake speech detection,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.809406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:49.402291Z digest=sha256:00a4caa89a65d99d68015cc715d0e87c4d7ea8e891d9c66a5c8c51ed67ea4de0

Observation 895e6283-bfc3-4fe6-a159-f39d9181dde1 · outbound

This paper cites Asvspoof 5: crowd- sourced speech data, deepfakes, and adversarial attacks at scale,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Asvspoof 5: crowd- sourced speech data, deepfakes, and adversarial attacks at scale,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.597029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:49.475227Z digest=sha256:b53bb193b661650ba507f981585acee454fd1435014539d5c5ca2045c8a788f8

Observation 92954748-fe91-46f9-8a87-5e3657184c02 · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:49.572150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:49.572150Z digest=sha256:9a83f75a832a3308bb16b07828615153f3da1ab6b27db27b82dfb87261928c0e

Observation c35826a2-ee87-426b-bcb1-1ba8811c9704 · outbound

This paper cites Is audio spoof detection robust to laundering attacks?.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Is audio spoof detection robust to laundering attacks?

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.425051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:49.650417Z digest=sha256:d29de8ccc39ab6e0e154356aff9424d57cffaa529ab402117f0fc2183c686e43

Observation 98d7743b-a51d-4fd3-90cc-2ec3146a1361 · outbound

This paper cites Dfadd: The diffusion and flow- matching based audio deepfake dataset,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Dfadd: The diffusion and flow- matching based audio deepfake dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.291982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:49.755413Z digest=sha256:90d68e6704b285da413bb27ce190b8a51e82bca924e471d8adbd8e4a57079b06

Observation 202f6993-73f5-4b3c-8b6d-764782f87ef3 · outbound

This paper cites The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:23:53.920068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:49.860090Z digest=sha256:620d828a7bf26930bc5d475d832b756d82ef4be9d74218d64e04fd45b021d1ca

Observation 59d04091-d3ea-4307-94a5-356266d1489a · outbound

This paper cites Mlaad: The multi- language audio anti-spoofing dataset,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Mlaad: The multi- language audio anti-spoofing dataset,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:55.120571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:50.027320Z digest=sha256:41c3c78b96a831e0d413e0fda5b1f5da936caac927f26f801a0ab37c5bc1704a

Observation 016cbdbd-4287-4066-bc7e-543d76fab3e5 · outbound

This paper cites Does audio deepfake detection generalize?.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Does audio deepfake detection generalize?

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.952316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:50.161862Z digest=sha256:9d60d723dbf7a6b2dba27b62c16441ca1206adcb1dc33f0225e316a647cbb579

Observation cd5d01b6-2b31-4925-a019-4ed5511686c0 · outbound

This paper cites Spoofceleb: Speech deepfake detection and sasv in the wild,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Spoofceleb: Speech deepfake detection and sasv in the wild,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.824184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:50.340067Z digest=sha256:3288be7dca650e7da883f9a85ea2b236b1b757513d1d9892185a9b043ceb92e8

Observation 285295d0-de4a-4fb1-8bc4-32d9917fa255 · outbound

This paper cites V oxceleb: Large-scale speaker verification in the wild,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges V oxceleb: Large-scale speaker verification in the wild,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:50.513280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:50.513280Z digest=sha256:05d4f99931b965f19719966c768ec9eb085a8a5a5d3c406614ef2e4cd85d2017

Observation 3d92ef0a-b118-4ee8-92f0-503f3c09a367 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style dif- fusion and adversarial training with large speech language mod- els,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Styletts 2: Towards human-level text-to-speech through style dif- fusion and adversarial training with large speech language mod- els,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:50.723911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:50.723911Z digest=sha256:945649514242673c9d60fe816b08e9ffc7d7c0575b9a8d9e5e49296f18f1e1d8

Observation 6edb5495-89ea-4980-8374-716e9984d724 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:50.893642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:50.893642Z digest=sha256:a2cc6c86248e32367f8bafd50ddc87e40d0300c5f4c8cb4a19c722f8be4f1dcc

Observation a5447dd1-93e1-46c0-a6c2-e656c2f2e369 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.063898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.063898Z digest=sha256:acf72209352488894602cc5284a793e535d9ddfa0a4029681d701b0f1f6a0256

Observation 298bbffc-0c0e-4256-9e0a-b1091923898c · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.634215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:51.260561Z digest=sha256:f7a9308695d922ef1244185801fd740acbd2319b2f0e4eb69993723e88c0fcba

Observation be5a0f09-b588-4d3b-be6d-6f1fa0693359 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.398868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.398868Z digest=sha256:6c0b840d27233082e3724cdf5a62c9c6c75d627ca37e6dc8b4637d1532976152

Observation 02eeabf4-9992-46e8-b95d-3e9fb890efc7 · outbound

This paper cites SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:23:53.712615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:51.548122Z digest=sha256:81f616fc1ba0db7a5ac609bbe2327f037205d77b62db980016b296fd44cc553d

Observation 824d47ee-60fe-4859-b56e-c26291a5af5e · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.703103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.703103Z digest=sha256:3d48cf56e0ff092a18006ba600b909a783e4242ef4b8d1dd2f545f759b62b28d

Observation 9ae2daa9-6aba-4670-8a54-0512751c1724 · outbound

This paper cites Cosyvoice 2: Scalable streaming speech synthesis with large language models,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Cosyvoice 2: Scalable streaming speech synthesis with large language models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.927221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.927221Z digest=sha256:7b8dbdb89c54deda0f155f3685844abbc3a5530dfd400fc450739f97fa70a173

Observation a0716cfd-1be5-4404-b1eb-a91a8fb0ecfa · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.244690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.244690Z digest=sha256:874d0a2ee90f0583046cd47828aeac615b5b9405215533f83356530f01bae3b1

Observation 09561bf6-690f-4de4-9bec-0780c6c14941 · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Tacotron: Towards End-to-End Speech Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.397089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.397089Z digest=sha256:60b1aa90388eb9ecbdd4c41c36b91dc47c2e137c8a6966dbdcd57b94f8094542

Observation 089e98bd-4114-4964-bf2b-ecf40630e816 · outbound

This paper cites Glow-tts: A genera- tive flow for text-to-speech via monotonic alignment search,.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Glow-tts: A genera- tive flow for text-to-speech via monotonic alignment search,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:54.495771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:52.542844Z digest=sha256:a744d4e91e0106270ee45cf5c4b21ea85f7c211d1b1556e242d5ca4968936fc5

Observation 416732da-24e4-4e09-9e70-d2114e6776df · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.752336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.752336Z digest=sha256:7ba255143fda0abff76f168cd56ce72b437fb1ba3af6970456d7a4c1823c47f9

Observation c8af0324-78ea-4ecb-b44b-104148aad9b4 · outbound

This paper cites UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.972604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.972604Z digest=sha256:300b9780180b4c8f2d5eff0805cc492283a4fd5b13b09a270c89b4759e3461bb

Observation 1dd4bc21-92bd-4520-bd50-7f7ad9e468a6 · outbound

This paper cites Deep Learning Based Assessment of Synthetic Speech Naturalness.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Deep Learning Based Assessment of Synthetic Speech Naturalness

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:23:53.453912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:23:53.153699Z digest=sha256:cfc1e8e769108e32e6f7698f1b9f6760ea624a40670d4ae9a5edd67d0707d556

Observation 94fc393c-e3fe-4e56-996a-d6340c2ed578 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:52.083979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:52.083979Z digest=sha256:77aa3a37c717d7270dbf16860c17cc8230461c0586e39756aac42b0f9ed95497

Pith citing papers

Observation f6badd89-b359-41b2-abf5-00ce9b129525 · inbound

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges cites this paper.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:45.554040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:45.554040Z digest=sha256:1549ce9739b2ecadec9a3ac39852704525069037e287eddf2072b023f2137cd0

Observation aa481225-ba01-4ef9-86ae-5445907a2a84 · inbound

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection cites this paper.

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.151421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:25:27.137423Z digest=sha256:dcf804a2bc92ea261a7c5107eda9eb7af6f2265150c67f7995153e5383602fcb