Pith. sign in

Paper Citation Record · LEDGER

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2506.01020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01020 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:56:29.213609Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:17:06.334514Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:17:07.028665Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0158a4b-d980-41a5-91d1-22158098bec0 · outbound

This paper cites Fgp-gan: Fine- grained perception integrated generative adversarial network for expres- sive mandarin singing voice synthesis,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Fgp-gan: Fine- grained perception integrated generative adversarial network for expres- sive mandarin singing voice synthesis,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.321756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:25.740423Z digest=sha256:472bbccb9265c54ab72d01940683db460e1a1675f9f07a2b9dff0b4cba28270f

Observation 012f4664-7b3e-46d6-82b4-8a58d7abfdee · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.311532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:25.836502Z digest=sha256:17f1b47e0e057c9cfd51666a371b2c03eface17cda89966edf945b4730f2865d

Observation 01affd88-9573-44b8-b6be-f81a9d7be7a4 · outbound

This paper cites Multilingual speech-to-speech translation system for mobile consumer devices,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Multilingual speech-to-speech translation system for mobile consumer devices,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.298268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:25.901514Z digest=sha256:46bc24beda73ceb0769fc27d15b547cc8c20d29b0f33eb61c6f89e0d0a5eaaa8

Observation 2768803a-2e0b-43fd-bc37-077c1401840e · outbound

This paper cites Multi-speaker and multi-dialectal catalan tts models for video gaming,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Multi-speaker and multi-dialectal catalan tts models for video gaming,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.072786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:25.978244Z digest=sha256:f97c1c325a3e380d8c88dc70b1c8b6077a5c8ca1d832524791d170e07a839181

Observation 49379baa-ebad-4c6f-b6e1-b861fcc75433 · outbound

This paper cites Neural voice cloning with a few samples,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Neural voice cloning with a few samples,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:33.926486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.046650Z digest=sha256:cd4a49e7e07f8f57fd3949d08a0051197bab7bfe68a0f064208aea6e52f0e899

Observation 54695d52-1902-4b97-ae2b-d9705a364e98 · outbound

This paper cites AdaSpeech: Adaptive Text to Speech for Custom Voice.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation AdaSpeech: Adaptive Text to Speech for Custom Voice

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:26.122977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:26.122977Z digest=sha256:922cf68e8f2c99e014c296d63c37066ff2bea27c9618e36699a2c8c7ccc8ab48

Observation 59ca98f9-385f-4dcf-a332-4bd389d5b627 · outbound

This paper cites Rapid speaker adaptation in low resource text to speech systems using synthetic data and transfer learning,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Rapid speaker adaptation in low resource text to speech systems using synthetic data and transfer learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:33.633705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.217438Z digest=sha256:ecfc4ce422ba9899cd8c4138d94faf5d49b684ab59783c87d9265e2fdae3c5c9

Observation 8468afde-e559-4552-9826-2348ea795914 · outbound

This paper cites Quantum target recognition enhancement algorithm for uav consumer applications,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Quantum target recognition enhancement algorithm for uav consumer applications,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:33.316786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.312211Z digest=sha256:bc76cda7f9d547fa7a11d2b503c84bff483903da44ff71a6562396050092b891

Observation 38df971a-f665-4f11-899b-126dece443b1 · outbound

This paper cites Dual channel based speech enhancement using novelty filter for robust speech recognition in automobile environ- ment,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Dual channel based speech enhancement using novelty filter for robust speech recognition in automobile environ- ment,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:33.138938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.383725Z digest=sha256:dda28b23e5818ee87bdd700d6ba1ef26c1e3eeaa8566975348f92f3714c52ec4

Observation 2fa65a51-a547-4c8f-bf03-7e1e6a88388b · outbound

This paper cites Transfer learning from speaker verification to multispeaker text-to-speech synthesis,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Transfer learning from speaker verification to multispeaker text-to-speech synthesis,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:26.464176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:26.464176Z digest=sha256:12eeb1df77f8d8da3b36dcbc257b0ffbe1bf62b9a8e782c60fc390be7c2432cc

Observation 1f58c910-faa2-4914-877a-f92329873ac3 · outbound

This paper cites In- vestigating on incorporating pretrained and learnable speaker represen- tations for multi-speaker multi-style text-to-speech,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation In- vestigating on incorporating pretrained and learnable speaker represen- tations for multi-speaker multi-style text-to-speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:32.876798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.557662Z digest=sha256:df9258e808272cda3bb910b2a5746b32610769482eb64452caea28bcd5f08392

Observation b838cdba-a696-4f64-a935-3a47f3e38f7b · outbound

This paper cites Mrmi-tts: Multi-reference audios and mutual information driven zero-shot voice cloning,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Mrmi-tts: Multi-reference audios and mutual information driven zero-shot voice cloning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:32.700632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.653690Z digest=sha256:598a34afceb75c120164c5b01dcb59a9a37f0234582db1ec95fb03369c1772f7

Observation 659de971-e51a-43af-b5de-7068704a42dd · outbound

This paper cites Meta-stylespeech: Multi- speaker adaptive text-to-speech generation,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Meta-stylespeech: Multi- speaker adaptive text-to-speech generation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:26.748637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:26.748637Z digest=sha256:562bc2bda1d4ebb8e9038cb5dffd9fbbdc23dbb9243d9046607ac1d9cf35129f

Observation a7333ae0-e69d-42ea-8692-bb4f34e4d6db · outbound

This paper cites Acfusion: Infrared and visible image fusion based on self-attention and convolution with enhanced information extraction,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Acfusion: Infrared and visible image fusion based on self-attention and convolution with enhanced information extraction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:32.535501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.847043Z digest=sha256:72943884f1a32fa8574198ba037a3b5d2a10abf9706465c6615f2516c41778f1

Observation bfc1ebe1-7ada-408f-8c20-8b558ec68dc0 · outbound

This paper cites Multi-feature fusion-based convolutional neural networks for eeg epileptic seizure prediction in con- sumer internet of things,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Multi-feature fusion-based convolutional neural networks for eeg epileptic seizure prediction in con- sumer internet of things,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:32.381858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:26.968284Z digest=sha256:46c6169d37776ecc91b1e7fe8655cad34fa089bc88f8799a31b02ff767e1f531

Observation 803e9747-1526-438f-a9f4-77f2c1cefae2 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:32.095237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:27.058490Z digest=sha256:2fad72adc8c15c11fa4edc3563ca070b1df187005c1eee78489c68e1ce92a021

Observation 4b222e0b-a11f-495d-a694-d51626e6e81c · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:31.852770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:27.219960Z digest=sha256:684b6b558ccccd42b584e5b34ba7e54d8ff5725772c8e652af81a097b430c6c1

Observation 3277402f-9674-4962-9bc9-c028d761e197 · outbound

This paper cites OpenVoice: Versatile Instant Voice Cloning.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation OpenVoice: Versatile Instant Voice Cloning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:27.291540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:27.291540Z digest=sha256:fa4da3b983b260585c4f515aa13d8eb36b5a1e40e17fc6997f4fde1b40a63e19

Observation e6601101-719e-4fd0-bb82-044aab76acee · outbound

This paper cites Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:31.656852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:27.358554Z digest=sha256:2a931af47b7129ea84e9628156de40180084662f769c6bcf9295800ab9eddb35

Observation e2c2095d-4b6f-4572-9596-c8f1568d129e · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Arbitrary style transfer in real-time with adaptive instance normalization,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:27.427288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:27.427288Z digest=sha256:0fdbb2a9e82c4bd3fee9b546926fc7e5e966b2947724be10932d9664c2c03133

Observation f1dbafcf-29d1-40c3-9372-de7635042253 · outbound

This paper cites One-shot voice conversion by separating speaker and content representations with instance normalization,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation One-shot voice conversion by separating speaker and content representations with instance normalization,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:31.436116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:27.485568Z digest=sha256:153a6278ee1f0be8b411900a1ee57f3fab4401ea9894f8c1f717ded61b28d491

Observation 5cb7ec09-c5ca-49ed-b241-1ab8528c266e · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:27.552487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:27.552487Z digest=sha256:be17bc18fb817b167d3292862a8d6f57de97d466753198cf9a680b245fc9cdc0

Observation a0dbcb50-1cfc-4f74-a214-b24f2add4081 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Film: Visual reasoning with a general conditioning layer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:31.269928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:27.623544Z digest=sha256:e9addcad12c2dbf1411d44f32b435b995a0127eca612c3121dc032cdcdebb460

Observation 60af8530-33fc-4522-a7ec-bf67ae9d6ee0 · outbound

This paper cites Dynamic neural networks: A survey,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Dynamic neural networks: A survey,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:30.995377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:27.688923Z digest=sha256:1ef14149660fa45e867715edf0da04a644fd35b746b649a11fcc8d6df67d0947

Observation 6a1dd9d0-1ef8-4662-9e56-43e48451c3aa · outbound

This paper cites Imagenet classification with deep convolutional neural networks,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Imagenet classification with deep convolutional neural networks,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:27.765026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:27.765026Z digest=sha256:9146a54c46b7a3eb81566dd8d2a5e3a9585d066c596f9f431dbafcffd3714cfd

Observation d30c6895-2136-41a5-8890-7a0ff3cc0d63 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:27.862469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:27.862469Z digest=sha256:7e1698f2a5d4260257d0f52f8c0809b722d93e1a537cb915b8f03b9cf7e26d3c

Observation e3135193-0559-453a-8050-9f402ef5ddee · outbound

This paper cites Going deeper with convolutions,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Going deeper with convolutions,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:27.925595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:27.925595Z digest=sha256:76ceb766de52e520f5bc6f9a980fc0c18e705c46d0f6f224e60d1ea4f2f5f22a

Observation 40bd5bdf-f2ac-4109-9019-45ff90c43ca0 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:27.996901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:27.996901Z digest=sha256:ac7250d727c0c6c4c6ef31849ce9dea1803b3df519259ed440e003da94d9710d

Observation 50369e01-2a93-42d3-a6ab-1f923d4c781b · outbound

This paper cites Language mod- els are few-shot learners,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Language mod- els are few-shot learners,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:28.096207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:28.096207Z digest=sha256:ae5a71ce503d44114f75ffdfb6a46094fb23612dca1de3b0605d8986e9e5cd1b

Observation 074e8646-13bf-4a84-a571-c18d9456c3a8 · outbound

This paper cites Skipnet: Learning dynamic routing in convolutional networks,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Skipnet: Learning dynamic routing in convolutional networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:30.800021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.187728Z digest=sha256:ded724cb03fa0cae0a4600131a913f6b9f8376670a5df55d63a46d7f505dabee

Observation 7d68e7f9-a106-4dea-918f-514c0e860226 · outbound

This paper cites Multi-scale dense networks for resource efficient image classification,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Multi-scale dense networks for resource efficient image classification,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:30.563055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.271714Z digest=sha256:8bae7c4f95684cdb7061b77c40d2e9d29040f6e5cc80de06ced2f379523d13d5

Observation d0fa8c20-3024-4eab-b738-0f81546160a0 · outbound

This paper cites Reso- lution adaptive networks for efficient inference,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Reso- lution adaptive networks for efficient inference,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:30.405648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.366684Z digest=sha256:4cde571e85f8ea81bce8a320c3a998cac42c4a525909546fb397ac117536fedf

Observation 49e93ffd-1a53-41b8-86cd-7bcbe28befa5 · outbound

This paper cites Any-precision deep neural networks,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Any-precision deep neural networks,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:28.437716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:28.437716Z digest=sha256:d6cbb0310abffe4303c323cbb3e70073a1f3fa4ca2bfc0526bdda9f3fcb06ce1

Observation 52513412-4504-4517-b367-7580bd9b2935 · outbound

This paper cites Big/little deep neural network for ultra low power inference,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Big/little deep neural network for ultra low power inference,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:30.211682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.508645Z digest=sha256:55a1650e1d153b598b7a5c9fbbb77997e218e09d2c609aa544022706e4fbd807

Observation 3018709a-42d2-4868-b453-ac0de11c140c · outbound

This paper cites A convolutional neural network cascade for face detection,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation A convolutional neural network cascade for face detection,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:30.046990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.624844Z digest=sha256:ee7404e13a4b5adaa9077ad65039bb2dd20363a6ad771e875839e5f9c0dd2e56

Observation b3b7e980-8f8e-43ea-8e91-b6a2364c3e8f · outbound

This paper cites Changing Model Behavior at Test-Time Using Reinforcement Learning.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Changing Model Behavior at Test-Time Using Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:56:29.433210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.692242Z digest=sha256:b1b3055db8d040fe19296581101d884a1d9abe1b8e8d58bd1d468651553a9c47

Observation 1c21de8c-1bb8-47d3-96cc-bad795bef5f9 · outbound

This paper cites Dynamic deep neural networks: Optimizing accuracy-efficiency trade-offs by selective execution,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Dynamic deep neural networks: Optimizing accuracy-efficiency trade-offs by selective execution,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:29.892679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.749420Z digest=sha256:a2f9e406b8e6af0058ca5d4086d79759158a490706d8862d183c5cfe24e39728

Observation ea821dfb-05bd-4e37-ba10-fe69715eea86 · outbound

This paper cites Speech emotion recognition with co-attention based multi-level acoustic information,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Speech emotion recognition with co-attention based multi-level acoustic information,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:28.833569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:28.833569Z digest=sha256:f4478da8955eb622aa9f4e90e55d72379531f19e21590873b957b251425cedf8

Observation f851d4cb-ec79-40c1-abe9-fcbb5b0f818e · outbound

This paper cites Task-adaptive neural process for user cold-start recommendation,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Task-adaptive neural process for user cold-start recommendation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:29.730362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:28.893990Z digest=sha256:0b7cf2adfc9cfce49a6ba62364fadb9262972dc83d960920d68312043244d9bc

Observation 0830f5e2-6acb-47d6-9d10-d3866d1d8747 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:28.950767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:28.950767Z digest=sha256:d3ffb0e14dfc6231d2919fdf79f83b262bb0204abc54ffbb0ce44e7a2049c3e6

Observation 89eac614-048e-425e-a3e7-8f120e1986cf · outbound

This paper cites Melgan: Generative adversarial networks for conditional waveform synthesis,.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Melgan: Generative adversarial networks for conditional waveform synthesis,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:29.604171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:56:29.017430Z digest=sha256:0b2364b2a7fb2c2f910a9eefd55cec9fdd91970a20219ae5002137085f7c492e

Observation 9a7b921a-32ce-4f59-9765-7ed4dbf538f6 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.091198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.091198Z digest=sha256:a4a6f3aaf8be0615d6d6bc83bf43964844b1724ce37c26ed5b58ff8bd735188c

Observation 7d902b6a-eab8-4e7b-bdbf-ab210c7bf1d1 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.155283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.155283Z digest=sha256:d94ed292c701dd1e01b1124bed82174aaa26f27648b03cd7daa9685925d08941

Observation 571a9844-a593-4402-ac6d-3aebec911138 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.213609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.213609Z digest=sha256:424139809e652e50b2d4e2d0bde440d12b553511e147db84c8531cd1b34828e3

Pith citing papers

Observation 39a25de5-0fe5-4e27-b05d-6ca480cb2e41 · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:17:07.107479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T05:17:06.334514Z digest=sha256:118cb7fa1118adcfa722339dbbbaa478b1d6e0f75e520e38cea64112683b6c19