Pith. sign in

Paper Citation Record · LEDGER

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation

As of 12 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2412.20048.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20048 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:41:01.512970Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy69
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c23cb365-eee1-4a1b-bb8a-07ca5bd05705 · outbound

This paper cites The amazing benefits of being bilingual,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The amazing benefits of being bilingual,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.269476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.917403Z digest=sha256:58e3c02e44402994f07830de626bfa04a5f67974d3a4ec92af4d1a720b070af1

Observation 5d15a981-3d6f-4e69-a3a7-70b09b9846fd · outbound

This paper cites Disentangled representation learning for multilingual speaker recogni- tion,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Disentangled representation learning for multilingual speaker recogni- tion,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.242873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.923994Z digest=sha256:15af93e168e4d55c756077d09e1aed55bcc4d3b850aa1eebf60de42626e49003

Observation 8216c3ff-4d1a-4620-a9b8-fd20ab69d79a · outbound

This paper cites Crosslingual and multilingual speech recognition based on the speech manifold,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Crosslingual and multilingual speech recognition based on the speech manifold,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.220288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.932355Z digest=sha256:7e5d5af064a0babfc1dbca2b334b30e3f734574737f54190b57f2d781d7cc091

Observation 2b16526b-174f-48db-b2f1-7ea9f1eaa3fc · outbound

This paper cites Distilling a pretrained language model to a multilingual asr model,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Distilling a pretrained language model to a multilingual asr model,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.199694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.938640Z digest=sha256:1aefd9b04513562a5067737a4c09c3abe6daf3ee8cc26a5099f8c4692e628d38

Observation 2b79a606-c8fd-4e75-98e3-43b4e84b1814 · outbound

This paper cites Joint ASR and language identification using RNN-T: An efficient approach to dynamic language switching,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Joint ASR and language identification using RNN-T: An efficient approach to dynamic language switching,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.174204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.946711Z digest=sha256:b3baccbf24a1c7e4d350aeddd4e26fb38fe84a5c8f870eb4b97b208603e9fd12

Observation eed4ff27-cd21-4963-ba27-574002257665 · outbound

This paper cites Joint unsupervised and supervised learning for context-aware language iden- tification,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Joint unsupervised and supervised learning for context-aware language iden- tification,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.151563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.953481Z digest=sha256:45c68988f0ecb829c1d76060014eb6bc3243f4745178246ef43113c284ed2c51

Observation 349dcebc-a7fa-433e-97cd-28dc3958c96b · outbound

This paper cites Fastpitch: Parallel text-to-speech with pitch prediction,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fastpitch: Parallel text-to-speech with pitch prediction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.131618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.961403Z digest=sha256:f6c3bf8be0bdfacc24082760159785a8332eaee5a8fb0c80c91f6bf8b1ab4d85

Observation bdb4b98c-6964-4637-9bcc-d30e1d0d7339 · outbound

This paper cites Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.107774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.969012Z digest=sha256:4187ea733c42f793b03f1d3688ec6ef5bbff3ac215613315405d899a6ff5b691

Observation a604f3c6-c4ef-4ebd-896c-415750cedae8 · outbound

This paper cites EfficientTTS 2: Variational end-to-end text-to-speech synthesis and voice conversion,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation EfficientTTS 2: Variational end-to-end text-to-speech synthesis and voice conversion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.085372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.976039Z digest=sha256:ef832b0211294511126bfc31701d7945efbc9305c98ada435b9b012a785e4611

Observation 05260672-3867-48f6-89a2-bf5c14eae1b5 · outbound

This paper cites Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.059788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.982674Z digest=sha256:1cfc68bd841009eb7968d0040d7b9d61bee5405e6b02264cab779b7e32106c97

Observation a11dc593-9a18-48e6-9cf7-05aff151bfdd · outbound

This paper cites Improve cross-lingual text-to- speech synthesis on monolingual corpora with pitch contour informa- tion,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Improve cross-lingual text-to- speech synthesis on monolingual corpora with pitch contour informa- tion,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.039156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.989110Z digest=sha256:9e3c360eaac0843333955f289dc537b6079d1ff9d0807449687e4d02048a2b70

Observation 8302a362-5913-4d46-8a29-daf86d124842 · outbound

This paper cites Language-agnostic meta-learning for low-resource text-to-speech with articulatory features,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Language-agnostic meta-learning for low-resource text-to-speech with articulatory features,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.015542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:00.994932Z digest=sha256:f88fa5c17d3ae8108894cde937b86d956c74fade4a977201f744d539b525d71b

Observation af47961d-8ebb-45e2-a4ea-6e4228fde64b · outbound

This paper cites Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.995913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.004849Z digest=sha256:d4448fb5d87aeabf1c3998ae7ce3c7e98063c1550fd639e18dadf9e02cc78273

Observation 8d41d9eb-99b6-4812-a974-2a26a56b0778 · outbound

This paper cites Disentan- gled speaker and language representations using mutual information minimization and domain adaptation for cross-lingual TTS,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Disentan- gled speaker and language representations using mutual information minimization and domain adaptation for cross-lingual TTS,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.973107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.013367Z digest=sha256:040279712ec05f169674a9ae01f55caea3d2d7993345620afb00acc0d13bd331

Observation c6f464f7-6ff1-4529-8ab7-2d9d93c1e1a1 · outbound

This paper cites GenerTTS: Pronunciation disentanglement for timbre and style generalization in cross-lingual text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation GenerTTS: Pronunciation disentanglement for timbre and style generalization in cross-lingual text-to-speech,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.954486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.020861Z digest=sha256:6bab2f179ad2ea0791f4de6bf49dfa458e4d443a9c7d3d052f2870a96ec7f7de

Observation cd8acda7-6487-45a4-82f8-9eba3e9e4432 · outbound

This paper cites DSE-TTS: Dual speaker embedding for cross-lingual text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation DSE-TTS: Dual speaker embedding for cross-lingual text-to-speech,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.929948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.026896Z digest=sha256:023b487920c13f0f5a0113d3615272412b0d02390e70ed44b1272c90db3e0c51

Observation 93e48c2c-23e7-41ff-8d96-4703c38e8ec7 · outbound

This paper cites ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:41:01.687794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.037296Z digest=sha256:5d84f36c56cceeb274c0016e2c9731ad80f9588fce44f596c05a38294a16b8d8

Observation 24b0fff1-493f-4e57-bbd5-7cbf40f7e594 · outbound

This paper cites Unit selection in a concatenative speech synthesis system using a large speech database,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Unit selection in a concatenative speech synthesis system using a large speech database,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.907351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.048078Z digest=sha256:8ddf63a4229dd946291d5f8ddbbba069c6a760ec8c378c185b9eb4e76e1f3f71

Observation 91598463-fd05-41fc-9cb8-7aa59d3722e9 · outbound

This paper cites Statistical parametric speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Statistical parametric speech synthesis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.884267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.054553Z digest=sha256:4b2bff779027e7398ea0ec4bb88673d0b1da4b738b22acefdc35b0bef5d27d31

Observation 6d1d7ac2-cff8-4500-a8f4-be36d07032a7 · outbound

This paper cites Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.864249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.061028Z digest=sha256:9ebd0fed809b56c8b821463277fa26f638c9c77e79d3a8a41cd0c990fa32a594

Observation 30070c5f-f1c9-4403-9a58-cf8fa98d6ec9 · outbound

This paper cites Harmonic-net: Fundamental frequency and speech rate controllable fast neural vocoder,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Harmonic-net: Fundamental frequency and speech rate controllable fast neural vocoder,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.845442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.067703Z digest=sha256:276e26fda03f1d0adcdf0de91de9153f4f940e8ccc6d0b360eabe838eef8ff11

Observation 17f12e01-95b9-44b2-8e78-61efbc5e51c5 · outbound

This paper cites Fregrad: Lightweight and fast frequency-aware diffusion vocoder,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fregrad: Lightweight and fast frequency-aware diffusion vocoder,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.824269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.074760Z digest=sha256:97a1c96614edc3aee12c463312e6bc0cf1de6993244ecfa604726354f5cdb4de

Observation 5b22e679-ee23-42cb-bb09-a32f3753b78f · outbound

This paper cites TriniTTS: Pitch-controllable end-to-end TTS without external aligner.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation TriniTTS: Pitch-controllable end-to-end TTS without external aligner

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.800273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.082362Z digest=sha256:1388abe322bdd808c8d77ed2332a2d0f213d9c96d0d0a3cede4d44810118cb9f

Observation 86d5654b-fd8c-4a89-b268-c97125f338ff · outbound

This paper cites Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.782213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.087731Z digest=sha256:894a67035f9bc07260fada45df33e7af3937393250ac36e430ea917e6e4d3bc1

Observation 614d5c98-5ce2-4392-8be2-08ff8d25809c · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation WaveNet: A Generative Model for Raw Audio

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.095472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.095472Z digest=sha256:08845f0675e1f0613be2aed6c210a722198fa199991555b485097f9961ca6c2d

Observation 4bceb40e-42d5-4061-a722-705378184bec · outbound

This paper cites Deep voice: Real-time neural text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Deep voice: Real-time neural text-to-speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.763678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.105039Z digest=sha256:cc37088fa815c92b4981ed0c1c5d295650760d8005103cb69e90e0df2bf1d56d

Observation c9318edb-cb90-4bc6-b523-4e1e1278648e · outbound

This paper cites Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.740586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.116179Z digest=sha256:afa41cc260e8f3951aa070a46e37e7915f7b3037b80166380abaea0ea3ab1d4e

Observation 57adf011-d731-4c2d-8505-c9b089076f6b · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.722246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.123118Z digest=sha256:0e644666845e78ded718c74ced8eb306153ca5ee56eac5af25750a4807752e48

Observation 199ac34f-19af-4db2-a874-85977630b03b · outbound

This paper cites Matcha- TTS: A fast TTS architecture with conditional flow matching,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Matcha- TTS: A fast TTS architecture with conditional flow matching,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.703668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.128890Z digest=sha256:f9040fdfb0aa7f4b6448ef3c808a1a898491b9b1df50454e2e5447ca4483fadd

Observation bb6e89df-9fc1-48d8-ac11-f7b31e9a642c · outbound

This paper cites Multispeech: Multi-speaker text to speech with transformer,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Multispeech: Multi-speaker text to speech with transformer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.685786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.135253Z digest=sha256:0d17ea46bd9099d661deff631d278e20bfbca53ee36f15b64245616a2d229d4e

Observation a10192b2-9bf5-40e9-b767-bef9dbf8dbc8 · outbound

This paper cites Lightspeech: Lightweight and fast text to speech with neural architecture search,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Lightspeech: Lightweight and fast text to speech with neural architecture search,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.666778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.141137Z digest=sha256:e339bd4fce6d1be724521e763e86a12600c60483e613b2a0010b89592e7d5093

Observation 92c323f5-604a-4040-ab6b-ff61414a0bcd · outbound

This paper cites Phonological features for 0-shot multilingual speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Phonological features for 0-shot multilingual speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.646583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.147694Z digest=sha256:5c7f865bd2d1019c78009f6a77a716729c90fe9943e3249a32aa5277f0474e5a

Observation 85d6b405-fd35-4a37-adff-08282372a955 · outbound

This paper cites Text-inductive graphone-based language adaptation for low-resource speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Text-inductive graphone-based language adaptation for low-resource speech synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.626269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.153949Z digest=sha256:ea70c4e014c49865938de3b4f0f24795f218f0756597864379c0befaa13dd7b1

Observation 7e7c046c-2cf5-42ac-b3f2-7cfea3e0e140 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.601319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.163504Z digest=sha256:a7109ef4489f842ae99ace53a9702c1bcd7387702e98cdc635d64f948fb93230

Observation 7af7583e-39dd-4fd3-acf9-644ea4cd4064 · outbound

This paper cites Domain-adversarial training of neural networks,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Domain-adversarial training of neural networks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.579706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.171886Z digest=sha256:7bef3084aa34e597f1f25a9eb76e4c052b5ab863af063fe7ff78d7dab247b1e9

Observation 785c3b1b-eafd-4333-8647-7bc47c845abd · outbound

This paper cites Learning disentangled representations via mutual information estimation,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Learning disentangled representations via mutual information estimation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.545882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.180485Z digest=sha256:3d09400606fedfe2e51388b5d88ef4b77a50893d49b0c6af7add71f46120bf73

Observation 36afed2e-9437-4b04-a4dc-48e4d2cfb84e · outbound

This paper cites SANE-TTS: Stable and natural end-to-end multilingual text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation SANE-TTS: Stable and natural end-to-end multilingual text-to-speech,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.506996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.190551Z digest=sha256:1a53b4577a3d44efde01577bd55d06962a18340aa7d283c6f2d5aed86a343ec8

Observation 97c9a134-6f77-4bdc-b621-9e75747d8f33 · outbound

This paper cites Crossspeech: Speaker-independent acoustic representation for cross- lingual speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Crossspeech: Speaker-independent acoustic representation for cross- lingual speech synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.489429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.198083Z digest=sha256:1e61d29bb29f1314e7efaf7bf32a5adb554ec64e731366352c6254eeb911352d

Observation a47dc9af-30fa-4784-a4a9-756411a563de · outbound

This paper cites Invariant risk minimization,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Invariant risk minimization,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.471418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.206074Z digest=sha256:09944421c6de25f70e9c810b88e11a36df6271e0a59b6fbe472c789781498a01

Observation dc79d5a0-1db8-437b-8f03-c3ec0ddb7f7d · outbound

This paper cites Domain generalization with mixstyle,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Domain generalization with mixstyle,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.451892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.213292Z digest=sha256:3d7734a5e8cfe9e84289bab7f1ecdcd0c104c2daa6d8a4529971d280c857919a

Observation b1a2639b-98e0-4e28-95e7-d75dd4702376 · outbound

This paper cites One TTS alignment to rule them all,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation One TTS alignment to rule them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.427354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.232108Z digest=sha256:94c5ae774c9ea744dea54eafa6bf007cc458509338fd83256b998dae6b32cb8e

Observation d0c5ad9d-09b4-4dde-967d-f6f1541e0898 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Conformer: Convolution-augmented transformer for speech recognition,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.407581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.240834Z digest=sha256:72ff01e047a7d40b2af37c56cd896df806e43a756a4a25e67ba0d8e9d939896e

Observation 3ba2697f-2546-41d8-ae46-6659e9c2772f · outbound

This paper cites Feature-critic networks for heterogeneous domain generalization,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Feature-critic networks for heterogeneous domain generalization,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.378049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.258299Z digest=sha256:28313bc2895a6fae569fa88997d007045463c7eed71b2e6432c39d56f85c4f39

Observation a66522b6-11f1-4ab5-b738-483249044bb8 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.349918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.267937Z digest=sha256:05c59ef9f4002f2baddb93b977b79f5aa532bff29799329b8c5abec61f80a11d

Observation c4ece250-45d7-4872-beb6-ebd9ae1e1680 · outbound

This paper cites PV AE-TTS: Adaptive text-to-speech via progressive style adaptation,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation PV AE-TTS: Adaptive text-to-speech via progressive style adaptation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.327619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.275765Z digest=sha256:17609789dad9a65a37d427c89d2868622607e5fe25ec6bdf100769e3355e329e

Observation fcfee05f-68af-4a65-998e-6b01c8c3a593 · outbound

This paper cites Style normalization and restitution for domain generalization and adaptation,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Style normalization and restitution for domain generalization and adaptation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.302559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.282956Z digest=sha256:db2b5e62cde3ad50611faa43c10059a9ca7c56cdd6fa808d893dabc855074ff4

Observation 602d7743-9ae7-4ad6-b30e-d0c9a3de8ea6 · outbound

This paper cites Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.277338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.290550Z digest=sha256:e4508911b928aa555ae3832b901e8b6b3287def155f7c08d6a7af6194442b302

Observation 01b8547f-fdc7-467c-a438-a8e4011503cd · outbound

This paper cites Diffprosody: Diffusion-based latent prosody generation for expressive speech synthesis with prosody conditional adversarial training,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Diffprosody: Diffusion-based latent prosody generation for expressive speech synthesis with prosody conditional adversarial training,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.252247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.296819Z digest=sha256:6353e646d46224410516116402d02d409c046f5100f958f3cf66ab4185429ad3

Observation 26cfbfa1-825f-4daa-a197-d4e51f33f2c6 · outbound

This paper cites pYIN: A fundamental frequency estimator using probabilistic threshold distributions,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation pYIN: A fundamental frequency estimator using probabilistic threshold distributions,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.228142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.302355Z digest=sha256:17ebdc6c75354c7920e2abe03b62463275edc3cd39711bebca106258ebc9b6f9

Observation f75cf353-7153-4d34-88dc-12db8c1bde0a · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representations,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Neural analysis and synthesis: Reconstructing speech from self-supervised representations,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.209008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.310006Z digest=sha256:1bbd6dd0c56af34422faab49fc51913b997a119e5e3a892b8b2d1ca4246b7fbe

Observation 760857da-4f2d-4ad3-b8d5-19ec24273660 · outbound

This paper cites Exploring wav2vec 2.0 on speaker verification and language identification,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Exploring wav2vec 2.0 on speaker verification and language identification,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.190483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.320135Z digest=sha256:4f748011f3bee1ccbf201e461eb128f74dc9577cc1d81447c76413c3f0905179

Observation 71799840-ac50-45c1-893c-407a7d1fd3a2 · outbound

This paper cites Let there be sound: Reconstructing high quality speech from silent videos,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Let there be sound: Reconstructing high quality speech from silent videos,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.169374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.326985Z digest=sha256:90ab87f19f602d5176cbbfd5d6a5decb176357c1f1d9897a1a0f2f2d7a344ab5

Observation 7cf0cd41-4470-4132-a999-417d4e9ebbae · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Scaling speech technology to 1,000+ languages,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.141990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.334513Z digest=sha256:a19b9625f951330a04d227d6c5dd0de460dccc9b6fc12c14c7bed43d9c8a8327

Observation d3ef05fd-f846-4cc0-a9b2-c146ec400b44 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.123668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.341458Z digest=sha256:f21584f97bb7101fe2ac8ea148f0569071268e0eb41f0effe8c560abb491c20c

Observation 0f4223bf-e28b-4eee-8e8c-47859fd98892 · outbound

This paper cites Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.101079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.347868Z digest=sha256:606ce93161a6e0fb41bcdbb5516fd19a5968924b326ea86cee966bbd759e7a90

Observation acaa876b-4f89-4927-9574-4d563ef18ac5 · outbound

This paper cites The LJ speech dataset,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The LJ speech dataset,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.082325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.356498Z digest=sha256:7a9311decb3a3019b1997e493742057cc09dc167ee88a265064bc39684e13c8f

Observation a57298f3-6db4-4c26-bd39-344a6ea1c185 · outbound

This paper cites CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.363187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.363187Z digest=sha256:c02606d966239167835a7670a8cb59a174517aa0ec6003def0f3d532649fa6b4

Observation f5c4ab24-b74c-4038-a6a3-f30616549ad2 · outbound

This paper cites The BIAOBEI dataset,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The BIAOBEI dataset,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.060733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.373919Z digest=sha256:875028541100a712f29efe9ead8a1135a8fe39e9f53ed550c4de1b111f7590c3

Observation 0e24fc1d-ff54-4eac-9a95-82c80d341ab7 · outbound

This paper cites AISHELL-3: A multi- speaker Mandarin TTS corpus,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation AISHELL-3: A multi- speaker Mandarin TTS corpus,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.040412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.384569Z digest=sha256:62c558b98d40df6deb53ff696bd50cc842af5997d61c3533d45da2f30238b398

Observation db8ce6fb-d726-4e9a-8b6f-d6c4797a14b3 · outbound

This paper cites CSS10: A collection of single speaker speech datasets for 10 languages,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation CSS10: A collection of single speaker speech datasets for 10 languages,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.022899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.393788Z digest=sha256:07c59a1776761212e7a511bebf1d4865a25be2109c1684e4564c64e209ac3c14

Observation b73612f4-d09b-4e05-8e46-eb109fdb70ce · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.404222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.404222Z digest=sha256:cbb93f21835cb26fb24e590fc557a2e4d30f1faee144538961843cdf8b22dd4d

Observation 4a2e9241-8696-4630-ae8c-87cb57ecb2ae · outbound

This paper cites Multi-speaker TTS data,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Multi-speaker TTS data,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.001771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.413718Z digest=sha256:cda7f6870211a4043704d529e8bafae9b31f9a637813d868b9fcdbe28285444d

Observation 3972753a-281c-48cd-8d64-58c05baed105 · outbound

This paper cites Phonemizer: Text to phones transcription for multiple languages in python,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Phonemizer: Text to phones transcription for multiple languages in python,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.981863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.419782Z digest=sha256:c7cf4c6eaf7dd30eec095bda719543e8c8d9f29c3934a5ba7a3fd8be3a85e243

Observation 83f93c23-0ea3-49da-bc02-620346aed726 · outbound

This paper cites NANSY++: Unified voice synthesis with neural analysis and synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation NANSY++: Unified voice synthesis with neural analysis and synthesis,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.958515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.426365Z digest=sha256:38e1f94a6d67fb0de8d6bd04f1d141b422532b9a739e8635fb2957382e1568c2

Observation eb81c13e-73ad-42df-bdce-8e2274db605a · outbound

This paper cites Language modeling with gated convolutional networks,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Language modeling with gated convolutional networks,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.934226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.439805Z digest=sha256:236d2b7e4525df5389e32054df1469bdfc5f3adaab7020074bad5e7f2ec695f7

Observation dd1b4f6f-e28d-4026-b08d-8bcceceaadab · outbound

This paper cites Fre-GAN: Adversarial frequency-consistent audio synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fre-GAN: Adversarial frequency-consistent audio synthesis,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.908739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.453287Z digest=sha256:4a3573d8a5d941fa9f07268b7449fa6ae15f21212bef37ee0263dce66f6df8a6

Observation 966db735-0d84-4e0b-bec3-82e608b80322 · outbound

This paper cites UTMOS: Utokyo-sarulab system for voicemos challenge 2022,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation UTMOS: Utokyo-sarulab system for voicemos challenge 2022,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.888958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.460239Z digest=sha256:ae4c3e1db80a9f31db20200a01c3f54285148d807d3aef8d65aa3d4a08c1c96a

Observation 3704482b-c728-455c-ad3e-694cd1fe8531 · outbound

This paper cites The blizzard challenge 2005: Evaluating corpus-based speech synthesis on common databases,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The blizzard challenge 2005: Evaluating corpus-based speech synthesis on common databases,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.862661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.467343Z digest=sha256:9d63960823eab95b414e6a62f717d2e5ac030d2b07f62a8a7acd5a0c05e05e75

Observation e17948f1-e53e-4637-a248-cd30562989d9 · outbound

This paper cites USAT: A universal speaker-adaptive text-to-speech approach,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation USAT: A universal speaker-adaptive text-to-speech approach,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.841392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.473731Z digest=sha256:c6c374f12c7385e76c7b446817ee91476ab964461f439002cab6daf31bf852bd

Observation d5973154-1774-4f7a-9d78-36add20b08ae · outbound

This paper cites Dual-branch modeling based on state-space model for speech enhancement,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Dual-branch modeling based on state-space model for speech enhancement,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.803661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.484734Z digest=sha256:446c51016da784fbf3f37e5a9c2c3fb65b1f79d15e6211cad239ad4de6536ef6

Observation 5b5b390d-242c-4881-949c-f96c4c20a761 · outbound

This paper cites V oicegrad: Non-parallel any-to-many voice conversion with annealed langevin dy- namics,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation V oicegrad: Non-parallel any-to-many voice conversion with annealed langevin dy- namics,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.779771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.490103Z digest=sha256:55250a966b1d771891feac3b93356fe495526c43f241e6acfaf8443f6120671a

Observation 7f978c48-dc69-4df5-913d-dc6478aadf7d · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Robust speech recognition via large-scale weak super- vision,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.737629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.495439Z digest=sha256:086dd14c50e1d786dff60aa7c7e1ce03dcd718b8b7da3e9937646f6303b6efe0

Observation 192c509f-3006-45f7-88c2-f350c5d28d0d · outbound

This paper cites Visualizing data using t-SNE,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Visualizing data using t-SNE,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.710005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T23:41:01.500859Z digest=sha256:f1c67a84c939bd5a9c386148e07784024f6a35cc337a0dd503b003aeee938478

Observation 96e5ca60-1ce3-47fe-b1c8-a37cf7fd8237 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.505486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.505486Z digest=sha256:c1c6c4b90e7a7b328103b07bda761b5332364a5a87bde2cf3f86083fa7964d9d

Observation 99f0a98a-20ff-4fe6-ac2b-0712939f2c70 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.512970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.512970Z digest=sha256:72d771a26abdc4e4e5c2874b8af033a1b5c86bbbe9e5af33bdcefc92c1b1b06a

Pith citing papers

No inbound Pith citation observations are available.