REVIEW 3 major objections 3 minor 159 references
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
T0 review · 3 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read SemBridge shows that supervising a continuous-latent autoregressive speech model with discrete semantic token labels during training, while keeping inference continuous, substantially improves content accuracy (WER/CER) without hurting…
desk verdict Novel training-only semantic anchoring for continuous AR speech; well-ablated and likely real, but single-run training leaves the effect size unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is semantic-token anchoring: a frozen semantic tokenizer (operating at 12.5 Hz, vocabulary 16,384) assigns each acoustic patch a discrete token ID and a continuous embedding; because the SA-VAE encodes 44.1 kHz audio into 25 Hz latent frames and two frames form one 12.5 Hz patch, patch $z_t$ and semantic token $s_t$ share a one-to-one causal alignment. A semantic prediction head attached after Transformer block $\ell^\star$ classifies $s_t$ from the hidden state $h_t^{(\ell^\star)}$, which sees only symbolic conditions and preceding patches, and this cross-entropy loss with weight $\lambda_{\mathrm{sem}}=0.1$ is added to the flow-matching and stop losses during training. The Semantic-Aligned Acoustic VAE (SA-VAE) complements this by aligning each patch with the tokenizer's embedding via a cosine-plus-$\ell^1$ loss, organizing the target space under the same semantic reference. At inference the semantic head is removed and generation stays entirely in the continuous PatchEnc-LM-LocDiT pathway, where LocDiT is the local diffusion transformer that models the distribution of each next acoustic patch.
What would settle it
Train the matched 0.8B/100K-hour generator several times with different seeds for the 'no anchor', '+Align', '+Anchor', and '+Align,+Anchor' configurations, and compare the spread of ZH CER / EN WER / ZH-Hard CER against the reported gaps (for example, 1.58 versus 1.01 on ZH CER). If the inter-run variation is comparable to or larger than those gaps, the central claim that semantic-token anchoring drives content fidelity loses support.
Extended reading notes
Core claim
On its own terms, the paper establishes that explicit semantic-token supervision of autoregressive states is what makes continuous-latent speech generation content-faithful, not just better target representations. With the same 0.8B backbone, the same 100K hours of data, and the same 300K-update budget, the unanchored baseline yields 1.58 CER on Mandarin, 2.43 WER on English, and 16.87 CER on the Chinese hard set; adding only target-space alignment gives 1.51, 2.30, and 15.94, while adding state-level semantic-token anchoring gives 1.21, 2.18, and 13.97, and the combination reaches 1.01, 1.87, and 11.87. The paper further shows that discrete token-ID classification outperforms continuous embedding regression using the same tokenizer, that anchoring at a middle layer (Anchor@24) best balances content against speaker similarity and quality, and that the same supervision improves lyric intelligibility in score-conditioned singing voice synthesis. The claimed discovery is that a training-only, discrete, token-level classification objective reorganizes continuous autoregressive states around linguistic content without changing the continuous-only generation interface.
Load-bearing premise
The load-bearing premise is that the reported WER/CER improvements are real effects of semantic-token anchoring rather than training noise: every configuration is trained once with a fixed seed (global seed 42), and the paper's own Limitations section states the evaluation does not characterize variation across independently trained models.
Editorial extensions
If this is right
- Continuous-latent autoregressive generators can gain content fidelity through a training-only auxiliary classification loss, without adding a discrete branch or quantization at inference time.
- The same frozen semantic tokenizer can supervise both the latent target space and the predictor states, and the two forms of supervision give complementary gains.
- Anchoring depth and loss weight trade content accuracy against speaker similarity and perceptual quality; system designers can pick the operating point, and the paper chooses Anchor@24 with $\lambda_{\mathrm{sem}}=0.1$.
- The benefit transfers from zero-shot TTS to score-conditioned singing, so the mechanism is not tied to one conditioning modality.
- Discrete token-ID classification is more effective than continuous embedding regression as a state-level semantic objective, suggesting categorical targets are a sharper training signal for LM states.
Reading between the lines
- The layer-following readability result suggests anchoring acts as an inductive bias that relocates where in the LM hierarchy linguistic content is linearly decodable; this could be tested as a general design rule for auxiliary heads in other continuous sequence models.
- If the causal claim survives seed variation, the same recipe of a frozen discrete tokenizer plus training-only classification on hidden states could improve content fidelity in other continuous-latent generative domains that have a discrete semantic tokenizer, such as music or sound-effects generation.
- The strong sensitivity to $\lambda_{\mathrm{sem}}$ and depth hints that an adaptive or layer-varying weight schedule might squeeze out further content gains while recovering the speaker-similarity losses observed at deeper anchoring.
- Because the system-level comparisons use unmatched external checkpoints, the controlled ablations, not the leaderboard, are the evidential core of the paper; a multi-seed rerun of the main ablation table would be the natural confirmation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SemBridge, a two-stage training framework for continuous-latent autoregressive speech generation. Stage I trains a Semantic-Aligned Acoustic VAE (SA-VAE) whose continuous latent patches are aligned with embeddings from a frozen GLM-4 semantic tokenizer. Stage II trains a continuous-latent autoregressive generator with a flow-matching objective and an auxiliary semantic-token classification loss applied to a selected hidden state ('semantic anchoring'). The semantic branch is removed at inference, so generation remains entirely continuous. Evaluations on zero-shot TTS (Seed-TTS-Eval, CV3-EVAL) and score-conditioned SVS (GMO-SVS, SoulX-Singer-Eval) show WER/CER reductions while maintaining comparable speaker similarity and perceptual quality, and controlled ablations on a fixed 0.8B/100K-hour setup attribute most of the gain to state anchoring, with additional gains from target-space alignment.
Significance. If the empirical claims hold, the paper makes a useful contribution by showing that discrete semantic-token labels can serve as an effective state-level regularizer for continuous-latent autoregressive speech models without changing the inference interface. The idea is simple and general, and the controlled ablations on the same backbone, data, and training budget are a methodological strength. The use of external ASR metrics (Whisper, Paraformer) avoids circularity with the GLM-4 tokenizer used for supervision. The convergence analysis at multiple checkpoints provides some internal consistency evidence for the anchoring effect. However, the lack of any training-seed variance estimates and a few unclear baseline definitions currently prevent the paper from fully supporting the generality claim that this direction is 'effective and general.'
major comments (3)
- [Limitations; Table 2; Table 4] The central controlled comparisons are each based on a single training run with global seed 42. The Limitations section correctly states that the evaluation 'does not characterize variation across independently trained models.' Because the headline claim is causal (the supervision objective changes content fidelity), the absence of any training variance estimate leaves open the possibility that part of the reported gaps (e.g., Table 2 ZH-Hard CER 16.87 to 11.87, EN WER 2.43 to 1.87; Table 4 Mandarin CER 9.18 to 8.32) is seed-dependent. The convergence curves in Figures 3 and 4 mitigate this concern but do not eliminate it, since they are also single runs. I request training at least three seeds for the key contrasts (no-anchor vs. full SemBridge, and the SVS w/o Anchoring contrast in Table 4) and reporting mean and standard deviation, or alternatively providing a clearly argued and empirically supported justification for why seed sensitivity is negligible for this architecture and 300K-update budget.
- [Table 2, 'Continuous-AR Baselines' rows; 'Acoustic Representation Control'] The rows labeled 'VoxCPM' and '+SA-VAE' are ambiguous. The published VoxCPM is 0.6B parameters and 1.8M training hours (Table 1), yet the caption states that all variants use the same 0.8B backbone, 100K hours, and 300K updates. If these are reimplementations of the VoxCPM architecture under matched settings, the paper should say so explicitly, describe the reimplementation, and state any differences from the SemBridge backbone besides the acoustic representation; if they are external reported numbers, the claim that 'Acoustic Representation Control' is a controlled comparison is invalid. This distinction is load-bearing for the conclusion that SA-VAE improves over the previous continuous-AR representation.
- [Method, Eqs. (1)-(3); 'Causally Aligned Semantic Prediction'] The one-to-one temporal correspondence z_t ↔ (e_t, s_t) is asserted from equal rates (12.5 Hz), but frame boundaries from a pretrained tokenizer and a causal VAE encoder are not guaranteed to align exactly without an explicit alignment procedure. Please provide an empirical verification of this correspondence (e.g., a frame-shift agreement analysis or a description of how tokenizer outputs are offset-adjusted before pairing). If the pairing is even slightly misaligned, both the SA-VAE alignment loss and the semantic anchoring objective could be supervising the wrong acoustic patches, which would undermine the method's core mechanism.
minor comments (3)
- [Section 'Semantic-Token Geometry at the Final Layer' vs. Method Eq. (8)] The indexing is inconsistent: the method states that the hidden state h_t predicts patch z_t and is supervised by semantic token s_t, while the geometry section says 'the hidden state at acoustic patch k is paired with semantic token k+1.' Please reconcile the notation to avoid an off-by-one confusion for the reader.
- [Figure 3 caption] The WER axis is limited to 25%, so the unanchored model's 58.38% WER at 20K updates is outside the displayed range; please add a note or an inset so that the early-training behavior is not hidden.
- [Table 2 caption] The caption states 'Tied results receive the same formatting,' but the table contains many tied values with formatting that is hard to parse in a monospaced rendering; please ensure the final typeset version makes the bolding/underlining unambiguous.
Circularity Check
No significant circularity: the content-accuracy evaluation uses external ASR systems independent of the GLM-4 semantic tokenizer that provides the supervision targets.
full rationale
SemBridge's central claim is that discrete semantic-token supervision of autoregressive states improves content accuracy. The measurement of content accuracy uses Whisper-large-v3 and Paraformer, which are independent of the GLM-4 tokenizer that supplies the semantic targets in Eqs. (1), (11), and (14). No fitted parameter is renamed as a prediction: the anchoring loss Lsem is a training objective, and the reported WER/CER gains come from matched ablations over the same backbone, data, and update budget (Table 2 and Table 11). The SA-VAE target-space alignment and the token-ID classification both use the same frozen tokenizer, but this is a deliberate shared-reference design choice rather than a definitional equivalence with the evaluation metric. The paper explicitly notes that cross-system comparison does not establish causality, and the Limitations section discloses the single-seed training and the absence of across-model variance estimates; that is a robustness limitation, not circularity. Self-citations such as SoulX-Singer appear as a baseline and evaluation protocol, but they are not load-bearing for the core derivation. The derivation chain is therefore self-contained against external benchmarks, and no circular step can be exhibited.
Assumptions & free parameters
free parameters (3)
- lambda_sem (semantic anchoring weight) =
0.1
- anchor depth (semantic head layer) =
24 (system-level), 32 (ablation default)
- lambda_align (SA-VAE alignment weight) =
100
assumptions (2)
- domain assumption GLM-4 semantic tokens and embeddings provide a reliable, temporally aligned proxy for linguistic content at 12.5 Hz.
- domain assumption The ASR metrics used for evaluation (Whisper-large-v3, Paraformer) accurately measure content fidelity.
Cite this review
Pith. "Pith review of SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation." pith.science (2026). https://pith.science/paper/BHIZKD5J
@misc{pith2026260807462,
author = {Pith},
title = {Pith review of: SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/BHIZKD5J}},
note = {Machine review of arXiv:2608.07462}
}
read the original abstract
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit token-level prediction tar- gets. Consequently, the autoregressive language model (LM) must acquire linguistic structure indirectly through acous- tic prediction, which can compromise the content fidelity of generated speech. We propose SemBridge, a training-only semantic-token anchoring framework for continuous-latent autoregressive speech generation. SemBridge uses discrete se- mantic tokens to directly supervise autoregressive LM states and employs a Semantic-Aligned Acoustic VAE to organize the continuous target space under the same semantic refer- ence. The semantic supervision is used only during train- ing, while inference remains entirely continuous. We evalu- ate SemBridge on zero-shot text-to-speech (TTS) and score- conditioned singing voice synthesis (SVS). Across multi- ple benchmarks, SemBridge improves content accuracy, as measured by word and character error rates (WER/CER), while maintaining competitive speaker similarity and percep- tual quality. Experimental results demonstrate that explicit semantic-token supervision for autoregressive state learning is an effective and general direction for continuous speech generation. Speech samples are available.1 The model code and checkpoints will be available at https://github.com/ASLP- lab/SemBridge
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Xia, Kangxiang and Zhu, Xinfa and Yao, Jixun and Tian, Wenjie and Li, Wenhao and Xie, Lei , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , doi =
2026
-
[2]
Kingma and Max Welling , title =
Diederik P. Kingma and Max Welling , title =. International Conference on Learning Representations , year =
-
[3]
International Conference on Learning Representations , publisher =
Ilya Loshchilov and Frank Hutter , title =. International Conference on Learning Representations , publisher =. 2019 , url =
2019
-
[4]
Advances in Neural Information Processing Systems , volume =
Rithesh Kumar and Prem Seetharaman and Alejandro Luebs and Ishaan Kumar and Kundan Kumar , title =. Advances in Neural Information Processing Systems , volume =. 2023 , doi =
2023
-
[5]
Moshi: a speech-text foundation model for real-time dialogue , journal =
Alexandre D. Moshi: a speech-text foundation model for real-time dialogue , journal =. 2024 , doi =. 2410.00037 , archiveprefix =
arXiv 2024
-
[6]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Yu Zhang and Rongjie Huang and Ruiqi Li and Jinzheng He and Yan Xia and Feiyang Chen and Xinyu Duan and Baoxing Huai and Zhou Zhao , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , doi =
2024
-
[7]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages =
Yu Zhang and Ziyue Jiang and Ruiqi Li and Changhao Pan and Jinzheng He and Rongjie Huang and Chuxin Wang and Zhou Zhao , title =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages =. 2024 , doi =
2024
-
[8]
International Conference on Learning Representations , publisher =
Xueyao Zhang and Xiaohui Zhang and Kainan Peng and Zhenyu Tang and Vimal Manohar and Yingru Liu and Jeff Hwang and Dangna Li and Yuhao Wang and Julian Chan and Yuan Huang and Zhizheng Wu and Mingbo Ma , title =. International Conference on Learning Representations , publisher =. 2025 , url =
2025
Show all 159 references
-
[9]
arXiv preprint arXiv:2512.04779 , year =
Junjie Zheng and Chunbo Hao and Guobin Ma and Xiaoyu Zhang and Gongyu Chen and Chaofan Ding and Zihao Chen and Lei Xie , title =. arXiv preprint arXiv:2512.04779 , year =. doi:10.48550/arXiv.2512.04779 , eprint =
-
[10]
Rix and John G
Antony W. Rix and John G. Beerends and Michael P. Hollier and Andries P. Hekstra , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2001 , doi =
2001
-
[11]
Taal and Richard C
Cees H. Taal and Richard C. Hendriks and Richard Heusdens and Jesper Jensen , title =. IEEE Transactions on Audio, Speech, and Language Processing , volume =. 2011 , doi =
2011
-
[12]
arXiv preprint arXiv:2606.06928 , year =
Yixuan Zhou and Guoyang Zeng and Xin Liu and Xiang Li and Renjie Yu and Jiancheng Gui and Jiaheng Wu and Ziyang Wang and Xudong Shen and Runchuan Ye and Zhisheng Zhang and Jiuyang Zhou and Bingsong Bai and Weiyue Sun and Mengyuan Deng and Qundong Shi and Zhiyong Wu and Zhiyuan...
-
[13]
arXiv preprint arXiv:2512.19090 , year =
Fan Yu and Tao Wang and You Wu and Lin Zhu and Wei Deng and Weisheng Han and Wenchao Wang and Lin Hu and Xiangyu Liang and Xiaodong He and Yankun Huang and Yu Gu and Yuan Liu and Yuxuan Wang and Zhangyu Xiao and Ziteng Wang and Boya Dong and Feng Dang and Jinming Chen and Jing...
-
[14]
IEEE Journal of Selected Topics in Signal Processing , volume =
Nguyen, Tu Anh and Sagot, Benoit and Dupoux, Emmanuel , title =. IEEE Journal of Selected Topics in Signal Processing , volume =. 2022 , doi =
2022
-
[15]
Proceedings of Interspeech , pages =
Saeki, Takaaki and Xin, Detai and Nakata, Wataru and Koriyama, Tomoki and Takamichi, Shinnosuke and Saruwatari, Hiroshi , title =. Proceedings of Interspeech , pages =. 2022 , doi =
2022
-
[16]
International Conference on Learning Representations , year =
Peng, Zhiliang and Yu, Jianwei and Wang, Wenhui and Chang, Yaoyao and Sun, Yutao and Dong, Li and Zhu, Yi and Xu, Weijiang and Bao, Hangbo and Wang, Zehua and Huang, Shaohan and Xia, Yan and Wei, Furu , title =. International Conference on Learning Representations , year =
-
[17]
arXiv preprint arXiv:2509.24650 , year =
Zhou, Yixuan and Zeng, Guoyang and Liu, Xin and Li, Xiang and Yu, Renjie and Wang, Ziyang and Ye, Runchuan and Sun, Weiyue and Gui, Jiancheng and Li, Kehan and Wu, Zhiyong and Liu, Zhiyuan , title =. arXiv preprint arXiv:2509.24650 , year =. doi:10.48550/arXiv.2509.24650 , eprint =
-
[18]
arXiv preprint arXiv:2605.16964 , year =
Wang, Huimeng and Lu, Hui and Deng, Jiajun and Xu, Haoning and Chen, Youjun and Chen, Xueyuan and Li, Zhaoqing and Peng, Shuhai and Kang, Shiyin and Liu, Xunying , title =. arXiv preprint arXiv:2605.16964 , year =. doi:10.48550/arXiv.2605.16964 , eprint =
- [19]
- [20]
-
[21]
arXiv preprint arXiv:2301.02111 , year =
Wang, Chengyi and Chen, Sanyuan and Wu, Yu and Zhang, Ziqiang and Zhou, Long and Liu, Shujie and Chen, Zhuo and Liu, Yanqing and Wang, Huaming and Li, Jinyu and He, Lei and Zhao, Sheng and Wei, Furu , title =. arXiv preprint arXiv:2301.02111 , year =. doi:10.48550/arXiv.2301.0...
- [22]
-
[23]
International Conference on Learning Representations , year =
Shen, Kai and Ju, Zeqian and Tan, Xu and Liu, Eric and Leng, Yichong and He, Lei and Qin, Tao and Zhao, Sheng and Bian, Jiang , title =. International Conference on Learning Representations , year =
-
[24]
Advances in Neural Information Processing Systems , volume =
Le, Matthew and Vyas, Apoorv and Shi, Bowen and Karrer, Brian and Sari, Leda and Moritz, Rashel and Williamson, Mary and Manohar, Vimal and Adi, Yossi and Mahadeokar, Jay and Hsu, Wei-Ning , title =. Advances in Neural Information Processing Systems , volume =. 2023 , doi =
2023
-
[25]
Proceedings of the IEEE Spoken Language Technology Workshop , pages =
Eskimez, Sefik Emre and Wang, Xiaofei and Thakker, Manthan and Li, Canrun and Tsai, Chung-Hsien and Xiao, Zhen and Yang, Hemin and Zhu, Zirun and Tang, Min and Tan, Xu and Liu, Yanqing and Zhao, Sheng and Kanda, Naoyuki , title =. Proceedings of the IEEE Spoken Language Techno...
2024
-
[26]
International Conference on Learning Representations , year =
Zhang, Xin and Zhang, Dong and Li, Shimin and Zhou, Yaqian and Qiu, Xipeng , title =. International Conference on Learning Representations , year =
-
[27]
International Conference on Learning Representations , year =
Ji, Shengpeng and Jiang, Ziyue and Wang, Wen and Chen, Yifu and Fang, Minghui and Zuo, Jialong and Yang, Qian and Cheng, Xize and Wang, Zehan and Li, Ruiqi and Zhang, Ziang and Yang, Xiaoda and Huang, Rongjie and Jiang, Yidi and Chen, Qian and Zheng, Siqi and Zhao, Zhou , titl...
-
[28]
International Conference on Learning Representations , year =
Yu, Sihyun and Kwak, Sangkyung and Jang, Huiwon and Jeong, Jongheon and Huang, Jonathan and Shin, Jinwoo and Xie, Saining , title =. International Conference on Learning Representations , year =
-
[30]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , year =
Wu, Chunyat and Deng, Jiajun and Liu, Zhengxi and Dai, Zheqi and He, Haolin and Kong, Qiuqiang , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , year =. doi:10.1109/ICASSP55912.2026.11463968 , url =
- [31]
-
[32]
arXiv preprint arXiv:2406.02430 , year =
Anastassiou, Philip and Chen, Jiawei and Chen, Jitong and Chen, Yuanzhe and Chen, Zhuo and Chen, Ziyi and Cong, Jian and Deng, Lelai and Ding, Chuang and Gao, Lu and Gong, Mingqing and Huang, Peisong and Huang, Qingqing and Huang, Zhiying and Huo, Yuanyuan and Jia, Dongya and ...
-
[33]
Proceedings of the International Conference on Machine Learning , series =
Dongya Jia and Zhuo Chen and Jiawei Chen and Chenpeng Du and Jian Wu and Jian Cong and Xiaobin Zhuang and Chumin Li and Zhen Wei and Yuping Wang and Yuxuan Wang , title =. Proceedings of the International Conference on Machine Learning , series =. 2025 , url =
2025
-
[34]
Meng and Furu Wei , title =
Lingwei Meng and Long Zhou and Shujie Liu and Sanyuan Chen and Bing Han and Shujie Hu and Yanqing Liu and Jinyu Li and Sheng Zhao and Xixin Wu and Helen M. Meng and Furu Wei , title =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol...
2025
-
[35]
arXiv preprint arXiv:2505.17589 , year =
Du, Zhihao and Gao, Changfeng and Wang, Yuxuan and Yu, Fan and Zhao, Tianyu and Wang, Hao and Lv, Xiang and Wang, Hui and Ni, Chongjia and Shi, Xian and An, Keyu and Yang, Guanrou and Li, Yabin and Chen, Yanni and Gao, Zhifu and Chen, Qian and Gu, Yue and Chen, Mengzhe and Che...
-
[36]
arXiv preprint arXiv:2602.07803 , year =
Qian, Jiale and Meng, Hao and Zheng, Tian and Zhu, Pengcheng and Lin, Haopeng and Dai, Yuhang and Xie, Hanke and Cao, Wenxiao and Shang, Ruixuan and Wu, Jun and Liu, Hongmei and Wen, Hanlin and Zhao, Jian and Jiang, Zhonglin and Chen, Yong and Yin, Shunshun and Tao, Ming and W...
-
[37]
Proceedings of Interspeech , pages =
Gao, Zhifu and Zhang, Shiliang and McLoughlin, Ian and Yan, Zhijie , title =. Proceedings of Interspeech , pages =. 2022 , doi =
2022
-
[38]
Proceedings of the International Conference on Machine Learning , pages =
Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya , title =. Proceedings of the International Conference on Machine Learning , pages =
-
[39]
IEEE Journal of Selected Topics in Signal Processing , volume =
Sanyuan Chen and Chengyi Wang and Zhengyang Chen and Yu Wu and Shujie Liu and Zhuo Chen and Jinyu Li and Naoyuki Kanda and Takuya Yoshioka and Xiong Xiao and Jian Wu and Long Zhou and Shuo Ren and Yanmin Qian and Yao Qian and Jian Wu and Michael Zeng and Xiangzhan Yu and Furu ...
2022
-
[40]
arXiv preprint arXiv:2510.01812 , year =
Tang, Yuxun and Liu, Lan and Feng, Wenhao and Zhao, Yiwen and Han, Jionghao and Yu, Yifeng and Shi, Jiatong and Jin, Qin , title =. arXiv preprint arXiv:2510.01812 , year =. doi:10.48550/arXiv.2510.01812 , eprint =
-
[41]
IEEE Transactions on Audio, Speech and Language Processing , volume =
Huang, Wen-Chin and Cooper, Erica and Toda, Tomoki , title =. IEEE Transactions on Audio, Speech and Language Processing , volume =. 2026 , doi =
2026
-
[42]
Advances in Neural Information Processing Systems , volume =
Yu Zhang and Changhao Pan and Wenxiang Guo and Ruiqi Li and Zhiyuan Zhu and Jialei Wang and Wenhao Xu and Jingyu Lu and Zhiqing Hong and Chuxin Wang and Lichao Zhang and Jinzheng He and Ziyue Jiang and Yuxin Chen and Chen Yang and Jiecheng Zhou and Xinyu Cheng and Zhou Zhao , ...
2024
-
[43]
Advances in Neural Information Processing Systems , volume =
Lichao Zhang and Ruiqi Li and Shoutong Wang and Liqun Deng and Jinglin Liu and Yi Ren and Jinzheng He and Rongjie Huang and Jieming Zhu and Xiao Chen and Zhou Zhao , title =. Advances in Neural Information Processing Systems , volume =. 2022 , doi =
2022
-
[44]
Proceedings of Interspeech , pages =
Wang, Yu and Wang, Xinsheng and Zhu, Pengcheng and Wu, Jie and Li, Hanzhao and Xue, Heyang and Zhang, Yongmao and Xie, Lei and Bi, Mengxiao , title =. Proceedings of Interspeech , pages =. 2022 , doi =
2022
-
[45]
Proceedings of the International Conference on Machine Learning , series =
Park, Se Jin and Salazar, Julian and Jansen, Aren and Kinoshita, Keisuke and Ro, Yong Man and Skerry-Ryan, RJ , title =. Proceedings of the International Conference on Machine Learning , series =. 2025 , url =
2025
-
[46]
International Conference on Learning Representations , year =
Jiang, Ziyue and Liu, Jinglin and Ren, Yi and He, Jinzheng and Ye, Zhenhui and Ji, Shengpeng and Yang, Qian and Zhang, Chen and Wei, Pengfei and Wang, Chunfeng and Yin, Xiang and Ma, Zejun and Zhao, Zhou , title =. International Conference on Learning Representations , year =
-
[47]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Liu, Jinglin and Li, Chengxi and Ren, Yi and Chen, Feiyang and Zhao, Zhou , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2022 , doi =
2022
-
[48]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Zhang, Yongmao and Cong, Jian and Xue, Heyang and Xie, Lei and Zhu, Pengcheng and Bi, Mengxiao , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2022 , doi =
2022
-
[49]
arXiv preprint arXiv:2601.03888 , year =
Li, Yunpei and Zhou, Xun and Wang, Jinchao and Wang, Lu and Wu, Yong and Zhou, Siyi and Zhou, Yiquan and Shu, Jingchen , title =. arXiv preprint arXiv:2601.03888 , year =. doi:10.48550/arXiv.2601.03888 , eprint =
-
[50]
Zico Kolter and Kaiming He , title =
Zhengyang Geng and Mingyang Deng and Xingjian Bai and J. Zico Kolter and Kaiming He , title =. Advances in Neural Information Processing Systems , year =
- [51]
-
[52]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , year =
Wang, Wei and Cao, Rong and Guo, Yi and Chen, Zhengyang and Chen, Kuan and Huo, Yuanyuan , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , year =
-
[53]
Advances in Neural Information Processing Systems , year =
Della Libera, Luca and Paissan, Francesco and Subakan, Cem and Ravanelli, Mirco , title =. Advances in Neural Information Processing Systems , year =
-
[54]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Zhen Ye and Peiwen Sun and Jiahe Lei and Hongzhan Lin and Xu Tan and Zheqi Dai and Qiuqiang Kong and Jianyi Chen and Jiahao Pan and Qifeng Liu and Yike Guo and Wei Xue , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2025 , doi =
2025
- [55]
-
[56]
arXiv preprint arXiv:2412.15115 , year =
An Yang and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chengyuan Li and Dayiheng Liu and Fei Huang and Haoran Wei and Huan Lin and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and Jingren Zhou and Junyang Lin and...
-
[57]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Kang, Wei and Yang, Xiaoyu and Yao, Zengwei and Kuang, Fangjun and Yang, Yifan and Guo, Liyong and Lin, Long and Povey, Daniel , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
-
[58]
Proceedings of the IEEE Spoken Language Technology Workshop , pages =
He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and others , title =. Proceedings of the IEEE Spoken Language Technology Workshop , pages =
- [59]
- [60]
-
[61]
arXiv preprint arXiv:2512.20151 , year =
Liu, Chengwei and Yan, Haoyin and Xue, Shaofei and Liang, Xiaotao and Chen, Xiaofu and Gong, Bin and Xue, Zheng and Song, Gang , title =. arXiv preprint arXiv:2512.20151 , year =. doi:10.48550/arXiv.2512.20151 , eprint =
-
[62]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =
Peng, Puyuan and Huang, Po-Yao and Li, Shang-Wen and Mohamed, Abdelrahman and Harwath, David , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =
-
[63]
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision , journal =
Kharitonov, Eugene and Vincent, Damien and Borsos, Zal. Speak, read and prompt: High-fidelity text-to-speech with minimal supervision , journal =. 2023 , doi =
2023
-
[64]
arXiv preprint arXiv:2512.14291 , year =
Jiayan Cui and Zhihan Yang and Naihan Li and Jiankun Tian and Xingyu Ma and Yi Zhang and Guangyu Chen and Runxuan Yang and Yuqing Cheng and Yizhi Zhou and Guochen Yu and Xiaotao Gu and Jie Tang , title =. arXiv preprint arXiv:2512.14291 , year =. doi:10.48550/arXiv.2512.14291 ...
-
[65]
arXiv preprint arXiv:2510.23541 , year =
Xie, Hanke and Lin, Haopeng and Cao, Wenxiao and Guo, Dake and Tian, Wenjie and Wu, Jun and Wen, Hanlin and Shang, Ruixuan and Liu, Hongmei and Jiang, Zhiqi and others , title =. arXiv preprint arXiv:2510.23541 , year =. doi:10.48550/arXiv.2510.23541 , eprint =
-
[66]
arXiv preprint arXiv:2512.18706 , year =
Liu, Zhanxun and Duan, Yifan and Wang, Mengmeng and Feng, Pengchao and Zhang, Haotian and Xing, Xiaoyu and Shan, Yijia and Zhu, Haina and Dai, Yuhang and Lu, Chaochao and others , title =. arXiv preprint arXiv:2512.18706 , year =. doi:10.48550/arXiv.2512.18706 , eprint =
-
[67]
arXiv preprint arXiv:2509.17765 , year =
Xu, Jin and Guo, Zhifang and Hu, Hangrui and Chu, Yunfei and Wang, Xiong and He, Jinzheng and Wang, Yuxuan and Shi, Xian and He, Ting and Zhu, Xinfa and others , title =. arXiv preprint arXiv:2509.17765 , year =. doi:10.48550/arXiv.2509.17765 , eprint =
-
[68]
arXiv preprint arXiv:2601.15621 , year =
Hu, Hangrui and Zhu, Xinfa and He, Ting and Guo, Dake and Zhang, Bin and Wang, Xiong and Guo, Zhifang and Jiang, Ziyue and Hao, Hongkun and Guo, Zishan and others , title =. arXiv preprint arXiv:2601.15621 , year =. doi:10.48550/arXiv.2601.15621 , eprint =
-
[69]
arXiv preprint arXiv:2510.08392 , year =
Ma, Guobin and Yao, Jixun and Ning, Ziqian and Jiang, Yuepeng and Xiong, Lingxin and Xie, Lei and Zhu, Pengcheng , title =. arXiv preprint arXiv:2510.08392 , year =. doi:10.48550/arXiv.2510.08392 , eprint =
- [70]
-
[71]
arXiv preprint arXiv:2509.18196 , year =
Jialong Mai and Jinxin Ji and Xiaofen Xing and Chen Yang and Weidong Chen and Jingyuan Xing and Xiangmin Xu , title =. arXiv preprint arXiv:2509.18196 , year =. doi:10.48550/arXiv.2509.18196 , eprint =
-
[72]
Proceedings of the International Conference on Machine Learning , series =
Sanyuan Chen and Yu Wu and Chengyi Wang and Shujie Liu and Daniel Tompkins and Zhuo Chen and Wanxiang Che and Xiangzhan Yu and Furu Wei , title =. Proceedings of the International Conference on Machine Learning , series =. 2023 , url =
2023
-
[73]
arXiv preprint arXiv:2507.06261 , year =
Comanici, Gheorghe and Bieber, Eric and Schaekermann, Mike and Pasupat, Ice and Sachdeva, Noveen and Dhillon, Inderjit and Blistein, Marcel and Ram, Ori and Zhang, Dan and Rosen, Evan and others , title =. arXiv preprint arXiv:2507.06261 , year =. doi:10.48550/arXiv.2507.06261...
-
[74]
2024 , url =
Jordan Darefsky and Ge Zhu and Zhiyao Duan , title =. 2024 , url =
2024
-
[75]
Proceedings of the International Conference on Machine Learning , series =
Park, Taejin and Medennikov, Ivan and Dhawan, Kunal and Wang, Weiqing and Huang, He and Koluguri, Nithin Rao and Puvvada, Krishna C and Balam, Jagadeesh and Ginsburg, Boris , title =. Proceedings of the International Conference on Machine Learning , series =. 2025 , url =
2025
-
[76]
arXiv preprint arXiv:2509.03959 , year =
Li, Longhao and Guo, Zhao and Chen, Hongjie and Dai, Yuhang and Zhang, Ziyu and Xue, Hongfei and Zuo, Tianlun and Wang, Chengyou and Wang, Shuiyuan and Li, Jie and others , title =. arXiv preprint arXiv:2509.03959 , year =. doi:10.48550/arXiv.2509.03959 , eprint =
-
[77]
arXiv preprint arXiv:2509.18004 , year =
Dai, Yuhang and Zhang, Ziyu and Wang, Shuai and Li, Longhao and Guo, Zhao and Zuo, Tianlun and Wang, Shuiyuan and Xue, Hongfei and Wang, Chengyou and Wang, Qing and others , title =. arXiv preprint arXiv:2509.18004 , year =. doi:10.48550/arXiv.2509.18004 , eprint =
- [78]
-
[79]
arXiv preprint arXiv:2505.09388 , year =
Yang, An and Li, Anfeng and Yang, Baosong and Zhang, Beichen and Hui, Binyuan and Zheng, Bo and Yu, Bowen and Gao, Chang and Huang, Chengen and Lv, Chenxu and others , title =. arXiv preprint arXiv:2505.09388 , year =. doi:10.48550/arXiv.2505.09388 , eprint =
- [80]
-
[81]
arXiv preprint arXiv:2507.09318 , year =
Zhu, Han and Kang, Wei and Guo, Liyong and Yao, Zengwei and Kuang, Fangjun and Zhuang, Weiji and Li, Zhaoqing and Han, Zhifeng and Zhang, Dong and Zhang, Xin and others , title =. arXiv preprint arXiv:2507.09318 , year =. doi:10.48550/arXiv.2507.09318 , eprint =
- [82]
- [83]
- [84]
- [85]
-
[86]
Advances in Neural Information Processing Systems , volume =
Leying Zhang and Yao Qian and Long Zhou and Shujie Liu and Dongmei Wang and Xiaofei Wang and Midia Yousefi and Yanmin Qian and Jinyu Li and Lei He and Sheng Zhao and Michael Zeng , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =
2024
- [87]
- [88]
-
[89]
arXiv preprint arXiv:2508.19205 , year =
Peng, Zhiliang and Yu, Jianwei and Wang, Wenhui and Chang, Yaoyao and Sun, Yutao and Dong, Li and Zhu, Yi and Xu, Weijiang and Bao, Hangbo and Wang, Zehua and others , title =. arXiv preprint arXiv:2508.19205 , year =. doi:10.48550/arXiv.2508.19205 , eprint =
-
[90]
Text to Spoken Dialogue Generation , year =
- [91]
-
[92]
arXiv preprint arXiv:2410.03930 , year =
Nishchal Bhandari and Danny Chen and Miguel Ãngel del RÃo Fernández and Natalie Delworth and Jennifer Drexler Fox and Migüel Jetté and Quinten McNamara and Corey Miller and OndÅ™ej Novotný and Ján Profant and Nan Qin and Martin Ratajczak and Jean-Philippe Robichaud , ti...
-
[93]
IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =
Changsheng Quan and Xiaofei Li , title =. IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =. 2024 , doi =
2024
-
[94]
Chandan K. A. Reddy and Vishak Gopal and Ross Cutler , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2022 , doi =
2022
-
[95]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Zhang, Shiliang and Lei, Ming and Yan, Zhijie and Dai, Lirong , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2018 , doi =
2018
- [96]
-
[97]
Advances in Neural Information Processing Systems , volume =
Sang. Advances in Neural Information Processing Systems , volume =
-
[98]
arXiv preprint arXiv:2503.20215 , year =
Jin Xu and Zhifang Guo and Jinzheng He and Hangrui Hu and Ting He and Shuai Bai and Keqin Chen and Jialin Wang and Yang Fan and Kai Dang and Bin Zhang and Xiong Wang and Yunfei Chu and Junyang Lin , title =. arXiv preprint arXiv:2503.20215 , year =. doi:10.48550/arXiv.2503.202...
-
[99]
Proceedings of Interspeech , pages =
Brecht Desplanques and Jenthe Thienpondt and Kris Demuynck , title =. Proceedings of Interspeech , pages =. 2020 , doi =
2020
-
[100]
International Conference on Learning Representations , publisher =
Sang. International Conference on Learning Representations , publisher =. 2023 , url =
2023
-
[101]
Findings of the Association for Computational Linguistics: EACL 2024 , volume =
Neil Shah and Saiteja Kosgi and Vishal Tambrahalli and Neha S and Anil Nelakanti and Vineet Gandhi , title =. Findings of the Association for Computational Linguistics: EACL 2024 , volume =. 2024 , doi =
2024
-
[102]
Proceedings of Interspeech , pages =
Ramanan Sivaguru and Vasista Sai Lodagala and Srinivasan Umesh , title =. Proceedings of Interspeech , pages =. 2023 , doi =
2023
- [103]
-
[104]
High Fidelity Neural Audio Compression , journal =
Alexandre D. High Fidelity Neural Audio Compression , journal =
-
[105]
arXiv preprint arXiv:2407.05407 , year =
Zhihao Du and Qian Chen and Shiliang Zhang and Kai Hu and Heng Lu and Yexin Yang and Hangrui Hu and Siqi Zheng and Yue Gu and Ziyang Ma and Zhifu Gao and Zhijie Yan , title =. arXiv preprint arXiv:2407.05407 , year =. doi:10.48550/arXiv.2407.05407 , eprint =
-
[106]
arXiv preprint arXiv:2412.10117 , year =
Zhihao Du and Yuxuan Wang and Qian Chen and Xian Shi and Xiang Lv and Tianyu Zhao and Zhifu Gao and Yexin Yang and Changfeng Gao and Hui Wang and Fan Yu and Huadai Liu and Zhengyan Sheng and Yue Gu and Chong Deng and Wen Wang and Shiliang Zhang and Zhijie Yan and Jingren Zhou ...
- [107]
-
[108]
arXiv preprint arXiv:1907.11692 , year =
Liu, Yinhan and Ott, Myle and Goyal, Naman and Du, Jingfei and Joshi, Mandar and Chen, Danqi and Levy, Omer and Lewis, Mike and Zettlemoyer, Luke and Stoyanov, Veselin , title =. arXiv preprint arXiv:1907.11692 , year =. doi:10.48550/arXiv.1907.11692 , eprint =
- [109]
-
[110]
arXiv preprint arXiv:2303.08774 , year =
Achiam, Josh and Adler, Steven and Agarwal, Sandhini and Ahmad, Lama and Akkaya, Ilge and Aleman, Florencia Leoni and Almeida, Diogo and Altenschmidt, Janko and Altman, Sam and Anadkat, Shyamal and others , title =. arXiv preprint arXiv:2303.08774 , year =. doi:10.48550/arXiv....
-
[111]
Advances in Neural Information Processing Systems , volume =
Kong, Jungil and Kim, Jaehyeon and Bae, Jaekyoung , title =. Advances in Neural Information Processing Systems , volume =
-
[112]
Generative Spoken Dialogue Language Modeling , journal =
Tu Anh Nguyen and Eugene Kharitonov and Jade Copet and Yossi Adi and Wei. Generative Spoken Dialogue Language Modeling , journal =
-
[113]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Li, Hanzhao and Zhu, Xinfa and Xue, Liumeng and Song, Yang and Chen, Yunlin and Xie, Lei , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2024 , doi =
2024
- [114]
- [115]
-
[116]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Xu, Yue and Li, Yong-Lu and Huang, Zhemin and Liu, Michael Xu and Lu, Cewu and Tai, Yu-Wing and Tang, Chi-Keung , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =. 2023 , doi =
2023
-
[117]
arXiv preprint arXiv:2503.01710 , year =
Wang, Xinsheng and Jiang, Mingqi and Ma, Ziyang and Zhang, Ziyu and Liu, Songxiang and Li, Linqin and Liang, Zheng and Zheng, Qixi and Wang, Rui and Feng, Xiaoqin and others , title =. arXiv preprint arXiv:2503.01710 , year =. doi:10.48550/arXiv.2503.01710 , eprint =
-
[118]
Proceedings of the IEEE International Conference on Multimedia and Expo , year =
Zhu, Xinfa and Li, Yuke and Lei, Yi and Jiang, Ning and Zhao, Guoqing and Xie, Lei , title =. Proceedings of the IEEE International Conference on Multimedia and Expo , year =. doi:10.1109/ICME57554.2024.10688322 , url =
2024
- [119]
-
[120]
arXiv preprint arXiv:1703.10135 , year =
Wang, Yuxuan and Skerry-Ryan, RJ and Stanton, Daisy and Wu, Yonghui and Weiss, Ron J and Jaitly, Navdeep and Yang, Zongheng and Xiao, Ying and Chen, Zhifeng and Bengio, Samy and others , title =. arXiv preprint arXiv:1703.10135 , year =. doi:10.48550/arXiv.1703.10135 , eprint =
-
[121]
Advances in Neural Information Processing Systems , volume =
Ren, Yi and Ruan, Yangjun and Tan, Xu and Qin, Tao and Zhao, Sheng and Zhao, Zhou and Liu, Tie-Yan , title =. Advances in Neural Information Processing Systems , volume =
-
[122]
Proceedings of the International Conference on Machine Learning , series =
Jaehyeon Kim and Jungil Kong and Juhee Son , title =. Proceedings of the International Conference on Machine Learning , series =
-
[123]
arXiv preprint arXiv:2309.16609 , year =
Bai, Jinze and Bai, Shuai and Chu, Yunfei and Cui, Zeyu and Dang, Kai and Deng, Xiaodong and Fan, Yang and Ge, Wenbin and Han, Yu and Huang, Fei and others , title =. arXiv preprint arXiv:2309.16609 , year =. doi:10.48550/arXiv.2309.16609 , eprint =
-
[124]
Minds and Machines , volume =
Floridi, Luciano and Chiriatti, Massimo , title =. Minds and Machines , volume =. 2020 , doi =
2020
- [125]
-
[126]
Attention is all you need , booktitle =
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser,. Attention is all you need , booktitle =
-
[127]
Communications of the ACM , volume =
Goodfellow, Ian and Pouget-Abadie, Jean and Mirza, Mehdi and Xu, Bing and Warde-Farley, David and Ozair, Sherjil and Courville, Aaron and Bengio, Yoshua , title =. Communications of the ACM , volume =
- [128]
-
[129]
IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =
Hsu, Wei-Ning and Bolte, Benjamin and Tsai, Yao-Hung Hubert and Lakhotia, Kushal and Salakhutdinov, Ruslan and Mohamed, Abdelrahman , title =. IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume =
-
[130]
Proceedings of the 20th Chinese National Conference on Computational Linguistics , pages =
Liu, Zhuang and Lin, Wayne and Shi, Ya and Zhao, Jun , title =. Proceedings of the 20th Chinese National Conference on Computational Linguistics , pages =. 2021 , url =
2021
-
[131]
Journal of Machine Learning Research , volume =
Raffel, Colin and Shazeer, Noam and Roberts, Adam and Lee, Katherine and Narang, Sharan and Matena, Michael and Zhou, Yanqi and Li, Wei and Liu, Peter J , title =. Journal of Machine Learning Research , volume =
-
[132]
IEEE Transactions on Knowledge and Data Engineering , volume =
Liu, Xiao and Zhang, Fanjin and Hou, Zhenyu and Mian, Li and Wang, Zhaoyu and Zhang, Jing and Tang, Jie , title =. IEEE Transactions on Knowledge and Data Engineering , volume =. 2023 , doi =
2023
-
[133]
Advances in Neural Information Processing Systems , volume =
Alexei Baevski and Yuhao Zhou and Abdelrahman Mohamed and Michael Auli , title =. Advances in Neural Information Processing Systems , volume =
-
[134]
Advances in Neural Information Processing Systems , volume =
Ho, Jonathan and Jain, Ajay and Abbeel, Pieter , title =. Advances in Neural Information Processing Systems , volume =
-
[135]
arXiv preprint arXiv:1609.03499 , year =
Aaron van den Oord and Sander Dieleman and Heiga Zen and Karen Simonyan and Oriol Vinyals and Alex Graves and Nal Kalchbrenner and Andrew Senior and Koray Kavukcuoglu , title =. arXiv preprint arXiv:1609.03499 , year =. doi:10.48550/arXiv.1609.03499 , eprint =
-
[136]
arXiv preprint arXiv:2303.18223 , year =
Zhao, Wayne Xin and Zhou, Kun and Li, Junyi and Tang, Tianyi and Wang, Xiaolei and Hou, Yupeng and Min, Yingqian and Zhang, Beichen and Zhang, Junjie and Dong, Zican and others , title =. arXiv preprint arXiv:2303.18223 , year =. doi:10.48550/arXiv.2303.18223 , eprint =
-
[137]
Proceedings of the International Conference on Machine Learning , series =
Li, Junnan and Li, Dongxu and Savarese, Silvio and Hoi, Steven , title =. Proceedings of the International Conference on Machine Learning , series =
-
[138]
Advances in Neural Information Processing Systems , volume =
Wenliang Dai and Junnan Li and Dongxu Li and Anthony Meng Huat Tiong and Junqi Zhao and Weisheng Wang and Boyang Li and Pascale Fung and Steven Hoi , title =. Advances in Neural Information Processing Systems , volume =. 2023 , url =
2023
-
[139]
Advances in Neural Information Processing Systems , volume =
Alayrac, Jean-Baptiste and Donahue, Jeff and Luc, Pauline and Miech, Antoine and Barr, Iain and Hasson, Yana and Lenc, Karel and Mensch, Arthur and Millican, Katherine and Reynolds, Malcolm and others , title =. Advances in Neural Information Processing Systems , volume =
-
[140]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Elizalde, Benjamin and Deshmukh, Soham and Al Ismail, Mahmoud and Wang, Huaming , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2023 , doi =
2023
-
[141]
arXiv preprint arXiv:2502.04128 , year =
Ye, Zhen and Zhu, Xinfa and Chan, Chi-Min and Wang, Xinsheng and Tan, Xu and Lei, Jiahe and Peng, Yi and Liu, Haohe and Jin, Yizhu and Dai, Zheqi and others , title =. arXiv preprint arXiv:2502.04128 , year =. doi:10.48550/arXiv.2502.04128 , eprint =
-
[142]
Advances in Neural Information Processing Systems , volume =
Aaron van den Oord and Oriol Vinyals and Koray Kavukcuoglu , title =. Advances in Neural Information Processing Systems , volume =. 2017 , url =
2017
- [143]
- [144]
- [145]
- [146]
- [147]
-
[148]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Ren, Yong and Wang, Tao and Yi, Jiangyan and Xu, Le and Tao, Jianhua and Zhang, Chu Yuan and Zhou, Junzuo , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
- [149]
- [150]
-
[151]
arXiv preprint arXiv:2510.16841 , year =
Chen, Wenxi and Wang, Xinsheng and Yan, Ruiqi and Chen, Yushen and Niu, Zhikang and Ma, Ziyang and Li, Xiquan and Liang, Yuzhe and Wen, Hanlin and Yin, Shunshun and others , title =. arXiv preprint arXiv:2510.16841 , year =. doi:10.48550/arXiv.2510.16841 , eprint =
-
[152]
Xtts: a massively multilingual zero-shot text-to-speech model , journal =
Casanova, Edresson and Davis, Kelly and G. Xtts: a massively multilingual zero-shot text-to-speech model , journal =. 2024 , doi =. 2406.04904 , archiveprefix =
2024 arXiv
-
[153]
Man-Machine Speech Communication , series =
Guo, Dake and Yao, Jixun and Ma, Lihan and Wang, He and Xie, Lei , title =. Man-Machine Speech Communication , series =. 2026 , doi =
2026
- [154]
-
[155]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Le Lan, Gael and Nagaraja, Varun and Chang, Ernie and Kant, David and Ni, Zhaoheng and Shi, Yangyang and Iandola, Forrest and Chandra, Vikas , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
-
[156]
Simple and controllable music generation , booktitle =
Copet, Jade and Kreuk, Felix and Gat, Itai and Remez, Tal and Kant, David and Synnaeve, Gabriel and Adi, Yossi and D. Simple and controllable music generation , booktitle =. 2023 , doi =
2023
-
[157]
Masked audio generation using a single non-autoregressive transformer , journal =
Ziv, Alon and Gat, Itai and Lan, Gael Le and Remez, Tal and Kreuk, Felix and D. Masked audio generation using a single non-autoregressive transformer , journal =. 2024 , doi =. 2401.04577 , archiveprefix =
2024 arXiv
-
[158]
arXiv preprint arXiv:2409.00750 , year =
Wang, Yuancheng and Zhan, Haoyue and Liu, Liwei and Zeng, Ruihong and Guo, Haotian and Zheng, Jiachen and Zhang, Qiang and Zhang, Xueyao and Zhang, Shunsi and Wu, Zhizheng , title =. arXiv preprint arXiv:2409.00750 , year =. doi:10.48550/arXiv.2409.00750 , eprint =
-
[159]
2024 , url =
Silero. 2024 , url =
2024
-
[160]
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Chen, Zhengyang and Chen, Sanyuan and Wu, Yu and Qian, Yao and Wang, Chengyi and Liu, Shujie and Qian, Yanmin and Zeng, Michael , title =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.