REVIEW 5 major objections 5 minor 1 cited by
A purely discrete flow matching model for text-to-speech claims state-of-the-art prosody cloning and 25.8x faster generation than current baselines.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 18:47 UTC pith:QHJ7NOQB
load-bearing objection A credible first application of discrete flow matching to zero-shot TTS, with real efficiency gains and honest ablations, but the headline SOTA numbers rest on an unreported evaluation protocol and small margins; deserves a serious referee but not acceptance as-is. the 5 major comments →
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that a masked-token discrete flow matching objective with factorized prediction heads is sufficient for high-quality zero-shot TTS. Using FaCodec to factor speech into content, prosody, acoustic details, and speaker embedding, the model maps phonemes to content tokens via a duration-predictor-aligned mapper, then regenerates prosody and acoustic tokens in parallel via discrete flow. The reported experiments show new best pitch and energy metrics, word error rate tied with the best baseline (0.05), UTMOS within 0.11 of ground truth, and competitive subjective MOS, all with far fewer parameters and much lower latency than prior systems.
What carries the argument
The central component is the Factorized Discrete Flow Denoiser (FDFD): a diffusion-transformer backbone that estimates the probability velocity over a discrete target space, with the target factorized into prosody and acoustic token streams under an assumed independence. Two parallel prediction heads, one per attribute, output categorical distributions over the token vocabulary, letting a single network learn aspect-specific distributions. The Phoneme-Content Mapper (PCM) generates content conditioning by predicting discrete content tokens from phonemes, using a duration predictor and length regulator to produce content embeddings used by the denoiser.
Load-bearing premise
The load-bearing premise is that FaCodec's factorization cleanly separates content, prosody, acoustic detail, and speaker identity in the training and reference audio; if content tokens leak prosody or acoustic tokens leak speaker traits, the claimed attribute cloning and clean WER would degrade.
What would settle it
A concrete check: evaluate DiFlow-TTS while swapping the reference prosody tokens with those from a different speaker but keeping acoustic tokens and speaker embedding fixed. If the synthesized F0 does not follow the reference F0, or if the voice still matches the original speaker, the factorization is entangled and the prosody metrics are not attributable to the separate prosody stream. A second check: measure correlation between predicted prosody and acoustic token distributions on held-out LibriTTS pairs; high correlation would directly contradict Definition 1's independence.
If this is right
- Purely discrete flow matching can replace continuous flow matching in TTS without sacrificing quality, simplifying optimization and enabling fully non-autoregressive decoding.
- Zero-shot voice cloning is feasible in compact models: the 164M-parameter architecture suggests DFM scales well for edge and low-latency deployment.
- Prosody and acoustic attributes can be generated in parallel with separate heads, enabling targeted pitch/energy cloning from a few seconds of reference audio.
- At 16 function evaluations, WER and speaker similarity are already near-saturated, so low-latency TTS can ship without a severe quality drop.
- The factorized two-stream prediction extends DFM to composite structured targets, a pattern applicable to other multi-attribute discrete generation tasks.
Where Pith is reading between the lines
- If FaCodec's factorization is not perfectly disentangled, the reported prosody gains may partly reflect content-conditioned prosody modeling rather than true attribute separation; targeted disentanglement tests would clarify credit assignment.
- The independence assumption between prosody and acoustic targets contradicts known correlations in real speech, yet the model still works; swapping reference acoustic tokens while fixing prosody tokens would reveal how much the shared backbone compensates for cross-attribute leakage.
- The method's data efficiency (470h versus 9K-100K hours for baselines) suggests the masking objective regularizes strongly; scaling to 1K+ hours could push naturalness past the current leader without architectural change.
- The same factorized-head design could be applied to other discrete speech factorizations (e.g., emotion, speaking rate, or language streams), enabling finer-grained control than prosody-plus-acoustic.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DiFlow-TTS, a zero-shot text-to-speech system built on purely discrete flow matching over factorized FaCodec tokens. The architecture combines a deterministic Phoneme-Content Mapper (PCM) for linguistic modeling and a Factorized Discrete Flow Denoiser (FDFD) that jointly generates prosody and acoustic token streams with separate prediction heads. The authors claim new state-of-the-art pitch and energy metrics (F0 RMSE 7.97 Hz, energy accuracy 0.73), a tied-best WER of 0.05, and runner-up UTMOS of 3.98 on LibriSpeech test-clean despite training on only 470 hours, at 164M parameters and RTF 0.066 at 16 NFE (up to 25.8x faster than listed baselines). The paper includes ablations of each component, an NFE trade-off study, and a MOS evaluation with confidence intervals.
Significance. If the empirical claims hold, DiFlow-TTS would be a meaningful contribution: it is, to the best of my knowledge, the first application of purely discrete flow matching to speech synthesis, and the factorized-head design is a clean way to model multiple token streams in one backbone. The paper also honestly reports that speaker similarity trails SOTA and that removing attribute embeddings slightly improves UTMOS. The ablations and the NFE study are useful. The main value is in showing that a compact, non-autoregressive discrete flow model can be competitive with much larger or slower systems. However, the evaluation protocol is under-specified and the supplementary material is absent, which currently blocks verification of the headline 'SOTA' claims.
major comments (5)
- [Evaluation Metrics / Table 1] The objective metrics in Table 1 are not self-contained. The 'Evaluation Metrics' paragraph defers all protocol details to a Supplementary Material section that is not present in the manuscript. Concretely, the paper does not specify the F0/energy extraction tool, the parameter ranges or accuracy thresholds, how prosody tokens are aligned with frames, which ASR model is used for WER, which speaker-verification model is used for SIM-O/SIM-R, or how the 3-second prompts are sampled. The claimed margins are small (WER tie at 0.05, F0 RMSE 7.97 vs 11.96, energy accuracy 0.73 vs 0.67), so the SOTA prosody claim and the WER tie are vulnerable to metric-version and evaluation-noise effects. Please include a complete protocol in the paper, ideally with a released evaluation harness.
- [Table 1 / Baseline comparability] Baseline numbers in Table 1 are a mix of author reproductions ([⋄]) and values inferred from official/unofficial checkpoints ([†]/[‡]), trained on corpora ranging from 9K to 100K hours versus 470 hours for DiFlow-TTS. No error bars, confidence intervals, or significance tests are reported for any objective metric. Since the baselines are not evaluated under an identical, frozen protocol, the differences are confounded with dataset and evaluation version. At minimum, report per-utterance distributions or bootstrap confidence intervals, and clearly separate reproduced from inferred numbers in the analysis.
- [Table 3 vs Table 1 and Table 5] The headline quality numbers (UTMOS 3.98, F0 RMSE 7.97, energy accuracy 0.73, WER 0.05) are reported at 128 NFE in Table 1, while the speed claim in the abstract and Table 3 uses 16 NFE, where Table 5 shows UTMOS drops to 3.86. The abstract's 'up to 25.8x faster' is not accompanied by the full metric set at the fast setting. Please report the complete Table 1 metrics at 16 NFE (or at least the prosody/energy metrics) so that the quality-efficiency trade-off is transparent for the headline claim.
- [Definition 1] Definition 1 assumes exact independence between prosody and acoustic detail sequences: q(x) = q_p(x^p) * q_a(x^a). In real speech these streams are strongly correlated (e.g., pitch range and voice quality depend on the same articulatory gestures). The paper does not validate this assumption or test how much the shared DiT backbone compensates for it. A simple diagnostic, such as measuring mutual information between the predicted prosody and acoustic distributions, or an ablation where the two heads share a joint prediction instead of factorized outputs, would address whether the factorization is a harmless modeling choice or a source of bias.
- [Eq. (3) / Speech Tokenization] The whole pipeline relies on FaCodec's factorization (Eq. 3) cleanly separating content, prosody, acoustic detail, and speaker identity. The paper cites NaturalSpeech 3 but provides no local check that this disentanglement holds on LibriTTS audio. If content tokens leak prosody or speaker identity, the PCM's linguistic modeling and the FDFD's attribute conditioning would be flawed, degrading the reported WER and speaker similarity. Please include a local disentanglement test (e.g., speaker classification or prosody regression from content tokens) or at least discuss the evidence for FaCodec's factorization on the training domain.
minor comments (5)
- [Figure 1 caption] The caption says 'DiLow-TTS' but the model is 'DiFlow-TTS'; please fix.
- [Table 2] Table 2 lists 'V ALLE-E' while the text and references use 'V ALL-E'. Please make the spelling consistent.
- [References] Reference 'Liu, Gong, and qiang liu 2023' has improper capitalization; should be 'Liu, Gong, and Liu'.
- [Table 5] The text says 'optimal audio quality observed at 64 NFE', but Table 5 shows UTMOS 3.958 at 64 and 3.978 at 128. Clarify what 'optimal' refers to, or correct the statement.
- [Supplementary material] Multiple sections (Dataset Details, Metrics Details, Implementation Details, Baselines Details) are promised in the Supplementary Material, but no supplementary file appears in the manuscript. This is a presentation issue only if the supplement is provided with the revision, but it is essential for reproducibility.
Circularity Check
No significant circularity: the derivation is externally anchored and the empirical claims are measured, not fitted.
full rationale
The paper's generative formalism (discrete flow matching, mixture path, probability velocity) is explicitly taken from Gat et al. (2024), an external prior work, and the speech tokenization/factorization is explicitly taken from Ju et al. (2024)/NaturalSpeech 3, also external. The same-group prior work OZSpeech appears only as a comparison baseline and as the target of a design critique; no load-bearing derivation step depends on OZSpeech's correctness. The training objective (Eq. 6) is a cross-entropy loss against ground-truth discrete codec tokens, not against the evaluation metrics, and the headline numbers (UTMOS, WER, F0 RMSE, energy accuracy) are measured on LibriSpeech test-clean rather than optimized or fitted. The independence assumption in Definition 1 is a stated modeling assumption, not a hidden restatement of the conclusion. Deferring detailed metric protocols to a supplementary that is not included is a reproducibility/verifiability concern, not circularity. No specific reduction of a claimed result to its own input could be exhibited.
Axiom & Free-Parameter Ledger
free parameters (5)
- Loss weights lambda_dur, lambda_c, lambda_FDFD =
not reported in main text
- Masking scheduler kappa_t =
not specified beyond kappa_0=0, kappa_1=1
- F0/energy 'accuracy' thresholds =
not specified
- NFE (sampling steps) =
16 for latency claim, 128 for quality claim
- Architecture hyperparameters (DiT depth/width, heads, hidden dim) =
not reported in main text
axioms (5)
- standard math Standard DFM formalism: mask-to-data mixture path, probability velocity formula (Eq. 1), denoiser posterior (Eq. 2)
- domain assumption FaCodec factorization yields cleanly separated content, prosody, acoustic, and speaker streams
- ad hoc to paper Prosody and acoustic target sequences are statistically independent (q = qp * qa)
- domain assumption Concatenating reference attribute tokens with corrupted targets enables in-context attribute cloning
- domain assumption Phoneme durations align to integer spans of content tokens
read the original abstract
Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregressive models suffer from high latency, while diffusion-based approaches are constrained by training-time configurations. Moreover, most flow-based methods operate in continuous space, which introduces optimization challenges because continuous token spaces are inherently more complex than discrete ones. To address these limitations, we propose DiFlow-TTS, a novel zero-shot TTS framework based on discrete flow matching. The model consists of a deterministic Phoneme-Content Mapper for linguistic modeling and a Factorized Discrete Flow Denoiser that simultaneously generates prosody and acoustic token streams. Experimental results demonstrate the effectiveness of our approach across multiple evaluation metrics.
Figures
Forward citations
Cited by 1 Pith paper
-
Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus
A bilingual TTS system for the Peruvian Constitution in Quechua and Spanish is developed with XTTS v2, F5-TTS, and DiFlow-TTS, releasing checkpoints and audio to support low-resource speech synthesis.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
D.; Ho, J.; Tarlow, D.; and van den Berg, R
Austin, J.; Johnson, D. D.; Ho, J.; Tarlow, D.; and van den Berg, R. 2021. Structured Denoising Diffusion Models in Discrete State-Spaces. In Beygelzimer, A.; Dauphin, Y.; Liang, P.; and Vaughan, J. W., eds., Advances in Neural Information Processing Systems
2021
-
[4]
Campbell, A.; Yim, J.; Barzilay, R.; Rainforth, T.; and Jaakkola, T. S. 2024. Generative Flows on Discrete State-Spaces: Enabling Multimodal Flows with Applications to Protein Co-Design. In ICML
2024
-
[5]
Chang, H.; Zhang, H.; Jiang, L.; Liu, C.; and Freeman, W. T. 2022. MaskGIT: Masked Generative Image Transformer. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[6]
Chen, S.; Liu, S.; Zhou, L.; Liu, Y.; Tan, X.; Li, J.; Zhao, S.; Qian, Y.; and Wei, F. 2024 a . VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers. arXiv:2406.05370
Pith/arXiv arXiv 2024
-
[7]
Chen, S.; Wang, C.; Wu, Y.; Zhang, Z.; Zhou, L.; Liu, S.; Chen, Z.; Liu, Y.; Wang, H.; Li, J.; He, L.; Zhao, S.; and Wei, F. 2025. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers. IEEE Transactions on Audio, Speech and Language Processing, 1--15
2025
-
[8]
Chen, Y.; Niu, Z.; Ma, Z.; Deng, K.; Wang, C.; Zhao, J.; Yu, K.; and Chen, X. 2024 b . F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. arXiv:2410.06885
Pith/arXiv arXiv 2024
-
[9]
Du, C.; Guo, Y.; Shen, F.; Liu, Z.; Liang, Z.; Chen, X.; Wang, S.; Zhang, H.; and Yu, K. 2024 a . UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding. Proceedings of the AAAI Conference on Artificial Intelligence, 38(16): 17924--17932
2024
-
[10]
Du, Z.; Wang, Y.; Chen, Q.; Shi, X.; Lv, X.; Zhao, T.; Gao, Z.; Yang, Y.; Gao, C.; Wang, H.; et al. 2024 b . Cosyvoice 2: Scalable streaming speech synthesis with large language models. arXiv preprint arXiv:2412.10117
Pith/arXiv arXiv 2024
-
[11]
E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.-H.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; et al
Eskimez, S. E.; Wang, X.; Thakker, M.; Li, C.; Tsai, C.-H.; Xiao, Z.; Yang, H.; Zhu, Z.; Tang, M.; Tan, X.; et al. 2024. E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts. In 2024 IEEE Spoken Language Technology Workshop (SLT), 682--689. IEEE
2024
-
[12]
Fuest, M.; Hu, V. T.; and Ommer, B. 2025. Maskflow: Discrete flows for flexible and efficient long video generation. arXiv preprint arXiv:2502.11234
Pith/arXiv arXiv 2025
-
[13]
Gat, I.; Remez, T.; Shaul, N.; Kreuk, F.; Chen, R. T. Q.; Synnaeve, G.; Adi, Y.; and Lipman, Y. 2024. Discrete Flow Matching. In Globerson, A.; Mackey, L.; Belgrave, D.; Fan, A.; Paquet, U.; Tomczak, J.; and Zhang, C., eds., Advances in Neural Information Processing Systems, volume 37, 133345--133385. Curran Associates, Inc
2024
-
[14]
Guan, W.; Su, Q.; Zhou, H.; Miao, S.; Xie, X.; Li, L.; and Hong, Q. 2024. Reflow-tts: A rectified flow model for high-fidelity text-to-speech. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 10501--10505. IEEE
2024
-
[15]
Han, B.; Zhou, L.; Liu, S.; Chen, S.; Meng, L.; Qian, Y.; Liu, Y.; Zhao, S.; Li, J.; and Wei, F. 2024. VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment. arXiv preprint arXiv:2406.07855
Pith/arXiv arXiv 2024
-
[16]
Hieu, N. H. N.; Nguyen, N. S.; Dang, H. N.; Vo, T.; Hy, T.-S.; and Nguyen, V. 2025. OZS peech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 21...
2025
-
[17]
Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising Diffusion Probabilistic Models. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems, volume 33, 6840--6851. Curran Associates, Inc
2020
-
[18]
Ji, S.; Jiang, Z.; Wang, H.; Zuo, J.; and Zhao, Z. 2024. M obile S peech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13588--13600. Bangkok, Thailand: Association fo...
2024
-
[19]
Jia, D.; Chen, Z.; Chen, J.; Du, C.; Wu, J.; Cong, J.; Zhuang, X.; Li, C.; Wei, Z.; Wang, Y.; and Wang, Y. 2025. Di TAR : Diffusion Transformer Autoregressive Modeling for Speech Generation. In Forty-second International Conference on Machine Learning
2025
-
[20]
Ju, Z.; Wang, Y.; Shen, K.; Tan, X.; Xin, D.; Yang, D.; Liu, E.; Leng, Y.; Song, K.; Tang, S.; Wu, Z.; Qin, T.; Li, X.; Ye, W.; Zhang, S.; Bian, J.; He, L.; Li, J.; and Zhao, S. 2024. N atural S peech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. In Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J....
2024
-
[21]
J.; and Yang, E
Kang, M.; Han, W.; Hwang, S. J.; and Yang, E. 2023. ZET-Speech: Zero-shot Adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models. In Interspeech 2023, 4339--4343
2023
-
[22]
J.; Badlani, R.; Santos, J
Kim, S.; Shih, K. J.; Badlani, R.; Santos, J. F.; Bakhturina, E.; Desta, M. T.; Valle, R.; Yoon, S.; and Catanzaro, B. 2023. P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[23]
Le, M.; Vyas, A.; Shi, B.; Karrer, B.; Sari, L.; Moritz, R.; Williamson, M.; Manohar, V.; Adi, Y.; Mahadeokar, J.; and Hsu, W.-N. 2023. Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[24]
W.; Kim, J.; Chung, S.; and Cho, J
Lee, K.; Kim, D. W.; Kim, J.; Chung, S.; and Cho, J. 2025. Di TT o- TTS : Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors. In The Thirteenth International Conference on Learning Representations
2025
-
[25]
Lipman, Y.; Chen, R. T. Q.; Ben-Hamu, H.; Nickel, M.; and Le, M. 2023. Flow Matching for Generative Modeling. In The Eleventh International Conference on Learning Representations
2023
-
[26]
Liu, X.; Gong, C.; and qiang liu. 2023. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. In The Eleventh International Conference on Learning Representations
2023
-
[27]
Lou, A.; Meng, C.; and Ermon, S. 2024. Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution. In Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J.; and Berkenkamp, F., eds., Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 328...
2024
-
[28]
Mehta, S.; Tu, R.; Beskow, J.; Sz\'ekely, E.; and Henter, G. E. 2024. Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 11341--11345
2024
-
[29]
M.; and Wei, F
Meng, L.; Zhou, L.; Liu, S.; Chen, S.; Han, B.; Hu, S.; Liu, Y.; Li, J.; Zhao, S.; Wu, X.; Meng, H. M.; and Wei, F. 2025. Autoregressive Speech Synthesis without Vector Quantization. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...
2025
-
[30]
Panayotov, V.; Chen, G.; Povey, D.; and Khudanpur, S. 2015. Librispeech: An ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5206--5210
2015
-
[31]
Peebles, W.; and Xie, S. 2023. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, 4195--4205
2023
-
[32]
Peng, P.; Huang, P.-Y.; Li, S.-W.; Mohamed, A.; and Harwath, D. 2024. V oice C raft: Zero-Shot Speech Editing and Text-to-Speech in the Wild. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12442--12462. Bangkok, Thailand: Association for Co...
2024
-
[33]
Qin, Y.; Madeira, M.; Thanou, D.; and Frossard, P. 2025. DeFoG: Discrete Flow Matching for Graph Generation. In Proceedings of the 42nd International Conference on Machine Learning (ICML)
2025
-
[34]
S.; Arriola, M.; Gokaslan, A.; Marroquin, E
Sahoo, S. S.; Arriola, M.; Gokaslan, A.; Marroquin, E. M.; Rush, A. M.; Schiff, Y.; Chiu, J. T.; and Kuleshov, V. 2024. Simple and Effective Masked Diffusion Language Models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[35]
Shaul, N.; Gat, I.; Havasi, M.; Severo, D.; Sriram, A.; Holderrieth, P.; Karrer, B.; Lipman, Y.; and Chen, R. T. Q. 2025. Flow Matching with General Discrete Paths: A Kinetic-Optimal Perspective. In The Thirteenth International Conference on Learning Representations
2025
-
[36]
Shen, K.; Ju, Z.; Tan, X.; Liu, E.; Leng, Y.; He, L.; Qin, T.; sheng zhao; and Bian, J. 2024. NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers. In The Twelfth International Conference on Learning Representations
2024
-
[37]
Shi, J.; Han, K.; Wang, Z.; Doucet, A.; and Titsias, M. 2024. Simplified and Generalized Masked Diffusion for Discrete Data. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[38]
Song, Y.; Chen, Z.; Wang, X.; Ma, Z.; and Chen, X. 2024. ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering. arXiv:2401.07333
Pith/arXiv arXiv 2024
-
[39]
P.; Kumar, A.; Ermon, S.; and Poole, B
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2021. Score-Based Generative Modeling through Stochastic Differential Equations. In International Conference on Learning Representations
2021
-
[40]
van den Oord, A.; Vinyals, O.; and Kavukcuoglu, K. 2017. Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, 6309–6318. Red Hook, NY, USA: Curran Associates Inc. ISBN 9781510860964
2017
-
[41]
N.; Kaiser, L
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017. Attention is All you Need. In Guyon, I.; Luxburg, U. V.; Bengio, S.; Wallach, H.; Fergus, R.; Vishwanathan, S.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc
2017
-
[42]
Wang, K.; Guan, W.; Jiang, Z.; Huang, H.; Chen, P.; Wu, W.; Hong, Q.; and Li, L. 2025 a . Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion. arXiv preprint arXiv:2505.24291
Pith/arXiv arXiv 2025
-
[43]
Wang, X.; Jiang, M.; Ma, Z.; Zhang, Z.; Liu, S.; Li, L.; Liang, Z.; Zheng, Q.; Wang, R.; Feng, X.; et al. 2025 b . Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens. arXiv preprint arXiv:2503.01710
Pith/arXiv arXiv 2025
-
[44]
Yadav, R.; Yan, Q.; Wolf, G.; Bose, A. J.; and Liao, R. 2025. RETRO SYNFLOW: Discrete Flow Matching for Accurate and Diverse Single-Step Retrosynthesis. arXiv preprint arXiv:2506.04439
arXiv 2025
-
[45]
Yao, J.; Yuguang, Y.; Pan, Y.; Ning, Z.; Ye, J.; Zhou, H.; and Xie, L. 2025. Stablevc: Style controllable zero-shot voice conversion with conditional flow matching. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 25669--25677
2025
-
[46]
Ye, J.; Cao, B.; and Shan, H. 2025. Emotional Face-to-Speech. In Forty-second International Conference on Machine Learning
2025
-
[47]
Ye, J.; and Shan, H. 2025. Shushing! Let's Imagine an Authentic Speech from the Silent Video. arXiv preprint arXiv:2503.14928
Pith/arXiv arXiv 2025
-
[48]
Yi, K.; Jamali, K.; and Scheres, S. H. 2025. All-atom inverse protein folding through discrete flow matching. In Forty-second International Conference on Machine Learning
2025
-
[49]
J.; Jia, Y.; Chen, Z.; and Wu, Y
Zen, H.; Dang, V.; Clark, R.; Zhang, Y.; Weiss, R. J.; Jia, Y.; Chen, Z.; and Wu, Y. 2019. LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech. In Interspeech 2019, 1526--1530
2019
-
[50]
Zhang, Z.; Zhou, L.; Wang, C.; Chen, S.; Wu, Y.; Liu, S.; Chen, Z.; Liu, Y.; Wang, H.; Li, J.; et al. 2023. Speak foreign languages with your own voice: Cross-lingual neural codec language modeling. arXiv preprint arXiv:2303.03926
Pith/arXiv arXiv 2023
-
[51]
Zuo, J.; Ji, S.; Fang, M.; Jiang, Z.; Cheng, X.; Yang, Q.; Liu, W.; Zhang, G.; Tu, Z.; Guo, Y.; et al. 2025 a . Enhancing expressive voice conversion with discrete pitch-conditioned flow matching model. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5. IEEE
2025
-
[52]
Zuo, J.; Ji, S.; Fang, M.; Li, M.; Jiang, Z.; Cheng, X.; Yang, X.; Feiyang, C.; Duan, X.; and Zhao, Z. 2025 b . Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vo...
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.