Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:58.490777Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2505.19462.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:58.490777Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T23:22:40.218077Z
A source-named dated measurement, never combined with another source.
Source: cited_works
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1d9a321e-26ae-4b4f-8f8e-d74b54358d5a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b9dbd1-e4e9-4bdf-a20d-6b0e93c64b32 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation High Fidelity Neural Audio Compression
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2460dc9a-5101-4a73-b0ba-b16d0f199337 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05b4db4-dc42-462e-93ba-1d2c9cada271 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 413b2487-d923-4f77-921b-443342a17858 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e854fbbc-f772-4500-8022-aa9bc2b23ef7 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9918ab20-a36c-4e57-8ef1-e28ca562c615 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation V oicecraft: Zero-shot speech editing and text-to-speech in the wild
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6404e8c6-308b-48ab-b2ff-92543293ef9a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation On generative spoken language modeling from raw audio.Transactions of the Association for Computational Linguistics, 9:1336–1354, 2021
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09cc527d-f973-43eb-9ce5-a5c6bd7309e5 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiolm: A language modeling approach to audio generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2022
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebbb6bb1-5a6c-4a0b-8b3c-05f49645f4d4 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation AudioGen: Textually Guided Audio Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986297ee-6df1-432f-954f-d02096739bf5 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.Transactions of the Association for Computational Linguistics, 11:1703–1718, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb56d6d4-287a-4e76-8944-dd5c35448e75 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SoundStorm: Efficient Parallel Audio Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c44eb73f-b23e-479f-822f-8d99bf5d7fe1 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03333264-800a-4980-89ee-16fd2c942e25 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1a758a-6b3e-4758-a03d-e2c1796f49f5 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27763106-efe7-4f7b-834b-59b89b87e191 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9382fd7f-03e0-478b-9daa-2c1b9d74e65d · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b032d13-40de-4694-b4d7-3fd94ebbce1d · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b84f76b2-5175-4208-9515-21c1bc861091 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c099402e-4fee-4705-b271-6f0c51e6f718 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474ee284-a658-4782-b777-5fa639db753e · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 669cdead-4f97-4776-88a7-4186167f048e · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Attention- constrained inference for robust decoder-only text-to-speech.2024 IEEE Spoken Language Technology Workshop (SLT), pages 630–637, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60758394-101e-43fd-8c43-d555c12e201d · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85484e04-365d-408c-a020-e216707c56b9 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e5db4f-f25a-4307-9d78-607b7fadd358 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed023195-40ab-4345-b43f-4eda5a5f4a1a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873d0146-bc69-4a97-8ac2-5e2a681475df · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f44ca9-e9d1-47d2-974d-62fe1ca1d699 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff82b7b-3459-4cce-a038-ca8c65736997 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb710981-e29d-40bd-8c39-1305b93c9d47 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Desta, Roy Fejgin, Rafael Valle, and Jason Li
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e1872e8-7f8e-468d-bf75-a5acd1a15bc1 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012daf19-969a-4d96-a9d1-43fba4969a32 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc3048c9-080b-4860-b527-1fcdc0f87fc5 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6279db10-5903-4d71-adf5-822e051483e3 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41a6ef4-5484-4196-b1e3-a003a49b9111 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7ab772c-e025-41bb-a22a-24ce39553164 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f6cce3-b598-495d-8db5-7f4ce98de632 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Metis: A foundation speech generation model with masked generative pre-training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 556d9d7d-cac4-4f08-b791-226968ee62bb · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Prompttts: Controllable text-to-speech with text descriptions.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2022
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc53d6c9-e5c1-4e46-8f80-0122508deb35 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ceec292-ba82-4ff9-8de3-dfdc53dd723a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ac4ea4-86e6-4c5e-976e-593653ca213a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a8ba9e7-f22b-493f-b6e8-46229dd6c244 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation PromptTTS 2: Describing and Generating Voices with Text Prompt
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0230389-f536-413d-b453-4a4d8395e27d · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ff2b8f-9d24-4841-84b3-62414e426ed8 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Scaling rich style-prompted text-to-speech datasets
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61f0d21b-5e1d-40cf-91ef-18199528b189 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Moshi: a speech-text foundation model for real-time dialogue
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91447976-ac9f-4a61-9ca5-c61009dce3b5 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Language Model Can Listen While Speaking
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca4f5b7-4a33-40d1-aa74-c22a6e2656c5 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c878640-8d68-4e7b-97c0-52e1752594f0 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832e6d0f-0bf6-40fd-bae9-48da9426e755 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Enabling Real-Time Conversations with Minimal Training Costs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1889e61-ab6e-4d2e-92cd-ac6d3819c761 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1dad2e6-6205-48be-adc3-86345098b25a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llm-enhanced dialogue management for full-duplex spoken dialogue systems
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c38d1d3-bd8b-427f-bff7-00677543bc84 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 992085bb-1e13-4e77-a897-c0c0e2e2511d · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5407115-8935-4156-98a6-2901d3393e33 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556b0628-1098-4b0f-8d76-a2ba4395e922 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Shih, Rohan Badlani, João Felipe Santos, Evelina Bakhturina, Mikyas T
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37fd8c4d-8a2b-4ab1-87ab-b21dca27f849 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Ditto-tts: Diffusion transformers for scalable text-to-speech without domain-specific factors
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 433d7614-33db-4942-975d-0ae475faa212 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Peebles and Saining Xie
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22457582-96a8-4eea-9a32-f5ecc442de3a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Dmospeech: Direct metric optimization via distilled diffusion model in zero-shot speech synthesis
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e2b3cef-62ad-48af-b47a-1ddb4dcb323e · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.2024 IEEE Spoken Language Technology Workshop (SLT), pages 682–689, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a77996d5-2d3b-4528-9b14-7bd57feddcde · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 235bff6a-1867-49b6-87de-bd2096abb91f · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62d1ea48-ddeb-422a-afc6-9cd1cf666d36 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Simple and Controllable Music Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d23f41-e7c4-4aef-9075-65944122b465 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural machine translation by jointly learning to align and translate, 2016
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3c1883d-8489-4d10-998e-13991ac28e25 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Roformer: Enhanced transformer with rotary position embedding, 2023
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation deee603a-0ce1-4795-8030-32fc1e0db0d1 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df4e2a8-decb-4308-9f13-f80961a67480 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4750667-7041-41f0-84fe-797038ed46c6 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Smith, and Mike Lewis
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032496ba-6547-4544-80ef-b8c1d4fbca3b · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.2024 IEEE Spoken Language Technology Workshop (SLT), pages 885–890, 2024
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99fa98a6-c588-438f-b2d3-69705bb847a2 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b8773f0-7299-4522-ab03-186b7534942e · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation eSpeak NG: Speech synthesiser.https://github.com/espeak-ng/espeak-ng
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d541a03-a710-477c-a4db-f2ce42905748 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9d4cbb-a552-4058-8c24-680bc8651cad · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Zipformer: A faster and better encoder for automatic speech recognition
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc47ff8e-9be7-4e0f-aaca-e4557fceaf5d · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8179d0d-46ab-486a-a30d-a4ba1206f51a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Utmos: Utokyo-sarulab system for voicemos challenge 2022, 2022
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ffd8915-540a-457e-ad2a-feebe127b0e5 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust Speech Recognition via Large-Scale Weak Supervision
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9774d91c-d74b-419e-80da-0d25f97b1e5a · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16:1505–1518, 2021
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fcdb7826-d3bc-4038-8c87-1cdf16295d18 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speech quality assessment
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d91192c-2adb-4c99-bb69-82eb63c45c81 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcfe4542-9f26-42fc-a7fd-831a2e9c5aa1 · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Fast and high-quality auto- regressive speech synthesis via speculative decoding
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fa78f88-7c67-4229-9457-1550633ce4dd · outbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation natural” in comparative naturalness task with “similar
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a70ccd9-9a71-43f5-a190-b1e16e566dd0 · inbound
Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.