Pith. sign in

Paper Citation Record · LEDGER

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2505.19462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19462 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:58.490777Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T23:22:40.218077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d9a321e-26ae-4b4f-8f8e-d74b54358d5a · outbound

This paper cites Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:49.841291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:49.841291Z digest=sha256:d01e6825ddd202be482c0e27fb9a19fb5d8b934be9e6b142aca8a7de8ef51261

Observation 70b9dbd1-e4e9-4bdf-a20d-6b0e93c64b32 · outbound

This paper cites High Fidelity Neural Audio Compression.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation High Fidelity Neural Audio Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:49.901237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:49.901237Z digest=sha256:fa7edf783563df222a0470450c328a29eb90af0ebe5a9fe1bcbfdf035626e364

Observation 2460dc9a-5101-4a73-b0ba-b16d0f199337 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.066833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.066833Z digest=sha256:33e3b986022c52cfae1ef5c6c461c2325b91ff8d0ef716119c9a8f991ef39079

Observation c05b4db4-dc42-462e-93ba-1d2c9cada271 · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.173827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.173827Z digest=sha256:eb74322fa573aefffcad0abc5a693863f9219fcc06df49a099d37e1b569855a9

Observation 413b2487-d923-4f77-921b-443342a17858 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.254673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.254673Z digest=sha256:249c7dffae196870cf123fe17ab7a4bff4df4ae893b8471c7f769511034848ab

Observation e854fbbc-f772-4500-8022-aa9bc2b23ef7 · outbound

This paper cites Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:05.424679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:50.374743Z digest=sha256:67ddd2a39077cbbf78cb302fe2e0c8e1b28220c7562657aeb9c3182fde93f71f

Observation 9918ab20-a36c-4e57-8ef1-e28ca562c615 · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation V oicecraft: Zero-shot speech editing and text-to-speech in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:05.056341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:50.481272Z digest=sha256:a93b3739a40eb5afd69ee61dd60e0903bc8de3dfa37fc751c8b451b18758423a

Observation 6404e8c6-308b-48ab-b2ff-92543293ef9a · outbound

This paper cites On generative spoken language modeling from raw audio.Transactions of the Association for Computational Linguistics, 9:1336–1354, 2021.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation On generative spoken language modeling from raw audio.Transactions of the Association for Computational Linguistics, 9:1336–1354, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:04.656485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:50.603358Z digest=sha256:e7dd5c85582b31c249741b60ac2e747ec14761e4ee10cfbe6c2e1ea474c7d02a

Observation 09cc527d-f973-43eb-9ce5-a5c6bd7309e5 · outbound

This paper cites Audiolm: A language modeling approach to audio generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiolm: A language modeling approach to audio generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:04.287314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:50.727860Z digest=sha256:054bd5b9f895ff2527fefb0443787a9310fe35866b506fd2defc3688d6480024

Observation ebbb6bb1-5a6c-4a0b-8b3c-05f49645f4d4 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation AudioGen: Textually Guided Audio Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.833721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.833721Z digest=sha256:ff5885e03650e9eb66e5468ca96273ca99c269584bc2024ce1869c2983972891

Observation 986297ee-6df1-432f-954f-d02096739bf5 · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.Transactions of the Association for Computational Linguistics, 11:1703–1718, 2023.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.Transactions of the Association for Computational Linguistics, 11:1703–1718, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.955411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:50.893333Z digest=sha256:3347a44f3459ab1b6931a71dc363b8af4d1df0903844cc6f684df2277f1c9dbc

Observation bb56d6d4-287a-4e76-8944-dd5c35448e75 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SoundStorm: Efficient Parallel Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.973655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.973655Z digest=sha256:8812b50c061f4e196adc099535b2436cf6738992fba319a374cad8b9f46e76ec

Observation c44eb73f-b23e-479f-822f-8d99bf5d7fe1 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.036153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.036153Z digest=sha256:3dd28cc93b90fe7fa1135064d1ea91189b1029aaab96b93ae307c6b0749bb71a

Observation 03333264-800a-4980-89ee-16fd2c942e25 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.135367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.135367Z digest=sha256:8855b4403a0de9443afe83a29ff8989c3820da6e8870a6d49bcdefa8b78c9ac8

Observation 3a1a758a-6b3e-4758-a03d-e2c1796f49f5 · outbound

This paper cites Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:19:58.963053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:51.283267Z digest=sha256:9742277590fc39d45b5ae051cbcb8d71cc1bf46fdd9099a828cd4ff0ab123cd8

Observation 27763106-efe7-4f7b-834b-59b89b87e191 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.454188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.454188Z digest=sha256:363b90336f673fed2e8bbf15c9d3d5186a0a41b3330d6b44a9d021905b01e62b

Observation 9382fd7f-03e0-478b-9daa-2c1b9d74e65d · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.609437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.609437Z digest=sha256:6cfa5edafae12ad6abaadb421995c562fbcdd18f41a08435dd3a1f0918c194cb

Observation 3b032d13-40de-4694-b4d7-3fd94ebbce1d · outbound

This paper cites Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.667748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:51.705529Z digest=sha256:4f20247ab1882875b2b856b901741ce42dd84b324a4a2fab3ebe9fa4136ce808

Observation b84f76b2-5175-4208-9515-21c1bc861091 · outbound

This paper cites Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.396435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:51.826778Z digest=sha256:b2a57e00e318d5964e3cf04253e4433a3615ffc9d5633e5f3f276da5a1c88eb9

Observation c099402e-4fee-4705-b271-6f0c51e6f718 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.952647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.952647Z digest=sha256:17d34bfabd67f0694467872e9dc6f4f3a795418687453669c73be59645cfc976

Observation 474ee284-a658-4782-b777-5fa639db753e · outbound

This paper cites Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.193256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:52.092794Z digest=sha256:76a62a489c7cac9be50a145320f343127e0e65a9768936972f738f3308862b89

Observation 669cdead-4f97-4776-88a7-4186167f048e · outbound

This paper cites Attention- constrained inference for robust decoder-only text-to-speech.2024 IEEE Spoken Language Technology Workshop (SLT), pages 630–637, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Attention- constrained inference for robust decoder-only text-to-speech.2024 IEEE Spoken Language Technology Workshop (SLT), pages 630–637, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.969811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:52.174655Z digest=sha256:5fc450074655359345f2d18428bce82b27619b4aabef2f94dadad70e15bd3d03

Observation 60758394-101e-43fd-8c43-d555c12e201d · outbound

This paper cites Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.306333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.306333Z digest=sha256:42bfae59b4ef8fdc6913a9fc5cdbb4118b331f0650b64ae3f4036a6c73249ce2

Observation 85484e04-365d-408c-a020-e216707c56b9 · outbound

This paper cites Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.424117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.424117Z digest=sha256:719363ce4a52f6cc406b5a8ef02bbcf188c98c06645716c3a2a7c9d2d8e86aea

Observation 32e5db4f-f25a-4307-9d78-607b7fadd358 · outbound

This paper cites Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.826188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:52.548815Z digest=sha256:2a8e7f7235ca528643eedf5a2169048118ab51c2b9c55cb77f81fd225b6c713c

Observation ed023195-40ab-4345-b43f-4eda5a5f4a1a · outbound

This paper cites SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.702338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.702338Z digest=sha256:3bbf951fab55ff5545dd14896400ee32931ef7aabe52da8413d8ac8f44e6c93f

Observation 873d0146-bc69-4a97-8ac2-5e2a681475df · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.899498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.899498Z digest=sha256:1b3a1cf2dc566875429de6c344fe0ecebad0eab6042045c8ca0d1556d1c1c290

Observation 35f44ca9-e9d1-47d2-974d-62fe1ca1d699 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.043874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.043874Z digest=sha256:9b2d86a03bdfc2ec10b071b50b2af1489c2f690272ae84205a8f5da805c2212e

Observation 4ff82b7b-3459-4cce-a038-ca8c65736997 · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.169474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.169474Z digest=sha256:f2faaa1801ccb82ee38323477cf9ab391344ee0fd0099ab3222d31bb5834a5a1

Observation cb710981-e29d-40bd-8c39-1305b93c9d47 · outbound

This paper cites Desta, Roy Fejgin, Rafael Valle, and Jason Li.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Desta, Roy Fejgin, Rafael Valle, and Jason Li

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.632742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:53.311490Z digest=sha256:bbb0d09a9362f1093b0abbb30591b0befcf9f7404f56080e8e9a20697d94aa25

Observation 7e1872e8-7f8e-468d-bf75-a5acd1a15bc1 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.462342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.462342Z digest=sha256:3b9d2cd6e68a3420c573ceb47798dfc84e9cdf76efc08f95ee97f718d4f70ad3

Observation 012daf19-969a-4d96-a9d1-43fba4969a32 · outbound

This paper cites Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.626191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.626191Z digest=sha256:45fef3ce50f28f60305079ebebfcdd943250f19685b1ab460b8ec4f515416892

Observation dc3048c9-080b-4860-b527-1fcdc0f87fc5 · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.771128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.771128Z digest=sha256:5371c6233e7d74d7467e334c65f1be1b787307388a66616068230e4d0d6a481e

Observation 6279db10-5903-4d71-adf5-822e051483e3 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.892566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.892566Z digest=sha256:c9ce35ea9cf9ad6321ce3cd2196c25342b7ecd4927b8a694663dc0e5ed3d8805

Observation f41a6ef4-5484-4196-b1e3-a003a49b9111 · outbound

This paper cites an unresolved cited work.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:20:02.462834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:54.016530Z digest=sha256:232c4ee7598dab5d4f6b0b6fe4903e450d62c547a071096fef11c62758329de8

Observation d7ab772c-e025-41bb-a22a-24ce39553164 · outbound

This paper cites SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.204415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.204415Z digest=sha256:049a19ab8e74e8f3bc1f4045355e7a65153d0808ba55ca625cd1f6f147c847a9

Observation 84f6cce3-b598-495d-8db5-7f4ce98de632 · outbound

This paper cites Metis: A foundation speech generation model with masked generative pre-training.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Metis: A foundation speech generation model with masked generative pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.247472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:54.327352Z digest=sha256:41e72348e946038baab51c54c45d1397e34c264ba732cf639466db73720dd2db

Observation 556d9d7d-cac4-4f08-b791-226968ee62bb · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Prompttts: Controllable text-to-speech with text descriptions.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.047892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:54.422588Z digest=sha256:d1239a7bfa1e839da301037d86b5cd7b309baf196fa74691f26842f1172da078

Observation fc53d6c9-e5c1-4e46-8f80-0122508deb35 · outbound

This paper cites InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.567317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.567317Z digest=sha256:f1abf43af4021bb839985fe6c157237a661cf9c09e71480afdd56c74afecb7a1

Observation 2ceec292-ba82-4ff9-8de3-dfdc53dd723a · outbound

This paper cites PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.715981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.715981Z digest=sha256:81874d509baed9817c364369600eefd934fa615f12e6b87926716b02fbf2959e

Observation a3ac4ea4-86e6-4c5e-976e-593653ca213a · outbound

This paper cites TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:19:58.766245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:54.931851Z digest=sha256:e32e7a69b366fec5dabe29ff64e7f7dab15074ffff30fa38545dae4fc9576d72

Observation 1a8ba9e7-f22b-493f-b6e8-46229dd6c244 · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.063985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.063985Z digest=sha256:aae7688d9a48f95ca1745dba35941278de02eb2dfca1fcbc98aeda468f3fdf45

Observation a0230389-f536-413d-b453-4a4d8395e27d · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.191205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.191205Z digest=sha256:1831be914c003d0be60726ab76bea627898565e041e2b446467e500e3189a72c

Observation 44ff2b8f-9d24-4841-84b3-62414e426ed8 · outbound

This paper cites Scaling rich style-prompted text-to-speech datasets.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Scaling rich style-prompted text-to-speech datasets

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.854731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:55.310667Z digest=sha256:de9597e4de736813e48919745a53b21b0c1dcb2dae31c0af576e2b1de9ef6c4a

Observation 61f0d21b-5e1d-40cf-91ef-18199528b189 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Moshi: a speech-text foundation model for real-time dialogue

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.465301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.465301Z digest=sha256:cbf94ad82e59205af48bf602213a92c5bfe3f73c708231ecec4616fbe16f92dc

Observation 91447976-ac9f-4a61-9ca5-c61009dce3b5 · outbound

This paper cites Language Model Can Listen While Speaking.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Language Model Can Listen While Speaking

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.581577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.581577Z digest=sha256:fcc85c16d8a39e9d2aaf6663dae11a7c86c199d4f393469b2590733da3d57533

Observation 9ca4f5b7-4a33-40d1-aa74-c22a6e2656c5 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.644063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.644063Z digest=sha256:7f1dc67d73cf87b84452867bf153a8d682c55217773b6721d36ce5d85363def2

Observation 0c878640-8d68-4e7b-97c0-52e1752594f0 · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.723227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.723227Z digest=sha256:9bc55f492645fc3d5c8e9375110aee6ca1f703f392c7e7a15d29e75c3e9b5f28

Observation 832e6d0f-0bf6-40fd-bae9-48da9426e755 · outbound

This paper cites Enabling Real-Time Conversations with Minimal Training Costs.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Enabling Real-Time Conversations with Minimal Training Costs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.833609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.833609Z digest=sha256:45e90b6ac638658b136bab237eb4077f51c6c97c8fdade7a3c138ea4f5e5ceea

Observation b1889e61-ab6e-4d2e-92cd-ac6d3819c761 · outbound

This paper cites IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.952954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.952954Z digest=sha256:c7b4a6996b7e94c920fb845adb37b82560b472deb4039eef99faa791092922f6

Observation c1dad2e6-6205-48be-adc3-86345098b25a · outbound

This paper cites Llm-enhanced dialogue management for full-duplex spoken dialogue systems.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llm-enhanced dialogue management for full-duplex spoken dialogue systems

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.644408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:56.035454Z digest=sha256:2aa37efc96f494de078579691940ce796f655ec13d9d4e35b2eba64b12839083

Observation 5c38d1d3-bd8b-427f-bff7-00677543bc84 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.106020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.106020Z digest=sha256:deac6d9a9e2e3b7b3d5847703f89d6c1c71642e074a30f7f8649b490f0ec686c

Observation 992085bb-1e13-4e77-a897-c0c0e2e2511d · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.181491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.181491Z digest=sha256:f28eaf771235b0bf03103bfd2113ed7fb4b4ce599b24352fe1f2fcd096082523

Observation d5407115-8935-4156-98a6-2901d3393e33 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.312162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.312162Z digest=sha256:1c993b92228acf76adbe35defc7310130a8c2120ab63faca5ee1861017679df0

Observation 556b0628-1098-4b0f-8d76-a2ba4395e922 · outbound

This paper cites Shih, Rohan Badlani, João Felipe Santos, Evelina Bakhturina, Mikyas T.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Shih, Rohan Badlani, João Felipe Santos, Evelina Bakhturina, Mikyas T

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.415277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:56.381071Z digest=sha256:b4d452ed9788a3ed07002a803f1ecbf955102659c68cf2bedfeb5e25a4a26fc4

Observation 37fd8c4d-8a2b-4ab1-87ab-b21dca27f849 · outbound

This paper cites Ditto-tts: Diffusion transformers for scalable text-to-speech without domain-specific factors.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Ditto-tts: Diffusion transformers for scalable text-to-speech without domain-specific factors

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.204599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:56.451386Z digest=sha256:2020997689bb53d94ac683c6f03c23eeef430106ba67f7a7899407c27c2f082e

Observation 433d7614-33db-4942-975d-0ae475faa212 · outbound

This paper cites Peebles and Saining Xie.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Peebles and Saining Xie

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.583065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.583065Z digest=sha256:5c3cb12593682581bfd84f4a569065f2de0ef830ef879749714149894f8d7b25

Observation 22457582-96a8-4eea-9a32-f5ecc442de3a · outbound

This paper cites Dmospeech: Direct metric optimization via distilled diffusion model in zero-shot speech synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Dmospeech: Direct metric optimization via distilled diffusion model in zero-shot speech synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.990413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:56.683947Z digest=sha256:96fc11669052fd734493c50319e944d83e2d8972a78f1d7018a934125b28a5c1

Observation 1e2b3cef-62ad-48af-b47a-1ddb4dcb323e · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.2024 IEEE Spoken Language Technology Workshop (SLT), pages 682–689, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.2024 IEEE Spoken Language Technology Workshop (SLT), pages 682–689, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.801253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:56.779894Z digest=sha256:aa836694b7f49c185cae7595e1192aefa43367e04e11efd5e7f8fc08f821a839

Observation a77996d5-2d3b-4528-9b14-7bd57feddcde · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.842801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.842801Z digest=sha256:6489ac8533127c4e6dc5fdab19bdac4bcbce516c5d01b70e4e1c8a6a3ce37cb2

Observation 235bff6a-1867-49b6-87de-bd2096abb91f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.932091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.932091Z digest=sha256:76fe5cb9a44bbaf712aee45d73747edaa5f3a64973de98dc55ed6f36459dfc9c

Observation 62d1ea48-ddeb-422a-afc6-9cd1cf666d36 · outbound

This paper cites Simple and Controllable Music Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Simple and Controllable Music Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.022040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.022040Z digest=sha256:22dbedb6bb28939c22d722cbf8b26270e96981b4262de1b1b75dac0a454ea453

Observation b5d23f41-e7c4-4aef-9075-65944122b465 · outbound

This paper cites Neural machine translation by jointly learning to align and translate, 2016.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural machine translation by jointly learning to align and translate, 2016

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.089308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.089308Z digest=sha256:04ea61b408fdbb8addb9ea75b1f3f24919e5cf30befecbd835bee86fa0081db9

Observation d3c1883d-8489-4d10-998e-13991ac28e25 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2023.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Roformer: Enhanced transformer with rotary position embedding, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.668512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:57.181904Z digest=sha256:59adcc626d4b3fb7b2385d8dea2118a894d0901163c8702f3fc4be771d6ceaa3

Observation deee603a-0ce1-4795-8030-32fc1e0db0d1 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.278975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.278975Z digest=sha256:742e996fb90c6a62b4c8e0254317e0c3a59692d98f0ab0db85a1900fdb4b8e6f

Observation 0df4e2a8-decb-4308-9f13-f80961a67480 · outbound

This paper cites an unresolved cited work.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.355945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.355945Z digest=sha256:ebdb63825c843706bf1224e4e945036ce7b608d991e8cbaef4a8df66ff90d553

Observation a4750667-7041-41f0-84fe-797038ed46c6 · outbound

This paper cites Smith, and Mike Lewis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Smith, and Mike Lewis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.434211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.434211Z digest=sha256:920d856913356b956f649022710cb15483e9851f2b11b363d098820840bfc7f4

Observation 032496ba-6547-4544-80ef-b8c1d4fbca3b · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.2024 IEEE Spoken Language Technology Workshop (SLT), pages 885–890, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.2024 IEEE Spoken Language Technology Workshop (SLT), pages 885–890, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.546516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:57.537109Z digest=sha256:b001950620f9ebcb46ce29f0cf02fb0381fae807b20f396be582fc7c905756b8

Observation 99fa98a6-c588-438f-b2d3-69705bb847a2 · outbound

This paper cites an unresolved cited work.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:20:00.387300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:57.594796Z digest=sha256:89ae3ae20035f506a0b18beb629c6d4014002576affc2cb30f76d54b7e31bd7a

Observation 0b8773f0-7299-4522-ab03-186b7534942e · outbound

This paper cites eSpeak NG: Speech synthesiser.https://github.com/espeak-ng/espeak-ng.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation eSpeak NG: Speech synthesiser.https://github.com/espeak-ng/espeak-ng

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.232493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:57.695586Z digest=sha256:ec5b77aca573a518ed44f9d73596a517d2d3eea9d8e29375949f86b1bc1381e8

Observation 8d541a03-a710-477c-a4db-f2ce42905748 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.780080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.780080Z digest=sha256:b979749bae24b029ac473625a211462cb873f72a90aa9a522445e64d1916d0bc

Observation fc9d4cbb-a552-4058-8c24-680bc8651cad · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Zipformer: A faster and better encoder for automatic speech recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.042837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:57.837727Z digest=sha256:0340c2122b50e7e02a46986d6e1ca792d1449b121df84d41ef3f729b0c619e2e

Observation cc47ff8e-9be7-4e0f-aaca-e4557fceaf5d · outbound

This paper cites Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.926601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:57.919375Z digest=sha256:922c514e5f002a655a4ff245e48a43a848abc1f7db9af0cd50c6cd376713f87b

Observation b8179d0d-46ab-486a-a30d-a4ba1206f51a · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Utmos: Utokyo-sarulab system for voicemos challenge 2022, 2022

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.759952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:58.017967Z digest=sha256:e5024177dfb5cc0fe47ca0b0a0ee08469233741f0b55294b5f3c7d7ac7abb8ce

Observation 7ffd8915-540a-457e-ad2a-feebe127b0e5 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:58.084628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:58.084628Z digest=sha256:69613b46e7d86838a7b95177aa201ae80a8581a87861eaf79766f242c611b015

Observation 9774d91c-d74b-419e-80da-0d25f97b1e5a · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16:1505–1518, 2021.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16:1505–1518, 2021

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.605196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:58.176788Z digest=sha256:d5cabaaec6c1f6d3d7f2d73107ee6ad374762be019a34a5de621468ab67a99c8

Observation fcdb7826-d3bc-4038-8c87-1cdf16295d18 · outbound

This paper cites Speech quality assessment.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speech quality assessment

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.455875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:58.249319Z digest=sha256:ba956a14b13ba74e04c546409173970e86234fa783ddfd6ec52263c5610c887d

Observation 8d91192c-2adb-4c99-bb69-82eb63c45c81 · outbound

This paper cites Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:58.330976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:58.330976Z digest=sha256:e8ba70299a2796a92afa6098d282383f16f20d6c5130ae7330af5931693c59bb

Observation bcfe4542-9f26-42fc-a7fd-831a2e9c5aa1 · outbound

This paper cites Fast and high-quality auto- regressive speech synthesis via speculative decoding.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Fast and high-quality auto- regressive speech synthesis via speculative decoding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.295479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:58.406514Z digest=sha256:89826ffeedbee913f47d64b3f2d01361671e181ee4626bea1ff7423cab2d4423

Observation 0fa78f88-7c67-4229-9457-1550633ce4dd · outbound

This paper cites natural” in comparative naturalness task with “similar.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation natural” in comparative naturalness task with “similar

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.140627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:19:58.490777Z digest=sha256:e4a5bd848b30ab1662f3869042d287dc34dad9df90de060b148f5fc4497a8489

Pith citing papers

Observation 4a70ccd9-9a71-43f5-a190-b1e16e566dd0 · inbound

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models cites this paper.

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T23:22:40.218077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:22:40.218077Z digest=sha256:82b8dbcd5f8cbab881181850ca8ba519767c10bbe5889ed7f5e17963de177133