Pith. sign in

Paper Citation Record · LEDGER

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

As of 6 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2510.04593.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.04593 v3

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:29:37.769261Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T23:22:40.218077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12a057a4-a5ff-44ae-8397-fb0cae38e85f · outbound

This paper cites GPT-4 Technical Report.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:29.467286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:29.467286Z digest=sha256:05fdb03f9bd270273c9076ab324a1f64367ecbf7479c84a7b7eb78fae04ce7bf

Observation 9f9bbc3c-dde2-4b00-a95d-8f2a57a21ed7 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:29.545439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:29.545439Z digest=sha256:e0b6715e0057b6060f995b2e73495bdcd408264d2e48e6e689c111c699d2e79a

Observation ad9c0fc2-96f3-42d0-8681-5988cbee1899 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:29.658196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:29.658196Z digest=sha256:6561ad8cea5c1341d4582c83f88179a26ea65df286642c5e9bc49b372712072b

Observation 4cc58584-f19d-4b4d-8286-78ba2c0b84b0 · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:29.783850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:29.783850Z digest=sha256:f7c800a35cc6cf2bf0037c624a7bae786ab4aa62706822dcd7119d6e3c719291

Observation e611af48-f6d8-4ec5-a2f4-30d2688536a1 · outbound

This paper cites Qwen Technical Report.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:29.906357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:29.906357Z digest=sha256:6ceb7ccc3f005025c3c10b2f122dcc1b03ed2a6ee384c2f86a4e00edfe1c9f37

Observation 326f4580-d5aa-41b9-b654-8a60e83ef953 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.053883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.053883Z digest=sha256:efeeb0f611373f58e54df24cfb5f2c0bc08f3282860b44b6d8cb3c728205dde6

Observation ca4f056e-b5cd-4220-ba98-417bf6a7d051 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.168637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.168637Z digest=sha256:57649c24c4d051c3abffce810aba69c8b56ddb341acbd950568bcda5ea42c22e

Observation b5b976f7-7195-4694-b649-e9204fc094b6 · outbound

This paper cites D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.306366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.306366Z digest=sha256:3d975cda1c191798997047cf48e0a33e7acd8447f96ed66dd78f711ae5801a5c

Observation 0a66343d-5dab-48a9-871d-109ba2cbdaee · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.396377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.396377Z digest=sha256:0e5399cae84d32d2b5253cecd773150f429024ec0697bebd31d3f95df2c921e0

Observation 73ca40a1-c953-4676-b024-e5b556234cfb · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.523441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.523441Z digest=sha256:d04ee59961d520a4a53feb643015b5149da1dc0663a60e12b2964b14b3d9f1fc

Observation be8b3450-5551-4e9d-97c3-ecd45d3e549a · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.654795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.654795Z digest=sha256:a10d06a159a8b13480094b4908fc6a869d0ef9001095f0e97dd6320496a350d0

Observation 81f83c4b-45f8-43b8-9d37-016b0153cecf · outbound

This paper cites SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.772401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.772401Z digest=sha256:e3ad98c3410117218f37383459fbd92fcdd37298b1f12d120d73c1eac4c42173

Observation 33ab5622-c8eb-4cf2-a821-eedc99c37674 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.853473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.853473Z digest=sha256:415b162fef5852adfde3b282d505f53768429880db9df6419deee6a976b33155

Observation 0acea275-804f-4114-b4a3-0986b44f2cc7 · outbound

This paper cites SpeechNet: A Universal Modularized Model for Speech Processing Tasks.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechNet: A Universal Modularized Model for Speech Processing Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.978703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.978703Z digest=sha256:ddcb17d41477901411108a8e6262b83353274674a6f8c5bb9627fe7da4b19886

Observation 721520c4-c719-475d-9685-3c36740ee176 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning for Speech Recognition.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unsupervised Cross-lingual Representation Learning for Speech Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.132461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.132461Z digest=sha256:790f1f6f389d0f78933a0bbd92b557b46ee736bfeb50a46c1a8f250e154ef7c3

Observation 60724a84-ac48-432f-8aa8-d76d4d74f85b · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.196577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.196577Z digest=sha256:96da356f2abf28a8dead92b75d42fd9d8bc0736782e43f8cc022d6db8ae58388

Observation e29b90aa-f9d3-471a-b68d-20275a3b971b · outbound

This paper cites Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.322714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.322714Z digest=sha256:34c4c961b647aa12ff2060af8c3aff9e92d1d193fcfa5eea8bcec8cb586146c3

Observation 0e0912cc-2990-488a-a4b5-b18f5fb2bda8 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.474390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.474390Z digest=sha256:c2f59b660dbcc0b6f1c56035b15836febb3b15e59f43bc0043a3bdb202867325

Observation 16ccd22c-94b2-4b0e-b467-4436db4b9cf0 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.571249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.571249Z digest=sha256:e70a640ae77e59423ae9f06b009cdc611d27b8cb12bf9112f02a44a8cb2337ba

Observation e3c490bd-0831-4eb7-b0f5-b81e63db752d · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.705443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.705443Z digest=sha256:696c0e3fbb387a80e9e0b4bb5461608d55c7fb90b99295e8c99521daf8cf6f81

Observation c517b810-88e0-407e-b179-ac0b573fca67 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.898662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.898662Z digest=sha256:b6f828754a4b4a1e72493199609f2036a0e38cbc5800d8947b65877706356a22

Observation 6dc1f721-c254-40e8-bbb1-550d02063719 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.996283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.996283Z digest=sha256:7656517a4b957866986961527c2ad786a12229637c74a4ec3cd18b4a6579cdbd

Observation 05ff4adb-62b4-4ea1-81ae-d5a2893e713f · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.118457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.118457Z digest=sha256:11bde97f7067888456bb2ff137d54af67abef233ca7f176982345cb932053473

Observation 8766751d-7fdd-461c-b2cb-2ca38762749b · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.222197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.222197Z digest=sha256:f4aadc3f8899dfa885381490a80c7acf21151b09cb0d3ced584c547d184c81f4

Observation bb09c3f4-f52a-4607-b5d4-37f811f7243b · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.338248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.338248Z digest=sha256:06f992408d7e98634c03d2cea1371efc691e28450bf77ec61b001fc72038c16d

Observation e26d98a2-a154-49b9-9b21-89d5298ec760 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.431418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.431418Z digest=sha256:744f1a1bfc89353f9e273508a4ab48948eadba02a311843c2ff566bcda5899d4

Observation 9aecc85a-55cb-424f-b3e8-f2e9e24b0887 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.540104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.540104Z digest=sha256:e5df70407ec321aa57de3e33befdcfcfc61a776dbb62aaae5908e52bb86fc649

Observation 67778385-97e8-4920-b666-a32c0c30aee5 · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.656117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.656117Z digest=sha256:1a35c8097aaa0291f298a73d66500cc587f9f2a9951bc206cb1f9677be17f992

Observation f3acd8b4-2027-457e-9e0f-24bc337b7baf · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.767944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.767944Z digest=sha256:51506d98d074a8ec05607574f36c4f4318f7eac855294493d22ba076f63dc144

Observation 735d6574-8dc5-446b-828b-1729141b3a67 · outbound

This paper cites H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.871599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.871599Z digest=sha256:b38a138d61a1ca544094d166b33d97939a7ecdcf66560093c57476037d094c0d

Observation e4ed716a-13e0-4bc3-a677-661395c9a330 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.995304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.995304Z digest=sha256:f8838db423746517fc94a4e1b09c8ce1d1a6e646377aeb7dd273640612227411

Observation 78e4abf9-eb85-44cf-9e23-1d155d9b1699 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.186280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.186280Z digest=sha256:c736738f7a718efb4c35f5891c7d815edbd06383306eb095a8b7ff0b6de604ef

Observation 1d62f73d-0a48-4179-8bfd-70020471dff4 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.361117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.361117Z digest=sha256:99723c1dc00008677dca54e2e223d37c56624269a02e0c933a5686a80d395ee6

Observation a84a6d7e-6345-447c-8fbb-e3bf3be8fac7 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.451397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.451397Z digest=sha256:e1c050df6ba426f05debbc36f67187ba990e13c2d0f04302f305d6c7a3bd6c38

Observation d77c3ef5-4916-47f1-9dac-172eeb7c6db4 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.561535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.561535Z digest=sha256:7d24d906c9ec3f20cf228061b2df1465200329c97af8995b9975ca066756570f

Observation 373248fe-be09-4514-ab30-ec207ad538f0 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.696265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.696265Z digest=sha256:38fad697c6a0f4c6ba0e074425308b8fbad0024b67aa826ac8ec2e327a7f13fc

Observation 0013b167-b28c-4451-acaf-ce1edadd2b33 · outbound

This paper cites Flow Matching for Generative Modeling.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Flow Matching for Generative Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.754234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.754234Z digest=sha256:f59afc84e54f62628fd97781c6066338b933b0114bede49f3a9c280a84273262

Observation 9ba08cfd-6e72-4426-9c96-74149d070b56 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.883026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.883026Z digest=sha256:ee22a0794d3709c77f274e7eea22911d6edf7ce275355891c9ef40335c820e0c

Observation 5941f89d-d6c2-4b1c-90f0-582dff6b0b01 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.994476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.994476Z digest=sha256:9080560ced22a8f15d8d8c50bc46de3b242384da714dc2310ca635dd5dd6a1c0

Observation 57f393f3-6b97-4470-8eb0-1ee4d4ea2ec9 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.131364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.131364Z digest=sha256:6e63eb851b692522e5070cc1a2145f4421e572d70b8f43dc93f2c11d94a44bf7

Observation f0f964f5-1a8a-4393-b402-e7f00e95610b · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.295228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.295228Z digest=sha256:de299d70fbcc7926116015246f900d84b0ea1af54e3d1abf3816edacb0be6d32

Observation 824117dd-ceae-4d70-8023-5e8857825b2a · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.414497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.414497Z digest=sha256:7063d13dccb300d94d24952ed66c6c672d0e03323d9b25b524b8dc04e292f03b

Observation 53c9b8f7-de86-422e-b20a-8907cae3e636 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Autoregressive Speech Synthesis without Vector Quantization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.568989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.568989Z digest=sha256:eeb37da58c204fe7dd4525fdaa50b6c30d15f1e9b3a30888d376bd3f988adce8

Observation c3e95c70-006b-4ae0-adea-c2799aed0a28 · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.725106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.725106Z digest=sha256:244b1fb965923983f76a570edb75a87811cd07593fc946d15b78a9dbb4524ef5

Observation 538ddc52-61fb-49c2-b2b9-d6b19698a0fd · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.824864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.824864Z digest=sha256:3522ae8bb344e106919dc6928a387e29b4a881178d5fbd6ad2875d3eede19c18

Observation d927ba39-53d1-4da2-9d32-d96732e04cc0 · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.882310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.882310Z digest=sha256:68275d897d289f68f913c3c45cfae58cbc1d3931ebf17e8091f695a65c3a6830

Observation 5e4547fd-349e-462e-a787-e147ed690b05 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:34.954596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:34.954596Z digest=sha256:789f5807713cb9801b106d8f724d82e0eb0ca8968ca97a0252144e3a108c0f94

Observation b88e8dbe-4465-4fa1-adda-ea000ed42111 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.013183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.013183Z digest=sha256:a84a7dac99b5f4e0217372d85c553e7c26719c40cc20826dcf65a4360746c31e

Observation 85a00b91-637e-463e-8539-58ba0c53a2ed · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.095216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.095216Z digest=sha256:4f9850d3202a8b577726aa9554ad94ffb3de2a11f67902b084feb8e9d222426a

Observation 5aa886f0-7a3f-46af-a0b0-bdfb212737bb · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.214098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.214098Z digest=sha256:c278c640dc156ef41afc33c25aa8b95563be967d82cffd4c6572465e2ab8d7be

Observation 34b8e233-0184-4db3-8a07-5e5d3014c484 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Score-Based Generative Modeling through Stochastic Differential Equations

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.347050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.347050Z digest=sha256:b094b83a3d87b6919727ee475567feac1202a3c5d1bc50000a28612587e99d52

Observation 8da660ea-1bfd-41f1-a7f3-410d61674243 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.497378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.497378Z digest=sha256:e2b2fa6c5dbce4684a01e2b4329e1d881bf775d4b7da2341faf994266c307498

Observation f7ac93b2-b13d-4169-bc04-40c1e4174723 · outbound

This paper cites OpusLM: A Family of Open Unified Speech Language Models.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models OpusLM: A Family of Open Unified Speech Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.787283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.787283Z digest=sha256:fdbb2a9148b8abebd4fa58690e03ca8ebc5f437317109364d3ca5dc16bd51f70

Observation f1829538-7131-4a4e-aa04-77826f46978f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.952638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.952638Z digest=sha256:4ce7184b050d2d855a60ea317e0520a3d17311f5266d19dedc7774833e34ee0f

Observation 2b3821e8-c4cb-4589-b362-5bfc4485c04c · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.075318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.075318Z digest=sha256:b83070a16c01c1869d20f9fbd3a18b973a2db4ecdc3a42a0667d4c77899b3fa5

Observation a9dd9d83-9aca-4892-bfef-27d9cf13f29b · outbound

This paper cites Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.222695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.222695Z digest=sha256:4fe91a2401e3b5c4ba258b8513757812d2ad7c788c43bc92c268e148f3e0637c

Observation 71e575a4-b2e1-4939-861f-3529f1edb453 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.344998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.344998Z digest=sha256:d517b19829244ba29e1ef625561fb1e3297af9b7cd43b91c9eea28565938efe0

Observation 06f231c7-7dd5-4e17-92c0-188fee700bf8 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.515375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.515375Z digest=sha256:33944fe45068cfac5e3d1c3b8ba9a6e0d45387b5e60a0c5a19b8a9d9b4f7697c

Observation f634f89e-bfc2-4b66-8aa0-d610f9a94a59 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.628124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.628124Z digest=sha256:37b254e7a7574bcccd94d19744b8e71aa6daccf24aa6ed604c2e19730a16691d

Observation 322652e2-702e-4bb2-9f40-afbf0934ccc1 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.756155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.756155Z digest=sha256:7fc1882350907ceda90a05415f30f09036bcde99360075cff5864c8029bec4ba

Observation a6ccd059-323c-4831-b26c-401f61f95b9c · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.870189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.870189Z digest=sha256:c32d998440267e927843956555e1e1a61f75b787b03efdee3e7b72848b2f7f69

Observation c6f16ced-7f86-4346-bbcb-62104c5b69fa · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Zipformer: A faster and better encoder for automatic speech recognition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.020821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.020821Z digest=sha256:7f2f84f7ba2cb9da51ad88265c34179c66f24b67cbd2937bc6af9795abbfe54f

Observation 5393d5e1-bc22-4581-a2f1-2f55b9d50dc0 · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.202373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.202373Z digest=sha256:a390119888e48cf24964b456f7dda6e71a25c40c1fd41e515cb3baf7366ef2b2

Observation a5daa850-7551-4dda-bb63-e5d0f2544a90 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.345020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.345020Z digest=sha256:045b662611ae19d9a8ef3471be0f92be989ac18a5f25a7a3db7d8d1519b87f1e

Observation 115c5972-fd02-4904-bdad-7c607d74e07a · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.454875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.454875Z digest=sha256:575e201af681c8ac1dfcc92feea2764e61400ac89abe6a0ebd5071942e700ed4

Observation 13a75cfc-a4ff-423f-baa5-31946f54146a · outbound

This paper cites an unresolved cited work.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.539911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.539911Z digest=sha256:2c3c7ab9864ad4abc3b47dccd91ccc13bc56ec5fc5cee3b7e32bf0c62505e463

Observation 47b089b6-57a7-46ab-8b68-20eebafca2cd · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.658307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.658307Z digest=sha256:8f006fe1f2d244701b9628d6810cbaa7385e781d0a1c6ac3418584b40b9a16cf

Observation 5cd6c1fe-12f6-46d5-9058-6cbcc06f8e41 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.769261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.769261Z digest=sha256:1ccc24f8ab0dcf47c5936a69656c183679646e1563742e0d6050844ec4321365

Pith citing papers

Observation 5f63ffd4-b4db-46af-bcd5-efd1c1592850 · inbound

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models cites this paper.

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T23:22:40.218077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:22:40.218077Z digest=sha256:33b7a63656ff89a8cab2eb7c7b8914e7f04af0338ed7702991870721f6c58a6d