Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:29:37.769261Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2510.04593.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:29:37.769261Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T23:22:40.218077Z
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 12a057a4-a5ff-44ae-8397-fb0cae38e85f · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9bbc3c-dde2-4b00-a95d-8f2a57a21ed7 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad9c0fc2-96f3-42d0-8681-5988cbee1899 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc58584-f19d-4b4d-8286-78ba2c0b84b0 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e611af48-f6d8-4ec5-a2f4-30d2688536a1 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326f4580-d5aa-41b9-b654-8a60e83ef953 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4f056e-b5cd-4220-ba98-417bf6a7d051 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b976f7-7195-4694-b649-e9204fc094b6 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a66343d-5dab-48a9-871d-109ba2cbdaee · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ca40a1-c953-4676-b024-e5b556234cfb · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8b3450-5551-4e9d-97c3-ecd45d3e549a · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f83c4b-45f8-43b8-9d37-016b0153cecf · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ab5622-c8eb-4cf2-a821-eedc99c37674 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0acea275-804f-4114-b4a3-0986b44f2cc7 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechNet: A Universal Modularized Model for Speech Processing Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721520c4-c719-475d-9685-3c36740ee176 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unsupervised Cross-lingual Representation Learning for Speech Recognition
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60724a84-ac48-432f-8aa8-d76d4d74f85b · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e29b90aa-f9d3-471a-b68d-20275a3b971b · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0912cc-2990-488a-a4b5-b18f5fb2bda8 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ccd22c-94b2-4b0e-b467-4436db4b9cf0 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c490bd-0831-4eb7-b0f5-b81e63db752d · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c517b810-88e0-407e-b179-ac0b573fca67 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc1f721-c254-40e8-bbb1-550d02063719 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ff4adb-62b4-4ea1-81ae-d5a2893e713f · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8766751d-7fdd-461c-b2cb-2ca38762749b · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb09c3f4-f52a-4607-b5d4-37f811f7243b · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26d98a2-a154-49b9-9b21-89d5298ec760 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aecc85a-55cb-424f-b3e8-f2e9e24b0887 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67778385-97e8-4920-b666-a32c0c30aee5 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3acd8b4-2027-457e-9e0f-24bc337b7baf · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 735d6574-8dc5-446b-828b-1729141b3a67 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ed716a-13e0-4bc3-a677-661395c9a330 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78e4abf9-eb85-44cf-9e23-1d155d9b1699 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d62f73d-0a48-4179-8bfd-70020471dff4 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a84a6d7e-6345-447c-8fbb-e3bf3be8fac7 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77c3ef5-4916-47f1-9dac-172eeb7c6db4 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373248fe-be09-4514-ab30-ec207ad538f0 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0013b167-b28c-4451-acaf-ce1edadd2b33 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Flow Matching for Generative Modeling
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba08cfd-6e72-4426-9c96-74149d070b56 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5941f89d-d6c2-4b1c-90f0-582dff6b0b01 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f393f3-6b97-4470-8eb0-1ee4d4ea2ec9 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f964f5-1a8a-4393-b402-e7f00e95610b · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824117dd-ceae-4d70-8023-5e8857825b2a · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c9b8f7-de86-422e-b20a-8907cae3e636 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Autoregressive Speech Synthesis without Vector Quantization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e95c70-006b-4ae0-adea-c2799aed0a28 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538ddc52-61fb-49c2-b2b9-d6b19698a0fd · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d927ba39-53d1-4da2-9d32-d96732e04cc0 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4547fd-349e-462e-a787-e147ed690b05 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88e8dbe-4465-4fa1-adda-ea000ed42111 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a00b91-637e-463e-8539-58ba0c53a2ed · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa886f0-7a3f-46af-a0b0-bdfb212737bb · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b8e233-0184-4db3-8a07-5e5d3014c484 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Score-Based Generative Modeling through Stochastic Differential Equations
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8da660ea-1bfd-41f1-a7f3-410d61674243 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7ac93b2-b13d-4169-bc04-40c1e4174723 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models OpusLM: A Family of Open Unified Speech Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1829538-7131-4a4e-aa04-77826f46978f · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3821e8-c4cb-4589-b362-5bfc4485c04c · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9dd9d83-9aca-4892-bfef-27d9cf13f29b · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e575a4-b2e1-4939-861f-3529f1edb453 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f231c7-7dd5-4e17-92c0-188fee700bf8 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f634f89e-bfc2-4b66-8aa0-d610f9a94a59 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 322652e2-702e-4bb2-9f40-afbf0934ccc1 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ccd059-323c-4831-b26c-401f61f95b9c · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f16ced-7f86-4346-bbcb-62104c5b69fa · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Zipformer: A faster and better encoder for automatic speech recognition
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5393d5e1-bc22-4581-a2f1-2f55b9d50dc0 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5daa850-7551-4dda-bb63-e5d0f2544a90 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 115c5972-fd02-4904-bdad-7c607d74e07a · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a75cfc-a4ff-423f-baa5-31946f54146a · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b089b6-57a7-46ab-8b68-20eebafca2cd · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd6c1fe-12f6-46d5-9058-6cbcc06f8e41 · outbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f63ffd4-b4db-46af-bcd5-efd1c1592850 · inbound
Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.