Pith. sign in

Paper Citation Record · LEDGER

Natural language guidance of high-fidelity text-to-speech with synthetic annotations

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2402.01912.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01912 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:51.142199Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 381e5d4b-ee52-4d40-b293-c6eadfcfbea1 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.598879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:c72edbbad90ed2ea13859b734f81076b0327b7cd067002238959bf814c1e2c06

Observation 1222f66f-e60a-4da6-a341-18a4d700dfa7 · inbound

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits cites this paper.

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:51.142199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:51.142199Z digest=sha256:c0828b22065ce7753833a78003637417612e053ccd24b816ef329affff0fefc6

Observation db7c9d04-f6fa-4db1-9bc1-def38a7a896a · inbound

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations cites this paper.

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:42.812459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:42.812459Z digest=sha256:7c95a0c49a9a6feda4e3720e85953c264a379241be042cac6b8149bffa5b73c5

Observation 1d643197-ab5f-49a1-96c9-e56bb18d5d6e · inbound

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis cites this paper.

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:04.837891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:04.837891Z digest=sha256:2b383dc528405e5daa49e927ff4b030a831963d7cef8cde36d0423f91980ee71

Observation a0230389-f536-413d-b453-4a4d8395e27d · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.191205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.191205Z digest=sha256:1831be914c003d0be60726ab76bea627898565e041e2b446467e500e3189a72c

Observation de9f7812-7aa8-4c1b-ba5a-0257ba1e96db · inbound

Optimizing Multilingual Text-To-Speech with Accents & Emotions cites this paper.

Optimizing Multilingual Text-To-Speech with Accents & Emotions Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:29.502416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:29.502416Z digest=sha256:765ae994e48cf8f32ddc4b6c94ee71e297eb51f1a6e3bc6104ec6a80a4e57cfb

Observation fb386c66-e20f-43ce-a4ca-d2c323e05703 · inbound

MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications cites this paper.

MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:13.177749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:13.177749Z digest=sha256:03451b4ed2d13b0c6bd242059f9c8912d272cce575e30dd0ff6868d2b24279e9

Observation 5c2cf6a8-f131-4464-984c-ec5e78ae53f9 · inbound

Multi-interaction TTS toward professional recording reproduction cites this paper.

Multi-interaction TTS toward professional recording reproduction Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.456914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.456914Z digest=sha256:73afbf5bda5214577dae3a5e3ddf5c4ac63e65d3e54bd315fb31e11639b85190

Observation 9d1545d8-6f18-4bb2-92e6-f9bf768217e7 · inbound

SecureSpeech: Prompt-based Speaker and Content Protection cites this paper.

SecureSpeech: Prompt-based Speaker and Content Protection Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:07.484115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:36:07.484115Z digest=sha256:958a83ea4e99b87a8d403569c4cbe7cdd0419b65ea1221adb0e1a32fdc82ac5c

Observation cdb68026-84dd-4f38-be9c-6b9913761e75 · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:33.855401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:33.855401Z digest=sha256:432021dfd1941ca80da33b4df0222b52ca61f34ab55856e2f29420f0563575b6

Observation aea71d72-4db1-4369-9af7-bc31544e2993 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.214444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.214444Z digest=sha256:541e849d2eb07c865b1ba1504257f4c6f09af308b7228af61dc099fae2fd50f6

Observation 1ce65cf1-7f38-4863-bf96-b730fdb28a2c · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:19.629984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:19.629984Z digest=sha256:5f79b7177ac8b4de64aaa1d7222dd83789687d7d52a18147b3ae97f6d733a6fc

Observation 604d43fb-5ff3-4590-b857-a3b1e3df899f · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.453562Z digest=sha256:c1ceb1875282589eda890625bb8a07acca4097e8a6b0320afe8a50b512baddbc

Observation aeae0d3f-9346-482b-8379-4d91f7bfc692 · inbound

Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates cites this paper.

Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:04.179780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:04.179780Z digest=sha256:924c9ceb7416a8df4cec56619e3bd0b0bd89c0b43c8b8b67f3734fdca0a631b5

Observation 160ea297-49c7-498b-8f39-1db404dc9e29 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.136038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:2d8d7c30afaed558c1e47d0909a866ff3b99a0f5090c95b905a4aabe2e08c185

Observation 949c6284-1e5f-489c-ab16-b3c6e92f45be · inbound

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus cites this paper.

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:30:09.664146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:29:10.899659Z digest=sha256:c4d13ae72703311f6b4e9d10bee34a58ffe0c1d22561ecad720027910b5c81e7

Observation 91288b67-7835-4021-acc9-f75b037dc95e · inbound

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus cites this paper.

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T19:24:56.523051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:24:56.523051Z digest=sha256:fbf93498dc0b61df77c4f96a15479fc48a501b33dfe82500b13e95d168dd60a9

Observation 5ff90eac-0152-49b9-a645-8644c62c3689 · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:04:47.244302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:b18365972e1bcf952a50539fb9d6ef02ac27ba25675aa984a482a5552b11879c

Observation 4b13ef46-b70c-4ac7-87e6-6a8f1720f315 · inbound

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation cites this paper.

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:50.107872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:01:06.094276Z digest=sha256:f9f3d8d4ee0e2d1e63b752a9a8d9b0556e63b6c893812abbc7237d75008285e0

Observation 8758a34a-8d0e-4f9b-b29b-0d694d21bec0 · inbound

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment cites this paper.

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.324177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:07:50.243903Z digest=sha256:2a123ecb7ef17ccdb05e4d1fc2d2691b63a60b38b0067f677b78f4107a616c32

Observation 12c7d0ed-e83f-4cc2-907b-800c1af311b4 · inbound

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs cites this paper.

Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:18.771185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:20:35.333172Z digest=sha256:38ba6b7b871e96afe47c4a092d2ec443a2aeaca65a0d9dd3a9f64a2f19380460

Observation b0b47817-d91b-4f24-b4e5-bb1081c9a2ee · inbound

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models cites this paper.

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T00:18:36.390365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:18:36.390365Z digest=sha256:d0e7a7833ff4dbad0334de7addf4b6523eb81f42162821185f11cbc94050fb35

Observation d9f6ae9c-89d1-4459-9a1f-92ba7f33e61a · inbound

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech cites this paper.

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:27:04.903956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T23:53:55.385445Z digest=sha256:9d43998a1d932153890cd0aedd2c86101aaafb18843ec6e6c4327e394f9cf092

Observation f00501ae-1501-41ad-b754-7230a5da2bd1 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.727005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:2c0691b96d97a27f3743058199614726a91ee60169e242cbb6b30c9dceedd7eb

Observation b681c0fc-bc11-41d0-9b8e-f75c196e9e85 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.790821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:6bea4fd1fd5dbf81cfe85b9898f3f2af605f8fc2009413ae15f3524e8908dd28

Observation 3ad237b5-80b3-46d7-9f7a-40afd9551ec8 · inbound

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech cites this paper.

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.577822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:09:33.605232Z digest=sha256:e946f63194952806e483bd8f8f0ac3683c829f3da0a723c2035cff8487a29e01

Observation 4655fea6-e4ec-4da8-8bdb-f9354cff3497 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.552835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:b9c8bb17b590dcaa3a6cdfd52bf94e18a93dd76785b34e0ef02ee198221c603b

Observation 7e2215f8-42e7-4a5b-9347-715b6a74b32e · inbound

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer cites this paper.

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:46:13.750265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:46:13.750265Z digest=sha256:57be105c6bc482ba967bc673c0d91a1a7ac5b3bf31d0628d7a707530d900cbba

Observation 69f255a2-f5d5-4f26-a467-9c911d7b2c6a · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:26.024843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:26.024843Z digest=sha256:d19cc16a898ada860930570e56533c2431b29bc39368d5645452318123949621

Observation d33c18b7-0caa-4bd3-9a9e-2e5011ebca0f · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.524756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.524756Z digest=sha256:4d9c0e229965c4f56c7fdef87fb4a9126301c136e2fff7e5120f6604fa9806e7

Observation 8851c956-d783-434b-9400-67f0ca5946d0 · inbound

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study cites this paper.

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:58:58.620852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:58:58.620852Z digest=sha256:001f6570725b010552e4e0fa9ad58e66b0ad9dd48d6f90aa1d039d1906228bd0