Pith. sign in

Paper Citation Record · LEDGER

Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2411.01156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.01156 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:11.377969Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.706029Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 14708426-a94d-4b89-b445-2ee36eab8058 · inbound

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch cites this paper.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.514869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.514869Z digest=sha256:ef4a0a06c1d7cf14ade93dd63a492b2016876f453cbff5d8df8c45dd1a2be71c

Observation 9c711ee6-7141-46e4-a003-909bbdfd5f41 · inbound

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling cites this paper.

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T10:37:03.657296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:37:03.657296Z digest=sha256:2f1c273467195f50f85173e12a43361ae9772b52c77b6d6691c6fb0cbbec18fe

Observation 56fb5224-3e15-411f-9401-d2b9b29b2629 · inbound

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System cites this paper.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.501225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.501225Z digest=sha256:966f53de0265e3a1bc162e2863e50c700e8d467d811a5d7f929d03e94de39ff0

Observation 6cfcd3fe-219d-4502-a5fa-f0bcea9aa9f3 · inbound

Position: It's Time to Act on the Risk of Efficient Personalized Text Generation cites this paper.

Position: It's Time to Act on the Risk of Efficient Personalized Text Generation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:07:36.486045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:07:36.486045Z digest=sha256:796478d7f2aca17bcdf3cbb3314df64b6d055967a7049a8230ae6a1e98ddd1cc

Observation 861424a8-255a-49e2-a223-0064ca68b613 · inbound

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget cites this paper.

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T06:05:11.377969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:05:11.377969Z digest=sha256:ec05f7488040acfce4b8560a0d302741a40227403bfeacb803130e46319c98df

Observation 9c0d5022-b7ef-4d7b-bd24-85b85705c8cd · inbound

Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning cites this paper.

Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:22:08.759883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:22:08.759883Z digest=sha256:7c358028e6a17b424c89d5764abc30987ed0da45d0521c5ace5cf3f05c39c7e1

Observation e6851a7f-c87b-480b-ab04-3ceaf0e81ce5 · inbound

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis cites this paper.

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:07.877528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:52:07.877528Z digest=sha256:9844d6eaf908752d36f7dd95f5663b14cec9624dfcf2e4d0f916dcb789519b88

Observation e8ed7c07-7ba5-4896-856a-ab21e0eb11b6 · inbound

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech cites this paper.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.165930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.165930Z digest=sha256:60029425cfd2ecc52f9ca5947d931819b65b4b31ee384447911e0be586d3d6de

Observation d8383ff8-ddd2-4fd7-8b0c-546f08809e16 · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.813372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.813372Z digest=sha256:a5e98cf2c8b0dab8c70cdb9c99fcd6a3d533177b0b3bf4df0bba9f39d8c5db1d

Observation 1bafbe8f-77ba-41e9-bc48-584777613c6e · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.316201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.316201Z digest=sha256:c7ea6f82a80196e865f6c7729b1b1e57898cfaf34bef95b53146c398aea89648

Observation f620d47a-d187-42f4-93e0-6e13081bcf3c · inbound

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis cites this paper.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:28.199944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:28.199944Z digest=sha256:d051f89cab228d7db967180eb344b31f5e37919afdae8547a1212a53a3922a84

Observation 47f8a2ca-2413-402c-a719-cf2fe2236001 · inbound

Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation cites this paper.

Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:33:51.609621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:33:51.609621Z digest=sha256:e9bf60868263e3654ffb12f0a55edf573399ef67fcec7d532e6b1621c6817488

Observation be5a0f09-b588-4d3b-be6d-6f1fa0693359 · inbound

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges cites this paper.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.398868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.398868Z digest=sha256:2125fdd2d05eea8a3329233490bd55840596c618bc63128cfdaa82b032f34f97

Observation f5804bb3-fddd-4e20-a57e-7810b4c9c2af · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.197119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.197119Z digest=sha256:31474b96a13a9e52c528083320a54d65991d8b888cb34d70481e815b31178421

Observation f21bb22d-6379-4d43-9121-f091d7b085d5 · inbound

Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models cites this paper.

Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:51.601861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:47:51.601861Z digest=sha256:f3f740a3ee51b4d0e185d6ff24088a20c74304cf35a346de4c61d2861cc9e954

Observation 80da568d-b212-4b0d-b9dc-962443201971 · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:19.352331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:19.352331Z digest=sha256:3760679387883f81dbc54fb34e97e1eb7e59a5f8e8030454fe0fd07f4f449337

Observation cada4e6f-d952-465e-8a7c-276f224b2f62 · inbound

MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening cites this paper.

MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T16:47:30.226454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:47:30.226454Z digest=sha256:b4da5d90fab83c63f83982da521f7e53ca039ee352e1367251d65c7415e2078b

Observation 5d7c9d51-f42b-4e92-8c76-cb2f56566c67 · inbound

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration cites this paper.

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:36:37.614517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:36:37.614517Z digest=sha256:a86fd800a0551a75cb490f45300c1b1cd129fa5b90a6bccaa697e45b3337185a

Observation 2b11f19d-93af-49a8-a8ea-d4fcb7aed126 · inbound

DarkStream: real-time speech anonymization with low latency cites this paper.

DarkStream: real-time speech anonymization with low latency Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T06:01:28.489246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:01:28.489246Z digest=sha256:52ebed45141a6d1017578c398714d82eb1820fe3f564995145eb1d8000ebc208

Observation 38140346-be0d-4da8-8620-30bce8c6ea37 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.747134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.747134Z digest=sha256:7673eb9fac5198a814902053ecde404ea95f8cd72ef1d302ff97a67d4ea82cac

Observation 9d3cb53e-2dbb-4ef6-8259-6ad3e65b75d7 · inbound

HISPASpoof: A New Dataset For Spanish Speech Forensics cites this paper.

HISPASpoof: A New Dataset For Spanish Speech Forensics Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:38:35.620052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:38:35.620052Z digest=sha256:b206ae6b89f87b8deebcce82fcfc3c0943e9600900ba4a5a383aefdedc9c3682

Observation a836ee88-0739-48bc-ad6b-83f6c2c22224 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.256180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:6b0b67672493cbbaa0312f22b9a5c8c7bb94fd0e2ff2af2527f530ccef588d97

Observation aa1182b9-20f7-4b6a-a44f-c777dfcd9268 · inbound

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection cites this paper.

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:02:23.354924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T05:01:57.044869Z digest=sha256:ef46bc2d0f7aa600d822f36c1d6d6f6a1a43157798e20e52a26acb44eb2ac44c

Observation d7f06d45-5136-43bd-a656-a0dbfa48de7d · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:24:56.131581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:5c274d799f6712f951ab5f48951c69dbfaf95c7bbbb7fcc12a4f02303b5d5f33

Observation 8f943109-eb0b-4c6d-ab35-2049e1e068d4 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:14.061855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:14.061855Z digest=sha256:6ae165c58593d759df882a87da0d4955d62c659ec4dacacaf59367a7c5d28cc2

Observation a0fb93cd-1be2-4fb5-b53c-ca71ed9541f1 · inbound

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus cites this paper.

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:30:09.642109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T16:29:10.899659Z digest=sha256:60f298deae83572bb5ed1af73e0e9c2db19305656e41ac6bb2b155094d422f34

Observation 074b5629-107c-4113-aa8c-b87eb816ccde · inbound

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus cites this paper.

When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T19:24:56.576910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:24:56.576910Z digest=sha256:f98cd042f33df256ddc4fc0320b07cf75b33dbe7c52f9fb4b6b8a1e8dd6d2c0e

Observation ca7150ad-4923-4f01-b7b0-f6e190f72502 · inbound

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis cites this paper.

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T18:56:10.569061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:56:10.569061Z digest=sha256:91dda3959e300210a67d8c571fb794cb2ba9043b106a4c301238217273c990c1

Observation bf9ea63e-d386-4fff-98af-102418eae34c · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.897141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:c01abc4d0d86130aca6681506e2a55d4971d6d42f192021f3b3105c32faf68a2

Observation 9bf6fc20-e0d0-4bbf-980b-5129030961ee · inbound

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative cites this paper.

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:46:12.407558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T02:27:05.776290Z digest=sha256:8750f58d790c38badeda20fd2c583510c3d7c1a9fa3cd24c5c765400d01e1812

Observation 47cbbba5-e1a2-460e-b72a-137d87a9c5b3 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.134673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:5d10c58a119a079938d051416fe99f16f108116a640be81620a4775821ae0141

Observation f348cb99-400d-4a35-b18c-dc165d39c950 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:cc359a971434381a8c03ff00d2ae182dd1309a3c136409b59c2d95239bbbe294

Observation e0349b6f-db83-4856-a82b-da36a5808452 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.969453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:41e57348032ee9d3c3811bb55055bdeb0fe9b331606b21f96637631d14cbc8e1

Observation bd7d2443-a996-4c83-a19f-54f5998998e0 · inbound

RTCFake: Speech Deepfake Detection in Real-Time Communication cites this paper.

RTCFake: Speech Deepfake Detection in Real-Time Communication Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:26:17.877875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T05:29:45.895667Z digest=sha256:9ae45f6b1d89a112e5c854a31e5ef4c4267c661b9a077da8084f7691b180dd4d

Observation e93d7b81-8cf5-4449-ade3-78e53c56bfe5 · inbound

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data cites this paper.

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:14.532711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T07:46:23.931059Z digest=sha256:adf34c3ee01ef42aa29623a08cd85ab5de3a5992fbe5f596504c0a7f1fec7f0c

Observation c12b94b0-311e-4bcf-8a7e-1a18febcd28d · inbound

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection cites this paper.

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:06:04.299894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T14:04:52.065878Z digest=sha256:a8d5fd9a40679dd11ac3bb6f44edaa9c1811f649f988a09362e28d344712c24a

Observation 8f7d2581-11fe-4319-930e-df8a9269f01e · inbound

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling cites this paper.

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:12:13.862135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T04:09:48.737629Z digest=sha256:41e4458e6eb56d96bc388392e4d146e57638a492abfe57897fc7849b24f164de

Observation 45966ad9-4f26-4fb6-92e8-e5cb3328bc24 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.104624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:2ed4be8b85d72f76bcd4edc4008718bb74f567b99e34be0a2af0aeb2bb6a3fab

Observation 5772ce05-c8ce-4ede-9068-4f767d606373 · inbound

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors cites this paper.

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.746809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T21:08:39.636966Z digest=sha256:693f1a7e28a8f2eaa4bc252008b4b890048114dac6b21bb95f68188980aa341c

Observation 456d7a1a-912c-428f-a5a1-474ab12e7993 · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:36.141295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:f56fd307fe7aea996a340feec15dc48f8d41b48f671ad04df1d0251af2962abb

Observation 8e9be2d8-7655-49ef-870b-50849499dc90 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.707667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:ee0e6a9fdba2bf43d9237a7e0c5bbb2ed376c6aca7d0731a23087ee04bb01d8a

Observation cf62a5da-8fb4-4b81-92b8-9318cda5d810 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.130806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:b2c9c242c7ad12312a94614c547f19687b62a18c100d719233ae742b683d8609

Observation 0efa5e6a-eb1c-4f4e-9f7b-b7be0635e344 · inbound

Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption cites this paper.

Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T17:26:24.553591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:26:24.553591Z digest=sha256:7451e96f4b024ae4989e6f985d5f30bcb45560d03e5c60c1fcf226f900314795

Observation 96e8430a-ea47-465d-8339-2ee17845339c · inbound

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs cites this paper.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.509915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.509915Z digest=sha256:cf417b882c4e691e71a989ed5a876127f19a2a11ad34e6c274bd73536d6b5d59

Observation 6310a5c4-2395-4954-bb9d-84ea5c281897 · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:38.049516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:38.049516Z digest=sha256:644e7d101075396443325bb2474e3cdc5dbe0485f9ee05a83ef7f54e5522fccf

Observation 5fda3f1f-e7e8-4272-9da1-7fd420f65fcc · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:20.277818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:20.277818Z digest=sha256:6052a9b0326c764fae4e887b2d1aed3f3e1a8af541a494e7423b1ac20c4dd1b3

Observation ee3559cc-902c-4a5e-8aa7-a43b9f23cc53 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.018619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.018619Z digest=sha256:2c0f50d2480c07a3ea986c561369aadfc724029fbbbc4c064cfd75d8828ff2e9

Observation 58f07b94-66e8-42f0-970a-a83355966612 · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:41.875811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:41.875811Z digest=sha256:422dff8bc6b2ba6232a5a2e5ce25959262d7520e1886101882e0076b54ffd2f2