Pith. sign in

Paper Citation Record · LEDGER

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024

As of 13 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2412.01100.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01100 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:46:23.404901Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:46:23.292385Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T04:46:23.543943Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1190a25e-e7b2-4c45-889b-8d3fc6a7410d · outbound

This paper cites On the one hand, TTS systems need to accurately replicate the target speakers’ voice, including their timbre, pitch, and prosody, using voice cloning techniques.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 On the one hand, TTS systems need to accurately replicate the target speakers’ voice, including their timbre, pitch, and prosody, using voice cloning techniques

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.787300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.285394Z digest=sha256:23ec006cdc81ea7966e3d62bf539089efc95e9d594b9cc63a8c8ae963976d2e1

Observation a5953255-227d-4271-9788-83a2cc98c164 · outbound

This paper cites The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T04:46:23.547969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.292385Z digest=sha256:ba62a20999c4004992f498dee54b749b0ec4056abc2eda6cc782e09104a1e23d

Observation 592e28c6-cd5f-4b2a-98ac-9cdff6417d90 · outbound

This paper cites Firstly, we will overview the text represen- tations and the discrete speech tokens, and then introduce the LLaMA-based codec language model with a delay pattern.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Firstly, we will overview the text represen- tations and the discrete speech tokens, and then introduce the LLaMA-based codec language model with a delay pattern

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.775974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.296336Z digest=sha256:1c0b2cc6e64cb2a664252881b62b15736ca2ab9dbb8253c96205a86a7af07412

Observation 6f26b3b5-fbf7-41de-9a15-3bd8a680e28b · outbound

This paper cites Model Configurations We use the open-source models and parameters of MT5-base, DAC and HuBERT, with both DAC and HuBERT configured for a sampling rate of 16 kHz.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Model Configurations We use the open-source models and parameters of MT5-base, DAC and HuBERT, with both DAC and HuBERT configured for a sampling rate of 16 kHz

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T04:46:23.764622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.300465Z digest=sha256:3abb6aff47ae8bed01f8df3502f6404b7fc35c318fd9afe1fb533abde232fdc9

Observation 01c61244-b2b2-41d0-90ba-e3291cb72429 · outbound

This paper cites 嗯” (En- glish translation: “um.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 嗯” (En- glish translation: “um

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.752707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.304430Z digest=sha256:a8cd74cea740a4e6d3774ebeec0bf564d9f2f908519ae05fff90305dbe259b86

Observation b7fa6797-065c-45fa-8d5d-8d8bf1dff17b · outbound

This paper cites We propose a LLaMA-based codec language model with a delay pattern for spontaneous style voice cloning.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 We propose a LLaMA-based codec language model with a delay pattern for spontaneous style voice cloning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.740832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.308086Z digest=sha256:45b48255aeb223bb2ca9c526f66cc46165642fc3733764172c3143b6b7c03d04

Observation 542d5b6a-0118-423b-9794-2f06bfeb2511 · outbound

This paper cites an unresolved cited work.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:23.727825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.311926Z digest=sha256:05aeea29cbb2cf01a73219e91d64d5baed1cfcccefeceb702092247222974e80

Observation 139d5379-45c2-4cf5-bd1c-874b58db1a35 · outbound

This paper cites Deep voice 3: Scaling text- to-speech with convolutional sequence learning,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Deep voice 3: Scaling text- to-speech with convolutional sequence learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.716321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.315475Z digest=sha256:0b17c5e92b3007cc1f90dde22734da1659caeaf2bd9c942e61498bf70811251e

Observation b710d84c-195f-4657-98c3-a383d04cd232 · outbound

This paper cites Neural voice cloning with a few samples,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Neural voice cloning with a few samples,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.705048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.319118Z digest=sha256:4c2ab83025aa6b65bf2c1fc898c4ebbd0d45d4e891933e1d524b1388541a3845

Observation ea60877d-f449-4e08-bac4-d82e9ed42ed0 · outbound

This paper cites Adaspeech: Adaptive text to speech for custom voice,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Adaspeech: Adaptive text to speech for custom voice,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.693828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.322596Z digest=sha256:23b753635d69cc506d830c7ceffb95ac6217967c7c32da19d886554625feb9e1

Observation 2be492a3-b375-428d-9fd9-c5743916fb77 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.326163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.326163Z digest=sha256:08f6a7f0e492fbdc96635a9f8647c47370a96ea8c05db1e814a8df650fe4105a

Observation d446060d-6a77-46c0-81dd-6940d5873c3c · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.330505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.330505Z digest=sha256:abef8c3181cadf48449946fb88523671cbe2769b536102cb5f083eed897d885e

Observation d293bd2d-c7d7-4380-9a81-fddc37b9b8b6 · outbound

This paper cites Speechx: Neural codec language model as a versatile speech transformer,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Speechx: Neural codec language model as a versatile speech transformer,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.682223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.334229Z digest=sha256:23e37d1251065eb84a5804ea17c477c6ca646832ac6cc8c1e9550667312b5d7b

Observation 52a4f702-c67b-4136-bb99-3b2fed255e00 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.337892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.337892Z digest=sha256:2f8ccb62d41c14f7894a1dee83fba2464ee9e1475500e8279d4487f23f815778

Observation 9536cd50-92d2-420f-b1a8-48d92887b299 · outbound

This paper cites High Fidelity Neural Audio Compression.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 High Fidelity Neural Audio Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.341898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.341898Z digest=sha256:eafd3fb071cc2f6a70dbc6c588c6cb4c9515c9ac60db041a4302135eb93da8e1

Observation 56d4e3c2-ab15-4282-8e69-e1e06b5951f9 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Soundstream: An end-to-end neural audio codec,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.345998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.345998Z digest=sha256:86ad3f83a35c0e913f2afa99a6700405ff934472697d7da34d9c7c1bd606c551

Observation 2336ffaf-b83f-4884-8c33-73fefbbd7960 · outbound

This paper cites Conversational end-to-end tts for voice agents,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Conversational end-to-end tts for voice agents,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.662670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.349799Z digest=sha256:5a50dfcfb346c6a9ae4f6f806c8ecf292f8fa3dfcd4e1fe9d7e0ec0814ad4811

Observation 9ace13ad-7a94-4da0-b69b-b61de73813ed · outbound

This paper cites End-to-end text-to-speech based on latent representa- tion of speaking styles using spontaneous dialogue,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 End-to-end text-to-speech based on latent representa- tion of speaking styles using spontaneous dialogue,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.651480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.353254Z digest=sha256:2d04bee904348649edc4a1d3dd30822adff527ce90ba277c055eb6a54013a1c6

Observation e941ab01-b81a-4b56-82c3-6611513a1fc5 · outbound

This paper cites Spontts: modeling and transferring spontaneous style for tts,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Spontts: modeling and transferring spontaneous style for tts,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.640293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.356699Z digest=sha256:01aafcd3ffbe8eddcc052206c424bfd691cc8b34b4f34577d4a25f50e1ac797b

Observation ab184b98-a3ca-4c99-96f8-ebee51f4f082 · outbound

This paper cites Controllable Context-aware Conversational Speech Synthesis.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Controllable Context-aware Conversational Speech Synthesis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:46:23.491867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.360982Z digest=sha256:418ebcd7e8545f76ba0b8eb7e918aa16c25850096462eabe2c482bd180d49e7c

Observation eef021a3-549c-473f-8d2f-0f8ff685f326 · outbound

This paper cites Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.364935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.364935Z digest=sha256:12224368aad84a54b4ccd723064aaa277930ec629fc5aa88bf411622cea6f579

Observation 8198aee0-864b-4a99-aae9-39ad1aece243 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.368773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.368773Z digest=sha256:6d84d6634a9c9b83646923907839c957f16cbe44fd7111c92a4a72434149a241

Observation 713e2b56-1c94-4773-bb5a-85dd25ea7cc5 · outbound

This paper cites Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.629094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.371902Z digest=sha256:bcc6d02863987d834afb07d5276786d49d0cb583e78a652768a1f36514afd350

Observation d7bf9a78-9f93-4ed1-baf8-69e0a9c4cdcc · outbound

This paper cites mt5: A massively multilingual pre-trained text-to-text transformer,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 mt5: A massively multilingual pre-trained text-to-text transformer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.617503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.374799Z digest=sha256:8a7b0d6a158e7c86eef2219b312895bacc885d01d97d50c2f3e88b2ed487d8ea

Observation a8c3f4cf-12b0-42e2-b03f-74b31a86af06 · outbound

This paper cites HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.377892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.377892Z digest=sha256:97906b3811a5605d9bc4bd0d26d56ec35dd91548f4493d139f2450b722d40337

Observation f66521be-d789-4b3c-8d13-68af8feaa5a6 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 High-fidelity audio compression with improved rvqgan,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.599558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.381246Z digest=sha256:782835d62336524cdc9965398f7758ebf3926462e7615cf2d195781b71c8892f

Observation 0f5ec975-e862-44f8-8c43-08b91f42b660 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.384154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.384154Z digest=sha256:320316afa850bbe8b9e7554e2673e4add3d860d059c9848e9e6ddcb43ef82b31

Observation d9cbc3d8-0b36-4985-b679-2ddd1768a059 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.589572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.387495Z digest=sha256:cc40baf10630e2c5f61f4b492213e78dd54d530fc8404b2a924b708299701314

Observation 30144310-756d-41df-ba16-52faf1b7a0e8 · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with min- imal supervision,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Speak, read and prompt: High-fidelity text-to-speech with min- imal supervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.390561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.390561Z digest=sha256:e98254c0718d8818eaeda00f3081f805cae0b1e78a561caa5594f4df23354fac

Observation b902a4c1-e7f8-48ac-8890-11e9b59feea2 · outbound

This paper cites Simple and controllable music gen- eration,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Simple and controllable music gen- eration,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.570739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.394037Z digest=sha256:324d98297d2744938b7e2b498de30a684ee0da94673e8f82ed18e6c6fcee389b

Observation d6439b49-a38e-4d9c-82ea-8e7aacd6f301 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.397620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.397620Z digest=sha256:f091c19a724c53fcb688ddd6248cdc3e9fcb9076e4271b7b879382be1a84dd08

Observation 39d733d5-23be-4d16-a63b-ccb8ab33b86a · outbound

This paper cites Stay on topic with Classifier-Free Guidance.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 Stay on topic with Classifier-Free Guidance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:23.401316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:23.401316Z digest=sha256:89271aa8996cc06853bb8e5399d9d4774676690d3f1fda229cac8680392473e7

Observation 724b1e42-e0be-4518-9a48-49440cf836ad · outbound

This paper cites V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:23.559034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.404901Z digest=sha256:cef155464bfa1a1cd0f440b8282b37b9f12f90c33c1d303f1afd9a151f6b41cf

Pith citing papers

Observation a5953255-227d-4271-9788-83a2cc98c164 · inbound

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 cites this paper.

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024 The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T04:46:23.547969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:46:23.292385Z digest=sha256:ba62a20999c4004992f498dee54b749b0ec4056abc2eda6cc782e09104a1e23d