Pith. sign in

Paper Citation Record · LEDGER

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2608.11590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11590 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:42:51.295730Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:42:51.095849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:42:51.460473Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved23
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98a2c2dc-f6b3-4b3c-8074-158163818324 · outbound

This paper cites CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:42:51.465765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.095849Z digest=sha256:8efea4deb8ed43d76ee81bc859331cae057a81693b4b4c64163ff354f47883b1

Observation b79a8942-3775-423d-b626-9630037b115d · outbound

This paper cites By decomposing human voice into style, content, and prosody, CookV oice successfully unifying multiple voice generation task within a unified model.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation By decomposing human voice into style, content, and prosody, CookV oice successfully unifying multiple voice generation task within a unified model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.930837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.101490Z digest=sha256:fb7d3a61f986766b1701a32f9d6f10f35fb04c51ccba81372d07c3fd4a988494

Observation a86d69a5-5434-405c-898d-0a53177e79d3 · outbound

This paper cites an unresolved cited work.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:42:51.914976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.106606Z digest=sha256:63e77fc62b781667756197dcf54b51857d928bca15579e7e3798a4ae9935a541

Observation be907e9c-246b-4088-b820-9c97f3426f48 · outbound

This paper cites an unresolved cited work.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:42:51.899631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.112317Z digest=sha256:f4acc128225f2c5ffb5a1c4e3dee8675119a222676d3ebc9622e9e62c11efad4

Observation 334ab11e-101c-4624-9871-3381d6496da4 · outbound

This paper cites Problem Formulation Human V oice Generation (HVG) is to generate waveform from different modality of control signal.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Problem Formulation Human V oice Generation (HVG) is to generate waveform from different modality of control signal

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.885249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.117568Z digest=sha256:cbaf6bb6ea8c35f02ce5f6bdcebf3f00c04f14662b3f9c64ddb5ed7953795347

Observation 69ae5454-0479-4350-b803-a2ce807068bc · outbound

This paper cites This section will present the architecture of CookV oice in detail.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation This section will present the architecture of CookV oice in detail

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.870567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.122723Z digest=sha256:2e5e4fd927f373ae2c55e050aa9012c67f149b1708e7a473b5011a4901f9404f

Observation 74e7d7a9-1b77-43f9-b4be-a010e1e8d6c6 · outbound

This paper cites an unresolved cited work.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:42:51.854597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.128595Z digest=sha256:b50010e8e0aadde3fd91146c6c1e1bc3128fb471570032c98d6de6761a5e20b4

Observation 09750361-4f73-4943-9c79-bbb85e092703 · outbound

This paper cites an unresolved cited work.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:42:51.838302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.134204Z digest=sha256:8ef58114c8a711a389c2fa1160c9fe3ce77bdd33fffd1c5fb13766cf926ffb7b

Observation a8d7a2b9-4130-4d71-8e6e-d6c6991cb19f · outbound

This paper cites an unresolved cited work.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:42:51.823399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.139569Z digest=sha256:7d66da3a8b2aece3ac3602202565b9eaada83454527831bbd9f6e4621564b1a6

Observation 6e6461fa-f1dd-4461-80f2-26fc76b36fa5 · outbound

This paper cites an unresolved cited work.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:42:51.807315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.145027Z digest=sha256:912f4f89b9cd1a2472e8dc29ccee6838738cab583ef7825963b1d7832020879e

Observation 5f72e243-2ba4-4986-88e2-d1ebb6485937 · outbound

This paper cites First, due to re- source constraints, CookV oice has not yet been scaled up.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation First, due to re- source constraints, CookV oice has not yet been scaled up

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.790639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.149789Z digest=sha256:71d19e35aa802b4dca49634ff1a18c01b8f5eb6d711d260a8deb212fc5ef9eff

Observation a37fe608-667b-48db-bff1-6fba55f0da50 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.154323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.154323Z digest=sha256:5038ff787f2c3afe6003abbe595941909cdbbb965b1fcc91311bd813f64a3b37

Observation fe8eb4b9-9524-4521-94c9-caab31cddeb3 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.159376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.159376Z digest=sha256:4d0e74d358c89fde6cb81096458559c4fc45325bd3c659ce165b70cb658546ce

Observation 66c686c9-89e8-42e3-85f1-9ec539400888 · outbound

This paper cites Parastyletts: Toward effi- cient and robust paralinguistic style control for expressive text-to- speech generation,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Parastyletts: Toward effi- cient and robust paralinguistic style control for expressive text-to- speech generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.773610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.165004Z digest=sha256:16a57dc1b81a0ccd8aa02fb5a3fecda91e4c7b6570ed9bf3064b42780263b368

Observation cf94d5fd-9273-42d3-bb02-07f2d22d7f1c · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.171355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.171355Z digest=sha256:3f40f0a053f1407378bd75d4f01282c873fb759cdd9286817b2a21f7fef3daa7

Observation fe291923-ad50-4d6d-ad76-f99a0b6b1633 · outbound

This paper cites Diffsinger: Singing voice synthesis via shallow diffusion mechanism,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Diffsinger: Singing voice synthesis via shallow diffusion mechanism,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.176021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.176021Z digest=sha256:89c325c42f3aeafc194fb9f0f1a84fa865a9a76bc6dab5019f097b00f8a73cd3

Observation 173cc171-d93c-4ca4-baaf-67182a6f41ac · outbound

This paper cites Stylesinger: Style transfer for out-of- domain singing voice synthesis,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Stylesinger: Style transfer for out-of- domain singing voice synthesis,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.181145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.181145Z digest=sha256:2a9b15bb928ad90dc9f71b98de285425411fc6c3d0b49f09a8ceb8c4fa1e7655

Observation 009b6443-3e49-41c7-b426-14c095605731 · outbound

This paper cites Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.725858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.185523Z digest=sha256:80515cb195fe99955e8996b01befd01a7b45aed25f2f4432f4b53ac4352ff420

Observation 974fac2a-68cb-4534-85db-f9e026e5b2e9 · outbound

This paper cites Vevo2: A unified and controllable framework for speech and singing voice generation,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Vevo2: A unified and controllable framework for speech and singing voice generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.708638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.190152Z digest=sha256:4ea6d4b02cfd43eeda1771f03f50a52f954c04897bc1f81af9a7ec5b4e248fbc

Observation 44a02b03-9794-4069-86ee-0e291628071a · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.194600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.194600Z digest=sha256:53ccb55e65e5f265c5ee26d701c7404398d706e5f9ba7bab1713a142d9e81a3b

Observation 45d1b6a3-090a-4493-9669-e6bdd0ee010c · outbound

This paper cites Param- eta: Towards learning disentangled paralinguistic speak- ing styles representations from speech,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Param- eta: Towards learning disentangled paralinguistic speak- ing styles representations from speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.681297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.199775Z digest=sha256:157a6e57d0252b9ea646a308783691c220df48c0cb04e10706d52570335b89f8

Observation 5d62aaf7-8d75-45a5-8d63-47ff93c4e04a · outbound

This paper cites Scalable diffusion models with transform- ers,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Scalable diffusion models with transform- ers,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.204364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.204364Z digest=sha256:8e2e580b1d7c03139b2f8c5513b7384287b2a64d45541dac8b7b6042f7d3f5a4

Observation 3c36f22b-0f7b-4694-a4fc-aaa26f3db153 · outbound

This paper cites Flow Matching for Generative Modeling.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Flow Matching for Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.209215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.209215Z digest=sha256:c42ce2d2680d02b9eff9af3e70defd63a5ffc25df69a974cad5aa853b9c609ce

Observation 0ccf258c-424b-447b-9257-f5ed5b413d3c · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Fastspeech: Fast, robust and controllable text to speech,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.214007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.214007Z digest=sha256:1fe8b41e13ece5345094570839fa2893eb8b21966180b2349eedfd0899459269

Observation 12131af5-38dc-4fe5-a1d8-9bd92cdd06e6 · outbound

This paper cites M4singer: A multi-style, multi- singer and musical score provided mandarin singing corpus,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation M4singer: A multi-style, multi- singer and musical score provided mandarin singing corpus,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.218977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.218977Z digest=sha256:b6473335cd173156765699ff015e158cba397709ba2bc794e508a9ce32f8d10b

Observation 6ad2d683-fcdf-47c6-9e7e-6f8e61f1bac7 · outbound

This paper cites Vevo: Controllable zero-shot voice imitation with self- supervised disentanglement,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Vevo: Controllable zero-shot voice imitation with self- supervised disentanglement,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.634876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.224026Z digest=sha256:81b820703396348f294805ffc0d182cea25a665dc9a1d896b5e5ebc799da04a4

Observation d17a0e1f-4fa3-4448-8e36-28f637dd63a7 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.228726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.228726Z digest=sha256:bcef3e887a2bac0e39fe770b1a1c280759beb16e8c4232de668eb409f68b9adb

Observation abd90a01-1665-4d68-b65a-076bb6eac9df · outbound

This paper cites Mpnet: Masked and permuted pre-training for language understanding,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Mpnet: Masked and permuted pre-training for language understanding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.233657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.233657Z digest=sha256:8dc7ecce573646f2b25d24e57840ce8ea95a0e51ce7e8be35c3b8ab42cfdfcba

Observation d83b9c17-b824-4ef4-926e-2b3fe6922816 · outbound

This paper cites Attention is all you need,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Attention is all you need,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.238525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.238525Z digest=sha256:d5390c98aa2e07e4344b12957a8634195c974a778d1b3a9cc8ed4522a5becb65

Observation 0c328557-4944-40be-9b3c-1d8dd6b3f201 · outbound

This paper cites Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:42:51.392027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.243406Z digest=sha256:83a0508fc3838e079062a3eb272a7764c23fe1c3757fe6a31928e9af3b2c4228

Observation 0ca79004-7c70-42aa-a609-fc8b51d69561 · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Tacotron: Towards End-to-End Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.248293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.248293Z digest=sha256:61e21d6750a24fbffcb1f2a2354edd5c4b4e00e8b8db802eb66ff7f3d463964f

Observation 217af388-708a-4e4a-b532-f2b313c347d5 · outbound

This paper cites Chinese mandarin female corpus,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Chinese mandarin female corpus,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.588447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.253351Z digest=sha256:8288555f00473ce0c05681b8917fda7514a37f0e05a46c5760a80669218cee05

Observation b8010839-144b-4a12-9e34-729c08864903 · outbound

This paper cites The lj speech dataset,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation The lj speech dataset,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.257760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.257760Z digest=sha256:5dd945793874a4f9a68156357f849990677510656bf975929d79f303b48833f4

Observation 276cbf09-14fd-4586-b6b0-33106e777b4a · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.262378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.262378Z digest=sha256:29d42b3aacb8a6f095b6d45f416467953f019b025749b0117b36a023d114a810

Observation b1a38fdb-4e3e-4635-8f11-f0333b8acf07 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Crema-d: Crowd-sourced emotional multimodal actors dataset,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.266868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.266868Z digest=sha256:1c85dcd9dccff276ae0aeb98cbf950c29ba989488d6fe299786200dc6d4cc698

Observation 020e2760-ffe3-464a-b42f-3870e233481b · outbound

This paper cites Common phone: A multilingual dataset for robust acoustic modelling,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Common phone: A multilingual dataset for robust acoustic modelling,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.541875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.271354Z digest=sha256:b12b3e81c575ca14b17710453f0c4f5168f55527b6928704af64001daa306dc4

Observation ebf432da-1285-4d7a-be81-5392174b4368 · outbound

This paper cites Genshin voice: A multi-lingual voice dataset from Genshin Impact,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Genshin voice: A multi-lingual voice dataset from Genshin Impact,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.526135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.276503Z digest=sha256:ddcafacfbe781125ae860314e07abd6f6012c8b248077c38aeff680af2aa9b07

Observation ed5c18f1-fcf9-42c2-b65d-1ae6fa2b0bcc · outbound

This paper cites Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:42:51.510189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.281545Z digest=sha256:8e70a423e85540cd36b9328fed142c06da9a539c7bbfc001bf2965ac3dd7d695

Observation 930af073-8f3b-4301-90d2-e8ca61305640 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Robust speech recognition via large-scale weak supervision,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:42:51.286306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:42:51.286306Z digest=sha256:e3a340fadc348d35f0adad65fa2f3c6687524f4a590c562024177fd72e4d985d

Observation 78b12c8d-717f-4cf9-adfd-358121caedba · outbound

This paper cites AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:42:51.345318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.290965Z digest=sha256:fb27c72817d0621f15fa252ad5c1d84930eea8d4755c9029f6a813f4f87dc71b

Observation b2f59396-0571-4e70-9455-1eb1fe06c189 · outbound

This paper cites Denoising diffusion probabilis- tic models,.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Denoising diffusion probabilis- tic models,

Reference 41

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:42:51.482911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.295730Z digest=sha256:53ada697303f7ddbc3b1cf27286107751fbe837a356c819b33daa8fdfb588c1a

Pith citing papers

Observation 98a2c2dc-f6b3-4b3c-8074-158163818324 · inbound

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation cites this paper.

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:42:51.465765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:42:51.095849Z digest=sha256:8efea4deb8ed43d76ee81bc859331cae057a81693b4b4c64163ff354f47883b1