Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:11:40.541402Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2507.00808.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:11:40.541402Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:11:26.568638Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T21:11:40.627696Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f5356663-a23d-46d0-be41-e81f4b445693 · outbound
Multi-interaction TTS toward professional recording reproduction Multi-interaction TTS toward professional recording reproduction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c49bf936-37ec-4847-a415-fd790f86d472 · outbound
Multi-interaction TTS toward professional recording reproduction During the recording of this dataset, a di- rector iteratively gave acting directions, and the voice actor then reflected the given directions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1c22324-83e5-42e7-bef6-bbd725d26ea2 · outbound
Multi-interaction TTS toward professional recording reproduction Speak more brightly,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d97edb3-756a-490f-b9ba-e537f8a8edae · outbound
Multi-interaction TTS toward professional recording reproduction In- sert a silent-pause after this word,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e59eaec2-69c1-42b3-8cb5-86b257961fb5 · outbound
Multi-interaction TTS toward professional recording reproduction Follow my example
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bc539be-d5a5-48d3-9873-1a8d7fd58242 · outbound
Multi-interaction TTS toward professional recording reproduction The Guideline for TTS Speaking Style Classification
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 589bc8ec-64dd-45fb-b977-05653a3c6f3c · outbound
Multi-interaction TTS toward professional recording reproduction Dataset We used three Japanese 22 kHz datasets: interactive, non- interactive, and large in-house datasets
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2255d797-b963-460e-befe-03f7b0d40e36 · outbound
Multi-interaction TTS toward professional recording reproduction Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53851e5c-e299-44a8-bb96-0a63fe1a2e24 · outbound
Multi-interaction TTS toward professional recording reproduction Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98649d8d-f586-442b-a582-b690c2f38fc5 · outbound
Multi-interaction TTS toward professional recording reproduction The loss function and learning rate were the same as in the previous step
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84dc560a-abc9-4044-bb90-bc8db27c61d3 · outbound
Multi-interaction TTS toward professional recording reproduction The training step was 10K steps with AdamW optimizer [41] with 4K warm-up steps
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62883a09-b4af-4998-81d9-34b4d500cd6f · outbound
Multi-interaction TTS toward professional recording reproduction Overall alignment only
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30701ae4-0d7b-4e3d-bc99-85fcfd78448c · outbound
Multi-interaction TTS toward professional recording reproduction at the beginning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11d0c593-3556-437e-8097-2fccaec969bf · outbound
Multi-interaction TTS toward professional recording reproduction Subjective evaluations demonstrated that our proposed method achieved iterative style refinement that matched the user’s di- rections to some extent
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 163bc918-16bc-40fe-a3db-a4eac5a8a983 · outbound
Multi-interaction TTS toward professional recording reproduction The relevance of trial-and-error: Can trial-and- error be a sufficient learning method in technical problem-solving- contexts?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8bf1404-3233-4bf1-9700-f8452d280a5c · outbound
Multi-interaction TTS toward professional recording reproduction From da Vinci’s flying machines to a theory of the creative process,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f4e05c7-a4fa-4026-a76d-176b02ac7acf · outbound
Multi-interaction TTS toward professional recording reproduction What are the stages of the creative process? What visual art students are saying
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af7aab1a-6d09-4f31-87d0-ce2ac2493414 · outbound
Multi-interaction TTS toward professional recording reproduction From page to stage: The director’s interpretation and picturization of a script,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e678791d-f81a-4aac-9e6b-68544a57c169 · outbound
Multi-interaction TTS toward professional recording reproduction Showing and telling—How directors combine embodied demonstrations and verbal descrip- tions to instruct in theater rehearsals,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cf3832f-cff7-4458-a084-a642fbc34abb · outbound
Multi-interaction TTS toward professional recording reproduction Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e58288b0-a4b1-4f92-a2ef-a146ededeeb0 · outbound
Multi-interaction TTS toward professional recording reproduction High-resolution image synthesis with latent diffusion models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aae4d33f-0de3-41cc-b023-a8f231788bf9 · outbound
Multi-interaction TTS toward professional recording reproduction Photorealistic text-to- image diffusion models with deep language understanding,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a6f5387-cb32-4563-8ea8-44b89c4a90a1 · outbound
Multi-interaction TTS toward professional recording reproduction Program Synthesis with Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ddfe0d6-1d71-498e-b9b0-ef9c9ae37a37 · outbound
Multi-interaction TTS toward professional recording reproduction CodeGen: An open large language model for code with multi-turn program synthesis,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d212c621-fc9d-4f65-a863-63aa3f829877 · outbound
Multi-interaction TTS toward professional recording reproduction Training language models to follow instruc- tions with human feedback,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75703a4b-a7e8-464c-a830-477c1c0ca0e4 · outbound
Multi-interaction TTS toward professional recording reproduction PaLM: Scaling language modeling with path- ways,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b928b82-e0bb-41eb-b8de-269c3720e067 · outbound
Multi-interaction TTS toward professional recording reproduction A Survey on Neural Speech Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229a5f78-d42d-43a9-9874-757d67685f5d · outbound
Multi-interaction TTS toward professional recording reproduction A review of deep learning techniques for speech processing,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b376339-a22d-4f5a-b305-0dcea196f62a · outbound
Multi-interaction TTS toward professional recording reproduction Model archi- tectures to extrapolate emotional expressions in DNN-based text- to-speech,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4da98607-1c1e-42f2-a736-10a0cfe4ae89 · outbound
Multi-interaction TTS toward professional recording reproduction V oice puppetry: Exploring dramatic performance to develop speech synthesis,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7d8a15b-bdbf-4a10-ad7b-6e9bb3c7339a · outbound
Multi-interaction TTS toward professional recording reproduction V oice puppetry with FastPitch,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a885f33-c101-4bd3-8e7c-a2a3de59c71d · outbound
Multi-interaction TTS toward professional recording reproduction Style Tokens: Unsu- pervised style modeling, control and transfer in end-to-end speech synthesis,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0858633e-7ff2-4f3a-909d-9a0cb74122cd · outbound
Multi-interaction TTS toward professional recording reproduction Robust and fine-grained prosody control of end-to-end speech synthesis,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82dd43d4-0a36-4e58-8ede-1133e2d234d4 · outbound
Multi-interaction TTS toward professional recording reproduction Fine- grained robust prosody transfer for single-speaker neural text-to- speech,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92a808f0-8fd2-4add-a03d-b91a9555accb · outbound
Multi-interaction TTS toward professional recording reproduction Daft- Exprt: Cross-speaker prosody transfer on any text for expressive speech synthesis,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e7cf286-77db-4070-8541-5dba1a3fde25 · outbound
Multi-interaction TTS toward professional recording reproduction PromptTTS: Controllable text-to-speech with text descriptions,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c2cf6a8-f131-4464-984c-ec5e78ae53f9 · outbound
Multi-interaction TTS toward professional recording reproduction Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e89fee7-5269-4fea-9a9c-9fcabc87d058 · outbound
Multi-interaction TTS toward professional recording reproduction VoiceCraft: Zero-shot speech editing and text-to-speech in the wild,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee147f85-f3dc-4ce7-ab12-bd22785346eb · outbound
Multi-interaction TTS toward professional recording reproduction V oice at- tribute editing with text prompt,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3f6fd9f-63d3-4cb1-8b9f-21642b264989 · outbound
Multi-interaction TTS toward professional recording reproduction The guidelines for TTS speaking style classifi- cation (IT-4012),
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c9e30ab-a281-45c9-baed-3c15a4a6478a · outbound
Multi-interaction TTS toward professional recording reproduction Zero-shot text-to-speech synthesis conditioned using self- supervised speech representation model,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edd0a729-303e-49e2-817c-e4fd0189a9cc · outbound
Multi-interaction TTS toward professional recording reproduction Why does self-supervised learning for speech recognition benefit speaker recognition?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90c31184-97b6-4947-8cf3-5ad8aea0dadf · outbound
Multi-interaction TTS toward professional recording reproduction Feed-forward networks with atten- tion can solve some long-term memory problems,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f58c8c9-6a46-4d3d-8bef-7e44cc592c25 · outbound
Multi-interaction TTS toward professional recording reproduction FiLM: Visual reasoning with a general conditioning layer,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a142c6f3-6c21-4885-9782-5769efddf13f · outbound
Multi-interaction TTS toward professional recording reproduction Learning alignment for multimodal emotion recognition from speech,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f8acf10-56df-4593-9a55-93ff0820cb56 · outbound
Multi-interaction TTS toward professional recording reproduction Multimodal cross- and self-attention network for speech emotion recognition,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30949dc0-96c0-4ef4-b9ea-9bc9b9e9bf48 · outbound
Multi-interaction TTS toward professional recording reproduction Hello GPT-4o,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2356dee1-9d92-443f-97df-de095f883371 · outbound
Multi-interaction TTS toward professional recording reproduction Rephrasing the web: A recipe for compute and data-efficient lan- guage modeling,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04d3e439-1ad4-410d-a1c7-fdd0af902c2f · outbound
Multi-interaction TTS toward professional recording reproduction FastSpeech 2: Fast and high-quality end-to-end text to speech,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24aeb924-e40f-458f-ac95-331827345ec7 · outbound
Multi-interaction TTS toward professional recording reproduction In- vestigating on incorporating pretrained and learnable speaker rep- resentations for multi-speaker multi-style text-to-speech,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bc8d84b-6de6-49fa-9a4b-4c6bfcf054b6 · outbound
Multi-interaction TTS toward professional recording reproduction HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc66171-3503-43ed-a7c8-4a85e6d20c9b · outbound
Multi-interaction TTS toward professional recording reproduction The Curse of Recursion: Training on Generated Data Makes Models Forget
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e60fafe-82e7-4392-80da-a62b9d224e36 · outbound
Multi-interaction TTS toward professional recording reproduction Multi-speaker modeling for DNN- based speech synthesis incorporating generative adversarial net- works,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac610f13-f5c7-4140-b63a-55e47da9bec4 · outbound
Multi-interaction TTS toward professional recording reproduction Variational discriminator bottleneck: Improving imitation learn- ing, inverse RL, and GANs by constraining information flow,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ddfa18e2-47f7-4b0b-8e0b-9cbe15bdf6c9 · outbound
Multi-interaction TTS toward professional recording reproduction Decoupled weight decay regulariza- tion,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f70cff-87dd-4cf0-8dfb-4b54da726799 · outbound
Multi-interaction TTS toward professional recording reproduction Expressive text-to-speech synthesis using text chat dataset with speaking style information,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79df0849-aaf6-4be3-923c-1d70f86d183f · outbound
Multi-interaction TTS toward professional recording reproduction NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55c5343b-e23d-4e33-94bd-556f25af4981 · outbound
Multi-interaction TTS toward professional recording reproduction Neural codec language models are zero-shot text to speech synthesizers,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 448d23c8-497b-4cee-9f0c-fb3ba3d0a156 · outbound
Multi-interaction TTS toward professional recording reproduction FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27e012b0-5b8d-47a2-a991-ebdd4448ff7e · outbound
Multi-interaction TTS toward professional recording reproduction Toward verifiable and repro- ducible human evaluation for text-to-image generation,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5356663-a23d-46d0-be41-e81f4b445693 · inbound
Multi-interaction TTS toward professional recording reproduction Multi-interaction TTS toward professional recording reproduction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.