Pith. sign in

Paper Citation Record · LEDGER

Multi-interaction TTS toward professional recording reproduction

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2507.00808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00808 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:11:40.541402Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:11:26.568638Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:11:40.627696Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5356663-a23d-46d0-be41-e81f4b445693 · outbound

This paper cites Multi-interaction TTS toward professional recording reproduction.

Multi-interaction TTS toward professional recording reproduction Multi-interaction TTS toward professional recording reproduction

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:11:40.632803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.568638Z digest=sha256:a38e0161e98f63ca717e0b4b19f29b873af661bf81daa713cdc2d3d3cdf89ca9

Observation c49bf936-37ec-4847-a415-fd790f86d472 · outbound

This paper cites During the recording of this dataset, a di- rector iteratively gave acting directions, and the voice actor then reflected the given directions.

Multi-interaction TTS toward professional recording reproduction During the recording of this dataset, a di- rector iteratively gave acting directions, and the voice actor then reflected the given directions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.243995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.576860Z digest=sha256:9e305f6f4f5da7a226bf42ed63082c0191dfee42355aff155013ea47850c20a2

Observation e1c22324-83e5-42e7-bef6-bbd725d26ea2 · outbound

This paper cites Speak more brightly,.

Multi-interaction TTS toward professional recording reproduction Speak more brightly,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.232486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.591023Z digest=sha256:a54732c197f697c427aa610da401f0ad1e22ab1f66aea69ba797f597f651bc4d

Observation 7d97edb3-756a-490f-b9ba-e537f8a8edae · outbound

This paper cites In- sert a silent-pause after this word,.

Multi-interaction TTS toward professional recording reproduction In- sert a silent-pause after this word,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.220883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.605668Z digest=sha256:ad1085609f54d21a721c3afdecd7b51100828b22ba820bd199a1a8d289617565

Observation e59eaec2-69c1-42b3-8cb5-86b257961fb5 · outbound

This paper cites Follow my example.

Multi-interaction TTS toward professional recording reproduction Follow my example

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.209689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.623467Z digest=sha256:cd91f8b440a8d44f865ce612cc190d7460a4b79d8d0e7b07abef7ba5f3efe93a

Observation 0bc539be-d5a5-48d3-9873-1a8d7fd58242 · outbound

This paper cites The Guideline for TTS Speaking Style Classification.

Multi-interaction TTS toward professional recording reproduction The Guideline for TTS Speaking Style Classification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.195100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.642547Z digest=sha256:677a8fcc35f32b698087af593133751358636a141b9e120c9d13ea6414c57413

Observation 589bc8ec-64dd-45fb-b977-05653a3c6f3c · outbound

This paper cites Dataset We used three Japanese 22 kHz datasets: interactive, non- interactive, and large in-house datasets.

Multi-interaction TTS toward professional recording reproduction Dataset We used three Japanese 22 kHz datasets: interactive, non- interactive, and large in-house datasets

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.182706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.663023Z digest=sha256:8e8fecca99d227644db51d69a4e95d1df9222f32db62732a28c78c322fe0dad3

Observation 2255d797-b963-460e-befe-03f7b0d40e36 · outbound

This paper cites an unresolved cited work.

Multi-interaction TTS toward professional recording reproduction Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:11:41.170322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.683475Z digest=sha256:01d7a5c959d6b70d9c4d477a19bc140feaa4eb0452615a23cfa14b35f339de89

Observation 53851e5c-e299-44a8-bb96-0a63fe1a2e24 · outbound

This paper cites an unresolved cited work.

Multi-interaction TTS toward professional recording reproduction Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:11:41.159357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.694379Z digest=sha256:321fce630cfc35e180779ce598867f8f0ef8b78ae7ef5012c68e9dff8e3a8d37

Observation 98649d8d-f586-442b-a582-b690c2f38fc5 · outbound

This paper cites The loss function and learning rate were the same as in the previous step.

Multi-interaction TTS toward professional recording reproduction The loss function and learning rate were the same as in the previous step

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.148255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.703922Z digest=sha256:45ca5ed80eebbcc1afb10995a5c933121218c4ee5bbcd8e9aa89e7193dfe88af

Observation 84dc560a-abc9-4044-bb90-bc8db27c61d3 · outbound

This paper cites The training step was 10K steps with AdamW optimizer [41] with 4K warm-up steps.

Multi-interaction TTS toward professional recording reproduction The training step was 10K steps with AdamW optimizer [41] with 4K warm-up steps

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.137813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.715018Z digest=sha256:12a6037c0404afb192c4155194a8a7fb8f2da5ac6fce1f6ca390cc8790a51e3d

Observation 62883a09-b4af-4998-81d9-34b4d500cd6f · outbound

This paper cites Overall alignment only.

Multi-interaction TTS toward professional recording reproduction Overall alignment only

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.125676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.725122Z digest=sha256:0230498a44a52a41bd8d01aafbd8af67e49bb57b5cd4f89d667d38eaaa2bead7

Observation 30701ae4-0d7b-4e3d-bc99-85fcfd78448c · outbound

This paper cites at the beginning.

Multi-interaction TTS toward professional recording reproduction at the beginning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.115057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.736074Z digest=sha256:fe37b275834308ca828282438d4f3c460b9810239bd5a0f3afa54502ad10ee65

Observation 11d0c593-3556-437e-8097-2fccaec969bf · outbound

This paper cites Subjective evaluations demonstrated that our proposed method achieved iterative style refinement that matched the user’s di- rections to some extent.

Multi-interaction TTS toward professional recording reproduction Subjective evaluations demonstrated that our proposed method achieved iterative style refinement that matched the user’s di- rections to some extent

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.104855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.749700Z digest=sha256:5bedb174b0841ff10af4a61cb3fab587b8355e1e1cd69f1c745eccf9dcc121ed

Observation 163bc918-16bc-40fe-a3db-a4eac5a8a983 · outbound

This paper cites The relevance of trial-and-error: Can trial-and- error be a sufficient learning method in technical problem-solving- contexts?.

Multi-interaction TTS toward professional recording reproduction The relevance of trial-and-error: Can trial-and- error be a sufficient learning method in technical problem-solving- contexts?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.092369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.754699Z digest=sha256:947fb61b95b5edfcb26eda939beaf058d09b0cfc7174571011f308f1fed992b8

Observation d8bf1404-3233-4bf1-9700-f8452d280a5c · outbound

This paper cites From da Vinci’s flying machines to a theory of the creative process,.

Multi-interaction TTS toward professional recording reproduction From da Vinci’s flying machines to a theory of the creative process,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.079650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.762988Z digest=sha256:91e7edf9fa5abe206d13d12ba0698f7792c8a4f15664b066a00fdc1b55123f95

Observation 1f4e05c7-a4fa-4026-a76d-176b02ac7acf · outbound

This paper cites What are the stages of the creative process? What visual art students are saying.

Multi-interaction TTS toward professional recording reproduction What are the stages of the creative process? What visual art students are saying

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.067214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.806058Z digest=sha256:e598d5c01c30b6af7e38f34fcefe79c41c4f63c179aacc6516a47cb40cdefe66

Observation af7aab1a-6d09-4f31-87d0-ce2ac2493414 · outbound

This paper cites From page to stage: The director’s interpretation and picturization of a script,.

Multi-interaction TTS toward professional recording reproduction From page to stage: The director’s interpretation and picturization of a script,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.056318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.863549Z digest=sha256:3385871f97d3514f33ff59cefa6d8ddfd859151b0411828ca71365f561571c89

Observation e678791d-f81a-4aac-9e6b-68544a57c169 · outbound

This paper cites Showing and telling—How directors combine embodied demonstrations and verbal descrip- tions to instruct in theater rehearsals,.

Multi-interaction TTS toward professional recording reproduction Showing and telling—How directors combine embodied demonstrations and verbal descrip- tions to instruct in theater rehearsals,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.043331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.952999Z digest=sha256:a8263fa082e9e7064bf2704a22ba797bf2a08603eabb9686f20537a0af1b6295

Observation 6cf3832f-cff7-4458-a084-a642fbc34abb · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Multi-interaction TTS toward professional recording reproduction Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:27.044864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:27.044864Z digest=sha256:1e8b7e0d46527ada1537afd6cce49dfe1ace2e00fbb7ba86687a60b651324a05

Observation e58288b0-a4b1-4f92-a2ef-a146ededeeb0 · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Multi-interaction TTS toward professional recording reproduction High-resolution image synthesis with latent diffusion models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.030672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:27.144346Z digest=sha256:4c3c35819e0b17ea43485a984fe78ba592c87b268eab2785e2f68700bef6c3fa

Observation aae4d33f-0de3-41cc-b023-a8f231788bf9 · outbound

This paper cites Photorealistic text-to- image diffusion models with deep language understanding,.

Multi-interaction TTS toward professional recording reproduction Photorealistic text-to- image diffusion models with deep language understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.019811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:27.271811Z digest=sha256:63edad50bf7508219c45699e17c5fd1b4bc38fced8e9819b5ed4638294f6b3be

Observation 0a6f5387-cb32-4563-8ea8-44b89c4a90a1 · outbound

This paper cites Program Synthesis with Large Language Models.

Multi-interaction TTS toward professional recording reproduction Program Synthesis with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:27.358467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:27.358467Z digest=sha256:558e5c0ee4f6b1a329b68edd7a1980385e72993201347353067773612ba3c5b0

Observation 5ddfe0d6-1d71-498e-b9b0-ef9c9ae37a37 · outbound

This paper cites CodeGen: An open large language model for code with multi-turn program synthesis,.

Multi-interaction TTS toward professional recording reproduction CodeGen: An open large language model for code with multi-turn program synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:41.006962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:27.459503Z digest=sha256:f1194997dad621581e68d5f77d3bd368603fb3be0858af0f09a39e7be293f484

Observation d212c621-fc9d-4f65-a863-63aa3f829877 · outbound

This paper cites Training language models to follow instruc- tions with human feedback,.

Multi-interaction TTS toward professional recording reproduction Training language models to follow instruc- tions with human feedback,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.993049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.323645Z digest=sha256:b2e0a1a965d7883bb1b54def434910e0a59b7f0b9e6438706de22fca10497bb3

Observation 75703a4b-a7e8-464c-a830-477c1c0ca0e4 · outbound

This paper cites PaLM: Scaling language modeling with path- ways,.

Multi-interaction TTS toward professional recording reproduction PaLM: Scaling language modeling with path- ways,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.980148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.341138Z digest=sha256:ba7d3a2b650168f83291b2a1cce3f8c218208d462551576d6d94b581f7d60e2b

Observation 8b928b82-e0bb-41eb-b8de-269c3720e067 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Multi-interaction TTS toward professional recording reproduction A Survey on Neural Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.364754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.364754Z digest=sha256:41e7a277d2970a23a9305da3a535fb2e7429ebd0799f0588e93bebc07195cedd

Observation 229a5f78-d42d-43a9-9874-757d67685f5d · outbound

This paper cites A review of deep learning techniques for speech processing,.

Multi-interaction TTS toward professional recording reproduction A review of deep learning techniques for speech processing,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.391082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.391082Z digest=sha256:660bd041a81703c1e7510af86f6d79ad2c166fdf975bcee59e0fbcdd4ce13af9

Observation 4b376339-a22d-4f5a-b305-0dcea196f62a · outbound

This paper cites Model archi- tectures to extrapolate emotional expressions in DNN-based text- to-speech,.

Multi-interaction TTS toward professional recording reproduction Model archi- tectures to extrapolate emotional expressions in DNN-based text- to-speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.961913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.408756Z digest=sha256:6846040a653350bf5e64ed05ff16fde805f8fd31f43005873827f8ca70d78636

Observation 4da98607-1c1e-42f2-a736-10a0cfe4ae89 · outbound

This paper cites V oice puppetry: Exploring dramatic performance to develop speech synthesis,.

Multi-interaction TTS toward professional recording reproduction V oice puppetry: Exploring dramatic performance to develop speech synthesis,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.948640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.417111Z digest=sha256:30c4a640c142870f4dc6be554447c45c7c3e649febcdaa43429a5f98578e0e08

Observation d7d8a15b-bdbf-4a10-ad7b-6e9bb3c7339a · outbound

This paper cites V oice puppetry with FastPitch,.

Multi-interaction TTS toward professional recording reproduction V oice puppetry with FastPitch,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.936595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.420350Z digest=sha256:3462b672058baec755981abda234c89e2ff4ab79c02b0a64f24ef399d056325a

Observation 6a885f33-c101-4bd3-8e7c-a2a3de59c71d · outbound

This paper cites Style Tokens: Unsu- pervised style modeling, control and transfer in end-to-end speech synthesis,.

Multi-interaction TTS toward professional recording reproduction Style Tokens: Unsu- pervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.925522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.423738Z digest=sha256:e715b076dba4be97a46b168317790842b3b79fc58917fa2cfb7705451c3a0096

Observation 0858633e-7ff2-4f3a-909d-9a0cb74122cd · outbound

This paper cites Robust and fine-grained prosody control of end-to-end speech synthesis,.

Multi-interaction TTS toward professional recording reproduction Robust and fine-grained prosody control of end-to-end speech synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.914140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.437462Z digest=sha256:0c408594f2d3ed9cfd340102c38a79ae55b92d91098ea751f2e99da1af8f1e9e

Observation 82dd43d4-0a36-4e58-8ede-1133e2d234d4 · outbound

This paper cites Fine- grained robust prosody transfer for single-speaker neural text-to- speech,.

Multi-interaction TTS toward professional recording reproduction Fine- grained robust prosody transfer for single-speaker neural text-to- speech,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.903959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.442586Z digest=sha256:02a4eb91429c4712b8b090d53629b0b873b48e9938096ee077d89633f1c70b45

Observation 92a808f0-8fd2-4add-a03d-b91a9555accb · outbound

This paper cites Daft- Exprt: Cross-speaker prosody transfer on any text for expressive speech synthesis,.

Multi-interaction TTS toward professional recording reproduction Daft- Exprt: Cross-speaker prosody transfer on any text for expressive speech synthesis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.892893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.448323Z digest=sha256:dbcf7739de2d5624056f89c3daa265ceadb5ed6ca444f8fc83464c5325af8336

Observation 3e7cf286-77db-4070-8541-5dba1a3fde25 · outbound

This paper cites PromptTTS: Controllable text-to-speech with text descriptions,.

Multi-interaction TTS toward professional recording reproduction PromptTTS: Controllable text-to-speech with text descriptions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.882082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.452192Z digest=sha256:a163196252324e19d94fccc720c95faa312c0d8456e34f2c4bcc739a8048391c

Observation 5c2cf6a8-f131-4464-984c-ec5e78ae53f9 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

Multi-interaction TTS toward professional recording reproduction Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.456914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.456914Z digest=sha256:73afbf5bda5214577dae3a5e3ddf5c4ac63e65d3e54bd315fb31e11639b85190

Observation 8e89fee7-5269-4fea-9a9c-9fcabc87d058 · outbound

This paper cites VoiceCraft: Zero-shot speech editing and text-to-speech in the wild,.

Multi-interaction TTS toward professional recording reproduction VoiceCraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.870810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.460077Z digest=sha256:74b908a57f03ef34fb68089d8fbe572c43c72c87f3fa48fbf151735faa7ef839

Observation ee147f85-f3dc-4ce7-ab12-bd22785346eb · outbound

This paper cites V oice at- tribute editing with text prompt,.

Multi-interaction TTS toward professional recording reproduction V oice at- tribute editing with text prompt,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.859744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.463626Z digest=sha256:73b956b69211bcfa9d51d557d04517c10230773713499d38c158929be342bc7c

Observation c3f6fd9f-63d3-4cb1-8b9f-21642b264989 · outbound

This paper cites The guidelines for TTS speaking style classifi- cation (IT-4012),.

Multi-interaction TTS toward professional recording reproduction The guidelines for TTS speaking style classifi- cation (IT-4012),

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.848015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.466780Z digest=sha256:34fe606ffccee9f3b2a7abbd7acd6f0131674f2e431904819045c8364f6ea8bb

Observation 1c9e30ab-a281-45c9-baed-3c15a4a6478a · outbound

This paper cites Zero-shot text-to-speech synthesis conditioned using self- supervised speech representation model,.

Multi-interaction TTS toward professional recording reproduction Zero-shot text-to-speech synthesis conditioned using self- supervised speech representation model,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.836114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.470345Z digest=sha256:06a0aaa44debc6aec81c8d2bc9ed0d1f5cd5f552427c60552239fee453335a56

Observation edd0a729-303e-49e2-817c-e4fd0189a9cc · outbound

This paper cites Why does self-supervised learning for speech recognition benefit speaker recognition?.

Multi-interaction TTS toward professional recording reproduction Why does self-supervised learning for speech recognition benefit speaker recognition?

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.824285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.474432Z digest=sha256:b7aa8c1c567fd8f397ef0ceee9e40276e4f76856d93521945186510fc64ddad8

Observation 90c31184-97b6-4947-8cf3-5ad8aea0dadf · outbound

This paper cites Feed-forward networks with atten- tion can solve some long-term memory problems,.

Multi-interaction TTS toward professional recording reproduction Feed-forward networks with atten- tion can solve some long-term memory problems,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.812387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.478099Z digest=sha256:7f9d11292dffad8243065407ab88b5bb0c0caf64eed73f10ed36dbec931d42b3

Observation 3f58c8c9-6a46-4d3d-8bef-7e44cc592c25 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer,.

Multi-interaction TTS toward professional recording reproduction FiLM: Visual reasoning with a general conditioning layer,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.799080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.481854Z digest=sha256:a07adefcdb5a8595a36d5cbb2fda6a79fa348c0da0e5ff4d02dcb6fb7737f4e3

Observation a142c6f3-6c21-4885-9782-5769efddf13f · outbound

This paper cites Learning alignment for multimodal emotion recognition from speech,.

Multi-interaction TTS toward professional recording reproduction Learning alignment for multimodal emotion recognition from speech,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.786499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.485594Z digest=sha256:6cd03b42791b37991c318fca87bb11ed7d08c9d1a0af2c87b77f425c14b2134d

Observation 6f8acf10-56df-4593-9a55-93ff0820cb56 · outbound

This paper cites Multimodal cross- and self-attention network for speech emotion recognition,.

Multi-interaction TTS toward professional recording reproduction Multimodal cross- and self-attention network for speech emotion recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.774085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.489331Z digest=sha256:f62f405143ed0a168811e1421f065d1a8764b7a4a725fc425dbf7a5991b12122

Observation 30949dc0-96c0-4ef4-b9ea-9bc9b9e9bf48 · outbound

This paper cites Hello GPT-4o,.

Multi-interaction TTS toward professional recording reproduction Hello GPT-4o,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.761622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.492151Z digest=sha256:7e1567edc7b0439e23d614f9ac651ebdb6f82b9e8605b0cc310b837c50a39a45

Observation 2356dee1-9d92-443f-97df-de095f883371 · outbound

This paper cites Rephrasing the web: A recipe for compute and data-efficient lan- guage modeling,.

Multi-interaction TTS toward professional recording reproduction Rephrasing the web: A recipe for compute and data-efficient lan- guage modeling,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.750736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.498128Z digest=sha256:1cc828fcdec82a4dd6549c8731474869a03be8f85b4ce2c34d4cf62abeade720

Observation 04d3e439-1ad4-410d-a1c7-fdd0af902c2f · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech,.

Multi-interaction TTS toward professional recording reproduction FastSpeech 2: Fast and high-quality end-to-end text to speech,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.740424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.502456Z digest=sha256:dd3efa1df26e2da34412c25fdccb0e52e8389f516f687daeacd558a559204ad7

Observation 24aeb924-e40f-458f-ac95-331827345ec7 · outbound

This paper cites In- vestigating on incorporating pretrained and learnable speaker rep- resentations for multi-speaker multi-style text-to-speech,.

Multi-interaction TTS toward professional recording reproduction In- vestigating on incorporating pretrained and learnable speaker rep- resentations for multi-speaker multi-style text-to-speech,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.729800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.505360Z digest=sha256:0fb2da2aea9b03b68a15acbb14cc389f99ee7cd767a071f16404616c3a509012

Observation 0bc8d84b-6de6-49fa-9a4b-4c6bfcf054b6 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Multi-interaction TTS toward professional recording reproduction HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.508676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.508676Z digest=sha256:e06b4b5f64ae9098387cdc98574db64ede314db2550945c175cc553d6d3747b3

Observation 3dc66171-3503-43ed-a7c8-4a85e6d20c9b · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Multi-interaction TTS toward professional recording reproduction The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.512618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.512618Z digest=sha256:71202d8274b25c7f49d11db0a7145a1f634fc18635cad13257f11da3866ba097

Observation 2e60fafe-82e7-4392-80da-a62b9d224e36 · outbound

This paper cites Multi-speaker modeling for DNN- based speech synthesis incorporating generative adversarial net- works,.

Multi-interaction TTS toward professional recording reproduction Multi-speaker modeling for DNN- based speech synthesis incorporating generative adversarial net- works,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.713292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.515996Z digest=sha256:dd5cc4000b7c61f63729f550c8235b4e602aa0452eb2ea85c7c5843d99305347

Observation ac610f13-f5c7-4140-b63a-55e47da9bec4 · outbound

This paper cites Variational discriminator bottleneck: Improving imitation learn- ing, inverse RL, and GANs by constraining information flow,.

Multi-interaction TTS toward professional recording reproduction Variational discriminator bottleneck: Improving imitation learn- ing, inverse RL, and GANs by constraining information flow,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.703451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.519705Z digest=sha256:fa94547e02bf86bb68bd328cb150df6c19808766fd692831cdff3e7ddc335c8f

Observation ddfa18e2-47f7-4b0b-8e0b-9cbe15bdf6c9 · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Multi-interaction TTS toward professional recording reproduction Decoupled weight decay regulariza- tion,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:40.523505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:40.523505Z digest=sha256:6d3687ab26e809396300a6293ca0fdbb9717b21a22848d4c33375a646844c463

Observation 03f70cff-87dd-4cf0-8dfb-4b54da726799 · outbound

This paper cites Expressive text-to-speech synthesis using text chat dataset with speaking style information,.

Multi-interaction TTS toward professional recording reproduction Expressive text-to-speech synthesis using text chat dataset with speaking style information,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.685250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.526777Z digest=sha256:1b89909e0edb74dbdafd2e6cdcdb4052b3bf3134c64901a3c91654841dd729d6

Observation 79df0849-aaf6-4be3-923c-1d70f86d183f · outbound

This paper cites NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Multi-interaction TTS toward professional recording reproduction NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.673622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.530589Z digest=sha256:b19179e6ef2e0eff2c9a5336294475a6d6708309ce38d48be38b383a125ce960

Observation 55c5343b-e23d-4e33-94bd-556f25af4981 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Multi-interaction TTS toward professional recording reproduction Neural codec language models are zero-shot text to speech synthesizers,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.663664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.534096Z digest=sha256:9bc6e713bfe0d7b7493ba998674861e4b86d9f80738f884021d7bcc0e7d1ab55

Observation 448d23c8-497b-4cee-9f0c-fb3ba3d0a156 · outbound

This paper cites FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation,.

Multi-interaction TTS toward professional recording reproduction FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.652925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.537401Z digest=sha256:1e0e4ef66abf8b0c81c1d587d553d662032b0bf3db139894758d2b18751a6c90

Observation 27e012b0-5b8d-47a2-a991-ebdd4448ff7e · outbound

This paper cites Toward verifiable and repro- ducible human evaluation for text-to-image generation,.

Multi-interaction TTS toward professional recording reproduction Toward verifiable and repro- ducible human evaluation for text-to-image generation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:11:40.643029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:40.541402Z digest=sha256:55d30485c002d26784a80927c68c85134560b958ff20f7c0137ac088f62ad9cd

Pith citing papers

Observation f5356663-a23d-46d0-be41-e81f4b445693 · inbound

Multi-interaction TTS toward professional recording reproduction cites this paper.

Multi-interaction TTS toward professional recording reproduction Multi-interaction TTS toward professional recording reproduction

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:11:40.632803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:11:26.568638Z digest=sha256:a38e0161e98f63ca717e0b4b19f29b873af661bf81daa713cdc2d3d3cdf89ca9