Pith. sign in

Paper Citation Record · LEDGER

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS

As of 6 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 3 inbound Pith citation observations for arXiv:2409.18512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.18512 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T20:31:24.494060Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T15:21:28.809697Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T00:47:30.552861Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact11
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad99742c-3ffc-4e72-bef6-190a5430f71d · outbound

This paper cites Improving language understanding by generative pre- training.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Improving language understanding by generative pre- training

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.960930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:8edf8e20a6d88db81d0ec6bebe994e35c6161c4cf534ce3a54eaf6a1e3c9e71f

Observation 30f9dfec-c267-459c-b604-d7a7e2cf491e · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.623334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:9330dd7443188d916810b54c852b6ea332b1357609ec8f5becfc394fbb3e3d99

Observation 41a2b1dd-079b-45c7-b714-8916ac2a2a2d · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Speak, read and prompt: High-fidelity text-to-speech with minimal supervision

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.957044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:9d7f0c29581287deecc673774011527dedcfa7bb59df0840de217ac2c2d25bbf

Observation 037ffe26-1e20-43f8-b9e4-ea3e935ba457 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.953253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:9a8d78b20e6a4c44a21bd1fcf198823348c0953ac146d27b55a7697b517701b6

Observation f01f4e79-4c5e-45b6-bed7-27dfaa3c9b4c · outbound

This paper cites Learn- ing speech representation from contrastive token-acoustic pretraining.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Learn- ing speech representation from contrastive token-acoustic pretraining

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.956374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:cdb8fb63224c62b854afa4783be5a1c27fab2323d40d37d4fbcfe767b04a4bff

Observation 25331a38-e140-4b57-858c-6ff85959fe2c · outbound

This paper cites High Fidelity Neural Audio Compression.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS High Fidelity Neural Audio Compression

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.618036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:bd06034cdf0e96c1b0e837a0d9d7c232ad637f91a98d850de1a1fcd39a65b02f

Observation ce977247-b194-45cc-aab6-fd1a9270f528 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.628224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:2084d50891d591c0482bb4a9b7bf07babea174285d85b0a928875064cec844eb

Observation 845d86b3-8597-459f-b362-021156237d04 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.633096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:e77461a9c256d0bf0f59dd67136d73651a447d27d75faa583d914dc788d44b7e

Observation b5c32be0-0f97-46bb-b4c3-b1ecb7d43a1d · outbound

This paper cites VioLA: Conditional language models for speech recognition, synthesis, and translation.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS VioLA: Conditional language models for speech recognition, synthesis, and translation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.949626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:705fdc6b07af9adef57d3108eafcb626ece85316d3e26399b3344892c9f42b57

Observation f44803bf-8ca9-4ce2-99ac-e5a902bf5f66 · outbound

This paper cites Large lan- guage models are zero-shot reasoners.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Large lan- guage models are zero-shot reasoners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.952837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:059b8b2bc4fa87c7710eec2a5edfbf7d07294043e79007d71c6725e3d0a8c3cb

Observation 8b51d302-f83c-4283-b395-0dec0540f6d8 · outbound

This paper cites Scaling instruction-finetuned language models.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Scaling instruction-finetuned language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.959598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:2d407d8366bf1f1bf4523aa02a286b7586a87f9be84300123b4bbee3239170cf

Observation d5625e6c-f2a4-4c3b-8f7c-1c155660c07b · outbound

This paper cites Retrieval-based prompt se- lection for code-related few-shot learning.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Retrieval-based prompt se- lection for code-related few-shot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.935925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:c79ac68a8df75da7a4f3c9ed4de32d3c8ee303d37dc4ff001ac6a31cd1e10774

Observation 8e68ddf2-4fd1-445b-880d-4961b4cf0ce5 · outbound

This paper cites Learning to prompt for continual learning.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Learning to prompt for continual learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.934275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:0448195973931a955295bee9e445aa02efcf11e516bec70c112ef58b72e9e1eb

Observation 19f305a6-b778-453d-a025-e067509cd618 · outbound

This paper cites Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.596984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:85749b534c0acbfdbdfcb0a4b74f1af312d2ce4e082e33c12201526a0b15c4be

Observation d8ba8b3f-8e7a-445f-b252-fe83809d15d0 · outbound

This paper cites Universal information extraction as unified semantic matching.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Universal information extraction as unified semantic matching

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.931247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:47591f42f690f631d96429bc8afe504760253c2a969e914be30f5e83453960b1

Observation ed3b68eb-c4ac-4634-99ea-8eb79382650d · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.928572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:c0d999de067b65f498b289f90c17900e83d37e8db272c4a7f2ed6e5f77007e97

Observation b83fb258-e51b-4a4c-8554-331dbcdcf692 · outbound

This paper cites Controlling Emotion in Text-to-Speech with Natural Language Prompts.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Controlling Emotion in Text-to-Speech with Natural Language Prompts

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.581696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:63814b6c0af7eb7e52721004dd7ebe8ef8ac8b4c49a92fcca10d91a6cd941ff1

Observation 041024c9-5f47-488e-9eb4-24126add3958 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.586949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:722317ab6b7f0ae067a271e244693be5591cf50867566a2bec569f1cf621e890

Observation 3be1f78d-e8b2-4032-8a1d-270876184055 · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.607564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:bc6ddb43cae3d71d8a9a1d581612dfc6a2b797b8a05c0a79a8ddeaa799112ba8

Observation 3d1c4d59-b651-4169-8a74-dc93eb7c4d49 · outbound

This paper cites Zmm-tts: Zero-shot multilingual and multi- speaker speech synthesis conditioned on self-supervised discrete speech representations.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Zmm-tts: Zero-shot multilingual and multi- speaker speech synthesis conditioned on self-supervised discrete speech representations

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.925502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:de78c029183aeb830c44345ce48141d6b45e9ba19986cad8486da31858031602

Observation d833eacc-5cb3-480c-84da-9624a9678e28 · outbound

This paper cites Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.972191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:a67931fc0e123fd37f30e0fdf652f62501c28c5dc34d8a9300261bb7fdba0c15

Observation 9e2f48f9-a5c9-4394-aa23-73f9e3ba91ee · outbound

This paper cites GPT-4 Technical Report.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS GPT-4 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.591281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:9495949b13224792c16ffaa120b8aeeea1ca3e9995b15cd2baab9a83b1319b46

Observation ee69e7e0-17f7-4555-b2e8-474e91aa36a8 · outbound

This paper cites Intonation and emotion: influence of pitch levels and contour type on creating emotions.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Intonation and emotion: influence of pitch levels and contour type on creating emotions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.922404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:c21f366c83cfaf647b1365f9776bdd5e92841b01339a8f529f6a9f1f3f224889

Observation 113cc4ef-a5d3-44af-8655-3f51c634c89c · outbound

This paper cites Analysis of emotionally salient aspects of fundamental frequency for emotion detection.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Analysis of emotionally salient aspects of fundamental frequency for emotion detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.919852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:718db364895ce5792e9fe72cf19ab435988fe60cd9d716001fbd05c1b04a7100

Observation c28df821-9ee5-4f54-800d-10863e76941a · outbound

This paper cites Pitch in emotional speech and emotional speech recognition using pitch frequency.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Pitch in emotional speech and emotional speech recognition using pitch frequency

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.917094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:c0b056ed7d08ec439549f6b3b7303bc5d9f03b69e22047f0f4cdff8186f8cc31

Observation 9efe4297-055b-4a9f-89ab-2dacc6fb56ca · outbound

This paper cites Communicating emotion: The role of prosodic features.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Communicating emotion: The role of prosodic features

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.914058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:12c2d682181e854b5ce4a70cf911a0c3d8b991d0c87c53810c48bb23bbbcf3c4

Observation 855181f3-0960-4072-91bf-c1a34991a917 · outbound

This paper cites The global k-means clustering algorithm.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS The global k-means clustering algorithm

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.910941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:04a3ffec6f3b3bb33e10610e46acbe3981e779a0b80a06a3253e7f8083036b6d

Observation 63f5efe2-bfd4-4afc-b7a7-6756df737bc0 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Generalized end-to-end loss for speaker verification

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.978874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:14358a0e1f186fd28137226c13665dbb95ae41c2b584cee1bd4f2b23366ac956

Observation 1ee941bb-36ce-4200-8d2a-43f8bc2d17a1 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Wavlm: Large-scale self-supervised pre- training for full stack speech processing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.975463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:6cb4168239af4b2b7d22045af2b0ec2bc980b8da79a640db4d2042ba387d4b3b

Observation 8c1e784a-8609-4a47-b1bb-6dc4731fc21d · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.612916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:3dcac72b61211e567ebd20ced413131a5e9cf89d8967cebbfad8d062ed8c7dd6

Observation 005b0f11-cf72-4691-8592-be398e32260c · outbound

This paper cites Improv- ing prosody for cross-speaker style transfer by semi-supervised style extractor and hierarchical modeling in speech synthesis.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Improv- ing prosody for cross-speaker style transfer by semi-supervised style extractor and hierarchical modeling in speech synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.968859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:0ca120d983a59c4ea3a65fa63267f2dd864227b539024ce06eecfaf848644144

Observation f89302da-3746-4e9b-b854-b0ed11be929d · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre- trained transformers.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Minilm: Deep self-attention distillation for task-agnostic compression of pre- trained transformers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:35:48.964677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:d8e8ba1da75ca2f72afb1c68041e31c88d2a14c78e8f7fb62d8cc2abf0786942

Observation 222b4700-42e0-42cf-a599-33585aa123cf · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.602164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:cd44df62fbc01fd776b58a3be1789ea058c87e329243131368b95890b4eaed92

Pith citing papers

Observation 04253d91-b32c-472f-8ddd-81d871d9c2c4 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:11:27.091266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:278d1de4aea13c2bee7f04fc2a34b3e19f871b913dc81b79ced6c20d7c8340eb

Observation 02415a3e-7d8b-468f-a09f-c4751d9e55d9 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.809697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.809697Z digest=sha256:7ca5a59823275672f3863990fe9bea7535cce9fcc6ef089f7ef220ce017287b8

Observation 234b4815-02f5-4a2b-9e88-cd6e0ca0ddb4 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:47:30.554212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:49a1a3b54490ae42d370dda9c5cabe02b1b45985ec8e2e701dea880baba82a6a